Review scraper comparison: 17 sources, side by side
Every review actor in the kestrel suite, with the facts that decide which one you need: what a single row carries, whether it has sub-ratings and the owner's reply, which languages you get, what one row costs, and — the number most listings leave out — how deep a single property or company actually goes. Everything here is measured and comes from each actor's own documentation.
| Source | What one row carries | Sub-ratings | Owner reply | Language coverage | Price per row | Ceiling per property or company |
|---|---|---|---|---|---|---|
| Agoda | 0-10 rating plus rating_5, rating word, title, text, positives/negatives kept apart, stay dates, stay_length, room_type, reviewer name and country, helpful votes, and the partner source the review came from. | No | Yes — responder, response, response_date | As written, tagged in language; original_text keeps the untranslated body. One Lisbon property returned 18 languages. | $0.005 / review | Every review; Agoda pages ~200 at a time. 1,457 unique reviews from one property was verified. |
| Booking.com | 1-10 rating plus rating_5, title, positives and negatives split, joined text, check-in/check-out, nights, room_type, traveller type, reviewer country, photos. | On the free hotel row — seven category scores (staff, facilities, cleanliness, comfort, value, location, wifi) | Yes — response | As written, tagged; language is a server-side filter and the hotel row lists the languages with counts. 24 languages across 892 reviews on one hotel. | $0.005 / review | Every review, 25 per call. Score-only reviews are real: 11 of 25 in one sample had no words (requireText drops them before billing). |
| TripAdvisor (hotels) | 1-5 bubbles, title, full text, six sub_ratings (value, rooms, location, cleanliness, service, sleep_quality), stay date, trip type, photos, reviewer hometown and contribution counts. | Yes — six per review | Yes — response, response_date | Up to 30 site languages; each domain shows its own language plus machine translations. A translation is a separate row unless includeTranslated: false. | $0.005 / review | Every review the language site lists — 1,456 on the English site against an all-language review_count of 1,977 in the verified example. Hotels only. |
| TripAdvisor (restaurants & attractions) | 1-5 bubbles, title, text, the four dining sub_ratings (food, service, value, atmosphere), visit date, trip type, the tips left for the next guest, photos, dish tags. | Yes — four on restaurants; attractions carry none | Yes — response, response_date | Per-language domains, with original_language and a translated flag on every row. | $0.005 / review | 15 reviews per page on a restaurant, 10 on an attraction. A busy city-centre restaurant holds 2,000-12,000 English reviews; a famous food market can pass 25,000. |
| Airbnb | rating (a string, as Airbnb returns it), the review as written in text, Airbnb's translation in text_localized, reviewer name, id and self-declared location, superhost flag, stay type, highlight. | No | Yes — host_reply, host_reply_date | Original always in text; text_localized only when Airbnb translated it. One Lisbon listing returned Portuguese, Spanish, German, Korean and English in one run. | $0.005 / review | Every review; Airbnb pages 50 at a time. A 441-review listing is 441 rows. |
| HRS | 0-10 rating plus rating_5, recommends, the HRS comfort word, positives and negatives apart, text, traveller type, age group, arrival and departure dates, nights. | On the free hotel row — HRS category averages (avg_by_category, avg_by_super_category) | Yes — hotel_reply | As written, tagged in language; HRS is German-founded, so most rows are German, then English. Nothing is translated. | $0.003 / review | Only commented ratings are delivered: a property may show rating_count 150 and comment_count 85. The whole list arrives in one answer — there is no paging on HRS's side. |
| Trip.com / Ctrip | 0-10 rating plus rating_5, rating word, four sub-scores (rating_location, rating_facility, rating_service, rating_room), text, check-in, room type, travel type, reviewer country and level, images. | Yes — four per review | Yes — response, response_date | As written, tagged. One Lisbon hotel returned nine languages and roughly a quarter of its corpus was Chinese. | $0.004 / review | Every review, 100 per request. total can sit a little above the pageable rows because Trip.com counts repeated reviews once. |
| Kurzurlaub (DACH short breaks) | score 1-6 (6 is best), category_scores per area (overall, hotel, room, service, food, leisure, location, value), title, text, reviewer, date. | Yes — per area | No | German. Titles, comments and the site's rating words come back in German; the column names and area keys are English. | $0.003 / review | Every rating; pages carry 10-20 and the actor follows the site's own "next" link. A 205-rating property is $0.62 in full. The site gives ratings no id, so review_id is a stable hash. |
| Hostelworld | 0-100 rating plus rating_5, seven sub_ratings (value, safety, location, staff, atmosphere, cleanliness, facilities), text, group type, age group, reviewer country and nickname. | Yes — seven per review | Yes — owner_comment | One machine-translated English set. Hostelworld accepts a languages parameter and then ignores it, so this actor offers no language filter rather than one that does nothing. | $0.003 / review | Every review, 50 per page (the API's cap). groupTypes, minRating, maxRating, requireText and sinceDate are applied actor-side, before billing. |
| Zoover (NL/BE) | score 1-10, aspects the guest rated 1-10 (room, food, service, location, hygiene, pool, childFriendly, priceQuality, general), title, text, travel date, who they travelled with, photos. | Yes — per aspect | Yes — response, response_author, response_position | Dutch, mostly locale: "nl" — the Dutch and Flemish complaints that rarely reach the English-language sites. | $0.004 / review | Shallow by nature: a sitemap sample ran 0, 0, 0, 0, 0, 0, 0, 1, 4, 4, 6, 11, 16, 23 reviews per property. Sweep lists of properties; empty ones are free. |
| Despegar (LatAm) | A hotel-level row: rating on Despegar's 10-point scale, rating_5, the label, six category_scores, the review summary, trip types, providers — plus the handful of comments the page publishes. | Yes — six (service, staff, location, internet, cleanliness, value) | No — management replies are not on the page | The country site answers in its own language (page_language, e.g. es), which is where LatAm guest opinion lives. | $0.004 / hotel | Four comments per hotel, maximum — the ceiling of the public page, reported exactly in reviews_shown. Dates are month-level (review_month). Residential proxy required; decolar.com is hard-blocked. |
| Trustpilot | 1-5 rating, title, text, review_date and experience_date, reviewer name, country and lifetime review count, the verified badge and verification_level, likes, location. | No | Yes — reply_text, reply_date | Every language, filtered server-side; the free company row carries review_languages with counts. One travel brand returned forty languages across 120,000 reviews. | $0.0007 / review | An anonymous visitor gets 200 reviews per filtered view. With deepPaging the five star bands are merged into roughly 1,000 reviews per company per run, and capped: true says when the wall was hit rather than the end of the reviews. |
| Indeed (employers) | 1-5 overall plus five sub-ratings (work/life balance, compensation and benefits, job security and advancement, management, culture and values), title, text, pros, cons, job title, location, employment status, helpful counts. | Yes — five per review | Yes — employer_reply, employer_reply_date | 44 Indeed country domains; the raw date string is kept in review_date_text in that domain's language. | $0.002 / review | Per domain: about 4,000 on indeed.com and about 6,250 across all domains for the company verified. A page past the end silently repeats page one, so the walk is bounded three ways and repeated rows are never billed. |
| Amazon | 1-5 rating, text, author, country (where the review was written), date, the verified-purchase flag, helpful votes, the variant reviewed — plus a free product row with the full rating summary. | No | No | Per marketplace domain. | $0.005 / review, $0.003 / product | 8-13 reviews per product. That is the public product page's ceiling: Amazon requires a signed-in session for the full archive and this actor deliberately does not attempt it. reviews_visible says exactly how many were on the page. |
| AliExpress | 1-5 rating (and rating_100), the buyer's words plus text_translated, buyer country, the exact SKU variant, buyer photos, logistics, and the follow-up review left later. | No — the free product row carries the whole star histogram | No | Original text plus a machine translation into the language you pick, with language on every row. | $0.002 / review | The whole corpus, but the endpoint is odd: page 1 always returns 20 rows whatever page size is asked for, later pages honour up to 500 at a plain offset, and the backend shuffles — repeats are deduped by review_id and counted in repeats. |
| Apple App Store | 1-5 rating, title, text, author, app version, the edited flag, the review date and Apple's developer_response — plus a free app row with the full star histogram. | No | Yes — developer_response, developer_response_date | One set of rows per storefront country; name every country you care about. | $0.004 / review | Apple pages ten reviews at a time by offset and keeps answering deep into an app's history (tens of thousands for the biggest apps). Only written reviews are pageable — written_reviews says how many the storefront holds. |
| Google Play | 1-5 rating, text, author and Google's opaque author_id, thumbs_up, app_version, the review date and the developer's reply_text. | No | Yes — reply_text, reply_date | Per hl/gl locale. Google decides which reviews it shows under a locale, so name every language the app has traffic in for a complete corpus. | $0.004 / review | The full corpus, 100 reviews per call. most_relevant is Google's ranking, not a date order — schedule most_recent. |
The ceilings that actually bite
Four of these sources cap what any anonymous visitor can read, and no scraper can honestly get past them:
- Amazon — 8-13 reviews per product. That is what the public product page renders. The full archive needs a signed-in session, which these actors deliberately do not attempt.
reviews_visiblereports the exact number. - Despegar — four comments per hotel. Not a setting and not a paging bug: the ceiling of the public page, reported in
reviews_shown. Use it for the hotel-level scores and the summary, not for a review corpus. - Trustpilot — 200 reviews per filtered view. Page eleven of any view redirects to the login screen. The five star bands are five separate allowances, so merging them yields about 1,000 reviews per company per run, and
capped: truesays when the wall stopped the walk. - Zoover — the catalogue is wide and shallow. Most accommodations hold no reviews at all; the busiest in a sitemap sample held 23. Sweep hundreds of properties rather than expecting depth from one.
Filters that run before billing
Every review actor applies its filters before the charge call, so the rows you drop cost nothing. Where the source supports a server-side filter the actor uses it — Trustpilot's stars=, dateRange, languages and keyword search all run on Trustpilot's servers, and Booking.com filters by language, traveller type and keyword on its own — and where it does not, as on Hostelworld, the pages are read and the rows are dropped before billing. A rating ceiling plus requireText is what makes a portfolio-wide complaints feed cost cents: one Agoda property with 324 visible reviews delivered 8 rows under maxRating: 5 with requireText on.
Joining rows across sources
The scales differ and that is the first thing to get right. Agoda, Booking.com, HRS and Trip.com rate 0-10; TripAdvisor, Airbnb, Amazon, AliExpress, Trustpilot, Indeed, the App Store and Google Play rate 1-5; Hostelworld rates 0-100; Zoover 1-10; Kurzurlaub 1-6, where 6 is the best. Every actor that uses a non-five-point scale also emits rating_5, the same score expressed out of five, so a cross-source average is a column read rather than a conversion you have to remember.
FAQ
Which review source gives the deepest single-property corpus?
Trip.com, Booking.com, Agoda, TripAdvisor and Airbnb all page an entire property: 1,457 reviews from one Agoda property and 1,456 on the English TripAdvisor site were verified. Amazon (8-13 per product) and Despegar (four comments per hotel) are the hard ceilings.
Which sources carry sub-ratings per review?
TripAdvisor hotels (six), TripAdvisor restaurants (four), Trip.com (four), Hostelworld (seven), Zoover (per aspect), Kurzurlaub (per area), Indeed (five) and Despegar (six, at hotel level). Booking.com and HRS publish their category scores on the free property row rather than per review.
Which ones return the property's reply?
Agoda, Booking.com, TripAdvisor (hotels and restaurants), Airbnb, HRS, Trip.com, Hostelworld, Zoover, Trustpilot, Indeed, Apple App Store and Google Play. Kurzurlaub, Despegar, Amazon and AliExpress do not publish replies on the pages these actors read.
Are reviews translated?
Reviews are delivered in the language they were written in and tagged. Airbnb and AliExpress add the platform's own translation in a second field; TripAdvisor serves machine translations as separate rows per language domain; Hostelworld only ever serves one machine-translated English set.