Why the number you see on AirDNA is too high
The one peer-reviewed measurement of vendor short-term rental data found occupancy overstated by 60% and revenue per available night by 179%. The cause is definitional, and here is how we define ours.
If you have underwritten a short-term rental in the last five years, you have used a vendor occupancy number. This page is about where that number comes from, what the one independent measurement of it found, and the definitions we use instead. It is not a claim that our numbers are more accurate. Nobody can prove that without ground truth, and we say so at the end.
What every vendor is actually doing
Since 2014 the platform's public calendar shows two states for a night: available or unavailable. It does not say whether an unavailable night is a booking or a host blocking the dates. Every occupancy figure built from public data is a guess at that one binary, wrapped in a revenue formula.
AirDNA describes its guess as "16 booking signals". AirROI describes its as "proprietary ML/AI-based techniques". Neither publishes the classifier, its features, or how often it is wrong. The only fully open method, Wang et al. in PLOS ONE in 2024, is explicit that it cannot resolve the ambiguity either, only bound it.
What the peer-reviewed measurement found
Agarwal, Koch and McNab at Cornell compared AirDNA's listing-level data for Virginia, 2015 to 2019, against the definitions the hotel industry uses to report occupancy. Their finding, in Cornell Hospitality Quarterly:
| Metric | AirDNA versus industry definition |
|---|---|
| Occupancy | overstated by 60% |
| Average daily rate | overstated by 78% |
| Revenue per available night | overstated by 179% |
The same vendor reconciles its national totals to the platform's own SEC filings within about one percent, and says so in its marketing. Both facts are true at once. A national total can be right while every local number is inflated, because the error is not random. It is built into the definitions.
Where the inflation comes from
AirDNA's own help centre gives the definitions. Occupancy is "reserved days" divided by "active listing nights", and a night counts as active only if it "is not blocked and is either reserved or available, provided there has been a reservation within the past 28 days."
Read that twice. Two things are removed from the denominator before dividing.
Blocked nights are removed. A host who blocks March for a renovation has no March nights in the denominator. The occupancy reported for that listing is occupancy of the nights the host chose to offer, not of the year. For an underwriter, the year is what matters. The mortgage is due in March.
Listings without a recent reservation are removed. A listing that has not taken a booking in 28 days stops being active and drops out of the average rather than being recorded near zero. This is a survivorship filter. The markets where this filter bites hardest are exactly the oversupplied ones where a buyer most needs to see the dead listings.
Neither choice is hidden. They are documented. They are also the reason the number is high.
The review-based alternative is worse
Inside Airbnb, whose open data underpins most academic work and several free tools, does not read the calendar at all. Its San Francisco model estimates bookings from reviews: reviews divided by an assumed review rate of 50%, times an assumed stay length, capped at 70% occupancy.
The 50% rate is contradicted by the platform's own researchers. Fradkin, Grewal, Holtz and Pearson, working inside the company with proprietary data, found that 67% of stays result in a review. Dividing by 0.50 instead of 0.67 inflates the estimate by about a third before the 70% cap claws any of it back. Inside Airbnb says plainly that the stay length and the cap are conventions rather than measurements, which is more disclosure than any commercial vendor offers.
What we do instead
We do not have a classifier that knows a booking from a block. We have a rule, and the rule is published.
A booked night is a night we watched flip. We read each listing's forward calendar on a schedule and count a night as booked only when it was available in one reading and unavailable in the next. A night we never saw available is never counted as booked. This is the conservative transition rule from Wang et al., and its error runs in the direction of undercounting, which for an underwriter is the safe direction.
Occupancy is served on the underwriting denominator. Booked nights over all nights in the window, blocks included. Where we show the vendor-comparable figure, blocks removed, we label it as such and note that it runs about ten points higher on the same bookings.
Nothing drops out for being dead. A listing with no bookings is a data point about the market and stays in the average.
Where we have not measured, we say so. Places without enough calendar readings show a modeled band from a small-area model with a published out-of-sample error, currently 21.8%, and the band is labelled modeled. Places where even that does not hold show nothing and the reason. Our Indio page, which served a revenue figure built from two houses until last week, is the case study for that rule.
The honest limit
Reading a calendar on a schedule misses bookings that happen between readings. Wang et al. measured a fortnightly cadence at less than half the booking activity that daily reading captures. Our current readings sit at that fortnightly end for most markets, so our measured pace understates activity, and we annualize it with a booking curve rather than serving the raw pace. The fix is cadence, and a denser reading schedule is the next thing we are building.
That is the trade. The vendors publish a number that is easy to use and, on the only independent test, badly high. We publish a band that is harder to use, states its denominator, and errs low. If you are borrowing against the revenue, you want the second one.
Sources
- Agarwal, Koch and McNab, "Airbnb's Success: Does It Depend on Who Is Measuring?", Cornell Hospitality Quarterly (2021)
- Agarwal, Koch and McNab, "Differing Views of Lodging Reality: Airdna, STR, and Airbnb", Cornell Hospitality Quarterly (2019)
- AirDNA help centre, how occupancy rate is calculated
- AirDNA help centre, how revenue is calculated
- Fradkin, Grewal, Holtz and Pearson, "Bias and Reciprocity in Online Reviews", EC '15
- Wang et al., PLOS ONE (2024), daily versus fortnightly calendar scraping
- Inside Airbnb, the San Francisco model