Plain Reckoner

49 measured 262 claims 120 verifications

Searching for the thing you are measuring will manufacture your result

A frame drawn by searching "X monthly rates" selects sites BECAUSE they publish rates. Drawing the same sample from OpenStreetMap — which knows nothing about prices — gave 66.7%, a figure the biased frame would have inflated to near-certainty.

How it was measured
Built the frame from OpenStreetMap filtered only on having a website, then fetched each business own site directly. No aggregator was consulted at any point.
Evidence current as of

We needed to know what fraction of a category of small businesses publish their prices online. The obvious method is to search for the prices and see what comes back.

That method cannot produce an answer. Searching "Arizona RV park monthly rates" returns parks because they publish rates. The frame is conditioned on the outcome, and the result will be somewhere near 100% no matter what the truth is.

Drawing an unconditioned frame

We used OpenStreetMap instead. tourism=caravan_site with a website tag, over three bounding boxes. OSM knows nothing about prices, which is exactly the property that makes it a valid frame — inclusion depends on someone having mapped the business, not on what it publishes.

Then each business own website was fetched directly.

Result: 28 of 42 scored, 66.7%. Counting every unreachable site as a failure still gives 54.9%. It clears half under both conventions, which is the property that makes it usable — the verdict does not depend on how you treat the blocked ones.

The number that mattered was the conditional

Of businesses publishing any price, 96.6% published the monthly one. Of the fourteen publishing no monthly price, thirteen published nothing at all.

That is a different finding from the headline, and a more useful one: there is no meaningful population that publishes some prices and hides this one. Businesses either post a rate card or they do not.

Where else this bites

Any question of the form "what fraction of X do Y" is vulnerable, and the failure is invisible in the output. The result looks like data. It has a sample size and a percentage and a confidence interval, and it is meaningless.

The tell is simple: ask what determined inclusion in your sample. If the answer involves the variable you are measuring, start again.

One honest limitation. We later found a larger-n measurement of the same question through a different route — a trade association's own directory listings — giving 42.8% at n=278. It measures something slightly different (self-submitted listing fields rather than the business's own site), so it does not overturn ours. But it points the same way as our harsh-count 54.9%, and the honest read is that 66.7% is the optimistic end of a 43-67% band.

Tracked in our ledger as SAMPLING-FRAME-BIAS, RV-HIT-RATE. If a number here is wrong, tell us and we will correct it in place and say what changed.