The Acquisition Index · Version 0.1
Methodology
What is counted, what is excluded, how each figure is derived, and what this dataset cannot tell you. Published before the first market figure, and versioned — when a definition changes, the version increments and figures published under an earlier version stay attributed to it.
What is counted
Publicly advertised businesses for sale in Canada, observed on the sites where they are listed. A listing enters the dataset when it is first observed and stays in it permanently, including after it disappears — a listing that vanishes is a data point, not a row to remove.
The record is append-only and enforced as such in the database. Observations cannot be edited or deleted; a correction is recorded as a new observation that supersedes the old one. The history is the asset, so it is not left to convention.
What is excluded
- Quebec listings, which follow a separate legal and linguistic market and are collected but not reported in the national cuts.
- Commercial real estate with no operating business — a building is not a business, and mixing them would distort every price distribution.
- Franchise territory offers and business-opportunity solicitations, which advertise a fee to start something, not a price to buy something.
- Sources whose data does not change. One aggregator in the corpus carries listings captured in February 2024 that have shown no departures, no new listings and no price movement since. Including it would contribute a structural 0% departure rate. A stale source that looks like data is more dangerous than a declared gap.
The distinction that governs every figure
Re-fetching a listing at its origin produces exactly three outcomes, and keeping them apart is the most important commitment in this document.
| Outcome | What it means | Counts as |
|---|---|---|
| Present | The source served the listing | Still on the market |
| Gone | The source explicitly said the listing no longer exists | Left the market |
| Unobserved | Blocked, rate-limited, timed out, or a server error | Nothing. A gap, recorded as a gap |
The third row is the safeguard. A firewall block and a withdrawn listing look identical if you only check whether a page loaded, and treating them alike would inflate the withdrawal rate — the single most damaging error this dataset could make. An ambiguous response is never a data point.
A listing’s absence from a site’s own sitemap is not evidence it is gone. This was measured: of eight listings missing from one source’s sitemap, four were genuinely removed and four were still openly for sale. On another, 25 of 25 “missing” listings were live. Sitemaps decide only which listings are worth re-fetching; nothing is retired without a direct answer from the origin. The converse was measured separately and points the other way: on one source, 40 of 40 listings sampled from its sitemap were live at the origin, so presence there is trustworthy even though absence is not.
What a disappearance does and does not tell us
A listing that vanishes has sold or been withdrawn, and the disappearance alone cannot say which. Only sources that explicitly mark an outcome allow the two to be separated. Until sale outcomes are observable at scale, this index reports a departure rate and does not describe it as a sale rate.
How “gone” is established, source by source
There is no shared convention for retiring a listing, so the departure rule is written and verified separately for every source. The assumption that a withdrawn listing returns an HTTP 404 is wrong on four of the ten: they serve the complete page with a success status and change one thing — a redirect target, a status field in the page data, an availability value, a banner. On those sources a status-only rule cannot report a single departure, and one of them was found to have lost 5 of 22 listings to sales while reporting none.
Every candidate signal is tested against listings known to be live before it is trusted, because the obvious signal is repeatedly the wrong one. Four measured cases: one source inlines its entire translation catalogue into every page, so “no longer available” and “expired” appear in the markup of 100% of pages, live ones included. Two others carry “expired” on every page as the name of a reCAPTCHA callback. A fourth ships a “Sold” ribbon element in the markup of all 22 of its listings and reveals it on 5. Any of them would have retired an entire source in one run.
Rules therefore read structured data rather than prose wherever a source publishes it, and they enumerate the states that mean departed rather than testing for “not available”. An unrecognised new status is treated as still present: a missed departure is corrected by the next run, whereas a wrongly recorded one is written into a log that cannot be edited.
Two sources cannot report departures at all. They publish no status markup, no availability field and no observed example of a listing leaving, and both phrase candidates fail the live-page test outright. They are marked as blind spots rather than left looking clean, because a source reporting zero departures is indistinguishable from a source that is not being measured — and any departure rate computed across all sources without excluding them is biased downward.
Two sources do state the reason a listing left. Where that happens it is recorded, so a sale is stored as a sale and not merely as a disappearance. It is not yet enough coverage to publish a sale rate, and it will be labelled by source when it is.
Observation dates, and the censoring problem
The date we observed something is recorded separately from every date a listing claims about itself. Everything downstream depends on when we saw it; the seller’s dates are evidence, not truth.
This matters most for duration. For 10,855 listings already in the corpus when observation began, the true listing date is earlier than anything we can see by an unknown amount — left-censored data, which can only ever yield a lower bound. For 2,877 listings observed appearing, the clock is exact. The two are stored with a flag distinguishing them and are never combined into one median.
How sector and geography are assigned
Sector is assigned from the source’s own taxonomy where it publishes one, and otherwise inferred from the listing title and description against a fixed sector list. Inferred assignments are marked as inferred.
Geography is taken at the finest granularity the listing states — city where given, otherwise region, otherwise province — and is never interpolated upward into a precision the listing does not support. Province is read from the source’s own taxonomy links in preference to the text, after an earlier text-first approach mislabelled listings by matching page navigation.
How prices are read
Price extraction is enabled per source only after being checked against that source’s real markup. A generic “find the labelled dollar amount” approach was written first and rejected: it read a listing priced at $225,000 as $5,000, having matched a related-listings sidebar and stopped at the space in “$5 200 000”. A missing price shrinks a sample; a wrong price corrupts a published median and cannot be found again once averaged in.
Two further facts about prices are recorded rather than assumed.
- Currency is not uniform. A substantial share of listings on Canadian sites are priced in US dollars. Currency is stored per listing and stated rather than converted.
- Not every “asking price” is the price of a business. At least one source advertises whole businesses, minority equity stakes and loan-seeking mandates through the same price field; on a sample of that source’s listings, half were partial stakes. The index filters to whole-business sales rather than averaging figures that are not comparable.
Sample size and suppression
Every published figure carries the number of listings it was computed from. Any cut with fewer than five is suppressed and shown as suppressed, rather than quietly omitted. A cell reading “n<5, not reported” builds more trust than a number nobody can stand behind.
Known biases, stated plainly
- Scraped listings over-represent brokered deals. Businesses sold privately, without ever being advertised, are invisible here and are believed to be a large share of all transactions.
- Asking prices are not sale prices. Nothing in this dataset is a sale price unless it was observed as one.
- Listing sites skew toward larger and more saleable businesses. The smallest owner-operated businesses change hands informally.
- Coverage is uneven by province. Ontario is heavily over-represented relative to its share of Canadian businesses, because the sources that publish openly are concentrated there.
- Some sources cannot be re-observed at all. Of 11,492 listings held, 9,914 are on sources we can re-fetch. Duration and outcome figures are computed over that subset, and the sample size shown beside each figure is that subset — never the full corpus.
Corrections
Errors are corrected by adding a superseding observation, never by editing the record. If a published figure changes materially as a result, the change is noted here with its date and the version increments.
Questions about the method
This document is intended to be argued with. If something here is wrong, unclear, or would not survive review, we would rather hear it than publish on it.