Skip to content

Site Selection Software Accuracy: How to Back-Test a Score

9 min read

Share

A site score can rank your worst store highly because it grades the ground, not the operation, and because it learned from locations your company already chose. Test any vendor by holding out stores you already know, sending addresses only, locking the scores in writing, then checking whether the rank order matches your actuals.

Every real estate leader who has sat through enough demos has had this moment. The vendor drops a pin on a store you already own, the score comes back an 87, and you know for a fact that store has missed its numbers for three years running. Nobody says anything. The demo moves on.

That moment is worth more than the rest of the meeting. It is a free validation test the vendor did not intend to run, and how they respond to it tells you most of what you need to know.

Why a model can rank your worst store first

There are a handful of ordinary reasons a score misses, and none of them require the vendor to be doing anything dishonest.

It learned from the sites you picked, not the ones you passed on

Most location models are trained on a retailer's existing stores. Those stores are the survivors of years of internal screening. Your team already threw out the bad corners, the wrong co-tenants, and the deals with impossible rent, so the model never saw them.

Statisticians call this range restriction, and the consequence is specific: a model trained only on sites that cleared your bar has very little information about what a failing site looks like. It can tell you which of your good sites is best. Asking it to recognize a bad one is asking it to extrapolate past the edge of everything it has seen.

Fleet size makes this worse. Paul Sill, who heads JLL's Visionary Insights Group, put the arithmetic plainly to Modern Retail in December 2025: "If I'm a 500-unit coffee shop, 500 observations is not enough data for AI to do anything intelligently." Most brands running an expansion program have far fewer than 500 units.

The score grades the ground, your P&L grades the business

A location model scores the site. Your store results carry the site plus everything else: the manager, the staffing level, the rent you negotiated, the build-out, the format you put there, whether the opening was timed well.

Two-column comparison of a site selection model's inputs, including population, income, foot traffic, competitor proximity, and access, against store-level drivers it usually cannot see, including manager tenure, rent and lease terms, build-out quality, opening timing, and inventory decisions.

This is the most common explanation for the demo moment, and it is partly good news. A high score on a bad store sometimes means the real estate is fine and the problem is elsewhere. That is a useful finding, but only if the vendor says it out loud instead of talking past it.

Everything comes back in the same narrow band

Look at the distribution, not the number. If every site the vendor evaluates scores between 70 and 85, the model is not separating anything. A score that never goes low cannot tell you what to avoid, which is the main thing you are buying.

Compressed output also breaks the validation math. When the range of an outcome is squeezed, any correlation you compute against real performance is suppressed, so the model can look defensible in a vendor's own testing and still fail to discriminate on your portfolio.

Some inputs are estimates wearing the clothes of facts

The layers underneath a score are often models themselves. Esri's own technical paper on its 2026 Business Locations data, published in February 2026, is direct about this: "Data Axle sales volume attributes are modeled, as verifiable sales volume figures are extremely difficult to obtain from private businesses. The primary source of information for the model is the U.S. Department of Commerce's data for sales per employee for each NAICS code."

Read that again with your competitor set in mind. A competitor's sales figure in your trade area may be an employee count multiplied by a national industry average. That is a reasonable estimate and a poor foundation for a decision worth several million dollars, and the difference matters when you are choosing which site selection data to trust.

Foot traffic panels are relative signals, not counts

Mobile location panels measure a sample of devices and extrapolate. A peer-reviewed audit of SafeGraph data published in PLOS ONE in January 2024 found the panel's sampling rate averaged 7.5% over five years, and that device counts tracked census population well at the county level while the association "at the tract and block group levels becomes weaker." The same study found that Hispanic populations, low-income households, and people with less education were underrepresented, and that the bias varied by place and over time.

Tract and block group is exactly the resolution a trade area cares about. Use foot traffic to compare one site against another. Do not treat the absolute number as a count of humans. Our guide to evaluating location data providers covers how to bake off the underlying data before it reaches a score.

Back-test the score against a portfolio you already know

You do not need a data science team to grade a vendor. You need stores whose results you already have.

Six-step back-test protocol for evaluating site selection software: hold out real stores, send addresses only, lock the scores in writing, compare rank orders, check the top and bottom of the list, and check the spread of scores.

Hold out real stores. Pull 20 or more locations from your own fleet, deliberately mixing your strongest, weakest, and unremarkable units. A test made only of good stores has nothing to separate.

Send addresses only. No sales, no rankings, no hints about which ones you like. Anything you volunteer is something the model no longer has to work out for itself.

Lock the scores in writing first. Get every score before you share your actuals, with a date on it. A score that arrives after the reveal is not evidence of anything.

Compare the rank orders. Line the vendor's score order against your actual performance order. The right test here is rank correlation, the same statistic you would use for any "did this list come out in the right order" question. You are not asking whether store 14 should have been an 82. You are asking whether the list is in roughly the right sequence.

Check the two ends. Middle-of-the-pack noise is forgivable. Your best stores appearing near the bottom, or your known underperformers near the top, is the failure that matters, because the ends are where site decisions actually get made.

Check the spread. Write down the highest and lowest score in the set. If the whole portfolio fits inside fifteen points, you have a model that agrees with everything.

A vendor worth continuing with will engage on all six and will tell you which inputs moved each score up or down. A vendor who explains the misses by telling you the method is a trade secret has answered the question.

What a good vendor answer sounds like

Push on four things, in any demo:

  • Which inputs moved this score, and in which direction. Not a feature list. This site, these variables, this much.
  • How the model was validated, and on whose data. Held out from training, or fit and reported on the same stores?
  • Whether you can change the weights. A drive-through brand and a mall inline store do not care about the same variables.
  • How it behaves on a format it has not seen. A model tuned on suburban strip centers has an opinion about urban street retail whether or not it has earned one.

Across the vendors a retail team typically shortlists, including Placer.ai, Buxton (now part of Audiense), SiteZeus (now Atlas), Kalibrate, Esri Business Analyst, CoStar, MRI Software, and Sitewise, published validation methodology is scarce. We went looking and found very little public documentation of how these scoring and forecasting models are tested. That is not an accusation of bad math. It means the burden of testing falls on you, which is why the back-test above is worth running yourself. If you are still assembling a shortlist, our comparison of retail site selection platforms covers what each one is built for.

Assumptions worth re-testing on your own portfolio

The back-test frequently turns up something more interesting than a vendor verdict. In custom model work for a veterinary group, the variables that predicted clinic performance were not the ones the industry assumed. Clinic maturity, staffing mix, and local competition mattered more than demographics, and household income correlated in the opposite direction from what everyone expected. Common yardsticks like revenue per square foot penalized the best clinics and rewarded weaker ones. Trade areas drawn from actual customer records looked nothing like the ring studies they replaced.

None of that is exotic. It is what happens the first time a fleet gets measured instead of assumed, and it is the same reason forecast accuracy work belongs next to site scoring rather than downstream of it. If your model says household income drives your business and your own stores say otherwise, the model is describing a category, not your brand.

How GrowthFactor approaches the same problem

The reason buyers run this test at all is the black box. A number arrives, nobody can explain it, and the real estate committee is asked to trust it.

GrowthFactor's site scoring is built the other way around. A site is scored across configurable lenses, every input that moved the score is visible with a written justification, and the weights belong to your team, so a car wash and a bookstore are not graded on the same assumptions. We call that the Site Scoring Glass Box internally. What it means in a committee room is that "why this site?" has an answer you can read out loud.

It does the analysis. You make the call. The landlord you can actually negotiate with, the manager you already have, the market you know because you grew up in it: none of that reaches a model, and the score is not trying to overrule it.

The honest position is that our score deserves the same back-test as anyone else's. Run it. If you already track cannibalization and market saturation across your fleet, you have most of the data assembled already.

Frequently Asked Questions about site selection software accuracy

Common questions from real estate teams evaluating scoring models.

Why does site selection software give a high score to a store I know is underperforming?

Usually because the model grades the trade area rather than the operation. Rent, staffing, build-out, format fit, and local execution move store results but rarely appear in a vendor's training data. A second common cause is that the model learned from sites your company already chose, so it never saw the locations you rejected and has little basis for recognizing a bad one.

How do you test the accuracy of site selection software?

Back-test it against stores you already know. Pick a mix of your strongest, weakest, and middling locations, send the vendor addresses only, and get every score in writing before you share your actual performance. Then compare the two rank orders. The question is whether the model puts your best stores near the top and your worst near the bottom, not whether any single score looks correct.

What is score compression in a site selection model?

Score compression is when nearly every site the model evaluates comes back in a narrow band, for example 70 to 85. A model whose output barely varies cannot separate good sites from bad ones, and the narrow range also suppresses any correlation you try to compute against real performance. If a vendor's scores never go low, the score is not doing the work you are buying it for.

What should I ask a site selection vendor about their model?

Ask which inputs moved this specific score and in which direction, how the model was validated and on whose data, whether you can change the weights, and how the model behaves on a category or format it has not seen before. Ask for the score distribution across your own portfolio. A vendor who can answer those in plain language is a different proposition from one who treats the method as a trade secret.

How does GrowthFactor compare to Placer.ai for validating a site score?

They solve different halves of the problem, and plenty of teams run both. Placer.ai sells foot traffic measurement, which is one input into a site decision. GrowthFactor scores the site across configurable lenses, shows which inputs moved the score, and keeps that score with the deal through the pipeline, so the back-test above is something you can run on your own portfolio inside the tool rather than as a one-off vendor exercise.

Share

Continue reading

Location Data Providers: How to Evaluate One Before You Buy

Vendor record counts tell you almost nothing. Here are the seven criteria that separate location data providers, and the bake-off you can run against your own store list in an afternoon.

Aug 3, 2026

Competitive Intelligence Reports: What Good Looks Like

A retail competitive intelligence report has six sections, and only one of them decides anything. Here is the template, where each section's data comes from, and how often to produce it.

Jul 29, 2026

7 Kalibrate Alternatives for Site Selection (2026)

Looking for Kalibrate alternatives? Compare 7 site selection platforms on forecasting depth, self-serve speed, and retail coverage beyond fuel and convenience.

Jul 23, 2026

Newsletter

This Week in Retail

Store closures, expansion tracking, and original market analysis. A five-minute read every other Thursday.

Ask GrowthFactor where to open next

Watch it pull the data, run the analysis, and explain the answer in maps and tables. It does the analysis. You make the call.