Skip to content

When to Trust Your Gut on a Site, and When to Trust the Data

8 min read

Share

Trust your gut on the physical property. Trust the data on the market. You've walked hundreds of sites and gotten corrected within minutes, every time. You've opened a few dozen stores and waited years for an answer that came back tangled up with everything else that happened in between. Your judgment is only ever as good as the feedback that trained it.

Most real estate teams turn this into a fight. Numbers people on one side, operators on the other, somebody cracking the same joke about the analyst who's never stood in a parking lot. But that framing treats a site decision like it's one judgment. It's two, and they don't have much in common.

Judgment gets trained by correction, not by tenure

The research on when to trust an expert is settled, more than you'd think. Kahneman and Klein spent years arguing about it and finally published their truce in 2009: judging whether an intuitive call is any good means looking at how predictable the environment is and how much chance the person had to learn its patterns (American Psychologist, 64(6), September 2009). Their sharpest line is the one nobody likes hearing: how sure you feel says nothing about whether you're right.

The other half of that picture is just as old. A meta-analysis stacking clinical judgment against mechanical prediction, across health and behavior, found formal statistical methods beat expert judgment by about 10% on average, and beat it substantially in 33% to 47% of the studies, depending how you slice it. Experts came out substantially ahead in only 6% to 16% (Grove, Zald, Lebow, Snitz and Nelson, Psychological Assessment, 12(1), 2000).

Put those two together and you get a practical rule. A simple model usually beats an expert on the average call. Someone working in a predictable environment with fast feedback still builds real skill there. So the question was never whether to trust your judgment. It's whether the judgment sitting in front of you belongs to the part of the job that's been correcting you all along.

The half your gut has been corrected on

Walk into any expansion meeting and the operators in the room call it right more often than the numbers people expect. There's decades of reps behind that instinct.

If you've been doing this fifteen years, you've stood on several hundred properties. Every one of them gave you an answer inside of ten minutes. The pad sits five feet below the road and the building's going to disappear. The sign won't read from the direction that matters. The left turn out of the lot is going to be miserable at five o'clock. The center's at ninety percent leased on paper and feels hollowed out on a Tuesday afternoon.

That's exactly what Kahneman and Klein are describing: stable regularities, hundreds of reps, feedback while you're still standing there. When operators say they need to walk a site in person before anyone signs, that's calibrated skill talking, not superstition, and any process that overrules them on that ground is throwing away the best instrument in the building.

Two-column comparison showing that judgments about the physical site are trained by hundreds of visits with feedback in minutes, while judgments about market and revenue get only a few dozen repetitions with feedback arriving years later.

The half it has never been corrected on

Now count the other kind of judgment.

How many times in your career have you predicted what a location would do in revenue? Not thrown out a guess in a meeting. Predicted it, written it down, then gone back and checked. For most people that honest number sits somewhere between zero and a few dozen, and the checking part almost never happened.

The feedback loop breaks in three places. The gap between the decision and the opening runs eighteen months to three years on a build. By the time real numbers show up, a manager's been hired and maybe replaced, a competitor's opened or closed, and a macro cycle has turned over. The signal's buried under noise you can't subtract. And on a construction timeline that long, whoever made the call has often moved to a new role, or a new company, before the answer ever comes in. Nobody's hiding a bad record here. The record was never scored to begin with.

Three judgments live in this half, and none of them ever gets graded:

  • Revenue. You find out what the store did. You never find out what it would've done at the site you passed on.
  • Cannibalization. The counterfactual is unobservable by definition. Sales at the older store dropped. You'll never know how much of that was the new unit and how much was the road work on the other side of town.
  • Market choice. You only see results from the markets you entered. The ones you skipped hand you no data at all, so a bias toward the markets you already know never gets corrected.

This is where writing down a real forecast earns its keep, and it's the same reason a committee wants to know where a number came from. A defensible forecast creates a record somebody can check later, and that record is the only way any of this half ever becomes learnable.

Where the data is the weaker instrument

The reverse is just as true, and it goes well past the physical property.

A model can't see the deal you can actually negotiate. It can't see which landlord returns your calls and which one's going to fight you over the TI allowance for six months. It can't see the road project that hasn't broken ground yet, or that the anchor two doors down is quietly shopping the space.

It also can't see who's going to run the store, and that one's worth more than most teams price it at. Research using store-level panel data from two multibillion-dollar retailers, where managers move between stores while the company's playbook stays fixed, found that individual managers explain a large share of the variance in store-level productivity (Metcalfe, Sollaci and Syverson, NBER Working Paper 31192, April 2023). A site model scores the ground. A meaningful piece of the outcome comes down to a person the model has never met.

Then there's the case where a standard input is measuring the wrong thing entirely. Walk-by counts predict a coffee shop pretty well. They say a lot less about a gym, a trampoline park, or a vet clinic, where the customer drives to you on purpose. If you run a destination category, your skepticism about foot traffic numbers is earned, and the fix is reweighting the input, not throwing out the whole exercise.

Back-test your gut the way you would back-test a vendor

Teams put a vendor's model through a validation run, watch it rank one known bad store too high, and drop it on the spot. Then they hand the exact same decision to an instinct that's never been through any validation at all.

That double standard is well documented. People lose confidence in an algorithm faster than they lose it in a person, after watching both make the identical mistake, and they'll pick a worse human forecaster over a better model because they saw the model screw up once (Dietvorst, Simmons and Massey, Journal of Experimental Psychology: General, 144(1), 2015). The model gets graded. The gut doesn't.

So grade it. The test takes an afternoon:

  1. Pull your last fifteen to twenty openings, far enough back that every one has a mature performance number.
  2. Dig up what you believed before each one opened. The pro forma, the committee deck, the email where somebody called it the best site in the market.
  3. Rank those sites the way you ranked them back then, most confident to least.
  4. Line that ranking up against actual performance and see how well the two agree.

If your pre-opening ranking tracks reality, you've found a real signal, and you should write down what you were reading. If it comes out close to random, your instinct was answering a question it had never been graded on, and that's worth knowing before the next committee meeting. Either way, run the same test on any model trying to sell you its service. The back-test protocol for a site score is that same procedure, pointed at a vendor instead of yourself.

The order does more work than the argument

Most teams that get this right never settled the philosophical question. They just fixed the sequence.

Data goes first, because it's the only thing that can look at every candidate parcel in a market, including the ones nobody in the room is championing. Judgment goes second, as a veto. You walk the shortlist and cross sites off it. What you don't do is add a site back on because it feels right, since that's exactly the case where your instinct is working the half of the job it was never trained for.

Run it backward and nothing looks broken from the inside. Somebody falls for a property, the analysis gets built around that property, and weak inputs get read charitably because the answer's already decided. Every step still looks like diligence. The list never had the site that would've won.

Two stacked workflow lanes contrasting a working order of operations, where market screening produces a shortlist and the site visit removes candidates, against a failing order where a chosen site gets analysis built around it after the fact.

This is the part of the site selection process worth being strict about, because the order's cheap to fix and the argument never ends.

Building the split into the workflow

A scoring tool only earns its keep here if you can see inside it. A number with no visible inputs gives your team nothing to push back on, which means the override happens quietly in somebody's head and leaves no trace behind.

GrowthFactor scores candidate sites across five configurable lenses, and every score opens up: the demographics, the foot traffic, the competitive proximity, and the trade area that moved it, weighted by your business format, not by us. Your team can overrule any of it. When they do, the ranking, the override, and the reason sit together on the deal record and travel to committee as one piece. Six months later, that record is what makes the back-test above even possible.

It runs the analysis. You make the call. The read you bring to a property, your take on a landlord, the manager you've already got in that city, the deal you can actually get: none of that is visible to a model, and it belongs to the people who've been collecting it for fifteen years.

Frequently Asked Questions about trusting your gut in site selection

Should I trust my gut or the data when choosing a store location?

Both, on different questions. Trust your instinct on the physical property, because you have walked hundreds of them and been corrected within minutes each time. Trust the data on market choice, revenue, and cannibalization, because those are judgments you have made a few dozen times at most and received an answer to years later, after weather, staffing, and the economy had all touched the result.

Why is gut feel unreliable for revenue forecasts but reliable for site visits?

Intuition is built by repetition plus fast, clear feedback. A site visit gives you both: you see the grade, the sightlines, and the traffic pattern, and you know within minutes whether the property works. A revenue forecast gives you neither. You make that call a handful of times in a career, and the answer arrives eighteen months to three years later, tangled up with every other thing that happened to the store.

What can a site selection model never see?

The deal you can actually negotiate, the landlord you would be signing with, the manager you already have in that market, the road project that has not broken ground, and the specific piece of ground under the building. Those are real inputs and they belong to your team. A model that pretends to price them is worse than one that leaves them out.

How do I test whether my own site judgment is any good?

Back-test it the way you would back-test a vendor. Pull your last fifteen to twenty openings, find your original pre-opening ranking of those sites, and compare it against actual performance. If the ranking was near random, your instinct was answering a question it had never been graded on. Run the same test on any model before you trust it.

How does GrowthFactor compare to Placer.ai when the call still comes down to judgment?

Placer.ai is a foot traffic data subscription and the category standard for visit trends at the property and brand level. It gives your team a strong input and leaves the analysis to them. GrowthFactor scores candidate sites across five configurable lenses, shows every input that moved the score, and keeps the score, the trade area, and the deal in one pipeline, so the point where a person overrides the ranking is recorded rather than lost. Neither one makes the decision. The difference is whether the reasoning behind your decision survives to the committee meeting.

Share

Continue reading

Does Foot Traffic Data Matter for a Gym or Trampoline Park?

Walk-by counts predict coffee shops. They do not predict a gym, a trampoline park, or a vet clinic. Here is where foot traffic data stops working for destination categories, and the four jobs it still does well.

Aug 21, 2026

Best Location Intelligence Software for Retail Teams (2026)

The best location intelligence companies ranked for retail and CRE teams in 2026, evaluated on data depth, scoring transparency, deal workflow, and setup speed.

Aug 20, 2026

7 Buxton Alternatives for Retail Site Selection (2026)

Looking for Buxton alternatives? Compare 7 platforms for retailers who want customer analytics depth without a months-long engagement, including self-serve scoring and deal management.

Aug 20, 2026

Newsletter

This Week in Retail

Store closures, expansion tracking, and original market analysis. A five-minute read every Thursday.

Ask GrowthFactor where to open next

Watch it pull the data, run the analysis, and explain the answer in maps and tables. It does the analysis. You make the call.