
Why Silent Customer is changing the traditional mystery shopping model

Rich Green | Commercial Director, Silent Customer
I’ve been on a bit of a journey over the last 18-odd years of my career. Like most people in market research, I didn’t exactly choose it as a career path; it kind of found me. Market research as a whole is going through a pretty significant period of change. I won’t use the word that starts with A and ends in I too much, but there is no getting away from the fact that technology is shaking up the industry and forcing us to rethink how research is delivered.
I’ve also had the joys of reaching mid-career and, for a period, stepping away from mystery shopping, something I’ve spent about 90% of my career doing, to work across different research methodologies. When I came back to mystery shopping, the thing that genuinely staggered me was how little had changed. The platforms looked a bit different. But fundamentally, the model was the same: a questionnaire, a set number of visits, a score, a league table, a dashboard and a report. Then do it all again next wave.
That genuinely fired me up. Customers have changed, workplaces have changed, technology has changed and the way people consume information has changed. So why should mystery shopping still be delivered in essentially the same way it was ten or twenty years ago?
That question has shaped a lot of what we are now building at Silent Customer. I don’t believe mystery shopping is broken. Far from it. I think it remains one of the most useful ways of understanding what is actually happening in a customer journey. But I do think the traditional model needs to change. And that starts with remembering what mystery shopping actually is.
Mystery shopping needs to remember that it is research
Over the years, mystery shopping has developed something of an identity crisis. It has been positioned as customer experience measurement, compliance, auditing, training measurement, benchmarking and plenty in between. In trying to make it do everything, we have sometimes lost sight of what it is particularly good at.
At its best, mystery shopping is a tactical observational research tool. It lets a business see what actually happens during any customer journey (be it online, calls, social media, or physical locations). Did the intended experience happen? Were the behaviours you expect from teams demonstrated? Where did the journey break down? Was an opportunity to help, reassure, recover or sell missed?
That means programme design is about much more than writing a questionnaire. Start with: what are we actually trying to understand or change? Then decide what behaviours to observe, what evidence will be useful, where to look and how often to measure it.
Crucially, the programme should not necessarily stay the same once you start learning. If one part of the journey is consistently working well, do you need to keep measuring it with the same intensity? If another part is creating the problem, shouldn’t more of the next wave focus there? That might mean changing the questions, sample, locations, journey or frequency.
If you want improvement, measurement needs to follow the issue.
Too many programmes repeat the same visits, at the same frequency, using the same questionnaire simply because that is how they were designed at the beginning. That isn’t really a research approach. It is a schedule. A better programme should learn: measure, understand where the journey is strongest and weakest, act on the evidence, then adapt the next measurement around what you are trying to improve. Sometimes that means a large ongoing programme. Sometimes it means 20 targeted journeys. Sometimes mystery shopping is not the right methodology at all. A good provider should be prepared to tell you that.
Your questionnaire is probably too long
This is one of my biggest frustrations with the industry. Mystery shopping sometimes feels like the only research methodology where adding another 20 questions is automatically treated as adding more value. Usually, the opposite is true.
We are asking a shopper to behave naturally while observing, remembering and accurately reporting what is happening around them. The longer the list, the greater the risk that evidence quality suffers. Long questionnaires also mean longer visits, more reporting, validation and clarification. It is entirely possible to end up paying more to collect poorer quality information.
The bigger problem comes afterwards. If a location receives feedback identifying 27 things it could improve, what exactly do we expect the manager to do with that on Monday morning?
If everything matters, nothing matters.
There is another consequence of over-designing these programmes: they can start to lose the element of mystery. The point of a mystery visit is not to catch teams out. It is to see the journey naturally, as a customer would experience it, without the measurement itself changing the behaviour.
When teams know a huge questionnaire inside out, spend time picking apart every individual answer and end up debating whether a visit should have scored 82% or 83%, the programme starts to become a test to be managed rather than evidence to learn from. A one-point movement might make a league table look different, or help them achieve a KPI but it rarely tells you what somebody should do differently tomorrow.
And this should not become an exercise in finding faults either. Understanding what teams did really well is just as important. Good behaviours need to be recognised, protected and repeated, particularly when the evidence shows they are shaping a stronger customer experience.
The goal is not to win an argument over a score. It is to understand what to keep doing, what to improve and what deserves more attention next time.
A better programme focuses on the moments and behaviours most likely to influence the outcome. Measure those properly, make the priorities clear, act on them and measure again. If your questionnaire contains 60 or 70 questions, there is a decent chance you are asking one customer journey to answer several business problems at once.
Your people don’t need more data. They need the point.
How we deliver insight needs to change too. Younger employees increasingly value practical, on-the-job learning and guidance. Deloitte’s 2025 Gen Z and Millennial Survey found 88% of Gen Z respondents valued on-the-job learning and practical experience, while 86% highlighted mentorship and guidance.
But this is not simply a Gen Z issue. Most of us are just incredibly busy. Microsoft’s 2025 Work Trend Index found that 80% of the global workforce reported lacking enough time or energy to do their work. So why give managers another portal, dashboard or long report to work through before they can even find the important bit?
Give people the facts, explain why they matter and make the next action obvious.
This is at the forefront of Silent Customers’ offering – A frontline manager might need one or two behaviours to discuss tomorrow. A regional manager may need to know which locations are driving a pattern. Learning and Development might need to know what requires coaching, while senior leaders may simply need the three findings that matter. The evidence can be the same. The way we deliver it should not be.
The industry’s operating model needs to change too
I’ve now worked across five mystery shopping businesses, and one thing they have all had in common is how people-heavy the industry is. People are involved at almost every stage: setup, allocation, shopper queries, validation, data processing, reporting and all the administration around the research itself.
Some of that human involvement is essential. Mystery shopping still needs experienced people making judgements, supporting shoppers and interpreting evidence. But much of the work surrounding it does not need to be as manual as it traditionally has been.
This is where technology and AI should play a bigger role. Not as a badge added to a report, but as part of the infrastructure of the business. At Silent Customer, we have built technology, automation and AI into the way we operate so repetitive work happens faster, more consistently and with less manual intervention.
There is an obvious commercial benefit. A more efficient operating model helps us remain highly competitive without simply adding overhead as we grow. More importantly, our people can spend more time designing the right programmes, challenging what is measured, interpreting the evidence and working with clients on what should happen next.
Less time administering mystery shopping. More time making it useful.
The score is a signal, not the outcome
Scores still have a role. They can show where performance differs, where something improved and where there may be a problem. But they should be a signal that helps us decide where to look, not the reason the programme exists.
That is the direction we are taking at Silent Customer. Start with the outcome, identify the moments and behaviours that shape it, capture reliable evidence, turn it into practical action and re-measure.
Crucially, the next measurement should respond to what you learned rather than simply repeat the previous wave for the sake of it.
So one final question is worth asking: has your mystery shopping programme materially changed in the last five years?
Not the dashboard/portal or the branding, the programme itself. Does your supplier challenge what you measure, help you decide what not to measure and adapt the next wave around what the evidence is uncovering? Are they helping teams focus on what to protect as well as what to improve? Are they making insight easier to use and using technology to make delivery faster and more efficient without losing the human judgement that matters?
Mystery shopping is not broken. It remains one of the best ways of understanding whether the customer experience you intended is the one actually being delivered in the real world. But the world around it has changed. If your programme has not changed with it, it is probably worth asking whether you are getting everything from it that you should be, and perhaps whether you have been doing your mystery shopping all wrong.
Measure behaviour. Change outcomes.