Who Is Allowed to Do Research Is the Wrong Question
Democratisation, expertise, and the case for governing evidence chains rather than access
The democratisation of design research is usually framed as a question of permission. One side argues that research should be opened because demand for insight comfortably exceeds the supply of researchers, and because teams working closest to users often make better decisions when they can gather evidence directly. The other side argues that research should be protected because, without appropriate training, errors can easily be presented as credible findings. Although these positions appear opposed, they are focused on the same variable: who is allowed to hold the research instrument.
That is an understandable question, but I believe it is close to the wrong one. The available evidence on research reliability does not resolve the permissions debate in either direction. Instead, it directs our attention to the system in which evidence is produced, interpreted and used, rather than to the credentials of the person who produced it. This distinction matters commercially as well as methodologically. McKinsey's design index research analysed more than two million pieces of financial data and over 100,000 design actions across three industries. It found that top-quartile performers achieved revenue growth 32 percentage points higher than industry peers over five years. One of the strongest correlations was with organisations that broke down functional silos, rather than concentrating design capability in one place (Sheppard et al., 2018). The commercial case therefore favours distribution, while the quality case is often presented as an argument for concentration. To reconcile the two, we need to be precise about what actually makes research reliable.
The uncomfortable evidence about expertise
In 2018, Raphael Silberzahn, Eric Uhlmann and a large group of collaborators published the results of an unusual experiment. Twenty-nine teams, involving 61 analysts, received the same dataset and the same research question: were football referees more likely to give red cards to players with darker skin tones? These were experienced researchers working in good faith. However, their estimated effects ranged from 0.89 to 2.93 in odds-ratio units, with a median of 1.31. Twenty teams found a statistically significant positive effect, while nine did not. Across the 29 analyses, the teams used 21 different combinations of covariates.
The most important finding is not simply that the results varied. It is that the expected explanations did not account for that variation. The authors found that the analysts' prior beliefs, their level of expertise and peer assessments of analysis quality did not explain the spread. In other words, expertise did not make the answers converge.
Figure 3. Effect estimates from 29 independent expert teams analysing one dataset to answer one question.
Four years later, Nate Breznau and colleagues conducted a larger version of the same experiment. They coordinated 161 researchers in 73 teams to test one hypothesis using identical cross-country survey data: that greater immigration reduces public support for social policy. Again, the results did not converge. They ranged from large negative effects to large positive effects, and the teams' statistical design choices explained very little of the variation. The researchers described what remained as a 'hidden universe of uncertainty'. Their recommendation was not tighter credentialing, but greater humility.
Neither study concerns inexperienced researchers. Both show what can happen when trained people make reasonable, defensible choices in complex analyses. The conclusion is not that experts are unreliable. It is that several defensible analytic paths can produce several defensible answers, and seniority alone does not eliminate that variation. If an organisation governs research mainly by asking whether the person conducting it is qualified, it is relying on a control that the strongest available evidence suggests cannot carry the full burden placed upon it.
What the evidence tells us about non-specialists
The evidence is equally challenging for a purely protectionist position. Aceves-Bueno and colleagues (2017) reviewed more than 1,300 comparisons between data collected by volunteers and data collected by professionals. Among studies that reported p values, citizen-science data differed significantly from professional data in 38% of cases. Importantly, accuracy was not random. It was associated with identifiable features of the research design, including whether participants received training, how long they had participated, the size of the group and whether they had a personal stake in the outcome.
Kosmala and colleagues (2016) reached a similar conclusion from a different direction. Volunteer projects that generate accurate data tend to use a recognisable set of methods: iterative project development, volunteer training and testing, expert validation, replication across volunteers and statistical modelling of systematic error. The authors recommended judging each dataset on its design and intended application, rather than assuming it is inferior because volunteers produced it.
There are limits to how far we should generalise from these studies. They concern ecological observation, where counting, identifying and measuring are structured tasks with an external reference point. That is not equivalent to interpreting forty hours of qualitative interviews. What does transfer tell us is the underlying mechanism: accuracy was associated with protocols, training, validation, and replication, not simply with the employment status of the person collecting the data.
Participation, however, is not automatically evidence of quality. Slattery, Saeri and Bragge at BehaviourWorks Australia, Monash University, conducted a rapid overview of reviews of research co-design in health. They found that co-design is widely used but inconsistently described, rarely evaluated in detail and seldom tested empirically or experimentally. The available evidence suggests benefits for researchers, practitioners, research processes and outcomes, but it has not yet established those benefits to the standard we would expect of other interventions. Co-design may be ethically appropriate, as it often is, but it is not automatically methodologically superior. Those are separate claims and require separate evidence. The cases for opening research completely and for locking it down are therefore both less settled than their strongest advocates sometimes suggest.
Why weak research rarely announces itself
If credentials are not the decisive control, where does research failure occur? Kara Pernice's account for Nielsen Norman Group identifies familiar problems: selecting a method that cannot answer the question, recruiting the wrong participants, asking leading or overly narrow questions, probing when observation is required, analysing poorly, and forcing quantitative measures onto qualitative work. She also identifies a less visible problem: when no one has explicit responsibility, accountability disappears. Studies are then duplicated across teams or quietly abandoned (Pernice, 2022).
Practitioner Carl Pearson captures why these failures can be difficult to detect: bad research does not stink. A poorly moderated session still creates a transcript. A biased sample still produces a chart. A weak analysis still results in a confident slide, and to a reader without methodological training, that slide may look indistinguishable from a robust one. Many forms of delegated work produce a visible error signal: Code throws an exception and broken interfaces generate complaints. Research can produce fluent, persuasive output even when the reasoning is unsound. Quality must therefore be built into the process, not inferred from the polish of the final deliverable.
This reframes the governance problem. The central risk is not that too many people are conducting research; it is that many organisations cannot reconstruct how a conclusion was reached. If a finding cannot be examined or challenged, its reliability remains unknown, regardless of who produced it.
Framework one: the evidence chain
A study is not a single act of judgement. It is a chain of five connected decisions, and each link can fail independently. Each also has a control designed to catch a particular type of failure. Governing this chain is far more practical than attempting to govern people through job titles alone.
Figure 1. The evidence chain, with the characteristic failure and the matching control at each link.
This framework separates two elements that organisations often conflate: skill and effort. Skill is most concentrated in links two and four, design and interpretation. Effort is most concentrated in links one and three, question framing and collection. When leaders talk about democratising research, they often mean distributing effort. When researchers raise concerns, they are usually protecting the quality of design and interpretation. Once the chain is visible, the discussion can be resolved link by link, rather than treated as a contest over professional status.
Every link should leave an artefact behind. These artefacts allow someone who was not in the room to reconstruct and challenge the reasoning later.
Table 1. The five artefacts of a traceable study.
A research degree is not required to produce these artefacts. Four of the five need little more than a clear template and the discipline to document decisions before, rather than after, the work. Interpretation is the exception because it carries substantial methodological weight. However, many-analyst studies also show that interpretation is vulnerable even in expert hands when conducted alone and left undocumented.
Framework two: the contestability matrix
If oversight is not allocated by job title, it needs another basis. I use two properties of the work itself. The first is the cost of reversing the decision informed by the research. The second is the degree of inference between the data and the conclusion: how many interpretive steps sit between what was observed and what is ultimately claimed.
Figure 2. The contestability matrix. Governance is allocated by the properties of the work rather than the seniority of the person.
Both axes describe the work, not the worker. This distinction leads to two conclusions that organisations may find uncomfortable, but for very different reasons.
First, a considerable amount of research can be opened up responsibly. A copy test supporting a reversible interface change, conducted through a shared template and stored where others can find it, belongs in the open zone. It does not need a researcher's approval to proceed. Sending that study through the same governance queue as a pricing study consumes scarce specialist capability on low-stakes work. That is precisely the kind of misallocation Pernice warns against.
Second, seniority does not provide an exemption. A principal researcher who analyses sensitive qualitative material alone, for a decision that will be expensive to reverse, is operating in the reserved zone. The same controls should apply to anyone working there: an analytic protocol documented before analysis and a second analyst involved in contested judgements. The lessons from Silberzahn and Breznau apply to experts as well. Many governance models exempt the very people whom the evidence suggests should remain within the control system.
Framework three: measure traceability, not throughput
Research-democratisation programmes are commonly assessed through volume: studies completed, participants recruited, teams trained and time from request to readout. These measures have an uncomfortable feature: they improve most quickly when quality controls are removed. Volume is also where failure can remain invisible. A more useful set of measures asks whether the evidence chain remained intact.
Table 2. Measures that track whether the evidence chain held.
Throughput measures still have value, but they belong in a capacity report rather than a quality report. Treating them as indicators of research health can reward the production of confident, unreconstructable findings at scale.
The obligations that remain
There is also an Australian dimension that in-house teams can easily overlook. Research obligations attach to the activity, not to the job title of the person performing it. AS ISO 20252:2019 establishes requirements for how research studies are planned, conducted, supervised and reported. The Research Society's Code of Professional Behaviour applies to the professional activities of members and corporate partners. Partner organisations are expected to ensure that the people they employ or engage comply, whether or not everyone is a member. The Privacy (Market and Social Research) Code 2021 also governs how personal information collected through research is handled.
A product manager conducting an unstructured customer interview is still conducting research. Consent, incentive handling, data retention and the obligation not to mislead participants about the purpose of the conversation continue to apply. In a democratised model, these responsibilities do not disappear; they are distributed. The organisation must therefore make them clear and workable for everyone who inherits them. The Research Society and the Australian Data and Insights Association began a joint review of the Code in 2025. This is a timely prompt for organisations to consider whether their internal research practices would withstand the same level of scrutiny.
What AI changes, and what remains the same
Until recently, the effort required to conduct research placed a natural limit on democratisation. That constraint has largely disappeared. Anyone can now generate a persona, thematic summary or set of research questions in less than a minute, and the output is fluent enough to circulate without immediate challenge. AI does not introduce an entirely new failure mode; it removes much of the friction that previously slowed an existing one.
Nielsen Norman Group takes a narrower position on these tools than many vendors, and it is a useful one. AI is most valuable in the planning and analysis stages of research, not as a substitute for the study itself. Synthetic users should be treated as sources of hypotheses that still require testing with real people, not as findings in their own right (Moran & Rosala, 2024; Rosala & Moran, 2024). When mapped against the evidence chain, a consistent pattern emerges.
Table 3. AI against the evidence chain.
The recurring weakness is provenance rather than accuracy alone. AI systems can produce a coherent answer, but they do not volunteer a reliable account of the alternatives they discarded along the way. That missing account is exactly the artefact required at link four. An organisation with a functioning evidence chain can adopt these tools relatively quickly because the control point already exists. Without that chain, it becomes difficult to distinguish genuine analysis from persuasive autocomplete.
Three practical moves
First, classify decisions before classifying people. Review the last twenty decisions informed by research and place each one on the two axes in Figure 2. Many organisations will find that substantial low-risk work sits in the open zone but is over-governed, while a smaller amount of high-risk work sits in the reserved zone and is assigned to whoever happens to be available. Correcting that allocation requires no new technology and can release capacity immediately.
Second, instrument the chain rather than the people: five artefacts, five owners and one shared location. Begin with the decision statement at link one and the alternative explanation at link four. These two artefacts provide most of the diagnostic value and can be introduced without new tooling.
Third, run one visible test of contestability. Select a finding that the organisation currently regards as settled and ask a second person to reconstruct it using only the recorded artefacts. Whether they reach the same conclusion is not the only point. The exercise demonstrates that findings are open to examination. That visible practice is more likely to build a strong research culture than another training module on its own.
The question beneath the debate
Democratisation and gatekeeping are both responses to a question about permission. The evidence points to a more useful question about traceability. Knowing whether a study was conducted by a principal researcher or a product manager tells us relatively little about whether its conclusions will hold. Knowing whether the reasoning can be reconstructed, challenged and, where necessary, overturned tells us considerably more.
I would therefore put one focused question to any leadership team drafting a research-democratisation policy: of the last five decisions your organisation made on the strength of research, how many could a capable colleague outside the project reconstruct from the written record, and how long would it take? That answer will reveal more about the health of the research function than the number of studies it completed.
References
Aceves-Bueno, E., et al. (2017). The accuracy of citizen science data: A quantitative review. Bulletin of the Ecological Society of America. https://doi.org/10.1002/bes2.1336
Breznau, N., Rinke, E. M., Wuttke, A., et al. (2022). Observing many researchers using the same data and hypothesis reveals a hidden universe of uncertainty. Proceedings of the National Academy of Sciences, 119(44), e2203150119. https://doi.org/10.1073/pnas.2203150119
Kosmala, M., Wiggins, A., Swanson, A., & Simmons, B. (2016). Assessing data quality in citizen science. Frontiers in Ecology and the Environment, 14(10), 551- 560. https://doi.org/10.1002/fee.1436
Moran, K., & Rosala, M. (2024, 27 September). Accelerating research with AI. Nielsen Norman Group. https://www.nngroup.com/articles/research-with-ai/
Pearson, C. J. (2025). Bad research does not stink. https://carljpearson.com/bad-research-doesnt-stink/
Pernice, K. (2022, 12 June). Democratise user research in 5 steps. Nielsen Norman Group. https://www.nngroup.com/articles/democratize-user-research/
Rosala, M., & Moran, K. (2024, 21 June). Synthetic users: If, when, and how to use AI-generated research. Nielsen Norman Group. https://www.nngroup.com/articles/synthetic-users/
Sheppard, B., Kouyoumjian, G., Sarrazin, H., & Dore, F. (2018, October). The business value of design. McKinsey and Company. https://www.mckinsey.com/capabilities/tech-and-ai/our-insights/the-business-value-of-design
Silberzahn, R., Uhlmann, E. L., Martin, D. P., et al. (2018). Many analysts, one data set: Making transparent how variations in analytic choices affect results. Advances in Methods and Practices in Psychological Science, 1(3), 337 to 356. https://doi.org/10.1177/2515245917747646
Slattery, P., Saeri, A. K., & Bragge, P. (2020). Research co-design in health: A rapid overview of reviews. Health Research Policy and Systems, 18(1), 17. https://doi.org/10.1186/s12961-020-0528-9
Standards Australia. AS ISO 20252:2019, Market, opinion and social research, including insights and data analytics.
The Research Society. Code of Professional Behaviour. https://www.researchsociety.com.au/code-of-professional-behaviour/
The Research Society and Australian Data and Insights Association (2025). Review of the Code of Professional Behaviour.
I am a research strategist who partners with businesses, technology organisations, and SaaS teams to turn research into clear strategic direction and measurable impact. I work hands-on across the full research lifecycle, spanning academic, consulting, industry, and policy research, with a strong focus on evidence-based decision-making.
My expertise brings together data-driven insights, UX and CX (Voice of Customer) research, strategic research planning, stakeholder engagement, and robust survey and measurement design. Using a mix of primary and secondary research methods, I help organisations move beyond surface-level insights to understand what truly drives customer behaviour, product adoption, and long-term value.
As a customer-centred, insights-led researcher, I focus on uncovering human behaviours, habits, motivations, and attitudes to help teams design products, services, and strategies grounded in real-world needs. I’m particularly drawn to emerging technologies and SaaS environments, where strong research can shape how people learn, work, and interact at scale.
Beyond UX and customer research, I bring deep experience in strategic research program design, vendor management, and cross-sector collaboration. I work closely with senior stakeholders and interdisciplinary teams to ensure research findings are translated into actionable strategy, product roadmaps, and policy-ready recommendations.
I actively contribute to the research and technology community through thought leadership, including writing on Medium and publishing a LinkedIn newsletter focused on research practice and emerging industry trends. I also partner with SaaS companies to evaluate user research platforms and capabilities, providing practical, real-world feedback that informs product innovation.
Forward-thinking and outcomes-focused, I bridge academic rigour, industry innovation, and strategic insight to help organisations build better products, make confident decisions, and deliver meaningful customer experiences.