User Research Methods: A Complete Guide to Choosing and Running the Right Ones
Published:
Reading time:
Category:

Content
Forrester's research puts the ROI of UX at 9,900 percent: every dollar invested in user experience design yields approximately $100 in return. Maze's 2025 Future of User Research Report offers an equally direct number: teams that integrate user research into product decisions see 2.7 times better business outcomes, including higher revenue and improved retention, compared to teams that rarely use it. And yet only 3 percent of organisations reach the highest stage of research maturity. (Nielsen Norman Group cautions against treating those aggregate ROI figures as precise, and that caution is worth carrying through this guide.)
The gap between knowing research matters and knowing which method to use, and how to run it well, is where most teams lose ground. They pick the familiar method, usually a survey, or the method that sounds most rigorous, usually a focus group, rather than the method best suited to the question they are trying to answer. This guide closes that gap.
What follows is a complete working guide to user research methods: the three-dimensional framework for choosing the right one, deep coverage of twelve methods across generative and evaluative categories, the six-step research process, the bias catalogue, and how to turn raw data into design decisions. It is grounded in Erika Hall's Just Enough Research, one of the most practical books written on the subject, and calibrated against Nielsen Norman Group's method framework and current industry practice. If you would rather start by auditing an existing product than researching a new one, the 47-point UX audit checklist is the companion piece to this guide.
What User Research Actually Is (and Is Not)
Erika Hall defines research simply: systematic inquiry. You want to know something, so you go through a process to find out. The type of process depends on what you need to know and what decisions that knowledge will inform.
The most important thing to get clear before picking a method is what research is not. Research is not asking people what they like. Like is a superficial, self-reported mental state with no reliable connection to behaviour, because people habitually engage in activities they claim to hate and do not do things they claim to love. Research is not a political tool to justify a decision that has already been made. And applied design research is not science: you do not need statistical significance in a qualitative interview, you need useful insights.
Most consequentially, research is not asking people what they want. Henry Ford's apocryphal faster-horses line contains a genuine insight. Users can accurately report their current behaviours and their frustrations with the present state of things. They cannot accurately predict what would solve those frustrations, because they can only want what they can imagine. The job of research is to gather the raw material, meaning behaviours, contexts, needs, and mental models, and let the designer do the imagining.
The Three-Dimensional Framework for Choosing a Method
Nielsen Norman Group's framework for UX research methods uses three axes, each representing a different dimension of what a method produces.
Attitudinal versus behavioural. Attitudinal methods measure what people say: their stated beliefs, preferences, and self-reports. Surveys, interviews, and focus groups all produce attitudinal data, and its limit is that people are unreliable reporters of their own behaviour. Behavioural methods measure what people actually do. Usability testing, contextual inquiry, analytics, and diary studies produce behavioural data, which is harder to collect but more reliable as a predictor of real-world performance. The most important research insights almost always come from the gap between what people say and what they do.
Qualitative versus quantitative. Qualitative research produces depth, working with small samples and answering why and how. Quantitative research produces breadth, answering how many, how often, and which is better. Neither is inherently superior. Hall's caution is worth repeating: resist arguments about statistical significance when doing qualitative work, because qualitative research is valid not for its sample size but because its questions are about understanding rather than measurement.
Context of use, scripted versus natural. Some methods observe people in a controlled, scripted environment, with usability testing the primary example. Others observe people in their natural context doing things they were already going to do, as with contextual inquiry and diary studies. The further from natural context, the more the data reflects the research environment rather than the real world.

The Four Types of Research and When to Use Each
Hall's taxonomy maps research to the stage of the design process and the type of question being asked. Each type uses different methods, but the boundary between them is permeable, because the same activity can serve generative or descriptive purposes depending on what you are trying to learn.
1. Generative research: what's up with this? The research you do before you know what you are designing. Its purpose is to surface problems worth solving and replace assumptions with real observations. Methods: in-depth interviews, contextual inquiry, diary studies, literature review. A generative interview with parents about weekend planning might reveal that the problem is not too few events but no reliable reminder mechanism, which is a completely different design target.
2. Descriptive research: what and how? You have identified the problem and now need to understand the context well enough to design the right solution. It is the homework phase: understanding enough about the domain and audience to design for them rather than for yourself. Methods: interviews, contextual inquiry, surveys as background data, existing literature.
3. Evaluative research: are we getting close? Once you have a sketch, prototype, or live product, evaluative research tests whether it works. This is the most common type in practice, with 84 percent of organisations running usability testing, making it second only to user interviews. Methods: usability testing, heuristic analysis, A/B testing, first-click tests. The key principle is early and often.
4. Causal research: why is this happening? Analytics tell you what happened, not why. Causal research investigates the relationship between a design decision and a measurable outcome, and it is what you do after noticing an unexpected pattern in your data. Methods: analytics analysis, A/B testing, session recording review, support ticket analysis.

The Twelve Methods: A Quick Reference
The table below summarises the twelve methods covered in detail through the rest of this guide. What follows each entry is the full treatment: the rationale, the preparation, the conduct, and the common failure modes.
Method | Type | Question it answers | Time required | Output |
|---|---|---|---|---|
User interviews | Qual / attitudinal-behavioural | Why do users behave this way? What are their goals and mental models? | 30–60 min per session | Quotes, patterns, mental models |
Contextual inquiry | Qual / behavioural | How do users perform tasks in their real environment? | 1–2 hours on-site | Observed behaviours, task flows |
Diary studies | Qual / behavioural | How does behaviour change over time and across contexts? | 1–4 weeks | Longitudinal logs, attitude shifts |
Usability testing | Qual / behavioural | Can users complete tasks? Where do they fail and why? | 60–90 min per session | Task success rates, pain points |
Unmoderated testing | Qual-quant / behavioural | Can users complete tasks at scale without a facilitator? | 15–30 min per participant | Completion rates, heatmaps |
Heuristic analysis | Qual / expert review | Does this interface violate established usability principles? | 2–4 hours, 2–3 evaluators | Ranked violations and fixes |
Surveys | Quant-qual / attitudinal | What do users think, prefer, or do at scale? | 10–20 min per respondent | Statistical data, satisfaction scores |
Card sorting | Qual-quant / attitudinal | How do users mentally organise concepts and labels? | 20–40 min per participant | Category groupings, label preferences |
Tree testing | Quant / behavioural | Can users find things within a proposed IA structure? | 10–20 min per participant | Success rates, path analysis |
A/B testing | Quant / behavioural | Which version performs better for a defined metric? | Days to weeks | Conversion rates, winning variant |
Analytics review | Quant / behavioural | What are users actually doing in the live product? | Hours to days | Funnel data, flow analysis |
Competitive audit | Qual / comparative | What do users already expect from this category? | 4–8 hours | Baseline conventions, opportunities |
Method 1: User Interviews
User interviews are the foundational method of design research. Hall calls them the most effective way to get inside another person's head and see the world as they do, and they are also the most-used, with 86 percent of research practitioners running them. Their advantage is depth: a good interview produces understanding of goals, habits, mental models, priorities, and barriers that no other method provides.
The right format for design research is semi-structured: you have prepared questions and topics, but not a rigid script. This lets you follow the threads that matter and surface things you did not know to ask.
The first rule: never ask what they want. When asked directly, people give answers shaped by what they think you want to hear, what makes them look good, and the limits of their own imagination. Instead ask about behaviour: what they do, how, when, and what happens when it does not work. Walk me through what happened the last time you tried to do that produces more useful data than what would make this easier?
Preparation. Prepare a guide, not a script: the study goal, two or three demographic questions, warm-up questions, and the core questions. Good core questions are open-ended, specific to behaviour rather than preference, and designed to invite storytelling. Do background work on the domain first, because a blank-slate interviewer cannot distinguish a routine response from an important revelation.
Structure, in three acts. Introduce yourself warmly, explain the goal without revealing specifics that would influence responses, and confirm consent to record. In the body, ask open questions and allow pauses, because silence is useful and the participant will fill it. To conclude, ask whether there is anything else, then stay quiet for a moment after the recording stops, because people often share something important when they think the session has ended.
How many participants? Five to eight per user type is a practical starting point. You reach thematic saturation, the point at which new participants confirm patterns rather than reveal new ones, faster than you might expect. Five genuinely representative interviews are worth more than fifteen poorly screened ones.

Method 2: Contextual Inquiry
Contextual inquiry takes the interview into the participant's natural environment: their office, home, or wherever the activity actually happens. Intuit's Scott Cook famously followed customers home from Staples to watch them install Quicken on their own computers, and learned more in one of those visits than months of focus groups would have produced.
The advantage over interviews is the visibility of unconscious behaviour. People do not reliably report habits so ingrained they have forgotten them. A workaround so automatic it has never been consciously noticed will appear plainly in a contextual inquiry session. Janky hacks, scribbled sticky notes, multi-app workflows: the field surfaces what the interview cannot.
To run it, arrange to observe the participant doing the specific activities you are studying, in the environment where they normally do them. Establish rapport first with a short interview about the activity and its context, then observe and take detailed notes. Ask brief questions when something interesting happens, but prioritise watching over talking. At the end, summarise what you observed and ask whether your summary is accurate, because their corrections are often as valuable as the observation itself.
Method 3: Diary Studies
Diary studies ask participants to log their own experiences over days, weeks, or months, capturing behaviour and attitudes at the moment they occur rather than in retrospect. They are particularly valuable for activities that happen sporadically, vary significantly across contexts, or involve habit formation over time.
The method's main strength is temporal coverage. A single interview can only capture a retrospective account, subject to all of memory's distortions. A diary study captures the actual moment a user first encounters a problem, the point at which frustration peaks, and the workaround they develop, in sequence as it happened. Slack used diary studies and prototype testing to improve team onboarding, reporting a 35 percent faster time-to-value and a 9-point NPS improvement.
The main challenge is participant compliance, because daily logging requires sustained motivation and low friction. Short, specific prompts sent at moments relevant to the research activity perform better than open-ended daily journals. The same sequencing logic applies when designing onboarding itself, which is covered in the 10 mobile app onboarding patterns guide.
Method 4: Focus Groups, and Why Not to Use Them
Focus groups appear often enough in briefs and client requests that it is worth addressing directly why the design research field largely rejects them.
The focused group interview was developed by sociologist Robert Merton, who later deplored how they came to be misused. The problem is structural: a focus group creates an artificial group dynamic that does not resemble any real-world context in which your product would be used. The conversation becomes a performance, subject to social desirability bias and group conformity effects, where powerful voices shape the discussion.
Hall is direct: focus groups are research theatre. They give the appearance of user insight without the substance of individual, contextualised observation. And a single bad participant can taint an entire session in a way that a bad one-on-one interview cannot. If stakeholders push for focus groups as what research looks like, redirect toward individual interviews. You will learn more, faster, from five thirty-minute semi-structured interviews than from a two-hour focus group of eight people.
Method 5: Usability Testing
Usability testing is a directed session with a representative user attempting specific tasks on a prototype or live product. Hall's framing is useful: usability is necessary but not sufficient. A product can be usable and still fail in the market, but if it is not usable it will fail regardless of its other qualities. Nielsen's five components of usability define what you are testing for: learnability, efficiency, memorability, errors, and satisfaction.
When to run it: early and often. The most expensive usability testing is the kind done right before launch, when problems are expensive to fix. The second most expensive is the kind your customers do for you after launch, via support tickets. Testing paper sketches catches problems cheaply; testing a launched product catches them expensively.
How many participants? Jakob Nielsen's foundational research established that five participants per user type uncover approximately 85 percent of usability problems. Adding more returns diminishing insights. The practical implication is to run tests frequently with small samples rather than occasionally with large ones.
No labs required. Labs give the illusion of control while removing the unpredictability of the real world: the distractions, the glare, the screaming children in the background. Those distractions are the conditions your product will actually be used in, so test in context wherever possible.
The task is everything. Good tasks are realistic scenarios, not instructions. You are planning a family visit to the science centre next weekend, so use this site to find what is on and buy tickets for two adults and one child is a good task. Click on the Events section and find a ticket purchasing option is an instruction that bypasses the comprehension challenge.
Facilitating: what not to do. Stay neutral, which is harder than it sounds when it is your design being tested. Do not help. Do not hint. Embrace the uncomfortable silence when a participant stares at the screen, because that silence is data. When a participant blames themselves, redirect gently: you are giving us exactly the information we need, so can you describe what you expected to happen?
Method 6: Heuristic Analysis
Heuristic analysis is an expert review of a design against established usability principles. Two or three evaluators independently assess a product using a checklist, then combine their findings. It requires no recruiting, which makes it fast and cheap, and therefore useful as a quick check between user studies or at the very beginning of a project. Nielsen and Molich's ten usability heuristics have been the standard since 1990 and remain valid: system status visibility, match between system and real world, user control and freedom, consistency and standards, error prevention, recognition rather than recall, flexibility and efficiency of use, aesthetic and minimalist design, help users recover from errors, and help and documentation.
Multiple heuristics concern error prevention and recovery, which remains the most neglected area of system design. Every unhelpful error message, every unknown-error code with no recovery path, is a heuristic violation someone should have caught before launch.
The limitation is that heuristic analysis is a proxy. Expert evaluators will find violations that do not bother real users, and miss problems invisible to the expert because they know the system too well. It is a complement to user research, not a replacement, which is exactly the argument made in what a $2,000 UX audit finds that your analytics can't.
Methods 7 to 9: Surveys, Card Sorting, and Tree Testing
Surveys. Surveys scale where interviews cannot, producing quantitative data about attitudes, satisfaction, and stated behaviours across large samples. Used well they validate patterns observed in qualitative research, measure satisfaction over time, and screen participants for other activities. Used badly they produce a stack of opinions with no actionable content. The failure modes are consistent: leading questions, response options that do not match how users think, questions about hypothetical future behaviour rather than actual past behaviour, and mistaking volume of responses for quality of insight.
Rules for better surveys: ask one thing per question; use closed options only when categories are mutually exclusive and exhaustive; avoid five-point scales for anything other than attitude measurement, because they feel precise and are not; ask about specific past behaviours rather than general tendencies; and pilot with five people before sending at scale, because you will find at least three questions that mean something different to a stranger.
Card sorting. Card sorting reveals how users mentally organise information: which items they group together and what labels they apply. It is the primary input to information architecture. In an open sort participants create and label their own categories, revealing the mental model without constraint; in a closed sort they place items into predefined categories, validating a proposed structure. The most useful output is not the categories themselves but the disagreements: where participants consistently split items your IA puts together, or group things your design separates.
Tree testing. Tree testing validates a proposed navigation structure using only text labels, without visual design or context, by asking participants to find specific items within it. The output is a success rate and path analysis per task, showing which branches confuse users and which labels do not communicate what they contain. Card sorting tells you how users would organise things; tree testing tells you whether they can find things once organised. The two together give a reliable foundation for navigation design.
Methods 10 to 12: A/B Testing, Analytics, and the Competitive Audit
A/B testing. A/B testing runs two or more versions of a design element simultaneously, splitting traffic randomly, and measures which produces the better outcome for a defined metric. Hall's framing matters: A/B testing is knob-twiddling, not strategy. It can tell you whether the green or orange button produces more signups at a measured confidence level. It cannot tell you whether signups are the right goal, whether a fundamentally different approach would work better, or why one version outperformed the other. Booking.com's date-picker improvement produced a 4 percent conversion increase, but that improvement came from UX research that identified confusion in the calendar interaction first. The A/B test validated and measured the fix; research found the problem.
Run it correctly: set a specific, quantifiable goal before creating any variants, and run the test until you reach 95 percent statistical confidence, not before. Cover at least two full weeks to account for day-of-week variation. Do not stop early because one variant is winning, because early stopping produces a high false-positive rate. Keep tests focused on elements where behaviour is expected to vary, such as landing pages, CTAs, and form designs, and avoid testing global navigation, which depends on consistency to function.
Analytics review. Analytics do not tell you why users behave as they do, but they tell you what they do with precision unavailable to any qualitative method. Funnel analysis identifies drop-off between steps; session path analysis reveals unanticipated navigation patterns; cohort analysis tracks whether retention is improving. The most useful analytical question is not what does this number mean, but what question does this number raise that I should investigate further? Analytics tell you where to look; user research tells you what you will find when you look there.
The competitive audit. Competitive research is systematic analysis of what competing products do and how users respond to them. It belongs in the toolkit because understanding what users are already familiar with is essential to designing something that fits their mental models. More useful than a SWOT grid is the competitive usability test: run the same tasks on a competitor's product that you plan to test on your own. You will discover what users already understand about the category, which becomes your baseline, where the competitor's design creates confusion, which is your opportunity, and which conventions are so established that departing from them creates unnecessary friction. The same comparative logic drives the 14 B2B SaaS website patterns analysis.
Key Takeaways
User research methods fall along three dimensions: attitudinal versus behavioural, qualitative versus quantitative, and scripted versus natural context of use.
The most important rule is never to ask people what they want, because users can report current behaviour accurately but cannot predict what would solve their frustrations.
Five participants per user type uncover roughly 85 percent of usability problems, so run tests frequently with small samples rather than occasionally with large ones.
Focus groups are research theatre: five thirty-minute individual interviews teach you more than a two-hour focus group of eight people.
Six biases corrupt research (design, sampling, interviewer, sponsor, social desirability, and the Hawthorne effect) and each has a specific technique that reduces it.








