How to Code Qualitative Design Research
Qualitative coding becomes flattening when a short label replaces the conditions that made an observation meaningful. A comment about “missing” a control may describe a navigation problem, an unfamiliar task, a device constraint, a preceding action, or an explanation the participant is still uncertain about. The words alone cannot decide among those possibilities.
A defensible approach to coding qualitative design research treats codes as provisional handles for retrieval and comparison. It stores context with each excerpt, separates description from interpretation, delays broad categorisation, and carries uncertainty into any design implication. The goal is not to eliminate judgment. It is to make judgment visible, revisable, and appropriately bounded.
Start with a context-preserving record
Before assigning a code, define what counts as one coding unit. It might be a complete answer, a turn in an interview, a sequence of observed actions, a moment of hesitation, or a short passage containing one coherent idea. The unit should be large enough to retain meaning but focused enough to compare with other material.
A useful record keeps the excerpt attached to the conditions around it:
- Participant: identifier and relevant description, without unnecessary personal information
- Source and location: interview, observation, usability session, transcript location, or clip reference
- Task or question: what the participant was doing or responding to
- Setting: device, environment, product state, or other conditions that may matter
- Sequence: what happened immediately before and after
- Observed material: the participant’s words or the researcher’s observable description
- Provisional descriptive code: a short label for what is present in the material
- Interpretive memo: what the researcher thinks may explain or connect it
- Uncertainty: what remains unknown, ambiguous, or dependent on context
- Design relevance: a possible implication, kept separate from the observation
This is more information than a code label alone, but it prevents later synthesis from depending on memory. It also makes the boundary between evidence and analysis inspectable. Qualitative coding is commonly described as assigning descriptive labels to aspects of data so the material can be organised for analysis. A context record extends that organising function; it does not turn a code into proof of a cause.
Apply descriptive codes before interpretive ones
A descriptive code stays close to what the material shows. Examples include “pauses before submitting,” “asks where to find saved work,” “describes the warning as unexpected,” or “repeats the same search.” These labels do not claim why the behaviour occurred.
An interpretive code makes a stronger move: “unclear system status,” “low confidence in recovery,” or “navigation hierarchy problem.” Such labels can be useful, but they are explanations rather than direct observations. Keep them in a separate memo field, and write the reasoning that connects the observation to the interpretation.
This separation matters because interpretations can become invisible once they enter a codebook. A researcher may later remember “navigation problem” as if the participant directly demonstrated that cause, even though the original material showed only a pause and a request for help. Keeping both labels preserves the point where analysis begins.
A code definition should state what belongs under the label, what does not, and what additional context must be retained. For example:
Provisional code: “searches for a previously used control”
- Include: a participant looks through the interface for a control they have used before.
- Exclude: a participant asks what an unfamiliar control does without searching for it.
- Retain: task, preceding action, interface state, device, and whether the control was eventually found.
- Do not establish: that the control is poorly placed or that the participant lacks product knowledge.
The definition gives the code a usable boundary without pretending that every instance has the same explanation.
Delay broad categorisation
Do not merge every similar phrase into a theme as soon as it appears. First code the material at a level that preserves differences. A commonly described thematic-analysis sequence moves from familiarisation through coding and then toward themes, as outlined by Nielsen Norman Group.
Delayed categorisation is a practical way to resist premature closure. One practitioner account recommends postponing coding or categorisation rather than forcing early labels as an alternative to immediate coding. Treat that advice as a tactic, not as proof that delay produces better analysis. Waiting longer can preserve uncertainty, but it can also make a project harder to manage without a working structure.
A workable compromise is to use narrow provisional codes early, then postpone the larger category. “Cannot locate saved work,” “checks another menu,” and “asks whether work was saved” may later relate to system status, navigation, or task expectations. Keeping them distinct allows the context to influence the eventual grouping.
When a category begins to form, ask what would make two excerpts meaningfully different:
- Do they occur during different tasks?
- Does the setting change the likely explanation?
- Are participants describing different outcomes despite using similar words?
- Is one observation direct while another is retrospective interpretation?
- Does the proposed category combine a behaviour with its suspected cause?
If the answer is yes, split the code or preserve a context qualifier. A broad label may be convenient for counting but weak for explanation.
Compare cases without treating them as interchangeable
Cross-case comparison is useful when it reveals a pattern and its boundaries. It becomes misleading when it removes the conditions that distinguish one case from another.
A context matrix can support comparison without replacing the original records. For each coded case, compare the same dimensions—participant position, task, setting, sequence, observed action, stated explanation, and outcome—while leaving room for differences. The purpose is not to produce a uniform dataset. It is to see whether a possible pattern persists, changes, or disappears under different conditions.
Frequency is only one signal. A code that appears often may describe a routine detail with little bearing on the decision. A rare observation may reveal a consequential failure condition, an accessibility concern, or a limitation on a proposed pattern. Treat frequency as a property of the collected material, not as a direct measure of importance or prevalence beyond that material.
A staged coding-and-theming process has been described as a way to structure analysis and support collaboration, but the available description does not establish that it improves validity or design outcomes in the reported qualitative analysis process. The practical test is whether the workflow makes disagreements easier to locate: which excerpt, context, code definition, interpretation, or comparison produced the difference?
Review contradictory and exceptional cases
A theme becomes more defensible when it explains both supporting and conflicting material. Search deliberately for cases that do not fit the emerging category, including observations from different tasks, settings, participant groups, or points in a sequence.
When a contradiction appears, check whether the cases differ in:
- the task participants were attempting;
- the product state or available information;
- prior experience or role;
- the timing of the observation;
- the researcher’s question or prompting; or
- the interpretation attached to otherwise similar behaviour.
The result may be a split theme, a narrower claim, or an explicit exception. “Participants struggle with navigation” might become “Some first-time participants searched for saved work after an interrupted task, while returning participants used a known path.” The second statement is less dramatic, but it preserves the conditions that determine where a design response might apply.
A negative case can also expose an overbroad code definition. If one excerpt appears to contradict a code, revisit what the code is meant to capture. The problem may be the case, or it may be that the code combines separate mechanisms such as discoverability, system feedback, and task memory.
Carry bounded implications into design
The final step is not to turn a theme directly into a requirement. Translate the coded material through a chain that remains traceable:
- Observation: what was said, done, or visibly encountered?
- Interpretation: what explanation is plausible, and what supports it?
- Scope: under which participants, tasks, settings, or states might it apply?
- Design implication: what change could address the condition?
- Open question: what would still need checking before the change becomes a broader rule?
Hypothetical example: several people “miss” a control
Suppose a researcher codes several comments about users “missing” a control. The initial record does not merge them immediately into one navigation problem. It retains the task, device or setting when known, participant description, preceding action, observed behaviour, participant explanation, researcher interpretation, and unresolved uncertainty.
One case might involve a participant searching after returning from another screen. Another might involve a participant overlooking the control while completing an unfamiliar task. A third might involve a visible control that resembles surrounding content. The shared phrase is not enough to establish a shared cause.
A bounded implication might be: “Review the control’s visibility and recovery path for the interrupted-task state observed in these sessions.” That differs from “Redesign navigation because users cannot find the control.” The first preserves scope and identifies a design check. The second turns an interpretation into a general diagnosis.
When the implication affects a reusable pattern or system rule, retain the evidence path. A related workflow for converting research notes into design decisions likewise separates observation, synthesis, interpretation, and system implication before a shared rule is introduced. Coding should give that later decision process better material, not bypass it.
Decide when the coding is sufficient
There is no universal number of codes, passes, or participants that makes a qualitative analysis complete. Set a practical stopping condition instead. You might stop a coding pass when new excerpts can be assigned to existing definitions without forcing important differences away, major contradictions have been reviewed, and each proposed implication can be traced back to specific records.
That stopping condition is not statistical saturation, and it does not prove that no further insight exists. It is a decision about whether the current analysis is sufficiently coherent for the question and design choice at hand. If a proposed implication still depends on an unresolved cause, narrow the implication, mark the uncertainty, or gather targeted material rather than strengthening the wording.
Context-preserving coding takes more time and produces a less tidy summary. The return is not a guarantee of better research. It is a clearer account of what the material shows, what the researcher inferred, where cases differ, and how far a design decision can responsibly travel.