The method comprehensibly demonstrates how users relate content to each other and which categories they intuitively form. It is primarily used in early phases when creating navigation, menu structures, and content categories for websites or apps.
The result is groupings from which a new structure can be derived or an existing structure can be verified.
Here's how a card sorting session proceeds
The research team creates the maps even before the meeting.
Each card bears a term, a content block, or a page heading, with exactly one element per card.
The selection of terms determines the expressive power: they should reflect the real content inventory, be phrased in the users' language, and not anticipate internal categories.
During an open session, blank cards and a pen will also be available so participants can create their own group names.
During the session, each participant sorts the cards individually, either physically on a table or digitally in a tool such as OptimalSort, UXtweak, Miro, Maze, or Lyssna.
The instruction for participants is deliberately open-ended: Group the cards in a way that makes the most sense to the participant.
In a facilitated session, the facilitator also asks participants to “think out loud”—that is, to comment on their sorting decisions as they go. Experience shows that it is precisely this kind of listening that constitutes the true value of the method: knowing why someone pairs two cards together reveals more about their mental model than the mere fact that they do so.
For 30 to 50 cards, depending on the variant and whether the session is facilitated, a realistic estimate is about 15 to 30 minutes per participant. Facilitated sessions that include thinking aloud tend to take closer to the upper end of that range, while unfacilitated online sorting sessions are often faster.
Three options: open, closed, or hybrid
The three variants differ in terms of who defines the categories.

In open card sorting, the participant creates the groups themselves and names them as well.
In the closed version, the categories are predetermined by the researcher; the participant simply assigns the cards to them.
The hybrid form combines both: predefined categories are available, but they can be supplemented or changed.
Offen is used in early, exploratory phases when a structure does not yet exist or is being completely reimagined. The method then yields raw categories from the user’s perspective, along with the terms that users actually use.
Closed is typically used as a follow-up step when a category draft is complete and the question is whether users can find the content in the proposed structure.
The hybrid approach is suitable when parts of the structure are fixed—for example, for legal or product-related reasons—but others are still open. The decision-making logic thus follows the project phase: generate with "open," validate with "closed," and combine with "hybrid.".
How many tickets and participants?
The Nielsen Norman Group recommends 30 to 50 cards; the German-language methodology catalog user-experience-methods.com specifies a maximum of 40. IBM researchers observed that, starting at around 40 cards, engagement begins to wane toward the end of a session. With 59 cards, a drop in performance occurred after the first 40.
For open sorts, things start to get tricky at 60 cards; closed sorts with simple mappings can handle even more. For websites, a practical upper limit of 70 to 100 information units in total applies.
Larger content inventories can be broken down into thematic sub-areas and tested separately.
The recommendation regarding the number of participants depends on the study objective. Based on data from Tullis and Wood (2004) there should be at least 15 participants.
With 5 participants, the correlation with the results of a full-scale study is only 0.75; with 15 participants, it is 0.90; and with 30 participants, it is 0.95. The returns begin to decline at 15 participants. The median number of participants in published open card-sorting studies is 34, with 34 cards as well, which empirically supports the recommendations.
By default, each participant conducts the card sorting on their own. This allows us to capture each person’s unbiased, „true“ way of thinking. In a group setting, dominant participants would dictate the structure, which would skew the overall results.
However, there are proven exceptions, depending on the goal we are pursuing:
1. Individual Work (Der Standard)
- How to Each user sorts the cards entirely on their own—either alone at the computer or at a table with a facilitator.
- Advantage: You will receive unbiased, statistically evaluable data about the mental models of your target audience.
2. Pair Sorting / Buddy Sorting (Pair Work)
- How to Two participants solve the task together and must agree on a structure.
- Advantage: Users need to discuss things with each other. That way, as an observer, you'll automatically hear the reasons why a card belongs where it does, without having to keep asking.
3. Workshop Format (Group Work)
- How to A group of 3 to 5 people works together.
- Advantage: This is well-suited for kicking off a project within a project team (e.g., with stakeholders or designers) to develop a shared understanding of the content. However, it is unsuitable for actual user testing.
Qualitative or quantitative?
Qualitative studies seek to answer the question "Why?" They are facilitated, involve thinking aloud, and typically involve about 15 participants, because the insights gained come from the participants' reasoning rather than from statistical patterns.
Quantitative studies, on the other hand, ask how often which cards are grouped together. They run unmoderated using a tool and require at least 30 to 50 participants for the clusters to be statistically reliable.
The drawback of the quantitative approach: Without the comments, the participants' motivations remain unclear, which makes interpretation more difficult.
In practice, this means: If the structure itself is in question, qualitative sorting with a facilitator is the more appropriate choice. If the goal is to validate a grouping that has already been developed using numerical data, a quantitative study with a larger sample size should follow.
From Sorting to Structure: Evaluation
The individual classifications are merged, usually using cluster analysis methods.
The result is usually a dendrogram, which is a tree-like diagram. It shows which cards were grouped together by how many participants, and how closely.
A taxonomy, which is a structured hierarchy of categories, or a folksonomy, which retains spontaneous user-generated tags, can be derived from this image.
The numbers alone carry little weight. One cluster shows that two cards were frequently grouped together, but not whether this is due to a thematic connection, linguistic similarity, or a shared usage context.
This is where the "think-aloud" notes come in: they provide the interpretation that turns a cluster into a justified design decision.
Anyone working quantitatively and forgoing moderation should at least incorporate open comment fields and read the results carefully.
Card Sorting or Tree Testing?
Card sorting and tree testing are often confused, even though they answer opposite questions.
Card sorting is generative: It helps find a structure by having users group content.
Tree testing, on the other hand, is evaluative: it tests an existing structure by having users search for content within an existing hierarchy.
Mantra: Card sorting generates ideas for an information architecture, tree testing evaluates them.
The practical consequence: If the question is „Does user X find Y in our navigation?“, card sorting is the wrong tool; tree testing answers that directly.
Card sorting is correct if the question is „How would users group these 40 content items?“.
Both methods complement each other in a typical sequence: open card sorting early on for category formation, closed card sorting for validation, and tree testing at the end for checking the final structure.
Where Card Sorting Reaches Its Limits
The limitation of card sorting lies in its method: a user-generated taxonomy is not automatically the best navigation structure.
Investigations of Schmettow and Summer as well as recent simulation studies show that deviations between user sortings and the actual navigation structure do not necessarily worsen findability.
Card sorting provides input, but does not replace testing the finished structure, for which tree testing is then responsible.
A second limit lies in the amount of content. With more than 70 to 100 units of information, the method becomes unmanageable for participants, and with approximately 60 cards per session, the sorting quality for open sorts noticeably declines.
The solution is to break it down into thematic sub-studies, not to inflate a single session.
Third: Unmoderated quantitative rankings without explanations are difficult to interpret because the participants' motives are missing. Relying solely on cluster diagrams risks drawing incorrect conclusions.
And finally, card sorting does not solve language problems: If the card labels are ambiguous or come from internal jargon, participants will sort by the label, not by the intended content.
FAQ
How long does card sorting take?
With 30 to 50 cards, 15 to 30 minutes per participant is within the usual range. Moderated sessions with think-aloud protocols tend to be at the higher end, while unmoderated online studies are usually faster.
So, with at least 15 participants, it takes 3.75 to 7.5 hours. With 30 participants, it would take 7.5 to 15 hours.
In addition, there's the time for preparation and evaluation.
What tools are suitable for remote card sorting?
The professional literature regularly mentions OptimalSort from Optimal Workshop, UXtweak, Maze, and Lyssna. For moderated remote sessions with visible sorting, whiteboard tools like Miro are also suitable. An independent evaluation is difficult because the feature sets and pricing models change frequently. The selection should be based on the study objective, in particular on whether the focus is on quantitative evaluation or moderated observation.
Can I do card sorting offline with paper cards too?
Yes, that is the original form of the method. Moderated on-site, it is useful because thinking aloud is directly observable and follow-up questions are possible. The effort lies in the merging: physical sorts must be photographed or typed out before cluster analysis can be performed.