Search Russell Research

Find a solution, case, or article

Enter at least two characters to search the site.

    Press Esc to close. Use Ctrl/⌘ K to search from anywhere.
    Research planning tools

    MaxDiff study planner

    How much should each person evaluate?

    Plan how often your items appear, how many questions to ask, and the sample to budget for each audience. Compare configurations before writing the questionnaire.

    Start with your item list

    For standalone benefits, claims, features, or fixed concepts, also called object-case best–worst scaling. The example starts with four items per question and a target of three appearances per item. All inputs are editable.

    View your plan
    01 What will you use the results for?

    Choose all that apply. Your objectives change the review guidance; they do not apply arbitrary sample multipliers.

    02 Describe the items and questions

    For example: 20 benefits, claims, or product ideas.

    The planner also compares three, four, and five.

    Three is a conventional starting point. Lower targets need sparse-design review.

    Your questionnaire constraint, not a universal maximum.

    Item assignment and absolute appeal

    An assignment is not a guarantee that every assigned item is displayed. Actual questionnaires need a coverage check.

    Anchoring adds judgments such as whether items are appealing at all. It does not establish purchase demand. Indirect anchoring needs review above four items per question.

    03 Who will answer?

    Changing this approach resets the example group entries.

    People remaining after quality review, not survey starts.

    Sample budgeting references

    Practitioner references are separate from the appearance calculation. They do not establish precision or statistical power.

    A sparse-MaxDiff rule of thumb. Repeated appearances are not distinct people. When included, the larger of the practitioner base and appearance reference is used.

    04 Allow for the whole exercise

    Holdouts, excluded from learning exposure. Two is an example.

    Excluded from learning; time is included in the setup allowance.

    Timing assumptions and survey limit

    Illustrative timings for a complete best-and-worst choice. Replace these with pilot observations, including slower completions. Larger sets may take longer; no automatic multiplier is applied.

    Your current configuration

    Questions to reach your appearance target

    15learning questions per person

    18 best–worst tasks including 2 validation and 1 practice.

    15 × 4item appearances20assigned items3per item

    Average exposure, not verified coverage or precision.

    Items that fit your task limit
    20at your target, within the assigned list
    Whole-survey time
    12.2–16.4minutes under your timing assumptions
    • Meets your average-appearance target
    • Within your 15-question limit
    • Within your 20-minute limit at the slower pace

    Candidate to test: 15 questions, 4 items each

    15 learning questions with 4 items each meet the selected arithmetic and presentation checks. Your preferred set size is retained. Confirm the actual design and respondent experience.

    Formula-based guidance. Design quality, ranking precision, and individual accuracy remain unassessed.

    Sample budgeting

    Your usable sample
    300
    Combined budgeting reference
    300
    Overall average appearances per item
    900

    Uses the current 15-question, 4-item configuration. The reference combines your 300 overall base; pooled appearance rules are shown separately.

    The total reaches the budgeting reference. This does not establish statistical adequacy.

    Edit your inputs
    Calculation details and your plan

    Questions: round up [target appearances × assigned items ÷ items per question]. Capacity reverses this calculation and rounds down.

    Pooled exposure: people × learning questions × items per question ÷ full-list items. Practice, validation, and anchoring judgments are excluded.

    Time: instructions and all practice + (learning + validation) × seconds per question ÷ 60 + anchoring time. Add the rest of the survey for the total. Timing ranges are scenarios, not confidence intervals.

    Sample: show 300 overall / 200 per group as editable practitioner references. If selected, combine each with the pooled-appearance reference by taking the larger value. Quota group requirements are added; natural-group requirements use the largest implied total. Natural yields are expectations.

    Candidate: retain the preferred size if it passes the disclosed checks, otherwise consider four, five, then three. More than half the assigned list, more than five items, five detailed items, or indirect anchoring above four require review. An absolute-appeal objective needs anchoring. These are planning safeguards.

    15 learning questions. 3 average appearances per assigned item. Within the question limit.

    Compare configurations

    Fewer items at once, or fewer questions?

    Each configuration aims for your selected exposure. The number of questions changes. Timing assumptions stay constant; test whether larger sets take longer.

    Alternative to compare

    3 items per question

    20 learning questions

    Appearances per assigned item
    3
    Whole survey (min)
    13.8–19.3
    Items within task limit
    15
    • Over the question limit
    • Meets the exposure target
    • Within the slower time allowance

    Current configuration

    4 items per question

    15 learning questions

    Appearances per assigned item
    3
    Whole survey (min)
    12.2–16.4
    Items within task limit
    20
    • Within the question limit
    • Meets the exposure target
    • Within the slower time allowance

    Alternative to compare

    5 items per question

    12 learning questions

    Appearances per assigned item
    3
    Whole survey (min)
    11.2–14.7
    Items within task limit
    25
    • Within the question limit
    • Meets the exposure target
    • Within the slower time allowance

    Sample and item coverage

    Check the people behind each result.

    Appearances include repeated views by the same people. Averages cannot reveal missing items, disconnected comparisons, or the precision of a group difference.

    Current configuration: 15 learning questions × 4 items
    Analysis populationUsable peopleAppearances per itemBudgeting baseGap to base
    Overall3009003000

    Sparse-design appearance references

    At this configuration, 500 appearances per item corresponds to 167 people in the population or group being analyzed; 1,000 corresponds to 334.

    The 500 and 1,000 refer to appearances, including repeat views. Neither reference establishes statistical power.

    How many people see each item in learning tasks?

    If assignment is balanced and every assigned item appears in learning tasks, each item would reach 300 people. This is a conditional scenario, not a design audit.

    Check actual unique reach, repetition, and links across questionnaire versions before fielding.

    What your intended results still need

    • Rank the whole list

      Check ranking stability, including near-ties in the middle of the list.

    From a plan to a study

    Check what the comparison must tell you.

    An overall shortlist, a customer–prospect difference, and a combination of benefits require different evidence. The calculator exposes planning assumptions. A generated design and a tested analysis must establish whether the study can support your decision.

    Discuss your MaxDiff study
    Why use item appearances instead of a fixed question count?

    Fifteen questions showing four items create 60 item appearances. That averages three per item for a 20-item list, but only one for a 60-item list. The same questionnaire length can provide very different coverage.

    Our three-appearance default is a conventional starting point for individual-level estimation. Meeting it does not verify the actual distribution of items or the accuracy of the resulting scores. Study-design guidance.

    When should we consider a different allocation?

    Sparse designs reduce repetition and rely more on pooled information. Express designs repeat a smaller assigned subset, leaving other items unseen by that person. The calculator can compare these exposure assumptions, but it does not create the designs.

    Bandit MaxDiff learns across respondents and allocates more attention to apparent favorites. It may suit a shortlist objective; it should not be treated as equally precise estimation of the whole list. Relevant-items designs use respondent eligibility and need explicit assumptions about missing items. MaxDiff variants.

    Can averages hide missing comparisons?

    Yes. Four questions showing five different items each can cover 20 items once while leaving four disconnected sets. The respondent has made no comparisons linking those sets.

    Check item frequencies, pairings, positions, and direct or indirect links within each version and across the pooled design. Different versions can connect a population-level analysis without creating the missing comparisons within each person. Design checks.

    Does a high rank mean people actually want the item?

    Ordinary MaxDiff orders items relative to one another. Even the highest-ranked item could have little absolute appeal. Anchoring adds a defined threshold, such as whether an item is appealing at all.

    Direct and indirect anchoring collect different evidence and add respondent work. Stated appeal still needs validation against the behavior you want to predict. Anchoring guidance; predictive-validity study.

    What should we test before fielding?

    Review the item list for near-duplicates, compound claims, inconsistent detail, and difficult wording. Pilot complete designs at alternative set sizes and lengths. Simply deleting later questions can remove important coverage.

    Evaluate time distributions, completion, prediction on withheld choices, and stability of the actual shortlist or portfolio decision. For segmentation, check both the recovered segments and individual assignment. Keep final validation separate from the data used to select the design. Sparse-segmentation evidence.

    Can we put MaxDiff scores into a margin-of-error calculator?

    A score rescaled to sum to 100 is not automatically a survey percentage. Uncertainty must come from the MaxDiff model and design. The sample references here do not calculate confidence intervals, power for group comparisons, or portfolio accuracy. Sample-planning guidance.

    Research behind the planner

    Coverage, effort, and the intended analysis.

    These sources inform the planning guidance. Timings and task limits are editable assumptions; the calculator has not evaluated a generated questionnaire or fitted model. Sources reviewed September 29, 2026.

    1. Sawtooth Software. Designing the Study: MaxDiff.

      Basis for exposure planning and the sparse-design appearance references. Actual designs also need frequency, position, and connectivity checks.

    2. Chrzan & Patterson (2006). Testing for the Optimal Number of Attributes in MaxDiff Questions.

      Three commercial studies examine set size, timing, and predictive performance. Supports comparing complete configurations rather than assuming more displayed items provide proportionately more information.

    3. Sawtooth Software (2024). MaxDiff Sample Size Calculation Best Practices.

      Source for the editable 300-person overall and 200-person per-group budgeting references. Formal precision and power require the intended analysis and design.

    4. Chrzan (2025). Segmenting With Sparse MaxDiff Data.

      A simulation with 36 items and four groups distinguished recovering the number of segments from assigning individuals accurately. Results from this one setting are not universal thresholds.

    5. Orme (2019). Making Sense of All Those MaxDiffs!

      Explains how sparse, Express, Bandit, relevant-items, and anchored approaches change allocation or interpretation. They are not interchangeable shortcuts.

    6. Sawtooth Software. Anchored MaxDiff.

      Distinguishes direct judgments from indirect follow-ups. Supports reviewing indirect anchoring when more than four items appear together.

    7. Incentive alignment in anchored MaxDiff yields superior predictive validity. Marketing Letters.

      A preregistered experiment with 448 participants examined consequential choices. Measuring a stated threshold does not by itself validate a demand forecast.