MaxDiff scores

Introduction

After collecting responses to your MaxDiff exercise, scores are estimated using Hierarchical Bayesian (HB) utility estimation. These scores provide powerful insights by ranking items and revealing their relative importance or preference.

Compared to traditional ranking questions, MaxDiff scores offer greater discrimination and avoid the scale use bias often seen in rating scales.

Interpreting scores

MaxDiff functions like a "beauty contest" among items. As respondents choose their most and least preferred items, Discover calculates preference scores for each item.

Each item receives a positive, ratio-scaled score that sums to 100. This scaling means an item with a score of 10 is twice as preferred as one with a score of 5.

MaxDiff scores summary charts in Discover

Anchored scores

MaxDiff scores - Anchored.png

When anchoring is enabled, a vertical reference line indicates the anchor point in the scores chart. Anchored scores are shown on a positive probability scale, with the anchor indexed at 100:

  • A score of 50 indicates the item is half as preferred as the anchor.
  • A score of 100 indicates the item has an equal likelihood of falling above or below the anchor.
  • A score of 200 indicates the item is twice as preferred as the anchor.

Confidence intervals

You can display 95% confidence intervals alongside MaxDiff scores. Assuming your respondents are representative of the population, you can be 95% confident that the true population score falls within this interval.

Confidence intervals also help determine if one item is preferred over another. If the intervals for two items do not overlap, you can be at least 95% confident that one item is preferred to the other.

Segmenting

Segmentation enables comparisons of MaxDiff results among respondent groups based on survey questions or variables. For example, if your MaxDiff compares music artists, you can segment by respondent location to compare preferences between North America and Europe.

Segmentation dropdown in MaxDiff analysis

When segmentation is active, results update to show scores for each group. Respondents without a defined value for the selected variable are excluded from segmented results.

Segmented MaxDiff results

Missing items behavior

This setting is only available for relevant items MaxDiff. When items are excluded from a respondent's design, you must specify how utility estimation handles those missing items.

Missing at random

Select this option when items are randomly excluded from a respondent's list, or when this assumption otherwise fits your data. Since there's no systematic reason for those items to be missing, they shouldn't be assumed to have better or worse scores than items the respondent saw.

The respondent provides no information about missing items, and the population's mean preference remains unaffected. If using HB estimation, HB will still estimate a utility value for the missing item based on population-level mean preferences and covariances.

Inferior

Choose this option if items are excluded because the respondent previously indicated they were less important or worse than the included items. Informing the estimation model about this distinction prevents bias.

Discover uses data augmentation to guide the estimation process. A threshold item is added to the design matrix — one task for each item in the source list. Each item is compared to the threshold: if the respondent saw the item, it is marked as "best"; if not, the threshold item is marked as "best." This encourages dropped items to have lower utility than included ones without pushing them to extreme negative values.

Examples

  • Express MaxDiff: A random subset of 30 items is chosen from 50 total items for each respondent. The correct assumption is missing at random — only respondents who saw an item should contribute information about its preference.
  • Movies seen: Only movies a respondent has seen are carried into their MaxDiff questions. If the researcher believes respondents can judge unseen movies from reviews and word of mouth, missing inferior may be appropriate. If only direct experience counts, missing at random is correct.
  • Top brands carried forward: Respondents rate a long list of brands and only their top 8 are carried into MaxDiff. The correct assumption is missing inferior, since excluded brands were implicitly ranked lower.

When in doubt, run utility estimation both ways and compare results.

Downloads

The download menu in the upper right corner of the settings panel provides six file downloads.

Scores

The Scores file summarizes MaxDiff utility scores, rescaled to match the table viewable in Discover.

MaxDiff scores summary table

When relevant items MaxDiff is used, an additional tab is included that reports the percentage of times each item was carried forward into respondents' MaxDiff tasks. Reviewing item frequency aids interpretation — for example, if only 2% of respondents saw an item: if missing at random is assumed, those 2% dictate the average preference for the item across the entire sample; if missing inferior is assumed, 98% of respondents are treated as having rejected that item.

Individual MaxDiff scores

Individual scores can be downloaded in two formats: raw and rescaled.

Both include a MaxDiff_Fit (RLH) column. This fit statistic, based on root likelihood, indicates how well a respondent's estimated utility scores predict their actual choices. RLH is the geometric mean of the probabilities of choice produced by the utility scores of what the respondent selected.

A Model fit relative quality column categorizes each respondent as Good, Moderate, or Poor, provided each item was shown at least twice. This rating accounts for task difficulty — predicting one answer from two options is easier than from five. Thresholds are determined by simulating random respondents across various MaxDiff exercise sizes and identifying cutoffs that maximize detection of random responding while minimizing false positives.

When relevant items MaxDiff is used, a Missing items export as blank fields setting appears before download. When off (default), imputed utility scores for missing items are included. When on, blanks replace imputed values for items not shown to a respondent.

Raw scores

Raw scores are regression weights from a multinomial logistic regression model, typically using HB MNL. They are centered around zero, unless anchored scaling is applied, in which case scores are positive for items preferred more than the anchor and negative for those less preferred.

Raw scores are valuable for advanced researchers using the logit equation to predict choice probabilities.

Rescaled scores

Rescaled scores are always positive, sum to 100, and follow a ratio scale. A score of 10 means the item is twice as preferred as an item with a score of 5. These are more intuitive for most audiences.

For details on the rescaling procedure, see Appendix K of the  CBC/HB manual,  specifically the section titled “A Suggested Rescaling Procedure.”

Individual Max Diff Scores Table

Counts

The Counts download provides detailed information on how often each item appeared and how respondents reacted. The data includes:

  • Number of times each item was shown.
  • Number of times each item was selected as “best” or “worst.”
  • Count proportions, expressing the likelihood an item was chosen when shown.
MaxDiff Counts Analysis Table

Count proportions

Count proportions indicate the likelihood of an item being selected when presented:

  • Best count probability: the likelihood respondents selected the item as “best” when it appeared.
  • Worst count probability: the likelihood respondents selected the item as “worst.”

Since these are probabilities, values are between 0 and 1. For example, if four items are shown per task, the chance of selecting any item is 25%. An item with a best count proportion of 50% is selected as “best” at twice the chance level.

To get a quick summary score highly correlated with HB results, subtract an item's worst count proportion from its best count proportion.

Chart

Exports a PNG of the Discover scores chart. This is useful for quickly sharing or presenting results without recreating the visual elsewhere.

Design & choices

Shows the specific design each respondent saw and the concepts they selected. The file is formatted for use in Sawtooth's desktop software tools. If the exercise uses a dynamic list or anchoring, additional tasks representing this information can be included.