How do questionnaires measure? More response options may not improve psychological measurement

3 Jul 2026 News Fresh studies For practitioners

No description

The public is generally interested in groundbreaking findings, and studies often examine phenomena that significantly impact society, human health, the mysteries of the universe, and the like. However, there are research topics that are quite important but lie completely outside the scope of interest for both the media and the general public. One such topic is research into how to accurately measure human characteristics. This is absolutely crucial—without high-quality measurement, it is practically impossible to conduct any further research. And one of INPSY’s research teams has long been focused precisely on measurement research in psychology.

Much of psychological research today is based on “measurement” using questionnaires, in which respondents assess a particular characteristic of their own using a large number of items. The most common format for these is the so-called Likert scale. It consists of a statement, such as “I am a good person,” and a scored scale with several options, such as “I disagree” (1), “somewhat disagree” (2), “somewhat agree” (3), “agree” (4). Finally, the responses are usually tallied, although more complex statistical methods can also be applied.

From the very beginning, however, researchers have been debating how many response options an ideal Likert scale should have. Four? Six? Ten? Or perhaps just two (yes/no)? Equally lively is the debate over the inclusion of a midpoint (“somewhat”) or exactly how the individual response options should be described.

And it is precisely the number of points that researchers at INPSY have focused on. Back in 2020, Petra Hubatka (now a doctoral student and researcher at the Faculty of Social Studies) defended her bachelor’s thesis under the supervision of Hynek Cígler; she and David Elek have now published a revised version of the text in the prestigious journal Assessment.

How Do We Evaluate the Quality of Measurement? Accuracy and Validity

Measurement in the social sciences is most often described using two main characteristics. Reliability tells us how accurately a questionnaire measures—regardless of what we are actually measuring. It simply refers to the stability of the measured score, or rather, the extent to which chance influences the measurement. Validity, on the other hand, describes whether the questionnaire actually measures what the researchers intended. There are many approaches to assessing validity. In this case, the researchers examined what is known as criterion validity—that is, whether the questionnaire results correspond to something measurable in the real world.

Reliability and criterion validity are often closely linked. If there are too many random errors in the measurement, the true result may be lost among them. In other words, as a general rule, the higher the reliability, the higher the criterion validity. We also know from previous research that the greater the number of response options on a Likert scale, the smaller the measurement error. Reliability generally increases as the number of response options rises from about two to six, after which it remains at the same level. This seems logical—a more detailed opportunity to express one’s characteristics should increase the scale’s discriminatory power and, consequently, its reliability.

Surprising Results

In this recent study, INPSY researchers fully replicated the findings regarding the relationship between the number of response options and reliability. However, they were the first to show that, paradoxically, increasing reliability does not necessarily lead to higher criterion validity. While reliability increased as the number of response options increased, criterion validity remained constant.

The Unique Height Questionnaire

The researchers reached their conclusions by using an innovative Height Questionnaire, which was developed specifically for such purposes at INPSY. Instead of asking directly how tall a person is, it asks questions such as “Am I tall enough to play basketball or volleyball?” or “Do I often have to stand on my tiptoes to see better?” As humorous as this may seem, this questionnaire actually measures a person’s height very accurately—in fact, you can try it out for yourself in the online demo version. This provided the researchers with an independent, objective criterion for the measurement tool (a person’s actual height), which is typically lacking in psychology.

Similar research findings may influence the nature of measurement in psychology. Since answering on a yes/no scale is faster than answering on a longer response scale, it may be more advantageous to improve measurement quality by adding more questions to the questionnaire rather than by increasing the number of response options. This, in turn, offers new, more entertaining response formats. In another, as-yet-unpublished study, our members Hynek Cígler and Adam Strojil presented a questionnaire in which respondents answer by “swiping,” similar to the dating app Tinder.


Recommended citation:

Hubatka, P., Cígler, H., & Elek, D. (2026). Spurious Reliability Increase?: The Number of Response Options in the Likert-Type Scale Influences Only Internal Consistency, Not Criterion Validity. Assessment. https://doi.org/10.1177/10731911261452890

Interested in the study? Contact its author!

Mgr. Petra Hubatka
Team Measurement in Psychology
petra.hubatka@mail.muni.cz

Read the study

Many of our publications follow the principles of Open Science. We want to ensure that our studies are reproducible by other teams and are free to read.

Our goal is to

Open Science at INPSY


More news

All articles

You are running an old browser version. We recommend updating your browser to its latest version.

More info