Numbers feel comfortable. We trust a thermometer to tell us the temperature and a weighing machine to tell us how heavy a sack of rice is. But what happens when a researcher needs to measure something that has no physical weight or length, like how satisfied a household is with a new water supply scheme, or how strongly people support a sanitation drive? You cannot place an opinion on a weighing machine. This is exactly the problem that scales were designed to solve. They give social researchers a structured way to turn feelings, attitudes, and perceptions into data that can be counted, compared, and analysed.
Table of Contents
- What scales mean in social research
- Why development work depends on good measurement
- Comparative scaling: judging items against each other
- Paired comparison
- Rank order and constant sum
- Non-comparative scaling: judging each item on its own
- The Likert scale
- Semantic differential and continuous rating
- Putting scales to work in programme evaluation
- Building a scale you can trust
What scales mean in social research
A scale is a measurement tool that assigns numbers or ordered categories to qualities that cannot be measured directly. Where a ruler measures distance, a scale measures the intensity or direction of an attitude, opinion, perception, or behaviour. It works as a continuum that runs from one extreme to another, with intermediate points in between, so that a response sitting higher on the scale clearly represents a stronger or more favourable position than one sitting lower.
The reason scales are necessary is that social behaviour is far more complex than physical quantities. As one academic resource on research methodology points out, measuring attitudes, perceptions, preferences, and opinions is difficult precisely because of this complexity, unlike the straightforward standardised instruments we use for height or volume. Scales bridge that gap. They impose a consistent rule so that two researchers measuring the same opinion arrive at comparable results, which is the foundation of reliable evidence.
Why development work depends on good measurement
In monitoring and evaluation, the questions that matter most are rarely about simple facts. Programme managers do not just want to know whether a microfinance scheme reached a village; they want to know whether participants feel their lives improved, by how much, and in which areas. They want to measure quality, satisfaction, dignity, empowerment, and trust. None of these can be counted like the number of bank accounts opened.
Scales make this kind of nuanced measurement possible. A well-built scale can capture not just whether an outcome occurred, but the degree to which it occurred. That sense of degree is what allows evaluators to compare one block against another, track change over time, and judge whether public money is producing real social value. The four basic levels of measurement that underpin every scale are nominal (simple categories), ordinal (ranked order), interval (equal gaps between points), and ratio (a true zero point). The level a scale operates at decides which statistical tests an evaluator can later apply, so the choice is made carefully at the design stage.
Comparative scaling: judging items against each other
Scaling techniques fall into two broad families. The first is comparative scaling, where respondents are asked to directly compare items against one another rather than rate each one on its own. The result tells you which option is preferred, but not by how much in absolute terms. Because of this, comparative scales usually produce ordinal data, capturing order or ranking rather than fixed values, as the Food and Agriculture Organization notes in its overview of comparative and non-comparative scales.
A simple way to picture this is a wage preference question. If a survey asks an agricultural labourer whether they would rather be paid a fixed weekly wage or a daily wage, the answer reveals a preference between two options placed side by side. The respondent is forced to choose one over the other, and that forced choice is the essence of comparative scaling. You learn the ranking of the two, but not how intensely the worker feels about either.
Paired comparison
The most common comparative method is the paired comparison scale. Here the respondent is shown two items at a time and asked to pick one according to a stated criterion. In a paired comparison, respondents are simply given the two extreme choices and asked to express a preference, with no rating scale within each item itself. If an evaluator wanted to know which of four welfare benefits villagers value most, they could present every possible pair and tally how often each benefit was chosen. The benefit picked most often ranks highest.
Rank order and constant sum
Two close relatives extend this idea. In a rank order scale, respondents see several items at once and arrange them from most to least preferred, which is quicker than comparing every pair separately. In a constant sum scale, respondents are given a fixed pool of points, say 100, and asked to distribute them across the items according to importance. If a respondent gives 60 points to clean drinking water and 20 each to electricity and roads, the allocation shows not only the ranking but a rough sense of relative weight. These techniques are valuable when a programme needs to understand community priorities before deciding where to spend limited funds.
Non-comparative scaling: judging each item on its own
The second family is non-comparative scaling, where each item is evaluated independently rather than against the others. Instead of asking which of two options is better, the researcher asks the respondent to place a single item somewhere on a defined continuum. This produces an absolute judgement and often yields interval-level data that is suitable for richer statistical analysis. Most attitude measurement in social research uses non-comparative scales because they let each respondent express how strongly they feel, not merely what they prefer.
The Likert scale
The Likert scale is the most widely used non-comparative technique in survey research. It presents a statement and asks the respondent to indicate their level of agreement, typically across five points running from “strongly disagree” to “strongly agree.” A statement such as “The local health centre treats patients with respect” can be answered on this five-point range, and the chosen point becomes a number that can be averaged across hundreds of respondents. Its popularity comes from being easy to understand for respondents and easy to score for analysts, which makes it a natural fit for large field surveys.
Semantic differential and continuous rating
Two other non-comparative tools are common in development research. The semantic differential scale places a pair of opposite adjectives at either end of a line, such as “unfriendly-friendly” or “inefficient-efficient,” and asks respondents to mark where their perception falls. It is well suited to capturing the overall feel of a service or institution. The continuous rating scale goes a step further by letting respondents place a mark anywhere along an unbroken line rather than at fixed points, which suits concepts that genuinely exist on a smooth continuum, like a household’s overall satisfaction with a programme.
Putting scales to work in programme evaluation
Scales come into their own when evaluators have to measure abstract but crucial concepts like quality and satisfaction. These ideas are multi-dimensional, meaning they are built from several underlying components rather than a single feeling, so a good scale breaks them into measurable parts.
Health programmes offer clear examples from the Indian context. A study in Uttar Pradesh developed a sixteen-item scale to measure patients’ perceptions of quality, which identified five distinct dimensions, including medicine availability, staff behaviour, doctor behaviour, and hospital infrastructure. By scoring each dimension separately, evaluators could pinpoint exactly where a public facility was falling short rather than receiving a single vague verdict of “poor service.” Similarly, a Hindi-translated scale tested among postnatal women in Chhattisgarh treated maternal satisfaction as a multi-dimensional phenomenon and proved reliable for use in Indian health facilities.
The same logic applies far beyond health. Researchers studying public service delivery have built multi-dimensional scales to gauge citizens’ satisfaction with government services across factors like reliability, responsiveness, and communication. When an evaluator can express something as slippery as “satisfaction” through a set of clear, scored dimensions, the programme finally has actionable feedback. It learns not just that beneficiaries are unhappy, but which specific aspect to fix.
Building a scale you can trust
A scale is only as good as its construction. Before relying on the numbers, researchers test a scale for reliability, meaning it produces consistent results when repeated, and validity, meaning it actually measures the concept it claims to measure. The statements that make up the scale must be clear, must point in a single direction, and must be able to distinguish between people who hold different views. A poorly worded item can quietly corrupt an entire dataset, which is why scale development is treated as a careful, tested process rather than a quick set of questions thrown together before fieldwork.
For anyone working in monitoring and evaluation, understanding the difference between comparative and non-comparative approaches is a practical skill, not an academic luxury. Choosing the right technique decides whether the final report rests on solid evidence or on numbers that look precise but mean very little.
What do you think? If you were evaluating a rural sanitation programme, would you reach first for a comparative scale to rank community priorities, or a non-comparative scale to measure how satisfied each household feels, and why? How might the choice change the story your data ends up telling?
References
- https://ebooks.inflibnet.ac.in/socp3/chapter/scaling-and-measurement/
- https://www.fao.org/4/w3241e/w3241e04.htm
- https://www.sciencedirect.com/topics/social-sciences/scaling-technique
- https://pubmed.ncbi.nlm.nih.gov/17012306/
- https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6352900/
- https://www.abacademies.org/articles/design-and-development-of-scale-of-measurement-for-effective-service-delivery-in-india-15693.html
Leave a Reply