Questionnaire Validity and Reliability: Types, How to Measure Them, and a Sample Write-Up
Validity says the questionnaire measures what it claims to; reliability says it gives the same number if you measure again. The types of each, the minimum expert votes for Lawshe's CVR, computing the CVI, Cronbach's alpha with a worked example, test-retest and split-half, and ready-to-adapt text for your methods chapter.
When a thesis examiner reaches the methods chapter, the first two words they look for are validity and reliability. If that section is two vague lines ("the questionnaire's validity was confirmed by professors and Cronbach's alpha was 0.8"), it will very likely come back for revision. This article shows exactly what each type of validity and reliability establishes, how it is measured and with how many people, and ends with a sample methods text you only need to fill with your own numbers.
If you are still writing the items, read the questionnaire design guide first; here we assume you have the questionnaire and want to demonstrate that it is sound.
Validity and reliability at a glance
Picture a scale that always shows two kilos too many. Every time you step on it you see the same number, so it is reliable; but it doesn't show your true weight, so it is not valid. The reverse is impossible: an instrument that gives a different number every time cannot be valid. That is why reliability is a necessary condition for validity, not a sufficient one.
| Type | Question it answers | Common method | Usual criterion |
|---|---|---|---|
| Face validity | Do the items look clear and relevant to respondents? | Feedback from 10–15 people from the population; item impact score | Impact score of 1.5 or more |
| Content validity | Are the items essential, and do they cover every aspect of the construct? | Expert judgement: CVR and CVI | CVR above Lawshe's critical value; I-CVI above 0.79 |
| Construct validity | Do the data reproduce the theoretical structure? | Exploratory or confirmatory factor analysis | Loadings of 0.4 or more; fit indices |
| Criterion validity | Does the score agree with a valid external criterion? | Correlation with another valid measure or a real outcome | Significant correlation in the predicted direction |
| Internal consistency | Do the items of one dimension move together? | Cronbach's alpha (or omega) | 0.7 or more |
| Stability (test-retest) | Same result if we measure again? | Second administration after 2–4 weeks; ICC | 0.7 or more |
| Split-half | Are the two halves of the instrument equivalent? | Correlation of the halves with the Spearman–Brown correction | 0.7 or more |
Face validity: ask the target population before any expert
Face validity is the simplest and cheapest step, and precisely for that reason it gets skipped. Give the questionnaire to 10–15 people from your target population (not classmates, not your supervisor) and ask which items they had to read twice, which words they didn't understand and which items seemed irrelevant. A nurse understands "self-efficacy" differently from a psychology student.
If an examiner wants a quantitative method, the "item impact score" is common: each person rates the importance of each item from 1 to 5, and for each item the frequency (the share who gave 4 or 5) is multiplied by the importance (the mean rating):
Impact score = frequency × importance = 0.8 × 4.1 = 3.28 (≥ 1.5 → keep)
Content validity: CVR and CVI, two indices for two different questions
Most theses report both, and many mix them up. CVR asks "is this item essential?" and CVI asks "is this item relevant, simple and clear?" So you need two questions for your experts: two separate forms or two separate columns in one form. You can build the expert form online too: each item is a row in a matrix question with three columns, "essential," "useful but not essential" and "not necessary"; the matrix report gives you the count in each column for every item.
CVR: is each item essential?
Each expert picks one of the three options for each item, and Lawshe's formula (1975) is:
CVR = (nₑ − N/2) / (N/2)
nₑ is the number of experts who said "essential" and N the total number of experts. CVR runs from −1 to +1; zero means exactly half the experts rated the item essential. An item whose CVR falls below Lawshe's critical value is dropped or rewritten.
Everyone quotes Lawshe's table, but what you actually need in practice is the third column: with this many experts, how many must say "essential"?
| Number of experts | Lawshe's minimum CVR | Minimum "essential" votes | CVR at that vote count |
|---|---|---|---|
| 5 | 0.99 | 5 of 5 | 1 |
| 6 | 0.99 | 6 of 6 | 1 |
| 7 | 0.99 | 7 of 7 | 1 |
| 8 | 0.75 | 7 of 8 | 0.75 |
| 9 | 0.78 | 8 of 9 | 0.78 |
| 10 | 0.62 | 9 of 10 | 0.80 |
| 11 | 0.59 | 9 of 11 | 0.64 |
| 12 | 0.56 | 10 of 12 | 0.67 |
| 13 | 0.54 | 10 of 13 | 0.54 |
| 14 | 0.51 | 11 of 14 | 0.57 |
| 15 | 0.49 | 12 of 15 | 0.60 |
| 20 | 0.42 | 15 of 20 | 0.50 |
| 25 | 0.37 | 18 of 25 | 0.44 |
| 30 | 0.33 | 20 of 30 | 0.33 |
| 40 | 0.29 | 26 of 40 | 0.30 |
Two things follow from this table. First, with 10 experts the famous 0.62 effectively means "at least 9": 8 votes give a CVR of 0.6, which is below 0.62, so an item with 8 "essential" votes out of 10 is rejected. Second, the "minimum votes" column matches the exact binomial calculation that Ayre and Scally (2014) performed when they revisited Lawshe's table; so if the decimals of the table are ever questioned, the vote count is the safer criterion.
CVI: is each item relevant, simple and clear?
The item-level CVI is usually computed by asking each expert to rate the "relevance" of each item on a four-point scale (1 = not relevant, 2 = somewhat relevant, 3 = quite relevant, 4 = highly relevant), often repeated for "simplicity" and "clarity" (the approach associated with Waltz and Bausell). For each item:
I-CVI = number of experts rating 3 or 4 ÷ total number of experts
A common decision rule: items with an I-CVI above 0.79 are kept, between 0.70 and 0.79 revised, below 0.70 dropped. The scale-level index (S-CVI/Ave) is the mean of the I-CVIs, and Polit and Beck (2006) recommend 0.90 or more.
Example: 10 experts rated four items in one dimension:
| Item | Ratings of 3 or 4 | I-CVI | Decision |
|---|---|---|---|
| A1 | 10 of 10 | 1.00 | Keep |
| A2 | 9 of 10 | 0.90 | Keep |
| A3 | 8 of 10 | 0.80 | Keep |
| A4 | 6 of 10 | 0.60 | Drop, or rewrite and ask the experts again |
With all four items, S-CVI/Ave is 0.825; after dropping A4 it becomes (1 + 0.9 + 0.8) ÷ 3 = 0.90 and reaches the recommended threshold. Note that "CVI" is used in two senses in the literature: Lawshe also called the mean CVR of the retained items a CVI. Say which one you mean in your report.
How many experts do you need?
For the CVI, Lynn (1986) considered at least 3 experts necessary and more than 10 usually unnecessary; with 5 experts or fewer, every item's I-CVI should be 1. For CVR, because the critical values are very strict with few experts (5 to 7 experts means unanimity), 10 to 15 experts is a more practical choice. An "expert" knows the construct (a faculty member in the field) and the field itself (say, an experienced manager in that industry); a mix of both beats ten professors from one department.
Construct validity: when the data testify
Face and content validity rest on people's judgement; construct validity comes from the data. The key question is whether the items you placed in one dimension actually move together in real responses and separate from the items of other dimensions.
- Exploratory factor analysis for researcher-made instruments, or instruments whose structure is unclear in your population. Prerequisites: KMO above 0.7 (0.6 at the very least) and a significant Bartlett's test of sphericity. An item loading below 0.4, or cross-loading on two factors with similar loadings, is a candidate for removal.
- Confirmatory factor analysis for standard instruments whose structure is already known. Common fit criteria: RMSEA and SRMR below 0.08, CFI and TLI above 0.90; the stricter criteria of Hu and Bentler (1999) are CFI of at least 0.95 and RMSEA of at most 0.06.
- Convergent and discriminant validity, mostly in structural equation modelling: average variance extracted (AVE) above 0.5 and composite reliability (CR) above 0.7 for convergence; for discrimination, the square root of each construct's AVE exceeding its correlations with other constructs (the Fornell–Larcker criterion, 1981), or HTMT below 0.85.
Factor analysis needs a large sample: the usual rule of thumb is 5 to 10 responses per item and at least 200 people. A 30-person pilot is not enough; report construct validity on your main data. How to size the sample properly is covered in the guide to Cochran, Morgan and power analysis.
Criterion validity
If another valid measure of the same construct exists, the correlation of your questionnaire's score with it (concurrent validity) or with a later outcome, such as next term's grades or staff turnover (predictive validity), is strong evidence. It is usually not required for a master's thesis, but if you have it, report it.
Reliability: three ways to ask "the same number?"
Cronbach's alpha with a complete example
Alpha measures the internal consistency of the items in one dimension and is computed for each dimension separately, not once for the whole questionnaire:
α = (k / (k − 1)) × (1 − Σsᵢ² / sₜ²)
k is the number of items, sᵢ² the variance of each item and sₜ² the variance of the dimension's total score. Suppose the dimension "satisfaction with online classes" has four items and the fourth is negatively worded ("I usually get bored in online classes"). Eight people's answers (hypothetical data, only to show the calculation) after reversing item 4:
| Respondent | I1 | I2 | I3 | I4 (reversed) | Total |
|---|---|---|---|---|---|
| 1 | 4 | 5 | 3 | 4 | 16 |
| 2 | 3 | 3 | 4 | 3 | 13 |
| 3 | 5 | 4 | 4 | 4 | 17 |
| 4 | 2 | 3 | 3 | 2 | 10 |
| 5 | 4 | 3 | 5 | 3 | 15 |
| 6 | 3 | 2 | 2 | 2 | 9 |
| 7 | 5 | 4 | 4 | 5 | 18 |
| 8 | 2 | 3 | 2 | 3 | 10 |
| Variance | 1.43 | 0.84 | 1.13 | 1.07 | 12.29 |
The item variances sum to 4.46, so α = (4 ÷ 3) × (1 − 4.46 ÷ 12.29) = 1.33 × 0.64 ≈ 0.85. Now compute the same data without reversing item 4: the variance of the total drops to 3.43 and alpha turns negative (−0.40). A negative or very low alpha in a pilot usually means a forgotten reverse-scored item, not a broken questionnaire.
You don't have to do this by hand: put the item data into the Cronbach's alpha calculator to see alpha and "alpha if item deleted." In SPSS the same output is under Analyze → Scale → Reliability Analysis; the step-by-step is in analyzing a Likert questionnaire in Excel and SPSS.
Three points of interpretation:
- Alpha is sensitive to the number of items. A two-item dimension struggles to reach 0.7, and a twenty-item dimension gets a high alpha even with loosely related items. If a dimension has two or three items, also report the mean inter-item correlation (roughly 0.2 to 0.5 is desirable).
- Alpha above 0.95 is not a compliment. It usually means redundant items reworded; the questionnaire can be shortened.
- A corrected item-total correlation below 0.3 flags a suspect item; but don't drop an item that covers part of the construct's content just to gain a few hundredths of alpha.
In recent years McDonald's omega is often reported alongside alpha, because it doesn't require alpha's strict assumption that all items carry equal weight. Recent SPSS versions and the free JASP compute it.
Test-retest: the same people, a few weeks later
Give the same questionnaire again to 20–30 of the same people after 2 to 4 weeks. The interval should be long enough that they don't remember their previous answers and short enough that the construct itself hasn't changed. It suits stable constructs (personality traits, attitudes); not this week's mood or pre-exam anxiety.
To compare the two rounds, the intraclass correlation coefficient (ICC) is better than Pearson's correlation: if everyone scores exactly one point higher the second time, Pearson is still 1, but an absolute-agreement ICC sees the shift. The common guidance of Koo and Li (2016): below 0.5 poor, 0.5–0.75 moderate, 0.75–0.9 good and above 0.9 excellent.
A practical tip: to pair each person's first and second answers, ask for a "personal code" only the respondent can rebuild (for example the last four digits of their phone plus their birth month). Anonymity is preserved and pairing the rounds is still possible.
Split-half and the Spearman–Brown formula
If a second administration isn't possible, split the items into two halves (usually odd and even), score each half and correlate them. Because each half is only half as long as the instrument, this correlation understates the reliability of the whole and is corrected with the Spearman–Brown formula:
r_SB = 2r / (1 + r) r = 0.70 → 1.40 / 1.70 ≈ 0.82
The same formula has another use: it predicts how much reliability would rise if you lengthened a dimension. A six-item dimension with an alpha of 0.65 would reach about 0.76 if extended to ten items of similar quality. For tests with right/wrong answers (0 and 1), the equivalent of alpha is KR-20.
How many people does each step need?
| Step | Who | Typical number |
|---|---|---|
| Face validity | People from the target population | 10–15 |
| CVR | Experts in the subject and the field | 10–15 |
| CVI | Experts | 3–10 |
| Pilot and alpha | A sample like the population, outside the main sample | About 30 |
| Test-retest | The same people, twice | 20–30 |
| Factor analysis | Main sample | 5–10 per item; at least 200 |
Does a standard questionnaire need validity and reliability too?
Yes, but less. If you use a validated translation of a standard instrument, you report validity by citing the validation paper and recompute reliability in your own sample; alpha is a property of the data, not of the paper. If you translated the instrument yourself or changed an item, it no longer counts as "standard" and needs at least content validity and reliability afresh. How to tell a trustworthy version is covered in what a standard questionnaire is and where to find one.
Sample text for your methods chapter
Fill this in with your own numbers and sources; the figures in it are only examples:
Validity. To assess face validity, the questionnaire was given to 12 members of [the target population], and three ambiguous items were rewritten. Content validity was assessed by 10 [experts in the field]. The content validity ratio (CVR) of every item except item [7] was 0.80 or higher, above the minimum of 0.62 in Lawshe's table for 10 experts; item [7] was removed. Item-level content validity indices (I-CVI) ranged from 0.80 to 1, and the S-CVI/Ave was 0.92.
Reliability. The questionnaire was piloted with 30 members of [the population] who were not in the main sample. Cronbach's alpha for the dimensions [A], [B] and [C] was 0.84, 0.79 and 0.81 respectively in the pilot, and 0.86, 0.80 and 0.83 in the main sample (n = [292]), all above the conventional minimum of 0.7.
And the table examiners like:
| Dimension | Items | Reversed items | Alpha (pilot, n = 30) | Alpha (main sample) |
|---|---|---|---|---|
| [A] | [5] | [3] | … | … |
| [B] | [4] | — | … | … |
| [C] | [4] | [2] | … | … |
Five mistakes that send a methods chapter back
- "The questionnaire's validity was confirmed by professors," with no number of experts, method or figures.
- One alpha for a whole questionnaire with four different dimensions.
- Quoting the original paper's alpha instead of computing it in your own sample.
- Dropping items just to raise alpha, until one aspect of the content disappears from the questionnaire.
- Factor analysis on a 30-person pilot.
Frequently asked questions
What is the difference between validity and reliability in one sentence?
Validity means you measure the right thing; reliability means you measure it consistently. A reliable instrument may not be valid, but an unreliable one is never valid.
What is the minimum acceptable Cronbach's alpha?
0.7 is conventional. Between 0.6 and 0.7 is accepted in exploratory research or short dimensions, with an explanation; below 0.6, revise the dimension.
Can I compute CVR with 5 experts?
You can, but every item then needs an "essential" vote from all five. If you can't reach more experts, make the CVI the centre of your report; it is common with 3 to 5 experts, provided each item's I-CVI is 1.
My pilot alpha came out low. What now?
In order: (1) check the reverse-scored items; (2) look at "alpha if item deleted" and the item-total correlations; (3) rewrite ambiguous items using respondents' feedback; (4) if the dimension has two or three items, add a content-equivalent item and pilot again.
Where in the thesis do validity and reliability go?
In the methods chapter, under "data collection instrument," right after you introduce the questionnaire and its dimensions. Put the full questionnaire and the item-by-item CVR table in the appendix.
Is your questionnaire ready? Build it in Porsino's online questionnaire builder, run the pilot with one link and take the Excel or SPSS export straight into your alpha calculation. The usual next step is sample size: how to determine sample size.