Questionnaire Validity and Reliability: Types, How to Measure Them, and a Sample Write-Up

Validity says the questionnaire measures what it claims to; reliability says it gives the same number if you measure again. The types of each, the minimum expert votes for Lawshe's CVR, computing the CVI, Cronbach's alpha with a worked example, test-retest and split-half, and ready-to-adapt text for your methods chapter.

When a thesis examiner reaches the methods chapter, the first two words they look for are validity and reliability. If that section is two vague lines ("the questionnaire's validity was confirmed by professors and Cronbach's alpha was 0.8"), it will very likely come back for revision. This article shows exactly what each type of validity and reliability establishes, how it is measured and with how many people, and ends with a sample methods text you only need to fill with your own numbers.

If you are still writing the items, read the questionnaire design guide first; here we assume you have the questionnaire and want to demonstrate that it is sound.

Validity and reliability at a glance

Picture a scale that always shows two kilos too many. Every time you step on it you see the same number, so it is reliable; but it doesn't show your true weight, so it is not valid. The reverse is impossible: an instrument that gives a different number every time cannot be valid. That is why reliability is a necessary condition for validity, not a sufficient one.

TypeQuestion it answersCommon methodUsual criterion
Face validityDo the items look clear and relevant to respondents?Feedback from 10–15 people from the population; item impact scoreImpact score of 1.5 or more
Content validityAre the items essential, and do they cover every aspect of the construct?Expert judgement: CVR and CVICVR above Lawshe's critical value; I-CVI above 0.79
Construct validityDo the data reproduce the theoretical structure?Exploratory or confirmatory factor analysisLoadings of 0.4 or more; fit indices
Criterion validityDoes the score agree with a valid external criterion?Correlation with another valid measure or a real outcomeSignificant correlation in the predicted direction
Internal consistencyDo the items of one dimension move together?Cronbach's alpha (or omega)0.7 or more
Stability (test-retest)Same result if we measure again?Second administration after 2–4 weeks; ICC0.7 or more
Split-halfAre the two halves of the instrument equivalent?Correlation of the halves with the Spearman–Brown correction0.7 or more

Face validity: ask the target population before any expert

Face validity is the simplest and cheapest step, and precisely for that reason it gets skipped. Give the questionnaire to 10–15 people from your target population (not classmates, not your supervisor) and ask which items they had to read twice, which words they didn't understand and which items seemed irrelevant. A nurse understands "self-efficacy" differently from a psychology student.

If an examiner wants a quantitative method, the "item impact score" is common: each person rates the importance of each item from 1 to 5, and for each item the frequency (the share who gave 4 or 5) is multiplied by the importance (the mean rating):

Impact score = frequency × importance = 0.8 × 4.1 = 3.28   (≥ 1.5 → keep)

Content validity: CVR and CVI, two indices for two different questions

Most theses report both, and many mix them up. CVR asks "is this item essential?" and CVI asks "is this item relevant, simple and clear?" So you need two questions for your experts: two separate forms or two separate columns in one form. You can build the expert form online too: each item is a row in a matrix question with three columns, "essential," "useful but not essential" and "not necessary"; the matrix report gives you the count in each column for every item.

CVR: is each item essential?

Each expert picks one of the three options for each item, and Lawshe's formula (1975) is:

CVR = (nₑ − N/2) / (N/2)

nₑ is the number of experts who said "essential" and N the total number of experts. CVR runs from −1 to +1; zero means exactly half the experts rated the item essential. An item whose CVR falls below Lawshe's critical value is dropped or rewritten.

Everyone quotes Lawshe's table, but what you actually need in practice is the third column: with this many experts, how many must say "essential"?

Number of expertsLawshe's minimum CVRMinimum "essential" votesCVR at that vote count
50.995 of 51
60.996 of 61
70.997 of 71
80.757 of 80.75
90.788 of 90.78
100.629 of 100.80
110.599 of 110.64
120.5610 of 120.67
130.5410 of 130.54
140.5111 of 140.57
150.4912 of 150.60
200.4215 of 200.50
250.3718 of 250.44
300.3320 of 300.33
400.2926 of 400.30

Two things follow from this table. First, with 10 experts the famous 0.62 effectively means "at least 9": 8 votes give a CVR of 0.6, which is below 0.62, so an item with 8 "essential" votes out of 10 is rejected. Second, the "minimum votes" column matches the exact binomial calculation that Ayre and Scally (2014) performed when they revisited Lawshe's table; so if the decimals of the table are ever questioned, the vote count is the safer criterion.

CVI: is each item relevant, simple and clear?

The item-level CVI is usually computed by asking each expert to rate the "relevance" of each item on a four-point scale (1 = not relevant, 2 = somewhat relevant, 3 = quite relevant, 4 = highly relevant), often repeated for "simplicity" and "clarity" (the approach associated with Waltz and Bausell). For each item:

I-CVI = number of experts rating 3 or 4 ÷ total number of experts

A common decision rule: items with an I-CVI above 0.79 are kept, between 0.70 and 0.79 revised, below 0.70 dropped. The scale-level index (S-CVI/Ave) is the mean of the I-CVIs, and Polit and Beck (2006) recommend 0.90 or more.

Example: 10 experts rated four items in one dimension:

ItemRatings of 3 or 4I-CVIDecision
A110 of 101.00Keep
A29 of 100.90Keep
A38 of 100.80Keep
A46 of 100.60Drop, or rewrite and ask the experts again

With all four items, S-CVI/Ave is 0.825; after dropping A4 it becomes (1 + 0.9 + 0.8) ÷ 3 = 0.90 and reaches the recommended threshold. Note that "CVI" is used in two senses in the literature: Lawshe also called the mean CVR of the retained items a CVI. Say which one you mean in your report.

How many experts do you need?

For the CVI, Lynn (1986) considered at least 3 experts necessary and more than 10 usually unnecessary; with 5 experts or fewer, every item's I-CVI should be 1. For CVR, because the critical values are very strict with few experts (5 to 7 experts means unanimity), 10 to 15 experts is a more practical choice. An "expert" knows the construct (a faculty member in the field) and the field itself (say, an experienced manager in that industry); a mix of both beats ten professors from one department.

Construct validity: when the data testify

Face and content validity rest on people's judgement; construct validity comes from the data. The key question is whether the items you placed in one dimension actually move together in real responses and separate from the items of other dimensions.

  • Exploratory factor analysis for researcher-made instruments, or instruments whose structure is unclear in your population. Prerequisites: KMO above 0.7 (0.6 at the very least) and a significant Bartlett's test of sphericity. An item loading below 0.4, or cross-loading on two factors with similar loadings, is a candidate for removal.
  • Confirmatory factor analysis for standard instruments whose structure is already known. Common fit criteria: RMSEA and SRMR below 0.08, CFI and TLI above 0.90; the stricter criteria of Hu and Bentler (1999) are CFI of at least 0.95 and RMSEA of at most 0.06.
  • Convergent and discriminant validity, mostly in structural equation modelling: average variance extracted (AVE) above 0.5 and composite reliability (CR) above 0.7 for convergence; for discrimination, the square root of each construct's AVE exceeding its correlations with other constructs (the Fornell–Larcker criterion, 1981), or HTMT below 0.85.

Factor analysis needs a large sample: the usual rule of thumb is 5 to 10 responses per item and at least 200 people. A 30-person pilot is not enough; report construct validity on your main data. How to size the sample properly is covered in the guide to Cochran, Morgan and power analysis.

Criterion validity

If another valid measure of the same construct exists, the correlation of your questionnaire's score with it (concurrent validity) or with a later outcome, such as next term's grades or staff turnover (predictive validity), is strong evidence. It is usually not required for a master's thesis, but if you have it, report it.

Reliability: three ways to ask "the same number?"

Cronbach's alpha with a complete example

Alpha measures the internal consistency of the items in one dimension and is computed for each dimension separately, not once for the whole questionnaire:

α = (k / (k − 1)) × (1 − Σsᵢ² / sₜ²)

k is the number of items, sᵢ² the variance of each item and sₜ² the variance of the dimension's total score. Suppose the dimension "satisfaction with online classes" has four items and the fourth is negatively worded ("I usually get bored in online classes"). Eight people's answers (hypothetical data, only to show the calculation) after reversing item 4:

RespondentI1I2I3I4 (reversed)Total
1453416
2334313
3544417
4233210
5435315
632229
7544518
8232310
Variance1.430.841.131.0712.29

The item variances sum to 4.46, so α = (4 ÷ 3) × (1 − 4.46 ÷ 12.29) = 1.33 × 0.64 ≈ 0.85. Now compute the same data without reversing item 4: the variance of the total drops to 3.43 and alpha turns negative (−0.40). A negative or very low alpha in a pilot usually means a forgotten reverse-scored item, not a broken questionnaire.

You don't have to do this by hand: put the item data into the Cronbach's alpha calculator to see alpha and "alpha if item deleted." In SPSS the same output is under Analyze → Scale → Reliability Analysis; the step-by-step is in analyzing a Likert questionnaire in Excel and SPSS.

Three points of interpretation:

  • Alpha is sensitive to the number of items. A two-item dimension struggles to reach 0.7, and a twenty-item dimension gets a high alpha even with loosely related items. If a dimension has two or three items, also report the mean inter-item correlation (roughly 0.2 to 0.5 is desirable).
  • Alpha above 0.95 is not a compliment. It usually means redundant items reworded; the questionnaire can be shortened.
  • A corrected item-total correlation below 0.3 flags a suspect item; but don't drop an item that covers part of the construct's content just to gain a few hundredths of alpha.

In recent years McDonald's omega is often reported alongside alpha, because it doesn't require alpha's strict assumption that all items carry equal weight. Recent SPSS versions and the free JASP compute it.

Test-retest: the same people, a few weeks later

Give the same questionnaire again to 20–30 of the same people after 2 to 4 weeks. The interval should be long enough that they don't remember their previous answers and short enough that the construct itself hasn't changed. It suits stable constructs (personality traits, attitudes); not this week's mood or pre-exam anxiety.

To compare the two rounds, the intraclass correlation coefficient (ICC) is better than Pearson's correlation: if everyone scores exactly one point higher the second time, Pearson is still 1, but an absolute-agreement ICC sees the shift. The common guidance of Koo and Li (2016): below 0.5 poor, 0.5–0.75 moderate, 0.75–0.9 good and above 0.9 excellent.

A practical tip: to pair each person's first and second answers, ask for a "personal code" only the respondent can rebuild (for example the last four digits of their phone plus their birth month). Anonymity is preserved and pairing the rounds is still possible.

Split-half and the Spearman–Brown formula

If a second administration isn't possible, split the items into two halves (usually odd and even), score each half and correlate them. Because each half is only half as long as the instrument, this correlation understates the reliability of the whole and is corrected with the Spearman–Brown formula:

r_SB = 2r / (1 + r)          r = 0.70  →  1.40 / 1.70 ≈ 0.82

The same formula has another use: it predicts how much reliability would rise if you lengthened a dimension. A six-item dimension with an alpha of 0.65 would reach about 0.76 if extended to ten items of similar quality. For tests with right/wrong answers (0 and 1), the equivalent of alpha is KR-20.

How many people does each step need?

StepWhoTypical number
Face validityPeople from the target population10–15
CVRExperts in the subject and the field10–15
CVIExperts3–10
Pilot and alphaA sample like the population, outside the main sampleAbout 30
Test-retestThe same people, twice20–30
Factor analysisMain sample5–10 per item; at least 200

Does a standard questionnaire need validity and reliability too?

Yes, but less. If you use a validated translation of a standard instrument, you report validity by citing the validation paper and recompute reliability in your own sample; alpha is a property of the data, not of the paper. If you translated the instrument yourself or changed an item, it no longer counts as "standard" and needs at least content validity and reliability afresh. How to tell a trustworthy version is covered in what a standard questionnaire is and where to find one.

Sample text for your methods chapter

Fill this in with your own numbers and sources; the figures in it are only examples:

Validity. To assess face validity, the questionnaire was given to 12 members of [the target population], and three ambiguous items were rewritten. Content validity was assessed by 10 [experts in the field]. The content validity ratio (CVR) of every item except item [7] was 0.80 or higher, above the minimum of 0.62 in Lawshe's table for 10 experts; item [7] was removed. Item-level content validity indices (I-CVI) ranged from 0.80 to 1, and the S-CVI/Ave was 0.92.

Reliability. The questionnaire was piloted with 30 members of [the population] who were not in the main sample. Cronbach's alpha for the dimensions [A], [B] and [C] was 0.84, 0.79 and 0.81 respectively in the pilot, and 0.86, 0.80 and 0.83 in the main sample (n = [292]), all above the conventional minimum of 0.7.

And the table examiners like:

DimensionItemsReversed itemsAlpha (pilot, n = 30)Alpha (main sample)
[A][5][3]……
[B][4]—……
[C][4][2]……

Five mistakes that send a methods chapter back

  • "The questionnaire's validity was confirmed by professors," with no number of experts, method or figures.
  • One alpha for a whole questionnaire with four different dimensions.
  • Quoting the original paper's alpha instead of computing it in your own sample.
  • Dropping items just to raise alpha, until one aspect of the content disappears from the questionnaire.
  • Factor analysis on a 30-person pilot.

Frequently asked questions

What is the difference between validity and reliability in one sentence?

Validity means you measure the right thing; reliability means you measure it consistently. A reliable instrument may not be valid, but an unreliable one is never valid.

What is the minimum acceptable Cronbach's alpha?

0.7 is conventional. Between 0.6 and 0.7 is accepted in exploratory research or short dimensions, with an explanation; below 0.6, revise the dimension.

Can I compute CVR with 5 experts?

You can, but every item then needs an "essential" vote from all five. If you can't reach more experts, make the CVI the centre of your report; it is common with 3 to 5 experts, provided each item's I-CVI is 1.

My pilot alpha came out low. What now?

In order: (1) check the reverse-scored items; (2) look at "alpha if item deleted" and the item-total correlations; (3) rewrite ambiguous items using respondents' feedback; (4) if the dimension has two or three items, add a content-equivalent item and pilot again.

Where in the thesis do validity and reliability go?

In the methods chapter, under "data collection instrument," right after you introduce the questionnaire and its dimensions. Put the full questionnaire and the item-by-item CVR table in the appendix.


Is your questionnaire ready? Build it in Porsino's online questionnaire builder, run the pilot with one link and take the Excel or SPSS export straight into your alpha calculation. The usual next step is sample size: how to determine sample size.

Tags:validity and reliabilitycontent validityCVRCVICronbach's alphathesisresearch methods