By the time responses are in, the quality of your data is already fixed. Nothing in the analysis stage can recover a construct that was never measured properly.
This is the part of a project where an hour of care saves weeks.
Start from constructs, not from questions
Before writing any item, list what you are measuring and how each construct is defined. Then map every item to a construct. Any item that does not map to one should be removed, however interesting it seems.
The question to apply to each item: what would I do differently in my analysis depending on the answer? If the answer is nothing, cut it. Long questionnaires lower response quality and completion rates.
Use a validated instrument where one exists
If a validated scale exists for your construct, use it. You inherit its psychometric evidence and you can compare your results to published studies.
- Search the literature for existing scales before writing your own.
- Cite the original source, and check the licence — some require permission or payment.
- If you adapt wording for your context, say so explicitly and report reliability for your sample.
- Do not mix items from several scales into a new one and treat it as validated. It is not.
Item wording: the recurring faults
- Double-barrelled — "The service was fast and friendly." Someone who found it fast but rude cannot answer. Split it.
- Leading — "How much did you enjoy the training?" presumes enjoyment.
- Loaded language — words carrying evaluation shift responses.
- Ambiguous quantifiers — "often" means different things to different people. Use concrete frequencies.
- Double negatives — "I do not disagree that…" produces measurement error rather than insight.
- Assumed knowledge — jargon or acronyms your respondents may not share.
- Recall beyond memory — "how many times in the past year" is usually guessed.
Response scales
- How many points? Five and seven are both well supported. Seven gives slightly more variance; five is easier on mobile.
- Midpoint or not? Include one if genuine neutrality is possible. Forcing a choice when someone truly has no view manufactures data.
- Label every point, not just the ends. Fully labelled scales are more reliable.
- Keep direction consistent across the questionnaire; reverse-code a few items deliberately to catch inattentive responding, and remember to reverse them before analysis.
- Offer "not applicable" separately from the midpoint. They mean different things.
Structure and flow
- Open with an easy, engaging, relevant question — never with demographics.
- Group items by topic so respondents are not switching context repeatedly.
- Put sensitive questions later, once some commitment has been built.
- Put demographics at the end, unless you need them for screening.
- Use clear section headings and a progress indicator for online surveys.
- Keep it as short as the research question genuinely allows.
Pilot it — properly
A pilot is not a formality. Give it to five to ten people from your actual population and watch what happens.
- Ask them to think aloud. Where they hesitate is where the wording is unclear.
- Time it, and compare that to what you told respondents it would take.
- Test on a phone. Matrix questions in particular often break on small screens.
- Export the pilot data and run your intended analysis on it. If the export cannot be analysed the way you planned, you have found that out at the cheapest possible moment.
Questions this raises
Run a power analysis for your planned test. As a rough guide, factor analysis wants at least five to ten responses per item, and regression at least ten to fifteen per predictor.
Only if you treat the pre-change responses as a separate sample or discard them. Data collected on different instruments cannot simply be pooled.
Incentives raise response rates but can bias the sample toward people motivated by the incentive. If you use one, keep it modest and report it.
Still stuck after reading this? That is usually the point at which it is worth asking someone. Describe your project or ask on WhatsApp.