Methodology
How Measure works
Where the questions come from, how a team’s answers are read, how the plan at the end is chosen, and what we don’t claim. Written so that you can check each part against the code’s behaviour and the sample report.
In one paragraph. Measure is a survey for one team at a time. The organiser picks seven to ten statements from a public question bank, adapts the wording, and sends a link. People answer anonymously on a five-point scale, and when they report holding back they are asked why. The report shows how answers are spread, never an average as the headline; counts the reasons people gave; and ends with a short plan chosen by fixed rules from that pattern. It is not a validated scale and does not benchmark. Everything below says how each part works, so you can check it against the sample report.
- What it is forOne team, one conversation; research is a different job
- Where the questions come fromSeven themes, seven items adapted from Edmondson, your own statements
- The follow-up questionsWhy people hold back, and what makes speaking up possible
- How answers are readSpread, lowest answer, theme scores, “couldn’t say”, the written picture
- How the plan is chosenThe rules, in order, with no AI in them
- Anonymity by designWhat is never stored, and when the report opens
- Measuring againWhat the second report compares, and one caution
- What we don’t claimAnd what we will test when there is enough data
- ReferencesWith DOIs
- Citing this pageVersion and date
What it is for
Measure helps one team understand and improve its own psychological safety. That shapes every design choice. The questions fit the team rather than staying fixed; the output is a conversation and a plan rather than a number; and nothing compares the team with anyone else.
Two things follow. First, a survey is an intervention as well as a measurement. Asking people whether they can raise problems teaches them what psychological safety is, makes it discussable, and often encourages the behaviour it asks about. We design for that rather than pretending the thermometer leaves the room unchanged. Second, if you need comparable, validated measurement, for research or across studies, this is the wrong tool: use Amy Edmondson’s original seven-item scale unchanged (Edmondson, 1999). Validation is exactly what adapting a scale gives up, and we would rather say so than imply otherwise.
Where the questions come from
The question bank is public. Its statements sit under seven themes: speaking up, learning from failure, innovation, team cohesion, inclusion and diversity, transparency, and voice and power. Each is written in the first person about the person’s own recent experience (“On this team I can raise problems and challenge ideas”), because people can place themselves more honestly than they can summarise a whole team. Statements about the team as a whole offer an extra option, “I don’t have enough information to say”, so that nobody is forced to guess.
Seven statements are adapted from the seven items in Edmondson (1999). They are reworded, all phrased so that agreeing means safety, and answered on our five-point agree–disagree scale rather than her seven-point accuracy scale. Each is set beside the original, with what changed. The rest were written by us from practice, and the voice and power items come from the calculus of voice, our framing of the research on silence and voice for practical use.
An organiser chooses seven to ten statements, can edit the wording to the team’s own language, and can add up to five statements of their own, each placed under a theme and read exactly like the rest. Fourteen is the sensible ceiling: beyond it, people stop finishing.
A survey can also be offered in Spanish, German, French, Korean and Easy Read English, and each person chooses on the first screen. Easy Read is the same questions in short sentences and everyday words, for anyone who finds it easier. Every version keeps the same statements, the same themes and the same five answer points, and all the answers go into one report. The translations and the Easy Read wording are still being reviewed.
The follow-up questions
A low score tells you where to look, not what is happening. So when someone answers at the quiet end of a statement about speaking up, the survey asks a short, optional follow-up: when you hold back on this team, what is closest to why? The options are the same throughout:
I can’t predict how it would be received. Ambiguity. The remedy is making reactions legible and consistent.
I can predict it, and it would cost me. Valence. The remedy is changing what happens next time, visibly.
It wouldn’t change anything. Futility. The remedy is closing loops, including when the answer is no.
It depends on who is present. Audience, the power gradient. The remedy works on who can say what, in whose presence. A second question asks which kind of power: formal authority, informal influence, expertise, who is like or unlike me, or the people I am closest to.
I’m not sure. It’s just how it feels. A real answer, counted as one. Silence often runs below awareness, and a team that cannot name its reason needs language before it needs a plan.
The same five reasons are asked of the high end too, as what makes speaking up possible here, because a team that knows why it works has something to protect. All of these are counted across the whole survey and shown as tallies. They are never linked to a person or to the answer that prompted them.
The theory behind the four named reasons: silence as an active choice rather than an absence (Morrison & Milliken, 2000; Milliken, Morrison & Hewlin, 2003); the beliefs people hold about what speaking up will cost them (Detert & Edmondson, 2011); voice as discretionary behaviour (Van Dyne & LePine, 1998); and the expectancy logic of effort, outcome and value that makes it a calculation at all (Vroom, 1964). The calculus of voice is a lens built from that work, not a validated scale.
How answers are read
Every statement is shown as a distribution. How many people chose each of the five points, one mark per person, with the mean, the lowest answer and a spread figure beside it. Where the spread is wide (1.1 or more on the five-point scale) the report names the split as the most interesting thing on the page. Why an average on its own misleads.
Theme scores are a summary, and labelled as one. A theme’s score is the mean of its statements’ means, with reverse-worded statements flipped so that higher always means safer. Themes are listed lowest first, and the report says to read the spreads before trusting the score.
“Couldn’t say” is counted, never averaged. Answers of “I don’t have enough information to say” are shown as their own count beside each statement and across the survey. They are never turned into a middle score, because several of them is a finding: parts of the team are working out of each other’s sight.
Predictability and price. Two statements, one about whether people can foresee how speaking up will land and one about what it tends to cost, place the team in one of four positions. The report’s note differs for each, because clarity fixes an unpredictability problem and does nothing for a consequences problem.
The written picture is built by rules, from numbers only. A short paragraph describes the team in ordinary words. It is generated from the counts by fixed rules, with no AI and none of anyone’s comments, and it is marked experimental and asks whether it rings true. Where fewer than four people answered a follow-up, it reports what the few said rather than claiming a pattern for the team.
Comments come back verbatim. Shuffled, unedited, never summarised, and detached from the ratings their author gave.
How the plan is chosen
The report ends with a small number of things to try. They come from a library of practices condensed from Psych Safety’s post-survey action guide and the voice-and-power practices, each tagged with the reasons for holding back it answers and an effort rating from one (try it this week) to three (a structural change). The same answers always give the same plan. There is no AI in it, and the rules are these:
- Take the two lowest-scoring themes. The lowest gets three suggestions, the next gets two, and a third theme, if there is one, gets one.
- Within a theme, rank the practices by fit: the share of the team’s stated reasons for holding back that the practice addresses. A team whose reasons are mostly futility sees loop-closing practices first.
- Break ties by effort, cheapest first. A lead who tries one small thing and sees it work tries another; a lead handed a three-session programme as step one tries nothing.
- Never name the same practice twice in one report, even where it legitimately belongs under several themes.
- Add two suggestions aimed squarely at the single most-reported reason for holding back, whatever theme it came from.
The plan is a starting place for the team’s next conversation, not a prescription. The sample report shows the result for one worked example.
Anonymity by design
Most tools promise anonymity. Measure is built so that the promise cannot be broken by anyone, including us. No name, email address, account or IP address is recorded with a response. Only the date of submission is kept, never the time. Nothing sealed opens until the survey closes with at least four responses, and once anyone has read the report the survey cannot be reopened for more answers, so nobody can compare the report at four responses with the report at five. The organiser reads first, with a few notes the team’s copy does not have, and then releases the identical report to everyone. An organisation licence shows what teams report across the organisation and never opens a team’s own report, including to the people paying for it. Why four, and how the data is handled.
Measuring again
Surveys of the same team join up automatically, and the second report carries a section on what changed since the first, theme by theme, written as a prompt for a conversation about what happened in between rather than a scoreboard. One caution comes with it: scores on the who’s-in-the-room statements can fall as awareness improves, because people start noticing a gradient they had stopped seeing. How often to measure.
What we don’t claim
It is not a validated scale. The questions change from team to team by design, and the survey changes what it measures by measuring it. Scores from Measure cannot be set against published results.
There are no benchmarks and no norms. A team’s only meaningful comparison is with its own past. Why we don’t benchmark.
A number tells you where to look. Not what is happening, and not what to do. The follow-up questions and the comments carry that weight, and the plan is a hypothesis to test.
Versions are pooled, not proven equivalent. Answers given in another language or in Easy Read are added to the same scores as the standard English. That keeps one team in one report and lets everyone take part, but nobody has yet shown that each version measures exactly the same thing: a reworded statement can shift how people answer it. In a team where several versions were used, read small differences between themes with that in mind. We record which versions a survey offered, never which one anyone chose.
No claim of predicting outcomes. We do not claim that a Measure score predicts performance, retention, incidents or anything else.
What we will test when there is enough data
People who take the free self-check can add their ratings, as numbers only with optional coarse context such as sector and team-size band, to an anonymous research pool. It is small so far, and we would rather say so than publish a weak result. When it holds a few hundred contributions we will publish, at population level only: how consistently the statements within each theme move together; how often “couldn’t say” is used, and on which statements; the distribution of reasons given for holding back; and whether wide spread within a team goes with particular reasons. Nothing in the pool is used to benchmark, rank or compare teams, and findings will appear on this page first.
References
- Edmondson, A. C. (1999). Psychological safety and learning behavior in work teams. Administrative Science Quarterly, 44(2), 350–383. doi:10.2307/2666999
- Edmondson, A. C. (2018). The Fearless Organization: Creating Psychological Safety in the Workplace for Learning, Innovation, and Growth. Wiley.
- Morrison, E. W., & Milliken, F. J. (2000). Organizational silence: a barrier to change and development in a pluralistic world. Academy of Management Review, 25(4), 706–725. doi:10.5465/amr.2000.3707697
- Milliken, F. J., Morrison, E. W., & Hewlin, P. F. (2003). An exploratory study of employee silence: issues that employees don’t communicate upward and why. Journal of Management Studies, 40(6), 1453–1476. doi:10.1111/1467-6486.00387
- Detert, J. R., & Edmondson, A. C. (2011). Implicit voice theories: taken-for-granted rules of self-censorship at work. Academy of Management Journal, 54(3), 461–488. doi:10.5465/amj.2011.61967925
- Van Dyne, L., & LePine, J. A. (1998). Helping and voice extra-role behaviors: evidence of construct and predictive validity. Academy of Management Journal, 41(1), 108–119. doi:10.2307/256902
- Vroom, V. H. (1964). Work and Motivation. Wiley.
- Geraghty, T. The calculus of voice. Psych Safety.
- Psych Safety. Principles of ethical measurement in organisations.
The question bank is licensed CC BY-SA 4.0; the seven statements adapted from Edmondson (1999) are hers to cite, not ours to license. The report, the follow-up design and the plan rules are © Iterum Ltd.
Citing this page
Psych Safety (2026). How Measure works: the methodology behind the psychological safety survey at measure.psychsafety.com. Version 1, October 2026. https://measure.psychsafety.com/methodology.html
This page is versioned. When the method changes, the version number and date change with it and the previous wording is noted here. A DOI for the versioned text will follow.
Questions
Is Measure a validated psychological safety scale?
No, and it cannot be. Validation needs the same questions, unchanged, every time; Measure is built to be adapted to each team, and it treats the survey as an intervention that changes what it measures. For research, use Amy Edmondson’s original seven-item scale unchanged. Use Measure to understand and improve one real team.
How are the suggested actions chosen?
By fixed rules. The two lowest-scoring themes are taken, practices under each are ranked by how well they fit the reasons the team gave for holding back and then by how cheap they are to try, no practice is named twice, and two more are added for the single most-reported reason. The same answers always give the same plan.
Does Measure use AI to read the answers?
No. The charts, the tallies, the written picture and the plan are all produced by fixed rules from the numbers. Comments are shown word for word and are never summarised or processed by AI.
Why is there no overall score?
Because one number hides the thing that matters. A team where four people feel safe and two feel silenced can have the same average as a team where everyone feels roughly fine, and they need opposite conversations. The report shows the spread, the lowest answer and where people disagree instead.
Can I cite Measure in research or a report?
Yes. Cite this page by its version and date, and cite Edmondson (1999) for the seven adapted statements. The question bank is CC BY-SA 4.0, so you can reuse and adapt it with credit and a link.
Who made this
Measure is made by Psych Safety, drawing on over ten years of practice with teams in technology, healthcare, aviation, heavy industry, financial services and more.
Tom Geraghty is co-founder of Psych Safety. He started out in ecology, where his first job title was “Experimentalist”, and has been running experiments of one kind or another ever since. He was a CIO and CTO before moving into organisational change, and founded Iterum Ltd, the company behind Psych Safety. He holds an MBA and a postgraduate diploma in Global Health and Humanitarianism, and is studying for a PhD. Tom’s full bio
Jade Garratt is co-founder of Psych Safety and leads the design of its courses, workshops and toolkits. She read Physics at Oxford and has spent 18 years across education, the charity sector and business. She holds a Master’s in Educational Leadership and is completing a PhD in Education at the University of Nottingham. Jade’s full bio
