Open any off-the-shelf 360 feedback tool and you will find items like these:
“Demonstrates strategic vision.” “Exhibits executive presence.” “Shows strong leadership potential.”
Ask yourself what a rater is actually observing when they score those items. Nothing. Those are constructs, not behaviors. They live inside someone’s head, not in observable reality. And when you build a 360 on items like that, you are not measuring performance, you are measuring perception of a concept. That is a meaningful distinction that vendors quietly ignore.
The Observable Behavior Standard
The core rule of defensible 360 item writing is simple: if a rater cannot see it, you cannot measure it.
This is not new thinking. Behavioral science has been clear on it for decades. The problem is that most item writers either do not know the rule or find it inconvenient, because observable behavior is harder to write than vague constructs. “Communicates clearly” is easy. “Adjusts the level of technical detail in explanations based on the audience’s background” is work. The second item is the one worth measuring.
The Five Failure Modes
After building and auditing competency libraries for 360 platforms, I have found that bad items fail in predictable ways. Here is the checklist I run every item through:
- Not observable: the item measures a trait, mental process, or construct a rater cannot directly see (“thinks strategically,” “remains resilient”).
- Double-barreled: one item contains two distinct behaviors, joined by “and,” “while,” or “without.” A rater who sees strong performance on one and weak on the other has nowhere to go.
- Redundant: two items measuring the same behavior in different words. The data looks richer than it is.
- Vague language: weasel words like “effectively,” “appropriately,” or “high standards” that let any behavior qualify. These items rate the rater’s tolerance more than the ratee’s behavior.
- Inference required: the item asks raters to guess at motivation or intent rather than observe action (“acts in the best interest of the organization”).
Run any item through those five tests. Most commercial 360 items fail at least one.
Why Simplicity Wins
The instinct in survey design is to add complexity: more nuance, more qualifiers, richer language. It usually makes items worse.
The best items are short, specific, and describe exactly one visible behavior. They do not require the rater to interpret, infer, or average two things together. When data is unambiguous at the item level, everything downstream, feedback reports, development planning, normative comparisons, gets cleaner.
Measurement elegance is simplicity. The item that looks almost too plain is usually the one doing the most honest work.
What This Means If You Are Evaluating a 360
You do not need to build your own library to apply this thinking. Before you buy or deploy any 360 instrument, pull ten items at random and ask: can my raters actually observe this behavior? If the answer is frequently no, the data you collect will not tell you what you think it is telling you.
Want a faster gut check? Take the free two-minute assessment: is your 360 tool working against your culture? Ten questions, an instant score, and a plain read on whether your instrument measures your leadership reality or someone else’s.