A number is not an instruction
An engineer told their specificity is 6.4 has learned nothing they can act on. An engineer handed the rewritten request has.
Most tools that look at prompts will tell you a score and stop, which leaves a leader holding a dashboard and no decision. This scores the work six ways, names the specific thing that went wrong, hands back the version that would have worked, and turns the patterns that keep winning into your organisation's default.
Six dimensions, scored per session and rolled up per engineer
Scoring is the easy part and it already exists. The four steps below are what turn it into something that changes output.
A single quality number tells you something is wrong without telling you which part.
What Tetriz does here
Plus an overall, per session, rolled up per engineer and per team.
A request that was perfectly clear and missing all its context is a different problem from a vague one.
Knowing specificity scored low is still not a thing anyone can do on Monday morning.
What Tetriz does here
An ambiguous directive, a missing acceptance criterion, internal context the agent could not have seen.
So the conversation is about the work rather than about the number.
| Issue | Severity | Sessions |
|---|---|---|
| Ambiguous directive | Medium | 61 |
| No acceptance criteria | High | 44 |
| Missing internal context | High | 38 |
This is where most tools stop. A diagnosis without a prescription leaves the engineer exactly where they were.
What Tetriz does here
Not a principle to apply — the actual improved version, specific to that task and that codebase.
The tokens the weaker version burned, so improvement is visible in spend as well as quality.
A fix that helps one engineer once is worth very little. The same fix as an org default is worth it every week.
What Tetriz does here
Requests that consistently produce shipped code are codified into your org's prompt library and rules files.
Someone joining starts from your best-performing pattern instead of assembling one over a year.
The gap between measuring prompt quality and improving it is the whole of this page.
An engineer told their specificity is 6.4 has learned nothing they can act on. An engineer handed the rewritten request has.
Generic prompting advice is freely available and does not work, because what makes a request answerable depends on your architecture, your conventions and your naming.
Individual improvement plateaus and walks out of the door. A pattern promoted into the org's standard applies to every engineer who arrives after it.
Clarity, specificity, context provision, actionability, completeness and efficiency, plus an overall score. Each is reported per session and rolled up per engineer and team.
The engineer sees their own scores first. Their manager and organisation leadership see them after that, because a score is something to coach on rather than to rank by, and it should never reach a manager before the person whose work it describes.
You get the improved version alongside the original. Nothing intercepts or edits what an engineer types, and adopting a recommendation is their choice or an admin's.
No. Length is not a dimension. What is scored is whether the request was described well enough to come back usable, which a longer one does not achieve on its own.
Detected, and handled as a security finding rather than a quality one. That surface is Session Security, under Governance.