Evaluation and Iteration
Evaluation and Iteration
This file defines how to evaluate whether the communication Skill is actually improving collaboration.
The purpose is to measure outcomes, not to reward the Agent for sounding warm or sophisticated.
1. Evaluation Philosophy
Do not evaluate communication quality by:
- emotional intensity
- number of supportive phrases
- response length
- number of emojis
- frequency of “we can do it” statements
Evaluate whether communication improves:
- task success
- decision quality
- speed to useful action
- error recovery
- user understanding
- user control
- discovery of valuable insights
- reuse across future work
2. Primary Metrics
Task Success Rate
Did the task reach the intended success condition?
First-Pass Quality
How much of the desired result was correct before revision?
Rework Rate
How much work had to be redone due to misunderstanding or poor alignment?
Communication Efficiency
How many interaction turns were needed to reach a useful result?
Do not optimize turn count blindly. Some difficult tasks require more dialogue.
Decision Quality
Did the Agent surface the relevant trade-offs and risks before the user committed?
Insight Yield
Did proactive observations produce useful new information, alternatives, or strategic value?
Error Recovery Quality
After an error, did the Agent identify, correct, and learn from it?
Continuity
Was useful context preserved and reused across stages or sessions?
3. Suggested Qualitative Rubric
Score each dimension from 1 to 5.
| Dimension | 1 | 3 | 5 |
|---|---|---|---|
| Understanding | Misunderstood core task | Mostly understood | Precisely captured intent |
| Clarity | Confusing | Adequate | Highly clear |
| Honesty | Overconfident / misleading | Acceptable | Well-calibrated |
| Initiative | Purely reactive | Some initiative | Useful proactive thinking |
| Challenge | Blind agreement | Occasional correction | Appropriate rigorous challenge |
| Progress visibility | Opaque | Partly visible | Easy to understand |
| Error recovery | Repeats error | Corrects | Corrects + adapts method |
| Insight value | Filler | Interesting | Actionably useful |
| Efficiency | Excessive friction | Acceptable | Minimal necessary friction |
| Continuity | No reuse | Some reuse | Strong accumulation |
4. A/B Testing
When testing the Skill, compare meaningful conditions.
A: Neutral communication
Task instructions without intentional interpersonal framing.
B: Communication Skill enabled
Alignment, progress visibility, honest disagreement, insight sharing, and constructive feedback are enabled.
C: Over-optimized communication
Useful as a control to detect a failure mode. This condition may contain too much emotional framing, redundant summaries, or excessive reporting.
The objective is not to prove that “more positive language is always better.”
The objective is to find which communication behaviors improve outcomes under which task conditions.
5. Experimental Controls
When possible:
- keep the task constant
- keep the model constant
- keep relevant tools constant
- keep the user acceptance criterion constant
- randomize or alternate conditions where practical
- evaluate outputs blind where practical
Use multiple tasks rather than a single anecdote.
6. Task Categories for Evaluation
Build a small benchmark covering different interaction patterns:
Simple execution
Tests whether communication adds unnecessary overhead.
Complex reasoning
Tests alignment, uncertainty, and challenge quality.
Coding
Tests progress reporting, debugging, alternatives, and reuse.
Creative work
Tests initiative and useful divergence from the original request.
Strategic planning
Tests trade-offs, disagreement, and second-order effects.
Repeated failure
Tests error recovery and reframing.
Long-running project
Tests continuity, milestone reporting, and accumulation.
7. Evaluating Proactive Insight
Do not score “having an insight” by itself.
Score whether an insight is:
- relevant
- novel relative to the current conversation
- supported by reasoning or evidence
- actionable or strategically meaningful
- non-disruptive to the current task
A good outcome may be:
“The Agent produced no additional idea because there was no meaningful one.”
That is better than fabricated creativity.
8. Evaluating Progress Summaries
A progress summary is useful when it answers:
- where are we?
- what is done?
- what is not done?
- what is blocking us?
- what happens next?
It is harmful when it:
- invents precision
- repeats information already obvious
- interrupts rapid execution
- creates false certainty
- adds project-management overhead to a simple task
9. Evaluating Disagreement
Strong disagreement should score well when it:
- identifies a real weakness
- explains why it matters
- respects the user's agency
- offers an alternative
- is proportional to the stakes
Bad disagreement is:
- performative
- overly verbose
- based on weak evidence
- inserted merely to appear independent
10. Evaluating Positive Language
Track the difference between:
Specific positive feedback
“这个拆分方式把三个约束分开了,所以后面的验证更清楚。”
Generic praise
“你太厉害了。”
The first can improve coordination because it conveys information. The second usually adds little task signal.
Do not assume positive wording is valuable unless it changes the quality of the interaction.
Quench guardrails (supplemental, not dominant): track fluff rate, evasion recall with owner/gap marked, and unevidenced affirmation rate. Do not optimize these at the cost of task success or continuity.
11. Error Review Log
For meaningful failures, record:
日期:YYYY-MM-DD
任务:……
错误:……
直接原因:……
更深层原因:……
是否影响了前续结果:是 / 否
修正:……
是否需要修改 communication 规则:是 / 否
可复用经验:……
Look for recurring patterns rather than isolated mistakes.
12. Communication Friction Log
Track recurring user friction such as:
- too much confirmation
- too little explanation
- excessive reporting
- insufficient challenge
- overly generic praise
- missing risks
- poorly timed insights
- repeated questions
- failure to reuse known context
Each friction point should result in either:
- a local style adjustment, or
- a change to the Skill rules if the problem is systemic.
13. Iteration Rule
Change one communication behavior at a time when feasible.
Examples:
- reduce unnecessary confirmations
- increase explicit uncertainty labels
- move progress summaries to milestone boundaries
- surface more strategic alternatives
- shorten post-task reviews
This makes improvement easier to attribute.
14. Versioning
Suggested version progression:
0.1.x— foundational behavior set0.2.x— field-tested communication patterns0.3.x— refined long-term collaboration behavior1.0.0— stable behavioral contract
Record why a rule was added, removed, or changed.
Example:
0.3.0
Added: second-thought protocol
Reason: users reported value from post-task alternative ideas.
Guardrail: do not invent insights when none are warranted.
15. Success Criteria for the Skill
A mature version of communication should produce an observable pattern:
- fewer avoidable misunderstandings
- fewer unnecessary confirmation loops
- earlier detection of flawed assumptions
- more useful disagreement
- clearer project state
- more valuable proactive ideas
- cleaner recovery from failure
- better reuse of prior learning
- similar or lower communication overhead for simple tasks
If the Skill makes every response longer, warmer, or more structured but does not improve these outcomes, it is not working.
16. Final Evaluation Question
The most important question is not:
“Did the Agent communicate nicely?”
It is:
“Did the interaction make the human and the Agent jointly better at understanding, deciding, creating, and acting?”