AI diagnoses exactly what's wrong, rewrites it, shows you the before-and-after difference.
The AI wasn't being bad. It was doing exactly what you asked — which was almost nothing. 'Professional email' has 10,000 interpretations. Without knowing the audience, context, tone, and length, the AI picks the blandest, safest default. Your prompt is a blank canvas and the AI is painting beige.
Refinement prompts fail when the original prompt was too vague. You're trying to steer a car that was never pointed in the right direction. The fix isn't better follow-ups — it's a better first prompt that lands close enough to refine, not overhaul.
AI defaults to a specific 'voice': formal, thorough, hedge-heavy, bullet-point-prone. That voice exists because most people never specify an alternative. Three phrases in your prompt can eliminate the robot tone entirely. You just don't know which three.
You look at the output and think 'that's not it' but can't articulate what's missing. Is it the tone? The structure? The depth? The angle? Without a diagnostic framework, you're guessing at fixes and wasting time on trial and error.