1 Comment
User's avatar
Promptslove's avatar

The comparison to a battlecard workflow is a useful test of model quality because it forces the prompt to earn its place in a real process. I’d love to see which Claude behavior changed the outcome—better instruction following, stronger synthesis, or simply fewer brittle assumptions. I’ve had the best results when the prompt also specifies how to flag uncertainty.