
Despite all the hoopla about pacing the frontier, LLMs keep evolving at breakneck speed. This got me thinking about the prompt workflows I previously created, so a few days ago I ran an experiment that I fully expected to win. I opened a fresh Claude session with no project files, no memory, and none of my skills, and I typed a single sentence: “Write me a prompt workflow for building a competitive battlecard for B2B sales.” Then I compared it with the Strategic Battlecard Prompt Workflow I published on Substack in July 2025, one of the roughly one hundred prompt workflows I have written over the past couple of years.
Claude’s version was better. It made win-loss notes and call recordings the core input, and it added landmine questions that expose a competitor’s weak spots without naming the competitor. It also ended with a list of claims reps should never say and a red-team review written from the perspective of the competitor’s VP of Sales. My 2025 version had none of those, and I copied two of them into my own battlecard skill that same afternoon. My battlecard was designed to be printed and fit on one page (back-to-back). Claude’s workflow asked for one page too, but in practice that’s one of the downfalls of LLMs: they generate verbose, copious amounts of text that sales reps can’t consume. In fact, no one can, because the outputs are too long.
Having been in this business for a while, I could see it was missing a critical ingredient. What was that, you ask? Well, it was the notion that key messaging or field validation from reps outranks everything the research says. If a rep loses credibility because you created an erroneous battlecard, they stop trusting it and stop trusting you. So after a hundred workflows, here’s what I’ve learned.
Prompts are cheap now, and they get better with every model release, but a library like mine still maps which GTM jobs feed which. Customer interviews feed the messaging framework, for example, and the messaging framework feeds both the launch plan and the cheat sheet a rep keeps open on a call. The library also holds rules you only pick up by doing the work, and the rule about field validation is one of them. AI GTM engineering is the work of turning those connections into systems that run.
A fresh model writes a better prompt than I did last year
Here’s how they compared, row by row, with the battlecard skill I run today added as a third column. The fresh model had nothing to work from except that one sentence, so you can paste it into any new chat window and repeat this cold-prompt test yourself in about two minutes. Rows marked “added after this test” are the two steps I copied from the model that afternoon.

The fresh model matched or beat my 2025 workflow on each of the top six rows. The bottom four rows tell a different story, because each one depends on something a chat window can’t see, such as the win-loss files, the deals in the pipeline, and the corrections reps have made along the way. Each one came out of building this process for clients.
In fairness, I pitted a 2026 model against a prompt workflow I wrote with 2025 tools, so the comparison says as much about how far the models have come as it does about my prompt. That’s exactly why it matters to anyone sitting on a prompt library, and I’m one of them. The step-by-step instructions improve on their own every time a new model comes out, so a folder of well-written prompts stops being much of an advantage, no matter how many hours went into it.
What the model couldn’t know
The rule that field observations outrank the research came from a battlecard system I built for a client earlier this year. Alongside the battlecards, I created a field-intelligence log, a running file of corrections from the client’s sales reps and product team. Every entry records the claim, who made the correction and when, how confident they are, and which assets have to change. When an entry conflicts with a competitor’s documentation, an analyst report, or my own research, the field-intelligence entry wins.
One correction shows how the log works in practice. The product team pointed out that the client’s materials listed four deployment options when there were only three, because one of the four was a product feature that had been mistaken for a way to deploy the software. That single entry had to reach every battlecard, feature matrix, and web page that mentioned deployment, and the log tracked which of them were still waiting on a fix.
“Field validation from reps outranks everything the research says. If a rep loses credibility because of an erroneous battlecard, they stop trusting it, and they stop trusting you.”
— David Sweenor, Founder/CEO, TinyTechGuides
Nobody writes a rule like that into a prompt until they have seen what happens without it. At another client, the person who built a battlecard against a large platform vendor told the team he checked every claim to make sure none overstated anything. He was happy to stand behind those claims on a call. He worried about anything stronger, such as claiming the competitor lacked a capability it had announced only two weeks earlier. The prospect’s own engineers often know the competitor’s product better than the seller does, and when they catch an overstated claim like that, it’s an argument a rep can’t win.
When a battlecard gets something wrong in front of a customer, fugettabout it. Reps go back to asking the sales engineer down the hall, and the deals the card should have helped win get lost for reasons that never appear in a pipeline report. For a CMO, that’s the real cost of a battlecard nobody maintains, and no model can see it from inside a chat window.
From a skill to a system
A Claude skill is one packaged job that you trigger with a slash command, and on its own, it waits for somebody to remember it exists. My battlecard skill is more useful than that because it depends on other jobs. It won’t build a card until my site monitor, a market research report, or a win-loss analysis has produced something about that competitor, and it reads the log of rep and product corrections last, so those corrections win.
Dependencies like that turn a prompt library into something more valuable than a set of prompts. Read my hundred workflows as a group, and they sketch a map of GTM work in which competitive intelligence feeds the battlecard, the battlecard feeds the rep, and the rep’s wins and losses feed the next round of competitive intelligence. When I built the B2B PMM Field Kit, the workflows sorted into five regions of that map: foundations, intelligence, competitive, launch and content, and the buying committee. Tool vendors tend to draw the map from what their own products touch, so it rarely matches the one a practitioner draws from the work.
A system also needs to know when to run. A battlecard goes out of date for predictable reasons, and my skill now watches for four of them. Two come from the market: a competitor pricing change caught by the site monitor and a new win-loss analysis, and two come from the field: a major deal won or lost and a high-confidence correction that names the card. Each is an event a system can act on without anyone’s calendar reminder, so the battlecard passes the test I set in What AI GTM engineering is: whether something you built still works the week you’re on holiday.
Three questions to ask of your own prompt library
Whether your prompt library is a shared Google Doc or a hundred Substack posts, three questions will tell you what it’s worth now. None of them takes more than an afternoon to answer, and together they show which workflows are ready to become part of a system.
Would a fresh model write it just as well? Run the cold-prompt test on your best workflow by asking a new chat session, with no context, for the same thing. If the model ties or wins, the prompt is your starting point rather than your advantage, so borrow whatever the model added and move on to the next two questions.
What does it read, and where does its output go? Name the inputs a workflow depends on and the place its output ends up. A workflow with neither is still a prompt someone has to remember, and connecting it to the jobs on either side makes it part of a system.
What should make it run again? Write down the events that would make its output wrong, such as a pricing change, a lost deal, or a correction from the field. Those events become the triggers, and the corrections become rules the system enforces every time it runs.
Over the next few posts, I’ll build these systems one at a time, starting with competitive intelligence, and show what each one reads, what it hands off, and what makes it run. For anyone who’d rather not wait, I’m happy to spend thirty minutes looking at your prompt library with you and picking the workflow most worth turning into a system first. A fresh model will keep writing better prompts every year, and it still won’t know which of your battlecards a rep stopped trusting last quarter.
Interested in learning more? Book a consultation.
Frequently asked questions
What is the cold-prompt test for prompt workflows?
The cold-prompt test compares your prompt workflow with one a fresh AI model writes with no project files or memory, given a single sentence of instruction, such as “Write me a prompt workflow for building a competitive battlecard for B2B sales.” When David Sweenor ran it in September 2026, a fresh Claude session wrote a competitive battlecard workflow that beat his July 2025 published version. The model’s version added win-loss inputs, landmine discovery questions, a list of claims reps should never say, and a red-team review. The test takes about two minutes and shows whether a prompt is still an advantage.
Are prompt workflow libraries still worth anything now that AI models write good prompts?
A prompt workflow library is worth less for its prompts and more for what it records about the work. A library of about a hundred B2B marketing workflows sketches a map of which go-to-market (GTM) jobs feed which, such as competitive intelligence feeding the battlecard and win-loss results feeding the next round of intelligence. It also holds rules learned from doing the work. AI GTM engineering uses that map and those rules as the raw material for systems that run.
What is the difference between a Claude skill and a GTM system?
A Claude skill is one packaged job that you trigger with a slash command, and it waits for someone to call it. A GTM system connects several skills so one job’s output becomes the next job’s input, and it runs when specific events happen. A battlecard skill becomes part of a system when it waits for site monitoring, market research, or win-loss analysis to produce something about a competitor, applies field corrections last, and rebuilds when a pricing change or major deal outcome makes the card wrong.
What is a field-intelligence log for competitive battlecards?
A field-intelligence log is a running file of corrections from sales reps and the product team that outranks a competitor’s documentation, analyst reports, and internal research. Each entry records the claim, who made the correction and when, how confident they are, and which sales assets have to change. When a product team corrects a claim, such as the number of deployment options a product supports, the log tracks the fix across every battlecard, feature matrix, and web page that mentions it.
When should a competitive battlecard be refreshed?
Refresh a competitive battlecard whenever an event makes it wrong; four events cover most cases. Two come from the market: a competitor pricing change caught by site monitoring and a new win-loss analysis, and two come from the field: a major deal won or lost against that competitor and a high-confidence correction from a rep or the product team. A GTM system can watch for all four and rebuild the card without anyone reminding it.
Where should a marketing team start turning prompt workflows into GTM systems?
Start by asking three questions of your best prompt workflow. First, would a fresh AI model write it just as well? If so, treat the prompt as a starting point. Second, what does the workflow read, and where does it send its output? Third, what events should make it run again? Workflows with clear inputs, clear destinations, and clear triggers, such as competitive battlecards, are the best first candidates for an AI GTM engineering system.
About David Sweenor
David Sweenor is the founder and host of the Data Faces podcast, where he talks with the people who are making data, analytics, AI, and marketing work in the real world. He is also the founder of TinyTechGuides and a recognized top 10 AI thought leader and international speaker who specializes in practical business applications of artificial intelligence and advanced analytics.
With over 25 years of hands-on experience implementing AI and analytics solutions, David has supported organizations including Alation, Alteryx, TIBCO, SAS, IBM, Dell, and Quest. His work spans marketing leadership, analytics implementation, and specialized expertise in AI, machine learning, data science, IoT, and business intelligence. David holds several patents and consistently delivers insights that bridge technical capabilities with business value.
Books
- Artificial Intelligence: An Executive Guide to Make AI Work for Your Business
- Generative AI Business Applications: An Executive Guide with Real-Life Examples and Case Studies
- The Generative AI Practitioner’s Guide: How to Apply LLM Patterns for Enterprise Applications
- The CIO’s Guide to Adopting Generative AI: Five Keys to Success
- Modern B2B Marketing: A Practitioner’s Guide to Marketing Excellence
- The PMM’s Prompt Playbook: Mastering Generative AI for B2B Marketing Success
Follow David on Twitter @DavidSweenor and connect with him on LinkedIn.


The comparison to a battlecard workflow is a useful test of model quality because it forces the prompt to earn its place in a real process. I’d love to see which Claude behavior changed the outcome—better instruction following, stronger synthesis, or simply fewer brittle assumptions. I’ve had the best results when the prompt also specifies how to flag uncertainty.