AI Practice

Stress-Testing a Strategy With AI: A Method for Claude and ChatGPT

What generic AI feedback on a strategy document actually catches and misses, four prompts that surface real gaps today, and what changes when the AI applies Richard Rumelt's documented kernel framework instead of generic strategic-review advice.

By Gareth Hoyle·26 September 2026·8 min read

Paste a strategy document into AI and ask for feedback, and you'll get something encouraging with a few caveats attached: "this is a strong foundation, consider clarifying the competitive positioning, and the timeline may be ambitious." Reasonable-sounding. Also close to useless, because a generic review prompt invites a generically agreeable response, and most models default to being helpful rather than adversarial unless you specifically ask them not to be.

The better question isn't "review this strategy." It's "does this document actually contain a strategy, or is it a list of goals wearing one's clothes."

Why is generic AI feedback on a strategy document always so encouraging?

Most strategy documents that get pasted into an AI tool get back the same shape of response: praise for the ambition, a few soft suggestions ("consider adding more detail on X"), and a generally validating tone. It reads like feedback. It rarely challenges the actual premise of the plan, because a generic review prompt doesn't ask it to.

This matters because the most common failure mode in real strategy documents, according to Richard Rumelt's own published research across hundreds of corporate strategy reviews, isn't a bad idea. It's the complete absence of a diagnosis: a list of goals and initiatives with no stated account of what specific obstacle the strategy is actually solving for. A generically encouraging review doesn't catch that, because it isn't looking for it.

What can AI actually check in a strategy document, and what can't it verify?

Unaided, it's genuinely good at checking internal logical coherence, do the stated actions actually follow from the stated goals, are there initiatives that contradict each other, is the resourcing described consistent with the ambition claimed. It's also useful for generating a competitor's-eye read: how would a rival reasonably respond to this plan.

It's weak at verifying whether your diagnosis of the actual market situation is correct. If your document confidently asserts "our customers churn because of price" and that's factually wrong, the AI has no independent way to know that unless you give it real data. A logically coherent strategy built on a false premise will read as strong to an AI reviewing only the document's internal logic.

What's worth pasting into Claude or ChatGPT before this review?

These four run in Claude, ChatGPT, or Gemini without any special setup, just paste them into a fresh conversation.

1. The diagnosis check. "Read this strategy document: [paste it]. Find the specific sentence or section that states the actual problem or obstacle this strategy is solving for. If you can't find one, say so explicitly rather than inferring one." Why it works: this directly tests for the single most common failure in real strategy documents, missing diagnosis, rather than asking a vague "is this good" question that a generic review would answer with praise.

2. The coherence audit. "List every initiative or action item in this document. For each one, does it follow logically from the stated goal, or does it look like it was added independently without connecting back to the diagnosis?" Why it works: strategy documents often accumulate initiatives added for unrelated reasons (a stakeholder's pet project, last year's leftover plan), and this surfaces which actions are actually load-bearing versus decorative.

3. The competitor's-eye read. "Read this plan as if you were a smart competitor who wants to beat it. What's the most effective response you could make, and what part of this plan does it exploit?" Why it works: strategy documents are usually written entirely from the inside, and forcing an adversarial outside perspective surfaces vulnerabilities the authors were structurally unlikely to notice themselves.

4. The resource-reality test. "Given the budget and headcount described in this document, is the stated ambition actually achievable, or is there a mismatch between what's promised and what's resourced?" Why it works: an unrealistic resource-to-ambition ratio is one of the most common and most avoidable strategy failures, and it's a purely arithmetic check an AI can run reliably once given the actual numbers.

What's still missing once the stress test is done?

These four checks will genuinely surface real gaps in this specific document. They won't give you a standing bar that every future strategy review has to clear, this quarter's plan, next year's pivot, the reorg memo landing on your desk in six months. Without that fixed bar, the same missing-diagnosis failure slips through again next time, simply because nobody thought to check for it.

How does a documented framework actually change the review?

Richard Rumelt, drawing on decades of consulting and research later collected in Good Strategy, Bad Strategy, documented a specific structure he calls the kernel: a diagnosis (what specifically is the challenge, stated plainly), a guiding policy (the overall approach chosen to deal with it), and coherent action (a set of steps that actually implement the policy without contradicting each other). His central, repeated finding is that most documents called a strategy contain goals and a wish list of initiatives, but no kernel at all.

Before, generic prompting: asked to review a strategy document that opens with "our goal is to become the market leader in the next three years," a generic AI response engages with the goal directly, suggesting it's ambitious but achievable with the right execution, effectively accepting the goal as the strategy.

After, Rumelt's framework applied: the framework refuses to engage with the goal at all until a diagnosis is located. It asks specifically: what is preventing this company from being the market leader today, stated as a real obstacle rather than restated ambition. If the document has no answer, the framework's documented verdict is blunt, this is a goal, not a strategy, and no amount of polishing the initiatives underneath it fixes a document that never diagnosed anything in the first place. Only once a real diagnosis exists does the framework move on to checking whether the guiding policy and the listed actions actually cohere with it.

A generic review prompt has no built-in reason to enforce that gate, no diagnosis, no strategy, regardless of how detailed or ambitious the rest of the document is. The framework does, as a hard requirement rather than one more item on a longer list of suggestions.

What's the actual mechanism for loading this into Claude, ChatGPT, or Gemini?

Claude reads it as a Skill: upload the .zip at Settings → Skills → Add skill, and it auto-invokes based on whether your question matches the skill's description, or gets forced directly with /richard-rumelt-framework. In Claude Code, the same package unzips into ~/.claude/skills for use everywhere, or into one project's own .claude/skills folder if you want it scoped there only.

ChatGPT has nothing built for Skills specifically, so the plain .md becomes a Custom GPT's Instructions instead: Create a GPT → Configure → Instructions, and paste it in directly. The regular Custom Instructions fields won't hold it, they're limited to 1,500 characters; a Custom GPT's Instructions field runs to roughly 8,000, enough for a full framework.

Gemini has no comparable upload option, so the .md content goes into the system prompt where one exists, or opens the first message of the chat instead: "Here is a thinking framework to apply throughout this conversation. Read it carefully, then answer using this framework's approach," with the framework pasted below it.

Where does this point next?

For the plan already sitting on your desk, Decision Brief is a $79 tool built to structure exactly this kind of high-stakes call before resources get committed. For the discipline behind it, Rumelt's method sits in the Business Strategist category alongside Michael Porter, Roger Martin, and others, downloadable as .md files for Claude, ChatGPT, or Gemini.

FAQ

Frequently asked questions

Can AI tell me if my strategy will actually work?

No, and be suspicious of any tool that implies it can. Whether a strategy works depends on competitor responses, execution quality, and market conditions the AI has no privileged access to. What it can genuinely do is check whether your document contains an actual diagnosis and a coherent set of actions, or whether it's a list of aspirations dressed up as a plan, which is a different and more answerable question.

Why does AI feedback on my strategy doc always sound so positive?

Because a generic "review this" prompt invites a generically encouraging response, and most models default to being helpful and validating unless explicitly told to be adversarial. The fix is in the prompt: ask it to argue against the plan, or to find the diagnosis it's missing, rather than to review it. A request for a review gets you praise with caveats; a request for a stress test gets you actual pushback.

What if my strategy doc doesn't have an explicit diagnosis section?

That's exactly the gap this guide's framework is built to surface. Most strategy documents skip straight to goals and initiatives without ever writing down what specific obstacle they're actually solving for, which is Richard Rumelt's core documented critique of bad strategy generally. If the AI can't find a diagnosis when asked directly, that's a genuine finding about the document, not a failure of the prompt.

Is it risky to paste a confidential strategy document into an AI tool?

Treat it the way you'd treat any sensitive internal document you're sharing with an external tool: fine for most individual and small-team use, worth checking your organization's data policy first if you're on an enterprise account or the document contains material non-public information. The stress-testing method here works the same whether you paste the full document or a summary with the sensitive specifics removed.

How is asking AI to 'poke holes' in my plan different from Rumelt's framework specifically?

'Poke holes' produces scattered objections, some useful, some not, with no organizing structure. Rumelt's kernel gives the AI a specific target to check for: is there a genuine diagnosis, is there a guiding policy that follows from it, and do the actions cohere with both. That's a sharper, more diagnostic question than generic hole-poking, and it tends to surface the same root gap, no real diagnosis, that most weak strategy documents actually share.

Does this only work for corporate strategy, or does it apply to smaller plans too?

The framework was built for corporate strategy but the underlying test, is there an actual diagnosis, and do the actions follow from it, applies to any plan with real stakes: a product roadmap, a marketing plan, a personal career strategy. The scale changes; the test for whether it's a genuine strategy or a wish list dressed up as one doesn't.

What's the honest limit of using AI to stress-test a strategy?

It can check the document's internal logic, diagnosis, policy, coherent actions, but it can't verify whether your diagnosis is factually correct about the market you're actually in. A perfectly coherent strategy built on a wrong diagnosis will pass every structural test and still fail in the market. Use AI to check the logic; use domain expertise and real data to check the diagnosis itself.

Should I use this before or after I've already committed resources to the strategy?

Before, ideally, while changing course is still cheap. The kernel test is most useful as a gate before a plan gets funded or staffed, not as a postmortem after the fact. That said, running it on an in-flight strategy is still worth doing: a missing diagnosis found midstream is a genuine, actionable finding, just a more expensive one to act on than it would have been earlier.

Written by Gareth Hoyle. Last updated 26 September 2026. Part of the authority.md guides library.

Keep reading

More guides.