Claude Code Runs Our Entire GTM Pipeline
115 skills. 3 autonomous agents. 30 minutes of human time per week.
Most agencies use AI to write cold email scripts. We did that too in 2025. Then we realized the bottleneck in outbound isn't script writing. It's everything else.
Monday morning audits. Which campaigns are performing? Which segments are saturated? Which clients have capacity sitting idle? What should we launch next?
Daily monitoring. Did bounce rate spike overnight? Is this campaign stalling? Should we reallocate volume from this underperformer to that winner?
Intelligence compounding. What did we learn from bakery campaigns for one client that applies to coffee shop campaigns for another?
All of that was manual. People checking dashboards, updating spreadsheets, making judgment calls campaign by campaign. Every single week.
So we turned Claude Code into the operations layer.
What Claude Code Actually Is
Quick context if you haven't used it. Claude Code isn't a chatbot. It's Anthropic's CLI tool - runs in your terminal, reads your files, executes commands, writes code. An AI team member that lives inside your codebase.
The key feature for GTM: skills. A skill is a reusable workflow you teach Claude Code once, then invoke with a slash command forever. /launch runs our full campaign launch. /delivery review runs our weekly performance review. /campaign-ideation generates ranked campaign ideas from client context.
We have 115 of these. Every repeatable workflow we've ever built is now a single command.
Skills are just the interface. The real power is what's underneath.
The Three Systems
Three agents running in production right now, all built with Claude Code.
Campaign Agent. Runs every Monday at 6am. A 6-phase launch orchestrator for every active client.
Phase 0 auto-pauses campaigns that are 95%+ complete. Frees up sender capacity. Flags senders that still have warmup on past 14 days - wasted capacity nobody catches manually.
Phase 1 pulls every campaign stat from EmailBison across the workspace. Classifies every segment by status.
Phase 2 scores idle segments using a 5-factor model. Interest rate at 35% weight, recency at 25%, saturation at 20%, seasonality at 10%, list availability at 10%. Every segment gets a composite score. Top picks get scraped first.
Phase 3 runs our local-enrich pipeline. Google Maps scraping, owner finding, email discovery, verification. For B2B SaaS clients, it routes through different enrichment APIs instead. All automated.
Phase 4 creates campaigns with the right copy auto-selected from an offers library based on segment and cooldown, attaches senders, sets send schedules, uploads verified leads. Full hands-off.
Phase 5 creates Linear issues with the full audit, actions taken, and recommendations for anything it couldn't handle autonomously.
The whole thing runs unattended. Monday morning we wake up to a Linear notification showing what it launched, what it paused, and what it needs a human decision on.
Campaign Checkpoint. Runs every 6 hours. The safety net and optimizer. Three layers.
Layer 1 is the circuit breaker. Hard auto-pause on any campaign exceeding 5% bounce rate. Flags campaigns between 3-5%. Our first dry run caught a campaign hitting 10.88% bounce on Microsoft domains. Would have tanked deliverability if it ran unchecked.
Layer 2 is the signal interpreter. Detects milestones - 500 sends, 1000 sends, first reply. Compares sibling campaigns against each other. "Campaign A has 3x the reply rate of Campaign B in the same segment" becomes an actionable signal, not a stat buried in a dashboard.
Layer 3 is a Thompson Sampling bandit. Bayesian modeling that dynamically allocates send volume across sibling campaigns. Winners get more volume. Losers get scaled back. 14-day rolling window. 15% floor so nothing gets completely starved. 200% ceiling so nothing runs away.
That last layer means we're not just monitoring. We're actively optimizing send distribution based on real-time data. Every 6 hours.
Delivery Scanner. Runs every Sunday night. Calculates ELR across every active campaign, flags underperformers with a specific diagnosis.
INFRASTRUCTURE - high bounce, sender or domain issues. OFFER - low interest despite sends, copy isn't landing. ANGLE - wrong audience angle. SEGMENT - exhaustion, TAM burned through. STALLED - campaign stuck, not sending.
Then it pulls the actual email copy of flagged campaigns and compares it side-by-side with the workspace's top performer. So you see exactly what's different.
Monday morning I type /delivery review in Claude Code. Review the recommendations. Approve or modify. Claude executes the actions.
Total time spent on campaign operations per week: about 30 minutes.
The Context Layer
Here's the part most people miss when they think about AI agents.
ChatGPT has no memory of your business. Every conversation starts from zero. You paste context in, it forgets it by next session.
Claude Code reads your files. That changes everything.
We have a git repo called revgrowth-context that stores everything Claude needs to make good decisions.
29 client profiles with ICP details, approved angles, and learnings from every past campaign. Cross-client intelligence patterns - "Google email infrastructure outperforms Microsoft 14x on reply rates" learned from one client, now applied to every client's segment launches. Per-client YAML configs defining segments, approved angles, seasonality, safety thresholds.
When the campaign agent scores a segment, it's not guessing. It's pulling from performance benchmarks, checking what worked for similar segments across other clients, and selecting offers based on what's been tested before.
Every campaign we run makes the next one smarter. That's intelligence compounding. It only works because Claude Code can read structured context at runtime - not because someone pasted a prompt into a chatbot.
Linear as Persistent Memory
One problem with AI assistants: they forget. Conversation context gets compressed. Windows fill up. Yesterday's analysis is gone today.
We solved this with Linear. Every campaign becomes a trackable issue with full history.
Launch data - lead count, campaign ID, senders attached. Performance metrics logged as comments at each checkpoint. Verdicts - scale, iterate, pause, kill - applied as labels. Version tracking where v2 links back to v1, so you can follow the full evolution of an angle from first attempt through scale.
The campaign agent writes TO Linear. The delivery scanner reads FROM Linear. Claude Code reads the issue history before making any recommendations.
Nothing gets lost between sessions. When I start a new Claude Code session and say "what's happening with client X campaigns" - it reads Linear and gives me full status in seconds. No dashboard hunting. No spreadsheet cross-referencing.
Every launched campaign gets a 14-day analysis deadline. Linear shows what's overdue. No more campaigns silently sending for 6 weeks with nobody checking results.
What a Normal Week Looks Like
Monday 6am. Campaign agent runs. Auto-pauses 3 saturated campaigns, launches 2 new ones - coffee shops v4 with a tariff savings angle, bakeries v3 with a social media angle. Creates Linear issues for both.
Monday 8am. I open Linear. See the agent's report. Review the 2 launches. Check the recommendations it couldn't auto-execute - budget cap hit on brewery segment, scrape cost needs manual approval.
Tuesday through Saturday. Checkpoint runs 4x/day. Campaign A hits 4.8% bounce - flagged. Campaign B crosses 500 sends with 0 replies - signal generated. Campaign C is crushing it at 6.2% reply rate - bandit starts reallocating volume from B to C.
Sunday 9pm. Delivery scanner generates the weekly review. Campaign A recommended for investigation. Campaign B recommended for pause. Campaign C recommended for scale.
Monday 8am. I type /delivery review in Claude Code. Approve pause on B, investigate A, scale C. Done.
30 minutes of my time. The rest was autonomous.
How to Start Building This
You don't need the full stack on day one. Here's the progression.
Start with Claude Code skills. Encode your repeatable workflows as slash commands. Lead enrichment, campaign creation, performance checking. Every workflow you do more than twice should become a skill. Claude Code skills are just markdown files with instructions - no code required to start.
Add a context repo. Store client profiles, campaign learnings, approved messaging in a git repo. Point Claude Code at it. This is what turns it from a stateless chatbot into a team member with memory.
Build your first agent. Start with the audit phase only - have Claude Code pull campaign stats from your sending platform, calculate key metrics, output a report. Run it manually first.
Add scheduled execution. launchd on Mac, cron on Linux, GitHub Actions for cloud. Make it run without you.
Layer in safety gates. Bounce ceilings, budget caps, minimum list sizes. The agent should make decisions, but it should know when to stop and ask a human.
Connect a project tracker. Make the agent write its results somewhere persistent. This turns a script into an operations system with a paper trail.
The system keeps producing without rebuilding anything. That's the whole point.
If you want the full playbook behind these systems - how we set up sending infrastructure, build lead lists, write scripts that actually convert, and orchestrate + everything end to end - it's all inside Outbound Secrets. 70+ training videos covering the workflows these agents automate.
Link here: https://outbound-secrets.com/
Peace,
Adam