Services-as-Software

For every dollar of software spend there are six or more dollars of services spend sitting next to it, and AI just made that services budget attackable by startups. The next trillion-dollar company sells the work, not the tool.

The Thesis

Sequoia partner Julien Bek's "Services: The New Software" thesis (2026): the software market is a fraction of the services market. QuickBooks costs ten thousand a year; the accountant who closes the books costs a hundred and twenty thousand. One to twelve. Most verticals run closer to one to six. Every software founder chases the tool budget. The real money is in the labor budget.

AI inverts the economics. A company that sells the outcome (meeting booked, NDA shipped, insurance policy closed) improves every time the model improves - cost of delivery drops, price stays put, margin widens, data moat deepens. A company that sells the tool is in a permanent race against the model and must stay upstream of it forever.

After Automation: Demand Moves To Judgment

Dan Shipper's "After Automation" is the best counterweight to the simplistic "agents replace the team" version of services-as-software. Every automated aggressively across coding, writing, design, support, and email; the company still grew from 4 to almost 30 humans. The mechanism is not sentimentality. Cheaper expert-shaped output increases the number of attempts, drafts, tickets, rewrites, support decisions, proposals, thumbnails, and product ideas that now need a human to frame, judge, route, and turn into a real operating process. Source: Every/Dan Shipper, 2026-05-21

The useful taxonomy is two modes of agent work. Agent employees are delegated actors in Slack or product workflows: support bots, proposal drafters, digest generators, research agents. Agent operating systems are Codex/Claude Code/Cowork-style environments where the human and multiple agents share one workspace. Both modes still need humans, but for different reasons: employee agents need maintenance and ownership; OS-style agents need a human "sandwich" around the task to set the frame, verify the output, and decide the next move. Source: Every/Dan Shipper, 2026-05-21

This is why services-as-software does not remove taste, skill, or management. It makes weak output abundant. Models commoditize the visible residue of old expertise, so default model output becomes easier to produce and easier to ignore. Differentiation moves to what is alive in the present: a specific customer, codebase, deadline, taste standard, risk surface, or blocked project. The human edge is not typing faster; it is knowing what should be asked, what should be rejected, and what should become a durable workflow. Source: Every/Dan Shipper, 2026-05-21

The benchmark lesson is the same. When a model saturates a benchmark frame, the cheap thing becomes attempting that framed work everywhere. If a "senior rewrite" benchmark becomes easy, more people request rewrites; the valuable work shifts to deciding whether to rewrite, what invariants to preserve, what rollback looks like, who reviews it, and whether the migration is worth the blast radius. In Kevin's stack, that means skills, doctors, evals, harnesses, and capsule routing are not optional garnish. They are the human judgment layer made executable. Source: Every/Dan Shipper, 2026-05-21

Copilot vs Autopilot

Two models that look similar but behave completely differently underneath.

Dimension Copilot Autopilot
Customer The professional (lawyer, banker, SDR) The buyer who wants the outcome
Risk holder The professional The vendor
AI role Assists the professional Replaces the professional
On model improvement Marginal productivity gain Margin expansion + moat deepening
Examples Harvey (law), Rogo (investment banking) Crosby (NDAs), WithCoverage (insurance), ColdIQ (meetings)
Vulnerability Next model release could turn the product into a feature Next model release makes delivery cheaper

The key insight: one model compresses your business, the other compounds it.

The ColdIQ Playbook (Zero to $7M ARR)

Liam Darmody left an $80K operations job at Worldcoin. First client paid $3K/month. Seven million ARR, four hundred B2B clients, thirty-plus people, bootstrapped. The six-step playbook:

  1. Pick one outsourced line item inside one industry. Three filters: already outsourced (existing budget line), mostly intelligence work (pattern recognition, not strategic judgment), services spend > software spend.
  2. Land first clients yourself. No website, no deck, no funnel. Record every sales call. The objections from the first ten calls become the sales page copy.
  3. Do the work by hand and document everything. Resist building the tool. The manual period is the training set. Four artifacts from day one: markdown SOP per task, Loom per cursor-based workflow, decision log per client, failed-campaigns file.
  4. Price like a service, report like a product. Sell the outcome. Build dashboards, telemetry, and reports the way a SaaS company would. Upfront setup fee + monthly retainer tied to outcome metric + performance bonus.
  5. Replace yourself on delivery before scaling anything else. Hire order: delivery operator → technical automator → head of delivery. No marketer, salesperson, or COO until delivery runs without you. You are the ceiling.
  6. Compound the data moat before the software. Save every input (raw + cleaned), every output (tagged with outcome), every reasoning trace from human judgment calls, every client objection and the response that closed it. Eventually agents run delivery while humans run judgment.

Why This Matters for Ideation

When evaluating new product ideas, this framework provides a core filter:

  • Is there a services budget 6x+ the software budget? If yes, there's an autopilot opportunity.
  • Is the work intelligence-heavy (pattern recognition + rule application)? If yes, AI can attack it.
  • Can you run it by hand first? The manual period produces the data moat no pure-SaaS competitor can buy.
  • Does model improvement help or hurt you? If you sell outcomes, better models = wider margins. If you sell tools, better models = existential threat.
  • What blocked project does this unblock? A buyer rarely pays because an abstract capability exists. They pay when an important project is stuck behind work, tokens, speccing, integration glue, or risk. See Blocked Projects Create Buyers. Source: User, 2026-06-29

This thesis pairs with Agent2Agent (A2A) (Agent-to-Agent protocol) as a current paradigm: A2A defines how agents coordinate; Services-as-Software defines what those agents should be doing - delivering outcomes, not assisting humans with tools.

Skills as Onboarding is the internal operating version of the same thesis. Anthropic's sales-ramp example packages top-rep behavior as an MCP connector plus five Claude skills, so new reps inherit procedure, tool access, and review criteria instead of waiting through a traditional ramp. That is not a SaaS feature in isolation; it is a service workflow compiled into executable memory. Source: X/@jasonlk, 2026-05-23

Long Live Hard SaaS is the counterweight to naive "AI kills software" thinking. Even when code generation gets cheaper, customers still pay vendors to absorb product-time, bugs, integrations, maintenance, and operational risk. Services-as-software expands the market by selling outcomes; hard SaaS persists by removing the maintenance burden for users. Source: raw/notes/saas-bug-moat-slack-thread-2026-06-22.md

The Convergence Point

Bek's long game: judgment turns into intelligence as the dataset grows. Three years of running the work gives you an asset no competitor can bootstrap. The transition from "humans do the work with AI assistance" to "agents run delivery while humans run judgment" happens quietly in the background for years before anyone outside the company notices.

The YC-founder anecdote is a thin but concrete org-design example of that transition: the company reportedly hires only product engineers and runs seven custom AI agents for the non-engineering functions, with a dashboard proving the work exists. Treat it as field evidence for the services-as-software operating shape, not a universal hiring recipe: product engineers still own the work definition, dashboards, integrations, and judgment loops. Source: X/@fin465, 2026-05-18


Timeline

49 pages link here

Aaref HilalyPeopleAdversarial Self-Audit PromptsConceptsAgency Agents - AI Agency GitHub RepoToolsAgent MachinesProjectsAgent Unit EconomicsConceptsAgent-Native CLIsConceptsAgent2Agent (A2A)ConceptsAgents as Service Workers (as a Service)ConceptsAustin LauPeopleBit HarmonyToolsBlocked Projects Create BuyersPhilosophiesBuild, Don't BuyPhilosophiesClickUp MCP ServerToolsClicky IRLInboxCodie SanchezPeopleConcept System MapConceptsContext Engine ArchitectureArchitectureDistribution Is the MoatPhilosophiesDistribution Wins on ImaginationPhilosophiesEarly Business ExplorationPhilosophiesFinn MalleryPeopleGraphedToolsGreg IsenbergPeopleHyperagentToolsImagination ChasmPhilosophiesLaunch Agent System (Matt Epstein)ToolsLive Process AuditingConceptsLLM Synthetic Panels (Purchase Intent)ConceptsLong Live Hard SaaSPhilosophiesLumachorProjectsOpen CoreConceptsOrigamiToolsOwnership Drives BuildingPhilosophiesPersistent Sandbox ThesisConceptsPersonal Principles (System Instructions for Myself)PhilosophiesPostHog CodeToolsPrompt-Caching EconomicsConceptsRecursive Agent OrchestrationConceptsRed-Team Your Business (Adversarial AI Audit)ConceptsSantiago Fernández de ValderramaPeopleServices-as-Software LensDecisionsShopify QuickToolsSkills as OnboardingConceptsWorld ModelsPhilosophiesX Bookmarks: AI Agents & Tools (Jan 2025 – Jun 2026)ToolsX Bookmarks: Career & Business (Mar–May 2026)CareerX Bookmarks: Dedalus (Oct 2025 – May 2026)ProjectsX Bookmarks: Dev Tools (Mar–Jun 2026)ToolsYC Verbs PositioningConcepts