Chapter 02
AI systems

Different machines. One direction.

AI is my passion, and I believe we are at the start of a new era in how work gets done. Research, marketing and software no longer need to be run by hand, so I don't run them that way. I'm a self-learner and an early adopter by nature, and right now I'm building several projects in this space at once. Every robot on this page stands for a real role in one of them.

AI Marketing AgentsVC CatalogAI Software Factory

My current team of ninjas

  • Orchestrator
  • Strategist
  • Researcher
  • Media operator
  • Analyst
  • Developer
  • Tester
  • Deployment engineer
Project 01

AI Marketing Agents

The know-how stays human. The working hands are agents.

Most of marketing can be automated today. Not the thinking: the strategy, the judgment about what a brand should say and to whom, the logic of what works and why. That stays the creative brain. What changes is who does the work, and in my setup the working hands are agents.

I spent my career building marketing teams, and now I build them again in a different material. Each agent owns one discipline, from research and creative to social, partnerships and performance analytics, and they all share one memory of the brand, the voice and what has worked before. The team runs at a pace no human team could sustain, and it gets a little smarter with every cycle.

  • 10 specialist agents
  • One shared brand memory
  • Human strategy, agent execution
  • 1,200+ creatives produced
  • Performance analytics built in

Technical deep dive · AI Marketing Agents

Narrow jobs, shared memory.

Every agent is built in the same order, from the bottom up. First the scripts that do the mechanical work, then the skills where the judgment calls happen, and only at the end the agent file that routes between them. It feels slow at first. It pays off the first time a platform changes and only one script needs fixing.

Pipeline

  1. 01 Brief A topic, the channels and the brand rules. For partnerships, a short description of who we're looking for.
  2. 02 Evidence Studies, public conversations or blind competitor searches. Every fact keeps a link back to where it came from.
  3. 03 Produce Several models and tools per asset: Claude writes, image and video models generate, code assembles to brand rules, then a proof-read pass.
  4. 04 Engage Social channels, influencers and affiliates run through agents within set limits, with pacing, hard stops and a full action log.
  5. 05 Learn Logs, playbooks and the rule base are append-only, so each run starts with whatever the last one learned.

The agents

  • Research to Ideas Content strategy Three sub-agents in a row. One finds recent studies, one turns them into 20 or so ideas, and one fact-checks every claim before the list goes anywhere.
  • Content Creator Creative production Turns ideas into carousels, single images, infographics and decks, combining language, image and video models. Around 1,200 creatives so far, each through a proof-read.
  • Social Media Manager Community Runs the brand's social presence in my voice, scoring what to engage with against my views. Every comment is checked for the usual AI tells first.
  • Influencers and Affiliates Partnerships Finds creators and partners who fit a brief, verifies each through four gates, and manages the pipeline from a tracked manifest.
  • Competitor Landscape Market intelligence Five research streams that can't see each other's results, so they can't all copy the same mistake. One run mapped 554 companies and 154 real direct competitors.
  • Social Listening Audience research Reads public conversations at scale to find what people actually worry about. One study went through about 345,000 posts and comments to find around 3,400 that mattered.
  • PR Kit Earned media Drafts press releases as designed PDFs, builds a journalist list with a confidence note on every address, and writes pitch templates.
  • Video Ads Paid social creative Makes short vertical ads, from AI-generated presenters to product demos assembled from automatically captured app screens.
  • Reporting Data engineering Pulls raw exports from ad, affiliate and analytics platforms, joins them only on declared keys, and computes every metric from fixed definitions.
  • Performance Analytics Analysis and recommendations Flags anomalies and outliers, applies a layered rule base from universal to client level, and ranks actions by expected impact. Rules I confirm join the base.

Stack

Agents
Claude Code subagents and skills, plus shared files for my views, voice and brand
Creative
Language, image and video models combined per asset; Remotion and ffmpeg for motion
Automation
Node.js and Playwright on persistent browser sessions; Apify for public data
Analytics
Python, layered metric definitions, declared join keys, a rule base that learns from feedback
Knowledge
Brand, voice and views files under version control; logs and playbooks that only grow

Human control points

  • The strategy, the brand and the final say on what gets published stay with me
  • I set how much autonomy each agent gets, from approving every action to fully automatic
  • A rate limit, a challenge or a logged-out session stops everything. No agent ever types a password.
  • Every outbound action is paced and logged
  • If a fact is unknown it stays blank. Guessing isn't allowed, and unclear cases wait for review.
Workflow

How the work moves between them

No agent skips a step. Each one gets a checked input and passes a checked output along.

  1. 01

    Research

    Agent: Researcher

    Everything starts with evidence rather than trends: recent research, public conversations at scale, and the competitors' own websites and pricing.

  2. 02

    Ideate and fact-check

    Agent: Strategist

    Findings get turned into angles and hooks. Then a separate agent, whose only job is to be suspicious, checks every claim against its source. If it can't find the source, the idea stops there.

  3. 03

    Create

    Agent: Media operator

    No single model does everything well, so every asset combines several models and tools: one writes, others generate images and video, and code assembles the result to the brand rules. Client material works the same way. Every client deck is personalized to that client's needs and profile.

  4. 04

    Source and engage

    Agent: Orchestrator

    Social media, influencers and affiliates are managed automatically, inside the brand's voice and the limits I set, with every action logged.

  5. 05

    Measure and learn

    Agent: Analyst

    Raw platform exports are normalized against layered metric definitions (universal, industry, client) and joined only on declared keys, never inferred. Derived metrics such as cost per acquisition (CPA) and return on ad spend (ROAS) follow fixed formulas. Before any interpretation, the analyst flags period-over-period swings above 20% and outliers beyond three standard deviations. A layered rule base then turns triggered patterns into findings, ranks the recommended actions by expected impact, and absorbs every rule I confirm into the next cycle.

Project 02

VC Catalog

The venture-capital market, in one place you can actually search.

VC Catalog is a research platform for the venture-capital market. It brings VC firms, their funds, the companies they back and the partners behind them into one connected database, built entirely from public sources.

Visitors can search and filter more than 13,000 firms, around 145,000 companies and about 61,000 partners, read a profile for each, explore market reports, and use Connect to find the shortest chain of investments and people between them and someone they want to meet.

  • 13,344 VC firms
  • ~145K companies
  • ~61K partners
  • ~215K investment links
  • Warm-intro routes

Technical deep dive · VC Catalog

One connected map of the venture market.

VC Catalog runs on two layers. A data pipeline collects firms, funds, companies and people from public sources, resolves them into single entities and links them into a graph of investments and team affiliations. A web platform on top turns that graph into search, profiles, market reports and Connect, which finds the shortest path between two people or organizations. Every change to the data ships as a versioned, checked release.

Pipeline

  1. 01 Discover Firms come from public databases and SEC Form D and Form ADV filings (US securities and adviser registrations). A Form D fund size is read as the amount sold, not the target.
  2. 02 Crawl Firm and startup websites are read for portfolios, teams and theses. The crawler looks for a site's own structured data first and never bypasses a bot check.
  3. 03 Link and dedupe Records join only on hard identifiers, like a domain or an SEC filer number. A matching name is never enough, and every rejected merge is remembered.
  4. 04 Classify Claude agents label sectors and business models offline. A local classifier handles volume where it reaches about 90% precision. Every label keeps its provenance.
  5. 05 Release Each release must pass 43 health checks. An import that sees an unexpected schema refuses to run, and no sync may delete more than 5% of a table.
  6. 06 Serve The release loads into PostgreSQL, the search and graph indexes rebuild, and the site picks it up.

Stack

Web
Next.js and React, the only public layer; ECharts for charts, Cytoscape for the graph
API
A private FastAPI service behind the web app
Data
PostgreSQL with a durable job worker, loaded from versioned SQLite releases
Search
PostgreSQL trigram index, faceted filters and a ranking formula
AI
Offline only. Claude agents and a local classifier; no model runs in production
Hosting
Google Cloud in the EU behind Cloudflare, with nightly backups and uptime checks from several regions
Sources
Public databases, SEC Form D and Form ADV, company websites

Human control points

  • Matches the system isn't sure about wait for review, and my merge decisions stick across reruns
  • I spot-check every batch of AI labels before it gets written
  • New features stay switched off in production until I approve them
  • Anyone listed can ask for a correction or for their profile to be removed
Project 03

AI Software Factory

From a one-line idea all the way to production.

The factory came out of a fairly humbling experiment. In an earlier version I had a team of AI agents build a product together, with roles, reviews and hand-offs. After two days they had one partial feature. A single strong agent, given the same job, built the whole product in about four hours.

That changed how I design these systems. The line starts with my Product Design Studio, an agent I built that turns a one-line idea into a full product package: requirements, UX, UI, architecture, security and acceptance tests. The factory takes it from there. It plans the work, builds it with coding agents in sandboxed containers, tests it, deploys previews and turns my feedback into new tasks, while plain Python code makes every decision about money, stopping and releasing. I talk to it through Slack, and it's built to ask me only when something is truly blocked or risky.

  • Starts from a one-line idea
  • Product Design Studio agent
  • 2 to 6 coding agents at once
  • Sandboxed containers
  • Journeys tested on 3 devices
  • One-time codes for risky approvals

Technical deep dive · AI Software Factory

Gates, not committees.

Models write the code. Everything else is decided by a deterministic controller: which model runs, how much it may spend, when a task gives up, and what is allowed to merge. An earlier version ran for eight days and about 700 agent runs before I stopped it, so nothing here retries forever. Every loop has a counter and an exit.

Pipeline

  1. 01 Intake The controller checks the Studio package first. Every file hash must match the manifest and every artifact must be approved. Then it sets up the repo and tasks, no model involved.
  2. 02 Plan Four planning passes on the strongest model, each followed by a code check for coverage and cycles. The first milestone is always a walking skeleton: the thinnest version that works end to end.
  3. 03 Build legs Tasks run as headless Claude Code sessions, capped in turns and budget. A router hands each one a short-lived key for a single model.
  4. 04 Integrate One candidate per project at a time: tests on the merge tree, a test canary, a secret scan, a review if the task is critical. Then the controller merges.
  5. 05 Milestone preview Build once, deploy a preview, replay every journey so far on three devices with accessibility checks, then tell me what I can try.
  6. 06 Release Scanners, red team, judge. The exact tested image deploys without traffic, is smoke-tested and approved, then goes live, with automatic rollback.

The agents

  • Product Design Studio Idea to product package Turns a one-line idea into requirements, UX, UI, architecture, security and runnable acceptance tests, stage by stage.
  • Planner Architecture and task graph Runs four planning passes and then reviews its own plan, on the strongest reasoning model available.
  • Developers Pool of 2 to 6 Cost-efficient coding models, chosen per run and never swapped halfway. How many run at once depends on how much memory the machine has free.
  • Diagnostician Repair after repeat failure Steps in when a task fails the same way twice. It gets one repair attempt. After that the task stops and waits for me.
  • Reviewer Specialist by exception Only looks at tasks marked as critical. Reviewing everything is how the committee problem creeps back in.
  • Slice Verifier Milestone demo check Clicks through the demo steps on every preview, roughly the way a new user would.
  • Red Team and Judge Release gate One tries to break the release candidate. The other, separately, decides go or no-go.
  • Health Monitor Controller, not a model Watches 16 warning signs and can pause the whole factory by itself. It can't unpause. That part is mine.

Stack

Controller
Python 3.12 daemon, CLI and Slack bridge, about 15,000 lines
Models
Claude for design, planning, review and release decisions; a routed pool of coding models, each run on its own budgeted key
Harness
Claude Code, version pinned, run headless
Isolation
Docker, a private network per project, an allow-listed egress proxy and a scope guard on every write
Quality
Playwright on three devices, axe, gitleaks, Semgrep, Trivy, OSV-Scanner, Nuclei, ZAP
State
A SQLite ledger per project; an intent is written before every external action, so a crash can be recovered
Delivery
GitHub, containerized deploys with Postgres, Sentry; Slack and ClickUp for running it

Human control points

  • I review every stage of the product package before the factory accepts it
  • Risky approvals, like a release, a destructive migration or deleting infrastructure, need a one-time code from an authenticator app (TOTP) in Slack
  • The first two production releases of every product need my sign-off
  • A blocked task waits for me to unblock it, drop it or answer its question, all from Slack
  • The health monitor can pause the factory. Only I can resume it.
Workflow

One line, with a gate between every stage

Each stage has an owner and an output, and the next one doesn't start until a check passes.

  1. 01

    Design Studio

    Agent: Strategist

    It starts with a single line describing an idea. The Product Design Studio, an agent I built, turns it into a full product package: requirements, UX, UI, architecture, security and acceptance tests that can actually run. I review each stage. This website was specified that way, and its package is the first real product the factory is building.

  2. 02

    Plan

    Agent: Orchestrator

    Planning takes four passes: architecture, milestones, tasks, and a self-review. After each pass, plain code checks the plan and refuses it if something is missing, circular, or too slow to reach the first thing a user can do.

  3. 03

    Build

    Agent: Developer

    Two to six coding agents work at once. Each gets its own copy of the repository inside a sandbox with no real passwords in it, so a confused agent can only damage its own copy.

  4. 04

    Integrate and test

    Agent: Tester

    Before anything merges, it's tested on the exact code that will land. Every required test also has to fail on the old code. A test that passes either way isn't testing anything, and coding agents write those more often than you'd think.

  5. 05

    Preview and release

    Agent: Deployment engineer

    Each milestone ends with a live preview and every user journey replayed on three devices. A release adds security scanners, a red team and an independent judge before the tested image goes out.

  6. 06

    Feedback

    Agent: Analyst

    My Slack messages and production errors become feedback. Each one is triaged, planned and checked on a preview, the same way the original work was.

Chapter 01

Where the systems thinking came from

Direct contact

Start a conversation.

Name, email and a message. It comes straight to my inbox and I read every one.

10 to 5,000 characters.

Your details are used only to reply to you. They are never added to analytics or a mailing list. Privacy

Message received. Thanks, I will reply to the email address you gave me.

Close