Pebble is the engine. Now I'm building the factory around it.
A few days ago I wrote about Pebble, the agentic harness I built for my software work. It is tuned the way I want, and it writes most of my code.
Pebble now writes most of my code. That moved the bottleneck, and it's me.
This post starts a new experiment: gFactory, my attempt at an autonomous software factory. I am building it in the open, and this is where it begins.
Where my time goes
When I start a new project, my work falls into three phases.
Direction. I write the PRD, the tech design, and the builder specs. This takes most of my effort, and I am happy with that. Specs and direction are where my judgment matters.
Execution. Once the specs are ready, I feed them to the harness one at a time. A big project can mean 30 to 40 specs, and each takes an hour or more to run. That is 30-plus hours of runtime, and I am the scheduler: starting each one, checking that it worked, starting the next.
Verification. Reviewing, testing, and end-to-end testing. This is still mostly unsolved, my harness included.
The pain was not in any one phase. It was the second one, where the person who should be thinking about direction is acting as a job queue.

It became undeniable when I started designing a context-management tool and the design grew to more than 40 specs. The harness could handle roughly 90% of the coding. I still could not run the project comfortably, because I was the bottleneck between the specs and the harness.
Testing the idea outside my own head
"I am annoyed by this" and "this is a common bottleneck" are different claims, and only one justifies months of weekend work. So I talked to a handful of people on LinkedIn and offline. The sample is small, and these were conversations, not a survey. But the reaction was consistent. Nobody described algorithmic difficulty as their problem. What they wanted was something that takes the work and produces code, without them hand-feeding it. It was the same shape as my frustration, from people with different stacks.
What gFactory is
Multiple projects feed work packets into one shared queue. Every project moves through the same five stages: plan, break down, build, test and review, ship. Idle agents claim the next packet as soon as they are free, so several can be working at once. If review sends a packet back, it loops through refinement before it ships.
The goal is not to make any single project faster. It is to stop unrelated projects from blocking each other because they share one person.
The reframe: context and trust
My first mental model was orchestration: which agent, which model, how they hand off. But I already know how I want to split work across models. The top tier takes specs and design. The mid tier takes implementation and review. The light tier takes documentation. That routing is a config choice for me, and the vendor is close to interchangeable.
The hard part is everything that has to be true before an agent's output is trustworthy enough to merge:
- a shared queue that projects feed into, so I stop being the scheduler
- packets an agent can pick up cold, so I stop re-explaining the world
- a review layer that decides what is good enough to ship
So the question moved from "which agent?" to "how does context travel, and what earns trust?"
Two constraints, and the pilot
It has to run on subscriptions, legally, using the supported headless modes of the tools, because API-only would make long, many-packet projects expensive.
The review gate is non-negotiable. Without review, testing, and end-to-end checks, a factory only produces unverified code faster. In practice, that means a project's first pieces of work need my sign-off, and that loosens project by project as trust is earned.
The first pilot is gFactory itself. The first real job I give the pipeline is building the thing that will later run it. If it cannot build its own v1, it should not build anyone else's.
What I am still working out
Direction drift is the one I think about most. If the factory's output slides away from my vision, I might not notice until it is expensive to undo. I want a way to catch that early, and designing it is part of the work. Deciding which agent takes which packet, and knowing when a result is really ready, are also still being tested. And anything visual or interface-heavy is harder to verify automatically than backend logic, so I expect to give those areas more manual attention first.
Follow along
This is week one. What exists today is the thinking: the thesis, the constraints, and the pilot. Next up is turning it into a scope I can commit to and writing the first spec, with gFactory writing about itself.
Every Sunday I publish the build journal at gFactory.dev: what moved forward, what is still open, and what comes next. The Week 1 entry is live now. Week 2 lands on October 11. Subscribe to get it in your inbox, or follow along on X at @rupreetg.
One question I would love your help with: if you build with coding agents, what does your version of "feeding specs one by one" look like? Where do you end up acting as the scheduler?
Found this useful? I do 1:1 sessions on AI architecture and strategy. → Book a session
Found this useful? I do 1:1 sessions on AI architecture and strategy. → Book a session