concepts · youtube · 9 min
Agent Sandboxes for Scaling Software Factories
IndyDevDan · Aug 24, 2026
Engineers… Your Software Factory NEEDS Agent Sandboxes to SCALE
Channel: IndyDevDan
Published: 2026-08-10
URL: https://www.youtube.com/watch?v=SEI_qIW4o2c
Engineers... if you're IN the loop, you ARE the bottleneck. And giving your agents a tiny corner of YOUR computer is why your software factory can't scale past one.
Video References
- Factory In A Box Codebase: https://github.com/disler/inkwell-agent-sandboxes-and-software-factory
- Super Simple Software Factory Video: https://youtu.be/haUfb1ievTE
- FORGET Loop Engineering Video: https://youtu.be/VQy50fuxI34
- Agent Sandboxes on Exe.dev: https://exe.dev/
- OpenRouter provisioning keys API
Transcript
What's up engineers? Indy Devdan here. It's time to focus. This is going to be a big video and a lot of engineers are going to miss the leverage I'm going to share in this one. I don't want you to be one of them. Your software factory can unlock an unprecedented level of engineering results. But there's a massive roadblock every agentic engineer runs into at some point. A roadblock you might have. Where should your software factory run? Where do your agents plus code actually run?
Most engineers allocate a tiny corner of their own computer for their agents, or they lean too heavily on CI/CD or some container. The best engineers aren't doing this. Engineers at big AI labs running the most insane workloads you can imagine do something else because they've realized one simple fact. The more autonomous and secure your agents are, the more they can do for you riskfree. On this channel, across hundreds of agentic engineering videos, you've heard me say scale your compute to scale your impact. Compute applies equally to GPU compute and CPU compute. Old school bare metal. An agent sandbox is the best space to place your agents.
Why Agent Sandboxes?
Why is that? Why can't we just use a container and call it solved? It's because agent sandboxes give you three key advantages: true isolation, insane scale, and agency. Here's the raw reality of the state of agentic engineering. If you are inside the loop, you are the bottleneck. In this video, I'm going to break down exactly how you can step outside the loop by putting your software factory inside an agent sandbox.
The Tech Stack
Here's the tech stack we're going to be using to put our factory inside a box. Don't focus on the tool. Focus on the purpose and choose the best tool you need. We'll use Claude and Fable as our top-level orchestrator inside of our software factory. We're going to use Pi as our agent SDK. Don't just limit yourself to one agentic coding tool. The tools you use directly limit what you believe is possible.
The Best of N Pattern
Let's go ahead and kick off the skill that activates our agents understanding. We're going to open up brand new VMs and run the best of N pattern. We're putting our software factories into an agent sandbox.
We're running five different agent configurations: Default, Frontier, Deepseek, Open Weights, and Top Speed, each with its own sandbox. And they're all going to solve the same problem. But we're not just launching agents. We're running an AI developer workflow. We're running the full software developer life cycle against our simple application.
Our orchestrator is kicking off not just agents inside sandboxes. Our agent is building our software factory and placing the entire factory inside our agent sandbox. And this unlocks a lot of really powerful things. When you use agent sandboxes, you're getting isolation, you're getting scale, and you're getting autonomy. Every single agent owns the computer. They're not just running on a little corner of yours interfering with your work, you interfering with their work. They have everything they need to own the outcome.
Agent Sandbox Tools
We are using exe.dev as my primary sandbox tool. One of the best, simplest ways to get started with agent sandboxes. Their tagline keeps it nice and simple for you: Computers for developers and agents. Durable sandboxes, fast, secure, and sharable. The agent sandbox is a developer device for your agents.
Let's really atomize everything. Let's say exactly what things are. There's a lot of hype and noise in the AI industry. I try to dehype things as much as possible. We're not just throwing an expensive Claude Opus at this. We are understanding what model needs to run where with what code surrounding it. That's the software factory. That's the next step.
Model Stack Architecture
Orchestrating other agents, orchestrating sandboxes does require state-of-the-art intelligence. I always track these models inside of my model stack. I delineate between three tiers of models:
- State-of-the-art: The newest, most powerful models
- Workhorse: Powerful models that aren't quite state-of-the-art
- Lightweight: Basically the floor, defined by their ability to be run on previous generation GPU nodes or directly on your Mac device if you have the unified memory for it
The brand new Deep Seek V4 Flash is absurd for its price. We're going to see a model configuration that I call Deepest Seek, which is just V4 Flash models. We're going to see how they compare to a bunch of other model plus harness configurations inside of our software factory.
If you're still fixated on models, you are behind. That's not the name of the game anymore. We're entering an age of abundance of compute. And now it's about focusing on what model plus what code do you need to do the job. And the software factory is that system.
Observing the System
We have five unique software factories running in their own agent sandboxes. We have discrete URLs for every sandbox: our default config, our frontier config, our deepest seek config, our open weights configuration, and our top speed. Each one inside an agent sandbox.
I'm not giving my agents a corner of my computer. I'm giving them an entire computer to do all the work they need to with isolation, with scale, and autonomy.
Beyond Single Agents
It's not just about a single agent anymore. One agent is not enough, just like one prompt was not enough. We then started running sub agents. We then started chaining agents. At some level you realize that agentic engineering is software engineering. We have a new primitive which is the agent autonomous software and we need to put them inside of our existing developer workflows that we ourselves would run. Hence the term AI developer workflow. We're adding AI to the work you and I used to do as engineers every single day.
When your agents have their own developer device just like you do, guess what they can do? They can perform and act and even outperform you. That's the whole goal here. A lot of engineers are afraid of these great models progressing further and further. I am not. This is a new tool in your toolbox. Don't run away from the fire. Run into the fire. These agents can do incredible work on your behalf for you, for your business, for your career, for your users, but you have to leverage it.
The Bottleneck Problem
Inside of every single sandbox, we have placed a set of agent configurations and we built an AI developer workflow that walks through the software developer life cycle and our agents are shipping work on our behalf. Here's a key word: without us. The key idea that I really want to communicate here is that if you are in the loop, you are always the bottleneck.
Now, there are exceptions to this. If you're doing hands-on agentic coding work that requires you, that requires your expertise—if you are building the system that builds the system, then of course get in the loop, do the work, go hands-on, prompt back and forth, babysit your agent. You need to do that at some level. But there's a lot of work you might be doing, and I know for a fact teams are doing and god forbid we talk about enterprises—there's a ton of work all these sets of engineers are doing that they don't need to be. Why not? Because they can have a team of agents do it for them. And not just a team of agents, a team of agents plus code plus engineers.
The Best of N in Practice
Say you have a great idea. Say you have multiple directions you want to go or even one single concrete direction. In our example here, we just pass in one singular prompt. We want a quiet room as our design principle for this application. The current version is cluttered. We want it to be simple and clear. So this is the prompt we pass into every agent sandbox and then into the software factory inside each agent sandbox.
I'm not just running one agent. I'm running a pipeline of agents plus code. We can see the request here: redesign the application. But then we can look at the plan workflow. We can see all the thinking, all the agent calls, the gate passes. We have deterministic gate checks on our nondeterministic systems.
In our super simple software factory, we are not just doing lightweight prompts. We are fully in control of the system. What do I mean by that? We are selecting a specific coding agent. Of course, the model, of course, the thinking level, of course, the tools. These tools came from our custom agent harness. We are harness engineering inside of our software factory. If you don't have custom harnesses, if you can't customize and control, you don't have a software factory. The software factory is your full control over the agents and your code.
The Next Leap
I'm going to start compounding some really heavy-hitting ideas here. We're going to lose a lot of engineers. A lot of engineers are going to look at this and just think, "I don't need this. I don't need to build this." But engineers that can focus right now: this is the time to focus because the next leap is available. It's your software factory. And then you can really leverage and scale your software factory by putting your factory in a box.
At a high level architecturally, it's relatively simple. The implementation details do make things a little more complex. If you're constantly going back and forth on what model you're using, you're missing the point. You need a model stack, not a single model. Combine compute. Don't select compute.
(capture appears truncated)