Megarepos, not Monorepos
Monorepos consolidate source code. A megarepo goes further by keeping the operating context of a business or domain in one place: code, docs, infrastructure, agents, environments, and contracts.
Most arguments for monorepos focus on source code. Related packages can share dependencies, refactors can cross boundaries, and one build graph can cover the tree.
Coding agents have pushed me to look past that boundary. Writing code is often easier than assembling the context around it. An agent needs to know what the business does, how the system should behave, which constraints are real, how a change reaches production, and how to prove that the result works.
I've started calling this kind of repository a megarepo, and I rely on the pattern heavily.
A megarepo may still be a monorepo mechanically. The name describes its operating scope. It holds the context for a business, business unit, product line, or another domain that changes as one system.
Match the repository to the domain
The repository boundary should sit where context, ownership, and change come together. It needs enough material for someone to understand, change, run, and audit the domain without chasing half the story through other systems.
That usually includes more than application code:
- services, frontends, workers, libraries, and command-line tools;
- architecture, business rules, decisions, and operator procedures;
- database schemas and ordered migrations;
- infrastructure, deployment definitions, and release workflows;
- API request collections and executable examples;
- brand assets, design tokens, and content conventions;
- build, test, lint, and repository-wide validation commands;
- agent instructions and repo-scoped skills;
- local development, test, preview, and production environment definitions; and
- references to important systems that must remain external.
Git does not need to hold every byte. Secrets, production values, generated data, and live state should stay in systems built for them. The repository should say where those things live, why they are external, and which contracts connect them to the rest of the domain. It is still the source of truth for context: an operator or agent can discover what exists, who owns it, how it changes, and how to verify it.
A megarepo is intentionally complete enough to operate its domain. That is what distinguishes it from a repository that is merely large.
Greenway Vault as an example
Greenway Vault is a warehouse management system for a collectibles business. It tracks physical inventory through purchasing, intake, storage, sales, and shipment. The entire business is implemented in a single megarepo, vault-os, which is arranged around the business's operations rather than around a language or framework.
Here is a simplified view:
vault-os/
├── apps/
│ ├── go/ # Single binary w/ sub-commands for API, Temporal Worker, etc
│ ├── frontend/ # warehouse operator application
│ ├── frontend-landing/ # public site
│ ├── python/ # catalog and provider data pipelines
│ └── etc...
├── docs/okf/
│ ├── intake/ # receiving and upload workflows
│ ├── wms/ # inventory, bins, labels, reconciliation
│ ├── catalog/ # product data, media, and pricing
│ ├── sales/ # channels, orders, and settlement
│ ├── reporting/ # metric definitions and profitability
│ ├── infrastructure/ # deployment and state runbooks
│ └── economics/ # throughput and unit economics
├── migrations/ # ordered database changes
├── bruno/ # executable API request collection
├── infra/ # OpenTofu-managed AWS infrastructure
├── brand/ # canonical identity and design assets
├── data/ # pinned or local data inputs
├── AGENTS.md # repository rules for coding agents
├── Makefile # shared build and verification contract
├── Dockerfile
└── docker-compose.yml # local backing services
The connections between those directories are what make the layout useful. An inventory lifecycle appears in backend services, migrations, frontend actions, scanner flows, API requests, and operator docs. Catalog work can touch provider ingestion, matching, media, pricing, and reporting. Infrastructure and release instructions live close enough to the applications to change in the same review.
The repository also carries an Open Knowledge Format (OKF) bundle. I treat that documentation as part of the system, not an archive downstream of it. The bundle records business rules, system design, workflows, external setup requirements, architectural decisions, and operating procedures. When behavior changes, the relevant knowledge should change in the same unit of work.
I want a vault-os change to match the level at which the business experiences it.
Change the business capability together
Conventional repository boundaries can turn one business change into a string of coordination tasks. The API lands in one repository, the frontend follows in another, and the migration belongs somewhere else. Infrastructure waits on a separate pull request. The runbook lives in a wiki, the request example sits in someone's local client, and the preview works only if a teammate remembers the steps.
Every team can move quickly while the complete change still moves slowly.
A megarepo lets one pull request include the domain behavior, migration, interface, tests, request example, infrastructure binding, operator docs, and preview configuration. Reviewers can see how the pieces fit, CI can test them together, and the commit history keeps one record of the capability as it changed.
Keeping those files together does not excuse bad module boundaries. A healthy megarepo still needs small applications, explicit dependencies, domain logic separated from transports, and native commands for each component. Small modules can live inside a broad repository. Colocation also makes existing coupling easier to see.
I prefer simple, colocated systems inside a megarepo. When a product needs several backend processes in the same compiled language, I would rather compile one binary and package one image. That avoids maintaining separate lint, validation, build, and image configuration for every process. It reduces overhead for me, coding agents, CI, and anyone else working in the repository.
Agents change the calculus
People bridge missing context with memory, conversation, and institutional habit. Agents are brittle when that context stays implicit, especially when the missing pieces cannot be inferred from the repository.
Drop an agent into a code-only repository and it has to reconstruct the surrounding system. It may find an implementation but miss the business rule behind it. It may change an API without updating the request collection, generate infrastructure that conflicts with the release model, or pass a package test while leaving the operation broken.
A predictable top-level layout helps the agent find applications, docs, infrastructure, migrations, and tools. Maintained domain docs explain intent and invariants. Root commands and repository instructions turn local habits into interfaces the agent can use. Local and ephemeral environments make a proposed change observable before release. Focused tests and a repository-wide check show whether the work is complete.
The agent can inspect the business context, make a bounded change, update the affected surfaces, and run the same checks used by people and CI.
Agents are especially useful when work crosses disciplines. The same change can update a service and its runbook, pair a migration with rollback notes, update a design token and its consumers, or trace a business rule from documentation through a UI guard to transactional enforcement. That only works when the agent can find and act on the relevant context.
Repo-scoped skills can teach an agent how to audit its architecture, establish a preview, add a service, evaluate a migration, or complete a release. Because those skills are versioned beside the system, the workflows can change with the system instead of living in one-off prompts.
Make the front door boring
At the root, I want a small set of predictable commands for building, testing, and running the global validation gate. Component commands should stay available for fast iteration. Local services need explicit start and stop commands, and CI should call the same entry points people use. Production configuration belongs in the environment rather than in a Makefile.
The Megarepo OKF calls these operating contracts. I do not care which tools implement them as long as the repository answers ordinary questions without archaeology.
Someone arriving fresh should be able to learn how to run and validate a component, where an API lives and how to call it, which migration introduced a field, and what creates a production resource. The same repository should identify the business rule that blocks an action, the steps that require a person or an external account, and the files that must change together.
When the repository answers those questions directly, people ramp up faster and agents spend less time guessing. Automation is safer because it has a clearer contract.
Know where to stop
Different owners, permission boundaries, regulatory constraints, release lifecycles, or operating models are good reasons to keep a system in another repository. Megarepos are not a silver bullet, but in my experience they make bootstrapping and day-to-day development faster when the operating boundary fits.
More than a source tree
The name "megarepo" is a little provocative, but the size is incidental. I use the term for a repository that acts as the versioned map of a business domain: its software, knowledge, infrastructure, workflows, and change process.
Monorepos made it easier to work across codebases. The megarepo idea applies the same benefit to the rest of the work required to run a system.
For agentic development, complete operating context matters more to me than the amount of code in the repository.