The Nvidia Central Bank: How Revenue-First Founders Break Free from AI's Hardware Monopoly
· Brandon Crenshaw
I was on a call the other week, just a quick sync with a founder I’m advising through the studio. They’re building an interesting AI-driven analytics product, pretty niche, and they’ve got some early traction. But the conversation quickly veered into a familiar territory: compute costs. Specifically, the line item for GPU access. They were crunching numbers, trying to figure out how to scale without their gross margins getting eaten alive by rented H100s.
It got me thinking, as it always does, about this strange financial reality we’ve built in AI. It feels a lot like we’re operating under a single, dominant central bank – only this one isn't issuing currency, it's minting the most powerful chips. And if you want to build anything significant in AI, particularly anything that touches large models, you’re often lining up at Nvidia’s window.
For a certain kind of AI startup, especially the ones flush with venture capital, this isn't seen as a problem. It’s a cost of doing business, a badge of honor to be running on the latest, greatest silicon. But for those of us who came up building products with revenue *before* funding, who understand that every dollar spent needs to justify itself immediately, this dependency feels like a chokehold. It’s an external, often unpredictable, factor directly impacting your unit economics. And it's a hell of a thing to build a sustainable business on.
The Mirage of Peak Performance (and the VC Echo Chamber)
I remember sitting in various WeWorks back in the day, pitching VCs on ideas. The conversation always, eventually, drifted to scale. To "going big." And in AI, "going big" has, for a long time, meant throwing the most powerful hardware at the problem. There's this almost unspoken assumption that if you're not on the bleeding edge of GPU technology, you're not serious. That if you're not talking about training models on thousands of H100s, you're not thinking big enough for venture scale.
This is the siren song that leads many AI startups down a path of unchecked expenditure. They burn through cash chasing benchmark numbers, believing that the fastest chip automatically translates into the best product or the most defensible moat. It’s a vanity metric, really, dressed up as technological superiority. The problem is, while the VCs are excited by the potential, it's the founders who are left holding the bill for an infrastructure that might be overkill for 80% of their actual product’s needs. This pervasive GPU dependency means a significant chunk of a startup's runway can evaporate into compute costs, often before they've truly validated a paying customer base.
When Pedro and I started JP Trading Capital, or later when I built other products before founding my studio, the mantra was always revenue-first. Every line of code, every feature, every piece of infrastructure had to contribute to a paying customer or a clear path to one. There was no room to fantasize about future capabilities that might be enabled by a chip that cost more than our entire quarterly operating budget. This fundamental difference in philosophy — revenue-first versus a burn-to-grow model — dictates entirely different approaches to something as critical as AI hardware and compute costs.
Radical Frugality: Building for Profit, Not Petaflops
So, how do you break free from being beholden to this "Nvidia Central Bank" when you’re building an AI-native product? It starts with a radical commitment to frugality, not as a temporary measure, but as a core architectural principle. It's about asking, "What's the absolute minimum compute power we need to deliver a valuable, revenue-generating outcome for our customer?" rather than "What's the most powerful chip we can get our hands on?"
This often means diving deep into **software optimization**. Forget just throwing more hardware at it. Can you quantize your models? Explore techniques like LoRA for fine-tuning instead of full model retraining. What about pruning or knowledge distillation? Can you refactor your inference pipelines to be more efficient? There are entire fields dedicated to making models run faster on less hardware. We use tools like Claude Code to iterate quickly on these optimizations, constantly looking for ways to squeeze more performance out of fewer resources. It’s a different kind of engineering challenge, one that prioritizes efficiency over raw power.
And sometimes, it’s about looking at **alternative AI compute**. Nvidia is dominant, no doubt. But are there specific tasks that can run effectively on AMD or even Intel GPUs? Can you offload certain parts of your pipeline to cheaper CPUs in the cloud, or even edge devices, if the latency requirements allow? It’s not about abandoning Nvidia entirely, which for many complex tasks is simply not feasible today. It's about strategically choosing where and when to use their hardware, and actively seeking out scenarios where you don't have to.
I’ll admit, it’s a tightrope walk. There are definitely times when the latest generation of GPUs provides a breakthrough that’s genuinely necessary for a specific type of research or an incredibly demanding real-time application. But for the vast majority of AI products aiming for revenue, an older generation GPU – say, a V100, or even a P100 for certain inference tasks – can be perfectly sufficient if your software stack is incredibly optimized. The cost difference between a current-gen and a previous-gen GPU can be astronomical, and that saving goes directly into your margins. It’s a question of product market fit for your compute, not just for your features.
The Studio Model: Shipping Revenue, Not Just Code
This philosophy of hardware independence and AI frugality is baked into how we operate at my creative studio. We partner with founders to build AI-native, revenue-first products. This means we're not just writing code; we're meticulously designing product architecture and business models simultaneously. The compute strategy isn't an afterthought, it's a foundational pillar.
My background in AI product engineering, alongside business development and sales strategy, has taught me that the fastest path to revenue often involves constraints, not unlimited resources. When you have a tight budget for compute, you're forced to be more creative. You're forced to ask harder questions about the necessity of every feature, every model parameter, every piece of data. This leads to leaner, more focused products that solve real problems efficiently.
We’re all about shipping fast. We build with React, Node.js, Vercel, Supabase, Airtable – a stack that allows for rapid iteration and deployment. But that speed isn't just about front-end development. It extends to the backend AI inference. If your models are bloated and your compute costs are high, that slows down everything, from development cycles to the actual user experience. Optimizing your AI infrastructure directly contributes to shipping faster and more cost-effectively, which in turn accelerates your path to revenue.
I’ve built products that had revenue before they ever saw a dollar of outside funding. This wasn't by accident; it was by design. It meant making pragmatic choices about technology, including where and how our AI models would run, and critically, how much it would cost. It meant pivoting from a fundraising-first mindset to a revenue-first one, focusing on paying customers rather than chasing the next valuation round. This approach isn't glamorous by traditional startup metrics, but it builds real, sustainable businesses.
Beyond the Central Bank: A Future of Distributed Intelligence
Ultimately, the goal isn't just to save a few bucks on GPUs. It's about building resilience. It’s about not having your entire business dependent on the pricing and availability of a single vendor’s hardware, however excellent that hardware might be. The current "Nvidia Central Bank" model is a single point of failure for a significant portion of the AI industry.
Looking ahead, I wonder how long this can truly last. The long-term viability of the AI ecosystem might depend on a more diversified and distributed compute infrastructure. We see glimmers of this in alternative hardware development, in open-source model optimization, and in the growing emphasis on efficiency. It’s a slow shift, and I'm not entirely certain when or if true hardware independence will become the norm. But the founders who are thinking about this today, who are actively building strategies to mitigate their GPU dependency, are positioning themselves for a more robust and sustainable future. They're not just building AI products; they're building businesses that can weather the inevitable shifts in the technology and economic landscape. And that, to me, is the real game.