
Over the last decade, compute, networking, and application delivery were transformed into automated code pipelines. The operational database didn't come along for the ride - it remained administrative: ticket-driven, manually guarded, and bound to legacy licensing.
Today, two unstoppable forces collide with this un-engineered layer: compounding cloud costs and the mandate for real-time AI execution. Enterprise AI initiatives stall the moment they hit an operational data estate still moving at "ticket speed." This series treats database infrastructure not as back-office maintenance, but as a P&L driver, a risk surface, and an executive priority for enterprise velocity. We put both forces in front of Jeff Carter, Tessell's Chief Strategy Officer, former VP of AWS Databases, and former CISO of Amazon's businesses outside of AWS.
This is Part 1 of a three-part series, in his own words, and it accompanies our video conversation with Jeff. This part covers the cloud performance trap, why decoupling compute from storage changes the economics, and why an administered database layer becomes the rate-limiter for real-time AI.
The question: When enterprises migrate heavy Oracle or SQL Server workloads to standard cloud infrastructure, they hit what's known as the cloud performance trap - where the only way to get more storage throughput (IOPS) is to scale up to massive, expensive compute instances. How does this forced coupling of compute and storage blow up core-based software licensing costs for Oracle and SQL Server, and why do traditional cloud migrations fail on ROI because of it?
I led the effort at Amazon.com to move our data infrastructure off on-premise hardware and into the cloud. I've reviewed thousands of these migrations since - first at Amazon, then at AWS leading databases and migrations for other companies, and now at Tessell. My experience across all three has been consistent: the cloud is less expensive. That holds at the macro level, across an entire data estate, and at the micro level, for each individual system.
But when people first move to the cloud, there are two patterns that can make that first implementation look more expensive than what they had before. I've seen both enough times to name them and both have a fix.
Systems configured for IO far beyond their worst-case peak. On-premise, that over-provisioning was effectively free - the hardware was already bought. In the cloud, you pay for exactly what you provision, so the habit shows up as a line item. The fix is measurement and right-sizing, and it's relatively easy to overcome.
This one shows up in databases that are relatively light on CPU but push hard on IO, or that are steadily growing their IO bandwidth needs over time. Cloud vendors offer a range of storage and instance options, but for a certain class of database - usually correlated with high-workload production systems pushing past your current instance's IO ceiling means moving to the next instance size up. That gets you more IO. It also gets you more CPUs, whether you need them or not.
Here's where it gets expensive. Larger instances cost more per core - that's real, but it's not the big number. The big number is what happens to your database licensing. Commercial database vendors, Oracle and SQL Server included, license by core. Jump from 16 cores to 32, or 32 to 64, and you're not looking at a proportional increase in your bill - you're looking at something that can double or triple it. Add a high-availability configuration on top, which duplicates the core count again, and it gets harder still to make the cost case for that individual system.

I call this the cloud performance trap: your storage throughput is tied to your instance size, and your instance size is tied to your license cost. Ask for more IO, and the bill you actually feel is the licensing bill.
The question: Tessell leverages patented NVMe-backed architecture to deliver up to 2 million IOPS without IOPS metering or compute penalties. Drawing on your experience managing multi-billion dollar AWS database P&Ls, how fundamentally does decoupling storage throughput from compute change the balance sheet for large database estates?
Tessell offers the same cloud-native block storage options as everyone else - but for IO-constrained databases, we also offer a second path.
The first part of it is instance types with large amounts of local NVMe storage. NVMe is dramatically faster than network-attached storage. The catch is what happens when a node fails: on a plain NVMe instance, that data is likely gone with it not something you can accept for a production database.
Tessell pairs the NVMe instance type with the commercial database engine you already run, plus patented software that closes that safety gap. Primary writes go straight to NVMe for the performance, while our software independently handles synchronization and logging across data centers, or across clouds. That gives you the high availability a production system requires, without giving up the performance the hardware is capable of.

We've published benchmark results of up to 2 million IOPS on a single instance.
What that number actually means: Tessell can bend, and in a lot of cases outright break, the link between how much IO your database needs and how large an instance you have to buy to get it.
That's the direct answer to the trap above. By changing the relationship between IOPS and instance size, Tessell lets you get more IO per instance - which means you don't need to keep growing instance size just to keep up with IO demand. That saves you on the compute side, and it saves you on the number that actually hurts: database licensing.
The question: Enterprise leaders tell us their data lakes were engineered years ago for analytical AI, but their operational, transactional databases are still run manually. When an enterprise is building real-time AI agents or RAG pipelines that need live production data from Oracle or SQL Server, why does an administered database layer become the primary rate-limiter for AI execution?
We hear the same thing from enterprise leaders constantly: their data lakes were engineered years ago for analytical AI, and work well for it. Their operational, transactional databases are still run largely by hand. That gap matters now in a way it didn't two years ago, because real-time AI agents and RAG pipelines need live production data, not a nightly export.
I've spent a lot of time on both sides of the argument this creates: DBAs holding the line on a production system, and application developers who feel like that line is slowing them down. I don't think it's a personality conflict. Most of the time, a DBA holding firm is reflecting something real - hundreds or thousands of downstream teams depending on that system behaving the way it always has.
I saw this up close at Amazon.com, where I was responsible for the operational systems running most of the site: Identity, the customer record; Orders, the entire history of every order ever placed; Payments, every financial instrument a customer has ever given us and every charge ever run against it. Systems like that have one overriding goal - keep the business running, at the response times customers expect. Nothing outranks that, and no new use case gets to degrade it.
Here's a simple example of why an administered database layer becomes the rate-limiter the moment you connect it to an agent.
You buy a pair of shoes. They arrive a size too small. You want to exchange them for a size up, and you're going to let a new AI agent on the retailer's site walk you through it. One step in that flow matters here: is the larger size actually in stock? The only place with accurate, current inventory data is the operational system itself, and the agent needs a real-time answer, not a snapshot from last night.
A well-behaved agent does exactly one lookup, and the inventory system barely notices. But an agent doesn't have to be well-behaved to be useful. If the exact size isn't available, a genuinely helpful agent might start searching the entire inventory for something similar. Great for the customer. Potentially disastrous for the operational database, which was never built to absorb that kind of unpredictable, exploratory query load.

That's the fundamental tension: a system built for predictable workloads, meeting a use case that is inherently not predictable. So how do you solve it? I've seen four approaches, and only one of them actually works at scale.
Run the agent's queries directly against the OLTP system, and hope for the best. Don't. This is the fastest way I know to break the fundamentals of your core business processes.
Clone the database on demand. A clone can take minutes, hours, or days to build, and by the time it's ready, the data it's built from is already stale.
Dump the data into your analytics environment on a schedule. Nightly, hourly, whatever cadence, the same problem applies: by the time the agent can see it, it's stale.
Build a read replica with continuous change data capture (CDC). A persistent clone that stays current in real time. Done in the operational environment rather than the analytic one, it can be transactionally consistent with the primary, not just eventually consistent.

The fourth option wins, and it's not close. It keeps all the RAG and GenAI query load off the primary instance entirely, so it can never degrade the workflows that actually run the business, and because it's a separate instance, you can size and adjust it independently, based purely on what the AI workload needs.
Tessell makes this option easy to run in practice. We provision the replica and keep it current using the database engine's own native replication - Data Guard, Always On, Group Replication, whatever the engine provides and place it in the same availability zone, a different one, a different region, or a different cloud entirely. Every replica we create is logged, which matters the moment your audit team asks where copies of production data actually live.
One clarification: Tessell's own documentation sometimes uses "CDC" to describe feeding operational data into a data lake, a different, also valid, use case. Here, I mean it specifically as syncing a primary operational system to a live replica.
None of this requires replacing your database engine or rearchitecting your applications. It requires an honest look at three things: whether your IO needs are quietly forcing you into bigger, more heavily licensed instances than the workload demands; whether decoupling storage throughput from compute changes your cost structure more than you've modeled; and whether your AI agents have a governed, real-time path to operational data, or whether they're one unpredictable query away from taking down the system that runs your business.
This is Part 1 of a three-part series. Part 2 turns to database consolidation, the CISO's view of data governance at AI speed, and how AI is being used inside the database platform itself. Part 3 covers multi-cloud strategy and the board-level blueprint for next-generation data infrastructure.
If you want to know where your own Oracle and SQL Server estate stands today - talk to us.