Why Your Cloud Architecture Can Scale Users But Not Costs

If you’ve built software for a while, you’ve probably seen a similar architecture diagram: database for transactions/relational data, Redis for caching, Kafka for events, OpenSearch for full-text and keyword search, object storage as a data lake, and maybe a vector database for AI integration.

Example of common cloud architecture, showing distributed system components.

There’s nothing wrong with these choices: each technology is good at its job. For example, you’ll find countless articles, papers, and engineering blogs extolling the virtues of Apache Kafka for asynchronous, service-to-service messaging.

But how much is it going to cost to run and maintain all of this? Before on-boarding new component, ask yourself if the problem can be solved using existing platform capabilities. Simple is good. Boring is good.

We optimize whiteboard architecture instead of the business system

Imagine building an activity feed for a SaaS app: on the feed, users generate events, we store activities somewhere, and retrieves the right ones for each user. It also must scale as users and events grow.

During design discussions, the activity feed system might begin with a few core components:

  • A RDBMS for relational records like users, preferences, activities
  • A message broker like Kafka to decouple application generated events from the service that builds activity feed
  • A cache like Redis to store frequently accessed feed items
  • A search engine like OpenSearch to index activity records for fast search and filtering

How much does the system cost to operate?

Not just how much the infra/cloud licensing costs, but how much it costs the engineering organization to understand, test, maintain, monitor, secure, upgrade, and troubleshoot the system.

Every component in a software system comes with this baggage. It creates operational work, consumes organization attention, and distracts engineers from building competitive, differentiating product features. Even for managed services, these costs don’t disappear (though they may be somewhat reduced).

This is particularly relevant for software businesses: architecture decisions eventually determine COGS (Cost of Goods Sold, essentially the metric of how much it costs to maintain and sell the software).

However, COGS is usually absent from architecture discussions. Whiteboard conversations focus on throughput, latency, scalability, and load projections. Costs and operational overhead become much clearer later.

At the extreme, I’ve seen systems where adding a new user actually costs more than the revenue that user generates.

Complexity compounds over time

An initial architecture can be perfectly manageable. But successful products evolve, new features are brought in, and new components are added to the stack.

In isolation, each decision may be technically sound, but the system may still be expensive and optimize the wrong things. A small engineering team can effectively operate one database platform, but it’s significantly more difficult to build expertise in separate caching, event streaming, text search, and vector database systems.

At some point, infrastructure, operations, and maintenance can become the engineering group’s main source of work. A system that requires complex infrastructure, significantly more operational staff, and a wide collection of specialized technologies may not have scaled economically.

The database can be more than a database

To avoid the complexity of gluing together independent products, we can look toward modern, multi-model databases like Oracle AI Database or PostgreSQL.

Instead of stitching together independent products, the platform provides integrated capabilities:

  • relational data
  • JSON/document storage
  • graph
  • spatial
  • vector search
  • analytics
  • AI capabilities
  • transaction processing

Take our activity feed, for example: each separate component could be mapped to in-database features:

And so forth. Most, if not all capabilities can be operated on a single database these days. This doesn’t mean you have to build everything on the same database, but it can be a dramatic simplification.

A note on testing

As the number of moving parts grows, it’s more and more difficult to effectively test the system without a fully deployed stack. I’ve written a lot of tests for these kinds of systems, and full-blown integration tests are challenging to build.

Interestingly, an integrated data platform like a multi-model database is often easier to test.

When I write an integration test, I don’t need to spin up and glue together a complex stack. I can create one (1) database container, run my schema setup, and connect my applications using the provided connection string. This makes integration tests simple and reproducible, with startup times measured in seconds.

This even allows complete local testing with frameworks like Testcontainers and Oracle AI Database Free containers. I have dozens of examples doing exactly this with various (and combined) database features here.

References

Leave a Reply

Discover more from andersswanson.dev

Subscribe now to keep reading and get access to the full archive.

Continue reading