Cloud & Data Lakes
Infrastructure evolved. Data still comes first.
Cloud changed how quickly we can provision, scale and automate technology. It did not remove the need to understand the data, the workload or the problem we are trying to solve.
My Starting Point
Cloud conversations have a habit of starting with the problem and ending with a shopping list of services.
I prefer to work the other way round.
Once those questions are understood, technology choices become considerably less mysterious.
What Cloud Actually Changed
The basic ingredients are familiar: compute, storage, networking and software. What changed is how they are obtained, combined, scaled and governed.
Minutes rather than procurement cycles.
Capacity follows demand rather than forecasts.
Infrastructure becomes programmable.
Cost moves towards what you actually use.
What Cloud Didn't Fix
Moving poor data into better infrastructure still leaves you with poor data — only now it may be distributed, elastic and generating a monthly consumption bill.
- Meaning still has to be understood.
- Quality still has to be measured.
- Ownership still has to be clear.
- Security still has to be designed.
- Costs still have to be controlled.
- Outputs still have to be trusted.
The Data-First View
I see cloud through a data-first lens. The useful question is not simply where the data lives, but how it moves from source to something people can safely use.
The Practitioner Test
Before choosing the platform, I want answers to a few less glamorous questions:
- What are we actually trying to achieve?
- What data do we have, and can we trust it?
- Batch, streaming or both?
- Who owns and governs it?
- Who needs access, and to what?
- What does failure look like?
- What will this cost when people actually use it?
Architecture Patterns & Working Notes
Once the problem is understood, the technology becomes useful. The material below explores reference patterns for Azure, AWS, Snowflake, Databricks and modern lakehouse architectures.
<<Content to be added >>.