The data warehouse has been around for at least three decades. It was created to address the shortcomings of the database. The database was the place for transactional data and it answered simple questions, like how many employees are there in the sales department. The data warehouse was created by aggregating historical data to answer analytical questions. Now you could ask the question – who was the best performing salesman over the last three years.
Data Warehouses are great if you want to understand the business. They are built to answer a set of known questions. The typical user is an analyst. These analysts are embedded in business teams.
There are, however three challenges with the current data warehouses.
- Freshness of data: Since data needs to be aggregated, complex transformations have to be performed before data is accessible
- Costs: Processing large amounts of data costs money. Things that are built need to be maintained as well.
- Scalability: The legacy softwares were not written for distributed computing
We tried to solve these problems technologically by creating data lakes. Instead of structuring and aggregating the data, the idea is that we dump all available data in its raw format in a lake and use distributed computing to process it. Since there is no structure, we could process multiple formats such as audio, video, text, logs etc.
Data lakes are great to test various hypotheses. They are used to answer unknown questions. For e.g. if we doubled our sales team, are we likely to increase sales by a factor of two? How much the sales revenue is due to the salesperson, the product and the economy? The typical user is a data scientist. These data scientists are usually part of a R&D team, away from the razzmatazz of daily business.
Data lakes have their own challenges in terms of cost, governance and lack of operationalisation.
All modern enterprises have analysts and data scientists in separate teams with separate goals – the analysts try to understand the business and the data scientists try to explore the business. This disconnect means that the business units miss out on the insights that the data scientists could deliver. And the data scientists ask their hypothetical questions but struggle to operationalise their work for ongoing delivery.
Cloud computing is changing that.
In the cloud, you do not have to pick between a data warehouse and a data lake. You can have a lake house.
For e.g. in Google cloud (the one I am most familiar with), the data warehouse product (Big Query) also offers storage API. Through the API, the data warehouse offers unlimited storage, a feature of the data lake. It also offers a series of connectors and data processing tools, making it a data warehouse. You can use its native BI engine to create views and dashboards to steer the daily business.
One cannot overstate the importance of culture in the organisation’s success. By bringing the data warehouse and the data lake together, organisations are juxtaposing the operational challenges of today with the hypothetical scenarios of tomorrow. In a world where data and insights are the competitive advantages, the lake house may end up being the differentiator between the digital winners and the losers.









