Kiran's Blog

About the cloud, technology and other fun things in life

  • Posts
  • Contact
  • About
  • Contact
  • Home

Database, Data Warehouse and the Data Lake

By Kiran Kulkarni on August 22, 2021

The data warehouse has been around for at least three decades. It was created to address the shortcomings of the database. The database was the place for transactional data and it answered simple questions, like how many employees are there in the sales department. The data warehouse was created by aggregating historical data to answer analytical questions. Now you could ask the question – who was the best performing salesman over the last three years.

Data Warehouses are great if you want to understand the business. They are built to answer a set of known questions. The typical user is an analyst. These analysts are embedded in business teams.

There are, however three challenges with the current data warehouses.

  1. Freshness of data: Since data needs to be aggregated, complex transformations have to be performed before data is accessible
  2. Costs: Processing large amounts of data costs money. Things that are built need to be maintained as well.
  3. Scalability: The legacy softwares were not written for distributed computing

We tried to solve these problems technologically by creating data lakes. Instead of structuring and aggregating the data, the idea is that we dump all available data in its raw format in a lake and use distributed computing to process it. Since there is no structure, we could process multiple formats such as audio, video, text, logs etc.

Data lakes are great to test various hypotheses. They are used to answer unknown questions. For e.g. if we doubled our sales team, are we likely to increase sales by a factor of two? How much the sales revenue is due to the salesperson, the product and the economy? The typical user is a data scientist. These data scientists are usually part of a R&D team, away from the razzmatazz of daily business.

Data lakes have their own challenges in terms of cost, governance and lack of operationalisation.

All modern enterprises have analysts and data scientists in separate teams with separate goals – the analysts try to understand the business and the data scientists try to explore the business. This disconnect means that the business units miss out on the insights that the data scientists could deliver. And the data scientists ask their hypothetical questions but struggle to operationalise their work for ongoing delivery.

Cloud computing is changing that. 

In the cloud, you do not have to pick between a data warehouse and a data lake. You can have a lake house. 

For e.g. in Google cloud (the one I am most familiar with), the data warehouse product (Big Query) also offers storage API. Through the API, the data warehouse offers unlimited storage, a feature of the data lake. It also offers a series of connectors and data processing tools, making it a data warehouse. You can use its native BI engine to create views and dashboards to steer the daily business.

One cannot overstate the importance of culture in the organisation’s success. By bringing the data warehouse and the data lake together, organisations are juxtaposing the operational challenges of today with the hypothetical scenarios of tomorrow. In a world where data and insights are the competitive advantages, the lake house may end up being the differentiator between the digital winners and the losers.

The cloud database as a proxy for digital maturity

By Kiran Kulkarni on August 13, 2021

Around 2010, I was running the analytics department for a large grocery retailer. We had one of the most sophisticated tools for analyzing the data. We had a fairly competent team. We had mature processes. In short, we had all the required ingredients for complex analyses.

And yet, my analysts took four hours to answer the simple question – “how many apples did we sell last week?”. if you asked two different analysts, you would get two – possibly three different numbers. And none of them were wrong.

Welcome to the challenges of database schema.

A schema represents that logical and physical structure of data in a relational database.  

Schemas decide how data is stored (and by extension, how the data is accessed) The data storage rules impact how the data is aggregated. All analyses are built on aggregated data. For example, you aggregate the millions of transactions by day or week, so that you don’t have to add it up every single time. 

Schemas provide structure and structure builds efficiency. At a time when computing was relatively expensive, schemas allowed you to do more with less.

Structure also builds rigidity. Database schema cannot be changed easily. Extracting data out of 2 different schemas is a nightmare. 

Back to the question of apples. The retailer had grown through acquisitions. There were two major databases and two schema. In one of them, apples were stored as a category. The sub-categories were fresh and frozen. In the second retailer, the category was Produce. And it had red apples and green apples. In the first retailer, a calendar week started on a sunday. In the 2nd retailer, it started on a saturday.

How much structure we need in a database is a function of efficiency vs. agility. The modern cloud-native databases are built with flexible schema. These databases are optimized for maximum fungibility of data. They are built knowing that computing is available on demand. They are built for today’s VUCA (volatile, uncertain, complex, ambiguous) world, where the biggest competitive advantage comes from staying nimble. 

The ratio of data sitting in structured databases vs. semi-structured or cloud-native databases is a proxy indicator of how advanced the company is in its digital journey. 

Business Model and Revenue Model in digital era

By Kiran Kulkarni on August 7, 2021

Business model refers to the value the firm provides to its customers. The business model of a pizza delivery company is to deliver pizza. 

Revenue model refers to the different sources of revenue. Revenue models can include different channels – B2B, B2C – or different sources, like advertising or subscription.

In the pre-digital era, the business model and revenue model were the same. A pizza delivery company made money making and delivering pizzas.

In the digital era powered by cloud computing and mobile phones, the business model and revenue model are diverging. 

Let’s take the example of Domino’s pizza. Domino’s was a winner in digitization till now. It did all the things right in the digital era so far.

First, it invested heavily in a successful omni-channel digital ordering platform. You can order your pizza on the web, the mobile app, with a phone call, sms etc. 

Second, it mined customer data for revenues. If a customer ordered pizza every friday, she got a SMS on Friday at 4 PM for a pizza at a discount. 

Third, Domino’s embraced social media for engagement. 

The result: Its stock price went from less than 4 USD in 2008 to more than 450 USD in 2021. That is a stellar performance. And yet, it faces an existential threat going forward. Not from a rival pizza delivery company. But from the likes of Doordash and Uber eats. They do not offer pizza. They offer an app through which you can order food from any restaurant on their network.

While Domino’s can expect one order per week from a given customer, Doordash expects to deliver food 3-4 times per week. They have much richer data about the customer preferences – data that they can easily mine for insights that help ghost kitchens. These are a single physical location operating multiple web storefronts on various delivery apps. 

Ghost kitchens are cloud technologies optimized for the commercial kitchen. Everything-as-a-service. They are connected to all major delivery apps, they have a payment gateway, order submission, pick-up and delivery – everything as a service.

Domino’s has a market cap of approximately 20 Billion USD on the back of revenues selling pizzas. It has the technology to make pizzas and a digital app to connect you. Doordash has a market cap of roughly 60 Billion on the back of a digital marketplace to order food. It gets a share of whatever you order. Its revenue model is the insights it derives from your data coupled with the agility that comes with cloud enabled ghost kitchens that can quickly adapt to your needs.

Consumer technology vs. enterprise technology

By Kiran Kulkarni on July 30, 2021

Traditionally, enterprise technology was way better than anything that the consumers had. Starting the mid 90’s, something extraordinary happened. Since the advent of the internet, consumer technology has outpaced enterprise technology. 

In the last two decades, the consumers have got realtime video, mail with AI/ML prompts and voice assisted smart appliances. And these products keep getting better every week. The updates happen so frequently and seamlessly, that we don’t even notice them until we are asked for a password because the device had to restart.

By contrast, enterprise IT still operates in much the same old, slow way it has always been. Why has it fallen behind in innovation? 

The web provided consumer technology providers with a single programmable interface – the HTML. Consumers got every application from every vendor over the web. The same thing happened to the smart phone with IOS and Android. In both the cases, the interface(s) are based on open standards and they auto update. Everyone is on the same version, moving with the same velocity. 

This allowed developers to spend more time developing technology. All they cared about was it should be delivered on HTML.

Gradual Rollouts:

Since the delivery platform was fixed, the delivery methodology became continuous. We know this as the CI/CD (continuous integration, continuous deployment) process. A new feature is released to a small percentage of users. Once successful, they just roll it out to everybody. 

Continuous testing:

They release the software to 1% of the users and they fix the bugs before it reaches the majority of the customers. They have removed the fear of upgrades from our minds

The enterprise customers lacked a single delivery platform until now.

Enterprises traditionally bought their datacenter hardware from the likes of HP, IBM or Dell, each with their own flavors of Unix, Linux or the Windows operating system. And they were locked-in with these vendors. Their choice of hardware and OS guided at times their choice of middleware and applications. 

So for the enterprise software vendors, every company represented a unique lego piece to integrate with. They had to worry about delivering their software and releases to multiple platforms and multiple versions of platforms. Even if the software had been installed hundreds of times, it was the first installation for that combination of applications and versions. 

The cloud is an opportunity to change all of that. With Anthos (GCP), EKS-A (AWS) or Arc (Azure), enterprises have the opportunity to port their legacy applications into a modern application management platform that works on-premises as well as the cloud. And at least with Anthos, it is based on open-source kubernetes framework and works on any datacenter or cloud platform. Now vendors can release one version of their software that is compatible with an open-standards based interface (Anthos, for example) and it will work fine for anything below it. This will allow enterprises to consume new releases much faster.

Conservatively speaking, 20% less friction from an open-source based common programming model over 25 years leads to a factor of 100.

The Edge-to-Cloud Continuum

By Kiran Kulkarni on July 24, 2021

The computing paradigm has swung between centralization and decentralization over the last 50 years. The mainframes represented the centralized computing era. Then came the era of PC’s. While the IT departments used mainframes, the typical office worker used the PC. 

I view the cloud as another era of centralized computing with two key differences to the mainframes. First, the computer is not on your premises; and second, the computer is infinitely scalable. It serves the digital era where everything is a service (hosted in the cloud). 

Edge computing is the equivalent of the decentralized PC of the cloud era. Edge computing is computing that’s done at or near the source of the data. This could be on the device itself, or a small server closeby.

Edge computing addresses two major limitations of cloud computing.

Latency: Cloud computing is constrained by the speed of light. When we ask our voice assistant for the weather forecast, the assistant sends a compressed file to the cloud for processing, which, in turn, calls the weather API (Application Programming Interface) to return the answer. As we get closer to everything-in-real-time, we would like the computing to be closer to us.

Bandwidth: Sometimes the internet bandwidth is not available or it is cost prohibitive. For example, an oil rig operating in the ocean. The sensors and robots generate lots of data which are better analyzed at the location or nearby. 

The exponential growth of IoT devices will need edge computing. They will build on cloud computing concepts such as microservices, containerization and APIs (Application Programming Interfaces). Newer programming frameworks will be built on top of the cloud computing frameworks to make it a cloud-to-edge continuum.

In this emerging cloud-to-edge computing continuum my fear is for the companies who do not see the future unfold right in front of them. The question is not if the current legacy, monolithic application serves its purpose. The question is – does it prepare me to take advantage of the new paradigm that is emerging.

The horse or the car

By Kiran Kulkarni on July 17, 2021

You either ride a horse or you drive a car. But why would you mount a horse on a car?

That is what I said to a colleague whose customer wanted to bring his relational database to the cloud.

Horses were the means of transport before the car was invented. They were used for everything – farmers used them in agriculture, armies used them in wars, traders used them to travel. We did not do intercontinental travel because horses did not fly. 

Then we had mechanized transport. We optimized the vehicle for a certain task. We used a car for work, tractors for our fields, tanks for the battlefield and airplanes to fly across continents.

Relational databases were like the horses – used for everything. They powered internal sales or finance departments reporting, the hosted retail Point-of-Sale (PoS) solutions and even the government social security organizations. 

Now we are in the digital era. The most successful companies now are – in a sense – database companies. Facebook is a huge graph database of all our relationships. It is optimized to derive value out of our networks. Google search is a database of the entire internet, optimized to scan web pages and deliver accurate results in milliseconds. Amazon is a database of every person’s buying pattern. It is optimized to recommend me kindle books based on my shopping preferences. These companies use pick and choose databases based on the task at hand. Usually it is a combination of multiple databases.

Everybody talks about the importance of data in the digital world. What is less understood is that every company needs to reimagine itself as a database company. For those that are ready, the cloud offers various options including relational, document, graph, key-value, in-memory, and data warehouse databases. 

It is time to quit the horse and pick the vehicle based on your journey.

Climate and the Cloud

By Kiran Kulkarni on July 11, 2021

Last week, western America witnessed an unprecedented heat wave. It is a poignant reminder of the fact that the climate is warming and our carbon footprint needs to reduce.

How does the cloud fit into that?

Computers emit heat as a by-product of their operation. Large data centers place huge demands on the grid for their cooling needs, which increases the carbon footprint. This demand is going to exponentially grow as the digital economy expands. 

Cloud computing helps our decarbonization efforts in two ways.

Decarbonizing of computing: 

Between 2010 and 2018 the amount of computing done in data centers increased by about 550%. But the amount of electricity consumed increased only 6%. I believe that in the coming decade, with the large-scale adoption of cloud computing, the total carbon footprint of computing services will actually reduce. The cloud providers are innovating to first, produce less heat and second, to make the cooling process more sustainable. For example:

Energy efficient Processing Units: Google Cloud Tensor Processing Unit (TPUs) (highly efficient compute chips to run machine learning applications) are designed with energy efficiency in mind, specifically to accelerate deep learning workloads at higher teraflops per watt compared to general purpose processors.

Innovative cooling techniques: Microsoft’s project Natick is experimenting with underwater servers to reduce the cooling needs. Initial results show a hardware failure rate of 1/8th that of a land based control group. 

Decarbonizing by computing: 

The cloud is – and will continue to transform our business to make it more sustainable.

Using data and AI/ML to reduce waste: It is well known by now that the use of AI/ML have driven efficiency across various industries. For example, predicting demand more accurately drives supply chain efficiency  among retailers, causing their carbon footprint to to reduce.

What caught my attention recently is the swiss agritech startup Gamaya. It uses cloud computing to capture light waves by deploying hyperspectral imaging. The technology was originally developed by NASA. Healthy plants reflect light differently than deceased ones. By capturing the last nuance of light, Gamaya develops granular diagnostic heat maps from Brazil to India. This data (and the algorithms built on top of it) has the potential to transform agriculture and make it more sustainable.

To put things in perspective, agriculture is responsible for roughly one quarter of all greenhouse gas emissions.

NFT

By Kiran Kulkarni on July 3, 2021
digital vs. paper currency

In March 2021, when Jack Dorsey’s first tweet auctioned as a NFT (non-fungible token) for approximately 2.9 million USD, it heralded a new era in digital economy – the market for digital collectibles had come of age. The transaction volumes of NFT’s in Q1, 2021 were more than 1.5 billion USD, representing a growth of more than 2600% over Q4 2020.

NFT’s are digital tokens. They are recorded in a global digital ledger so everybody knows who owns which token. What makes them unique is that each token has a digital asset assigned to it. The digital asset can be an image, an audio or a video to name a few. Our creativity is the sole limitation on the type of digital asset.

The NFTs allow artists to create collectibles. Their fans can own these collectibles. Investors buy and sell them at appropriate prices. 

But most important for me, they allow new marketplaces to emerge. Companies like OpenSea and rarible have emerged as digital marketplaces where NFTs are bought and sold. 

All these companies were formed in the last few years and of course, they use cloud computing technologies and services. At the core, NFTs are cryptoart that are transacted in cryptocurrencies. And cryptocurrencies are built on the blockchain database. The technology and the pace of innovation dictates that it runs in the cloud.

I am sure the digital natives will soon follow suit. Companies like soundcloud and spotify that are already providing digital marketplaces for artists, who have the tech infrastructure as well as the skills to innovate and bring digital collectibles as additional offerings on their platform.

What I do not understand is that some players in the entertainment ecosystem are still arguing about why they should invest in cloud computing infrastructure because apparently their current infrastructure serves their current business just fine. I can see Blockbuster written all over their future.

Success in the digital era requires one to build digital muscles. Cloud computing is the bench press that helps you build that muscle.

Bundling and Unbundling

By Kiran Kulkarni on June 23, 2021

In the late 90’s, the founder of Ebay sold a broken laser pointer in an auction. He asked the buyer if he understood that he was buying a broken product. The buyer said, he is a collector of broken laser pointers!

Soon, Ebay connected all the collectors of the world into one giant marketplace. It went on to become arguably the largest online auction platform. Millions of customers were trading billions of articles amongst themselves. 

About 20 years later, a collector of sneakers ordered Jordan sneakers over Ebay. They turned out to be fake. He realized that Ebay does not have a process to verify the products that individuals buy or sell on its website. He went ahead and created goat.com. That company vets every sneaker for authenticity.

The two stories best illustrate how technology has changed the business model.

Ebay’s business model is optimized to be everything for everybody. It is the hallmark of the industrial era applied to an internet platform.

On the contrary, goat.com is built for specific categories (sneakers & apparel). It is the hallmark of the digital era, where everything is customized using data.

In the late 1990’s, the web brought us together into a global marketplace. The geographic chains of the physical world had shackled us until then. Platforms like Ebay rode the wave. We became buyers, sellers or both on one global auction site.

These days, the cloud is democratizing the technology that removes the vulnerabilities of the previous era where we were users, not individuals. Our footprints across multiple digital touch points are being analyzed to serve us as individuals. An undifferentiated platform serving the lowest common denominator is no longer good enough.

They say there are two ways to make money – Bundling and unbundling (of products and services). Welcome to the era of unbundling.

Efficiency Vs. Effectiveness

By Kiran Kulkarni on June 19, 2021

Efficiency refers to the resource utilization in getting to the desired outcomes. It is unidimensional. Absolute. The mileage on a car, for instance. 

Effectiveness refers to how useful something is. It is contextual. A highly efficient car may be completely ineffective in winning you a formula 1 race. A fast car, however, is ineffective during an off-road camping trip.

A bottom-up approach improves efficiency.  If you give the middle managers an efficiency target they will achieve it. 

On the contrary, a top-down approach drives effectiveness. It starts with the vision. Whether the company wants a car for camping weekends or adrenaline filled, high-speed drives is not a decision that the middle management is paid to take.

So, what has all this got to do with the cloud?

Until recently, organizations bought cloud computing for efficiency. Now, they need to buy it as part of their effectiveness on digital transformation. And that’s the hard part. Digital transformation means different things to different companies. Each journey is unique. There is no single metric for it. There are, however, some hallmarks to identify the digitally mature companies.

For instance, one of the hallmarks of a digital leader is the use of data to drive its decisions at all levels. This means it has a huge appetite for machine learning use cases. By extension, it is testing hundreds of hypotheses. The tools required for such a test-and-learn mindset are only available in the cloud. 

Take any digital transformation journey. You will find the tools required for an effective transformation invariably in the cloud.

The cloud spend is a proxy for the digital transformation journey of the company.

Previous
Next
© 2026 Kiran's Blog. tru Theme by SPYR
✕
The Navigation Area? is not yet configured. Simply set a menu to have a Display Location of Navigation Area Menu.
Loading Comments...