Most executives that I speak with – from large enterprises to startups – mention that experienced data science talent is the biggest bottleneck to their machine learning efforts. A lot of them imagine the process of building a model as a complicated, intellectually challenging process that requires very experienced folks. They see movies like matrix and visualize their brilliant data scientists typing away fervently at their computer terminals.
They would be right, if this were 2012.
In the intervening decade we got the cloud. And while it has improved a lot of things, I would argue that its biggest impact has been the democratization of the machine learning process.
Take the google cloud, for example. Google cloud has a petabyte scale analytics data warehouse called BigQuery. It offers a service called BigQuery ML. Users can create and execute machine learning (ML) models in BigQuery using standard SQL queries.
One can build a model with a simple query in 3 lines of code.
CREATE MODEL numbikes.model
OPTIONS
(model_type='linear_reg', labels=['num_trips']) AS
WITH bike_data AS
(
SELECT COUNT(*) a num_trips,
...
One can make a prediction with another query.
SELECT predicted_num_trips, num_trips, trip_date
FROM
ml.PREDICT(MODEL 'numbikes.model...
In just a few lines of code, your data analysts can build a prototype. I used linear regression in the example above, but it could be logistic regressions, clustering algorithms, deep neural networks – probably 90% of all the algorithms that we have available in the public domain.
Using service like BigQuery ML, teams can focus on building prototypes instead of writing machine learning algorithms from scratch. The product managers can evaluate the prototypes to identify the most promising ones. The data scientists can then focus on fine tuning the most promising prototypes.
They say all good writing is average writing that was rewritten. All good models are prototypes that were fine-tuned.

Leave a Reply