Data Council Blog

30/09/17 01:51 | by Pete Soderling | in Data Engineering, Event Updates, Databases

Rolling Your Own Distributed Column Store

When solving your customers' technical challenges push you to break the rules

A re-wording of one of the key maxims for startup success could be "KISS" - "keep it simple, stupid." If you've ever run your own startup, you also know the mantras of "focus" and "fail fast," and the critical reminder of how your product should be a "pain-killer not a vitamin."

14/09/17 08:00 | by Pete Soderling | in Data Science, big data, Data Visualization, disaster management

How Big Data Can Help Improve the Meteorological Risk Models That Are Out of Date

According to a recent article published in The New York Times, water damage from hurricane Harvey extended far beyond flood zones. Now that the rescue efforts are underway, it’s clear that much of the damage occurred outside of the typical boundaries drawn on official FEMA flood maps.

17/04/17 09:07 | by Pete Soderling | in Data Science, Data Engineering, Speaker Spotlight

A Day in the Life: What's it like Being an Engineer at Stripe?

Alyssa Frazee tells us about the unicorn data skills she's honed on the job.

One thing that Alyssa Frazee loves about her work at Stripe is that, like someone with traditional data science skills, she gets to build machine learning models. "Oh, the rapture," cries Alyssa the data scientist!

12/04/17 12:21 | by Pete Soderling | in Data Engineering, Event Updates

Rebuilding Open Source Analytics @ Airbnb

How open source allowed Airbnb to rebuild their expensive BI tool in less than one developer year

Granted Maxime Beauchemin isn't your average data engineer. As any Bay Area engineer worth their salt knows, anyone who worked on data at Facebook receives (deserves) a certain outsized respect from their peers.

06/04/17 12:45 | by Pete Soderling | in Data Engineering, Speaker Spotlight

Pushing Kafka to the Limit at Heroku

How Everyone's Favorite PaaS Operates Kafka at Scale

Scale presents unique challenges for engineers, particularly those at companies who have the largest number of users throwing off the most data exhaust, resulting in the fattest data pipelines with the gnarliest problems. For example, Heroku, arguably the most popular platform as a service (PaaS), who last year decided to offer Apache Kafka to their customers as a hosted service, quickly realized they would need to support a large number of distinct users, each with varying use cases. This put them on a challenging path to attempt to minimize the operational headaches that come inherently with running this kind of infrastructure at scale.

28/03/17 14:05 | by Pete Soderling | in Data Science, Speaker Spotlight

Fighting Fraud in Cryptocurrency using Machine Learning

Coinbase is on the front-lines of discovering advanced cryptocurrency and payment fraud techniques. Hear about how they use machine learning to help them fight the war.

23/03/17 12:07 | by Pete Soderling | in Data Engineering, Speaker Spotlight

Building a Column-Oriented, Distributed Data Store for Analytics - The Story of Druid

Druid is a modern data store built for analytics use-cases. As the volume of data has exploded, and companies have sought deeper insights from their data, ad-hoc analytics have become difficult as more data is buried in distributed systems like Hadoop & Spark. The query model for these systems can result in long latencies making them sub-optimal for interactive analytics applications.

21/03/17 11:15 | by Pete Soderling | in Data Engineering, Speaker Spotlight

How to Build a Data Pipeline That Handles Hundreds of Different Inputs

How many different file formats does your ETL system need to parse? For many data pipelines, several well-defined formats will suffice. Things break, and at times require manual intervention, but not so often that a couple engineers can't keep tabs on the system and keep things running relatively smoothly.

Pete Soderling

Rolling Your Own Distributed Column Store

When solving your customers' technical challenges push you to break the rules

How Big Data Can Help Improve the Meteorological Risk Models That Are Out of Date

A Day in the Life: What's it like Being an Engineer at Stripe?

Alyssa Frazee tells us about the unicorn data skills she's honed on the job.

Rebuilding Open Source Analytics @ Airbnb

How open source allowed Airbnb to rebuild their expensive BI tool in less than one developer year

Pushing Kafka to the Limit at Heroku

How Everyone's Favorite PaaS Operates Kafka at Scale

Fighting Fraud in Cryptocurrency using Machine Learning

Building a Column-Oriented, Distributed Data Store for Analytics - The Story of Druid

How to Build a Data Pipeline That Handles Hundreds of Different Inputs

Subscribe to Email Updates

Fresh Posts

Categories