connectedcloud

What Every Marketer Needs to Know About Hadoop

  |  August 30, 2012   |  Comments

With "big data" on everybody's lips, here's all you need to know to keep up your end of the conversation.

"Big data." There's no escaping it.

It's catchy. It's generic enough that everybody is using it for everything. It's a one-size-fits-all phrase.

It's so all-encompassing that the best definition I've seen recently is from Stephané Hamel who put it this way:

bigdatatweet

So with "big data" on everybody's lips, here's all you (the marketing executive) need to know to keep up your end of the conversation.

A. Disk drives got cheaper so we can store more data. The ways and means of collecting all sorts of data have proliferated faster than Twitter traffic or TSA lines at the airport. We have more of data, more types of data, and it's coming at us faster (real time) than ever dreamed possible. That's what makes up "volume, variety, velocity."

So, the ability to replace big, honkin' disk drives with many smaller, cheaper drives that we can wire together is the first, significant technical advance.

B. We can split up the processing. The second advance is the ability to augment the big, honkin' processors with many smaller, cheaper servers. We have distributed the processing to the data instead of waiting for the data to rocket back and forth from disk farm to processor.

connectedcloud

So What?

So, there are two things to keep in mind when your marketing budget is being allocated to what seems like pure IT projects.

  1. The more data you throw into the pot, the more likely you are of finding some sort of relationship (correlation) to act on. More on that can be found in a July column I called "Consilience - The Intrinsic Value of Big Data."
  2. This practice of splitting up the data, solving smaller problems, and bringing it back together (MapReduce) is very useful for some specific types of processing. Getting this under your belt gives you voting rights when discussing options.

Big, honkin' analytics processors are very good at finding hidden pieces in a hurry. (Show me all the customers who have bought in the past three months after clicking on these special offers and abandoning their shopping carts.)

But those types of questions are known unknowns. You know the things you're going to ask and the entire database is set up that way. You know you'll want to see things by date, by region, by product line, etc. That is what gives these enterprise data warehouses their power: they are designed in advance to answer the questions you know you might ask, and they can answer them very quickly so you can refine your questions - as long as you have deep knowledge about what data you have and how it is structured in the database.

hal-seye

But the other data - the messy data - is chock-full of unknown unknowns. We know the information might be valuable, but we don't know what to ask.

MapReduce is great as a low-cost storage medium for unstructured data and for refining that data into a more structured form for heavy analysis. Social media data, call center transcripts, clickstream data, website content, and sensor data all start out unstructured.

MapReduce is ideal for pre-processing text, turning all those tweets into numerical models of opinion (sentiment analysis), which can then be fed to the big, honkin' analytics machines for correlation discovery and problem solving. It's great for asking slower questions of larger amounts of data. It's great for finding a representative sample of data so the big, honkin' processors don't have to juggle all of the bits at once.

So the next time somebody throws "Hadoop" into the conversation, you'll know more than the fact that it was named after Doug Cutting's son's stuffed elephant.

Connected Cloud and Hal's Eye images via Shutterstock.

Tags:

ClickZ Live San Francisco This Year's Premier Digital Marketing Event is #CZLSF
ClickZ Live San Francisco (Aug 11-14) brings together the industry's leading practitioners and marketing strategists to deliver 4 days of educational sessions and training workshops. From Data-Driven Marketing to Social, Mobile, Display, Search and Email, this year's comprehensive agenda will help you maximize your marketing efforts and ROI. Register today!

ABOUT THE AUTHOR

Jim Sterne

Jim Sterne is an international consultant focused on measuring the value of the online marketing for creating and strengthening customer relationships. Sterne has written eight books on using the Internet for marketing, produces the eMetrics Marketing Optimization Summit and is co-founder and current chairman of the Digital Analytics Association.

COMMENTSCommenting policy

comments powered by Disqus

Get the ClickZ Analytics newsletter delivered to you. Subscribe today!

COMMENTS

UPCOMING EVENTS

Featured White Papers

BigDoor: The Marketers Guide to Customer Loyalty

The Marketer's Guide to Customer Loyalty
Customer loyalty is imperative to success, but fostering and maintaining loyalty takes a lot of work. This guide is here to help marketers build, execute, and maintain a successful loyalty initiative.

Marin Software: The Multiplier Effect of Integrating Search & Social Advertising

The Multiplier Effect of Integrating Search & Social Advertising
Latest research reveals 68% higher revenue per conversion for marketers who integrate their search & social advertising. In addition to the research results, this whitepaper also outlines 5 strategies and 15 tactics you can use to better integrate your search and social campaigns.

WEBINARS

Jobs

    • Internet Marketing Campaign Manager
      Internet Marketing Campaign Manager (Straight North, LLC) - Fort MillWe are looking for a talented Internet Marketing Campaign Manager to join the...
    • Online Marketing Coordinator
      Online Marketing Coordinator (NewMarket Health) - BaltimoreWant to learn marketing from the best minds in the business? NewMarket Health, a subsidiary...
    • Call Center Manager
      Call Center Manager (Common Sense Publishing) - Delray BeachWanted: Dynamic Call Center Manager with a Proven Track Record of Improving Response...