Search This Blog

Showing posts with label BIg Data. Show all posts
Showing posts with label BIg Data. Show all posts

Thursday, December 19, 2013

Can there be to "Perfect" of Data


As we deal with the deluge of data that is being collected and presented everyday do we spend too much time striving for perfection of said data when we could be spending our efforts better else where?

That is what I want to talk about.  The idea that perfection in data is necessary in some instances but, getting past the potential issues with a set of data can help us to move towards understanding the big picture.  What might this big picture be?   Sudden drop off in sales of a specific product or product area.  Increasingly negative feedback about a product, company or service.
These things do not require that we have perfect data.  We rarely need to dig down into the details so much that it matters.  What matters is you see a trend.  People are not buy that item any more.  Why?

You might also notice via the analysis of big data that sales in a certain region are very low compared to another.  If you in the snow chain business you will likely sell a lot more in Michigan than say Florida or South California.   Now you might see other trends that indicate that there is little movement in certain border snow areas.  That might leads to evaluation of additional marketing in those areas.  Maybe even evaluation of pricing or product placement.  The details come following the identification of a trend.  You might even identify untapped market areas or complimentary products you will want to sell.

Big data by its nature does not allow for detailed analysis.  You have to much data your dealing with.
So that is where the idea of trying to be to perfect in the data might make you slow to react to a trend.
If your spending to much time worrying about the "quality" of the data you might be missing out on the chance to find a trend or insight.  if you delay to long you might even miss the window of opportunity with a product.
One issues brought up was the merging together of multiple databases or sets of data from say 2 companies that just merged.  Do you want to wait 6-12 months while you work out the details of getting the data merged together properly.  Take the data set independent of each other, find the trends then go dig in the details when you notice some overlap.

Take the rainbow loom for example.  If someone just now is deciding they need to jump on this band wagon that might be 6 -12 months late to the show and the opportunity likely has been missed.
As you might start to see anecdotal information indicating its popularity is weaning.  New products launch with that information at hand can make a big difference.
You also have to make a concerted effort to look at the origin of the data in question.  Are you compiling a bunch of customer data from say web page visits or are you tagging into internal sales figures.  This will also help you decide how much time you need to spend making the data "perfect".

A great case example was given of Amazon.  They compile piles and piles of data on related searches and click throughs on products.  They have become very good at offering complimentary products and related items that might be of interest to the person on the site.  They have learned you don't have to be perfect in the suggestions provided.  Depending on which way it is being offered you can be fairly general and have a bunch of misses in the pile.  Not so bad.  On the other hand bundles need to be more on target.




"Amazon is good at this because they don’t worry about everybody. They develop a model where they’re eventually going to get a consistent model of the world, but at the moment they need to do it, they don’t care that they can’t roll it out for everyone. They’ve got hundreds of millions of clicks a day, and they figure, why don’t we just look at 20% of them? The key thing is to do it quickly and to make sure that whatever we conclude, there are many observations for it."  
"This is when the term “analytics” becomes interesting. Analytics doesn’t have to be based on super-precise data. That doesn’t again mean wrong data, but it might mean some outcome that wins for the customer. If you profile a jazz CD that people didn’t know they wanted, and some people buy it, great. The fact that some of the 100,000 people that you showed it to didn’t buy that CD is irrelevant."
Sid Probstein:Why Companies Have to Trade “Perfect Data” for “Fast Info”


The whole concept here is whether you worry about getting that last 5-10% of the data right (so you can have some potential increase in sales for example) or just live with what you have and go for the 5-6% increase (generally lesser increase) in sales that can come from that data.  The trade off is how much time and effort do you want to spend on "fixing" or getting right the last little bit of data for some small even minute additional increase?

Everything has to be in context.  You need perfection in say medical records, and even in medical billing but, do you need that perfection before you can do something.  To take action you have to decide at what level is this stuff have to be accurate.  Can it be 70-80% or does it have to be 95-100%.  Granted due to the ever changing issues with data and the dynamic nature of things you will find a "perfect" set of data may be 98% accurate.

IN some ways are we blessed by all the additional information at our finger tips.  Have we gotten away form the big picture?
Think of our ancestors some 300-400 years ago.   They may not have had the advantage of satellite imagery to help deal with a storm.  They always looked at the big picture and then when they saw a trend they would go for the details so that could properly plan.

We definitely need to get back to this type of analysis.  And not be overly jumpy to move on things if just a little bit more time or patience might be in order!




articles worth reading

Why Companies Have to Trade “Perfect Data” for “Fast Info”
Engineering The Perfect Big-Data Bra 
Obamacare: a lesson in data entry design
5 Reasons Why Consumer Collaboration and Big Data Are a Perfect Match
Predictive Data Delivering The Perfect Pricing Fit
Healthcare Data’s Perfect Storm
Digging through the Data: How to Separate Quality from Quantity


Buaidh - NO -Bas

Thursday, December 5, 2013

Consequences of Data Quality or Lack there of!

Over the last few years more and more has been brought up about data quality.  And now as they talk about big data it's coming up again and again.

Data quality is a bit hard to define as I've found from some posts in Linkedin Groups.  I means different things to different people.  Sometimes to the downfall of the people involved.
It really depends on the objective of the data in question.  Are they sales figures, contacts/customers or employee information, general information on response times, or something else of equal  importance?

Each of these has a different level of said "QUALITY".   Imagine if you have input some sales figures and they are off by a factor of 10 or seem to be replicated.  That would cause a lot of havoc in sales projection and other strategies.  You might even go off and order extra product that later becomes a front page listing at overstock.com.


I think though this ever excessive quantity of data could be causing some other unintended consequences.  We already know about the issues with a wrong phone or address and other typos.  We know about the wrong information being entered on what a customer is interested in or has bought.  Bad scans on products because the one in his hand did not come up in the system properly etc.
All those things people understand and can understand in general terms.  Sometimes these issues are hard to find in the system.  They can come up at a multitude of points along the process.




But, there is a bigger problem.  To much data and lack of coordination of that large amount of data.

Let me give two examples that will help to illustrate the problem.

One person related the story of how they received multiple phone calls from the same company on the same day but, from different offices.  You might say what the heck does this have to do with data quality.  It is rather simpler then you might think.  In this case there may not be duplication in the data (IE. the person entered twice in the database), there is duplication in the access to the data and the use of the data.
It is very apparent that this company does not have their people cross checking to make sure they are not selling to the same person from multiple offices.  Goes back to the old rule of sales regions.  If you have the "South" region here in the united states you'd be contacting only people in states like Florida, Georgia, Tennessee, Mississippi, etc.  You'd not be selling to a company based in Seattle Washington because it is out of your region.  Things can get complicated with multiple locations or offices though you'd technically not have the same guy listed for Seattle and Atlanta.  Also maybe the sales force is small for said company and guys in different offices sell all over the place.  However, proper protocol should have them at least entering some kind of contact date or sales information such that other sales reps know they are being contacted.
It did not sound like the reporter of the story has all that common of a name so it should be easy to figure out your going to contact the same person or maybe not.  And there in lies the problem.

See more on this story at Data Quality & Daily Life.

Bottom line in this situation is that multiple people had access to the data and had no clue anyone else was using it or how they were using it..   So that can be labeled a data quality issue just as much as anything else.  The data wasn't quality since no one knew how to use it properly or did not not have proper information associated to the record.

The second example happened to me.  I'm a member of an auto club that provides roadside assistant and related services.  I've been a member of this group for over 15 years.  I've also been at my current location for over 4 years.  I've also requested auto insurance quotes from them in the past.  No big deal.  The sad part was when an offer to become a member was sent in the mail to me.  As I stated previously I've been a member for a long time.  Now you'd think that what ever data bases they have they'd at least check to see if there is some overlap in the system (dedup) or at least try and verify if a person they want to send membership information to is not already a member.  You'd think they would want to at least review before taking further action anything that came up as a potential duplicate or current member.

This could in reality highlight one of the problems we now find ourselves in.  We have no much data on hand.  Could Big Data be causing us problems.  Do we have to much information to accurately and appropriately deal with?

I know from experience with the data I deal with there are what seem like a million pieces of information available or that has been compiled on people.  Routinely, we append consumer enhancements in a bundled form that includes things like credit cards, owner/renter, presence of child and related items.  However, we did do some work for the Obama campaign the first go around for the state of South Carolina.  You can append information on gun ownership, hunting, RVs, newly weds, or new parents, cellphone owner, and related tech topics, buying habits, ethnic information, age, income and housing information, and any number of interests.  Imagine how targeted your advertising can become if you have this information.  Also imagine how cumbersome your database may become if you included information on all these areas for each and every records or household in your system?

The great thing about a database you can split it into smaller chunks and still centrally link everything via your primary key.




This still does not justify the need for 300 different fields of data on 1 individual in all circumstances.
When we start to talk about big data we start to need to discuss these issues.  Have we due to so much data being available at our finger tips gotten away from doing due diligence in our sales endeavors, or customer service.  Have we justified the need for something or just taken things for granted since we have what seems like unlimited storage space due to the cloud?



So as we are deluged with extreme quantities of data we may very well need to say "is this information really necessary for our operations".  Those who have the responsibilities of designing databases, or maintaining them should be asking that question all the time.  Is this information adding value to what we are doing on a regular basis.?  Is there a better way to deal with this Big Data?

"And big data is really, really big. According to industry think tank IDC, the world now generates more data in 24 hours than existed in the history of planet earth up to the year 2000. If you want to put a figure on it, it was 2.5 quintillionbytes daily."

"Clearly this amount of data can be used as a goldmine of information for businesses around the world. But in its raw, unprocessed state, data can be more of a hindrance than a help, and the single customer view is no use if it doesn’t assist each department in achieving its objectives."
How Big Data Can Help You Close the Deal

Because as we get to the point of having terabytes of data being archived on a regular basis we need to consider the amount of space that takes, the amount of energy consumed moving it around and maintaining it, and the amount of time wasted as we do said actions.
We have to consider data corruption, and other related information along with the usual things related to data input errors.

We really need to change our way of thinking.  Bigger may not be better, but it can be if we use what we are given in an intellectual and prudent way.




So let us not fall into the trap of claiming - data is knowledge or power.  Unless, we put the resources at hand to proper use we may just end up with a bunch of offended customer and a soon to be "out of business" company.

"The success of its [Dell's] big data experiment proves that information is power, but that information on its own is useless. In order to harness the power of data – whether it’s big data, or data from your own CRM –  the data needs to be processed, cleansed, merged and combined.
But more importantly, it needs to be used selectively so that each department gains useful insight and can benefit from the single customer view."
How Big Data Can Help You Close the Deal




Here are couple of additional pages you can check out on data quality I think are worth the read

Big Data Can Tell Big Lies Through Fifth Normal Form (5NF)
“We hold these truths to be self-evident” (or do you trust your data?) 
How Big Data Can Help You Close the Deal
Data Quality & Daily Life 
Data Quality and the Blemishing Effect 


So now you have a few more things to think about in the data quality realm.  Are we doing ourselves a disservice as we head for the holy grail and embrace big data at full warp speed?

So look for the not so obvious as you work to ensure the data you are using is of the quality you want and need..... anything else would be yourself and others a disservice.  But it might just help out godd old overstock.com!

Buaidh - NO - Bas