Friday, January 25, 2013

How big is Big Data?

Big data continues to be the topic of much discussion and hype, but “big” is really a red herring. Oil companies, telecommunications companies, and other data-centric industries have had huge datasets for a long time. And as storage capacity continues to expand, today’s “big” is certainly tomorrow’s “medium” and next week’s “small.” The most meaningful definition I’ve heard: “big data” is when the size of the data itself becomes part of the problem.

We’re discussing data problems ranging from gigabytes to zetabytes of data. At some point, traditional techniques for working with data becomes inadequate. What are we trying to do with data that’s different?  We’re trying to build information platforms or dataspaces. Information platforms are similar to traditional data warehouses, but different. They are designed for capturing and understanding the data rather than for traditional need for immediate analysis and reporting. They accept variety of data formats, including the apparently incomprehensible ones, and their structures (schemas) evolve as the understanding of the data improves.

Most of the organizations that have built data platforms have found it necessary to go beyond the
relational database model. Traditional relational database systems cease to be effective at the stream of such volume(scale). Managing sharing and replication across a stack of database servers is difficult, slow and costly. The need to define a schema in advance conflicts with reality of numerous, unstructured data sources, in which you are not aware what’s important until the data is analyzed. 

Data capture and storage is only part of forming a data platform, though. Data is only useful if you can do something with it, and enormous datasets present computational problems. Machine learning is another essential tool for the data scientist. We now expect web and mobile applications to incorporate recommendation engines, and building a recommendation engine is an artificial intelligence problem.
Building statistical models plays yet another important role in any data analysis. Statistics is the “grammar of data science.” It is crucial to “making data speak coherently.” We’ve all heard the joke that eating pickles causes death, because everyone who dies has eaten pickles. That joke doesn’t work if you understand what correlation means. More to the point, it’s easy to notice that one advertisement for R in a Nutshell generated 2 percent more conversions than another. But it takes statistics to know whether this difference is significant, or just a random fluctuation. Data science isn’t just about the existence of data, or making speculation about what that data might mean; it’s about testing hypotheses and making sure that the conclusions drawn from the data are valid. Thus Statistics plays a vital role in everything from traditional business intelligence to contemporary data analytics. It isn’t superseded by newer techniques from machine learning and other disciplines; it complements them.That's where different technologies evolve in addressing the challenges in modern day data analytics including Big Data.

Tuesday, January 8, 2013

Big Data - A synopsis

Every day we create about few quintillion bytes of data.  90% of the data in the world today has been created in the last 2 to 3 years alone. The sudden burst in growth of data can be attributed to: posts to social media sites, digital pictures and videos, purchase transaction records, cell phone GPS signals, sensors used to gather climate information to name a few. This data is BIG DATA.
BIG DATA is a collection of data sets so large and complex that it becomes difficult to process using on-hand database management tools or traditional data processing applications thus it outgrows your current ability to process it, store it, and cope with it efficiently.
BIG DATA is 4D in IT spatial: Size, Acceleration, Form, Accuracy (SAFA)
Size does matter after all.
Sometimes big data is measured in terabytes, petabytes, exabyte, zettabyte or more. In real word, it's usually measured in frustration, annoyance, anxiety, and money down the drain.
The challenges include capture, curation, storage, search, sharing, analysis, and visualization. The trend to larger data sets is due to the additional information derivable from analysis of a single large set of related data, as compared to separate smaller sets with the same total amount of data, allowing correlations to be found to "spot business trends, determine quality of research, prevent diseases, legal citations, combat crime, and determine real-time roadway traffic conditions
Acceleration
For time critical processes such as fraud detection in trade events, predict customer churn etc. BIG DATA must be used as it flows into your enterprise in order to maximize it's value. It's just not a race against time rather to derive potential insights that provides innovative ways of doing things.
Form
BIG DATA is varied type of data - structured and unstructured data such as text, video, audio, sensor data, click streams, log files and much more. New insights are found when these data types are analyzed together.
Accuracy
1 in 3 business leader do not trust the information they use to make decisions. How can one act upon information that they don't trust. Establishing trust in BIG DATA presents a huge challenge and the variety and number of source grows.
BIG DATA is more than simply a matter of size; it is opportunity to find insights in new and emerging types of data and content.