Sunday, January 25, 2015

Module 1 Musings



Big data is the next evolutionary step for statistical analysis and predictive models. It is made possible by the large amount of data made available by the internet of things. This data can be used to “answer what, not why, and often that’s good enough” (The Rise of Big Data, 2). This is the future of business in healthcare, humanitarian efforts, civic planning, and all industries. There is so much data now that there are new job roles being created to manage all this data - the advent of the data scientist is bringing together diverse skills like programming, statistics and specific business understanding, among others, to be able to use the copious piles of data – much of it not in a usable state yet. Big data is all about handling its sheer volume, velocity of creation and variety of formats that are stored.

After completing the class reading and googling for “healthcare and big data” where I found titles like “NHS director argues big data in healthcare is a ‘moral obligation’”, “Why Health Care May Finally be Ready for Big Data” and “How Big Data Will Help Save Healthcare” I was left with the sense that big data is seen as being able to do no wrong – it will save us all! Most of the articles in in “Business Report on Big Data Gets Personal” focuses on the advantages of sharing personal data, how said data is bought and sold today, and how, well, you could identify people from the data from multiple data points but no one ever would. Those assumptions feel a little hollow after a year of many big companies – Home Depot, Chase, JP Morgan, and Sony, just to name a few, have suffered security breaches. Healthcare data is extremely popular on the black market (Monegain) and to effectively utilize big data there must be more thorough sharing of data between healthcare organizations. Security breaches are nearly inevitable and privacy concerns must be addressed.

Taking it in another direction, let’s assume there are no nefarious agendas, the use of big data could still lead to negative unintended consequences. For example, there is concern in the healthcare industry that predictive data could lead to health insurance agencies denying coverage for conditions that do not exist, may never exist, but a predictive model suggests could exist. Predictive models are just that – predictive and not infallible. We that in 2009, when Google used big data to find a correlation between internet searches between 2003 and 2008 and flu outbreaks. The idea being that if they could find a predictive model flu outbreaks could be analyzed in real-time. This mathematical analysis did the trick in that they found 45 search terms which often coincided with an outbreak. With a billion searches a day they wouldn’t have been able to predict these terms but by using such a large sample they felt confident in the results. The study concluded, though, with “[…] it seems that Google’s system may have overestimated the number of flu cases in the United States. This serves as reminder that predictions are only probabilities and are not always correct […]” (Cuckler, 3-4).

There are other limitations that must be conquered. Some of the industries that could benefit the most from utilizing big data are ill equipped to use the existing data. For example, in healthcare, from the paper patient records, electronic health records, lab results and all of the other data there is endless data that could be used to help drive diagnostic decisions. In an NPR interview with Amy Standen and Jenny Frankovich, they relate a story about a girl with lupus who might have been at risk for a blood clot. They faced the dilemma of proactively administering an anti-coagulant which might cause more problems or waiting for a potential blood clot to form before administering the anti-coagulant. One of the physicians used their electronic database with patient records to pull up the number of lupus patients and the number that had a clot. The course of action was made based on those results and was considered successful. So is this method in use by the hospital today? No. Why? Standen says "The system just isn't ready, the hospital decided. What if Frankovich had used the wrong search terms or the engine itself had bugs? What if the records has been mis-transcribed? Even Frankovich agrees that it's just too risky." The potential is there but this is an industry where the uncertainty, and messy data is a problem. Even if we assume the records are correct there is little to no standardization across organizations which means the results could be skewed if the data samples in each organization were insufficient. For this kind of clinical decision support healthcare really need the advantages of “N=All” sample sizes.

I do not want to come off as being anti-big data. In fact, I think the potential for good is limitless and has the potential to transform and how we interact with each other and other nations. In healthcare we already see Accountable Care Organizations (ACOs) are using big data to find trends to reduce the need for patient readmits and inefficient practices which will reduce healthcare costs (Burg). Another great example is from Canadian researchers who are using big-data to identify possible problems with premature babies long before the problem is evident. “By converting 16 vital signs, including heartbeat, blood pressure, respiration, and blood-oxygen levels, into an information flow of more than 1,000 data points per second, they have been able to find correlations between very minor changes and more serious problems” (The Rise of Big Data, 3). Even if they are not right 100% of the time those premature babies will get extra care and attention so the ones who are at risk can be saved at higher rates.

One of my favorite articles from this week was “Big Data from Cheap Phones” (Business Report - Big Data Gets Personal, 17-21). In developing countries where there is not existing government infrastructure, like a census, they are starting to use simple cell phone data points to under more about the nation and deal with problems. That article focuses on malaria outbreaks and how with just the data points of cell phones moving between cell towers they are able to track the outbreak and its flow. These seemingly simple data points will give them the information to focus their limited resources on the areas that will have the most impost. This is similar to the story in the “Rise of Big Data” about how Manhattan is using data analysis to deal fire-prevention in overcrowded buildings. By studying previous cases they found data points, which weren’t immediately obvious, suggested which buildings were more at risk. Their inspectors could focus their time on those buildings and have more impact (Cuckler, 4-5).

I’ll admit, these examples are much more motivating to me than how businesses can increase the impact of their marketing. Big data is here to stay – how can it not with the massive amount of data generated daily – and while the use and sharing needs to be controlled I look forward to seeing the improvement it will bring to lives across the globe.


Citations
Burg, Natalie. 'How Big Data Will Help Save Healthcare'. Forbes.com. 10 Nov. 2014: Web. 23 Jan. 2015.
Business Report - Big Data Gets Personal. MIT Technology Review. PDF. 20 Jan. 2015.
Cuckler, Kenneth Neil, and Viktor Mayer-Schoenberger. The Rise Of Big Data. ForeignAffairs.com. May/June 2013. PDF. 20 Jan. 2015.
Monegain, Bernie. "6 steps to keep security issues at bay." Healthcare IT News. 25 Apr. 2013. Web. 23
Jan. 2015.
Standen, Amy, and Jenny Frankovich. Big Data Not A Cure-All In Medicine. 5 Jan. 2015. Radio. 24 Jan. 2015.