Showing posts with label data science. Show all posts
Showing posts with label data science. Show all posts

Saturday, 12 August 2017

Why Data Visualization matter now?

Data Visualization is not new, it has been around in various forms for more than thousands of years. 

Ancient Egyptians used symbolic paintings, drawn on walls & pottery, to tell timeless stories of their culture for generations to come.

Human brain understands the information via pictures more easily than writing sentences, essays, spreadsheets etc. You must have seen traffic symbols while driving…why do they have only 1 picture instead of writing a whole sentence like school ahead, deer crossing or narrow bridge? Because you as driver can grasp the image faster while keeping your eyes on the road.

Over last 25 years technology has given us popular methods like line, bar, and pie charts showing company progress in different forms, which still dominate the boardrooms.

Data visualization has become a fundamental discipline as it enables more and more businesses and decision makers to see big data and analytics presented visually. It helps identify the exact area that needs attention or improvement than leaving it to the leaders to interpret as they want.

Until recently making sense of all of that raw data was too daunting for most, but recent computing developments have created new tools like Tableau, Qlik with striking visual techniques, especially for use online, including the use of animations.

There is a wealth of information hiding in the data in your database that is just waiting to be discovered. Even historical complicated data collected from disparate sources start to make sense when shown pictorially. Data Scientists do a fantastic job of analyzing this data using machine learning, finding relationship but communicating the story to others is the last milestone.

In today's Digital age, we as consumers generate tons of data every day and businesses want to use that for hyper-personalization, sending right offers to us by collecting, storing & analyzing this data. Data Visualization is the necessary ingredient to bring power of this big data to mainstream.

It is hard to tell how the data behaves in the data table. Only when we apply visualization via graphs or charts, we get a clear picture how the data behaves. 

Data visualization allows us to quickly interpret the data and adjust different variables to see their effect and technology is increasingly making it easier for us to do so. 

The best data visualizations are ones that expose something new about the underlying patterns and relationships contained within the data. Data Visualization brings multiple advantages such as showing the big picture quickly with simplicity for further action.

Finally as they say “A picture is worth a thousand words” and it is much important when you are trying to show the relationships within the data.

Data is the new oil, but it is crude, and cannot really be used unless it is refined with visualization to bring the new gold nuggets.

Sunday, 6 August 2017

Do you want to hire a Data Scientist?

As mentioned by Tom Davenport few years back, Data Scientist is still a hottest job of century.

Data scientists are those elite people who solve business problems by analyzing tons of data and communicate the results in a very compelling way to senior leadership and persuade them to take action.

They have the critical responsibility to understand the data and help business get more knowledgeable about their customers.

The importance of Data Scientists has rose to top due to two key issues:
·     Increased need & desire among businesses to gain greater value from their data to be competitive
·     Over 80% of data/information that businesses generate and collect is unstructured or semi-structured data that need special treatment

So it is extremely important to hire a right person for the job. Requirements for being a data scientist are pretty rigorous, and truly qualified candidates are few and far between.

Data Scientists are very high in demand, hard to attract, come at a very high cost so if there is a wrong hire then it’s really more frustrating. 

Here are some guidelines for checking them:
·     Check the logical reasoning ability
·     Problem solving skills
·     Ability to collaborate & communicate with business folks
·     Practical experience on collaborating Big Data tools
·     Statistical and machine learning experience
·     Should be able to describe their projects very clearly where they have solved business problems
·     Should be able to tell story from the data
·     Should know the latest of cognitive computing, deep learning

I have seen smartest data scientists in my career, who do the best job at analytics, but cannot communicate the results to senior leaders effectively. Ideally they should know the data in depth and can explain its significance properly. Data visualizations comes very handy at this stage.

Today with digital disrupting every field it has an impact on data science also.

Gartner has called this new breed as citizen data scientists. Their primary job function is outside analytics, they don’t know much about statistics but can work on ready to use algorithms available in APIs like Watson, Tensor flow, Azure and other well-known tools.

The good data scientist can make use of them to spread the awareness and expand their influence.

It has become more important to hire a right data scientist as they will show you the results which may make or break the company.


Saturday, 3 December 2016

Digital Transformation helping to reduce patient's readmission

Digital Transformation is helping all the corners of life and healthcare is no exception.

Patients when discharged from the hospital are given verbal and written instructions regarding their post-discharge care but many of them get readmitted in 30 days due to various reasons. 

Over last 5 years this 30 days readmission rate is almost 19% with over 25 billions of dollars spent per year.

In October 2012 the Centers for Medicaid and Medicare Services (CMS) began penalizing hospitals with the highest readmission rates for health conditions like acute myocardial infarction (AMI), heart failure (HF), pneumonia (PN), chronic obstructive pulmonary disease (COPD) and total hip arthroplasty/total knee arthroplasty (THA/TKA).

Various steps to reduce the readmission:

·        Send the patient home with 30-day medication supply, wrapped in packaging that clearly explains timing, dosage, frequency, etc
·        Have hospital staff make follow-up appointments with patient's physician and don't discharge patient until this schedule is set up
·        Use Digital technologies like Big Data & IoT to collect vitals and keep up visual as well as verbal communication with patients, especially those that are high risk for readmission.
·        Kaiser Permanente & Novartis are using Telemedicine technologies like video cameras for remote monitoring to determine what's happening to the patient after discharge
·        Piedmont Hospital in Atlanta provides home care on wheels like case management, housekeeping services, transportation to the pharmacy and physician's office         
·        Use of Data Science algorithms to predict patients with high risk of readmission
·        Walgreens launched WellTransitions program where patients receive a medication review upon admission and discharge from hospital, bedside medication delivery, medication education and counseling, and regularly scheduled follow-up support by phone and online.
·        HealthLoop is a cloud based platform that automates follow-up care keeping doctors, patients and care-givers connected between visits with clinical information that is insightful, actionable, and engaging.
·        Propeller Health, a startup company in Madison has developed an app and sensors track medication usage and then send time and location data to a smartphone
·        Mango Health for iPhone and wearables like Apple Watch makes managing your medications fun, easy, and rewarding. App feature include: dose reminders, drug interaction info, a health history, and best of all - points and rewards, just for taking your medicines.

These emerging digital tools enable health care organizations to assess and better manage who is at risk for readmission and determine the optimal course of action for the patients. 

Such tools also enable patients to live at home, in greater comfort and at lower cost, lifting the burden on themselves and their families.

Digital is helping mankind in all ways !!

Saturday, 15 October 2016

Using Data Science for Predictive Maintenance

Remember few years ago there were two recall announcements from National Highway Traffic Safety Administration for GM & Tesla – both related to problems that could cause fires. These caused tons of money to resolve.

Aerospace, Rail industry, Equipment manufacturers and Auto makers often face this challenge of ensuring maximum availability of critical assembly line systems, keeping those assets in good working order, while simultaneously minimizing the cost of maintenance and time based or count based repairs.

Identification of root causes of faults and failures must also happen without the need for a lab or testing. As more vehicles/industrial equipment and assembly robots begin to communicate their current status to a central server, detection of faults becomes more easy and practical.

Early identification of these potential issues helps organizations deploy maintenance team more cost effectively and maximize parts/equipment up-time. All the critical factors that help to predict failure, may be deeply buried in structured data like equipment year, make, model, warranty details etc and unstructured data covering millions of log entries, sensor data, error messages, odometer reading, speed, engine temperature, engine torque, acceleration and repair & maintenance reports.

Predictive maintenance, a technique to predict when an in-service machine will fail so that maintenance can be planned in advance, encompasses failure prediction, failure diagnosis, failure type classification, and recommendation of maintenance actions after failure.

Business benefits of Data Science with predictive maintenance:
  • Minimize maintenance costs - Don’t waste money through over-cautious time bound maintenance. Only repair equipment when repairs are actually needed.
  • Reduce unplanned downtime - Implement predictive maintenance to predict future equipment malfunctioning and failures and minimize the risk for unplanned disasters putting your business at risk.
  • Root cause analysis - Find causes for equipment malfunctions and work with suppliers to switch-off reasons for high failure rates. Increase return on your assets.
  • Efficient labor planning — no time wasted replacing/fixing equipment that doesn’t need it
  • Avoid warranty cost for failure recovery – thousands of recalls in case of automakers while production loss in assembly line

TrainItalia has invested 50M euros in Internet of Things project which expects to cut maintenance costs by up to 130M euros to increase train availability and customer satisfaction.

Rolls Royce is teaming up with Microsoft for Azure cloud based streaming analytics for predicting engine failures and ensuring right maintenance.

Sudden machine failures can ruin the reputation of a business resulting in potential contract penalties, and lost revenue. Data Science can help in real time and before time to save all this trouble.


Sunday, 18 September 2016

What is Cognitive Computing?


Although computers are better for data processing and making calculations, they were not able to accomplish some of the most basic human tasks, like recognizing Apple or Orange from basket of fruits, till now.

Computers can capture, move, and store the data, but they cannot understand what the data mean. Thanks to Cognitive Computing, machines are bringing human-like intelligence to a number of business applications.

Cognitive Computing is a term that IBM had coined for machines that can interact and think like humans.

In today's Digital Transformation age, various technological advancements have given machines a greater ability to understand information, to learn, to reason, and act upon it. 

Today, IBM Watson and Google DeepMind are leading the cognitive computing space.

Cognitive Computing systems may include the following components:
·      Natural Language Processing - understand meaning and context in a language, allowing deeper, more intuitive level of discovery and even interaction with information.
·      Machine Learning with Neural Networks - algorithms that help train the system to recognize images and understand speech
·        Algorithms that learn and adapt with Artificial Intelligence
·        Deep Learning – to recognize patterns
·        Image recognition – like humans but more faster
·        Reasoning and decision automation – based on limitless data
·        Emotional Intelligence

Cognitive computing can help banking and insurance companies to identify risks and frauds. It analyses information to predict weather patterns. In healthcare it is helping doctors to treat patients based on historical data.

Some of the recent examples of Cognitive Computing:
·  ANZ bank of Australia used Watson-based financial services apps to offer investment advice, by reading through thousands of investments options and suggesting best-fit based on customer specific profiles, further taking into consideration their age, life stage, financial position, and risk tolerance.
·   Geico is using Watson based cognitive computing to learn the underwriting guidelines, read the risk submissions, and effectively help underwrite
·    Brazilian bank Banco Bradesco is using Cognitive assistants at work helping build more intimate, personalized relationships
·       Out of the personal digital assistants we have Siri, Google Now & Cortana – I feel Google now is much easy and quickly adapt to your spoken language. There is a voice command for just about everything you need to do — texting, emailing, searching for directions, weather, and news. Speak it; don’t text it!

As Big Data gives the ability to store huge amounts of data, Analytics gives ability to predict what is going to happen, Cognitive gives the ability to learn from further interactions and suggest best actions.

Sunday, 7 August 2016

How to evaluate Data Science models ?

In today’s Digital age,  insights received from data science are extremely important to deliver the best customer experience. 

Data Scientists use various techniques such as Regression, SVM, Neural network, Nearest neighbor, Naive Bayes, Decision Tree and Ensemble models.

These algorithms help to identify previously unrecognized patterns and trends hidden within vast amounts of structured and unstructured information. These patterns are used to create predictive models that try to forecast future behavior.

These models have many practical business applications: predicting patients at risk, they help banks decide which customers to approve for loans, and marketers use them to determine which leads to target with campaigns.

But how to determine if the predictive models you create are accurate, meaningful representations that will prove valuable to your organization?

There are various methods used by data scientists to measure the accuracy of the model:
  • Lift Charts & Gain Charts: These are widely used in campaign targeting problems, to determine which decile can we target customers for a specific campaign. Also, it tells you how much response you can expect from the new target base.
  • ROC Curve: The ROC curve is the plot between false positive rate and True Positive rate.
  • Gini coefficient: This is the ratio of area between the ROC curve and the diagonal line & the area of the above triangle
  • Cross Validation: splitting the data into two parts, where one part is used for "training" your model, and the second part is used to make predictions. By this you can test the model on the data that was "not seen" by it previously, and check how it could possibly behave with external data.
  • Confusion Matrix: A table showing the number of predictions for each class compared to the number of instances that actually belong to each class. This is very useful to get an overview of the types of mistakes the algorithm made. This method shows accuracy, true positive, false positive, Sensitivity & specificity of the model.
  • Root Mean Squared Error: This is the average amount of error made on the test set in the units of the output variable. This measure helps you get an idea on the amount a given prediction may be wrong on average. This is most popular in regression techniques.
In general, the assessment used should be closely matching the business objectives. Using the right metric can have more influence on you model performance than the algorithm you use.


There are so many data points generated by Internet of Things, Mobiles, Social Media and all the Omni-Channels used for customer interactions. Only storing this data is useless , unless it is used by data scientists for generating insights that is used for next actions. 

Saturday, 30 July 2016

How to deploy Data Science projects?

In this Digital age today, data science has become the top skill and sexiest job of the century. 

Data science projects do not have a nice clean life-cycle with well-defined steps like software development lifecycle (SDLC), but they are non-linear, highly iterative and cyclical between the data science team and various others teams in an organization.

SAS Institute, the leader in Analytics developed its own method called SEMMA (Sample, Explore, Modify, Model & Assess) for data mining.

However, many of the companies have adopted a standard workflow of a data science called CRISP-DM (CRoss Industry Standard Process for Data Mining). It was developed by a consortium of companies like SPSS, Teradata, Daimler and NCR Corporation in 1997.

With any method the process is similar which involves following steps:
  • Business Understanding: This is the basic and first step as understanding business problem is extremely important for data scientist to move forward.
  • Data Acquisition: Based on the business problem the next step is to understand and acquire the data which is needed. Identify the sources from where it is available, who are responsible to provide that data. It can come from various data sources like customer data, demographic data, third party data, weblogs, social media data, streaming data like sensor data, audio or video data. Main challenge is to decide whether data is up-to-date and clean for model consumption. With Internet of Things in full swing, data acquisition into Big Data platform is important step.
  • Data Preparation: This is also called as data wrangling phase which takes almost 60% of overall project time. Collected data has to be formatted, treated for any missing values, any abnormalities or seasonality from the data and make it ready for model consumption.
  • Modelling: This is the core activity of a data science project that requires writing, running and refining the programs to analyse and derive meaningful business insights from data. Often open sources tools like R, Python and commercial tools like SAS, IBM SPSS are used to create the statistical models. Various machine learning techniques are applied to data based on the business problem.
  • Evaluation: There are several methods to compare the developed models and then use the best model for deployments. Typical comparison methods are AUC – area under curve, Confusion matrix, Gain/Life charts, Root Mean Squared Error etc.
  • Deployment: Once the most suitable model is identified above, it is further tested with live data and then deployed into production environment.

There are further steps as well such as monitoring the live model performance, observe any degradation and new models are developed which are again compared with live model.

Data Science has evolved beyond normal predictive modeling into recommendation engines, text mining, deep learning, Artificial Intelligence. The foundation still remains the same of data gathering, data cleaning and then applying various algorithms.

360TotalSecurity WW