Showing posts with label Machine Learning. Show all posts
Showing posts with label Machine Learning. Show all posts

Tuesday, May 9, 2017

Published 10:03 by with 0 comment

Predicting the onset of Diabetes

One of the applications of Machine Learning is to predict likelihood of occurrence of an event from the past data. With our learning from exploring data to testing hypothesis, lets see if we can put together the learning to predict likelihood of onset of diabetes from diagnostic measures.

Data Source

We would be working on Pima Indians Diabetes Data Set from UCI Machine Learning Repository.

Data Set Information

This data set is originally from the National Institute of Diabetes and Digestive and Kidney Diseases. This data set contains diagnostic measures for female patients of at least 21 years of age who belong to Pima Indian Heritage.

Please note: This is a labelled data set.

Number of Records                  :        768
Number of Attributes               :            8

Attribute Information

  1. Number of times pregnant 
  2. Plasma glucose concentration a 2 hours in an oral glucose tolerance test 
  3. Diastolic blood pressure (mm Hg) 
  4. Triceps skin fold thickness (mm) 
  5. 2-Hour serum insulin (mu U/ml) 
  6. Body mass index (weight in kg/(height in m)^2) 
  7. Diabetes pedigree function 
  8. Age (years) 
  9. Class variable (0 or 1)

Taking a quick glance at the data


Please note: The column names are shortened for convenience

The action begins

Now we have sourced the data, using our learning lets start exploring the data.


Summarizing the data set is a simple but powerful step to start understanding the data. As we can see here, there are a few measures with zero values which are biologically impossible (example would be blood pressure). Now this leads us to the next action of handling the missing/invalid data. And its also a good idea to visualize the data.

A Quick Look

As the visuals suggest, Glucose, Blood Pressure, Skin Thickness, Body Mass Index are normally distributed while are Pregnancies, Insulin, DBF, Age exponentially distributed.





Also please have in mind that this visuals are rendered without handling for the missing values. The next step is to handle or treat the data for missing values.

Data Preparation


Missing Values

There are zero value entries for the following measures. We ll impute them with mean value for corresponding measure.

  • Glucose
  • Blood Pressure
  • Skin Thickness
  • Insulin
  • Body Mass Index

Pregnancies

Quick tabulation on this column shows that minimum value being zero (permissible) and maximum value being 17 (still theoretically possible even though this is an outlier).


Just to cross-check the feasibility lets look at the record and check the Age (which is 47). This suggests that this can be possible (at least theoretically).


As the values of Pregnancies seem to be biologically valid, no data handling is needed.

Visualizing Missing Values

Before Imputing

After Imputing
With cleaner data, we are geared up for building models and do the prediction.

Building a model

We try to build a logistic regression model on the data as the Outcome is categorical.


This is easy, isn't it?

The summary gives us interesting information (listed below).

Out of eight features that we have, four stand out to be significant on influencing the person to be diabetic or not (Outcome)

  • Pregnancies
  • Glucose
  • BMI
  • DPF 

Lets looks at the Odds Ratio to see to what extent these features are influencing the Outcome.

Odds Ratio


Each pregnancy is associated with an odds increase of 13.31%
Each unit increase of plasma glucose concentration is associated with an odds increase of 3.80%
An unit increase in BMI increases the odds by 9.76%
For every unit increment in DPF, the odds are up by 137.76%

Prediction


Here comes the interesting part. Having built a model that can predict the likelihood of onset of diabetes, lets try to do some prediction and see how well does this model work.

Here is the confusion matrix from prediction.

Confusion Matrix Visualized


Closing Notes


This model has got an accuracy of 77.73% but the point to be noted here is the model used the same data to do the prediction.

Now a few questions that linger our mind are
Are there any disadvantages in building the model and testing the model on the same data set?
How can we apply this model in a new data with confidence?
How can I ensure this is a good model to predict?
Can I build a better model?
and so on.

Lets try to answer these questions (and more) in the subsequent posts.


P.S.: The complete code that was used for this article is here.

Read More
      edit

Tuesday, January 31, 2017

Published 18:13 by with 2 comments

The ABCs of Machine Learning

It would be a good idea to spend some time reading the basics and this page would be dedicated for that.

Why do we care about variable type (in the context of Machine Learning)?

Understanding the variable types gives us the power to treat them appropriately. For instance, it would help us in avoiding mistakes like doing an average on postal codes or taking ratio of two pH values. It also helps us choosing appropriate operation on the variable based on the context. For example, in a psychological study of perception, different colors would be regarded as nominal. In a physics study, color is quantified by wavelength, so color would be considered a ratio variable.

Feature

Feature (aka Variable) is a measurable property or attribute of an observation. An example would be while studying about performance of cars the possible features would be
  • Number of Cylinders
  • Miles Per Gallon (or Kilometers Per Liter)
  • Horse Power
  • Weight
  • Number of Gears
  • Type of Transmission
Feature is also referred as 'Explanatory Variable' or 'Independent Variable'

Label

Label represents the outcome of the or the output whose variation is being studied. An example would be the learning problem shown below. Here the examples are labeled '1' and '0'. In this case, animals are marked with '1' and others with '0'. screen-shot-2017-01-25-at-12-09-44-pm
Label is also referred as 'Explained Variable' or 'Dependent Variable' The dependent variable responds to the independent variable and for this reason it is called 'dependent' variable.

Dataset

Dataset (or Data Set) is simply a collection of observations. Typically its the data collection in rows and columns where columns correspond to feature/label and rows correspond to the observations. Most popular dataset format used is spreadsheet which is powerful for quick analysis.

Data (Variable) Type

Data is often classified as below based on the nature of the data.
Please Note: This is different the 'datatype' defined in database realm.

Quantitative Variable

Quantitative Variable is expressed in numerical form and therefore arithmetic operations can be performed on them. Quantitative Variable can be further classified as follows
  • Continuous Variable
A continuous variable can take values between two numbers. This variable can take infinitely many values.
Example: Time taken by top five athletes to complete 100m in Rio Olympics: 9.81, 9.89, 9.91, 9.93, 9.94.
    • Interval Variable
      Interval Variables take numeric values and they can be measured along continuum. The intervals between the values of the interval variable are equally spaced.
      Example: Temperature measured in degrees Celsius or Fahrenheit. The difference between 20C and 30C is the same as 30C to 40C
    • Ratio Variable
      Ratio Variables are Interval Variables but with the different that they posses the clear definition of zero (0) which indicates that there is none of that variable.
      Example: Temperature measures in Kelvin as 0 Kelvin (also called as absolute zero) indicates that there is no temperature. And for the very same reason temperature measured in Celsius and Fahrenheit are NOT Ratio Variables. Other examples include height, mass, distance, etc.
  • Discrete Variable
A discrete variable does not admit intermediate values between two specific numbers. It is represented by whole integer values.
Example: Total medals by the USA at last three summer Olympics: 121 (Rio 2016), 103 (London 2012), 110 (Beijing 2008).

Qualitative Variable

Qualitative Variable takes non-numeric value. It describes data that fit into categories. This is also referred as Categorical Variable. Qualitative Variable can be further classified as follows
  • Nominal Variable
A nominal variable is one that has two or more categories, but there is no intrinsic ordering to the categories.
Example: The blood type of a person has multiple categories A, B, AB or O but there is no intrinsic order.
    • Dichotomous (Binary) Variable
      Dichotomous variables are nominal variables which have only two categories or levels. Example: If we ask a person if s/he owns a car. The response can be either 'Yes' or 'No'
  • Ordinal Variable
An ordinal variable is similar to a nominal variable but the difference being there is a clear ordering (or ranking) of the variables.
Example: Clothing Size having values like S, M, L, XL where there is an order (S < M)

In a nut shell

screen-shot-2017-01-29-at-11-48-40-am

Read More
      edit

Wednesday, January 18, 2017

Published 13:32 by with 5 comments

Introduction to Machine Learning

What is Machine Learning?

Machine Learning is a branch of Computer Science that deals with making a machine (i.e. computer)  learn without explicitly being programmed. Essentially, it is a method of teaching computers to make and improve predictions or patterns based on data.

A common example used to explain Machine Learning is ‘Digit Recognition’ - where a machine is taught to understand how different digits look like (10 of them - zero included) using images that contain handwritten single digit. The algorithm is then made to ‘recognize’ new set of images of handwritten digit.

Machine Learning can be broadly classified into the following categories

Supervised Learning

In Supervised Learning, the machine is taught using example inputs and their corresponding outputs. The objective of the machine is to learn the “rule” that derives the output (for the given set of input).

Digit Recognition is an example for Supervised Learning as the labelled example data (handwritten image and the corresponding digit) is used to train the machine. The objective of this method is to Predict the output given an unknown/new input data.

Unsupervised Learning

In Unsupervised Learning, the machine is *not* provided with labelled examples and the objective is to identify hidden patterns, outlier detection, clustering, etc.

An everyday example for Unsupervised Learning is Google News - news articles of same/similar content are sourced from various sites and grouped together.

Reinforcement Learning

Reinforcement Learning is where a machine operates in a dynamic environment and makes decisions and it is penalised or awarded for making such decision periodically. The objective of this method is to maximize the performance or efficiency of the whole system.

Self driving cars or computers playing games against human players are notable examples of Reinforcement Learning.

Footnote: As humans, we inherently perform all the three ways of learning and here are the instances of them

  • Humans learning to identify shapes, colors, alphabets, etc are Supervised as we are taught with prior examples


  • Doctors diagnosing medical issues, identification of bugs in software, etc are examples of humans doing anomaly detection. Grouping items based on similarities is another example of humans doing clustering. For the same sample items, the groups/clusters can be very different for different humans


  • The process of humans learning to speak a language is an example of Reinforcement Learning where based on the reaction or response of the other person(s) involved (and over a period of time) the 
    mastery on the language improves



Why Machine Learning?

Because of its application in the real world.

Following are some of the use cases for Machine Learning

  • Fraud Detection
  • Sentiment Analysis
  • Recommendation Engine
  • Self-Driving Cars

Why now?

As collection and processing of data becomes cheaper, applying Machine Learning becomes more and more practical and effective. For some organizations, Machine Learning applications are the game changer. Examples are Amazon, Google, Netflix, etc

Steps involved in Machine Learning

  • Collection of Data
  • Preparation of Data
  • Building (or Training) a Model
  • Evaluating the Model
  • Improving the Performance or efficiency




Read More
      edit