Showing posts with label Machine Learning Steps. Show all posts
Showing posts with label Machine Learning Steps. Show all posts

Tuesday, May 23, 2017

Published 14:21 by with 0 comment

Predicting the onset of Diabetes - II

We concluded our previous attempt to predict the onset of diabetes with a few questions. Lets try to answer the questions in this post.

Recap


The model that we built previously had an accuracy of 77.73% - with a caveat that we used the same data for training the model and do the prediction. And we had a few questions.

Is there a better model to predict? If so, how do we build and measure it?

Yes. There can be many models that we can build on the data set. Lets try to build one more and try to measure it.

Are there any disadvantages in building the model and testing the model on the same data set?

Certainly, its like looking at the questions that ll be asked before taking an examination. Any average person can memorize the answers and can increase the chance of scoring better marks in the exam. Same is the case with the model that we build. It stands a better chance and accuracy if we work on the same data set to predict.

A different model - Decision Trees


Lets try to build a decision tree (without going through the math behind them) on the data set . We ll use the model to make prediction (as we did earlier). We ll also measure the accuracy of the model.

Here we ll build a Decision Tree using Recursive Partitioning to classify the data. The technique is to classify members of the population by splitting it into sub-populations based on several dichotomous independent variables. The process is termed recursive because each sub-population may in turn be split an indefinite number of times until the splitting process terminates after a particular stopping criterion is reached.

A major advantage with this method is this is very simple and more intuitive.

rpart.fit <- rpart(Outcome ~ .,data=pimaimpute)
prp(rpart.fit, faclen = 0, box.palette = "auto", branch.type = 5, cex = 1)

Doing that produces this beautiful Decision Tree.



How to read this?


The model splits the data set into two sections based on Glucose Level and then on Age (for the section where Glucose level is <128) and on BMI and so on until it finds stopping criteria.

As we have build the model, lets use this model for predict (on the same data set).

rpart.pred <- predict(rpart.fit, data=pimaimpute, type = "class")


Lets predict the accuracy of this model.





Confusion Matrix Visualized



Closing Notes

This model has got an accuracy of 83.98% (versus our Linear Model's accuracy of 77.73%).

One important question that we ll address in the next post is - 
Are there any disadvantages in building the model and testing the model on the same data set?


P.S.: The complete code that was used for this article is here.




Read More
      edit

Wednesday, January 18, 2017

Published 13:32 by with 5 comments

Introduction to Machine Learning

What is Machine Learning?

Machine Learning is a branch of Computer Science that deals with making a machine (i.e. computer)  learn without explicitly being programmed. Essentially, it is a method of teaching computers to make and improve predictions or patterns based on data.

A common example used to explain Machine Learning is ‘Digit Recognition’ - where a machine is taught to understand how different digits look like (10 of them - zero included) using images that contain handwritten single digit. The algorithm is then made to ‘recognize’ new set of images of handwritten digit.

Machine Learning can be broadly classified into the following categories

Supervised Learning

In Supervised Learning, the machine is taught using example inputs and their corresponding outputs. The objective of the machine is to learn the “rule” that derives the output (for the given set of input).

Digit Recognition is an example for Supervised Learning as the labelled example data (handwritten image and the corresponding digit) is used to train the machine. The objective of this method is to Predict the output given an unknown/new input data.

Unsupervised Learning

In Unsupervised Learning, the machine is *not* provided with labelled examples and the objective is to identify hidden patterns, outlier detection, clustering, etc.

An everyday example for Unsupervised Learning is Google News - news articles of same/similar content are sourced from various sites and grouped together.

Reinforcement Learning

Reinforcement Learning is where a machine operates in a dynamic environment and makes decisions and it is penalised or awarded for making such decision periodically. The objective of this method is to maximize the performance or efficiency of the whole system.

Self driving cars or computers playing games against human players are notable examples of Reinforcement Learning.

Footnote: As humans, we inherently perform all the three ways of learning and here are the instances of them

  • Humans learning to identify shapes, colors, alphabets, etc are Supervised as we are taught with prior examples


  • Doctors diagnosing medical issues, identification of bugs in software, etc are examples of humans doing anomaly detection. Grouping items based on similarities is another example of humans doing clustering. For the same sample items, the groups/clusters can be very different for different humans


  • The process of humans learning to speak a language is an example of Reinforcement Learning where based on the reaction or response of the other person(s) involved (and over a period of time) the 
    mastery on the language improves



Why Machine Learning?

Because of its application in the real world.

Following are some of the use cases for Machine Learning

  • Fraud Detection
  • Sentiment Analysis
  • Recommendation Engine
  • Self-Driving Cars

Why now?

As collection and processing of data becomes cheaper, applying Machine Learning becomes more and more practical and effective. For some organizations, Machine Learning applications are the game changer. Examples are Amazon, Google, Netflix, etc

Steps involved in Machine Learning

  • Collection of Data
  • Preparation of Data
  • Building (or Training) a Model
  • Evaluating the Model
  • Improving the Performance or efficiency




Read More
      edit