Showing posts with label Data Exploration. Show all posts
Showing posts with label Data Exploration. Show all posts

Tuesday, April 18, 2017

Published 11:50 by with 0 comment

Lets Explore Data - II

In the previous post we understood the "mtcars" data set and did some univariate analysis to understand the features from the data set.

The objective of this post is to understand how multiple features (aka variables) are related. For example, one might be interested in understanding how weight of the car affects its fuel efficiency.

Correlating Weight of the Car to its Fuel Efficiency

Here is the visual representation of the relationship between weight of the car to its fuel efficiency (mpg).



From this plot, its evident that weight and mpg are inversely related which means lower the weight of the car more the miles per gallon is (i.e. more is the fuel efficiency of the car).

Testing Statistically

To statistically validate this relation and to ensure that this is not by chance, we ll perform correlation test. This can be achieved by cor.test() command in R.


Couple of things to look at the results from Correlation Test is p-value and correlation coefficient.

The smaller the p-value, the more significant the correlation, so here we can be very confident that a correlation exists.

Correlation Coefficient is a number between +1 and −1 calculated so as to represent the linear interdependence of two variables. +1 represents a perfectly linear positive relationship while -1 represents a perfectly linear negative relationship. 0 indicates that the two variables are not correlated.

Inference 

p-value from the above correlation test is a extremely small number which suggests a strong correlation exists.

Correlation Coefficient being -0.87 suggests that the relationship between these two variables are negative (i.e. they are inversely related).

Both the p-value and for suggest that the variation that is observed in the data set is not by chance and there is a interdependence between the variables.

This is a powerful way for understanding correlation between two variables. Now, can we extend this to all the numeric variables in mtcars data set and does it still make sense?

Lets see.

Extension


Well, this chart gives us the relationship between multiple variables but its has a few problems such as it is cluttered and also its redundant (the plots are mirrored across the diagonal).

Lets make it pretty by tweaking this plot a little bit.



This looks much clearer and also gives us the correlation coefficient in the upper half of the plot with text size proportional to the value to improve readability.

We observe stronger negative correlation between the following variable pairs

  • mpg & wt
  • mpg & disp
  • mpg & hp
  • drat & disp

and stronger positive correlation between the following variable pairs

  • disp & wt
  • disp & hp


A Better Visualization? May be!




This visualization gives a sense of which variables are in relation along with their type or relationship (as indicated by the color/correlation coefficient value). This can be an effective tool to communicate the correlation between variables. In this example, the reader can understand quickly that wt & disp are positively correlated while wt & mpg are negatively correlated.

Now a question hits our mind - if the reader is interested in building an understanding of relation between multiple variables (as opposed to just two variables) is there a way to do it?

Visualize correlation of many variables at once

Parallel Coordinate Visualization is the answer.

Parallel Coordinate is an effective visualization technique to visualizing high-dimensional geometry and analyzing multivariate data.




Note: The number of lines correspond in the plot above corresponds to number of cars in each class (highlighted by the color).

This visualization technique gives an awesome way to understand relationship between multiple variables.

With cars categorized based on number of cylinders (4-cyl cars represented by Blue lines, 6-cyl cars by Red lines and 8-cyl cars by Dark Green lines), we are able to explain the relationship between variables much quicker.

For instance - say if you follow the blue lines, its evident that mpg enjoys an inverse relationship with disp, hp & wt. This not only helps us to understand the relationship between different variables/features but also in understanding few interesting facts like cars with four cylinders are fuel efficient in general and higher horse power is obtained by eight-cylinder cars.

We are now equipped with techniques to explore and understand data.

In the next post, lets see if we can use the data to build a model to predict missing variables from similar data.


Read More
      edit

Saturday, February 25, 2017

Published 22:16 by with 0 comment

Lets Explore Data - I(a)

I thought it would be a good idea to visually explain the inferences from the previous post because a picture is worth a thousand words. Visualization is one of the most effective ways to communicate the findings from Data Analytics.

This post would be a shorter one as we are going to visually represent the analysis we did in the previous post.

I am using R to explore the data. Please feel free to use tool of your choice for data exploration.

Descriptive Statistics on "mpg"

Summary Statistics - Miles Per Gallon


The above "box plot" would give the five number statistics (minimum, lower quartile, median, upper quartile, maximum) and Mean.

Distribution of cars based on number of cylinders




Distribution of cars based on number of forward gears




Distribution of cars based on transmission and on engine type




We can also combine different aspects/parameters and chart them together


In the next post, we ll see about multivariate data analysis.
Read More
      edit

Tuesday, February 14, 2017

Published 13:12 by with 0 comment

Lets Explore Data - I

Having acquainted with The ABCs of Machine Learning, its time to do some data exploration to understand the nature of data to graduate to the next level.

After we acquire data, the immediate next step would be to make sense of the data, quickly. This step would be of paramount importance as doing this right would help in saving lots of time while we build model/draw conclusion.

Data Set

The data that we are going to explore is sourced from here.

Lets get started.

As defined, the data was extracted from the 1974 Motor Trend US magazine, and comprises fuel consumption and 10 aspects of automobile design and performance for 32 automobiles (1973–74 models).

I am using R to explore the data. Please feel free to use tool of your choice for data exploration.

Quick Tip: As the data is available in csv format, Microsoft Excel would also help us in doing the data exploration.

A quick look at (part) of what is inside the data set.

screen-shot-2017-02-08-at-3-54-40-pm

Here is the metadata for the data set.

Variable Name Description Variable Type
mpg Miles/(US) gallon Quantitative (Continuous)
cyl Number of cylinders Qualitative (Nominal)*
disp Displacement (cu.in.) Quantitative (Continuous)
hp Gross horsepower Quantitative (Continuous)
drat Rear axle ratio Quantitative (Continuous)
wt Weight (1000 lbs) Quantitative (Continuous)
qsec 1/4 mile time Quantitative (Continuous)
vs V Engine/S Engine (0=V, 1=S) Qualitative (Nominal)
am Transmission (0 = automatic, 1 = manual) Qualitative (Nominal)
gear Number of forward gears Qualitative (Nominal)
carb Number of carburetors Qualitative (Nominal)


* Depending on the analysis that we do using "cyl" this can be classified as Quantitative (Discrete) too.

I have also mapped the variable name and description with the variable type. Please have a quick look to review the same. If there are any questions, please ask them in the comments section.

After we skim through the metadata and understand the variable properties, next step is to get descriptive statistics on the data set. Descriptive Statistics include frequencies (counts), ranks, measures of central tendency (e.g., mean, median and mode), and measures of variability (e.g., range and standard deviation). These help in developing good and quick understanding of the data.

Remember we would want to do numeric operations only on the Quantitative Variables even though some of the Qualitative Variables that have numeric values (e.g., vs, am, etc)

Descriptive Statistics on "mpg"

Average Miles Per Gallon







Median Miles Per Gallon






Miles Per Gallon Range






screen-shot-2017-02-08-at-5-35-40-pm
Summary Statistics - this captures all of that we calculated earlier at one shot












Inference on mpg

The mpg (for all the vehicles in the data set) ranges from 10.40 miles per gallon to 33.90 miles per gallon with an average of 20.09 miles per gallon and median 19.20 miles per gallon.

Please note: summary() in R gives another interesting measure "IQR" (interquartile range) which we ll read about in a future post.

Extending the Descriptive Statistics to the other variables of interest (mpg, weight, gear, am, vs, cyl) in the data set yields us the following.

Inference on mtcars data set

  • Total number of cars in the data set - 32
  • Distribution of cars based on number of cylinders is as below
    Number of Cylinders Number of Cars
    Four 11
    Six 7
    Eight 14

  • Distribution of cars based on number of forward gears is as below
    Number of Forward Gears Number of Cars
    Three 15
    Four 12
    Five 5

  • Distribution of cars based on transmission is as below
    Transmission Type Number of Cars
    Automatic 19
    Manual 13

  • Distribution of cars based on engine type is as below
    Engine Type Number of Cars
    V-Engine 18
    Straight Engine 14

  • Miles Per Gallon ranges between 10.40 mpg to 33.90 mpg
  • Weight of the cars range between 1513 lbs to 5424 lbs

Conclusion

We did univariate analysis on the variables in the "mtcars" data set to understand more about the content of the data set. Descriptive Statistics helps us get a good understanding about the data, quickly. We also reinforced our learning about the variable types and we are able to apply different operations for different variable type (e.g., we did not do an average on a Qualitative Variable).

We would continue to do more analysis (involving multiple variables) to make more sense on these data set in the next post.

Gist

A short hand way in R to get descriptive statistics of a data set.

screen-shot-2017-02-08-at-8-43-09-pm
Summary Statistics on "mtcars" data set


Use summary() command with caution as it treats all the variables that appear to be number as Quantitative irrespective of whether they are Quantitative or not.

Based on the type of the variable, summary() function gives different output.
  • Gives the range, quartiles, median, and mean & number of missing values (if any) for Quantitative Variables
  • Gives a table with frequencies & number of missing values (if any) for Categorical Variables

How to handle Qualitative Variables

The results of the summary() has given Quantitative statistics on Qualitative data - how to handle this?

Change the type of the variable (as a categorical) on the fly and we can compute frequency on the Qualitative variable.


 




Please Note: The above is an example of changing the type of one variable from the data set. We can also do this treatment as we do preprocessing before we do the descriptive statistics on the data set.

Note #1 : Based on the feedback (in the comments section), I would plan to write a post detailing on these measures (such as mean, range, etc) and explain the significance and on which variable types/when/where these should be used.

Note #2 : I also think of a post on "getting started with R" - if there are takers. Please leave comments if you are interested.

Note #3: Visualizing the data would be another powerful and significant method in exploring data. Please watch out for a post on the same.
Read More
      edit