In this article we are discussing some of the most important statistical concepts that you should know, if you have decided to go ahead with data science. We use these concepts in data science, for analyzing the data, getting insights from it, and using those insights to create a model that would solve our problem. This article ‘Statistics for data science – Descriptive Statistics’, will be helpful to beginners.
Take a look at our article on introduction to Data Science for getting a better understanding of Data Science and also check out our other Data Science articles for understanding the major tools you should be familiar with for becoming a Data Scientist.
Lets discuss each concepts one by one….
The two main branches of statistics are Descriptive statistics and Inferential statistics.
Let’s discuss descriptive statistics first.
But, before that we must know the most important terms in statistics, Population and Sample.
Learn about the providers of online masters in data science by clicking here
Population
A Population is a group of values of anything, which is under statistical study. A population can be a list of heights of whole people of a country, list of weights of all products produced in a shift in a factory, list of marks of whole students in a school, list of values obtained after continuously tossing a coin for some time, etc.
I think now you got an idea of Population.
Sample
A sample is a subset of population, which contain values that are collected through sampling techniques from population. List of heights of 100 peoples from a country, list of weight 10 products sampled from a 1000 products produced in a factory, etc. are examples of sample.
In descriptive statistics we analyze the sample data and in inferential statistics we will try to relate the results of sample data analysis, to the population.
Descriptive Statistics

In descriptive statistics, we use certain mathematical tools to describe the data in the form of numbers or graphs, or tables. For example, if we give a statistician a set of data about the patient’s detail of a hospital, and tell him to find out interesting information from it, then what he needs to do is descriptive statistics. He would learn about the data and come up with interesting information that he derived from the data. The results of this would be information like 50% of the patients has shown certain symptoms, out of 180 patients 70 of them had spent more than a week in the hospital, etc.
Following are some of the important concepts in descriptive statistics which will help to extract required information from the sample data and let’s discuss each in detail
- The measure of Central Tendency
- Measure of Spread
- Covariance & Correlation
1. The measure of Central Tendency
The measure of central tendency is a value that helps us identify the central position of the data.
Mean
Mean or the average of a set of numbers is the result of summing up all the values in a dataset and dividing it by the total number of values. Taking the mean of a dataset would help us to understand the performance of a system in general. For example, the average of a batsman in cricket can be used to understand how well his performance is. If he often scores more 100s than 0s then his average would be getting higher and higher. Also, it can be used to calculate the popularity of a product by taking the average buys in a season or a region.
The formula for calculating the mean is simple:

Here x1, x2, etc are the values in the data set and n is the no of values in the data set.
Continue reading with KIE Premium
Unlock this article and all KIE premium articles.
Now or Never
We’ve got your back on your manufacturing journey — Stay in touch
Follow us for step-by-step guidance, templates, and insights that save time and reduce mistakes.
Know Industrial Engineering Platform – Helping manufacturing industry professionals worldwide since 2019