Statistics is the branch of mathematics that helps us collect, organize, analyze, interpret, and present data. Whether you are studying mathematics, science, economics, business, or simply trying to understand information in everyday life, basic statistics provides useful tools for making sense of numbers.
At first, statistical formulas may look complicated because they use symbols and mathematical notation. However, most basic formulas are based on simple ideas. Once you understand what each quantity represents and when to use a particular formula, statistics becomes much easier to learn.
This guide explains the most important basic statistics formulas for beginners, including mean, median, mode, range, variance, standard deviation, quartiles, percentages, and other commonly used measures. Each formula is explained in a simple way with examples so that you can understand how it works and when to apply it.
What Are Statistics Formulas?
Statistics formulas are mathematical rules used to calculate useful information from a collection of data. A set of numbers is called a data set, and each individual value in the data set is called an observation or data value.
For example, consider the following data:
5, 7, 8, 10, 10
Using statistical formulas, we can find its mean, median, mode, range, and other characteristics.
Statistics formulas help answer questions such as:
What is the average value?
Which value occurs most often?
How spread out are the data?
What is the difference between the largest and smallest values?
How much does each value differ from the average?
What proportion or percentage of observations belongs to a particular group?
Important Symbols Used in Statistics
Before learning the formulas, it is useful to understand some common statistical symbols.
x = an individual data value
n = number of observations in a data set
Σ = summation of values
x̄ = arithmetic mean of the data
μ = population mean
σ = population standard deviation
s = sample standard deviation
N = population size
Q₁ = first quartile
Q₂ = second quartile or median
Q₃ = third quartile
The exact symbols can vary slightly between textbooks, but the underlying concepts remain the same.
Arithmetic Mean Formula
The arithmetic mean, commonly called the average, is one of the most frequently used statistical measures.
To calculate the mean, add all the data values and divide the total by the number of observations.
Formula
Mean = Sum of all observations ÷ Number of observations
Using symbols:
x̄ = Σx / n
Example
Suppose the data is:
4, 6, 8, 10, 12
The sum is:
4 + 6 + 8 + 10 + 12 = 40
There are 5 observations.
Therefore:
Mean = 40 ÷ 5 = 8
So, the mean of the data set is 8.
The mean is useful when you want a single value that represents the central level of a data set.
Weighted Mean Formula
Sometimes every value does not have the same importance. In such cases, a weighted mean can be used.
Formula
Weighted Mean = Σ(wx) / Σw
where:
x = value
w = weight assigned to the value
Example
Suppose a student’s marks are:
Test = 70, weight = 2
Assignment = 80, weight = 1
Exam = 90, weight = 3
Then:
Weighted Mean = (70 × 2 + 80 × 1 + 90 × 3) / (2 + 1 + 3)
= (140 + 80 + 270) / 6
= 490 / 6
≈ 81.67
The weighted mean is approximately 81.67.
Median Formula
The median is the middle value when the observations are arranged in ascending or descending order.
For example:
3, 5, 7, 9, 11
The middle value is 7, so the median is 7.
Median Formula for an Odd Number of Observations
If there are n observations and n is odd:
Median position = (n + 1) / 2
Example
For 7 observations:
2, 4, 5, 7, 9, 12, 15
Median position:
(7 + 1) / 2 = 4
The fourth value is 7, so the median is 7.
Median Formula for an Even Number of Observations
When the number of observations is even, the median is the average of the two middle values.
Example
Consider:
2, 4, 6, 8, 10, 12
There are 6 observations. The two middle values are 6 and 8.
Median = (6 + 8) / 2 = 7
Therefore, the median is 7.
The median is particularly useful when a data set contains unusually large or small values because it is less affected by extreme observations than the mean.
Mode Formula
The mode is the value that occurs most frequently in a data set.
For example:
2, 3, 3, 4, 5, 3, 6
The number 3 occurs three times, more often than any other value.
Therefore:
Mode = 3
A data set can have:
One mode, called unimodal
Two modes, called bimodal
More than two modes, called multimodal
No mode when no value occurs more frequently than the others
Unlike the mean and median, the mode does not require a mathematical calculation. You simply identify the most frequently occurring value.
Range Formula
The range measures the difference between the largest and smallest values in a data set.
Formula
Range = Maximum value − Minimum value
Example
Consider:
5, 8, 12, 15, 20
Maximum value = 20
Minimum value = 5
Therefore:
Range = 20 − 5 = 15
The range is 15.
The range is simple to calculate, but it depends only on two observations: the maximum and minimum values. Therefore, it does not provide as much information about the overall spread as variance or standard deviation.
Variance Formula
Variance measures how far data values are spread around their mean.
A small variance means the observations tend to be close to the mean. A large variance means they are more widely spread out.
There are separate formulas for population and sample variance.
Population Variance Formula
σ² = Σ(x − μ)² / N
where:
σ² = population variance
x = individual observation
μ = population mean
N = population size
Sample Variance Formula
s² = Σ(x − x̄)² / (n − 1)
where:
s² = sample variance
x̄ = sample mean
n = sample size
The difference between dividing by N and n − 1 is important. Population variance is used when the entire population is being analyzed, while sample variance is commonly used when a sample is used to estimate characteristics of a larger population.
Standard Deviation Formula
The standard deviation is one of the most important measures of variability in statistics. It tells us how much the observations typically vary around their mean.
Standard deviation is the square root of variance.
Population Standard Deviation
σ = √[Σ(x − μ)² / N]
Sample Standard Deviation
s = √[Σ(x − x̄)² / (n − 1)]
Example
Suppose a data set has a mean of 10. If the observations are generally close to 10, the standard deviation will be small. If many observations are far away from 10, the standard deviation will be larger.
Because standard deviation is expressed in the same units as the original data, it is often easier to interpret than variance.
Quartile Formulas
Quartiles divide an ordered data set into four parts.
The three main quartiles are:
Q₁ = first quartile
Q₂ = second quartile or median
Q₃ = third quartile
Q₁ represents the point below which approximately 25% of the observations fall, while Q₃ represents the point below which approximately 75% of the observations fall.
For a simple position-based method:
First Quartile
Q₁ position = (n + 1) / 4
Second Quartile
Q₂ position = (n + 1) / 2
Third Quartile
Q₃ position = 3(n + 1) / 4
Different statistical methods may use slightly different conventions for calculating quartiles, especially when the position falls between observations.
Interquartile Range Formula
The interquartile range, or IQR, measures the spread of the middle 50% of the data.
Formula
IQR = Q₃ − Q₁
Example
If:
Q₁ = 20
and
Q₃ = 35
then:
IQR = 35 − 20 = 15
The interquartile range is 15.
The IQR is useful because it is less affected by extreme values than the range.
Percentage Formula
Percentages are commonly used in statistics to describe part of a whole.
Formula
Percentage = (Part / Whole) × 100
Example
If 25 students out of 40 students passed an examination:
Percentage = (25 / 40) × 100
= 62.5%
Therefore, 62.5% of the students passed.
Percentage Change Formula
Percentage change measures how much a value increases or decreases compared with its original value.
Formula
Percentage Change = [(New Value − Original Value) / Original Value] × 100
Example
Suppose a quantity increases from 50 to 60.
Percentage Change = [(60 − 50) / 50] × 100
= 20%
Therefore, the value increased by 20%.
If the result is negative, it represents a percentage decrease.
Frequency Formula
Frequency tells us how many times a particular value or category occurs.
For example, in the data:
2, 3, 3, 4, 4, 4, 5
the frequency of 4 is 3 because it occurs three times.
There is no special calculation required for simple frequency. It is obtained by counting the occurrences of each value.
Relative Frequency Formula
Relative frequency expresses the frequency of an observation or category as a proportion of the total number of observations.
Formula
Relative Frequency = Frequency / Total Frequency
To express it as a percentage:
Relative Frequency (%) = (Frequency / Total Frequency) × 100
Example
If 15 out of 60 observations belong to a particular category:
Relative Frequency = 15 / 60 = 0.25
As a percentage:
0.25 × 100 = 25%
Therefore, the category represents 25% of the observations.
Probability Formula
Probability is closely related to statistics and is used to describe the likelihood of an event.
For equally likely outcomes:
Formula
P(E) = Number of favorable outcomes / Total number of possible outcomes
Example
For a fair six-sided die, the probability of rolling a 4 is:
P(4) = 1 / 6
The probability can be expressed as a fraction, decimal, or percentage.
Z-Score Formula
A z-score indicates how many standard deviations an observation is above or below the mean.
Formula
z = (x − μ) / σ
For sample-based calculations, the corresponding mean and standard deviation may be used.
Example
Suppose:
x = 70
μ = 60
σ = 5
Then:
z = (70 − 60) / 5
= 2
The z-score is 2, meaning the observation is two standard deviations above the mean.
Mean Absolute Deviation Formula
The mean absolute deviation, or MAD, measures the average distance between each observation and the mean.
Formula
MAD = Σ|x − x̄| / n
The absolute value is used so that negative and positive deviations do not cancel each other out.
Example
Suppose the mean of a data set is 10. For every observation, calculate its distance from 10, ignore the sign, add the distances, and divide by the number of observations.
This gives the average absolute distance from the mean.
Coefficient of Variation Formula
The coefficient of variation, or CV, compares the standard deviation with the mean.
Formula
CV = (Standard Deviation / Mean) × 100
For a population:
CV = (σ / μ) × 100
For a sample:
CV = (s / x̄) × 100
The coefficient of variation is often useful when comparing variability between data sets that have different means or scales.
Summary of Basic Statistics Formulas
The following formulas are among the most useful for beginners:
| Statistical Measure | Formula | ||
|---|---|---|---|
| Mean | x̄ = Σx / n | ||
| Weighted Mean | Σ(wx) / Σw | ||
| Median Position for Odd n | (n + 1) / 2 | ||
| Range | Maximum − Minimum | ||
| Population Variance | σ² = Σ(x − μ)² / N | ||
| Sample Variance | s² = Σ(x − x̄)² / (n − 1) | ||
| Population Standard Deviation | σ = √[Σ(x − μ)² / N] | ||
| Sample Standard Deviation | s = √[Σ(x − x̄)² / (n − 1)] | ||
| Interquartile Range | Q₃ − Q₁ | ||
| Percentage | (Part / Whole) × 100 | ||
| Percentage Change | [(New − Original) / Original] × 100 | ||
| Relative Frequency | Frequency / Total Frequency | ||
| Probability | Favorable Outcomes / Total Outcomes | ||
| Z-Score | (x − μ) / σ | ||
| Mean Absolute Deviation | Σ | x − x̄ | / n |
| Coefficient of Variation | (Standard Deviation / Mean) × 100 |
How to Choose the Right Statistics Formula
Choosing the correct formula depends on the question you are trying to answer.
If you want to find the average, use the mean. If the data contains extreme values and you want a central value that is less affected by them, the median can be useful. If you want to identify the most common value, use the mode.
For a simple measure of spread, use the range. For a more detailed measure of variability, use variance or standard deviation. If you are interested in the middle 50% of the data, use the interquartile range.
When comparing a value with the mean in terms of standard deviations, use the z-score. When comparing relative variability between different data sets, the coefficient of variation can be useful.
Tips for Using Statistics Formulas
Learning formulas is easier when you understand the meaning behind them rather than memorizing them without context.
First, identify the type of data you are working with. Then determine what the question is asking you to calculate. Write down the relevant formula before substituting values. Keep track of whether you are working with a population or a sample, especially when calculating variance and standard deviation.
It is also important to arrange data in order when finding the median or quartiles. Check your arithmetic carefully, particularly when calculating deviations from the mean. Finally, always interpret the result in the context of the original data rather than treating the numerical answer as the complete conclusion.
Conclusion
Basic statistics formulas provide a foundation for understanding and analyzing data. Measures such as the mean, median, mode, and range help describe the center and basic spread of a data set, while variance, standard deviation, and interquartile range provide more information about variability. Other formulas, including percentage, relative frequency, probability, z-score, and coefficient of variation, help us describe and compare data in different ways.
The key to learning statistics is not simply memorizing formulas. It is understanding what each formula measures, recognizing when it should be used, and interpreting the result correctly. Once these basic concepts become familiar, more advanced statistical methods become much easier to understand.
FAQs
1. What is the mean in statistics?
The mean is one of the most common measures of central tendency in statistics. It represents the average value of a data set. To calculate the mean, add all the observations and divide the total by the number of observations. The formula is x̄ = Σx / n, where x̄ is the mean, Σx is the sum of all observations, and n is the number of observations. For example, for 4, 6, 8, 10, and 12, the sum is 40 and there are five values. Therefore, the mean is 40 ÷ 5 = 8.
2. What is the difference between mean, median, and mode?
Mean, median, and mode are three measures used to describe the center of a data set. The mean is calculated by adding all values and dividing by the number of values. The median is the middle value after arranging the data in order. If there are two middle values, their average is the median. The mode is the value that occurs most frequently. For example, in the data 2, 3, 3, 5, and 7, the mean is 4, the median is 3, and the mode is 3. Each measure describes the data from a different perspective.
3. What is the formula for the median?
The median is the middle value of an ordered data set. When there is an odd number of observations, the median position can be found using (n + 1) / 2, where n is the number of observations. For an even number of observations, there are two middle values, and the median is their average. For example, in the ordered data 2, 4, 6, 8, and 10, there are five values, so the median position is (5 + 1) ÷ 2 = 3. The third value is 6, making 6 the median.
4. What is the formula for range in statistics?
The range is a simple measure of the spread of data. It shows the difference between the largest and smallest observations in a data set. The formula is Range = Maximum value − Minimum value. For example, consider the data 5, 8, 12, 15, and 20. The maximum value is 20 and the minimum value is 5. Therefore, the range is 20 − 5 = 15. A larger range generally indicates that the data values cover a wider interval. However, because the range depends only on the two extreme values, it may be strongly affected by unusually large or small observations.
5. What is variance in statistics?
Variance is a statistical measure that describes how widely observations are spread around their mean. A small variance means that the values tend to remain close to the mean, while a larger variance indicates greater spread. For a population, the formula is σ² = Σ(x − μ)² / N. For a sample, the formula is s² = Σ(x − x̄)² / (n − 1). To calculate variance, find the difference between each observation and the mean, square those differences, add them, and divide by the appropriate denominator. Variance is useful for understanding the variability within a data set.
6. What is standard deviation and why is it used?
Standard deviation measures how much observations typically vary from the mean. It is calculated by taking the square root of the variance. The population standard deviation formula is σ = √[Σ(x − μ)² / N], while the sample formula is s = √[Σ(x − x̄)² / (n − 1)]. A small standard deviation indicates that observations tend to be close to the mean. A large standard deviation indicates greater variability. Standard deviation is especially useful because it is expressed in the same units as the original data. For example, if the data represents centimeters, standard deviation is also measured in centimeters.
7. What is the interquartile range formula?
The interquartile range, commonly abbreviated as IQR, measures the spread of the middle 50% of a data set. Its formula is IQR = Q₃ − Q₁, where Q₁ is the first quartile and Q₃ is the third quartile. Q₁ represents the lower quartile, while Q₃ represents the upper quartile. For example, if Q₁ = 20 and Q₃ = 35, then IQR = 35 − 20 = 15. The IQR is useful for describing the central spread of data and is less affected by extreme observations than the range. It is commonly used when data contains unusually high or low values.
8. What is the percentage formula in statistics?
The percentage formula is used to express a part of a whole as a value out of 100. The formula is Percentage = (Part ÷ Whole) × 100. For example, if 25 students out of 40 students pass an examination, the percentage is (25 ÷ 40) × 100 = 62.5%. Percentages make it easier to compare proportions between groups of different sizes. They are commonly used in surveys, examinations, business reports, scientific studies, and data analysis. Before calculating a percentage, identify the part being measured and the complete total to make sure the correct values are used.
9. What is a z-score in statistics?
A z-score indicates how far an observation is from the mean when measured in standard deviations. The formula is z = (x − μ) / σ, where x is the observation, μ is the mean, and σ is the population standard deviation. For example, suppose a value is 70, the mean is 60, and the standard deviation is 5. The z-score is (70 − 60) ÷ 5 = 2. This means the observation is two standard deviations above the mean. A positive z-score indicates a value above the mean, while a negative z-score indicates a value below the mean.
10. Which basic statistics formulas should beginners learn first?
Beginners should first understand formulas that describe the center, spread, and proportions of data. Important formulas include the mean, median, mode, and range. After becoming comfortable with these, beginners can learn variance, standard deviation, quartiles, and interquartile range. Percentage, relative frequency, probability, z-score, mean absolute deviation, and coefficient of variation are also useful. Rather than memorizing every formula immediately, it is better to understand what each one measures and when it should be used. Practicing with small data sets can make the formulas easier to remember and helps develop confidence in basic statistical calculations.

















