Basic Statistics Formulas for Beginners

Realistic 3D illustration of basic statistics formulas, charts, graphs, and mathematical data

Statistics is the branch of mathematics that helps us collect, organize, analyze, interpret, and present data. Whether you are studying mathematics, science, economics, business, or simply trying to understand information in everyday life, basic statistics provides useful tools for making sense of numbers.

At first, statistical formulas may look complicated because they use symbols and mathematical notation. However, most basic formulas are based on simple ideas. Once you understand what each quantity represents and when to use a particular formula, statistics becomes much easier to learn.

This guide explains the most important basic statistics formulas for beginners, including mean, median, mode, range, variance, standard deviation, quartiles, percentages, and other commonly used measures. Each formula is explained in a simple way with examples so that you can understand how it works and when to apply it.

What Are Statistics Formulas?

Statistics formulas are mathematical rules used to calculate useful information from a collection of data. A set of numbers is called a data set, and each individual value in the data set is called an observation or data value.

For example, consider the following data:

5, 7, 8, 10, 10

Using statistical formulas, we can find its mean, median, mode, range, and other characteristics.

Statistics formulas help answer questions such as:

  • What is the average value?

  • Which value occurs most often?

  • How spread out are the data?

  • What is the difference between the largest and smallest values?

  • How much does each value differ from the average?

  • What proportion or percentage of observations belongs to a particular group?

Important Symbols Used in Statistics

Before learning the formulas, it is useful to understand some common statistical symbols.

  • x = an individual data value

  • n = number of observations in a data set

  • Σ = summation of values

  • x̄ = arithmetic mean of the data

  • μ = population mean

  • σ = population standard deviation

  • s = sample standard deviation

  • N = population size

  • Q₁ = first quartile

  • Q₂ = second quartile or median

  • Q₃ = third quartile

The exact symbols can vary slightly between textbooks, but the underlying concepts remain the same.

Arithmetic Mean Formula

The arithmetic mean, commonly called the average, is one of the most frequently used statistical measures.

To calculate the mean, add all the data values and divide the total by the number of observations.

Formula

Mean = Sum of all observations ÷ Number of observations

Using symbols:

x̄ = Σx / n

Example

Suppose the data is:

4, 6, 8, 10, 12

The sum is:

4 + 6 + 8 + 10 + 12 = 40

There are 5 observations.

Therefore:

Mean = 40 ÷ 5 = 8

So, the mean of the data set is 8.

The mean is useful when you want a single value that represents the central level of a data set.

Weighted Mean Formula

Sometimes every value does not have the same importance. In such cases, a weighted mean can be used.

Formula

Weighted Mean = Σ(wx) / Σw

where:

  • x = value

  • w = weight assigned to the value

Example

Suppose a student’s marks are:

  • Test = 70, weight = 2

  • Assignment = 80, weight = 1

  • Exam = 90, weight = 3

Then:

Weighted Mean = (70 × 2 + 80 × 1 + 90 × 3) / (2 + 1 + 3)

= (140 + 80 + 270) / 6

= 490 / 6

≈ 81.67

The weighted mean is approximately 81.67.

Median Formula

The median is the middle value when the observations are arranged in ascending or descending order.

For example:

3, 5, 7, 9, 11

The middle value is 7, so the median is 7.

Median Formula for an Odd Number of Observations

If there are n observations and n is odd:

Median position = (n + 1) / 2

Example

For 7 observations:

2, 4, 5, 7, 9, 12, 15

Median position:

(7 + 1) / 2 = 4

The fourth value is 7, so the median is 7.

Median Formula for an Even Number of Observations

When the number of observations is even, the median is the average of the two middle values.

Example

Consider:

2, 4, 6, 8, 10, 12

There are 6 observations. The two middle values are 6 and 8.

Median = (6 + 8) / 2 = 7

Therefore, the median is 7.

The median is particularly useful when a data set contains unusually large or small values because it is less affected by extreme observations than the mean.

Mode Formula

The mode is the value that occurs most frequently in a data set.

For example:

2, 3, 3, 4, 5, 3, 6

The number 3 occurs three times, more often than any other value.

Therefore:

Mode = 3

A data set can have:

  • One mode, called unimodal

  • Two modes, called bimodal

  • More than two modes, called multimodal

  • No mode when no value occurs more frequently than the others

Unlike the mean and median, the mode does not require a mathematical calculation. You simply identify the most frequently occurring value.

Range Formula

The range measures the difference between the largest and smallest values in a data set.

Formula

Range = Maximum value − Minimum value

Example

Consider:

5, 8, 12, 15, 20

Maximum value = 20

Minimum value = 5

Therefore:

Range = 20 − 5 = 15

The range is 15.

The range is simple to calculate, but it depends only on two observations: the maximum and minimum values. Therefore, it does not provide as much information about the overall spread as variance or standard deviation.

Variance Formula

Variance measures how far data values are spread around their mean.

A small variance means the observations tend to be close to the mean. A large variance means they are more widely spread out.

There are separate formulas for population and sample variance.

Population Variance Formula

σ² = Σ(x − μ)² / N

where:

  • σ² = population variance

  • x = individual observation

  • μ = population mean

  • N = population size

Sample Variance Formula

s² = Σ(x − x̄)² / (n − 1)

where:

  • s² = sample variance

  • x̄ = sample mean

  • n = sample size

The difference between dividing by N and n − 1 is important. Population variance is used when the entire population is being analyzed, while sample variance is commonly used when a sample is used to estimate characteristics of a larger population.

Standard Deviation Formula

The standard deviation is one of the most important measures of variability in statistics. It tells us how much the observations typically vary around their mean.

Standard deviation is the square root of variance.

Population Standard Deviation

σ = √[Σ(x − μ)² / N]

Sample Standard Deviation

s = √[Σ(x − x̄)² / (n − 1)]

Example

Suppose a data set has a mean of 10. If the observations are generally close to 10, the standard deviation will be small. If many observations are far away from 10, the standard deviation will be larger.

Because standard deviation is expressed in the same units as the original data, it is often easier to interpret than variance.

Quartile Formulas

Quartiles divide an ordered data set into four parts.

The three main quartiles are:

  • Q₁ = first quartile

  • Q₂ = second quartile or median

  • Q₃ = third quartile

Q₁ represents the point below which approximately 25% of the observations fall, while Q₃ represents the point below which approximately 75% of the observations fall.

For a simple position-based method:

First Quartile

Q₁ position = (n + 1) / 4

Second Quartile

Q₂ position = (n + 1) / 2

Third Quartile

Q₃ position = 3(n + 1) / 4

Different statistical methods may use slightly different conventions for calculating quartiles, especially when the position falls between observations.

Interquartile Range Formula

The interquartile range, or IQR, measures the spread of the middle 50% of the data.

Formula

IQR = Q₃ − Q₁

Example

If:

Q₁ = 20

and

Q₃ = 35

then:

IQR = 35 − 20 = 15

The interquartile range is 15.

The IQR is useful because it is less affected by extreme values than the range.

Percentage Formula

Percentages are commonly used in statistics to describe part of a whole.

Formula

Percentage = (Part / Whole) × 100

Example

If 25 students out of 40 students passed an examination:

Percentage = (25 / 40) × 100

= 62.5%

Therefore, 62.5% of the students passed.

Percentage Change Formula

Percentage change measures how much a value increases or decreases compared with its original value.

Formula

Percentage Change = [(New Value − Original Value) / Original Value] × 100

Example

Suppose a quantity increases from 50 to 60.

Percentage Change = [(60 − 50) / 50] × 100

= 20%

Therefore, the value increased by 20%.

If the result is negative, it represents a percentage decrease.

Frequency Formula

Frequency tells us how many times a particular value or category occurs.

For example, in the data:

2, 3, 3, 4, 4, 4, 5

the frequency of 4 is 3 because it occurs three times.

There is no special calculation required for simple frequency. It is obtained by counting the occurrences of each value.

Relative Frequency Formula

Relative frequency expresses the frequency of an observation or category as a proportion of the total number of observations.

Formula

Relative Frequency = Frequency / Total Frequency

To express it as a percentage:

Relative Frequency (%) = (Frequency / Total Frequency) × 100

Example

If 15 out of 60 observations belong to a particular category:

Relative Frequency = 15 / 60 = 0.25

As a percentage:

0.25 × 100 = 25%

Therefore, the category represents 25% of the observations.

Probability Formula

Probability is closely related to statistics and is used to describe the likelihood of an event.

For equally likely outcomes:

Formula

P(E) = Number of favorable outcomes / Total number of possible outcomes

Example

For a fair six-sided die, the probability of rolling a 4 is:

P(4) = 1 / 6

The probability can be expressed as a fraction, decimal, or percentage.

Z-Score Formula

A z-score indicates how many standard deviations an observation is above or below the mean.

Formula

z = (x − μ) / σ

For sample-based calculations, the corresponding mean and standard deviation may be used.

Example

Suppose:

x = 70

μ = 60

σ = 5

Then:

z = (70 − 60) / 5

= 2

The z-score is 2, meaning the observation is two standard deviations above the mean.

Mean Absolute Deviation Formula

The mean absolute deviation, or MAD, measures the average distance between each observation and the mean.

Formula

MAD = Σ|x − x̄| / n

The absolute value is used so that negative and positive deviations do not cancel each other out.

Example

Suppose the mean of a data set is 10. For every observation, calculate its distance from 10, ignore the sign, add the distances, and divide by the number of observations.

This gives the average absolute distance from the mean.

Coefficient of Variation Formula

The coefficient of variation, or CV, compares the standard deviation with the mean.

Formula

CV = (Standard Deviation / Mean) × 100

For a population:

CV = (σ / μ) × 100

For a sample:

CV = (s / x̄) × 100

The coefficient of variation is often useful when comparing variability between data sets that have different means or scales.

Summary of Basic Statistics Formulas

The following formulas are among the most useful for beginners:

Statistical MeasureFormula
Meanx̄ = Σx / n
Weighted MeanΣ(wx) / Σw
Median Position for Odd n(n + 1) / 2
RangeMaximum − Minimum
Population Varianceσ² = Σ(x − μ)² / N
Sample Variances² = Σ(x − x̄)² / (n − 1)
Population Standard Deviationσ = √[Σ(x − μ)² / N]
Sample Standard Deviations = √[Σ(x − x̄)² / (n − 1)]
Interquartile RangeQ₃ − Q₁
Percentage(Part / Whole) × 100
Percentage Change[(New − Original) / Original] × 100
Relative FrequencyFrequency / Total Frequency
ProbabilityFavorable Outcomes / Total Outcomes
Z-Score(x − μ) / σ
Mean Absolute DeviationΣx − x̄/ n
Coefficient of Variation(Standard Deviation / Mean) × 100

How to Choose the Right Statistics Formula

Choosing the correct formula depends on the question you are trying to answer.

If you want to find the average, use the mean. If the data contains extreme values and you want a central value that is less affected by them, the median can be useful. If you want to identify the most common value, use the mode.

For a simple measure of spread, use the range. For a more detailed measure of variability, use variance or standard deviation. If you are interested in the middle 50% of the data, use the interquartile range.

When comparing a value with the mean in terms of standard deviations, use the z-score. When comparing relative variability between different data sets, the coefficient of variation can be useful.

Tips for Using Statistics Formulas

Learning formulas is easier when you understand the meaning behind them rather than memorizing them without context.

First, identify the type of data you are working with. Then determine what the question is asking you to calculate. Write down the relevant formula before substituting values. Keep track of whether you are working with a population or a sample, especially when calculating variance and standard deviation.

It is also important to arrange data in order when finding the median or quartiles. Check your arithmetic carefully, particularly when calculating deviations from the mean. Finally, always interpret the result in the context of the original data rather than treating the numerical answer as the complete conclusion.

Conclusion

Basic statistics formulas provide a foundation for understanding and analyzing data. Measures such as the mean, median, mode, and range help describe the center and basic spread of a data set, while variance, standard deviation, and interquartile range provide more information about variability. Other formulas, including percentage, relative frequency, probability, z-score, and coefficient of variation, help us describe and compare data in different ways.

The key to learning statistics is not simply memorizing formulas. It is understanding what each formula measures, recognizing when it should be used, and interpreting the result correctly. Once these basic concepts become familiar, more advanced statistical methods become much easier to understand.

FAQs

1. What is the mean in statistics?

The mean is one of the most common measures of central tendency in statistics. It represents the average value of a data set. To calculate the mean, add all the observations and divide the total by the number of observations. The formula is x̄ = Σx / n, where x̄ is the mean, Σx is the sum of all observations, and n is the number of observations. For example, for 4, 6, 8, 10, and 12, the sum is 40 and there are five values. Therefore, the mean is 40 ÷ 5 = 8.

2. What is the difference between mean, median, and mode?

Mean, median, and mode are three measures used to describe the center of a data set. The mean is calculated by adding all values and dividing by the number of values. The median is the middle value after arranging the data in order. If there are two middle values, their average is the median. The mode is the value that occurs most frequently. For example, in the data 2, 3, 3, 5, and 7, the mean is 4, the median is 3, and the mode is 3. Each measure describes the data from a different perspective.

3. What is the formula for the median?

The median is the middle value of an ordered data set. When there is an odd number of observations, the median position can be found using (n + 1) / 2, where n is the number of observations. For an even number of observations, there are two middle values, and the median is their average. For example, in the ordered data 2, 4, 6, 8, and 10, there are five values, so the median position is (5 + 1) ÷ 2 = 3. The third value is 6, making 6 the median.

4. What is the formula for range in statistics?

The range is a simple measure of the spread of data. It shows the difference between the largest and smallest observations in a data set. The formula is Range = Maximum value − Minimum value. For example, consider the data 5, 8, 12, 15, and 20. The maximum value is 20 and the minimum value is 5. Therefore, the range is 20 − 5 = 15. A larger range generally indicates that the data values cover a wider interval. However, because the range depends only on the two extreme values, it may be strongly affected by unusually large or small observations.

5. What is variance in statistics?

Variance is a statistical measure that describes how widely observations are spread around their mean. A small variance means that the values tend to remain close to the mean, while a larger variance indicates greater spread. For a population, the formula is σ² = Σ(x − μ)² / N. For a sample, the formula is s² = Σ(x − x̄)² / (n − 1). To calculate variance, find the difference between each observation and the mean, square those differences, add them, and divide by the appropriate denominator. Variance is useful for understanding the variability within a data set.

6. What is standard deviation and why is it used?

Standard deviation measures how much observations typically vary from the mean. It is calculated by taking the square root of the variance. The population standard deviation formula is σ = √[Σ(x − μ)² / N], while the sample formula is s = √[Σ(x − x̄)² / (n − 1)]. A small standard deviation indicates that observations tend to be close to the mean. A large standard deviation indicates greater variability. Standard deviation is especially useful because it is expressed in the same units as the original data. For example, if the data represents centimeters, standard deviation is also measured in centimeters.

7. What is the interquartile range formula?

The interquartile range, commonly abbreviated as IQR, measures the spread of the middle 50% of a data set. Its formula is IQR = Q₃ − Q₁, where Q₁ is the first quartile and Q₃ is the third quartile. Q₁ represents the lower quartile, while Q₃ represents the upper quartile. For example, if Q₁ = 20 and Q₃ = 35, then IQR = 35 − 20 = 15. The IQR is useful for describing the central spread of data and is less affected by extreme observations than the range. It is commonly used when data contains unusually high or low values.

8. What is the percentage formula in statistics?

The percentage formula is used to express a part of a whole as a value out of 100. The formula is Percentage = (Part ÷ Whole) × 100. For example, if 25 students out of 40 students pass an examination, the percentage is (25 ÷ 40) × 100 = 62.5%. Percentages make it easier to compare proportions between groups of different sizes. They are commonly used in surveys, examinations, business reports, scientific studies, and data analysis. Before calculating a percentage, identify the part being measured and the complete total to make sure the correct values are used.

9. What is a z-score in statistics?

A z-score indicates how far an observation is from the mean when measured in standard deviations. The formula is z = (x − μ) / σ, where x is the observation, μ is the mean, and σ is the population standard deviation. For example, suppose a value is 70, the mean is 60, and the standard deviation is 5. The z-score is (70 − 60) ÷ 5 = 2. This means the observation is two standard deviations above the mean. A positive z-score indicates a value above the mean, while a negative z-score indicates a value below the mean.

10. Which basic statistics formulas should beginners learn first?

Beginners should first understand formulas that describe the center, spread, and proportions of data. Important formulas include the mean, median, mode, and range. After becoming comfortable with these, beginners can learn variance, standard deviation, quartiles, and interquartile range. Percentage, relative frequency, probability, z-score, mean absolute deviation, and coefficient of variation are also useful. Rather than memorizing every formula immediately, it is better to understand what each one measures and when it should be used. Practicing with small data sets can make the formulas easier to remember and helps develop confidence in basic statistical calculations.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top