FRM Part 1Quantitative AnalysisChapter QTA 5

Sample Moments

MidhaFin20 min readUpdated

Reading tools
Table of contents
  • Video Lecture
  • |
  • PDFs
  • |
  • List of chapters

Learning Objectives

  1. Estimate the mean, variance, and standard deviation using sample data.
  2. Explain the difference between a population moment and a sample moment.
  3. Differentiate between an estimator and an estimate.
  4. Describe the bias of an estimator and explain what the bias measures.
  5. Explain what is meant by the statement that the mean estimator is BLUE.
  6. Describe the consistency of an estimator and explain the usefulness of this concept.
  7. Explain how the Law of Large Numbers (LLN) and Central Limit Theorem (CLT) apply to the sample mean.
  8. Estimate and interpret the skewness and kurtosis of a random variable.
  9. Estimate quantiles, including the median, using sample data.
  10. Estimate the mean of two variables and apply the CLT.
  11. Estimate the covariance and correlation between two random variables.
  12. Explain how coskewness and cokurtosis are related to skewness and kurtosis.

Every distribution in the previous chapters was described by its moments: a mean, a variance, a skewness, a correlation. All of those are properties of the true, underlying distribution, and in the real world none of them can be observed directly. What a risk manager actually has is data, a finite sample of returns or losses, and the job is to work backward from that sample to a best guess of the moments that generated it. This chapter is about that guessing, done properly.

The central idea is the sample analog: wherever a population moment uses an expectation, the sample version replaces it with an average over the data. That single move turns the true mean into the sample mean, the true variance into the sample variance, and so on. The rest of the chapter is about how good these guesses are, whether they are centered on the right answer, how their uncertainty shrinks as more data arrive, and why the sample mean in particular is as good as a linear estimator can be.

Key Takeaways

  • An estimator is a formula that turns data into a guess; an estimate is the number it produces. The estimator is a random variable, the estimate one of its realizations.
  • Sample moments estimate population moments by replacing the expectation with an average over the data.
  • Bias is the average gap between an estimator and the truth; the sample mean is unbiased, and dividing the sample variance by n minus 1 makes it unbiased too.
  • The standard error, the standard deviation of the data divided by the square root of n, measures the uncertainty of the mean estimate and shrinks as data grow.
  • The sample mean is BLUE for iid data: the best linear unbiased estimator of the population mean.
  • The Law of Large Numbers drives the sample mean to the truth; the Central Limit Theorem makes its distribution normal, which is what confidence intervals rely on.

Estimators, Estimates, and Sample Moments

Three words get used loosely in conversation but mean distinct things here. A population moment is a true property of the underlying distribution, such as its mean; it is a fixed number, and it is unknown. An estimator is a formula that takes data and returns a guess at that number. An estimate is the single value the estimator produces once real data are plugged in. The estimator is the recipe; the estimate is the dish.

The distinction that trips people up is that an estimator is itself a random variable. Feed it a different sample, drawn from the same distribution, and it returns a different value. So the sample mean is not just a number; it has its own distribution, its own mean, and its own variance, and the whole chapter is really about the behavior of that distribution. The hat notation marks the difference: the true mean is written as the Greek letter mu, and its estimator carries a hat.

μ̂ = 1n Σi Xi

where μ̂ (“mu-hat”) is the sample mean estimator, n is the number of observations, and the Xi are the sampled random variables. Replacing the random variables with observed data values turns the estimator into an estimate.

Estimating the Mean, Variance, and Standard Deviation

The sample mean is the average of the data. The sample variance follows the same sample-analog recipe: average the squared deviations from the sample mean. There is one wrinkle in the divisor, taken up in the next section, but the shape is exactly the population definition with the expectation swapped for an average. The sample standard deviation is the square root of the sample variance.

A small worked sample runs through the chapter. A fund reports five monthly returns: 3%, −2%, 5%, 1%, and 3%. Exhibit 1 lays out the arithmetic of the mean and the squared deviations.

Exhibit 1. Five monthly returns: mean and squared deviations
MonthReturn x (%)x − mean(x − mean)²
1311
2−2−416
3539
41−11
5311
Sum10028
Worked Example 1: mean, variance, and standard deviation

From the five returns in Exhibit 1, estimate the mean, the variance, and the standard deviation.

Step 1. The sample mean is the sum over n.

μ̂ = 105 = 2%

Step 2. The unbiased sample variance divides the summed squared deviations by n − 1.

s²= 285 − 1 = 284 = 7.0 (%²)

Step 3. The standard deviation is the square root.

s = √7.0 ≈ 2.65%

Answer: an estimated mean of 2% per month, a variance of 7.0, and a standard deviation of about 2.65%. These three numbers are estimates: a different five months would give different values, which is exactly why the estimator has a distribution of its own.

Every sample moment is built the same way, by taking the population definition and swapping the expectation for an average. Exhibit 2 lines up each population moment against its sample-analog estimator, which is the single pattern behind the whole chapter.

Exhibit 2. Population moments and their sample analogs
QuantityPopulation (true, unknown)Sample estimator (from data)
Meanμ = E[X](1/n) Σ xi
Varianceσ² = E[(X − μ)²](1/(n−1)) Σ (xi − x̄)²
Standard deviationσs = √(sample variance)
Skewnessμ3 / σ³μ̂3 / σ̂³
Kurtosisμ4 / σ⁴μ̂4 / σ̂⁴
CovarianceE[(X−μX)(Y−μY)](1/(n−1)) Σ (xi−x̄)(yi−ȳ)

Bias and the n Minus 1 Correction

An estimator’s bias is the average gap between the estimate and the truth: the expected value of the estimator minus the true parameter. A bias of zero means the estimator is unbiased, that is, correct on average across all possible samples, even though any single estimate will miss.

Bias(θ̂) = E[θ̂] − θ

where θ̂ (“theta-hat”) is the estimator and θ the true parameter. The sample mean is unbiased, since its expected value is exactly the population mean.

The sample variance needs care. If the squared deviations are divided by n, the estimator comes out slightly too small on average. The reason is subtle but important: the deviations are measured from the sample mean, not the true mean, and the sample mean is, by construction, the point that fits its own data most snugly. That snug fit shrinks the deviations a touch, and estimating the mean uses up one degree of freedom. Dividing by n − 1 instead of n corrects for exactly that, and the result is unbiased.

s² = 1n − 1 Σi (Xi − μ̂)²,   E[s²] = σ²

where dividing by n − 1 rather than n makes the estimator unbiased. The version that divides by n is biased downward by a factor of (n − 1) / n, a bias that fades as n grows.

Common Mistake

Dividing the sample variance by n out of habit. On the five returns, dividing by n gives 28 / 5 = 5.6, while the unbiased estimate divides by n − 1 to give 28 / 4 = 7.0. For a five-point sample the difference is large; for a sample of a thousand returns it is trivial. Know which one a question wants: the unbiased s² (divide by n − 1) is the default in financial statistics.

Standard Error and Scaling

Because the sample mean is a random variable, it has a standard deviation of its own, and that quantity has a special name: the standard error. It is not the spread of the data; it is the spread of the estimate. Where the standard deviation says how scattered individual returns are, the standard error says how far the estimated mean is likely to sit from the true mean. It is the data’s standard deviation divided by the square root of the sample size, so it shrinks as more data arrive, while the data’s own standard deviation does not.

SE(μ̂) = σ√n

where σ is the standard deviation of the data (estimated by s) and n is the sample size. Quadrupling the data halves the standard error, since the square root of 4 is 2. This is the diminishing return of collecting more data.

Two more scaling rules matter in practice, both flowing from the iid results of the previous chapter. To move a mean or a standard deviation from one frequency to another, the mean scales with the number of periods while the standard deviation scales with the square root of that number. A daily mean is annualized by multiplying by the number of trading days; a daily standard deviation is annualized by multiplying by the square root of that number.

Worked Example 2: standard error and annualizing

Using the five-month sample (mean 2%, standard deviation 2.65% per month), find the standard error of the estimated mean, and annualize the monthly mean and standard deviation.

Step 1. Standard error: the standard deviation over the square root of n.

SE= 2.65%√5 ≈ 2.65%2.236 ≈ 1.18%

Step 2. Annualize: the mean by 12, the standard deviation by the square root of 12.

Annual mean= 2% × 12 = 24% Annual SD= 2.65% × √12 ≈ 9.2%

Answer: a standard error of about 1.18%, an annualized mean of 24%, and an annualized standard deviation of about 9.2%. Note the two different roles of the square root: risk scales with the square root of the horizon, and the uncertainty of the mean shrinks with the square root of the sample size.

These behaviors are worth holding side by side, because the exam likes to test whether a candidate keeps them straight. Exhibit 3 collects them.

Exhibit 3. Spread of the data versus uncertainty of the estimate
QuantityWhat it measuresHow it behaves
Standard deviation, sSpread of the data around its meanRoughly stable as more data arrive
Standard error, s/√nUncertainty of the estimated meanShrinks as one over the square root of n
Annualized meanMean scaled to a one-year horizonScales with the number of periods, k
Annualized standard deviationRisk scaled to a one-year horizonScales with the square root of k

BLUE and Consistency

Two properties tell you the sample mean is not just adequate but excellent. The first is that it is BLUE, the Best Linear Unbiased Estimator of the population mean when the data are iid. Unpack the acronym: among all estimators that are both linear, a weighted sum of the observations, and unbiased, correct on average, the sample mean has the smallest variance. No other linear unbiased rule squeezes more precision out of the same data. It does not claim that no biased or nonlinear estimator can ever do better; it claims the sample mean wins within the linear, unbiased class.

The second is consistency. An estimator is consistent if it homes in on the true value as the sample grows without limit. Consistency needs two things: any bias must shrink to zero, and the variance must shrink to zero as well. Consistency is why more data genuinely help: it guarantees that the estimate can be made as close to the truth as desired simply by collecting enough observations. It is a large-sample promise, distinct from bias, which is a fixed-sample statement.

Key Insight

Bias and consistency answer different questions. Bias asks, “at this sample size, is the estimator right on average?” Consistency asks, “as the sample grows, does the estimator close in on the truth?” The two can disagree. The variance estimator that divides by n is biased at every finite sample size, yet its bias vanishes as n grows, so it is still consistent. An estimator can be unbiased but not consistent, or biased but consistent; the properties are independent.

The Law of Large Numbers and the Central Limit Theorem

Two results describe what happens to the sample mean as the sample grows, and together they are the backbone of statistical inference. The Law of Large Numbers, or LLN, says the sample mean converges to the true population mean: pile up enough independent observations and the estimate settles onto the target. It is the formal reason a large sample gives a trustworthy average.

The Central Limit Theorem, or CLT, describes the shape of the estimate along the way. For a large sample, the distribution of the sample mean is approximately normal, centered on the true mean with a standard deviation equal to the standard error, and this holds whatever the shape of the underlying data. Returns can be skewed and fat-tailed, yet the distribution of their average is close to a bell curve. That is what lets a confidence interval be built with the familiar normal critical values.

μ̂ → μ (LLN),   μ̂ ≈ Normal(μ, σ²/n) (CLT)

where the LLN gives convergence of the estimate to the truth and the CLT gives the approximate normal shape of the estimator, with variance equal to the squared standard error, σ²/n.

Worked Example 3: a confidence interval for the mean

Using the estimated mean of 2% and the standard error of 1.18% from earlier, build an approximate 95% confidence interval for the true monthly mean return, applying the CLT.

Step 1. A 95% interval spans the estimate plus and minus 1.96 standard errors.

95% CI= μ̂ ± 1.96 × SE = 2% ± 1.96 × 1.18% = 2% ± 2.32%

Answer: an interval from about −0.32% to 4.32%. With only five observations the interval is wide, and it even straddles zero, so this sample cannot rule out a true mean of zero. Collect more data and the standard error shrinks, the interval tightens, and the estimate sharpens, exactly as the LLN promises.

The practical payoff is large. Because the Central Limit Theorem holds whatever the shape of the raw data, a risk manager can attach a normal-based margin of error to an estimated mean, an estimated covariance, or the estimated mean of a combined pair of series, without first proving that the underlying returns are normal. This is why so much of statistical inference leans on the normal even though returns themselves are not normal: it is the average that turns normal, not the data. The one caveat is that the approximation needs a reasonably large sample to be trusted, and with a handful of observations, as in the five-return example, the interval is only a rough guide.

Estimating Skewness and Kurtosis

The same sample-analog recipe estimates the higher moments. Sample skewness averages the cubed deviations and divides by the cubed standard deviation; sample kurtosis averages the fourth-power deviations and divides by the fourth power of the standard deviation. Both are standardized, so they carry no units and can be compared across assets. Because the cube keeps the sign of a deviation, skewness reads asymmetry; because the fourth power magnifies large deviations, kurtosis reads tail heaviness.

skewnesŝ = μ̂3σ̂³,   kurtosiŝ = μ̂4σ̂⁴

where μ̂3 and μ̂4 are the sample third and fourth central moments (the averaged cubed and fourth-power deviations). Excess kurtosis subtracts 3, the kurtosis of a normal, so a positive excess kurtosis flags fatter-than-normal tails.

For most financial return series the estimates come out with a recognizable signature: a mildly negative skewness, because large drops arrive more often than equally large jumps, and a kurtosis well above 3, because extreme moves of either sign are more frequent than a normal would allow. A risk manager who estimates a near-zero skewness and a kurtosis of 3 on a long return series should suspect the data rather than trust the calm, since real markets rarely behave that neatly.

Two cross-moment cousins extend the idea to pairs of variables. Coskewness and cokurtosis are to skewness and kurtosis what covariance is to variance: they measure whether one variable takes an extreme value at the same time as another. Negative coskewness between two returns, for instance, says that when one falls sharply the other tends to be turbulent too, which is precisely the joint tail behavior that matters for a portfolio in a crisis.

Estimating Quantiles and the Median

Quantiles are estimated by sorting the data and reading off the value at the right position. The median, the 0.5 quantile, is the middle value of the sorted sample; the 0.95 quantile sits 95 percent of the way up the sorted list. A rough position rule places the p-quantile at rank p times n in the sorted data, with interpolation between neighboring points when that rank is not a whole number. Quantile estimates have one great virtue for financial data: they are robust to outliers. A single monstrous return drags the mean around but barely moves the median, which is why quantiles are trusted when the data are heavy-tailed, the usual case in markets, and why value at risk, itself a quantile of losses, is often estimated directly from the sorted data.

Worked Example 4: the sample median

Find the median of the five monthly returns: 3%, −2%, 5%, 1%, 3%.

Step 1. Sort the returns from smallest to largest.

−2%,   1%,   3%,   3%,   5%

Step 2. With five values, the median is the third, the middle one.

Answer: the median is 3%, above the mean of 2%. The single weak month of −2% pulls the mean down but leaves the median untouched, which is the robustness that makes quantiles useful when a few extreme observations would distort the average.

Estimating Covariance and Correlation

With two data series, the sample analog estimates their covariance and correlation. Sample covariance averages the product of the two variables’ deviations from their own means, dividing by n − 1 for the same unbiasedness reason as the variance. Sample correlation then divides that covariance by the product of the two sample standard deviations, landing between minus one and plus one.

Cov̂(X, Y) = 1n − 1 Σi (Xi − X̄)(Yi − Ȳ),   Corr̂ = Cov̂(X, Y)sX sY

where X̄ and Ȳ are the two sample means and sX, sY the two sample standard deviations. As with the two-variable mean, the Central Limit Theorem applies, so a covariance or correlation estimated from a large sample is itself approximately normally distributed around its true value.

Worked Example 5: sample covariance and correlation

Two funds report returns over four periods. Fund A: 2, −1, 4, 3 (mean 2). Fund B: 1, 0, 3, 4 (mean 2). Estimate the covariance and correlation.

Step 1. Multiply the paired deviations and sum. Deviations of A are 0, −3, 2, 1; of B are −1, −2, 1, 2.

Σ products= (0)(−1) + (−3)(−2) + (2)(1) + (1)(2) = 0 + 6 + 2 + 2 = 10

Step 2. Divide by n − 1 = 3 for the covariance.

Cov̂(A, B) = 103 ≈ 3.33

Step 3. Divide by the two standard deviations, sA ≈ 2.16 and sB ≈ 1.83.

Corr̂ = 3.33(2.16)(1.83) ≈ 0.84

Answer: a sample covariance of about 3.33 and a correlation of about 0.84, a strong positive relationship. The two funds move together closely in this sample, though with only four periods the estimate is noisy, and a longer history would pin it down more tightly.

Check Yourself

A sample of four losses is 2, 4, 4, 6 (in millions). Estimate the unbiased sample variance.

Show answer

The mean is (2 + 4 + 4 + 6) / 4 = 4. The deviations are −2, 0, 0, 2, and their squares are 4, 0, 0, 4, summing to 8. Divide by n − 1 = 3: the unbiased variance is 8 / 3 ≈ 2.67 (millions squared), so the standard deviation is about 1.63 million. Dividing by n instead would have given 2.0, which is the biased estimate.

Check Yourself

A mean is estimated with a standard error of 4%. Roughly how much data, relative to what you have, is needed to cut the standard error to 2%?

Show answer

The standard error falls with the square root of the sample size, so halving it requires quadrupling the data. Cutting the standard error from 4% to 2% means dividing it by 2, which needs 2² = 4 times as many observations. This square-root law is why the first few observations sharpen an estimate quickly, while later ones add precision ever more slowly.

Check Yourself

A daily return series has a standard deviation of 1.5%. Under an iid assumption with 252 trading days, what is the annualized standard deviation?

Show answer

The standard deviation scales with the square root of the number of periods, so the annualized figure is 1.5% × √252 ≈ 1.5% × 15.87 ≈ 23.8%. The mean would scale by the full 252, but risk scales only by the square root, because variance, not standard deviation, adds across independent days.

Check Yourself

A colleague says an estimator that divides the sample variance by n is useless because it is biased. Why is that too harsh?

Show answer

Because it is still consistent. The bias of the divide-by-n estimator is a factor of (n − 1) / n, which approaches 1 as n grows, so the bias vanishes in large samples and the estimate converges to the true variance. For the large datasets common in finance the difference between dividing by n and by n − 1 is negligible. Bias at a fixed sample size and behavior as the sample grows are separate questions.

Chapter Summary

  • A population moment is a true, unknown property of a distribution; a sample moment estimates it by replacing the expectation with an average over the data.
  • An estimator is a formula and a random variable; an estimate is the single number it produces from one sample.
  • The sample mean is unbiased. The sample variance is unbiased only when the squared deviations are divided by n minus 1, which corrects for the degree of freedom used in estimating the mean.
  • The standard error, σ over the square root of n, is the standard deviation of the mean estimate and shrinks with more data; the standard deviation of the data does not.
  • Means scale with the number of periods and standard deviations with the square root of that number, the square-root-of-time rule.
  • The sample mean is BLUE for iid data: the minimum-variance choice among linear unbiased estimators. Consistency, a separate property, guarantees convergence to the truth as the sample grows.
  • The Law of Large Numbers drives the sample mean to the population mean; the Central Limit Theorem makes its distribution approximately normal, which is what confidence intervals use.
  • Skewness and kurtosis are estimated by the same sample analog; coskewness and cokurtosis extend them to the joint tail behavior of two variables.
  • Quantiles and the median are read off the sorted data and are robust to outliers; sample covariance and correlation estimate the co-movement of two series.

Frequently Asked Questions

What is the difference between an estimator and an estimate?

An estimator is the rule or formula that turns data into a guess of an unknown quantity, for example the formula that averages the observations to estimate the mean. An estimate is the single number that rule produces once actual data are plugged in. The estimator is a random variable, because it would give a different value for a different sample; the estimate is one realized value of it. Loosely, the estimator is the recipe and the estimate is the dish.

What is the difference between a population moment and a sample moment?

A population moment is a fixed but unknown property of the whole distribution, such as its true mean or variance. A sample moment is a quantity computed from a finite set of observed data that is used to estimate the corresponding population moment. The population moment is the target; the sample moment is the data-based guess at it. Sample moments are built by replacing the expectation in the population definition with an average over the sample, which is called the sample analog.

Why is the sample variance divided by n minus 1 instead of n?

Dividing by n gives a variance estimator that is biased slightly downward, because the deviations are measured from the sample mean rather than the true mean, and the sample mean fits its own data a little too well. Estimating the mean uses up one degree of freedom, and dividing by n minus 1 instead of n exactly corrects for that, producing an unbiased estimator whose expected value equals the true variance. The correction matters most for small samples and becomes negligible as the sample grows.

What is the difference between a standard error and a standard deviation?

A standard deviation measures the spread of the data itself, how far individual observations sit from their mean. A standard error measures the uncertainty of an estimator, how far the estimate is likely to sit from the true value. For the sample mean, the standard error is the standard deviation of the data divided by the square root of the sample size, so it shrinks as more data are collected, while the standard deviation of the data does not. The standard error is what turns an estimate into a confidence interval.

What does it mean for the mean estimator to be BLUE?

BLUE stands for Best Linear Unbiased Estimator. When the data are independent and identically distributed, the sample mean is BLUE for the population mean, which means that among all estimators that are both a linear combination of the observations and unbiased, the sample mean has the smallest variance. In plain terms, no other linear unbiased rule uses the data more efficiently to estimate the mean. It does not rule out biased or nonlinear estimators being more accurate in some settings.

What is the difference between the Law of Large Numbers and the Central Limit Theorem?

The Law of Large Numbers says that as the sample size grows, the sample mean converges to the true population mean, so with enough data the estimate lands on the target. The Central Limit Theorem says something about the shape along the way: for a large sample, the distribution of the sample mean is approximately normal, centered on the true mean with a standard deviation equal to the standard error, whatever the shape of the underlying data. The Law of Large Numbers gives convergence; the Central Limit Theorem gives the bell-shaped distribution used to build confidence intervals.

Can an estimator be biased but still consistent?

Yes. Bias is about the average error at a fixed sample size, while consistency is about what happens as the sample size grows without limit. An estimator is consistent if it converges to the true value as the sample grows, which requires any bias to shrink to zero and the variance to shrink to zero. The biased version of the sample variance, which divides by n, is a good example: it is biased for any finite sample, yet its bias vanishes as the sample grows, so it is still consistent.

Go to Syllabus

Courses Offered

image

FRM® Part-1 Sample Course

Instructor · Micky Midha

  • 9 Hrs of Videos

  • Available On Web, IOS & Android

  • Access Until You Pass

  • Lecture PDFs

  • Class Notes

image

FRM® Part-2 Sample Course

Instructor · Micky Midha

  • 12 Hrs of Videos

  • Available On Web, IOS & Android

  • Access Until You Pass

  • Lecture PDFs

  • Class Notes

image

FRM® Part-1 Self Paced Course

Instructor · Micky Midha

  • 257 Hrs Of Videos

  • Available On Web, IOS & Android

  • Access Until You Pass

  • Complete Study Material

  • Quizzes,Question Bank & Mock tests

image

FRM® Part-2 Self Paced Course

Instructor · Micky Midha

  • 240 Hrs Of Videos

  • Available On Web, IOS & Android

  • Access Until You Pass

  • Complete Study Material

  • Quizzes,Question Bank & Mock tests

image

PRM Exam 1

Instructor · Shubham Swaraj

  • Lecture Videos

  • Available On Web, IOS & Android

  • Complete Study Material

  • Question Bank & Lecture PDFs

  • Doubt-Solving Forum

No comments on this post so far:

Add your Thoughts:

    Chat with MidhaFin on WhatsAppJoin MidhaFin on Telegram