Sampling and confidence

Clustered, Stratified, and Weighted Samples

Understand why survey and grouped designs change estimator construction, effective information, and variance beyond ordinary simple-random formulas.

Direct answer

Stratification samples within defined subgroups, clustering samples groups of related units, and weighting changes each observation’s contribution; all can alter precision and require design-aware estimation.

Visual explanation

Sampling design changes represented information

strataclustersweights
Strata, clusters, and weights must remain part of estimation and uncertainty.

What this calculation tells you

Complex sampling designs determine who can be observed and how sample records represent a target population. Their structure is part of the estimator, not a detail added after analysis.

Record strata, clusters, selection probabilities, stages, finite-population information, and weight meaning. Use software and methods that carry those design variables into variance estimation.

Where it is used

Public surveys

Explain why released estimates include design-aware standard errors.

Experiments

Account for treatment delivered by site, class, or household clusters.

Market research

Document weighting and stratified allocation choices.

Education

Visualize representation versus raw record count.

Common situations

  • Planning cluster sampling.
  • Allocating a sample across strata.
  • Interpreting survey weights.
  • Explaining effective rather than literal sample size.

Start with the statistical question

Record strata, clusters, selection probabilities, stages, finite-population information, and weight meaning. Use software and methods that carry those design variables into variance estimation.

One population is sampled three ways: scattered simple units, colored strata with within-stratum samples, and whole clusters. Weight sizes then change represented population mass.

Worked example

If respondents come in similar clusters of average size m with intraclass correlation ρ, the equal-cluster approximation DEFF=1+(m−1)ρ shows why 100 clustered observations may carry less information than 100 independent ones.

Assumptions that carry the result

The simple design-effect formula assumes equal cluster size and one correlation approximation. Kish effective sample size addresses weight variation only and is not a complete survey variance estimator.

Interpret the result without overreaching

These arithmetic summaries do not replace replicate weights, Taylor linearization, multistage design variables, nonresponse adjustment, calibration, or specialist survey analysis.

  • Analyzing clustered records as independent.
  • Treating frequency weights as survey weights.
  • Reporting effective sample size as the literal respondent count.

Choose the right tool

Practical questions

Frequently asked questions

Are weighted results automatically representative?

No. Weight construction, coverage, nonresponse, calibration, and the target population all matter.

Does clustering always reduce precision?

Positive within-cluster similarity commonly does; the effect depends on correlation, cluster sizes, allocation, and estimator.

Can I use n_eff in every ordinary formula?

Not safely. Effective-size summaries are approximations and do not reproduce every feature of a complex variance estimator.

Further reading

Authoritative sources

Use these primary and professional resources to check definitions, conventions, or requirements that may extend beyond this guide.