Statistics for Researchers: A Complete Guide

Most researchers don't need to become statisticians. They need to choose the right analysis, run it correctly, interpret it honestly, and report it the way their field expects. This guide is the map for that work: statistics for researchers, organized in the order a real study moves through them.

Everything here links to a dedicated guide that works the topic out in full, with formulas, worked examples, and reporting conventions. Use this page to find where you are and what comes next, then follow the link that matches the decision in front of you.

Starting a study? Begin with populations and samples, then sampling method, then sample size. Those three decisions shape everything downstream and can't be fixed later.

Have data and don't know what to run? Skip to the test-selection table below. It routes by what you're comparing and what kind of data you have.

Writing up results? Go to the reporting section. It covers effect sizes, APA notation, and the errors reviewers flag most.

Want the concepts first? Read the two sub-guides: descriptive statistics for summarizing data, and inferential statistics for generalizing from it.

Which Test Do You Need?

The most common question in applied statistics is which test to run. It comes down to what you're comparing and what kind of data you have.

What you want to do Your data Test Guide
Compare one or two group means Continuous outcome T-test T-tests
Compare three or more group means Continuous outcome ANOVA ANOVA
Test whether two categories are related Categorical counts Chi-square Chi-square tests
Measure the strength of a relationship Two continuous variables Correlation Correlation coefficients
Predict an outcome from one variable Continuous outcome Simple linear regression Simple linear regression
Predict an outcome from several variables Continuous outcome Multiple regression Multiple regression
Predict a yes-or-no outcome Binary outcome Logistic regression Logistic regression
Estimate a range rather than test a difference Any Confidence interval Confidence intervals
Any of the above when assumptions fail Skewed or ordinal Rank-based alternative Non-parametric tests

The Shape of the Field

Statistics divides cleanly in two, and knowing which half you're in resolves most confusion about what a number is allowed to say.

Descriptive statistics describe the data in front of you. A mean, a standard deviation, a histogram: each one summarizes your sample and claims nothing beyond it. If you surveyed 200 students and found men reporting higher risk tolerance, descriptive statistics establish that this was true of those 200.

Inferential statistics take the next step. They use the sample to make a claim about a population you couldn't measure, and they quantify how uncertain that claim is. That second part is what makes inference more than guessing. A p-value, a confidence interval, and a power calculation are all ways of stating how much the sample could be misleading you.

Almost every quantitative paper uses both, in that order. You describe before you infer, because the shape and spread of your sample determine which inferential tools are valid on it.

Foundations: Who You're Studying and What You're Measuring

Before any analysis, two questions need clear answers. Who is this study about, and what exactly are you measuring?

The first is the distinction between the group you want to draw conclusions about and the smaller group you actually measure. Getting that relationship right is what licenses any claim beyond your own data. See the guide to population vs sample in research. The companion distinction is notational: a number describing a population is a parameter, one describing a sample is a statistic, and they use different symbols. The guide to parameter vs statistic covers the notation and where it trips people up.

The second question is about variables. Which factor are you treating as the cause, and which as the outcome? For what a predictor is and how it functions in a design, see the guide to independent variables. For the practical tests that tell the two roles apart in your own study, see identifying independent and dependent variables.

Sampling: Choosing Who to Study

How you select participants determines what your results can support. A biased selection method can't be rescued by a large sample, so this decision matters more than most researchers expect.

Methods that give every member of the population a known chance of selection are the ones that support formal inference. The overview of probability sampling covers the family. The four methods within it each suit a different situation:

Once the method is settled, the question is how many. That's a power analysis, and it's the one calculation that has to happen before data collection rather than after. See the guide to calculating sample size, which covers both the quantitative power analysis and the saturation logic qualitative studies use instead.

Describing Your Data

Every results section starts here. Descriptive statistics tell readers who was in your sample and what their values looked like, and they come before any test.

Three questions organize them. What's the typical value, how spread out is the data, and what shape is the distribution? For central tendency, see mean, median, and mode, which covers when each measure is the right choice. For spread, see standard deviation and variance. For shape, see skewness and kurtosis, which determines whether the mean or the median describes your data honestly.

Numbers alone compress a lot of information, so visualization belongs at this stage too. The guide to frequency distributions and histograms covers how to see a distribution's shape directly. For everything else, box plots, scatter plots, and choosing the right visualization gives a decision framework. The full treatment of this half of the field is in the descriptive statistics guide.

Distributions and Standardized Scores

One distribution underlies most of inferential statistics. Its properties are what let a test attach a probability to a result, which is why it's worth understanding before the tests themselves. See the guide to the normal distribution.

Standardized scores follow directly from it. Expressing a value as its distance from the mean in standard deviation units makes results from different scales directly comparable. The same standardization feeds into nearly every test statistic. See the guide to z-scores.

The Logic of Inference

Inference is the move from your sample to a claim about a population. Every specific test is a variation on one underlying procedure, so learning the procedure once makes the rest much easier.

That procedure is hypothesis testing. You state a null hypothesis of no effect, assume it's true, and ask how surprising your data would be under that assumption. See the guide to hypothesis testing and setting up the null and alternative.

The number that procedure produces is the p-value, and it's the most misread quantity in research. It's the probability of data at least as extreme as yours if the null were true. It is not the probability the null is true. See p-values explained for the misinterpretations that draw reviewer comments.

Every test can be wrong in two directions. A false positive reports an effect that isn't there. A false negative misses one that is. The guide to Type I and Type II errors covers both, along with the power that governs the second.

Testing isn't the only form of inference. Estimation gives a range of plausible values rather than a yes-or-no verdict, and many journals now prefer it. See confidence intervals. The full framework for this half of the field is in the inferential statistics guide.

Working on a dissertation or manuscript?

The statistical writing is where a subject-matter editor earns their place, because a reviewer who spots an overstated claim will question everything around it. Start with a free sample edit of your first 300 words and choose an editor in your field.

Choose Your Editor

Write the Analysis Plan Before You Collect Data

The single habit that prevents the most problems is deciding your analysis in advance and writing it down. Reviewers increasingly ask for it, and some journals require pre-registration.

A usable plan answers five questions. What are the hypotheses, stated as null and alternative? Which variables are predictors and which are outcomes? Which test will you run, and why does it suit your design and data type? What alpha level and power are you assuming, and what sample size do they require? And what will you do if an assumption fails?

Answering these in advance closes off the choices that damage credibility when made afterward. Consider picking a test once you've seen which one gives a significant result. Or switching from two-tailed to one-tailed to clear a threshold. Or deciding on subgroups after the fact. All three inflate your false positive rate. None of them look like misconduct while you're doing them, which is exactly why the plan matters.

The plan also makes the write-up faster. When the methodology section states the test, the inputs, and the reasoning before data collection, the results section becomes a matter of filling in numbers.

The Tests Themselves

With the framework in place, each test becomes a matter of matching the analysis to the question and the data.

Comparing groups

For one or two groups on a continuous outcome, see t-tests, which covers the one-sample, independent, and paired versions. For three or more groups, running repeated t-tests inflates your error rate, so see ANOVA instead. For categorical data, where you're comparing counts rather than means, see chi-square tests.

Measuring relationships

Correlation measures how strongly two continuous variables move together. For what the coefficient actually measures and how to read it, see correlation coefficients explained. For choosing between Pearson and Spearman and reporting the result in a manuscript, see using correlation coefficients in research papers.

Predicting outcomes

Regression extends correlation into prediction. Start with simple linear regression for one predictor, which introduces the slope, intercept, residuals, and R squared. Move to multiple regression when several predictors act at once, which is the usual case in real research. When the outcome is binary rather than continuous, see logistic regression.

When assumptions fail

The tests above assume things about your data, including normality and equal variances. When those assumptions don't hold, rank-based alternatives make fewer demands and remain valid. See when to use non-parametric tests.

Reporting Your Results

A correct analysis reported badly still gets sent back. Three things matter most here.

First, magnitude. Significance says an effect is detectable, not that it matters, and most journals now require an effect size alongside every test. See the guide to effect sizes for which measure pairs with which analysis.

Second, notation. Italics, leading zeros, degrees of freedom, and decimal conventions are things reviewers know by heart, and deviations get noticed immediately. See reporting statistical results in APA format.

Third, the errors themselves. Most statistical problems in published work are errors of interpretation and reporting rather than calculation, and they repeat predictably. See common statistics mistakes in academic papers for the full list organized by stage.

One practical matter often gets overlooked. Statistical notation in a manuscript has to be typeset correctly, and equation formatting is its own small skill. See how to write equations in Word.

One Question, All the Way Through

To see how the pieces connect, follow a single research question through the whole sequence. Fisher and Yao (2017) studied gender differences in financial risk tolerance using the Survey of Consumer Finances.

Design. Define the population, decide on the sampling method, and run a power analysis to set the sample size.

Describe. Report means and standard deviations for symmetric variables like age. Use medians and interquartile ranges for skewed ones like net worth. Check the shape before choosing.

Test. State the null and alternative, set alpha in advance, and run the analysis the data calls for. Comparing two group means calls for a t-test.

Interpret. Read the p-value for what it says about surprise under the null, and the effect size for whether the difference is large enough to matter.

Report. Write the result with its degrees of freedom, exact p-value, effect size, and confidence interval, in the notation your style guide expects.

Every stage in that sequence has a guide above. The order rarely changes, and skipping a stage is where most problems start.

Getting Your Statistics Reviewed

Statistical writing is where a second reader adds the most value, because the errors that matter are errors of claim rather than calculation. Software returns the right number. What goes wrong is what the researcher says it means.

Editor World's academic editing, dissertation editing, and journal article editing services include review of statistical reporting. You choose your own editor by field, so the person reading your results knows your discipline's conventions. A certificate of editing confirming human-only native English editing is available as an optional add-on for journals that require an AI-use disclosure.


Frequently Asked Questions

What statistics do researchers actually need to know?

Most researchers need four things rather than a full statistics education. Describe your data with the right measures. Choose an analysis that matches your design and data type. Interpret the output without overstating it. Report it in the notation your field expects. The specific tests vary by discipline, but that sequence doesn't.

What is the difference between descriptive and inferential statistics?

Descriptive statistics summarize the data you collected, using the mean, the standard deviation, and the shape of the distribution. They make no claims beyond your sample. Inferential statistics use that sample to reach conclusions about a population you didn't fully measure, with an account of how uncertain those conclusions are.

How do I choose the right statistical test?

It depends on what you're comparing and what data you have. One or two group means call for a t-test, three or more for ANOVA. Categorical counts call for chi-square. Relationships call for correlation or regression. When assumptions fail, use rank-based alternatives. The table above routes each case.

What is the most common statistical mistake in research?

Misinterpreting the p-value. It's the probability of data at least as extreme as yours if the null were true. It isn't the probability the null is true, and it isn't a measure of how large an effect is. The close relative is treating significance as importance, which is why an effect size now accompanies every test.

When should sample size be decided?

Before data collection. A power analysis uses the expected effect size, the power you want, and your alpha level to set the minimum number of participants. Running it afterward from observed results is circular and tells you nothing. A study that recruited too few participants can't be repaired by any later analysis.

Do I need to report an effect size with every test?

Most journals in psychology, education, and health now expect one, and many want a confidence interval too. Significance depends heavily on sample size, so a difference too small to matter can still reach significance in a large study. The effect size reports magnitude, which is what lets a reader judge whether the finding means anything. Name the specific measure you used.


Page last reviewed: September 2026. Content reviewed by Editor World editorial staff. Editor World, founded in 2010 by Patti Fisher, PhD, graduate of The Ohio State University, provides professional editing and proofreading services for academic researchers, doctoral candidates, faculty, business professionals, and authors worldwide. 100% human editing, no AI at any stage. BBB A+ accredited since 2010 with 5.0/5 Google Reviews and 5.0/5 Facebook Reviews. 16 years in business with 140 million+ words edited for over 8,000 clients in 65+ countries. Multiple Gold and Bronze Stevie Award winner. Native English editors from the United States, the United Kingdom, and Canada. Less than 5% of applicants are accepted to the editor panel. Recommended by the Boston University Economics Department, University of San Diego, University of Michigan, UCLA, University of Missouri, and more.