ACADEMIC WRITING GUIDE

How to Analyse Quantitative Data: From Raw Numbers to Meaningful Results

By Acadelyra Editorial TeamPublished

You have collected your quantitative data.

Perhaps you distributed a questionnaire and received 350 responses.

You open SPSS, Excel, R, Stata, Jamovi, or another statistical package.

Then comes the difficult question:

What analysis am I actually supposed to do?

This is where many students make a costly mistake.

They begin clicking through statistical tests until something produces a p-value.

Quantitative data analysis should work in the opposite direction.

A stronger process is:

Research question → variables → research design → data quality → descriptive analysis → appropriate inferential analysis → interpretation

The statistical test comes after you understand what you are trying to answer and what kind of data you have.

University statistical guidance similarly emphasises that choosing an analysis depends on factors such as the research purpose and the nature of the outcome and predictor variables rather than on a universally preferred test.

The goal is not to use the most complicated analysis available.

It is to use an analysis that produces evidence capable of answering your research question.

Start with the research question

Before analysing anything, return to the question your study was designed to answer.

Compare these questions:

What proportion of first-year students report high academic stress?

Is academic stress different between first-year and final-year students?

What relationship exists between weekly study time and academic performance?

To what extent do study time, attendance, and academic self-efficacy predict academic performance?

All four could use quantitative data.

But they require different analytical reasoning.

The first asks for a description or estimate.

The second asks about a difference between groups.

The third asks about a relationship between variables.

The fourth asks about an outcome in relation to multiple predictors.

So do not begin with:

“Should I run an ANOVA?”

Begin with:

“What exactly am I trying to estimate, compare, associate, or predict?”

Your research question should provide that direction.

Map each research objective to an analysis need

Your objectives can help turn the broad research question into specific analytical tasks.

Suppose your study has these objectives:

  1. To describe students' levels of academic self-efficacy.

  2. To compare academic self-efficacy between students in two programmes.

  3. To examine the relationship between academic self-efficacy and academic performance.

These objectives imply different analysis needs.

Objective 1 requires descriptive analysis.

Objective 2 requires some form of group comparison, depending on the variables, design, assumptions, and exact question.

Objective 3 requires an analysis of association or relationship.

Your research aims and objectives should therefore connect directly to the analysis plan.

If an analysis cannot be connected to a research question, objective, hypothesis, or justified exploratory purpose, ask why you are running it.

Understand your variables before choosing tests

Statistical procedures make assumptions about the data they analyse.

One of the first tasks is therefore to identify what each variable represents and how it is measured.

Common distinctions include:

  • categorical variables;

  • ordinal variables; and

  • quantitative or continuous variables.

The exact terminology differs somewhat across textbooks and statistical traditions, but the principle remains important:

The nature of your variables constrains which analyses make sense.

UCLA's statistical guidance, for example, organises possible analyses according to the nature and number of dependent and independent variables and other assumptions.

Categorical variables

Categorical variables place observations into groups or categories.

Examples include:

Programme of study

  • Nursing

  • Engineering

  • Business

  • Education

or:

Employment status

  • Employed

  • Not employed

The numbers used to code categories in statistical software do not necessarily make the variable quantitative.

If you code:

0 = Not employed 1 = Employed

the numbers are labels representing categories.

It would not make sense to calculate an “average employment status” and interpret it as though employment were measured on a continuous numerical scale.

Ordinal variables

Ordinal variables have categories with a meaningful order.

For example:

Level of satisfaction

  • Very dissatisfied

  • Dissatisfied

  • Neither satisfied nor dissatisfied

  • Satisfied

  • Very satisfied

The categories have an order, but the distance between adjacent categories is not automatically known to be equal.

How ordinal responses should be analysed depends on the research design, measurement model, number of categories, analytical tradition, and assumptions being made.

Do not assume that every numbered response scale can automatically be treated as continuous merely because software stores the responses as 1, 2, 3, 4, and 5.

Quantitative variables

Quantitative variables represent numerical amounts for which arithmetic differences have substantive meaning.

Examples might include:

  • age in years;

  • examination score;

  • study hours per week;

  • income;

  • reaction time; or

  • number of completed assignments.

Even here, you still need to inspect how the variable behaves before selecting summary statistics or inferential procedures.

A numerical column in a dataset is not automatically ready for analysis.

Understand what your scores represent

Questionnaire studies often create scores from multiple items.

Suppose academic self-efficacy is measured using eight items.

Before calculating one total or average score, establish:

  • which items belong to the scale;

  • whether any items require reverse scoring;

  • how missing item responses are handled;

  • what higher scores mean;

  • whether the scoring procedure follows the instrument's guidance; and

  • whether combining the items is methodologically defensible.

Do not invent a scoring rule after seeing which version produces the most favourable result.

Article #29's questionnaire-design guide emphasised planning measurement and analysis before collecting responses. This is where that planning becomes essential.

Clean the data before testing hypotheses

Raw data should not go directly from collection into inferential analysis.

First inspect it for problems.

Depending on the study, data cleaning may involve checking for:

  • duplicate cases;

  • impossible values;

  • incorrect coding;

  • inconsistent category labels;

  • missing responses;

  • data-entry errors;

  • incorrect reverse scoring;

  • out-of-range values;

  • implausible dates;

  • unexpected units; or

  • cases that should not have been included.

Suppose age is restricted to participants aged 18–65, but one record says:

Age = 650

Do not simply include it in the mean because the software accepts the number.

Investigate whether it is:

  • a typing error;

  • a coding problem;

  • a genuine value entered in the wrong unit; or

  • evidence of another data-quality issue.

Cleaning decisions should be systematic and documented.

Do not silently “fix” inconvenient data

Data cleaning is not permission to change observations until the results look sensible.

If an observation is unusual, first determine why.

For example, a very high study-hours value could be:

  • a genuine extreme case;

  • a misunderstanding of the question;

  • hours entered per month rather than per week; or

  • a data-entry error.

Those possibilities require different responses.

Never delete a case simply because it weakens statistical significance.

Any exclusion or correction should have a defensible methodological reason.

Examine missing data

Missing data can affect both descriptive and inferential results.

Ask:

  • How much data are missing?

  • Which variables contain missing values?

  • Are particular participants missing many responses?

  • Does missingness appear concentrated in particular groups or questions?

  • How does the planned analysis handle missing cases?

Do not assume software automatically makes the best decision.

For example, some procedures may exclude any case missing one of the variables used in an analysis. UCLA's statistical examples illustrate how listwise deletion can reduce the effective sample size when values are missing.

Always check the actual N used in the analysis, not merely the number of people originally recruited.

Describe the sample before testing relationships

Before asking whether variables differ or relate, understand who is in the dataset.

Depending on the study, you might describe:

  • participant numbers;

  • age;

  • programme;

  • year of study;

  • gender where relevant and appropriately collected;

  • employment status;

  • response rate; or

  • other characteristics important to the research design.

The purpose is not to produce a table containing every variable you collected.

Describe characteristics that help the reader understand the sample and evaluate the evidence.

Begin with descriptive statistics

Descriptive statistics summarise what is present in the dataset.

For categorical variables, useful summaries often include:

  • frequencies; and

  • percentages.

For quantitative variables, you may consider measures of:

  • central tendency;

  • spread;

  • distribution; and

  • range.

Penn State's introductory statistics guidance similarly begins quantitative description with measures of centre and variability and recommends graphical examination of distributions and potential outliers.

Descriptive analysis is not a minor step to rush through before “real statistics.”

It helps you understand what you are actually analysing.

Frequencies and percentages

Suppose 300 students respond to a questionnaire.

Their programme distribution is:

  • Nursing: 120

  • Engineering: 90

  • Business: 60

  • Education: 30

The frequencies tell you how many respondents fall into each category.

Percentages make the distribution easier to interpret:

  • Nursing: 40%

  • Engineering: 30%

  • Business: 20%

  • Education: 10%

But always check what denominator the percentage uses.

If some participants did not answer the programme question, the percentage among valid responses may differ from the percentage among all participants.

Your table should make that clear.

Mean and median answer different questions

The mean is the arithmetic average.

The median is the middle observation after values are ordered.

Penn State describes both as measures of central tendency but distinguishes the mean as the numerical average and the median as the centre of the ordered distribution.

Consider these monthly incomes:

20, 22, 24, 25, 150

The mean is strongly influenced by the unusually high value.

The median better represents the centre of the ordered observations in this small example.

That does not mean:

Median is always better than mean.

It means the choice should reflect:

  • the distribution;

  • measurement properties;

  • research purpose; and

  • analytical conventions.

Do not report the mean automatically just because SPSS prints it.

Measures of spread matter

Two groups can have the same mean but very different variability.

Consider:

Group A: 48, 49, 50, 51, 52

Group B: 20, 35, 50, 65, 80

Both have a mean of 50.

But the observations in Group B are much more dispersed.

Measures such as:

  • standard deviation;

  • variance;

  • range; and

  • interquartile range

describe different aspects of spread.

Penn State's quantitative-data guidance explicitly pairs measures of centre with measures of spread rather than treating an average as a complete description.

A mean without information about variability can hide important features of the data.

Look at the distribution

Do not understand a variable only through one summary number.

Use appropriate visualisations and diagnostics.

Depending on the variable and purpose, these might include:

  • histograms;

  • boxplots;

  • dotplots;

  • bar charts;

  • scatterplots; or

  • other appropriate displays.

Visual inspection can reveal:

  • skewness;

  • unusual clusters;

  • outliers;

  • data-entry errors;

  • ceiling or floor effects;

  • non-linear relationships; or

  • unexpected gaps.

Penn State's statistics materials emphasise graphical displays because they reveal centre, variability, shape and potential outliers that a single statistic may conceal.

Investigate outliers rather than automatically deleting them

An outlier is an observation that appears unusually distant from the rest of the data under some criterion.

But an outlier can represent very different things:

  • an error;

  • a genuine rare observation;

  • a participant from a different population;

  • an unusual but important case; or

  • evidence that your assumed statistical model does not fit well.

Do not use:

“It was an outlier.”

as a complete justification for deletion.

Ask:

Why is it unusual, and what does my analysis assume about such observations?

Depending on the study, you may need sensitivity analyses, robust methods, transformations, model diagnostics, or a clearly justified exclusion.

Descriptive results do not automatically generalise

Suppose 62% of your sample reports high academic stress.

You can accurately say:

62% of respondents in the analysed sample reported high academic stress.

Whether you can infer that approximately 62% of the wider student population has high stress depends on issues such as:

  • sampling design;

  • representativeness;

  • non-response;

  • measurement quality; and

  • uncertainty.

Article #28's sampling guide explains why a large sample does not automatically guarantee population representativeness.

Analysis cannot repair a fundamentally inappropriate sampling design.

Move from description to the analytical question

Once you understand the data, return to what the study needs to test or estimate.

A useful starting distinction is whether you are primarily asking about:

  • a difference;

  • an association;

  • a prediction;

  • an estimate; or

  • a more complex model.

Then consider:

  • the outcome variable;

  • predictor or grouping variables;

  • whether observations are independent or paired/repeated;

  • variable measurement;

  • distributional/model assumptions;

  • sample size; and

  • the research design.

This is why statistical-test selection cannot be reduced to memorising:

“Two groups = t-test.”

UCLA's statistical-test guide shows that even apparently similar questions can lead to different procedures depending on the type of outcome, number and nature of predictors, dependence between observations, and assumptions.

Group comparisons

Suppose your question is:

Do mean academic-performance scores differ between students who are employed and those who are not employed?

You have:

  • one grouping variable with two groups; and

  • a quantitative outcome.

An independent-samples t-test may be one possible analysis under suitable conditions.

But now change the design:

Do students' stress scores differ before and after an intervention?

The measurements come from the same students at two time points.

Those observations are paired.

That is a different analytical structure.

UCLA's statistical guidance distinguishes independent-group comparisons from paired or repeated measurements for exactly this reason.

Do not choose a test merely by counting groups.

Understand how the observations were generated.

Association between categorical variables

Suppose you ask:

Is programme of study associated with whether students use the university counselling service?

Both variables are categorical.

Depending on the data and design, a chi-square test of association might be considered.

But again, the name of the test is not the starting point.

First ask:

  • What are the categories?

  • Are observations independent?

  • Are expected cell counts adequate for the intended procedure?

  • Is another method more appropriate for sparse data?

  • What effect measure will communicate the relationship?

The p-value is only one part of the result.

Correlation

Correlation measures the direction and strength of an association between variables under the assumptions of the particular correlation coefficient.

For a Pearson correlation, the coefficient ranges from −1 to +1.

Broadly:

  • positive values indicate that higher values of one variable tend to accompany higher values of the other;

  • negative values indicate that higher values of one tend to accompany lower values of the other; and

  • values near zero indicate little linear association.

But do not interpret a correlation coefficient without looking at the data.

A scatterplot can reveal:

  • non-linearity;

  • influential outliers;

  • separate clusters; or

  • restricted ranges

that make one correlation coefficient misleading.

UCLA describes Pearson correlation as an analysis of linear relationships between suitable variables and distinguishes it from non-parametric alternatives such as Spearman correlation when different measurement/assumption conditions apply.

Correlation does not prove causation

Suppose study time and academic performance are positively correlated.

You cannot automatically conclude:

Studying longer causes higher grades.

Other explanations may exist.

For example:

  • motivated students may both study more and perform better;

  • prior academic ability may influence both;

  • course difficulty may affect the relationship;

  • students performing poorly may change how much they study; or

  • measurement error may affect the observed association.

Causal claims require research designs and assumptions capable of supporting causal inference.

Statistical association alone does not provide that guarantee.

Regression asks a different question from correlation

Correlation can describe the strength and direction of an association between two variables.

Regression can address a broader set of questions.

Suppose your study asks:

To what extent do study time, attendance, and academic self-efficacy predict academic performance?

You now have:

  • an outcome: academic performance; and

  • several predictors: study time, attendance, and self-efficacy.

A regression model may allow you to examine how the outcome relates to those predictors simultaneously.

For example, you may want to know whether study time remains associated with performance after accounting for attendance and self-efficacy.

That is different from calculating three separate correlations.

Regression does not automatically establish causation

The word predictor can cause confusion.

In statistical modelling, a variable can be called a predictor without proving that it causes the outcome.

Suppose a regression finds that attendance predicts academic performance.

That does not automatically establish:

Increasing attendance will cause grades to improve by the estimated amount.

The interpretation depends on:

  • study design;

  • timing;

  • measurement;

  • confounding;

  • model specification;

  • assumptions; and

  • the broader causal reasoning.

Be particularly careful with observational cross-sectional data.

A sophisticated regression model does not magically transform an observational association into a causal effect.

Choose the regression model for the outcome

“Regression” is not one single statistical test.

The appropriate model depends partly on the outcome and research design.

For example, different approaches may be considered when the outcome is:

  • continuous;

  • binary;

  • count-based;

  • ordinal;

  • time-to-event; or

  • otherwise structured.

Multiple linear regression may be suitable for some continuous outcomes under appropriate assumptions.

Binary outcomes may require a different model, such as logistic regression.

Do not select linear regression simply because it is the regression procedure you recognise.

Check assumptions rather than merely naming them

Students sometimes write:

“The assumptions of the test were met.”

without showing what they checked.

Assumptions vary by statistical procedure.

Depending on the analysis, relevant issues might include:

  • independence;

  • linearity;

  • distribution of residuals;

  • homoscedasticity;

  • influential observations;

  • multicollinearity;

  • expected cell counts; or

  • other model-specific conditions.

Do not copy an assumption checklist from an unrelated statistical test.

Ask:

What assumptions does my chosen analysis actually rely on, and how can I assess whether they are reasonable for these data?

Normality is often misunderstood

Students frequently hear:

“Your data must be normally distributed.”

That statement is too vague.

Which data?

The raw outcome?

Each variable?

The residuals?

Within which groups?

And for which statistical procedure?

Different analyses make different distributional assumptions, and some procedures are more robust to particular departures than others.

Do not transform data, delete observations, or abandon an analysis simply because one generic normality test produces a significant p-value.

Examine the assumptions that matter for the actual model, using appropriate diagnostics and substantive judgment.

Statistical significance is not practical importance

Suppose a study with a very large sample finds that a predictor is associated with a 0.2-point difference on a 100-point academic-performance scale, with:

p < .001

The result may be statistically significant.

But is a 0.2-point difference meaningful?

That requires another question.

Statistical significance concerns evidence relative to a statistical model and null hypothesis.

It does not automatically tell you whether the observed difference or association is:

  • large;

  • educationally important;

  • clinically meaningful;

  • practically useful; or

  • theoretically important.

Always interpret the magnitude and context of the result, not only its p-value.

What does a p-value tell you?

A p-value is commonly misunderstood.

It is not:

the probability that the null hypothesis is true.

Nor is:

p = .03

equivalent to:

“There is a 97% probability that my hypothesis is correct.”

Broadly, under the statistical model and assuming the null hypothesis and relevant assumptions, the p-value reflects how incompatible the observed data—or something at least as extreme according to the test statistic—are with that null model.

That is much narrower than many students assume.

The American Statistical Association has specifically cautioned against scientific conclusions based only on whether a p-value crosses a particular threshold.

Do not worship the .05 threshold

You may encounter:

p < .05 = significant p ≥ .05 = not significant

That convention is common, but it should not turn statistical reasoning into a binary ritual.

Compare:

p = .049

and:

p = .051

Those results are extremely close.

It would be misleading to describe the first as a major discovery and the second as proof that nothing exists.

Interpret p-values alongside:

  • effect estimates;

  • uncertainty;

  • study design;

  • sample size;

  • assumptions;

  • prior evidence; and

  • substantive importance.

Confidence intervals communicate uncertainty

A point estimate gives you one estimated value.

A confidence interval provides a range constructed using a specified statistical procedure.

Suppose the estimated mean difference between two groups is:

4.2 points

with a 95% confidence interval:

1.1 to 7.3

The interval communicates uncertainty around the estimated difference.

Confidence intervals can help you assess:

  • the direction of plausible effects under the model;

  • the precision of the estimate; and

  • whether substantively small or large values remain compatible with the data and procedure.

A narrow interval generally indicates greater precision than a wide one, all else equal.

But avoid the common statement:

“There is a 95% probability that the true value lies inside this calculated interval.”

Under the usual frequentist interpretation, the 95% refers to the long-run performance of the interval-generating procedure, not a probability assigned to the fixed parameter after this particular interval has been calculated.

Effect sizes help describe magnitude

An effect size describes the magnitude of a difference, relationship, or effect in a form appropriate to the analysis.

Examples may include:

  • mean differences;

  • standardised mean differences;

  • correlation coefficients;

  • odds ratios;

  • risk ratios;

  • regression coefficients; or

  • other model-specific quantities.

There is no single universal effect-size statistic.

And labels such as:

small, medium, and large

should not be applied mechanically without considering the research context.

A “small” effect could still matter greatly in some settings.

A statistically large effect could be practically irrelevant in another.

Interpret magnitude substantively.

Non-significant does not mean “no effect”

Suppose your analysis produces:

p = .18

Do not automatically conclude:

“There is no relationship.”

A non-significant result may reflect:

  • a genuinely small or absent association;

  • limited statistical power;

  • imprecise measurement;

  • substantial variability;

  • a small sample;

  • model misspecification; or

  • other sources of uncertainty.

Look at the estimated effect and its confidence interval.

A wide interval may indicate that the study is compatible with a range of potentially meaningful effects and simply lacks precision.

A better conclusion may be:

The study did not provide sufficiently strong evidence of an association under the specified analysis.

The exact wording should reflect the design, estimate, uncertainty, and statistical framework.

A significant result does not prove your hypothesis

Suppose your hypothesis predicts a positive relationship between self-efficacy and academic performance, and the analysis gives:

p = .02

That does not mean the hypothesis has been “proven.”

Statistical evidence can be consistent with a hypothesis without establishing it as certain truth.

Likewise, a hypothesis should not be judged only through the p-value while ignoring:

  • effect direction;

  • magnitude;

  • confidence interval;

  • measurement quality;

  • study design; and

  • assumptions.

Acadelyra's research hypothesis guide emphasises that hypotheses should specify relationships the study can genuinely test. Analysis completes that logic—it does not convert uncertainty into proof.

Multiple testing increases the chance of misleading findings

Suppose you collect 30 variables and test every possible pair until something gives:

p < .05

Even when no meaningful relationships exist, repeated testing increases the chance that some results will cross a conventional significance threshold by chance.

This is one reason the analysis should be planned around the research questions and hypotheses rather than becoming a search for significant results.

Depending on the research context, approaches to multiple comparisons or multiplicity may need consideration.

The appropriate response is not simply:

“Never run more than one test.”

It is to distinguish:

  • planned analyses;

  • justified secondary analyses; and

  • exploratory analyses,

and interpret them appropriately.

Do not change the analysis because you dislike the result

Suppose your planned analysis gives a non-significant result.

You then try:

  • a different outcome definition;

  • several transformations;

  • different exclusion rules;

  • multiple subgroup analyses; and

  • several alternative tests

until one gives p < .05.

That creates a serious risk of misleading inference.

Analytical flexibility should be driven by methodological reasoning, diagnostics, sensitivity analysis, or clearly labelled exploration—not by a desire to obtain significance.

Where possible, define important analytical decisions before examining the final results.

Report exploratory analyses honestly

Exploratory analysis can be valuable.

Unexpected patterns may generate important new questions.

The problem arises when an analysis discovered after looking at the data is presented as though it had been planned from the beginning.

Be transparent.

You might distinguish:

Primary analysis: specified to address the original research question.

Secondary analysis: planned but not primary.

Exploratory analysis: developed after examining the data.

Transparency helps readers judge the evidence appropriately.

Interpret software output selectively

Statistical software can produce pages of output.

Your dissertation does not need every number.

Suppose a regression package prints:

  • model summaries;

  • ANOVA tables;

  • coefficients;

  • residual diagnostics;

  • collinearity statistics;

  • case diagnostics;

  • covariance matrices; and

  • several optional plots.

Do not copy the entire output into your findings chapter.

Identify the information needed to answer the research question and demonstrate that the analysis was conducted appropriately.

Software output is evidence for your analysis, not the analysis itself.

Know what each reported number means

Never report a statistic simply because it appears in the output.

If you report:

β = .37

you should know:

  • whether that is standardised or unstandardised;

  • what the predictor represents;

  • what the outcome represents;

  • how the coefficient should be interpreted;

  • what uncertainty accompanies it; and

  • whether the model makes that interpretation defensible.

Likewise, if you report an odds ratio, correlation, F statistic, t statistic, or chi-square statistic, understand what it contributes.

If you cannot explain a statistic in ordinary language, you probably need to understand it better before putting it in the dissertation.

Tables should communicate—not dump output

A useful results table should help readers understand the evidence efficiently.

It might present:

  • variable names;

  • sample sizes;

  • descriptive statistics;

  • estimates;

  • confidence intervals;

  • effect measures; or

  • relevant inferential statistics.

But do not reproduce software screenshots or enormous tables containing irrelevant decimal places.

Use consistent formatting.

Label variables clearly.

Explain abbreviations.

Report enough precision to support interpretation without pretending your measurements are more precise than they really are.

Visualisations should answer a question

Charts are not decorations.

A good visualisation can reveal patterns that a table makes difficult to see.

For example:

  • a scatterplot can show the form of a relationship;

  • a boxplot can compare distributions across groups;

  • a histogram can show the shape of a quantitative variable; and

  • an appropriate interval plot can communicate estimates and uncertainty.

Choose a visual because it helps answer or diagnose something.

Avoid unnecessary:

  • 3D effects;

  • decorative colours;

  • misleading axes;

  • overcrowded labels; or

  • chart types that distort comparisons.

The reader should understand the data more clearly after seeing the figure.

Write results as evidence, not as a list of tests

A weak results section may read:

A t-test was conducted. A correlation was conducted. A regression was conducted.

That tells the reader what buttons you pressed.

A stronger structure follows the research questions or objectives.

For example:

Objective 1: Describe academic self-efficacy

Report the relevant descriptive evidence.

Objective 2: Compare self-efficacy between programmes

Report the appropriate comparison, effect estimate, uncertainty, and interpretation.

Objective 3: Examine the relationship between self-efficacy and performance

Report the association or model relevant to that objective.

The statistics should serve the research argument.

Separate results from unsupported explanation

Suppose students with higher attendance have higher average grades.

Your results may establish an association.

They do not automatically explain why it exists.

Avoid writing:

Students with higher attendance performed better because attending lectures increased their understanding.

unless your design and evidence actually support that mechanism.

A more defensible results statement may be:

Higher attendance was associated with higher academic-performance scores in the analysed sample.

Possible explanations can then be discussed carefully in relation to theory, prior research, design limitations, and alternative interpretations.

A worked quantitative-analysis example

Suppose your research question is:

What relationship exists between weekly study time and academic performance among first-year undergraduate students?

Step 1: Identify the variables

Study time: hours studied per week.

Academic performance: assessment score measured on a 0–100 scale.

Step 2: Inspect and clean the data

Check:

  • impossible study-hour values;

  • assessment scores outside 0–100;

  • duplicate cases;

  • missing values;

  • coding errors; and

  • unusual observations.

Document any corrections or exclusions.

Step 3: Describe the sample

Report the analysed sample size and relevant characteristics needed to understand the study.

Step 4: Describe the variables

For each variable, inspect:

  • centre;

  • spread;

  • range;

  • distribution; and

  • potential outliers.

Use appropriate numerical and graphical summaries.

Step 5: Visualise the relationship

A scatterplot of study hours against assessment scores can help reveal:

  • direction;

  • approximate form;

  • outliers;

  • clusters; and

  • whether a linear relationship appears plausible.

Step 6: Choose the analysis

If the research question, measurement, design, and assumptions support examining a linear association between the two quantitative variables, Pearson correlation may be one possible analysis.

If those conditions are not appropriate, another method may be required.

The test is chosen after evaluating the analytical structure—not simply because there are two numerical columns.

Step 7: Report the estimate and uncertainty

Do not report only:

p = .01

Report the correlation coefficient, sample size, relevant uncertainty where appropriate, and p-value according to the reporting conventions required by your discipline.

Step 8: Interpret cautiously

Suppose the analysis indicates a moderate positive association.

A defensible interpretation might be:

Students reporting more weekly study time tended to have higher academic-performance scores in this sample.

Do not automatically write:

Increasing study time causes higher grades.

unless the research design supports that causal claim.

Step 9: Connect the result to the research question

Finish by explaining what the analysis contributes to answering the original question.

The result should close the loop:

Research question → variables → analysis → evidence → interpretation

Common quantitative-analysis mistakes

Choosing a test before understanding the question

Statistical procedures should follow the analytical problem.

Treating coded categories as continuous numbers

Numerical labels do not automatically create quantitative measurement.

Skipping data cleaning

Errors discovered after analysis can invalidate results and waste substantial work.

Reporting only means

Centre without spread can hide important differences in distributions.

Deleting outliers automatically

Unusual observations need investigation and methodological justification.

Ignoring missing data

The analysed sample may be much smaller or systematically different from the recruited sample.

Assuming normality without understanding the model

Check the assumptions relevant to the actual procedure.

Treating p < .05 as the entire result

Magnitude, uncertainty, design, and practical meaning matter.

Treating p ≥ .05 as proof of no relationship

Absence of sufficient evidence is not automatically evidence of absence.

Confusing correlation with causation

Association alone cannot establish a causal mechanism.

Running many tests until something is significant

This increases the risk of misleading findings.

Reporting every number software produces

Include statistics that help answer the research question and evaluate the analysis.

Copying software screenshots into the dissertation

Create clear tables and figures using the reporting conventions appropriate to your field.

Interpreting before checking assumptions

A model that poorly represents the data may produce misleading estimates.

Using complicated statistics to look advanced

Complexity is not quality.

The simplest analysis that appropriately answers the research question is often preferable to an unnecessarily elaborate model.

A quantitative-data-analysis checklist

Before finalising your analysis, ask:

  1. Does every major analysis connect to a research question, objective, hypothesis, or clearly identified exploratory purpose?

  2. Do I understand what each variable represents?

  3. Have questionnaire scales been scored correctly?

  4. Have I checked for data-entry and coding errors?

  5. Have I investigated missing data?

  6. Do I know the actual sample size used in each analysis?

  7. Have I described the sample appropriately?

  8. Have I examined relevant distributions visually and numerically?

  9. Have I investigated unusual observations rather than automatically deleting them?

  10. Does the selected statistical procedure match the outcome, predictors, and research design?

  11. Have I considered whether observations are independent, paired, clustered, or repeated?

  12. Have I checked assumptions relevant to the chosen analysis?

  13. Am I reporting effect estimates rather than only p-values?

  14. Am I communicating uncertainty appropriately?

  15. Have I interpreted statistical significance separately from practical importance?

  16. Have I avoided treating non-significance as proof of no effect?

  17. Have I considered multiplicity where many tests were performed?

  18. Have I distinguished planned analyses from exploratory analyses?

  19. Do I understand every statistic I report?

  20. Are tables and figures clear rather than copied directly from software?

  21. Have I avoided causal claims that the study design cannot support?

  22. Does the final interpretation actually answer the research question?

If several answers are no, the analysis probably needs more work before you begin writing conclusions.

Final takeaway

Quantitative data analysis is not a competition to find the most advanced statistical test.

It is a structured process for turning numerical observations into evidence that can address a research question.

Start with the question.

Understand the variables.

Clean the data.

Investigate missingness and unusual observations.

Describe the sample and distributions.

Then choose an analysis that matches:

  • what you are trying to answer;

  • how your variables are measured;

  • how the observations were generated;

  • the research design; and

  • the assumptions of the statistical procedure.

Do not stop at the p-value.

Examine the size and direction of the result.

Communicate uncertainty.

Consider whether the result matters practically as well as statistically.

Distinguish association from causation.

Be transparent about missing data, exclusions, exploratory analyses, and analytical decisions.

And use statistical software as a tool for performing calculations—not as a substitute for understanding them.

The most useful question during quantitative analysis is therefore not:

“Which statistical test will give me a significant result?”

It is:

“What analysis will give me the most defensible evidence for answering my research question?”

When you can explain why you chose the analysis, what its results mean, what uncertainty remains, and what conclusions the research design can reasonably support, you have moved beyond simply producing statistical output and into meaningful quantitative analysis.