Enrollment for the ISI-CMI is Live! Apply Now →
Back to Master Directory
Student analysing data distributions and regression output
Special Courses Rs 900 / session

Statistics: Basic to Advanced

EG

EduGlobal Masterclass

1-on-1 Remote Mentorship

Share:
For
Class 9 - Undergraduate
Level
Beginner to Advanced
Duration
24 weeks
Each week
4 hours
Batch
Group class, maximum 8 students
Mode
Online

The subject everyone uses and few interpret correctly

Statistics is unusual in that the computational side is widely taught and the interpretive side is widely misunderstood — including by people who have passed statistics examinations.

This course covers the mathematics properly, from descriptive measures to generalised linear models, and spends unusual time on what each result licenses you to say. That second part is where the value is.

What a p-value is not

A p-value is the probability of observing data at least as extreme as yours, assuming the null hypothesis is true.

It is not the probability that the null hypothesis is true. It is not the probability your result arose by chance. It is not a measure of effect size — a tiny, meaningless difference will produce a minuscule p-value given a large enough sample.

These three misreadings are everywhere, including in published work, and a student who holds the correct interpretation reads research very differently. We treat this as core content rather than as a cautionary footnote.

Confidence intervals say less than you think

A 95 per cent confidence interval is not the range in which the true value probably lies, though that is how nearly everyone reads it.

It is produced by a procedure that captures the true parameter in 95 per cent of repeated samples. The particular interval you computed either contains the parameter or it does not. The interpretation people actually want — a probability statement about the parameter given this data — is the Bayesian credible interval, which is one good reason to do the Bayesian unit.

Sampling distributions: the idea everything rests on

If a student does not understand the sampling distribution of a statistic, nothing after it makes sense; they are following procedures.

The move is this: a statistic computed from a sample is itself a random variable, with its own distribution across possible samples. Standard error, confidence intervals and every hypothesis test are consequences of that single idea. Students habitually rush this unit to get to tests, and then find the tests arbitrary. We do not let them.

Regression, with the assumptions checked

Simple and multiple linear regression, least squares, residual analysis, R-squared and its limits.

The emphasis is on diagnostics, because fitting a model is easy and checking whether it was appropriate is the part that gets skipped. A regression with excellent R-squared and patterned residuals is telling you the model is wrong, and a student who never plots residuals never hears it. We also take the causation question seriously: what correlation does not establish, and what kind of design or argument would.

Bayesian inference, as a genuine alternative

Prior, likelihood and posterior; conjugate priors including beta-binomial and normal-normal; credible intervals.

We teach this as a coherent framework rather than as an advanced extra, and we are explicit about where the two approaches genuinely differ. The Bayesian posterior answers the question people instinctively ask — what should I now believe about this parameter — at the cost of requiring a prior. That trade is worth understanding rather than inheriting a tribal position about.

Why so much published research does not replicate

The final unit is the one students remember, and it follows directly from techniques taught in every introductory course.

Test twenty hypotheses at the 5 per cent level and you expect one significant result from pure noise. Decide which analysis to run after seeing the data — the garden of forking paths — and you can reach significance almost at will without any conscious dishonesty. Add publication bias, where null results are never submitted, and the published literature becomes a biased sample of what was actually found.

We cover Bonferroni and false discovery rate control, pre-registration, and what a reader should look for. A student who has done this unit evaluates evidence better than most graduates.

Who this suits

Students from Class 9 upward through undergraduate study with school algebra and basic probability, which is reviewed within the course. It suits students heading for economics, psychology, medicine, data science or any field that argues from data — which is most of them now.

Classes are live and online in groups of at most eight, with written work corrected individually.

What students will be able to do

  • Describe data honestly, including when a mean is the wrong summary to report
  • Understand sampling distributions, which is the idea the whole subject turns on
  • Build and interpret confidence intervals, and state precisely what they do not claim
  • Run hypothesis tests and interpret a p-value correctly, which most graduates cannot
  • Apply maximum likelihood estimation and judge estimators by bias, consistency and efficiency
  • Fit and diagnose regression models, including checking the assumptions rather than assuming them
  • Work Bayesian inference with conjugate priors and understand where it differs from the frequentist account
  • Recognise multiple testing, p-hacking and the reasons many published results do not replicate

Course structure

  1. BASIC 1: Describing data
    Types of data; frequency distributions and histograms; mean, median and mode and when each misleads; range, variance, standard deviation and the interquartile range; skewness; outliers and what to do about them.
  2. BASIC 2: Sampling and study design
    Populations and samples; random, stratified and cluster sampling; sampling bias and non-response; the difference between an observational study and an experiment, and why that difference governs what you may conclude.
  3. BASIC 3: Probability for statistics
    The distributions statistics depends on: binomial, Poisson, normal, t, chi-squared and F; the standard normal and z-scores; a review rather than a first course.
  4. INTERMEDIATE 1: Sampling distributions
    The sampling distribution of a statistic; the standard error; the central limit theorem applied to inference. This is the idea the entire subject rests on and the one students most often skip past.
  5. INTERMEDIATE 2: Estimation
    Point estimation; bias, consistency and efficiency; maximum likelihood estimation; the method of moments; confidence intervals for means and proportions and what their coverage actually means.
  6. INTERMEDIATE 3: Hypothesis testing
    Null and alternative hypotheses; Type I and Type II errors; significance level and power; z-tests, t-tests, chi-squared tests and ANOVA; the correct interpretation of a p-value and the three standard misreadings.
  7. INTERMEDIATE 4: Correlation and regression
    Covariance and correlation; simple and multiple linear regression; least squares; residual analysis and the model assumptions; R-squared and its limits; why correlation does not establish causation and what would.
  8. ADVANCED 1: Bayesian inference
    Prior, likelihood and posterior; conjugate priors including beta-binomial and normal-normal; credible intervals and how they differ from confidence intervals; where Bayesian and frequentist answers diverge and why.
  9. ADVANCED 2: Advanced modelling
    Logistic regression; generalised linear models; model selection with AIC and BIC; cross-validation; the bias-variance tradeoff; regularisation through ridge and lasso.
  10. ADVANCED 3: Multiple testing and the replication problem
    The multiple comparisons problem; Bonferroni and false discovery rate control; p-hacking and the garden of forking paths; publication bias; pre-registration. Why a large share of published findings do not replicate.

Who this is for

School algebra and basic probability. The probability this course needs is reviewed within it.

Students join this programme from India, United States, United Kingdom, Singapore and United Arab Emirates.

Common questions

What does a p-value actually mean?

The probability of observing data at least as extreme as yours, assuming the null hypothesis is true. It is not the probability that the null hypothesis is true, not the probability your result occurred by chance, and not a measure of effect size. Those three misreadings are extremely common among people who have passed statistics courses, which is why we spend real time on this rather than a footnote.

Is a confidence interval the range the true value probably lies in?

Not in the frequentist framework, though almost everyone reads it that way. A 95 per cent confidence interval is produced by a procedure that captures the true parameter in 95 per cent of repeated samples. The interval you actually computed either contains the parameter or does not. The interpretation people want is the Bayesian credible interval, which is one reason the Bayesian unit is worth doing.

Why cover the replication problem?

Because it is the most consequential statistical story of the last two decades and it is a direct consequence of techniques taught in every introductory course. Test enough hypotheses and some will reach significance by chance; choose your analysis after seeing the data and you can reach significance nearly at will. A student who understands multiple testing and the garden of forking paths reads published research very differently, and more accurately.

Is this a mathematics course or a data analysis course?

A mathematics course with its applications kept in view. We derive where deriving teaches something and interpret throughout, because a student who can compute a t-test but cannot say what it licenses has learned the less useful half. It is not a software course, though the methods are standard and transfer directly to any statistical package.

How does this fit with the probability course?

Probability is the mathematics of uncertainty given a model; statistics is the problem of inferring the model from data. The probability this course needs is reviewed within it, so you can start here. Students who take both find the foundations considerably firmer, and the probability course is the better starting point for olympiad work.

End of Syllabus. Apply for Admission

Book a trial

Three sessions with the mentor who would teach the full course. Nothing is charged until your slot is confirmed.

INR 599 from, by class
Request a trial slot Browse other courses
  • A diagnostic, a taught class and written feedback
  • Taught by the mentor who leads the course
  • You pick the slot from our live calendar
  • Pay only after the slot is confirmed
For
Class 9 - Undergraduate
Level
Beginner to Advanced
Duration
24 weeks
Each week
4 hours
Batch
Group class, maximum 8 students
Mode
Online
Talk to us

Ask a question, or book a trial

Tell us about the student and we will reply with an honest view of whether this programme fits. If you would like to see the teaching first, add a trial class.

  • No obligation - send the enquiry without booking anything
  • Three sessions if you do book: a diagnostic, a taught class and feedback
  • Taught by the mentor who would lead the full course
  • You choose the slots from our calendar after payment
Trial fee by class, if you book
Class 1 to 5 3 sessions INR 599
Class 6 to 8 3 sessions INR 799
Class 9 and 10 3 sessions INR 899
Class 11 and 12 3 sessions INR 999
Graduation and above 4 sessions INR 1,099

Sending an enquiry is free. The fee applies only if you tick the trial box below.

Send an enquiry

A parent or guardian should fill this in. We reply within one working day.

Not ready to pay yet? Leave the box unticked and just press Send enquiry. We will still receive your details and reply within one working day, and you can book a trial later whenever you are ready.

Free resources, straight to your inbox

Problem sets, strategy guides and olympiad registration deadlines — sent when they matter, never more than twice a month.