Victoria · school-age population by area · 2001 to 2025
Simple Victorian school demand forecasting from public data
We forecast the number of children aged 5 to 14 in each of 505 Victorian areas, using only ABS population estimates, and checked every forecast against what happened. Across 7,463 five-year forecasts (each area, each starting year), a cohort-survival model missed by a median 5.5% and a trend line by 7.8%. That is a 29% lower median error for the cohort model. In new housing estates the cohort model failed badly. No public population series shows an estate before the families move in.
Deeper Than Data · 9 October 2026
What the forecast is for
School planners need to know how many children will live in each area, and in which year levels, five to fifteen years from now. That is the lead time needed to buy land and build a school. The usual chain is: forecast the children in each small area, then multiply by a capture rate (the share of those children who will enrol at a given school or sector), then split by year level.
We tested only the first link, using public data. We measured how accurately the school-age population of each area can be forecast, and which method does best. We did not forecast enrolments at any school, because that needs enrolment data that is not public.
Data used
Estimated resident population (ERP) by area, age and sex, 2001 to 2025. ABS, Regional population by age and sex, 2025 (cat. no. 3235.0), from the ABS Data API, dataflow ERP_ASGS2021, retrieved 9 October 2026. Licensed under Creative Commons Attribution 4.0. Source: Australian Bureau of Statistics.
Areas are Statistical Areas Level 2 (SA2) on the ASGS Edition 3 (2021) boundaries. Victoria has 522 SA2s. Most hold between 3,000 and 25,000 people, about the size of a suburb or a country town.
Boundaries for the map are the ABS ASGS Edition 3 (2021) SA2 digital boundaries for Victoria, GDA2020, from abs.gov.au (file SA2_2021_AUST_SHP_GDA2020.zip), retrieved 9 October 2026. Licensed under Creative Commons Attribution 4.0. Simplified for the web, so borders are approximate.
Ages come in five-year bands at this level: 0 to 4, 5 to 9 and 10 to 14 for children, and women aged 15 to 44 for the ten-year forecasts. The ABS does not publish single years of age for SA2s in this series.
The method in plain words
Three forecasts were made for every area, plus a blend of two of them.
Naive. The number of children stays where it is now. A useful method has to beat this.
Trend. A line fitted to the area's own history of children aged 5 to 14, from 2001 to the forecast year, starting from the latest count. The slope is damped, so it flattens out over time instead of growing forever.
Cohort-survival. Children age into the next band. If an area had 1,000 children aged 0 to 4 five years ago and has 1,050 aged 5 to 9 today, the ratio of 1.05 sums up everything that happened to that group: families moving in and out, and a few deaths. The model applies the same ratio to today's 0 to 4 year olds to forecast the 5 to 9 year olds in five years, and the same for 5 to 9 into 10 to 14. This is the Hamilton-Perry method, a standard tool for small areas. Ten years out, the 0 to 4 band does not exist yet, so it comes from the area's ratio of young children to women aged 15 to 44, applied to a forecast of those women.
Combined. The simple average of the cohort and trend forecasts.
For the cohort model, forecasts one and three years ahead sit on a smooth path between today's count and the five-year cohort forecast. Every setting was fixed before any forecast was scored, so nothing was tuned to the test results.
How the back-test works
A back-test picks a past date, makes a forecast using only the data available up to that date, and scores it against what happened. Here the forecast date moves one year at a time from 2006 to 2024. From each date, each method forecasts each area's children aged 5 to 14 at 1, 3, 5 and 10 years ahead. That gives 30,347 scored forecasts across 505 areas.
Error is the absolute percentage error: how far the forecast was from the published figure, as a share of the published figure. The headline compares each method's median error across all the forecasts at a given horizon. At five years that is 7,463 forecasts (each area, each starting year), so half did better and half did worse.
Small areas are left out. An area is scored only when it had at least 100 children aged 5 to 14 on the forecast date. That drops parkland, industrial and airport areas, and new estates before their first families arrive.
Same cases for every method. A forecast is scored only when all four methods could make one, so the comparison is like for like.
Growth groups use no hindsight. Each area is grouped by how fast its school-age population grew in the five years before the forecast date: growing fast (3% a year or more), growing (0.5% to 3%), or stable or falling.
The cohort model has the lowest error up to five years out
Median absolute percentage error of forecasts of children aged 5 to 14, by how far ahead the forecast was made. 505 Victorian SA2s, forecasts made each year from 2006 to 2024. Lower is better.
Each point is the median error over thousands of forecasts; the table gives the count. Ten-year forecasts can only start from 2006 to 2015, so they rest on fewer cases. Source: ABS, Regional population by age and sex, 2025 (ERP_ASGS2021), CC BY 4.0. Analysis by Deeper Than Data.
Results
The cohort model beat the trend at every horizon on the median. One year ahead its median error was 1.5% against 2.0%. Five years ahead it was 5.5% against 7.8%, which is 29% lower. Ten years ahead it was 10.7% against 12.7%, 16% lower. Forecast by forecast, it beat the trend about six times in ten up to five years out, and in just over half of ten-year forecasts. The trend barely beat the naive forecast: its median error was within a few tenths of a point of "no change" at every horizon.
The trend line also leans low. Its median forecast was 2.7% under the actual figure at five years and 7.2% under at ten, because it damps growth that kept going. The cohort model's median miss was within about 1% either way.
Ten years ahead, the simple average of the two did best, at 9.9%. Averaging a method that overshoots in growth areas with one that undershoots cancels some of both errors.
Cohort errors are small in established areas and large in fast-growing ones
Median absolute percentage error by growth group, measured over the five years before each forecast. Choose a horizon.
Growing fast: the school-age population grew 3% a year or more in the five years before the forecast. Growing: 0.5% to 3% a year. Stable or falling: less than 0.5% a year. Source: ABS, Regional population by age and sex, 2025 (ERP_ASGS2021), CC BY 4.0. Analysis by Deeper Than Data.
In established areas, growing slowly or not at all, the cohort model cut five-year error by about a third: 4.6% against 6.9% for the trend. Those are most areas and most forecasts.
In fast-growing areas the cohort model loses its lead. At five years the two methods tie on the median, near 16%. At ten years the cohort model's median error is 44% against 29% for the trend, and some of its misses are enormous. When an estate goes from a handful of toddlers to hundreds in five years, the ratio between the two counts is huge, and the model projects that surge forward as if it will repeat. In Point Cook South, a forecast made in 2010 expected about 104,000 children in 2015. There were about 2,100.
Those few large misses dominate any total. Weighted by the size of each area, the five-year cohort error was 24% of all children, against 16% for the trend and 12% for no change at all. A planner who used the raw cohort model in growth corridors would have been badly wrong in the fastest-growing areas. Practitioners do not apply cohort ratios where the starting numbers are tiny. Those areas need a different method, set out below.
The map shows where this happens. The trend did better in pockets right across the state, but the areas where both methods failed sit almost entirely in Melbourne's new estates, to the west, north and south-east.
Where each method did better, five years ahead
Median error of five-year forecasts of children aged 5 to 14 in each SA2, trend minus cohort, in percentage points. Forecasts made each year from 2006 to 2020. Blue: the cohort model was more accurate. Orange: the trend was. Tap or hover an area for its figures.
Victoria
Greater Melbourne
Most areas lean blue. Outside the outlined areas, the cohort model was at least a point more accurate in 284 of 467 areas and the trend in 104, with 79 within a point. Taking each area's median error, the middle value across those areas was 5.2% for the cohort model and 7.3% for the trend. In the 38 outlined areas both methods missed by more than 25% on the median. Nearly all are new estates on Melbourne's fringe, such as Point Cook South, Rockbank and Tarneit North, plus a few inner-city apartment areas. Grey areas had fewer than 100 children aged 5 to 14, or none in some years, and are not scored. Source: ABS, Regional population by age and sex, 2025 (ERP_ASGS2021), CC BY 4.0. Boundaries: ABS ASGS Edition 3 (2021) SA2 boundaries (CC BY 4.0), simplified for the web. Analysis by Deeper Than Data.
Three areas, forecast from 2015
Children aged 5 to 14, each area on its own scale. Grey is the published estimate. The coloured lines are what each method forecast in 2015, for 2016 to 2020 and for 2025. The dotted line marks the forecast date.
We chose these three by hand to show one success and two ways of failing. They show what the errors look like. The evidence is the back-test above. Tarneit North had fewer than 100 children in 2015, so it is not in the scored results. Source: ABS, Regional population by age and sex, 2025 (ERP_ASGS2021), CC BY 4.0. Analysis by Deeper Than Data.
What public data cannot see
Estates not yet built. Tarneit North had 8 children aged 5 to 14 in 2015 and about 3,050 in 2025. Every method based on past population forecast roughly the same small number. Only a dwelling pipeline (land releases, subdivisions and building approvals) can see this coming.
Where children go to school. Population is not enrolment. The share of children at government, Catholic and independent schools, and which school they choose, varies by area and changes over time. School zones, new schools and travel time all move it. None of that is in ABS population data.
Year levels. The ABS publishes five-year age bands at this level, so this test cannot split forecasts by year level. Single-year detail needs Census counts or enrolment records.
Policy changes. Changes to school zones, school starting age, migration settings or housing policy can shift demand quickly. A model fitted to the past will not know about them until they show up in the data.
Hindsight in the history. The ABS revises its estimates after each Census. This test uses today's revised series, so each past forecast started from better data than a planner had at the time. Real forecasts would have done somewhat worse. Estimates after the 2021 Census will be revised again after the 2026 Census.
Small-area noise. SA2 estimates for single years are themselves modelled by the ABS. Small areas swing a lot from year to year, and the smallest were left out of the scores.
How this scales with agency data
With an education agency's own data, each gap listed above has a known fix. Each fix can be back-tested the same way as the methods on this page.
School enrolment census. Enrolments by school, year level and student home area give capture rates for each area and each school, and let the cohort model run on year levels directly: this year's Year 3 becomes next year's Year 4, scaled by the ratio seen in past years.
Dwelling pipeline. Lot releases, subdivision approvals, building approvals and precinct structure plans tell you where homes will be built and when. In new estates, forecast children from new dwellings, using the number of children per home observed in earlier estates at the same age. Hand over to the cohort model once the estate has settled.
Capture rates and travel time. Allocate each area's children to schools using observed capture rates, school zones and travel time, so a new school or a boundary change can be tested before it is made.
Accuracy testing every year. Store every forecast as it was made, on the data available at the time. Each year, score last year's forecasts against the new enrolment census and report the error by area type and horizon. If a method gets less accurate, the yearly score shows it within a year.
The back-test on this page is the template for that last step. It takes minutes to rerun when new data arrives.
How this was built
Target. ERP aged 5 to 14 (the 5 to 9 and 10 to 14 bands) for each SA2, at 30 June each year, 2001 to 2025.
Trend model. Least-squares log-linear slope fitted to every year from 2001 to the forecast date, anchored on the latest value and damped with a factor of 0.85 a year. With a damping factor of 0.75 the five-year median error is 7.5%; with no damping it is 8.9%. The cohort model wins either way.
Cohort model. Hamilton-Perry cohort-change ratios over the five years to the forecast date. Averaging the ratios over the last three forecast dates instead of using the latest one gives a five-year median error of 5.7%, against 7.9% for the trend on the same cases.
Scores. Median absolute percentage error, median signed error, and error weighted by area size. 30,347 forecasts across 505 SA2s. The other 17 Victorian SA2s are not scored. Thirteen are parkland, airports, industrial land or alpine areas that never had 100 children aged 5 to 14. Four (Craigieburn North West, Burnside Heights, Truganina North and Port Melbourne Industrial) recorded no children in some early years, which the trend model cannot fit, so they drop out for every method. Three of those are new estates, so the test slightly under-counts the hardest cases.
A guard tested after the fact. Switching to the trend wherever the 0 to 4 or 5 to 9 band had fewer than 100 children five years before the forecast cuts the size-weighted five-year cohort error from 24% to 12%, without changing the median. That rule was chosen after seeing the failures, so its score is flattering. It shows the direction a real system would take. It is not a tested result.
An earlier version. An earlier demonstration of this work, on 2016 boundaries and data to 2021, tested only the trend model. A claim from that period that a cohort model was 39% more accurate was not backed by a cohort back-test and was withdrawn. The figure measured here, on current boundaries and data to 2025, is 29% at five years.
The map. For each SA2, the median absolute percentage error of its five-year forecasts from each forecast date, for the cohort and the trend model, on the same cases as the tables. An area counts as about even when the two medians are within one percentage point. Outlined areas had a median error above 25% for both methods. The boundaries were simplified with mapshaper to about 190 KB.
Code. The back-test is about 300 lines of Python with no dependencies beyond the standard library, and runs in about a second. The code and the ABS extract are available on request.
An earlier version of this analysis was first published on plwp.net on 19 July 2026.
Deeper Than Data takes on forecasting and planning work as a data partner. Email go@deeperthandata.com.au.