City digital twin · Part 6 · How it works

Inside the Newcastle twin: how a synthetic city is built, run and checked

This post opens up the Newcastle twin stage by stage, from open data to a scored run: the study area, the road and path network, the timetables, the traffic signals and level crossings, the synthetic population and their days, the simulation that lets them choose how to travel, how each run is checked against the real city, and what it takes to run one.

· ~14 min read · Part 6 of a series · with Pranav Dhoolia

The pipeline

Previously I showed what the twin looks like and how close it gets to the real city. Underneath is a pipeline of scripts. Each stage reads what the stage before produced, and each can be rerun from the raw downloads. Nothing in the chain is edited by hand.

SOURCES BUILT RUN CHECKED SEEN OpenStreetMaproads, paths, places timetables (GTFS)bus, train, tram, ferry 2021 censustables for 1,500 areas licence countsby council and age travel surveytrips, purposes, lengths ridership, countsOpal, road counts network368,230 links timetables mapped1,270 routes, all mapped signals, crossingsSCATS logic, 405 closures population612,667 residents daily planstours to real places,no mode given MATSim everyone plans,travels, scores,adjusts; 250rounds, 25% ofthe population 12 targets each on its ownbasis, a 20%gate, 143 moreheld back viewer3D map,scoreboard,run records every stage is a script; every file it writes is hashed and listed with its source and licence
FIG 1 - The pipeline, left to right. The sources are public; everything in the middle is built by scripts from them; the run is checked against observed figures the build never used as inputs.

Every file the pipeline writes, 968 of them, is listed in a manifest with its hash, row count, the script that produced it, its source, retrieval date and licence. Files built from OpenStreetMap carry its share-alike licence (ODbL); the rest are CC-BY. A check fails the build when a file changes without its manifest entry changing.

The study area

The twin covers five council areas: Newcastle, Lake Macquarie, Maitland, Cessnock and Port Stephens, 4,086 km² in all, made of 1,500 census areas (SA1s). The map extent is derived from those council boundaries with a 5 km margin, never typed in. That rule came from a mistake: early on, a study area typed into a script as a rectangle left 87 of the 1,500 census areas outside the road network for three stages of the build before anyone noticed. Now no coordinate may appear in the code, and a check counts any that do.

The network

OpenStreetMap is downloaded in ten themed layers (roads, rail, footpaths, buildings, places and so on) over tiles covering the extent, merged, and converted into a MATSim network with pt2matsim, in the local map projection so every distance is in metres.

Walking and cycling have their own network. 40,203 footpaths, cycle paths, stairs and tracks are added as walk and bike links, and each path's own OpenStreetMap tags (foot=, bicycle=, access=) override the default for its class: 4,484 paths were rewritten and 2,252 dropped that way. The network grew from 181,892 to 368,230 links. Speed limits come from the transport department's speed-zone data, and every link carries a gradient from the Copernicus 30-metre elevation model, which slows walkers and cyclists going uphill.

Close-up of the twin's network in central Newcastle: roads in dark lines, a dense web of footpaths, stairs and cycle paths in green alongside and between them, and the rail, light rail and ferry links in purple along the harbour.
FIG 2 - Central Newcastle, street by street, from the network the run used. The green is footpaths, paths, stairs and cut-throughs, which let a walking route be shorter than the driving one.

The timetables

Public transport runs from the transport department's open timetable feeds (GTFS). Fifteen feeds - the 2026 timetable, four earlier eras back to 2014 for the light rail question, and the scenario variants - are mapped onto the network in one pass, so every stop sits on a link and every route follows real streets and tracks, with no stop left unmapped.

Mapping runs once per network build, on purpose. The mapper is not reproducible: two mappings of the same inputs place nearly one route in five on different streets. Weekday, Saturday and Sunday timetables and the scenario variants are cut from the one mapped schedule instead of being mapped again, so any two runs being compared share the same routes. On a weekday that schedule has 1,448 bus trips, 332 train trips, 252 light rail trips and 107 ferry crossings, and each vehicle type carries its seated and standing capacity.

Line chart of timetabled departures per hour on a weekday: buses peak near 170 an hour around 07:00 and 230 around 15:00, with the afternoon school peak; trains, light rail and ferries run a steady 5 to 20 an hour through the day.
FIG 3 - Every timetabled departure on a weekday, by hour. The 15:00 bus peak is school services, which carry the afternoon school run in the simulation too.

Signals and level crossings

The fourteen signalised intersections on the light rail corridor are modelled as signal systems. New South Wales runs adaptive signals (SCATS), and the transport department would not release the operated phase plans or offsets. Rather than guess a fixed timing, the twin runs the adaptive logic SCATS publishes: after each cycle an intersection measures how saturated each approach was, moves its cycle length by 6 seconds towards a target saturation of 0.90, within 30 to 150 seconds, and re-splits the green time between approaches. The offsets that coordinate neighbouring intersections are not adapted, because they were not obtained, and that is recorded rather than filled in.

Freight trains are handled the same way: by what the road network actually experiences. Putting coal trains on the passenger model's tracks would invent interactions nobody observed, so the two level crossings on the network close their boom gates at the times trains pass. The closures come from the mapped passenger timetable plus the freight movements in a published survey of the coal chain: 405 closures on a weekday, written as 2,478 timed changes to the road links.

The people

Households and people are drawn for each census area so that, added up, they reproduce what the census published there: household size, dwelling type and cars per household; age and sex; work, study and income. The result is 612,667 synthetic residents against the 611,915 the 2021 census counted, with the household-size distribution within 0.6 percentage points of the census in every band.

Two things are measured rather than assumed. Driver licences come from the transport department's licence statistics divided by the resident population, for each council area and age band: 78% of 18 to 24 year olds hold one, 94% of 25 to 34 year olds, all of those aged 35 to 44, and 51% of those over 85. And households own the cars the census says they own, as named vehicles: 81,384 households, a third, have more licensed drivers than cars, so their drivers compete for them during the day. Each person's income, from the census band, sets how much a dollar matters to them when they choose.

Two charts. Left: an age and sex pyramid of the 612,667 synthetic residents, widest between 20 and 64. Right: the share holding a driver licence by age, rising from about 64% at 17 to 19 to near 100% from 35 to 75, then falling to about 50% over 85.
FIG 4 - The synthetic population by age and sex, and licence holding by age as the transport department's counts set it.

Their days

Each person gets a day built from home-based tours: out to work or study and back, a stop at the shops, taking a child to school, a work trip in the middle of the day. The number of trips is solved so that the population makes 3.473 trips per person per day, the travel survey's figure; it realises 3.470. Destinations are real buildings and places from OpenStreetMap, chosen with a distance model whose decay is solved so that trip lengths match the survey for each purpose and home council. People without a car choose nearer destinations: their decay per kilometre is 1.25 times faster, solved so that both groups together match the survey.

Some trips are tied to other people. A child's trip to school is bound to a parent with a licence and a car, and a lift is bound to a driver who owns one, so a passenger in the simulation can only ride with a real driver who is making that trip. What nobody is given is a mode.

Horizontal bar chart of a weekday's 2,275,503 planned trips by purpose: home and other errands 706,252, home and work 413,430, dropping off or picking up someone 349,846, home and shopping 345,341, between two places away from home 286,477, home and school or study 126,161, work-based trips 47,996, each with its average straight-line distance.
FIG 5 - A weekday's planned trips for the study area's residents, by purpose, with their average straight-line distance. Commutes are the longest; school trips the shortest.

The simulation

The engine is MATSim. Everyone's plans are executed together on the network, in a queue simulation where every mode takes up space: walkers and cyclists on paths and roads, cars, taxis and trucks in traffic, buses, trains, trams and the ferry on their timetables, stopping for signals and boom gates. Nothing teleports.

Each executed plan gets a score. Time spent at activities earns 6 utility points an hour; time spent travelling costs, with values of time from the literature by purpose ($18.60 an hour for a commute); waiting and walking to a stop cost more than riding, and each transfer costs the equivalent of 8 minutes, a value swept from 3 to 15. Money counts too: Opal fares from the published fare schedule, and parking priced by how many jobs are nearby, up to $3.20 an hour, at 7,710 parking facilities.

Between rounds, people revise their plans: 70% of revisions pick among remembered plans by score, 15% try a new route, 10% a new mode for a tour, 5% a new departure time. After 80% of the 250 rounds, only remembered plans are used, so the last rounds show where the population settled.

PLANS up to 8 remembered per person EXECUTE every mode in one queue simulation, real timetables SCORE time at activities, minus travel, waits, fares, parking REPLAN 70% pick a remembered plan 15% route · 10% mode 5% departure time 250 iterations innovation stops at 80% mode shares emerge; none is typed in
FIG 6 - The loop MATSim runs 250 times. The mode shares the scoreboard reads come out of it; none is an input.

Several behaviours needed code of their own, written as additions to MATSim in Java. A passenger getting a lift boards the driver's actual car, and the driver detours to pick them up. Taxis are a finite fleet - 800 at full scale - so a request can be refused and the traveller has to choose again. Household cars are shared, so a driver may find the car already out. Crowded buses and trains feel worse, cycling in heavy traffic feels worse, and hills slow walkers and cyclists. The constant each mode carries in the choice model is kept at its prior value, so that a mode which misses its target shows a missing behaviour instead of being tuned to fit.

A run simulates a quarter of the population, drawn by household so families travel together, with road capacity, seats and standing room scaled to match.

Checking a run

Each of the twelve modes has a target on the basis its real figure is published on: shares of residents' trips for car, lift, walk, bike, motorbike, taxi and bus, from the travel survey and, for taxis, industry trip counts; weekday Opal boardings at the train stations, on the light rail and at the ferry wharves; road counts for trucks; and the timetable and coal-chain survey for freight trains. Every 100 rounds the run reads all twelve, and a mode 20% or more off its target is the signal to stop and find its cause. Of 210 validation targets, 143 are held back and have never been read by the calibration code; they are opened once, at the end.

Runs are grouped into families: any change that can move the results starts a new family, and runs are never compared across one. Within a family, differences are read as bands. MATSim runs are not bit-reproducible - three runs of one build, same seed, diverge by the second round because parallel threads draw random numbers in a different order - though their mode shares came out the same.

DEVIATION FROM THE REAL FIGURE, EVERY RESULT SO FAR car ride walk taxi bike m'bike bus train light rail ferry inside 9 SepF32 +11% -41% -26% +202% +113% +12% +45% +221% -59% +40% 0 / 12 12 SepF35 +10% -42% -12% +131% +202% -6% -16% +55% -74% +66% 2 / 12 15 SepF35 +9% -41% -12% +143% +198% -6% -12% +60% -71% +45% 2 / 12 16 SepF35 +10% -40% -12% +133% +188% -5% -18% +55% -69% +56% 2 / 12 26 SepF37 +4% -14% -21% +130% +131% -7% -11% +188% -58% +119% 2 / 12 29 SepF38 +6% -25% -27% +169% +144% +466% -30% +118% -72% +116% 1 / 12 29 SepF39 +8% -25% -25% +161% +149% -4% -31% +119% -74% +95% 2 / 12 inside 10% (the goal)10 to 20%20% or more: the run stops, or the cause is fixed each row is a different model version, re-read today against today's targets
FIG 7 - Every run that has reached its last round since 9 September, re-read today by the twin's own reporter. Families differ from row to row, so this is a history of versions, not a comparison. The 29 September F38 row shows a change that made motorbike far worse; the next version fixed it.

Running it

Everything runs on one Windows workstation with 64 GB of memory: 16 threads for the simulation, a 48 GB Java heap for a quarter-sample run, and about 28 hours for 250 rounds. Before a full run, a four-round probe of about 45 minutes measures what the change costs, and the full run starts only with an approved, measured cost.

250 RUNS, BY HOW EACH ENDED short probes that price a change139 failed before finishing50 aborted, stopped by hand36 full runs that reached their last iteration12 full runs stopped at a gate or a ceiling12 left without a record1 41% of the machine-hours spent on runs went to runs that died
FIG 8 - Every run directory on disk, by how it ended. Most are probes; twelve full runs reached their last round.

The workstation was the least reliable part. Runs were killed by a forced Windows Update restart at round 237 of 250, by a disk controller crash, and by a shutdown from the Start menu. Each became a rule in the run launcher: it now refuses to start unless updates are paused past the run's expected end, it launches runs detached from any session, and it can resume a dead run from its last checkpoint.

The record

Every value a run can be sensitive to - 600 of them - is declared in a registry, with its units and where it came from. Values that are assumed or taken from the literature carry a range to sweep them over, or a stated rule for holding them fixed. A check counts any number a script decides for itself, and it stands at zero. Every decision is written into a dated record that is never rewritten, only added to.

600 DECLARED INPUT VALUES, BY SOURCE observed 39 measured 43 derived 56 literature 87 definition 152 assumed 223 blue: from this city's own data · grey: from published studies, or a definition such as a unit or a switch pink: assumed, and therefore swept over a stated range or held fixed with a stated reason
FIG 9 - The 600 declared input values by source. More than a third are still assumptions, each with a range. The registry is published as a reference.

What is still missing

  • The walk to a parked car. The model charges nothing for it, so people drive 59% of trips under a kilometre.
  • Public transport routes for everyone. Almost half of the attempts to plan a public transport journey find no route.
  • Choices for people without a car. On a trip no driver is bound to and no transit serves, they have walking, cycling or a taxi, and they cycle and take taxis far more than real people do.
  • The run-to-run band. How far readings of the same model spread across seeds is declared but not yet measured.
  • The signal offsets that coordinate intersections, which were not released.

The next post adds a second city on the same framework, and describes how both were built with AI coding agents.

Sources & anchors

  1. Figures 2 to 5: drawn from run 20260929T072135_250it_25pct's network, timetables and plans, and from the synthetic population and weekday trips files · figure 7: report_mode_ridership.py on each completed run · figure 8: the run index · map data © OpenStreetMap contributors (ODbL)
  2. Network, timetables and the one-mapping rule - network-and-inputs.md · footpaths - walk-and-bike.md · signals and crossings - signals-and-crossings.md
  3. Population and daily plans - population-and-demand.md · choices and replanning - seed-and-choice-set.md · lifts - ride-and-pairing.md · taxis - taxi-and-rideshare.md
  4. Targets, gates and families - public-transport-and-yardsticks.md · monitoring-and-gates.md · sampling-and-families.md · running costs - runs-and-economics.md
  5. MATSim - matsim.org · Horni, Nagel & Axhausen (eds.), The Multi-Agent Transport Simulation MATSim, Ubiquity Press (2016) - doi.org/10.5334/baw · pt2matsim - github.com/matsim-org/pt2matsim

City digital twin · Part 6 of the series · ← A synthetic Newcastle Two cities, one framework →