Most forecasts of humanoid robots model the robot. When will it be cheap enough, when will it be dexterous enough, when will it walk without falling over. Then they draw a curve.
That gets you a forecast for factories. It gets you nothing useful for an aged care home, because a robot that is ready in a lab and has no safety standard, no insurer, no funding line and no nurse willing to hand over the medication round is not ready. It is a demo.
So the engine models the chain instead of the robot. Here is how it works, and here is exactly where you should distrust it.
The chain
For each job we want a robot to do, we list everything that has to be true first. Not the interesting things. Everything. For aged care that comes to thirty-nine prerequisites across six areas: the technology itself, how reliable and safe it is in the field, the economics, the rules and who carries the liability, whether anyone is willing to delegate the job, and the environment the robot has to work in. A dementia ward is an environment. So is a private home with a cat and a loose rug.
Some prerequisites are shared. Battery life matters to every job. Some are specific. The ability to take a frail person’s weight matters only to the mobility job. Some gate a job outright. If there is no safety standard for a machine that touches a person, the transfer robot is not being piloted in an Australian facility no matter how cheap it is. We mark those.
How a prerequisite is scored
Every prerequisite gets five scores, 0 to 9, on the same scale, so a number means the same thing everywhere.
| 0 | 3 | 5 | 7 | 9 | |
|---|---|---|---|---|---|
| Maturity | does not exist | works in a lab | could be piloted | deployable with supervision | commodity |
| Reliability | unknown, or fails | demo grade | pilot grade | industrial grade | boring |
| Cost pressure | no business case | weak | plausible | strong | overwhelming |
| Regulatory readiness | vacuum | grey zone | clear pathway | standard adopted | routine |
| Adoption | none | first pilots | early commercial | mainstream | everywhere |
Higher is always more favourable to adoption, including for the awkward rows. A regulatory vacuum scores low. A dementia ward scores low as an environment. That takes a moment to get used to and then it is obvious.
Every score comes with a written reason and its sources, and with a stated uncertainty. Where we could not find evidence, that absence is the finding and the score is low. Nobody publishes reliability data for any humanoid. That is not a gap in our research. It is a two out of nine.
Why the weakest link sets the pace
The obvious way to combine thirty-nine scores is to average them. We do not, because an average lets a cheap robot compensate for a missing safety standard, and in the real world it cannot.
Instead the combination leans on the weakest links. Each job’s score is part weighted average and part the average of its three lowest prerequisites, with the mix set by how unforgiving the job is. Medication reminders can tolerate a weak link, because a missed prompt is recoverable. A transfer cannot, because a single failure is an injury. And any gating prerequisite that is below pilot grade caps the whole job at its own score.
Today that cap is the safety standard, for both physical jobs. Fix that and the next weakest link becomes the cap, which is the robot’s ability to safely touch a person. Fix that and it is the reliability record. The working shows the queue.
Speeds, and the simulation
Where things are today is fairly well known. How fast they move is not, and that is where the uncertainty in the timeline lives.
Each prerequisite has an expected speed in score points per year and an uncertainty on that speed. Progress slows as a score gets near the top of the scale. Then the whole thing is run forward, year by year, a few thousand times, with every starting score and every speed drawn at random from its range, under a mix of scenarios: the base case, an accord between industry and unions, an unmanaged default where cheap robots arrive and the rules do not, and a technology stall. Chance events fire inside the runs. A publicised injury sets regulation back three years in about one run in ten per year.
What comes out is not a date. It is a band: the year by which a job reached pilot grade in a tenth of runs, half of runs and nine tenths of runs. If most runs never got there by 2045, the table says so. It does not invent a year.
One thing about those bands. The median year trails the single best-guess path by about four years. That is not a mistake. When many uncertain things all have to be true, the typical outcome is later than the best guess for each of them. It is also the most useful thing the model says: reducing uncertainty moves the timeline as much as improving the robots.
What the model cannot do
It treats scores as if they were measurements. They are judgements on a 0 to 9 scale, and a weighted average of judgements that are not really commensurable. We keep the weights visible for that reason.
It does not know that adoption feeds back into speed. A hundred robots in facilities generate the reliability data that unlocks the next hundred. The model has a crude version of this and no more.
It assumes the prerequisites move independently, which they do not. We correct for that with a shared shock, a quarter of the uncertainty on every prerequisite is common to all of them, and another quarter is common within its area. Without that correction the model guarantees that one prerequisite stalls and pins everything to it for twenty years. That was an artefact of the maths and we removed it. The correction is a knob and it is worth turning.
And it can be wrong in the way any model with starting estimates can be wrong. Bad starting numbers produce confident bad answers. Every number is therefore shown, with its reason and its source, so that you can find the one you disagree with.
Betting in advance
At the end of the report we name specific things that might happen next and commit, now, to what each one will do to the model. If Standards Australia adopts the personal-care robot standard by the end of 2028, the regulatory score on that prerequisite goes to five. If no vendor has published a reliability record by the end of 2027, the speed on that prerequisite gets cut.
Once you know how something turned out it is trivially easy to explain why it was always the likely outcome. Writing the update down beforehand makes it mechanical. When a signpost resolves, the number is applied and the timeline is re-run, whether or not the new answer flatters the old one. Every resolution goes on the record with what the model said and what happened.
A forecast without a scorecard is an opinion with a chart.
Who does what
The research is AI-assisted and we are not coy about it. An AI agent did the desk research, proposed every score with its reasoning and sources, and built the engine. We set the structure, argued with the scores, and decide what gets published. A script does the arithmetic, so that neither the AI nor we can quietly round the answer toward the one we prefer.
The engine is a folder of plain text and Python that any capable model can run. It is not welded to whichever AI company is winning this year.
What this is not
It is not a prediction that a robot will help your mother out of bed in 2039. It is a statement about what has to be true first, where each of those things stands, how fast they are moving, and what that implies if nothing changes. The point of publishing it is that things should change, and the working says which things, in which order.
You get the conclusion in plain English. If you do not believe it, you get the working.