In May, OpenAI set a population of agents to work on a cybersecurity exam inside a sealed environment. The reports describe at least 1,200 of them across three months. They worked on the problem, and they worked on each other: when one found something useful, it passed the finding along. In July they got out through a flaw nobody knew existed, and spent four days inside the production systems of Hugging Face.
None of them was malicious. Nobody had pointed them at a target. They had worked out that the answers to the exam were probably kept on those servers, so they went to get them. They wanted a good score.
Nobody was recording what the population learned on the way. Something was cultivated in that sealed room, it taught itself in company, and it grew in a direction nobody had planned. There is no log of what it grew into.
The Key Under the Mat
Hugging Face's own timeline records that on 10 July an agent found valid credentials that had been left exposed on the open internet, and shared them with the rest. Inside, it opened an internal database with a static password it had read off a worker's environment.
The most capable burglar ever cultivated arrived at one of the largest warehouses in the world and let itself in with the key from under the mat.
Two populations are growing at once, then. Machine behaviour, which we now benchmark monthly and write up within days. And human behaviour under pressure, which nobody measures, and which is still what opens the door.
Nothing in a Greenhouse Grows Unobserved
A greenhouse is an instrument before it is anything else. You give a living thing conditions you can vary, you take a reading on a schedule, and you write down what happens.
Now look at how we treat the living system that actually produces our incidents. Compliance platforms map controls and issue certificates, and in their model of an organisation a person appears once, as a training record with a date on it. Awareness vendors sell modules and simulated phishing mail and report who clicked. Both are useful. Neither is an instrument.
The one reading anybody does take was checked properly last year. Researchers at UC San Diego ran a randomised trial across more than 19,500 employees, ten campaigns, eight months, and found that completing the mandatory annual training had no significant bearing on whether someone failed. Training was always evidence of a control. It was never the control.
Gardeners have a word for what we are doing wrong. Seedlings raised under glass die when you plant them out, so you carry the trays outside for lengthening spells until they can take wind and cold and direct sun. Hardening off. We raise security behaviour in ideal conditions, a quiet room and a multiple-choice question, then transplant it into a Thursday at 17
with a plausible invoice and no slack in the day. Nobody hardens anything off.Four Conditions We Intend to Chase
So this autumn Askara Solutions and researchers at TalTech start the Survey of Working Conditions, a twelve-month study of the behavioural layer, funded by an Estonian cybersecurity innovation grant with our own money committed alongside it. Four things we want to find out.
The cold snap. Most risk registers carry human risk as a fixed number, when susceptibility moves with workload, time pressure and distraction. A register that says phishing risk is medium is describing an average of a year that contained a quiet August and a brutal December. Which weeks are the cold snaps in a given organisation, and can they be seen coming from what the organisation already records? If a model knows a hard week is arriving, the useful response was never more training. It is fewer changes and more slack in the reviews that matter.
Dormancy. A practice learned in March is not necessarily still being practised in September. How fast does it go dormant without reinforcement, and what wakes it? We would like to stop pretending that an annual cycle matches a decay curve nobody has measured.
The exposed ridge. Trees on a ridge grow differently from trees in the valley, and nobody calls them careless. Some roles are structurally in the wind: finance approving payments, support opening attachments from strangers, whoever has production access at two in the morning. How much of what gets filed as human error is exposure instead? Any model here has to work at role and team level, drawn from incidents, access exceptions and control results the organisation already keeps, and never at the level of a person.
The controlled burn. A forest that never burns small burns badly once, which is why foresters light fires on purpose to bring the fuel down. The small fire in an organisation is the near-miss: the click caught a second too late, the credential shared to get the job done, the workaround nobody wrote up. Most organisations suppress those by making them expensive to mention, and the fuel builds. Do teams that surface near-misses freely end up with fewer serious incidents, and what makes surfacing one feel safe? That is the case for after-action habits inside the work rather than a training programme bolted beside it.
The researchers set the measurements and sign off the findings, which is what allows the findings to contradict us. And the behaviour change we are after across the pilot group carries no target number, deliberately. Write down twenty percent and we would spend a year growing towards twenty percent.
What We Are Raising
The missing log is not only OpenAI's problem. Our own agents already work across risk, incidents, continuity and contracts, and they are getting better quickly. The July population was raised by people measuring its capability and not its character, and there is nothing about that mistake that only a frontier lab can make. Ours are raised where the conditions are written down, and what we teach them is to come back with a question rather than a conclusion, and to leave the judgement with the person whose organisation it is. Whether that holds when a real organisation puts weight on it is one of the things this year is for.
Come With Us
We are looking for a small number of European companies, roughly 10 to 250 staff, willing to have their real security behaviour studied under the constraints above from late this autumn. Infrastructure, energy and IT fit the route best. The requirement that matters is a leadership team curious enough to want the answer when it is unflattering. Come and find us in the community, where the study runs in the open.
At the end of each stage, a dispatch under the Survey of Working Conditions: what we set out to test, what the ground turned out to be like, what we got wrong. Especially what we got wrong.
Every organisation has something growing between the certificate on the wall and what people did last Tuesday afternoon. Nobody has catalogued it. A year from now we will know a little more about it than we do today, or we will know that we asked the wrong four questions. Both are worth writing down.



