Marco Combetto

AI & Digital Transformation — Public Sector — Data Science

Can AI agents conduct open-ended AI research? Early evidence from two case studies

Six days and $3,000 in API credits. That is what was given to frontier AI agents to try and complete original research for two unpublished NeurIPS submissions. The researchers used a “shadow evaluation” method (giving the agents the same research questions as the human authors) to see if the machines could actually do the work. The result is a perfect example of why we need to be very pragmatic about the current hype around autonomous systems. I have to admit that the engineering side of things is now much simpler and the AI finished the coding without help, but it failed to actually conduct the research.

Moreover, it struggled with judgement (the “bar” for what is actually useful). It could not backtrack effectively when hitting a dead end, and it suffered from instruction drift. In my experience with digital transformation in Italian public administration, I have seen many tools that are technically capable but lack the logic to handle the messy, non-linear reality of governance.

The machines can do the heavy lifting of the technical “how,” but they still cannot figure out the “why.” We need to slow down the search for a “set and forget” solution for public services. What is actually required is a model where the human remains the architect of logic and the AI acts as a high-speed executor. The decision on which path serves the citizen best must remain a human responsibility.

#AI #PublicSector #DigitalTransformation #Governance #MachineLearning

🔗 Read the original article