Deep Dives
Google’s AI takes on the NHS breast screening bottleneck — two new studies, real results
Google Research just dropped two companion studies in Nature Cancer on using AI in NHS...
Testing LLMs on Superconductivity Research Questions
Google researchers tested six LLMs on expert-level high-temperature superconductivity questions. NotebookLM and a custom system...
ConvApparel: Why Your AI User Simulator Is Probably Lying to You
Google's ConvApparel dataset exposes how LLM-based user simulators fail to mimic real humans—they're too patient,...
Google’s New Framework Puts LLM Personality Tests on the Couch
Google Research introduces a framework that adapts psychological questionnaires into situational judgment tests to measure...
How many raters do you actually need for AI benchmarks? Google has answers
Google Research challenges the standard 1-5 rater approach in AI benchmarks, showing that depth over...
ReasoningBank: Giving AI Agents a Memory That Actually Learns from Failure
Google's ReasoningBank framework lets agents distill generalizable reasoning strategies from both successes and failures, moving...
Simula: A Smarter Way to Generate Synthetic Data by Designing Datasets, Not Just Samples
Google Research's Simula framework treats synthetic data generation as mechanism design, using reasoning to build...
MoGen: How Google Is Using Synthetic Neurons to Speed Up Brain Mapping
Google Research's new MoGen model generates synthetic neurons to train AI, cutting reconstruction errors by...