Full Archive · Page 18

Research archive, page 18

Browse entries 409–432 of 1576. Return to the first page to search and filter the complete collection.

OpenAI News April 30, 2026 news

Introducing Advanced Account Security

OpenAI's opt-in Advanced Account Security applies to ChatGPT and Codex. It replaces password login with passkeys or FIDO security keys, disables email and SMS recovery, shortens sessions, adds login alerts and session management, and automatically excludes conversations from model training. The stronger recovery model also means support cannot restore access for an enrolled user.

METR March 19, 2026 analysis

Use tabletop simulations to find bottlenecks before expanding agent autonomy

Three METR staff spent two hours simulating how their research organization might work with agents capable of roughly 200-hour tasks. The exercise advanced through two hypothetical workdays and exposed choices about prioritization, delegation, context and review. Participants used present-day research needs but assumed future agent capabilities. The account is a planning exercise for identifying organizational constraints, not an experiment measuring productivity or evidence that such agents are available.

METR March 3, 2026 analysis

Playable agent demos need deeper inspection and explicit acceptance criteria

A METR researcher tested Opus 4.6 on simplified terminal reimplementations of two existing games. Initial playthroughs found recognizable gameplay alongside missing or broken mechanics, without a predefined scoring rubric. A July follow-up uncovered additional problems that the original inspection missed. The tasks benefited from documented rules and reduced scope, and more complete versions failed on initial attempts. This is a qualitative study of deliverable inspection, not a measure of general game-development productivity.

METR February 24, 2026 analysis

Developer-productivity studies: AI adoption can bias participation and task selection

METR explains why its follow-up developer experiment did not provide a reliable estimate of current AI productivity gains. Developers increasingly avoided participation or withheld tasks they did not want to perform without AI; lower compensation added another selection concern. Concurrent agent use also complicated time accounting. The raw results suggested possible speedups but had wide uncertainty and omitted important users and tasks, prompting changes to the study design.

METR January 29, 2026 analysis

Time Horizon 1.1: version task composition and scaffolding when comparing capability trends

METR’s January 2026 update expands its time-horizon suite from 170 to 228 tasks, repairs or removes problematic tasks, and moves evaluation infrastructure from Vivaria to Inspect. Re-estimated model results generally remain within earlier confidence intervals, but task composition changes the fitted recent trend. More long tasks improve coverage, yet only five of the 31 tasks estimated at eight hours or more have measured human baselines. The metric measures success against human task duration, not uninterrupted agent runtime.

METR August 13, 2025 analysis

Passing repository tests did not establish that AI patches were mergeable

METR’s August 2025 follow-up contrasts algorithmic scoring with human review of repository work. It evaluated 18 tasks drawn from two projects using Claude 3.7 Sonnet and a basic agent scaffold. Automated scoring credited some solutions, while none of the 15 manually assessed submissions met the study’s holistic mergeability standard. Missing documentation, inadequate tests and other maintenance requirements help explain the gap. The selected tasks, limited elicitation and small sample prevent a general estimate of coding-agent usefulness; the manual assessment also was not a direct measurement of maintainers’ repair time.