/

Amazon Science

25 stories

Amazon Science

Amazon ScienceResearchDeveloping provably correct Rust code with Verus How the Verus "program verifier", which automatically checks code against a mathematical specification of its functionality, helps increase security assurance in software projects.

August 31
Amazon Science

Amazon ScienceResearchWhen LLM judges agree, should we believe them? Discounting the opinions of LLM judges with highly correlated outputs ensures that panels of judges reflect a true diversity of perspectives.

August 26
Amazon Science

Amazon ScienceResearchSOP-Bench: A new benchmark for evaluating AI agents on real business procedures Extendable framework enables testing agents on the full set of capabilities required to successfully complete a procedure, not isolated proxy tasks.

August 21
Amazon Science

Amazon ScienceResearchA decade of mathematical certainty: Reflections on the Automated Reasoning Group Ten years after we founded the Automated Reasoning Group, mathematical logic has moved from academic research into production services that secure millions of customer workloads — demonstrating that systems can be provab

August 11
Amazon Science

Amazon ScienceResearchAWS Trainium Frontier competition: Co-design models and kernels on purpose-built AI chips A competition with a finalist ceremony during NeurIPS 2026, challenging researchers to train language models from scratch on Trainium, exploring what optimal architectures look like when the hardware changes.

August 10
34 Amazon Research Awards Build on Trainium recipients announced

Amazon ScienceResearch34 Amazon Research Awards Build on Trainium recipients announced Amazon announces 34 recipients of the Build on Trainium program, a $110 million credit initiative supporting AI research at 30 universities including Stanford, UC Berkeley, UIUC, UCLA, CMU, and MIT, with a focus on Respo

August 5
Amazon Science

Amazon ScienceResearchHow controllers from industrial machinery can coordinate multitask machine learning Instead of compromising among parameter updates dictated by different training objectives, ControlG allocates computational capacity to objectives sequentially and dynamically.

July 30
Amazon Science

Amazon ScienceResearchA new benchmark for evaluating patient-facing health AI agents PatientAgentBench generates a synthetic patient health record, a realistic clinical vignette, and a patient agent that converses with the AI system under evaluation, to capture what a patient-facing agent actually has to

July 29
Amazon Science

Amazon ScienceResearchAmazon is investing in the Lean Focused Research Organization As AI agents take on higher-stakes decisions, Lean programming language makes it possible to mathematically prove they will behave safely.

July 26
Amazon Science

Amazon ScienceResearchAmazon and University of Michigan give robots a sense of touch HydroShear, a new physics-based simulator, teaches robots how to use their sense of touch to perform complex manipulation tasks, in a way that transfers seamlessly to the real world.

July 10
Amazon Science

Amazon ScienceResearchCapturing token IDs during agentic interactions for better reinforcement learning A new Rust proxy called Turnstile sits between the model backend and the agent harness to capture information lost in mere text transcripts.

July 9
How Amazon tracks carbon intensity across its operations

Amazon ScienceResearchHow Amazon tracks carbon intensity across its operations Amazon is developing precise, sector-specific approaches to measuring decarbonization progress — starting with emissions per unit shipped.

July 1
Amazon Science

Amazon ScienceResearchThe fuel of the future is already here: Why TRISO matters Millimeter-scale particles of nuclear-reactor fuel are encased in four layers of different materials that act as a “miniature containment system”.

June 24
Amazon Science

Amazon ScienceResearchEC2’s formally verified “isolation engine” provides mathematical assurance of virtual-machine isolation Splitting the “separation kernel” off from the rest of the Nitro security system and using only a subset of the Rust programming language to code it enabled its formal verification.

June 10
Graviton5’s improved design increases speed and energy efficiency — beyond Moore’s law

Amazon ScienceResearchGraviton5’s improved design increases speed and energy efficiency — beyond Moore’s law A new chiplet architecture, custom die-to-die connectivity, and support for DDR5-8800 memory and the latest PCIe gen6 interconnects improve performance by 25% for general-purpose and agentic AI workloads.

June 10
Amazon Science

Amazon ScienceResearchReal-world grounding in agentic AI Four approaches can dramatically improve the performance and trustworthiness of AI agents in operational environments.

June 8
Amazon Science

Amazon ScienceResearchBridging intent and execution in agentic systems The harnesses that mediate between models and tools in agentic systems are becoming their own performance bottleneck, but a few simple design principles can fix what ails them.

June 8
Amazon Science

Amazon ScienceResearchGround truth is a process, not a dataset Automatically fact-checking long, AI-generated research reports poses new challenges — including benchmarking.

June 3
Amazon Science

Amazon ScienceResearchHow flat is replacing fat in AWS data center networks “Quasi-random” network topologies and new passive optical components called ShuffleBoxes make more-efficient flat networks as practical as traditional “fat-tree” networks.

May 28
Amazon Research Awards recipients announced

Amazon ScienceResearchAmazon Research Awards recipients announced Awardees represent more than 49 universities in 11 countries. Recipients have access to Amazon public datasets, along with AWS AI/ML services and tools.

May 27
Amazon Science

Amazon ScienceResearchDiverse reasoning traces teach LLMs to make better decisions How to train language models to generate diverse, accurate reasoning paths using tokens that control distinct reasoning strategies.

May 26
Amazon Science

Amazon ScienceResearchMaking LLMs faster without sacrificing accuracy A new scaling law that relates particular architectural choices to loss helps identify models that improve throughput by up to 47% with no loss of accuracy.

May 15
Amazon Science

Amazon ScienceResearchPromptimus: Improving already good LLM prompts with zero manual engineering By focusing on specific failure points and suggesting targeted solutions, a new automated prompt-engineering framework improves prompt performance without compromising existing functionality.

May 14
Amazon Science

Amazon ScienceResearchNavigating uncertainty in Amazon's middle-mile network Amazon engineers and scientists have created new tools to optimize delivery networks under uncertainty — and keep them adapting without missing a beat.

May 6