跳到正文
10月1日周四
  1. NVIDIA Blog

    Productive, Durable, Fungible: How NVIDIA AI Factories Maximize Return on Investment

    AI factories are built by the megawatt, even by the gigawatt. Each megawatt factory costs roughly $60 million, and AI factory operators will only commit capital on that scale with a clear view of the return on investment. Three key things shape AI factory returns: Earning capacity: What the factory could earn in a year […]

  2. Apple Machine Learning Research

    RLTL;DR: Self-Improvement by Internalizing Self-Generated Feedback

    The common paradigm of reinforcement learning with verifiable rewards (RLVR) is to let agents make multiple attempts at a task, and optimize towards the successful ones. This becomes problematic in the realms of self-improvement, where tasks are so difficult that the agent has a low or even no chance of success, and where there are no teacher models or example solutions to distill from. In this paper, we introduce RLTL;DR. After each failed attempt, we show the policy the verifier outputs and let it write its own feedback, in the form of a single TL;DR insight. The next rollout is conditioned…

  3. Apple Machine Learning Research

    How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?

    Recent autonomous machine learning engineering (MLE) agents have made significant progress on public leaderboards. Often motivated by progress stagnation over long-horizon cycles and limited Large Language Model (LLM) primitives, modern MLE agents are deployed on top of increasingly elaborate machinery: multi-agent orchestrators, dedicated retrieval subagents, and more. While such harnesses expand, the use of more primitive but improved coding agents—where LLMs have direct access to the execution environment through read, write, and bash primitives—has received little attention in the field…

  4. Claude Code:GitHub Releases

    v2.1.286

  5. NVIDIA Blog

    NVIDIA Opens Applications for 2027–2028 Graduate Fellowships With Awards Up to $60,000

    Bringing together the world’s brightest minds and the latest accelerated computing technology leads to powerful breakthroughs that help tackle some of the biggest research problems. To foster such innovation, the NVIDIA Graduate Fellowship Program provides grants, mentors and technical support to doctoral students doing outstanding research relevant to NVIDIA technologies. The program, in its 26th […]

  6. Pragmatic Engineer

    Distributed databases with Peter Mattis

    Cockroach Labs co-founder Peter Mattis discusses building reliable distributed systems and how AI helps him write more code without sacrificing quality.

  7. Microsoft Research

    Forecasting space weather risks on power grids

    Extreme space-weather events can damage power systems on Earth and degrade GPS accuracy and satellite operations. A new machine learning system can predict where damage is likely to occur 30-60 minutes before a storm arrives. The post Forecasting space weather risks on power grids appeared first on Microsoft Research .

9月30日周三
  1. Google DeepMind

    Introducing SynthID Bio

    Proof of concept for watermarking AI-generated proteins while preserving biological function.

  2. NVIDIA Blog

    From Training to Production, NVIDIA and CoreWeave Close the Loop on Agentic AI

    Building on nearly a decade of co-engineering, CoreWeave has built NVIDIA compute, networking and software into a cloud purpose-built for AI that’s still returning on investment across multiple generations of deployment. Now, CoreWeave is bringing the next generation of NVIDIA infrastructure to production. At CoreWeave Fully Connected, running this week in San Francisco, CoreWeave announced […]

  3. Apple Machine Learning Research

    SCLATE: A Substrate for Continual-Learning Agent Training and Evaluation

    Continual-learning agents are systems of models, harnesses, and memory operating over long multi-session horizons. Evaluating and training them requires interleaving tasks with agent-side events such as session stop and start, crons, and memory consolidation. Yet existing benchmarks and training frameworks schedule only the benchmark’s own events, leaving each benchmark and agent pair to build a custom scheduling loop. We present SCLATE, an execution substrate where benchmarks and unmodified agents each add their events to one open event scheduler through an adapter. A hybrid simulated clock…

  4. Apple Machine Learning Research

    On the Effectiveness-Fluency Trade-Off in LLM Conditioning: A Systematic Study

    Controlling the output of Large Language Models (LLMs) is a central challenge for their reliable deployment, yet a clear understanding of the involved trade-offs remains elusive. Current approaches to conditioning are often evaluated with a narrow focus on their effectiveness at injecting or removing a target concept, neglecting generation quality. We systematically investigate a range of conditioning methods in both injection and removal scenarios. We find that efficient steering methods frequently achieve conditioning at a steep cost to fluency. Furthermore, we identify a critical yet…

  5. Claude Code:GitHub Releases

    v2.1.285

9月29日周二
  1. Pragmatic Engineer

    Why has Shopify dropped React Native?

    It’s only been a year since the e-commerce platform declared it was very happy with React Native, but now Shopify is dumping it – and the reason, unsurprisingly, is AI

  2. Microsoft Research

    Introducing Quine: An AI research system designed for the complexity of biology

    Biology doesn't operate in silos, and neither should the AI representation of it. Quine is an early-stage research effort to create a multimodal world model of biology. By connecting insights across biological scales and modalities, Quine helps scientists computationally search a space far larger than intuition allows and prioritize hypotheses before they reach the lab. Experimental results provide important feedback, helping researchers sharpen future research directions. The post Introducing Quine: An AI research system designed for the complexity of biology appeared first on Microsoft Research .

  3. MIT Technology Review · AI

    Making AI an asset, not an expense

    When customers talk about AI costs, the conversation usually starts with token prices and ends with access to the latest, most capable model in the cloud. Do they always need that level of capability? Not necessarily. But that is often where the conversation goes. As AI moves from experimentation to production, model choice is only…

  4. Apple Machine Learning Research

    The Communication Bottleneck: A Round-Trip Study of Tree-Structured Expression Serialization in Language Models

    When language models reason in chain-of-thought or exchange free-text intermediates, they serialize structured information into natural language. How much tree-structured compositional content survives this bottleneck? We propose a round-trip protocol that answers this question empirically for tree-structured expressions. A generator converts a procedurally generated arithmetic expression into a word problem, a separate extractor recovers the expression from the word problem alone, and symbolic equivalence provides an exact oracle. Evaluating all pairwise combinations of sixteen models yields…

  5. Microsoft Research

    One year in: How Microsoft Research Asia – Singapore is advancing research, partnership and talent for real-world impact

    Since launching a year ago, the Microsoft Research Asia — Singapore lab has established a strong foundation, deepened collaboration across government, academia, and industry, and explored how frontier AI research can create real-world value. The post One year in: How Microsoft Research Asia – Singapore is advancing research, partnership and talent for real-world impact appeared first on Microsoft Research .

9月28日周一
9月27日周日
9月26日周六
  1. Azeem Azhar:Exponential View

    🔮 Safety in numbness

    On originality, intellectual bravery and why disagreeing with LLMs is a good sign

  2. GitHub Blog · AI & ML

    GitHub Copilot app for Beginners: How to build custom workflows with canvases

    Describe the interface you need in plain English, then let the agent build a live surface you can both use and update—so you spend less time adapting to tools and more time getting work done. The post GitHub Copilot app for Beginners: How to build custom workflows with canvases appeared first on The GitHub Blog .

9月25日周五
  1. Azeem Azhar:Exponential View

    🤖 Your agent, whose interests?

    Meta’s Muse shows that users want a digital butler, but there may be a conflict of interests

  2. GitHub Blog · AI & ML

    When chat is the wrong UI

    What is a developer to do when they need something more tangible than a chat box? Enter canvases. The post When chat is the wrong UI appeared first on The GitHub Blog .

  3. GitHub Blog · AI & ML

    AI-powered fuzzing with the GitHub Security Lab Taskflow Agent

    In this blog post, I explain how to use the new fuzzing taskflow based on the GitHub Security Lab Taskflow Agent AI framework. The post AI-powered fuzzing with the GitHub Security Lab Taskflow Agent appeared first on The GitHub Blog .

9月24日周四
  1. GitHub Blog · AI & ML

    Rendering huge pull requests in the GitHub Copilot app

    How we rebuilt the diff surface in the GitHub Copilot app to open a million-line pull request with hundreds of inline review comments. The post Rendering huge pull requests in the GitHub Copilot app appeared first on The GitHub Blog .