全部 AI 动态
Codex:GitHub Releases Apple Machine Learning Research Negotiating Ontological Boundaries in User-Authored Personal Sensing Systems
Designed artifacts are ontological, shaping, and at times limiting, what becomes possible or imaginable. One path toward mitigating such foreclosures is giving people power over how systems are designed and built. Despite decades of scholarship around systems that enable such authorship, these systems are often evaluated on whether or not they are usable, useful, or technically feasible, leaving questions of ontological boundary negotiation, unexamined. We design two open-ended probes that utilize a Wizard of Oz technique to enable the experience of training a personalized machine learning…
Simon Willison Qwen3.8 27B addition in words
Research: Qwen3.8 27B addition in words Colin Frasier posted on Bluesky about an experiment he ran over two years ago using GPT-4o to see how well it could "compute the sum but return the answer in words" across increasingly large numbers. Here's the chart he shared of those results: I'm confident GPT-4o didn't cheat and use a calculator, especially since it got so many of the calculations wrong, but I was inspired to run the experiment again on local hardware (a DGX Spark) to explore the effect in a fully controlled environment. I pasted his image into a Codex Remote session (GPT-6 Astra) and had it run the same experiment using Qwen3.8-27B-Q4_K_M.gguf . Here's the result for a run of 30 attempts per combination with reasoning disabled: Then I ran it again with reasoning enabled. This took a lot longer per pair, so instead of running 30 samples per square I ran just one - which results in a much less visually appealing heatmap since each square is either 100% or 0%: It got the right answer in 167 out of 169 attempts, and since these were one-shot I'm confident a second run would produce different results here. Here's a version of the report that includes the reasoning traces from some of those larger calculations, which include text like this: Wait, let me redo this more carefully. 4,299,366,105,622 6,088,794,067,970 Let me align them: 4 2 9 9 3 6 6 1 0 5 6 2 2 6 0 8 8 7 9 4 0 6 7 9 7 0 Adding from right to left: Position 1 (units): 2 + 0 = 2 Position 2 (tens): 2 + 7 = 9 Position 3 (hundreds): 6 + 9 = 15, write 5, carry 1 Tags: mathematics , ai , generative-ai , local-llms , llms , qwen , llm-reasoning , dgx-spark
Codex:GitHub Releases 0.162.0-alpha.13
Codex:GitHub Releases 0.162.0-alpha.12
Azeem Azhar:Exponential View 🔮 The transition is hiding in plain sight #604
AI consciousness, electrification’s hidden progress, and the coming agentic bank run++
Simon Willison We're going to need default hard budget caps on pretty much everything
Here's a product feature which the world is going to need a whole lot more of over the coming months and years: default hard budget caps . I'm talking about the feature of pay-by-usage services and APIs that lets you say "after $X/month, cut this thing off and return errors". These need to be hard limits. Soft caps, "after $X/month, send me a warning email", will not cut it. Coding agents, and personal agents (coding agents wrapped in a less threatening UI), greatly reduce the friction of spinning up code that can do useful things. Sometimes those things cost money - calls to paid APIs, or hosted web applications, or systems that can bill for additional storage and compute. Nobody wants to wake up to an email sent at midnight warning about a budget limit and find that, while they slept, their rogue service had consumed several hundred (or several thousand) more dollars of usage. An argument against this is that businesses don't want their hosted applications to start throwing errors because some budget was exceeded. I expect that most businesses and individuals would prefer errors to a surprise $10,000+ bill. I think hard budget caps need to be the default. If someone wants to live dangerously they should be able to do that, but it needs to be on an opt-in basis. Have a nice, clear checkbox somewhere prominent: Remove the budget cap. My application will not be shut down if I exceed the configured budget limit, and I will be responsible for subsequent charges. The service I most want to see this from is AWS. I've heard plenty of stories from people who refuse to use AWS for personal projects out of (justified) fear that a runaway service might bankrupt them. I've also heard stories from people who didn't anticipate this and ended up seriously burned. ... and it turns out AWS finally launched spending limits a few weeks ago! From their announcement New AWS experience helps builders get started and ship faster on 16th September: When you're ready to upgrade to a paid p
Claude Code:GitHub Releases v2.1.289
Azeem Azhar:Exponential View 🔮 Quick weekend reads: the big AI questions
A short set of readings on some of the biggest questions.
404 Media Our Solar System Is Terminally Unstable and Will Be Completely Destroyed, Study Finds
The estimated lifespan of the outer solar system has been downgraded from 100 billion years to just a few billion years, according to a study that probed “terminal instability” during the Sun’s death.
Latent Space [AINews] not much happened today
a quiet day.
Ars Technica · AI Apple changes full-disk access permissions to curb abuse from AI agents
Meta says FDA isn't sufficient to Muse reading messages. Apple begs to differ.
404 Media Federal Judge Rules a Flock Search Was ‘Indiscriminate Mass Surveillance’ and Unconstitutional
Flock’s nationwide network is quickly “approaching dragnet-type law enforcement practice” and the cop should have got a warrant, the judge wrote.
Ars Technica · AI Amazon’s $1B plan to combat data center backlash draws more backlash
Amazon praised for ending NDAs but slammed for downplaying data center pollution.
Claude Code:GitHub Releases v2.1.288
MIT News Computational tools for society’s most complex challenges
Associate Professor Cathy Wu uses reinforcement learning to help map out improvements to transportation and other multifaceted systems.
Ars Technica · AI US arrests tech CEO accused of smuggling $300M in Nvidia chips into China
Nvidia’s chip-smuggling problem won’t go away as arrests continue.
MIT News Documenting the tech worker movement
Writing as a participant and researcher, PhD student JS Tan SM ’22 has co-authored a new book about the rise of tech worker protests and the employer backlash that followed.
Newcomer 新闻长文 Anthropic's IPO Can't Wait Forever
Plus, Vinod Khosla fires off at Factory founder that Khosla Ventures backed
OpenAI News A model guide for the GPT-6 family
Learn how startups can choose GPT-6 models, tune reasoning effort, improve prompts and skills, coordinate tools, and prepare workflows for production.
MIT Technology Review · AI Redefining enterprise intelligence with autonomous AI
Enterprise AI is no longer a future ambition. It is in full operational flight. Model capabilities are advancing faster than most organizations can absorb, while the cost of performance continues to fall. Globally, AI investment is set to reach $2.5 trillion in 2026, up 44% from the previous year. For many enterprises, this investment has…
GitHub Blog · AI & ML AI is changing developer work. Here are three skills to strengthen.
Learn to direct AI agents, critically review their output, and keep technical judgment at the center of your workflow. The post AI is changing developer work. Here are three skills to strengthen. appeared first on The GitHub Blog .
Latent Space Inside-Out AI: Rebuilding Airbnb Behind the Scenes and Across the Guest Experience
After leading Meta’s Llama models, Ahmad Al-Dahle is now transforming Airbnb with AI — from how its teams develop products to how it serves guests.
NVIDIA Blog NVIDIA DGX Spark 64GB Gives Developers More Ways to Build and Scale Local AI
Local AI is becoming more useful by the token. As AI agents move from experiments into everyday development, increasingly capable open models are shrinking to fit on more devices, giving builders more to run locally. Coming this month, NVIDIA DGX Spark will be available with 64GB of unified memory from top manufacturer partners — Acer, […]
MIT Technology Review · AI Don’t be fooled—LLMs don’t reason
On an afternoon in Seoul in March 2016, I watched a program I helped build put a stone on the fifth line of a Go board in what looked like a gift to its human opponent. Move 37 in game two of the five-game match looked so absurd that some commentators thought it was a…
Latent Space [AINews] Pi 1.0, Pi Durable, and AIE NYC
the minimalist harness goes stable... and TypeScript!
MIT News 3 Questions: A new resource to empower young entrepreneurs
Martin Trust Center Managing Director Bill Aulet introduces Dear Dreamer, a free platform for middle and high school students who want to learn about entrepreneurship.
Latent Space Academia is for Ambition — Alex Zhang, MIT
We catch up with RLM first author Alex Zhang, MIT PhD, on Jev, PhD masxing, and the future of harnesses.
Apple Machine Learning Research Limits of Confidence in Diffusion
Discrete diffusion, including remasking and uniform-state samplers, generate a sequence by writing multiple token positions per step, drawing each from a per-position distribution and choosing which positions to write from those same distributions. For domains of general interest (pixels, phonemes, or words) there are inherent dependencies between tokens. We show that a step matches the training distribution only when the positions it writes are conditionally independent given the tokens already fixed, that no product of per-position distributions can match a dependent group, and that…
Apple Machine Learning Research Language Discrimination Improves Linguistic Learning in Multilingual Speech Models
Multilingual self-supervised speech models can benefit from sharing information across languages, but under a matched total pretraining data budget they still fall short of monolingual models. We show that strengthening the model’s ability to discriminate languages during pretraining reduces and, on some measures, closes this multilingual gap on continuous phonetic and higher-level linguistic measures, while preserving substantial cross-language sharing. Using a controlled English/French HuBERT setting, we test two interventions which strengthen language discrimination: an auxiliary language…
OpenAI News Chatham scales its capital markets expertise with OpenAI
Chatham Financial uses Codex and GPT-5.6 to build technology and redesign workflows, cutting trade validation from 30 minutes to under 4.
NVIDIA Blog How NVIDIA GPUs Help Accelerate OpenAI’s GPT-6 Astra Ultrafast
GPT-6 Astra Ultrafast, running on NVIDIA Blackwell GPUs, is available now in the OpenAI API and to eligible ChatGPT Work and Codex users. Accelerated by inference optimizations through OpenAI’s models that tap into the capabilities of the NVIDIA Blackwell architecture, Ultrafast offers up to 8x faster token generation than the Astra Standard mode. For developers, […]
MIT News New tool lets users repair AI-generated 3D models, then fabricate them just the way they want
“InstructMesh” can generate designs for everyday objects that are easy to edit and fabricate for both experts and newcomers to 3D modeling.
Newcomer 新闻长文 Igor Babuschkin + Founders of Legora, Factory, Fireworks AI, Etched & Crosby Join the Cerebral Valley AI Summit as Speakers
New speakers will join the CEOs of Anthropic, Databricks & Palo Alto Networks
Claude Code:GitHub Releases v2.1.287
OpenAI News The eternal complement
Advanced AI may matter most for the routine work behind breakthrough ideas. Explore why execution could shape the next economy and the pace of progress.
Pragmatic Engineer The Pulse: Firebase’s global outage & poor response
Also: OpenAI’s platform play that has similarities to AWS, more data on companies moving to open models, and more
Newcomer 新闻长文 SCOOP: Benchmark Backs Early-Stage Startup Tendrils Compute as Chip Momentum Builds
Plus, a market map on the chip alternatives ecosystem
OpenAI News How Albertsons Companies is reimagining retail from the inside out
Albertsons Cos. is using ChatGPT Enterprise and the OpenAI API to help teams work faster and make grocery shopping easier for millions of customers.
NVIDIA Blog Fall Into 25 New Games on GeForce NOW This October
Spooky season is streaming in. Alongside falling leaves, pumpkin spice and everything nice, 25 new games are joining GeForce NOW throughout October, including six ready to play this week. From a new CONTROL Resonant reward for Performance and Ultimate members to The Witcher 3: Wild Hunt – Remastered joining the cloud, this GFN Thursday is […]