NewslettersThe Tech Siblings Podcast — powered by Tech Spindle
The Tech Siblings Podcast

Tech Spindle Weekly Newsletter 08.22.2026: NVIDIA hits 100% on ARC-AGI-3, and what it means

Harika (the Pragmatist) vs. Rohan (the Tinkerer) — this week's stories, debated.

NVIDIA AVO Reaches 100% on ARC-AGI-3, Demonstrating a Frontier-Level General-Purpose Architecture for Long-Horizon Autonomous Agents
NVIDIA Developer Blog

Rohan: Okay, so NVIDIA just hit 100% on ARC-AGI-3, which is wild because this benchmark was literally designed to test reasoning that LLMs can't just memorize their way through. The clever bit is their agent architecture — basically letting the model think in steps, backtrack, and self-correct over long tasks instead of just spitting out one answer.

Harika: Right, and the practical piece here is that this isn't just about benchmarks — it's about whether AI can actually handle complicated, multi-step workflows without human babysitting. Think coding across multiple files, research tasks, planning — stuff that takes hours, not seconds.

Rohan: Exactly! And NVIDIA's showing that the architecture matters as much as the model itself. You can't just throw a bigger model at these problems — you need the scaffolding that lets it actually think like an agent.

Harika: Which means we're probably going to see a wave of companies retrofitting their LLMs with agent frameworks. The race just shifted from 'who has the biggest model' to 'who can build the best autonomous systems.'

“Agent architecture just became the new frontier — 100% on ARC-AGI-3 proves that how you let a model think matters more than raw scale.”
Michael Polansky is training an AI model on skin that’s still alive
TechCrunch

Rohan: Okay, this is genuinely wild — they're training AI on living skin tissue that they're keeping alive outside the body for weeks. Like, not just cells in a dish, actual functioning skin!

Harika: Right, and the practical play here is skincare discovery — basically finding new compounds that actually work on real human skin instead of relying on animal testing or theoretical models. That's a massive upgrade for the cosmetics industry.

Rohan: The technical challenge of keeping that tissue viable long enough to get useful data is insane though. Most labs can barely keep it alive for days, so if they've cracked weeks, that's legitimately impressive.

Harika: Sure, but it's still early-stage startup territory — cool science doesn't always mean viable business. I want to see if they can actually scale this and partner with big skincare companies.

Rohan: Fair, but come on, you have to admit training AI on literal living human skin is the coolest biotech application we've covered in a while!

Harika: Oh totally! I'm just keeping us grounded while you geek out over it. *laughs*

“Genuinely novel science with real commercial potential, but needs to prove it can scale beyond the lab.”
From Atari to EVE Online: Building on 15 Years of AI Research in Games
Google DeepMind

Rohan: Okay, DeepMind going from Atari to EVE Online is actually wild — EVE is this massive multiplayer economy with like, spreadsheet-level complexity. We're talking agents that have to navigate politics, trade, alliances, not just 'shoot the thing.'

Harika: Right, and that's why this matters beyond just gaming. If AI can handle EVE's chaos, you're looking at real applications in supply chains, market simulation, any system where you have tons of agents making decisions at once.

Rohan: Exactly! It's the messiness that makes it interesting — Atari was clean, predictable physics. EVE is humans being weird and unpredictable, which is way harder to train for.

Harika: Though I will say, partnering with actual game studios is the smart move here. They get real environments, studios get better NPCs, and we get research that might actually ship in products people use.

“DeepMind's leveling up from arcade games to actual complex worlds, and the messiness is the point.”
AWS Releases Aws-Bench to Evaluate Agents on Cloud Tasks
InfoQ

Rohan: Okay, this is actually pretty clever — AWS is letting AI agents loose on real AWS infrastructure in throwaway accounts to see if they can actually complete tasks. Like, can your AI agent spin up an EC2 instance or configure S3 buckets without breaking everything?

Harika: Right, and the practical angle here is that if you're building agents that interact with cloud infrastructure, you finally have a standardized way to test them. No more just hoping your agent doesn't delete production databases when it goes live.

Rohan: Though honestly, it's a bit narrow — it's AWS-specific, and we've seen a million benchmarks this year. Is this really moving the needle or just AWS marketing their own services?

Harika: Little bit of column A, little bit of column B! But hey, if it means fewer catastrophic cloud misconfigurations from AI agents, I'll take the incremental progress.

“Useful tooling for a real problem, even if it's not exactly groundbreaking.”
AI Model Routing: The Missing Infrastructure Layer for Multi-Model AI Applications
Dev.to

Harika: Okay, so this is basically a smart traffic cop for AI models — you've got GPT-4, Claude, Llama, whatever, and something needs to decide which one gets each request based on cost and how reliable you need it to be. If you're running a production app hitting multiple models, you were probably building this yourself until now.

Rohan: Right, and the clever bit is it's not just round-robin or random — it's actually looking at your requirements in real time. Like, 'this query is simple, send it to the cheap fast model,' versus 'this one's complex, worth paying for the expensive one.'

Harika: Exactly, and that matters because companies are burning so much money on overkill — sending every single request to the most expensive model when half of them could've been handled by something way cheaper.

Rohan: It's like… okay, imagine if every time you wanted to move something you rented a semi-truck instead of just using your car for the small stuff. That's what's happening without routing!

Harika: Ha! Yeah, and now someone's finally building the layer that picks the right vehicle. About time this became infrastructure instead of everyone reinventing it.

“Model routing is the unsexy plumbing that'll save companies a fortune and make multi-model AI actually practical.”
Nvidia AVO scores 100% on the ARC-AGI-3 interactive reasoning benchmark
Hacker News

Rohan: Okay, so Nvidia just dropped AVO, and it's the first model to hit 100% on ARC-AGI-3 — which is basically the benchmark that's supposed to test if AI can actually reason like humans do, not just memorize patterns. This is kind of a huge deal!

Harika: Right, but let's be real — this is an *interactive* version where the model gets to try multiple times and learn from mistakes. It's impressive, sure, but it's not like it nailed every puzzle on the first shot.

Rohan: Fair, but even with the back-and-forth, the fact that it can adapt and figure out completely novel visual reasoning puzzles? That's way closer to actual general intelligence than just throwing more training data at a problem.

Harika: I mean, yes — and if this translates to real-world tasks where trial-and-error is allowed, like robotics or scientific research, then Nvidia just moved the goalposts. But I'm waiting to see if this works outside the benchmark lab.

Rohan: Ha, the classic 'demo versus deployment' gap. Though coming from Nvidia, they've got the hardware and the research chops to actually push this into products.

Harika: Exactly — and honestly? If AVO becomes the brain behind their next-gen robotics or simulation tools, this isn't just a benchmark flex, it's a market play.

“Nvidia just proved AI can learn to reason interactively — now we wait to see if it escapes the lab and actually ships.”

Harika: And that's a wrap on this week's top curated news from Tech Spindle!

Rohan: If you want these stories delivered straight to your inbox, subscribe to our Daily and Weekly Newsletters—

Harika: And hey, if you've got your own take on tech news, you can actually publish your own blogs on Tech Spindle too. Just head to techspindle.ai and Register.

Rohan: Thanks so much for listening to The Tech Siblings Podcast, brought to you by TechSpindle.ai. Catch you next week!

Your vote shapes what gets covered next

Save and vote on stories in the Feed — that's the same signal Harika and Rohan pull from when picking next week's list.

Request an invite →
Read this week's newsletter

Install Tech Spindle

Add it to your home screen for a faster, app-like experience.