← Back to archive

Friday, July 17, 2026

OpenAI x Jony Ive building a smart speaker

OpenAI's GPT-Red is making GPT-5.6 6x more robust through automated security testing, while Bloomberg reports they're building a ChatGPT smart speaker with Jony Ive for 2027 (bold move after Humane's struggles). Meanwhile, ReactBench drops some sobering data: top AI coding agents are failing 57% of realistic React tasks—yikes. Ready to trust AI agents with your production code?

Top Stories

1
GPT-Red for Safety Testing

OpenAI's GPT-Red is an automated red-teaming model that uses self-play reinforcement learning to find AI vulnerabilities at scale, successfully making GPT-5.6 6x more robust to prompt injection attacks. This creates a safety flywheel where today's models help secure tomorrow's systems, addressing the scalability limitations of human red-teaming alone.

openaisafetyred-teamingprompt-injection
2
Supply Co. x Work Louder

Supply Co. and Work Louder released a $230 specialized keyboard with RGB feedback and physical controls designed specifically for managing AI agent workflows and ChatGPT Codex interactions. The hardware represents emerging demand for purpose-built peripherals as agentic AI becomes embedded in developer workflows.

agentshardwaredeveloper-toolschatgpt
3
Simon Willison on Kimi K3

Simon Willison's Blog

Moonshot AI's Kimi K3 becomes the largest open-weight model at 2.8T parameters, achieving top-tier benchmark performance but at significantly higher pricing ($3/$15 per million tokens) than previous Chinese models. The model shows strong reasoning capabilities but consumes extensive reasoning tokens even for simple tasks.

open-sourcellmmoonshotbenchmarks
4
ChatGPT grew ears

Bloomberg

OpenAI plans to release a portable, screenless ChatGPT smart speaker in 2027 as its first major hardware device, featuring camera-based environmental awareness and mechanical elements for human-like interaction. The device is being developed with Jony Ive's design company and represents OpenAI's broader push into consumer hardware.

openaismart-speakerhardwarechatgpt
5
ReactBench v1

ReactBench

ReactBench v1 evaluates coding agents on realistic React work using both behavioral tests and 400+ deterministic quality rules, revealing that even leading models like GPT-5.6 Sol achieve only 43% success rates while introducing hundreds of production-critical bugs. This benchmark addresses the growing risk of AI-generated React code causing real-world outages and performance issues as models write more frontend code at scale.

benchmarkscoding-agentsreactllm

Keep Reading

Enjoyed this issue?

Get daily AI intel delivered to your inbox. No fluff, just the stories that matter.