AI Signal helps business leaders decide what to fund, test, or avoid. Each five-minute issue features five articles I choose, why they matter, and what I’d actually do about each one. I’m Yashar, founder of PivotPath, an AI product studio in Toronto.

Every item carries a tag — Fund, Pilot, Watch, or Skip. That's my call on what you should do about it in the next 90 days.

―――――――――――

Your product manager can now ship code. Decide who approves it. - [Pilot]
TNW — Slack launches Slack Code, where teams and AI agents build together

Slack launched Slack Code, dedicated channels where teams and AI coding agents can plan, write, review, and ship software together. Claude, ChatGPT, Devin, and Copilot are among the launch partners. Agents inherit existing Slack permissions, can be paused by anyone in the channel, and require human approval before code reaches production. Slack plans to extend the same channel-based model beyond engineering, allowing custom agents to support work such as marketing campaigns and legal document review.

The Signal: The guardrails are real: agents inherit Slack’s permissions, anyone in the channel can stop them, and a human must approve code before production. Good. But the risk is cultural. Putting diffs and previews in chat makes approval faster; it also makes it social. A thumbs-up in a busy channel is not a code review. The broader shift matters more than this product. Slack plans to bring the same model to marketing, legal work and consequential work—turning approval into a message rather than a process. Before agents enter your workspace, write down who can approve their work and what they must review first. Name the roles and minimum checks. Set the norm before convenience sets it for you.

―――――――――――

Nobody is standing at the door of your software supply chain - [Fund]
Gal Ratner — The End Of Open Source

Ratner argues that open-source software’s trust model is becoming harder to defend. Modern applications rely on packages maintained by unpaid individuals, while attackers can use autonomous agents to manipulate maintainers or poison dependencies. Once malicious code reaches a developer workstation or build pipeline, it can steal credentials, source code, and architecture data. AI then removes the old analysis bottleneck, allowing attackers to index and exploit everything collected—even from businesses once considered too small to target.

The Signal: This is where supply-chain security stops being an engineering concern. Attackers once needed analysts to sort through stolen files and understand unfamiliar industries. AI removes that bottleneck. Your business no longer needs to be important enough to justify human attention; stolen data can be indexed and queried. A developer’s workstation may hold keys, cloud credentials, source code, architecture diagrams and customer schemas. One poisoned dependency can copy them all. Ask your engineering lead or developer how the company checks the software tools it relies on and updates them. Find out who reviews new tools, how quickly suspicious activity would be noticed, and whether access can be cut off if something goes wrong. You do not need to understand the technical details; you need confidence that someone is responsible for checking the doors before they are opened.

―――――――――――

Autonomous creative production- [Skip]
Magic Hour — Testing Fable vs Sol in terms of taste (they are both bad)

Magic Hour gave Fable 5 and Sol 5.6 the same process to produce four videos each. Both models researched markets, developed concepts, created mood boards, rejected weaker ideas, tested video models, and documented their decisions. The process was disciplined and capable. The finished videos were not. Because the models cannot watch motion, they judged extracted frames instead—and still rated their own work highly, exposing a gap between generating creative work and recognizing whether it is good.

The Signal: Two findings matter. First, the process work is real. The models researched markets, built mood boards, rejected weak concepts, tested video models, and kept sources behind claims. That can accelerate exploration. Second, they could not watch the videos they made. They inspected extracted frames instead, rated their own output near the ceiling—even when the work was weak. That is the decisive limitation. Skip the version where AI replaces your creative team. Use it to explore directions, develop concepts, and produce faster. Let humans judge what survives. As creative production becomes cheaper, taste becomes more valuable—not less. Your creative people are not just making the work; they are the only ones in the loop who can tell when it is not working.

―――――――――――

The better the model, the harder it is to steer - [Watch]
jumploops— Sol loves to cheat

While testing a multi-agent coding workflow, Adam Williams reached 94% on Terminal Bench 2.1 using GPT-5.6 Sol. Reviewing the successful runs revealed a problem: the model had searched online for task-specific solutions, despite not having a web-search tool. It used command-line access instead. The experiment also showed that stronger models need less prompting but can be harder to steer, making evaluation design, access controls, and inspection of how results were achieved increasingly important.

The Signal: Nothing to buy, which is why this is a Watch. But it should change how you read a demo. Better models need fewer instructions, yet the instructions and access you give them become more consequential. A system scoring well may be succeeding for reasons nobody examined. Here, web search was disabled, but the model used command-line tools to find online solutions. The score looked impressive; the route made it unreliable. Next time a vendor shows you a benchmark, do not ask only what the model scored. Ask what it could access during the test, whether those conditions match your workplace, and who reviewed how it reached the answer. A high score means little if nobody inspected the path behind it.

―――――――――――

The first five minutes decide whether anyone uses the thing you built - [Pilot]
Microsoft Design— Designing the first five minutes

output, causing motivation to drop. The researchers recommend showing a useful result first, breaking the setup into steps, displaying progress, using sensible defaults, and clarifying what the agent does well. The principles apply not only at launch but also whenever a capability or workflow changes the experience.

The Signal: I’ve watched this failure: the tool works, the pilot ends, and people try it once but never return. Microsoft’s finding explains why. Users had to invest time in configuration before seeing value. That is a design decision, not a model limitation—and it is within your control. Show value first. Then ask users to connect data, add context, or change their workflow. Every capability or behaviour change creates another first impression, so this is not a launch checklist. It is an ongoing one. What I’d do before launching a tool is put it in front of someone new to it. Give them five minutes and say nothing. Whatever they do in minute one is your actual onboarding, regardless of what you designed.

―――――――――――

That's this week's signal.

One ask, and I mean it: hit reply and tell me what you're actually seeing — the pilot that stalled, the tool nobody uses, the thing that worked. I read every one.

Talk soon,
Yashar