Aug 2026
Video as a tool call for agents.
Text was the easy medium. That is changing.
Agents already call search, code interpreters, and browsers. Video should be another tool output, not a separate app the user has to open.
The interface is simple: submit a prompt or upstream model answer, poll job status, receive a video_url. Cheap enough that a tutoring bot can afford to call it.
That is why we care about unit cost and near-real-time latency. If a render costs too much or takes too long, agents will never make video the default.
The next wave of LLM products will not only answer. They will show.