Chat used to be the place where the user asked. Work used to be the place where somebody did. ChatGPT Work collapses that separation into one surface, and that is why the product matters beyond the launch itself.

OpenAI's RSS description says ChatGPT Work is an agent that can “take action across your apps and files,” “stay with a project for hours if needed,” and “turn a goal into finished work.” That sentence is the shift. The product is no longer only a conversational assistant with better answers. It is being positioned as a worker that can hold a project, operate inside connected systems, and produce an outcome.

Once chat becomes the worker, the prompt box is no longer the main design object. The job boundary is.

A normal chat interface can survive a little ambiguity. The user asks, clarifies, and moves on. A work agent cannot safely live in that looseness. If the same box can answer a casual question, edit files, inherit project instructions, use work plugins, touch connected apps, and continue for hours toward a finished output, the interface has to show whether the worker's role is visible enough to manage.

Chat is a conversation surface. Work is a role surface.

The public reaction to ChatGPT Work was revealing because much of the confusion was not about model intelligence. People were trying to understand what kind of worker had appeared inside a familiar chat product.

On Hacker News, one commenter described the move as OpenAI “catching up to Anthropic's Cowork.” Another called it “just retiring the ChatGPT app and replacing it with Codex for the masses.” Others asked, “How is this different than Codex?” and answered dryly, “It has a new name!” One person reduced the visible distinction to the affordance itself: “The button changes from green to blue.”

That sounds like ordinary launch snark until it hits the operating problem. A user wrote, “I'm so confused. I'm still so unclear when I'm supposed to use one or the other.” Another asked, “what happens when your workflow involves both coding and general knowledge work?” Another question was even more basic: “where am I supposed to casually chat?”

Those are role-boundary complaints. The user is trying to determine whether the system is a conversational partner, a coding worker, a file operator, a general work agent, or some mixture that changes authority based on context the user cannot see. Software has always had modes, but AI modes carry labor authority. When the mode determines whether a language-driven system can act across apps and files, continue for hours, or inherit hidden instructions, ambiguity becomes an operations risk.

The hidden role contract has to move into the product.

Every useful worker operates inside a contract, whether the company writes it down or not. A human employee has a job, a manager, a workspace, a scope of authority, tools, standards, memory, and a definition of done. An AI agent inherits instructions, context, permissions, connected files, app access, tool affordances, model behavior, and recent conversational residue. Some properties may be obvious. Some may be buried in a workspace, a plugin, a developer instruction, a folder, a project setting, or a previous thread. The user still experiences one box.

That is the dangerous compression. The worker's real role may be assembled from layers the human did not consciously assign at the moment of delegation. One HN commenter pointed toward a “global AGENTS.md file” and hidden developer instructions. Another said, “I want my chat isolated from work contexts.” People do not only want more capability. They want to know when context has become authority.

A work agent therefore needs a visible job contract at the moment of use. Not a legalistic permission wall. Not a generic safety disclaimer. A practical operating boundary that tells the user what job the worker is doing, what context it inherited, what systems it can touch, what artifact it is expected to produce, what state it is in, and what condition stops the work.

If those properties remain hidden, the human has to manage the agent by intuition. That may work in a demo. It does not work when the agent is inside client work, finance work, hiring work, delivery work, code work, or professional-service operations where a wrong inherited assumption can travel through the artifact.

The button is less important than the boundary.

A live Work button is a useful signal because it shows the interface trying to name a new mode of labor. But a button alone cannot carry the role.

A button can tell the user that work mode exists. It cannot tell the user what job is active, what files were pulled into scope, which apps are reachable, whether the agent is allowed to write or only read, whether it is still working or waiting, or whether the result should be treated as a draft, a decision, a deliverable, or a request for review. Operators should therefore ask what the product shows before and during the action: the job, inherited context, separation from casual chat, authority, and expected artifact.

A good work interface should let the user answer a simple question at any time: what is this worker doing right now?

If the answer requires reading system prompts, remembering plugin behavior, guessing which folder is loaded, or trusting that a blue button means the right thing, the product may be powerful, but the work is not yet legible.

Finished work requires a visible artifact.

OpenAI's phrase “turn a goal into finished work” is the right ambition. It also raises the standard for evidence. Finished work is not the same as generated output. A professional-service leader cannot run a business on fluent completion language. They need to know what artifact was produced, what evidence supports it, and what remains unresolved.

The Ethereum Foundation's article on running AI agents against Ethereum protocol code gives the adjacent proof. The memorable line was not that agents found bugs. It was that “The agents found real bugs,” and then the article sharpened the real bottleneck: “Agents finding bugs wasn't the surprise. The surprise was how little of the work went into finding them, and how much went into telling the real bugs from the ones that just looked real.”

That distinction belongs far beyond code security. A candidate is not a finding. A summary is not a decision. A draft is not client-ready work. Ethereum's article names the proof requirement cleanly: “A candidate isn't a finding until there's a self-contained artifact that reproduces the failure against the real code.”

For operators, the same rule applies to every agent role. A sourcing agent should produce a source-linked comparison that makes the decision inspectable. A client-prep agent should produce unresolved commitments, likely risks, and source evidence. A file-work agent should show the changed artifact, the assumptions used, and the review condition.

The old bottleneck was whether AI could produce enough. The new bottleneck is whether the organization can trust what was produced. As the Ethereum piece puts it, “The bottleneck didn't go away. It moved from finding bugs to trusting the results.” ChatGPT Work and similar products will make that sentence operational across normal business work.

The visible job has six parts.

A practical agent boundary does not need to be complicated. It needs to be visible. Before a founder or operator lets a chat-based worker into real work, the interface and the operating process should expose six things.

First, the job. The worker should have a named role for the current assignment, not a vague promise to help. “Prepare the partner meeting from the last thirty days of account activity” is a job. “Help with the client” is fog.

Second, inherited context. The agent should show what project, thread, files, instructions, memory, or work context it is carrying into the task. This is where many failures begin. The human thinks they are asking a fresh question. The agent may be acting with the weight of previous work, global instructions, workspace assumptions, or connected folders.

Third, authority. The user should know what the agent can read, write, send, change, and escalate for approval. Authority should be part of the active job display.

Fourth, artifact. Every serious work agent should know what object should exist when the job is done: memo, diff, spreadsheet, message draft, triage report, issue list, source-linked recommendation, meeting brief, reproduction script, client update. The artifact turns labor into something reviewable.

Fifth, state. The agent should make clear whether it is chatting, working, waiting on the human, reviewing its own output, blocked, or stopped. State matters because human attention depends on it. If the user cannot tell whether the worker is still pursuing the goal or merely answering conversationally, supervision becomes guesswork.

Sixth, stop condition. Work needs a clean ending. The agent should know when to stop because the artifact is complete, because authority ran out, because evidence was insufficient, because the cost or time budget was reached, or because the next step requires human judgment. Without a stop condition, “stay with a project for hours” can mean persistence, drift, or expensive ambiguity.

These six properties are not bureaucracy. They are the minimum visible shape of delegated work.

The next interface is managerial.

ChatGPT Work points to a larger product direction. Chat is becoming a role interface. The user is no longer only typing requests into a model. The user is assigning work to a bounded artificial worker that may operate across systems, hold context, and return an artifact.

That means the human's job changes too. The operator is not merely prompting. The operator is managing the boundary between conversation and labor. They have to know when they are exploring, when they are delegating, when they are reviewing, and when they are stopping the worker.

The companies that get value from agents will not treat every new work button as magic. They will define jobs clearly, expose inherited context, limit and name authority, require artifacts, show state, and teach agents when to stop.

The naming matters because the mechanism will repeat. A chat surface can feel familiar while the system underneath becomes a worker. Once that happens, the organization cannot manage the product as a better chat app. It has to manage it as labor entering the operating system.

The worker may live in the chat box. The job cannot stay hidden there.

Sources

Stephen Nickerson.
Built for operators who need AI agents they can test, trust, and improve.