Chat and agent: a similar window, a different task
A conversation helps clarify an idea. An agent is meant to carry out a bounded action and show evidence of the result.
- Goal
- Explain the difference between asking a chat a question and assigning work to an agent.
- Tools
OpenAI Agents · agent work with tools
- Process
- Question → clarifying the need → bounded assignment → execution → checking a visible result.
- Result
- A table that distinguishes an answer, an action, and verification, plus a prompt for a small assignment.

In one conversation, I ask what I need to learn in order to launch a team of AI agents. In another, I want files on my Mac to open in a convenient program. Both begin with an ordinary message. In the first, I need to understand the topic better. In the second, I expect a change on the computer and a check that it works.
From a beginner’s perspective, the difference can be invisible. There is still a conversation window, my sentence, and an answer. But when working with an agent, that sentence can be followed by specific actions: reading a file, changing a setting, trying to open a document, making a correction, and checking again.
The easiest way to show it is through my conversations.
When a conversation helps establish what I am actually looking for
In the chat “Building a swarm of agents,” I ask how people organise this kind of work, what needs to be learned, whether ready-made solutions exist, and whether they run on a computer or in the cloud. I am interested in whether an application can be built in one or two days.
After the answer, I clarify that I probably already know the basics of coordinating work, and that I am looking for a way to parallelise it. In other words, a division of work that lets several workers act at the same time.
That is a concrete benefit of conversation: the first answer helps me name the need more precisely. I can say which parts I already understand and where I want to go next. But a description of a team of agents does not mean that such a team has been launched. The goal of “an application in two days” remains an experimental goal until there is execution and a checked result.
What you can take from this: when an answer goes back to the basics, clarify your starting point. For example: “I already understand how to divide roles. Now I need a first attempt in which two people or two processes really work at the same time.” This is a new example of an instruction, not a quote from my chat.
When I say: “set it up that way”
In a conversation on 22 September, I ask whether an .md file can open in a note-taking app by default. Markdown, the format of such files, is used to write text with simple markers for headings, lists, and links. After the options are explained, I write:
“set it up that way”
Here, the expected result changes. The agent is meant to apply the agreed setting on the computer.
First, a solution with TextEdit appears. I look at the effect and change my preference: I would rather use Obsidian or be able to open files in Codex. After my confirmation, the agent moves to Obsidian.
The next step is exactly what shows why completing a task means more than clicking a setting. The program launches, but the correct document is not yet in it. After opening the relevant folder as a collection of notes, the agent also checks an Obsidian link that leads to a specific file. The final readout of the window shows an open README document.
This attempt confirms opening a file through a link. The final report still contains a limitation, though: an ordinary double-click in Finder only launches Obsidian and does not always pass the document to it. The agent proposes an additional solution, but this record does not contain its execution. The original need to open files conveniently by double-clicking therefore cannot be considered fully handled.
What you can take from this: describe the end of a task in terms of what you will be able to do yourself. “I will double-click this document in Finder and see its text in the chosen application” is a much better criterion than “the setting was changed.” It also indicates how to reach the result.
An answer, execution, and verification are three different things
This situation can be used to set out a simple distinction. The first column describes the help I could ask for in conversation alone. The second corresponds to the type of actions from my example.
| Stage | I ask for an explanation | I assign the agent an action |
|---|---|---|
| Need | How do I open files like these? | Set the selected way of opening them. |
| Material for the work | My description and possibly a screenshot | Access to the right setting, application, and file |
| Execution | I receive instructions and carry them out myself | The agent carries out steps in the available environment |
| Unsuccessful attempt | I describe the problem in the next message | The agent can read the result, find the obstacle, and continue |
| Evidence | I understand the instructions | README opened through a link; the limitation of ordinary double-clicking is described |
This is not a fixed division between brands. ChatGPT can have tools for taking action, and I can ask Codex only for an explanation. What matters is what I assigned, which tools are available, and what was actually done. OpenAI describes agent work as repeating model and tool steps until a task is completed or stopped. How agents run.
How to move from a question to an assignment
My short “set it up that way” made sense because the preceding conversation had established the subject of the action. In a new chat, the same three words are not enough. The instruction needs to carry over the result, place of work, and method of checking.
Here is a template based on that situation. It is a proposal for your own use, not a record of the entire historical conversation:
I want Markdown files on my computer to open in [the chosen, installed application]. Change the default application for this file type. Use [a specific file] to check it. After the change, open this file with an ordinary double-click in the file manager and check that its text is visible, not only the program window. Do not change the document’s content or location. If only another method works, such as a special link, describe it separately. Do not treat it as confirmation that double-clicking works. If an additional configuration step is needed, explain its scope. At the end, state: what was changed, which file was checked, and whether it opened.
The same approach can describe a task involving a spreadsheet, document, or website. Instead of “take care of this,” state what condition should result and what you will use to check it. Access and consent must still match the specific task—a prompt alone does not give an agent access to the computer.
I look at what changed, not only whether someone said “done”
In the file example, it was necessary to distinguish launching the program, opening the document through a link, and opening it with an ordinary double-click. Only the last attempt answered the original need. A similar mistake is easy to make in other tasks: a file was saved, but in the wrong place; a page opens, but a button does not work; a report exists, but it rests on incomplete data.
That is why, after execution, I ask three things:
- What exactly changed? A file, setting, document, or only a proposal?
- How was it checked? What was opened, compared, or run?
- What is still unknown? Was one case checked, or the entire needed path?
An agent can carry out further steps and corrections within the assignment. I still decide whether the effect meets my need—just as I chose Obsidian after seeing TextEdit.
To start, choose one small, reversible action and set a visible completion condition. Then the difference between conversation and working with an agent stops being a definition. You can see it in a result you can check yourself. For an example of an agent working in live panels, see the domain-move story using Computer Use.
