Computer Use at work: do it and show the result
How I ask AI to act in a program's interface, then check the result on a specific screen or in a file.
- Goal
- Show how I combine AI acting in an interface with conversation, checking the result, and making a correction.
- Tools
Computer Use — OpenAI · operating an interface; Playwright · checking page elements
- Process
- Goal and scope → action in an interface → showing the state → human feedback → correction or recording a limitation.
- Result
- A practical way to delegate small actions in a program, with a prompt that requires the result to be shown and verified.

Open the page and show the transition between views. Change the time of an event in the calendar. Set which app Markdown files should open in. These are tasks I have genuinely delegated to AI in my conversations.
They share a simple need: I want an agent to do work in a program I use and then show me the result. Computer Use lets it act through the interface—opening views, selecting buttons, entering data, and checking what changed.
The most interesting part of how I use this feature is combining action with conversation. The agent shows something, I see a problem, clarify what I expect, and we can move on to a correction. You can see this both when building a website and when doing ordinary computer housekeeping.
What I mean here by Computer Use
I mean operating applications through their interface. The agent reads the available state, performs an action, and checks the screen again or the information returned by the tool. My records include operations in a browser as well as in Mac applications.
Technically, this does not have to mean clicking image coordinates alone. Some browser operations used Playwright—a tool that can, among other things, find buttons and fields and operate a page. Alongside that, the agent used screenshots and descriptions of interface elements. These different ways of controlling an interface fit within the broader approach described in the OpenAI Computer Use documentation.
This distinction is useful when describing results. In a website task, an agent may change code in files and use Computer Use to view and check the running version. I do not attribute the entire process to one tool.
The website: show me how it works
While working on Adrian Lab, I asked for the local version to be opened beside the conversation, for subpages to be shown, and for transitions to be checked. I wanted to see the behaviour and respond to a specific place.
In a conversation on 1 October, I also described how I run a Live Demo. I ask the agent to show a transition, talk by voice, and stop it when I have feedback. I can ask for that feedback to be recorded for later or for a small correction right away. After the demo, a list of feedback remains for organised follow-up. I want to see larger changes again.
This describes my way of working. The same conversation also contains concrete tool actions: opening sections of the website, following a link to the demo description, returning to the previous route, and checking the view at different window widths.
A very ordinary problem appeared then: a light strip above the footer. I also reported that the content of a tile about the podcast and newsletter had changed. The agent reproduced the view, checked the layout, and asked me to clarify which strip I meant. I pointed to the one above the footer.
The work then combined several things: a code correction, another look at the page, and a check of the tile's behaviour. After clicking, the readout showed the right view with a choice of podcast or newsletter. Checking a narrower window also revealed clipped text, so the correction covered that specific case too. Finally, I looked at the result and accepted it.
That is what I use this kind of demo for: talking about something that can be seen. “There is an empty strip here” or “this tile opens the wrong content” gives a much more concrete starting point for work than a general “fix the website.”
This does not mean that one demo confirms the correctness of the entire product. It shows whether the checked fragment behaves as expected. Other tests still have their own role.
Prompt to use—a proposal based on this way of working:
Open the local version of the project and show me how [a specific route] works. Go through it in sequence, as a user would. After each important step, show the result and let me give feedback. When I say “stop”, pause the demo. Record the feedback along with its location and the expected behaviour. You may make small corrections within the agreed scope. After the demo, organise the feedback, make the corrections, and check them. Show the changed fragments again. This stage concerns a local demo; we will agree on publication separately.
Geometry editor: go through the process on a floor plan
I also have a more striking example. During a voice conversation on 10 September, I asked the agent to download a sample floor plan and show me, in the interface, the process of working it up.
The record of this attempt shows actual actions: opening the file picker, uploading the plan, clicking successive outline points, and dividing up the interior. The editor showed the created room and information about changes awaiting saving. Later it returned confirmation that the room changes had been saved in the draft.
This is stronger evidence than an agent's answer alone that it can draw. An effect was created in the interface. The description must remain precise, however: confirmation that the draft was saved does not check scale, dimensions, or correspondence with the plan. Nor is it evidence that the whole process or final import was completed.
I treat this case as an example of checking a tool by using it. In another attempt like this, it would be worth evaluating two things separately: whether the agent can operate the editor and whether the drawn geometry is correct. Only the second stage would allow conclusions about the quality of the technical work.
Calendar: change this one appointment
I have also delegated much less striking work to Computer Use. In September, I provided a schedule of match dates to add to my private calendar. A few days later, the time of the first match changed. I asked for that entry to be corrected while leaving the other dates unchanged.
The agent opened Google Calendar, found the existing events, and went to the right entry. The search results showed six events. That matters: when updating something, you first need to recognise what already exists, so you do not add another copy of the same appointment.
Saving produced an error. The tool reported that the indicated interface element was already out of date or no longer existed. A first attempt to click the button was not enough.
After reading the interface again, the agent repeated the save. The calendar displayed a message that the event had been moved to its new time. Such a message is evidence from the application; an intention to click is not. For an important appointment, it is worth opening the saved entry as well and checking its details.
This case also shows a limitation of work through an interface. The page state can change between reading and acting. An agent must be able to recognise an unsuccessful attempt and establish what is now visible. Without that, it is easy to confuse an issued instruction with an achieved result.
A practical prompt for a similar correction:
In my calendar [identify the correct calendar], find the event [name and date] and change its time from [old] to [new]. Keep its current duration and all other details. Do not create a new event. If more than one entry matches, show the difference before editing. After saving, open the event again and check the date, time, time zone, and calendar. State what actually changed.
This is a new template to use, not a verbatim record of my historical instruction. It adds an explicit condition to check after saving.
Mac: a setting is only the beginning
Another example concerned .md files, which are text files in Markdown format. I asked whether it was possible to set a convenient program as their default opener.
The agent went through Finder and the “Open With” option. It first set TextEdit and checked that a sample file opened. After seeing the result, I decided I would prefer Obsidian or the ability to open files in Codex. In the end, I accepted Obsidian.
There was another change to the setting and another test. Launching Obsidian alone did not yet mean that I had the right document in front of me. The appropriate folder also had to be opened as a vault, meaning a collection of files managed by Obsidian. The final readout of the window indicated that a specific README file was open in that application.
It is a small example, but it captures the point of working with an agent well. I can see the first solution, change my mind, and check the second. The completion criterion is practical: I can open and read the file as I wanted.
This test confirmed that it worked for the checked file and folder. It is not a promise that every Markdown document, in every location, will always open the same way.
What these tasks have in common
The same order recurs in these stories: I define a need, the agent acts in an interface, and then we check the result. Sometimes I clarify what I mean on the way. Sometimes an unsuccessful attempt needs fixing.
| Task | What needs to be identified before acting | What to check at the end |
|---|---|---|
| Website demo | The project version and the route to go through | Whether the view, content, and transitions match the agreement |
| Geometry-editor attempt | The plan, the test environment, and the scope of the work | The saved draft and the correctness of the geometry, separately |
| Event correction | The right calendar and existing entry | The saved event details, not just clicking “Save” |
| Opening files | The file type, selected application, and scope of the setting | Whether a sample document actually opens in the expected way |
That is why I suggest describing a task in three sentences: what must change, where action is allowed, and how we will know it worked. This helps both the agent and me during acceptance.
I have not measured from these conversations how much time Computer Use saved me. Nor do I draw effectiveness statistics from a few selected cases. What I do see is a concrete use: the agent performed actions in applications, and I could respond to their effects.
Access to a program needs a clear scope
In practice, you need to identify the right account, document, or environment. In the calendar, it mattered that this was my private calendar. For the website, it was a demo of the local version and a separate decision about publication. For files, it was selecting the default application for a given document type.
It is also worth distinguishing what an agent can see from what it is authorised to do. Text on a page or in a document should not extend the assignment on its own. The official Computer Use documentation recommends limiting access to the necessary scope, supervising actions with significant consequences, and checking the real result. A prompt does not replace the tool's permissions and safeguards.
For me, the practical boundary is this: viewing a local screen, changing a specific entry, and publishing content are different actions. The agent should know which of them was assigned. If it encounters an access block or cannot confirm a result, I need clear information about that point.
Computer Use makes the most sense to me when I can pair an instruction with seeing the effect: take this step, show the result, and I will say whether that is what I meant. That is how I worked on the website, corrected the calendar, and adjusted how files open on my Mac.
