Quickstart

Create a Computer, take a screenshot, act once, and screenshot again. This is the loop every computer-use agent runs, whether it is headless or watched by a person.

1. Create and start a Computer

start() returns once the desktop is running and accepting actions. Computer state is provisioning, running, stopped, recovering, error, or deleted.

2. The first loop

Capture, decide, act once, capture again. Do not issue several blind actions in a row.

Coordinates are screen pixels with the origin at the top-left. Read the current resolution with get_screen_size() if you need to scale a model’s output.

3. Watch it live

The same running desktop can be shown to a browser. On the platform, open the Computer and use the desktop button. In your own app, mint a short-lived stream URL on your backend and embed it.

See Embedding & streaming for the token flow, iframe policy, and expiry handling.

4. A full agent loop

import anthropic, os
from miosa import Miosa

miosa_client  = Miosa(api_key=os.environ["MIOSA_API_KEY"])
claude_client = anthropic.Anthropic()

computer = miosa_client.computers.create(name="browser-agent", template_type="miosa-desktop", size="small")
computer.start()
computer.launch("firefox")

for _ in range(10):
    b64 = computer.screenshot_base64()

    response = claude_client.messages.create(
        model="claude-opus-4-5",
        max_tokens=1024,
        messages=[{
            "role": "user",
            "content": [
                {"type": "image", "source": {"type": "base64", "media_type": "image/png", "data": b64}},
                {"type": "text", "text": "Navigate to miosa.ai and screenshot the homepage."},
            ],
        }],
    )

    # Parse an action from the model output, then act once.
    text = response.content[0].text
    if "left_click" in text:
        computer.left_click(500, 500)
    elif "type" in text:
        computer.type("https://miosa.ai\n")

computer.stop()

Next

Was this page helpful?