SDK examples

The Python and TypeScript SDKs expose the same desktop contract. In TypeScript the desktop methods also live on computer.desktop. Method names are case-sensitive.

Method reference

GroupPythonTypeScriptPurpose
Screenscreenshot()screenshot()Full desktop as PNG bytes
screenshot_base64()screenshotBase64()Same, base64, ready for a vision model
screenshot_region(x, y, w, h)screenshotRegion(x, y, w, h)Region as PNG bytes
Clickclick(x, y, button="left")click(x, y, button)Click at pixels; button left|right|middle
left_click(x, y)leftClick(x, y)Left-button click
right_click(x, y)rightClick(x, y)Right-button click
-middleClick(x, y)Middle-button click
double_click(x, y)doubleClick(x, y)Double-click
Mousemove_cursor(x, y)moveMouse(x, y)Move the pointer without clicking
mouse_down(x, y, button="left")mouseDown(x, y, button)Press and hold a button
mouse_up(x, y, button="left")mouseUp(x, y, button)Release a held button
drag(from_x, from_y, to_x, to_y)drag(fromX, fromY, toX, toY)Click-drag between two points
Keyboardtype(text)type(text)Type literal text at the focus
key(key)key(key)Press one key or a + chord
hotkey(*keys)hotkey(...keys)Simultaneous combo
key_down(key) / key_up(key)keyDown(key) / keyUp(key)Hold or release a key
Scrollscroll(direction, clicks)scroll({ direction, clicks })Scroll up|down|left|right by N notches
scroll_up/down/left/right(clicks)-Convenience scrolls
Clipboardget_clipboard() / set_clipboard(text)getClipboard() / setClipboard(text)Read or write clipboard text
Screen infoget_screen_size()screenSize()Resolution { width, height }
get_cursor_position()cursor()Cursor { x, y } in screen pixels
Windowswindows()windows()List open windows
launch(app)launch(appName)Open an installed app
focus_window(id)focusWindow(id)Bring a window to the front
get_window_size(id) / set_window_size(id, w, h)windowSize(id) / resizeWindow(id, w, h)Read or set window size
get_window_position(id) / set_window_position(id, x, y)windowPosition(id) / moveWindow(id, x, y)Read or set window position
maximize_window(id) / minimize_window(id) / close_window(id)maximizeWindow(id) / minimizeWindow(id) / closeWindow(id)Change window state
Environmentget_desktop_environment()environment()DE name and version
set_wallpaper(path)setWallpaper(path)Set the background from a VM file
get_accessibility_tree()accessibilityTree()AT-SPI element tree
Shell and filesbash(cmd) / python(code)bash(cmd) / python(code)Run inside the VM
write_file(path, content) / read_file(path)writeFile(path, content) / readFile(path)Read or write a VM file

First loop

Feed a screenshot to a vision model

Screenshot_base64 returns a string you can pass straight through as image data.

b64 = computer.screenshot_base64()

import anthropic
client = anthropic.Anthropic()
response = client.messages.create(
    model="claude-opus-4-5",
    max_tokens=1024,
    messages=[{
        "role": "user",
        "content": [
            {"type": "image", "source": {"type": "base64", "media_type": "image/png", "data": b64}},
            {"type": "text", "text": "What is on screen, and where should I click?"},
        ],
    }],
)

Keyboard and mouse

Clipboard, scroll, and windows

Shell, Python, and files

These run inside the VM with no GUI interaction and are usually the fastest way to do multi-step work.

print(computer.bash("ls /workspace/Desktop"))
print(computer.python("print(1 + 1)"))

computer.write_file("/workspace/script.py", "print('hello')")
computer.bash("python3 /workspace/script.py")
print(computer.read_file("/workspace/.bashrc"))

CLI

miosa api GET "/computers/$COMPUTER_ID/desktop/screenshot" > before.png
miosa api GET "/computers/$COMPUTER_ID/desktop/accessibility-tree"
miosa api POST "/computers/$COMPUTER_ID/desktop/click" -d '{"x": 500, "y": 500}'
miosa api GET "/computers/$COMPUTER_ID/desktop/screenshot" > after.png

See also

Was this page helpful?