Perplexity Portable Computer: local agents, zero token meter

Key takeaways

  • Perplexity Portable Computer launched Aug. 25, 2026 with Nvidia — a local version of Perplexity’s agentic “Computer” that packages model, harness, tools, and sandbox into one app.
  • On-device work burns no subscription token credits. Tasks start local by default; any cloud step asks permission first (with a PII preview before data leaves).
  • Day-one target: Nvidia DGX Spark and Linux boxes with RTX ≥24GB VRAM (~3090-class floor). Pro / Max / Enterprise tiers. Windows in September. No Apple silicon roadmap yet.

Cloud AI sold chat as a meter. Agents want to run for hours — reviewing folders, verifying themselves, iterating — which turns the meter into a bill. Perplexity Portable Computer’s pitch is blunt: park that counter at zero by running the agent on hardware you already bought.

Mode: rotate · Category: AI

What Perplexity Portable Computer is

Perplexity Computer is the company’s agent platform for multi-step knowledge work — documents, data, tools, connectors. Portable Computer is the same UI and harness, rewritten to run fully on-device first.

Nate, Perplexity VP of engineering for infrastructure and enterprise, on the Aug. 25 briefing: they brought “the exact same UI to a fully local app,” including the agent harness and inference needed to work offline.

Coverage: VentureBeat on the Nvidia partnership launch, The Verge on on-device Portable Computer.

Local-first: files stay, meter sits at zero

Company claims for local runs:

  • Model weights and user files can stay on the machine
  • Completed local work consumes no billing credits
  • Demo detail that landed: while a 27B model chewed a folder of 1099s / investment docs on a DGX Spark, the usual cloud-credit tally “is just parked at zero”

That is the product screenshot worth remembering. Chat was bursty. Agents are always-on token sinks. Nvidia’s Nader framed it as “insatiable demand for tokens” — local inference flips the economics because you are not paying per token for the long loop.

Privacy angle for finance / legal / health desks: the 1099 demo exists because people will not upload tax packs to a random cloud agent. Same instinct as locking down accounts with modern auth — what is a passkey — applied to document folders instead of logins.

Hardware bar: DGX Spark and 24GB RTX

Launch footprint:

  • Primary showcase: Nvidia DGX Spark desktop AI box
  • Also: Linux PCs with Nvidia RTX GPUs and at least 24GB VRAM (roughly GeForce RTX 3090 or newer)
  • Subscribers: Pro, Max, Enterprise Pro, Enterprise Max
  • OS: Linux now; Windows in September
  • Not on the map yet: Apple silicon — Nate said the focus is Nvidia hardware

24GB is a high floor. Most laptops fail it. This is not “local AI for everyone.” It is local AI for people who already bought (or will buy) workstation-class Nvidia iron — which is also why Nvidia cares. DGX Spark needed a killer app that is not “assemble Ollama yourself.” Related chip tape: NVIDIA stock into earnings week.

Models, harness, and the sandbox rule

At launch users can run Qwen 3.8 27B or PPLX 27B (Perplexity post-trained on its harness), with Nvidia Nemotron 3.5 Lightning said to be coming soon. Inference under the hood uses vLLM; advanced mode can point at your own endpoint.

The research claim Perplexity is pushing: small local models need a co-designed harness, not a frontier-sized tool surface. Their stack uses a short system prompt, a small core tool set, on-demand “skills,” and CLI-style connectors instead of fat MCP servers that burn context. They say Qwen’s advertised huge context still degrades past ~100K tokens in practice, so the harness stays lean on purpose.

Security detail that matters: always-on OS sandbox. If the sandbox is unavailable, the harness disables itself rather than running tools with full user permissions — the opposite of many DIY agent setups that happily shell out as you.

Internal “Local Knowledge Work Bench” scores (company’s own bench, planned to open-source): Computer + Qwen 27B on Spark 82.6%; PPLX 27B 85.4%; open harnesses with the same model trailed. Treat vendor benches as directional, not gospel.

Hybrid escalate: when the cloud is allowed

Local-first does not mean air-gapped forever. Demos showed:

  • Analyze a funnel CSV locally, then push results to Slack via connectors
  • Connectors also named for Google Drive, Gmail, GitHub
  • Escalate a hard step to a frontier “advisor” (Claude Opus cited in one bench) after user permission

On Terminal Bench 2.1 (company figures): local Qwen ~59.6% at ~$0 marginal; escalate to cloud advisor ~73% at ~$0.42/task; frontier alone ~82% at ~$0.65/task. Escalation closes part of the gap cheaper than full cloud — and before anything leaves, a PII classifier plus a preview of outbound context. The remote model returns text guidance only; it does not get local files or tools.

That is the adult version of agent permissioning — closer to “confirm before act” than a one-word SMS cancel. Different product class, same week as finance agents: Rocket Money Rowan.

Not Ollama — the agent layer is the product

Asked how this differs from Ollama-style local inference, Nate drew the line at the agent harness. Getting a model to answer is solved-ish. Getting a 27B model to run a long tool loop without collapsing is the product.

Nvidia’s Nader compared DIY agent stacks to the ocean: “The deeper you go, the deeper it gets.” Portable Computer sells the appliance path — model + harness + sandbox + connectors — for buyers who found DGX Spark easier to purchase than to operate.

Who this is actually for

  • Pros / enterprises with sensitive document folders and Nvidia iron
  • Teams that already burn Max/Enterprise budgets on long agent runs
  • Engineers who want a packaged local Computer, not a weekend of harness glue

Who should not expect magic:

  • Mac-only users (not on roadmap)
  • RTX 8–16GB laptops (under the 24GB floor)
  • Anyone who needs frontier-only reasoning on every step without paying the escalate tax

The limits Perplexity still admits

  • Compact models still trail frontier reasoning; escalation “narrows but does not fully close the gap”
  • Headline benches are mostly Perplexity’s own
  • Linux-first; Windows later; Apple absent
  • Hardware cost is the real subscription — tokens go free, GPUs do not

Perplexity Portable Computer is the clearest Aug. 25 statement yet that agent economics want to leave the data center for the desk — when the desk has 24GB of Nvidia VRAM and a Pro badge. Local work at zero credits is the feature. Permissioned cloud escalate is the escape hatch. Everything else is still a workstation product wearing a consumer headline.

Product explainer only — not investment advice. Confirm hardware requirements, plan tier, and privacy settings in Perplexity’s own docs before buying silicon for this stack.

Leave a Comment