Key takeaways
- Perplexity Portable Computer launched Aug. 25, 2026 with Nvidia — a local version of Perplexity’s agentic “Computer” that packages model, harness, tools, and sandbox into one app.
- On-device work burns no subscription token credits. Tasks start local by default; any cloud step asks permission first (with a PII preview before data leaves).
- Day-one target: Nvidia DGX Spark and Linux boxes with RTX ≥24GB VRAM (~3090-class floor). Pro / Max / Enterprise tiers. Windows in September. No Apple silicon roadmap yet.
Cloud AI sold chat as a meter. Agents want to run for hours — reviewing folders, verifying themselves, iterating — which turns the meter into a bill. Perplexity Portable Computer’s pitch is blunt: park that counter at zero by running the agent on hardware you already bought.
Mode: rotate · Category: AI
What Perplexity Portable Computer is
Perplexity Computer is the company’s agent platform for multi-step knowledge work — documents, data, tools, connectors. Portable Computer is the same UI and harness, rewritten to run fully on-device first.
Nate, Perplexity VP of engineering for infrastructure and enterprise, on the Aug. 25 briefing: they brought “the exact same UI to a fully local app,” including the agent harness and inference needed to work offline.
Coverage: VentureBeat on the Nvidia partnership launch, The Verge on on-device Portable Computer.
Local-first: files stay, meter sits at zero
Company claims for local runs:
- Model weights and user files can stay on the machine
- Completed local work consumes no billing credits
- Demo detail that landed: while a 27B model chewed a folder of 1099s / investment docs on a DGX Spark, the usual cloud-credit tally “is just parked at zero”
That is the product screenshot worth remembering. Chat was bursty. Agents are always-on token sinks. Nvidia’s Nader framed it as “insatiable demand for tokens” — local inference flips the economics because you are not paying per token for the long loop.
Privacy angle for finance / legal / health desks: the 1099 demo exists because people will not upload tax packs to a random cloud agent. Same instinct as locking down accounts with modern auth — what is a passkey — applied to document folders instead of logins.
Hardware bar: DGX Spark and 24GB RTX
Launch footprint:
- Primary showcase: Nvidia DGX Spark desktop AI box
- Also: Linux PCs with Nvidia RTX GPUs and at least 24GB VRAM (roughly GeForce RTX 3090 or newer)
- Subscribers: Pro, Max, Enterprise Pro, Enterprise Max
- OS: Linux now; Windows in September
- Not on the map yet: Apple silicon — Nate said the focus is Nvidia hardware
24GB is a high floor. Most laptops fail it. This is not “local AI for everyone.” It is local AI for people who already bought (or will buy) workstation-class Nvidia iron — which is also why Nvidia cares. DGX Spark needed a killer app that is not “assemble Ollama yourself.” Related chip tape: NVIDIA stock into earnings week.
Models, harness, and the sandbox rule
At launch users can run Qwen 3.8 27B or PPLX 27B (Perplexity post-trained on its harness), with Nvidia Nemotron 3.5 Lightning said to be coming soon. Inference under the hood uses vLLM; advanced mode can point at your own endpoint.
The research claim Perplexity is pushing: small local models need a co-designed harness, not a frontier-sized tool surface. Their stack uses a short system prompt, a small core tool set, on-demand “skills,” and CLI-style connectors instead of fat MCP servers that burn context. They say Qwen’s advertised huge context still degrades past ~100K tokens in practice, so the harness stays lean on purpose.
Security detail that matters: always-on OS sandbox. If the sandbox is unavailable, the harness disables itself rather than running tools with full user permissions — the opposite of many DIY agent setups that happily shell out as you.
Internal “Local Knowledge Work Bench” scores (company’s own bench, planned to open-source): Computer + Qwen 27B on Spark 82.6%; PPLX 27B 85.4%; open harnesses with the same model trailed. Treat vendor benches as directional, not gospel.
Hybrid escalate: when the cloud is allowed
Local-first does not mean air-gapped forever. Demos showed:
- Analyze a funnel CSV locally, then push results to Slack via connectors
- Connectors also named for Google Drive, Gmail, GitHub
- Escalate a hard step to a frontier “advisor” (Claude Opus cited in one bench) after user permission
On Terminal Bench 2.1 (company figures): local Qwen ~59.6% at ~$0 marginal; escalate to cloud advisor ~73% at ~$0.42/task; frontier alone ~82% at ~$0.65/task. Escalation closes part of the gap cheaper than full cloud — and before anything leaves, a PII classifier plus a preview of outbound context. The remote model returns text guidance only; it does not get local files or tools.
That is the adult version of agent permissioning — closer to “confirm before act” than a one-word SMS cancel. Different product class, same week as finance agents: Rocket Money Rowan.
Not Ollama — the agent layer is the product
Asked how this differs from Ollama-style local inference, Nate drew the line at the agent harness. Getting a model to answer is solved-ish. Getting a 27B model to run a long tool loop without collapsing is the product.
Nvidia’s Nader compared DIY agent stacks to the ocean: “The deeper you go, the deeper it gets.” Portable Computer sells the appliance path — model + harness + sandbox + connectors — for buyers who found DGX Spark easier to purchase than to operate.
Who this is actually for
- Pros / enterprises with sensitive document folders and Nvidia iron
- Teams that already burn Max/Enterprise budgets on long agent runs
- Engineers who want a packaged local Computer, not a weekend of harness glue
Who should not expect magic:
- Mac-only users (not on roadmap)
- RTX 8–16GB laptops (under the 24GB floor)
- Anyone who needs frontier-only reasoning on every step without paying the escalate tax
The limits Perplexity still admits
- Compact models still trail frontier reasoning; escalation “narrows but does not fully close the gap”
- Headline benches are mostly Perplexity’s own
- Linux-first; Windows later; Apple absent
- Hardware cost is the real subscription — tokens go free, GPUs do not
Perplexity Portable Computer is the clearest Aug. 25 statement yet that agent economics want to leave the data center for the desk — when the desk has 24GB of Nvidia VRAM and a Pro badge. Local work at zero credits is the feature. Permissioned cloud escalate is the escape hatch. Everything else is still a workstation product wearing a consumer headline.
Product explainer only — not investment advice. Confirm hardware requirements, plan tier, and privacy settings in Perplexity’s own docs before buying silicon for this stack.