How private LLM inference actually works

Presented by

James Wilson
James Wilson

Technology Editor

In this podcast episode James Wilson chats with Tinfoil co-founder Tanya Verma about how you can run a powerful LLM in the cloud without the inference provider seeing your prompts.

Tanya talks James through how private inference works, from trusted execution environments and hardware attestation, to TLS termination and GPU isolation. Customers can verify the exact code and model processing their data, while Tinfoil and its infrastructure providers remain locked out.

That’s clever engineering… but who really needs it? Is private inference only useful if you’re doing something bad, or will it become a privacy baseline like TLS? James and Tanya discuss the costs and trade-offs, and how open weights make private inference more transparent and trustworthy.

How private LLM inference actually works
0:00 / 83:49