Running agents on local models has one classic failure mode: the framework quietly assumes one specific provider, and then your local agent is an API client with extra steps. The demo works. The architecture lies. - Pick a framework with a pluggable model layer: develop against an API, test against a local model, deploy wherever privacy or budget demands. - Local models change the money math. Agents loop and retry, and a hundred internal calls cost zero locally versus terrifying on metered pricing. That difference compounds fast. - Pair with a security review anyway. Local inference shrinks the attack surface, but your prompts, tools, and permissions still need auditing. - Checklist: pluggable model layer, real open-weight support, sane small-context handling, examples showing actual local setups. PrivateLLM deploy puts the private LLM underneath your agents on AWS for $50 plus usage. The framework database stays free.