Integration
Run Pinchy agents against local Ollama for fully offline operation. Or against Ollama Cloud for EU-friendly managed models. Same provider, different deployment — switched per agent.
Self-hosting used to mean giving up capability. Frontier models were closed, local models were rough, the quality gap was real. That's changed. Modern open-weight models — Qwen, Llama, DeepSeek — are good enough for a large fraction of enterprise work, especially the inbox-type work Pinchy agents specialise in.
Pinchy supports Ollama as a first-class provider. You can wire up a local Ollama server and never talk to an external API again, or use Ollama Cloud for the convenience of managed hosting without leaving the Ollama ecosystem.
Two Deployments
Install Ollama on a server you control — a workstation, a GPU box in your datacentre, a dedicated VM. Pinchy connects via a base URL. No external calls. Suitable for air-gapped and regulated environments.
Managed hosting by Ollama with the same API contract. EU-friendly alternative to US frontier providers when you want managed convenience without the CLOUD Act overhead.
Mix and match. Sensitive finance agent → local Ollama. Customer-facing draft agent → Claude. Internal helper → Ollama Cloud. The provider is a property of the agent, not the installation.
For EU companies navigating the US CLOUD Act and GDPR reality, Ollama — especially local Ollama — is the cleanest story. No US-based cloud provider in the loop, no transatlantic data transfer, no CLOUD Act exposure. The architecture is compliant by construction, not by DPA.
Pair that with Pinchy's self-hosted deployment, the audit trail, and scoped permissions, and you have a stack that can reasonably be deployed in banking, healthcare, legal, and government environments.
Learn More
FAQ
Yes. Pinchy supports Ollama as a first-class model provider, for both local Ollama installations and Ollama Cloud. You pick the model per agent — a local Qwen or Llama for sensitive work, a hosted frontier model for another agent, the choice is per-agent.
Yes. Pair Pinchy with local Ollama and nothing leaves your network. No external API calls, no cloud dependency, no telemetry leak. For regulated industries or air-gapped environments, this is the whole point.
Local Ollama runs models on your own hardware — your CPU or GPU, your latency, your limits. Ollama Cloud hosts models on Ollama's infrastructure with an API contract very similar to the local runtime. Pinchy talks to both through the same provider; you switch by changing a base URL.
Yes. Different agents can use different providers. A finance agent might use local Qwen via Ollama, a customer-facing agent might use Claude, an internal writing helper might use GPT. Pinchy is model-agnostic per agent.
Self-host Pinchy yourself in minutes, or book a call to talk it through. Your choice.
Or email us: info@heypinchy.com