Models
Package optimized artifacts with explicit runtime, device, memory, and license compatibility.
convert → optimize → benchmarkThe deployment platform for on-device AI
Deliver private AI models to mobile and edge devices. CactusBrain handles distribution and updates. cellm runs inference on-device.
No token fees. No inference servers. App data stays on-device.
The CactusBrain stack
Text, vision, audio, embeddings.
Prepare, distribute, and manage releases.
Resolve, verify, install, and update.
Run inference on the device.
CactusBrain handles the journey to the device. Your SDK loads the model, and cellm runs inference locally. Explore each step below.
Example model ocr-lite
Total $42.90
The model travels.
Your data stays with you.
One platform connects model preparation, secure distribution, device-aware releases, and local execution without putting an inference server between your app and its users.
Package optimized artifacts with explicit runtime, device, memory, and license compatibility.
convert → optimize → benchmarkSign, encrypt, distribute, target, update, and observe model releases across a device fleet.
publish → deliver → updateThe open-source runtime that loads compatible models and executes inference on the user's hardware.
CPU · Metal · offline inferenceClear boundary: CactusBrain manages production delivery while cellm runs inference on the device, so prompts and model inputs never need to reach CactusBrain.
Target releases by device class, runtime support, memory budget, and rollout cohort. Keep one stable API while the model underneath improves.
Each SDK owns its workflow, validation, and developer experience. cellm remains the execution layer underneath.
On-device visual identification for apps, with matching that stays local.
On-device language model inference for private text features.
On-device personalization and temporal ranking without user data leaving the device.
On-device audio tagging and speech enhancement from microphone streams.
cellm runs optimized AI models directly on mobile and edge devices. CactusBrain SDKs handle the product workflows; inference stays in one open-source engine on the device.
Every release declares its task, supported targets, runtime requirements, access state, and licensing before it reaches a device.
Showing 6 of 29 models
| Model | Task | Targets | License | Access |
|---|---|---|---|---|
| fal.aiAuraFace Face EmbeddingBeta | Face embeddingVision | iOS / macOS / AndroidCPU | Apache 2.0Permitted | Entitlement requiredSign in |
| GoogleBlazeFace Back (Document and Group Range)Beta | Face detectionVision | iOS / macOS / AndroidCPU | CC-BY 4.0Permitted | Entitlement requiredSign in |
| GoogleBlazeFace Front (Selfie Range)Beta | Face detectionVision | iOS / macOS / AndroidCPU | CC-BY 4.0Permitted | Entitlement requiredSign in |
| Bryce Beattie / PiperCori High · British EnglishBeta | Speech synthesisAudio | iOScpu | MIT (voice repository); public-domain source audio; Apache-2.0 dictionaryPermitted | Entitlement requiredSign in |
| Bryce Beattie / PiperCori · British EnglishBeta | Speech synthesisAudio | iOScpu | MIT (voice repository); public-domain source audio; Apache-2.0 dictionaryPermitted | Entitlement requiredSign in |
| RikoroseDeepFilterNet3Beta | Speech enhancementAudio | iOS / macOS / AndroidCPU | MIT / Apache 2.0Permitted | Entitlement requiredSign in |
CactusBrain is building a measurement layer for real device performance: model size, peak memory, load time, throughput, energy use, and quality tradeoffs.
Unlimited on-device inference. Run ten requests or ten billion. CactusBrain does not charge for tokens, prompts, inference requests, CPU time, GPU time, or monthly active users because that work happens on your users' devices.
Build and validate your first on-device experience
≈ 640 downloads of an 80 MB model
Ship a production app with room to iterate
≈ 3,800 downloads of an 80 MB model
Operate multiple models across a growing fleet
≈ 15,000 downloads of an 80 MB model
Custom governance, capacity, and infrastructure
≈ 66,000 downloads of an 80 MB model
Builder includes this delivery volume.
Every plan runs models locally with no token, request, compute, or active-user meter.
| Included | Explorer$0 / month | Builder$15 / month$150 / year | Launch$49 / month$490 / year | ScaleFrom $199 / month |
|---|---|---|---|---|
| Limits | ||||
| On-device inference | Unlimited | Unlimited | Unlimited | Unlimited |
| Private models | 1 | 5 | 20 | Unlimited |
| Storage | 5 GB | 50 GB | 250 GB | 500 GB+ |
| Encrypted delivery / mo | 50 GB | 300 GB | 1.2 TB | 5 TB |
| Catalog downloads | Use delivery allowance | Use delivery allowance | Use delivery allowance | Use delivery allowance |
| Shipping | ||||
| Delta updates | ||||
| Version history & rollback | ||||
| A/B testing & staged rollouts | ||||
| BYO S3 / R2 bucket | ||||
| Support & governance | ||||
| Support | Community | Priority | Dedicated Slack | |
| SSO / SAML | ||||
| SLA & audit logs | ||||
Zero metering for inference. cellm stays free and open source. Prompts, tokens, CPU, GPU, and active users are never billed. Delivery is capped rather than billed without limit: paid plans deliver up to 125% of their included allowance, then model delivery pauses until the next period instead of invoicing on, so there is no runaway bill. Explorer gets no headroom and stops at its included 50 GB. Plans are billed in USD by Lemon Squeezy, our merchant of record, which accepts cards and PayPal worldwide and handles sales tax and VAT at checkout.
Access the developer workspace and available SDKs.
Attach a compatible, signed model artifact to your project.
The SDK resolves, delivers, verifies, installs, and loads it on-device.