CactusBrain

The deployment platform for on-device AI

Ship AI models to any device.

Deliver private AI models to mobile and edge devices. CactusBrain handles distribution and updates. cellm runs inference on-device.

No token fees. No inference servers. App data stays on-device.

From a model to a working feature.

CactusBrain handles the journey to the device. Your SDK loads the model, and cellm runs inference locally. Explore each step below.

Interactive walkthrough Sample data
Receipt readerOne delivery flow. A different capability in your app.
Managed delivery

CactusBrain

Example model ocr-lite

Your deviceRuns locally
Receipt reader
Sample input
DESERT GENERAL STORE
Field notebook $18.50
Canvas tote $19.40
Cold brew $5.00

Total $42.90

Recognized textTotal: $42.90Processed on this device
Your input stays here.

The model travels.
Your data stays with you.

Network for model delivery. Local execution for inference.

Explore model delivery

One platform. From model to device.

One platform connects model preparation, secure distribution, device-aware releases, and local execution without putting an inference server between your app and its users.

01 / Prepare

Models

Package optimized artifacts with explicit runtime, device, memory, and license compatibility.

convert → optimize → benchmark
02 / Operate

CactusBrain

Sign, encrypt, distribute, target, update, and observe model releases across a device fleet.

publish → deliver → update
03 / Execute

cellm

The open-source runtime that loads compatible models and executes inference on the user's hardware.

CPU · Metal · offline inference

Clear boundary: CactusBrain manages production delivery while cellm runs inference on the device, so prompts and model inputs never need to reach CactusBrain.

CactusBrain selects the right artifact for each device.

Target releases by device class, runtime support, memory budget, and rollout cohort. Keep one stable API while the model underneath improves.

Device classSelected artifactBackend
Older phoneAssistant · smallCPU
Modern phoneAssistant · mediumMetal / GPU
Desktop classAssistant · largeAccelerated

Choose a capability. Build on one runtime.

Each SDK owns its workflow, validation, and developer experience. cellm remains the execution layer underneath.

01Developer Preview

CactusBrain Vision

On-device visual identification for apps, with matching that stays local.

Explore capabilities
  • Image embeddings
  • Custom visual collections
  • Multiple reference images
  • Similarity search
  • Unknown-object rejection
  • Fully local matching
Explore CactusBrain Vision →
02Available

CactusBrain Text

On-device language model inference for private text features.

Explore capabilities
  • Local text generation
  • Qwen 2.5 and LFM packages
  • CPU and Metal execution
  • Offline after model download
Explore CactusBrain Text →
03Developer Preview

CactusBrain Personalize

On-device personalization and temporal ranking without user data leaving the device.

Explore capabilities
  • Temporal profiles
  • On-device ranking modes
  • Local state persistence
  • Zero server dependencies
Explore CactusBrain Personalize →
04Developer Preview

CactusBrain Signal

On-device audio tagging and speech enhancement from microphone streams.

Explore capabilities
  • Local audio tagging
  • Speech enhancement
  • Signal input contract
  • Low-latency edge processing
Explore CactusBrain Signal →

The engine inside your app.

cellm runs optimized AI models directly on mobile and edge devices. CactusBrain SDKs handle the product workflows; inference stays in one open-source engine on the device.

CPU + Metal
Make use of the hardware already in your users’ hands.
Local execution
Process model inputs on-device, with no inference server.
RustOpen sourceQuantized inferenceC ABI / FFICLI
CactusBrain / core3D illustration
Rooted in your device.Local inference, illustrated

Find the right model for your app.

Every release declares its task, supported targets, runtime requirements, access state, and licensing before it reaches a device.

Showing 6 of 29 models

ModelTaskTargetsLicenseAccess
fal.aiAuraFace Face EmbeddingBetaFace embeddingVisioniOS / macOS / AndroidCPUApache 2.0PermittedEntitlement requiredSign in
GoogleBlazeFace Back (Document and Group Range)BetaFace detectionVisioniOS / macOS / AndroidCPUCC-BY 4.0PermittedEntitlement requiredSign in
GoogleBlazeFace Front (Selfie Range)BetaFace detectionVisioniOS / macOS / AndroidCPUCC-BY 4.0PermittedEntitlement requiredSign in
Bryce Beattie / PiperCori High · British EnglishBetaSpeech synthesisAudioiOScpuMIT (voice repository); public-domain source audio; Apache-2.0 dictionaryPermittedEntitlement requiredSign in
Bryce Beattie / PiperCori · British EnglishBetaSpeech synthesisAudioiOScpuMIT (voice repository); public-domain source audio; Apache-2.0 dictionaryPermittedEntitlement requiredSign in
RikoroseDeepFilterNet3BetaSpeech enhancementAudioiOS / macOS / AndroidCPUMIT / Apache 2.0PermittedEntitlement requiredSign in

Measure model size, memory, load time, speed, energy, and quality.

CactusBrain is building a measurement layer for real device performance: model size, peak memory, load time, throughput, energy use, and quality tradeoffs.

ModelSize on disk
MemoryPeak working set
LoadCold and warm start
SpeedPrefill and decode
EnergyWork per result
QualityMeasured degradation

Plans meter model delivery while local inference stays unlimited.

Unlimited on-device inference. Run ten requests or ten billion. CactusBrain does not charge for tokens, prompts, inference requests, CPU time, GPU time, or monthly active users because that work happens on your users' devices.

Billing cycle
Paid plans include 25% delivery headroomExplorer stops at 50 GBNo surprise bills
Explorer
$0$0/mo/year

Build and validate your first on-device experience

Private models
1
Model storage
5 GB
Model delivery
50 GB / mo

≈ 640 downloads of an 80 MB model

  • Unlimited local inference
  • Public catalog access
  • Community support
Create account →
Launch
$49$490/mo/year

Operate multiple models across a growing fleet

Private models
20
Model storage
250 GB
Model delivery
1.2 TB / mo

≈ 15,000 downloads of an 80 MB model

  • A/B testing and staged rollouts
  • Priority support
Create account →
Scale
From $199From $199/mo/mo

Custom governance, capacity, and infrastructure

Private models
Unlimited
Model storage
500 GB
Model delivery
5 TB / mo

≈ 66,000 downloads of an 80 MB model

  • Audit logs and SSO
  • Bring your own bucket
  • SLA support
Talk to us →
Private modelsNew private-model registration stops at the plan limit.
StorageUsage is measured and reported; uploads are not blocked today.
DeliveryExplorer stops at 100%; paid plans pause at 125% until the next period.
Delivery estimator

Size a plan from your model and monthly downloads.

Estimated private delivery78.1 GB / month

Builder includes this delivery volume.

Plan comparison

Choose delivery capacity with unlimited local inference.

Every plan runs models locally with no token, request, compute, or active-user meter.

IncludedExplorer$0 / monthBuilder$15 / month$150 / yearLaunch$49 / month$490 / yearScaleFrom $199 / month
Limits
On-device inferenceUnlimitedUnlimitedUnlimitedUnlimited
Private models1520Unlimited
Storage5 GB50 GB250 GB500 GB+
Encrypted delivery / mo50 GB300 GB1.2 TB5 TB
Catalog downloadsUse delivery allowanceUse delivery allowanceUse delivery allowanceUse delivery allowance
Shipping
Delta updates
Version history & rollback
A/B testing & staged rollouts
BYO S3 / R2 bucket
Support & governance
SupportCommunityEmailPriorityDedicated Slack
SSO / SAML
SLA & audit logs

Zero metering for inference. cellm stays free and open source. Prompts, tokens, CPU, GPU, and active users are never billed. Delivery is capped rather than billed without limit: paid plans deliver up to 125% of their included allowance, then model delivery pauses until the next period instead of invoicing on, so there is no runaway bill. Explorer gets no headroom and stops at its included 50 GB. Plans are billed in USD by Lemon Squeezy, our merchant of record, which accepts cards and PayPal worldwide and handles sales tax and VAT at checkout.

Publish a signed model release and load it on-device.

  1. 01Create an account

    Access the developer workspace and available SDKs.

  2. 02Publish a release

    Attach a compatible, signed model artifact to your project.

  3. 03Ship the capability

    The SDK resolves, delivers, verifies, installs, and loads it on-device.

Open a developer workspace