Pay less for AI. Earn from your GPU.

Managed production inference for developers. Connect your GPU and earn the majority of token revenue on inference-only catalog jobs.

  • API OpenAI-compatible endpoints, plug and play with our API tokens
  • Pricing Per-token catalog pricing with prepaid credits
  • Compute Managed routing across a distributed GPU network with vetting and security tiers

Open models on the Scalattice network, tailored for personal, production, and development inference.

  • Qwen3 8B / 14B / 32B
  • Llama 3.3 70B
  • DeepSeek R1 7B
  • Qwen2.5 Coder 7B
7 Curated open models
$0 Platform fee to connect a machine
128K Max context on select models
Live Per-token usage and payouts

Developers

Ship on published per-token pricing.

OpenAI-compatible API with built-in vetting, security tiers, and regional policy.

API

Ship on day one.

Keys, docs, and billing in one place. Prototype to production without rewrites.

Why Scalattice
Setup

Plug and play.

Set base URL and key. Optional vetting, security tiers, and region headers on every request.

How to connect
# OpenAI-compatible: swap base URL + key
from openai import OpenAI

client = OpenAI(
    base_url="https://api.scalattice.cloud/v1",
    api_key="slt_…",
)

client.chat.completions.create(
    model="qwen-3-8b",
    messages=[{"role": "user", "content": "Ship it."}],
)
Three lines to your first request
Models on the network Qwen Llama DeepSeek Gemma

Explore developer docs

Providers

Idle GPU time is money on the table.

Connect the hardware you already own. Scalattice routes catalog inference jobs when your schedule allows. Majority token revenue share, $0 connection fee.

Hardware

Any GPU welcome.

Gaming rig, home lab, or small rack. You set when each machine is available.

Income

Jobs in. Payouts out.

Scalattice routes inference and handles billing. Curated models and demand ratings tell you what to host.

Become a provider Open-source agent
Gaming PC and workstation GPUs ready to serve inference
Consumer and datacenter cards welcome
Uptime Offer nights, weekends, or 24/7 on your schedule.
Payouts Track earnings live on Scalattice Cloud.

Network

Capacity where your users already are.

Route inference across regions from one API. Lower latency for users worldwide.

Europe

EU pools

Keep workloads inside approved jurisdictions.

Americas

US pools

Serve North America from nearby hosts.

Asia Pacific

APAC edge

Cut round-trip time for real-time apps.

Resilience

Capacity fallback

Backup capacity when provider GPUs are unavailable.

Global network of distributed inference capacity
Distributed GPU network across regions

Explore the network

Pricing

Transparent rates per-token.

Per-million-token pricing published upfront. Usage dashboards on Scalattice Cloud update as you use the API.

7 Curated open models
$0 Platform fee to connect a GPU
Live Usage and payout tracking

See full pricing · View offers

# Per-million tokens (live catalog)
qwen-2.5-coder-7b   in $0.029  out $0.077
deepseek-r1-7b      in $0.040  out $0.100
gemma-3-27b         in $0.065  out $0.129
qwen-3-32b          in $0.066  out $0.242
qwen-3-1.7b         in $0.083  out $0.083
qwen-3-8b           in $0.086  out $0.173
qwen-3-14b          in $0.091  out $0.192
llama-3.3-70b       in $0.098  out $0.306

# Full catalog → Scalattice Cloud
Live rates · open Scalattice Cloud for the full catalog

Trust

Keep data close. Run models nearby.

Set region policy on each request so sensitive data stays where your team needs it.

Inference-only jobs

Every inference job is metered and logged. Operators run models locally inside the agent.

Global capacity

Serve users worldwide from regional inference capacity.

Your policies

Set region, vetting, and security tier per request. Scalattice routes automatically.

Compliance Geo-fence requests to approved regions and security tiers.
Audit Per-token usage tracked in your Scalattice Cloud dashboard.