Early access: Claude and Gemini available now
One API for every model. The cost of every call.
Feldspar is an AI deployment platform. Send the same request to any model and change one parameter to switch. Every response reports its tokens, latency, and cost.
curl $FELDSPAR_URL/api/v1/chat/completions \
-H "Authorization: Bearer $FELDSPAR_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic/claude-sonnet-5",
"messages": [
{ "role": "user", "content": "Summarize this ticket." }
]
}'input_tokensoutput_tokenscost_usdlatency_msA workspace for the models you already pay for
Bring a Claude or Gemini key. You get chat, multi-agent workflows, and a record of every call, all through the same gateway.
- https://feldspar.gneisslabs.com/app/
Chat
Switch models inside one conversation. Each reply records the tokens sent, the tokens received, and the estimated cost.
Open chat → - https://feldspar.gneisslabs.com/app/workflows/
Agent workflows
Connect agents into a graph and run it one step at a time. Branch from any checkpoint. Earlier steps are copied, not called again.
Build a workflow → - https://feldspar.gneisslabs.com/app/metrics/
Metrics
Median and 95th-percentile latency, failures, tokens, and cost across your recent calls. Open any call to see the exact prompt and result.
View metrics → - https://feldspar.gneisslabs.com/app/account/
Your own keys
Connect a Claude or Gemini key. It is encrypted before storage and never displayed again. You see the last four characters. Your provider bills you directly.
Connect a key →
Same request, different model
The request shape does not change between providers. Switch the model field. Nothing else in your code changes.
| Model ID | Provider | Status |
|---|---|---|
anthropic/claude-sonnet-5 | Anthropic | Available |
google/gemini-3.5-flash-lite | Available | |
meta-llama/Llama-3.3-70B-Instruct | Open source | Phase 1 |
Three ways to run an open model
You choose the trade-off between cold starts and idle cost. Feldspar rents you a dedicated GPU only when you ask for one.
| Mode | Behavior | You pay |
|---|---|---|
| Shared | Popular models stay loaded on shared GPUs. No cold start. | Per token |
| Serverless dedicated | The model loads on demand and scales to zero. Cold starts occur. | Per GPU second |
| Always-on dedicated | The GPU stays on. For steady volume or privacy requirements. | GPU cost plus markup |
LoRA adapters run against a shared base model, so many adapters share one GPU. Rates are not published yet.
What happens to a request
Plugins run inside the request pipeline, not next to it. You can see, time, and turn off each step.
- GatewayAuth, rate limit, balance
- Input guardrailsPlugin
- RAGPlugin
- ModelOpen or closed
- ToolsPlugin
- Output guardrailsPlugin
Third-party plugins run in a sandbox and declare their permissions: network access, data access, and cost per call. The dashboard shows the latency and cost each one adds.
Attach it to the deployment, not your app
Most teams rebuild the same parts: a PII filter, a retrieval connector, an eval. Feldspar attaches them to a deployment so they stay out of your application code.
Guardrails
PII filter, prompt-injection filter, content safety, topic limits.
Tools
Web search, database query, code execution, MCP servers.
RAG connectors
Google Drive, Notion, S3, websites, SQL.
Evaluators
Hallucination check, groundedness score, regression tests.
Harness templates
A prompt, tools, guardrails, and RAG, preassembled into one setup.
First-party plugins ship before the marketplace opens to other publishers. Tools use MCP, not a proprietary interface.
Three tiers, metered per unit
Inference is metered per token, GPU time per second, plugins per call, and RAG per GB of storage and embeddings.
BYOK
Connect your own provider key and pay the provider directly. Feldspar charges a flat monthly fee for the software layer.
Flat monthly feeManaged
Buy prepaid credits. Feldspar supplies model access and meters what you use.
Prepaid creditsEnterprise
Invoiced billing, a custom contract, and private deployment.
Invoiced
Rates are not set yet. We will publish them when they are final, not before.
What exists, and what does not
This is the full roadmap, including the parts that are not built yet.
- Phase 1Building now
Done. One API for Claude and Gemini. Keys you bring, encrypted at rest. Chat, agent workflows, and per-call metrics.
Remaining. 5 to 10 open-source models on serverless GPUs. First-party guardrails and tools. Prepaid credits.
- Phase 2Next
Managed RAG as plugins. Shared GPUs for the most-used models. Custom Hugging Face model IDs. Spend caps.
- Phase 3Planned
Open marketplace: plugin SDK, review process, publisher payouts. Harness templates. Fine-tuning presets.
- Phase 4Planned
Invoiced billing. SOC 2. Private deployment. Team roles.
Try it with the key you already have
Connect a Claude or Gemini key and start a chat. Your provider bills you directly. Saving a key makes no model call.