Skip to main content
FA

Run and fine-tune open models with usage-based pricing

Fireworks AI is a model hosting and inference platform for teams building with open and proprietary models. It covers serverless inference, fine-tuning, embeddings, speech-to-text, and on-demand GPU deployments.

Agenticness = how independently a tool can take action, scored across 9 dimensions. Scored independently by David Kooi, Skylark Creations — see full rubric →

See pricing ↓

Paid
Enterprise
API
B2B
Usage-Based
Cloud Hosted
Visit Fireworks AI

Is this your tool? Claim this listing to manage your content and analytics.

Recent activity

What's happened with Fireworks AI lately

  • Score change
    Rubric upgrade v3_0 → v3.1: score 3/32 → 4/3634/36(+1)

    Rubric upgrade: agenticness v3.0 (8 dims, /32) → v3.1 (9 dims, /36). Adds Dim 9 (Operator Sovereignty), splits Dim 6 into 6a/6b lenses, tightens Dim 4 autonomous-retry distinction. Not a product change — score shift reflects new dimension + recalibrated rubric, not a change in the tool. Fanout suppressed.

    See the news that prompted this

News mentions sourced from our news feed; score changes from periodic re-evaluations.

Ask about Fireworks AI

Get answers based on Fireworks AI's actual documentation

Try asking:

About

What It Is

Fireworks AI is a cloud platform for developers and teams that need to serve, tune, and deploy AI models through an API. Based on the pricing and docs pages, it focuses on model inference infrastructure rather than a chat-style assistant or autonomous agent.

What to Know

The platform is strongest if you want managed model infrastructure with pay-as-you-go billing and multiple deployment modes. It is not a general-purpose agent that plans tasks or takes actions on your behalf; it is infrastructure for calling models...

Key Features
Serverless inference with per-token pricing
Fine-tuning for open models using your own data
On-demand GPU deployments billed per GPU second
Speech-to-text inference priced per audio second
Image generation support with multiple model families
Use Cases
Serving LLM and vision models through an API for application backends
Fine-tuning open models on proprietary data for domain-specific tasks
Deploying dedicated GPU endpoints for higher throughput workloads
Agenticness: Reactive Tool

Responds to prompts but takes no autonomous action.

High evidence
Last evaluated: May 23, 2026

Dimension Breakdown

Action Capability
Autonomy
Adaptation
State & Memory
Safety

Categories

Pricing
  • Free: $1 in free credits for serverless inference onboarding.
  • Usage-based: Serverless inference is billed per token; STT is billed per audio second; image generation, embeddings, fine-tuning, and on-demand deployments each have published usage-based rates.
  • Enterprise: Contact sales for enterprise deployments, faster speeds, lower costs, and higher rate limits.
Details
AddedApril 1, 2026
RefreshedApril 1, 2026
Agenticness
Quick Facts
DeploymentCloud-hosted
AutonomyCopilot (human-in-loop)
Model supportMulti-model
Open sourceNo
Team supportEnterprise
Pricing modelUsage-based
Interfaceweb, api, cli
Sources

Lineation is an agent security platform that monitors, governs, and defends AI workflows across tools and providers. It meters pricing by protected agent actions, with plans for prototypes, teams, and enterprise deployments.

Enterprise
Chrome Extension
B2B
+4

A small open-source collection of production patterns for running LLM agents more safely. It focuses on redacting secret-shaped text, checking agent memory integrity, and testing prompt or skill changes before they regress.

Open Source
Chrome Extension
File Access
+4

A Rust-based memory layer for AI agents that stores, recalls, and consolidates memory locally with SQLite/FTS. It is built for developers who want explainable recall, explicit forgetting, and portable agent memory without a required cloud service.

File Access
Memory
B2B
+4

ThumbGate sits in front of coding-agent tool calls and blocks risky repeats before they run. It turns thumbs-up and thumbs-down feedback into reusable rules, with local use for individuals and paid options for dashboards, sync, and team visibility.

Chrome Extension
Code Execution
Integrations
+4
Stay in the loop

Get the weekly agentic AI briefing

New tools, top picks, and trends — delivered every Thursday.

I use AI for: