About

Together AI is a cloud platform for running open-source models through serverless inference, batch inference, dedicated inference, and dedicated container inference. It also includes accelerated compute with GPU clusters, sandbox environments, managed storage, and fine-tuning and evaluation tools for model shaping.

The pricing page shows usage-based pricing rather than a free plan: serverless inference is priced per 1M tokens, dedicated inference starts at $3.99 per hour for 1x H100 80GB, GPU clusters start at $3.49 per hour, sandbox compute is billed per vCPU and GiB RAM, managed storage is $0.16 per GiB/month, and fine-tuning is priced per 1M tokens with model-specific minimum charges.

  • Serverless inference APIs
  • Batch inference workloads
  • Dedicated GPU endpoints
  • GPU clusters on demand
  • Sandbox development environments
  • Managed storage for model data
  • Fine-tuning and evaluations

Free Tier Value

0
FTV score
Credit cardNot stated
Feature parity0%

This pricing page does not show any concrete free tier, free credit, or trial, so there is no defensible monthly value to assign. The cheapest visible paid offering is usage-based serverless inference, but without a stated free quota there is nothing to compare against, and a credit card requirement is not explicitly waived.

What's included in the free tier

Verified

Details aren't itemized for this entry yet - check the pricing page below for the latest.

Paid plans

Serverless Inference

Usage-based
Price per 1M tokens
input tokens
per 1M tokens
output tokens
per 1M tokens
  • Chat
  • Vision
  • Image
  • Audio
  • Video
  • Transcribe
  • Embeddings
  • Rerank

GPU Clusters

Contact sales
  • Accelerated compute
  • Reliable GPU clusters at scale
  • GB300
  • GB200
  • B200
  • H200
  • H100

Sandbox

Contact sales
  • Developer environments
  • Build development environments for AI

Managed Storage

Contact sales
  • Storage
  • Store model weights & data securely

Fine-Tuning

Contact sales
  • Shape models with your data
  • Model shaping

Provisioned Throughput

Contact sales
Token-based capacity with SLAs
tokens
capacity-based
  • Token-based capacity
  • SLAs
  • Dedicated capacity

Pricing extracted from Together AI's pricing page. Always verify current pricing before committing.