The AWS Bedrock Pricing and Guide

AWS Bedrock is Amazon Web Services’ serverless AI solution, providing developers with access to cutting-edge foundation models (FMs) from leading suppliers like Amazon Titan, Anthropic, Mistral AI, Meta AI, Stability AI, and others.

Building complex generative AI applications is easier with AWS Bedrock, a fully managed, serverless AI solution. Bedrock offers high-performance foundation models (FMs) from Amazon Titan, Anthropic, Mistral, and Stability AI vendors.

Its serverless design expands automatically, allowing developers to concentrate on application development while AWS manages the infrastructure. However, efficient cost management necessitates comprehension of Bedrock’s pricing mechanisms as well as utilization optimization.

What is AWS Bedrock?

AWS Bedrock is Amazon Web Services’ serverless AI solution, which provides developers with access to cutting-edge foundation models (FMs) from leading suppliers like as Amazon Titan, Anthropic, Mistral AI, Meta AI, Stability AI, and others. Bedrock is designed to make machine learning integration easier for businesses by allowing users to use pre-trained models in a variety of applications such as natural language processing (NLP), image generation, virtual assistants, and data analysis without having to build or maintain complex AI infrastructure.

Bedrock offers a flexible, plug-and-play approach to machine learning. Developers may choose from a variety of pre-trained foundation models, each suited for a certain job, such as text summarization, sentiment analysis, or content development. Bedrock also supports model fine-tuning on private data, allowing organizations to customize the model’s replies or outputs to meet unique industry standards or brand tone.

AWS Bedrock stands out for its serverless design, which dynamically scales resources on demand and charges customers just for model consumption. This robustness is especially crucial for workloads with changing AI needs, as it allows for maximum performance during peak periods without the effort or expense of manually scaling resources. Bedrock also works smoothly with other AWS services, such as S3 for storage, Lambda for event-based functionality, and Sagemaker, to improve its end-to-end machine learning operations. The capability of its end-to-end machine learning operations.

Microsoft Azure OpenAI vs AWS Bedrock

In contrast, Azure OpenAI Service focuses on OpenAI-developed models like as GPT, Codex, and DALL-E. It integrates deeply with Microsoft’s Azure environment, allowing businesses to safely run OpenAI models while accessing Azure’s security, compliance, and identity services. Azure OpenAI is especially useful for organizations dedicated to the Microsoft ecosystem since it offers a closely integrated experience with Azure products such as Azure Cognitive Services, allowing for easy data integration and administration.

In summary, AWS Bedrock provides a broader selection of foundation models as well as a serverless infrastructure that is ideal for multi-model use, whereas Azure OpenAI provides specialized access to OpenAI models with strong integration into Azure’s ecosystem, making it ideal for enterprises that have invested in Microsoft’s cloud and compliance offerings.

AWS Bedrock Pricing Models

Feature

AWS Bedrock On-Demand

AWS Bedrock Provisioned Throughput

AWS Bedrock Batch

Use Case

Variable or unpredictable workloads

Large, consistent workloads needing guaranteed performance

Large-scale, one-time or periodic batch processing

Pricing

Pay-per-use, charged per token processed

Hourly rate based on reserved model units

Discounted rate per token for supported models

Throughput

Scales dynamically with demand

Guaranteed throughput with dedicated model units

Processes multiple prompts in a single input file for efficient high-volume processing

Custom Model Access

Not available

Required for using custom models

Limited; custom models may not be supported

Commitment Terms

None

1-month or 6-month commitments

None

Ideal For

Ad hoc analysis, exploratory or variable workloads

Production applications with predictable, high-demand usage

Bulk data processing, cost-efficient handling of large datasets

Bedrock’s price is usage-based and may change depending on workload. Here’s an overview of the main price options:

Upon demand pricing

With AWS Bedrock’s On-Demand pricing model, you are only billed for your actual consumption, with no upfront or time-based obligations. This model costs per token or each image created, depending on the sort of AI model being used.

  • Text-Generation Models: For text-based models, costs are determined by the number of tokens processed. Charges are applied to both input tokens (what the model processes from user input) and output tokens (what the model produces in response). A token, which represents a short portion of text such as a few letters or a word fragment, is the fundamental unit of cost. This granular approach allows expenses to closely match consumption, making it an adaptable option for varying workloads.
  • Embedding Models: Pricing depends only on the number of input tokens. This implies that the expenses are restricted to the text data you need to generate vector embeddings, which are helpful in search and recommendation engines.
  • picture-Generation Models: Charges are applied to each picture created. This uncomplicated pricing technique allows you to easily estimate expenses depending on the quantity of photographs required.

Cross-Region Inference: On-demand pricing also offers cross-region inference for chosen models, allowing you to manage traffic spikes and increase performance by distributing processing across AWS regions. This enables robust, scalable operations across several areas at no additional expense, as charges are still based on the pricing rates of the originating (source) region.

This approach offers great flexibility because you are only invoiced for resources utilized rather than committing to certain use levels or time limits, making it perfect for unexpected or dynamic workloads.

On-demand example (text generating)

Assume you’re running AWS Bedrock to power a product suggestion chatbot during peak shopping hours. Here’s the breakdown of prospective costs:

  • Token Usage: Each user inquiry requires 50 input tokens, with the result generating around 200 output tokens.
  • Rates: Input tokens cost $0.001 per 1,000, while output tokens cost $0.003 per 1,000.
  • Usage Volume: The chatbot processes 2,000 inquiries each day, resulting in 100,000 input tokens and 400,000 output tokens.

Daily Cost Calculation:

  • Input cost is $0.10 per day, calculated as 100,000 tokens / 1,000 x $0.001.
  • The output cost is $1.20 per day, calculated as 400,000 tokens divided by 1,000 and $0.003.
  • $0.10 + $1.20 = $1.30 every day.
  • Monthly cost estimate: $1.30 multiplied by 30 is $39.

This on-demand strategy provides flexibility by charging just for the tokens processed. This makes it suited for dynamic workloads.

Provisioned throughput pricing

AWS Bedrock’s Provisioned Throughput pricing model provides committed capacity for enterprises with consistent, large-scale inference workloads that demand assured performance. In this mode, businesses buy model units for a given foundation or bespoke model, each with a set throughput measured in tokens per minute. Tokens are units of text that are processed, such as input or output tokens, therefore having more tokens allows you to handle larger or more complicated data inputs faster.

This architecture is especially beneficial for production-grade AI systems with known usage patterns since it allows businesses to reserve throughput capacity, ensuring that models can manage consistent workloads without interruption. This reserved capacity helps enterprises that require low latency and dependable performance for tasks such as customer assistance, content development, and real-time data processing.

1. Guaranteed Throughput: Model units provide a fixed rate of throughput, allowing businesses to achieve service-level agreements for applications that require continuous performance.

2. Custom Model Access: Custom-trained models are only available through Provisioned Throughput, making it critical for enterprises that demand specialized or fine-tuned models.

3. Hourly Pricing with Flexible durations: Organizations are paid hourly based on the number of model units reserved, and they may pick between 1-month and 6-month commitment durations based on their expected usage and budget. Longer commitment lengths are often more cost-effective.

Overall, AWS Bedrock’s Provisioned Throughput model provides a solid option for businesses with big, predictable AI workloads. It combines predictable cost and assured performance, making it ideal for corporate applications that require reliable, scalable AI processing.

Provisioned throughput pricing example

Let us walk through an example calculation for AWS Bedrock Provisioned Throughput price.

Assume a firm wants constant, high-throughput access to a model and wishes to set aside 5 model units for a bespoke model. Each model unit generates a specific amount of tokens each minute, and pricing is based on an hourly rate per model unit.

Let us assume:

  • Each model unit costs $0.50 per hour.
  • The company decides on a one-month commitment (about 730 hours per month).

1. Hourly Cost for Model Units:

5 model units × $0.50/hour = $2.50 per hour

2. Monthly Cost:

$2.50/hour × 730 hours = $1,825 per month

Batch Processing

Batch processing allows you to submit several prompts in a single input file, making it an effective approach to handling huge datasets all at once. The price strategy for batch processing is identical to on-demand, but with one significant advantage: you may frequently gain a 50% discount on supported foundation models (FMs) as compared to the on-demand cost.

This reduction is an excellent method to save money but bear in mind that batch processing is not available on all models. The official documentation includes a list of supported models.

Model Evaluation

Model evaluation pricing allows you to pay based on token consumption while evaluating multiple models without large-scale usage, making it excellent for early-stage projects or model performance comparisons.

AWS Bedrock Pricing is based on the model chosen and suppliers.

Amazon Bedrock offers a selected set of foundation models (FMs) from leading AI providers such as Anthropic, Meta, and Mistral AI. The pricing varies according to the AI provider and the area.

Let’s look at how much Anthropic, Meta, and Mistral AI cost in North Virginia. If you want to see more rates from different AI providers, go here.

Anthropic

Anthropic models prioritize safe and aligned AI, stressing responsible language processing and strong safeguards for user involvement.

Anthropic models

Price per 1,000 input tokens

Price per 1,000 output tokens

Price per 1,000 input tokens (batch)

Price per 1,000 output tokens (batch)

Claude 3.5 Sonnet**

$0.003

$0.015

$0.0015

$0.0075

Claude 3.5 Haiku

$0.001

$0.005

$0.0005

$0.0025

Claude 3 Opus*

$0.015

$0.075

$0.0075

$0.0375

Claude 3 Haiku

$0.00025

$0.00125

$0.000125

$0.000625

Claude 3 Sonnet

$0.003

$0.015

$0.0015

$0.0075

Claude 2.1

$0.008

$0.024

N/A

N/A

Claude 2.0

$0.008

$0.024

N/A

N/A

Claude Instant

$0.0008

$0.0024

N/A

N/A

Provisioned Throughput Pricing – Regions: US East (North Virginia) and US West (Oregon)

Anthropic models

Price per hour per model with
no commitment

Price per hour per model unit for 1-month commitment

Price per hour per model unit for 6-month commitment

Claude Instant

$44.00

$39.60

$22.00

Claude 2.0/2.1

$70.00

$63.00

$35.00

Meta Llama

Meta LLaMA is Meta’s high-performance language model, built for large-scale natural language interpretation and generating workloads.

Llama 3.2

On-Demand and Batch Pricing – Region: US East (N. Virginia)

Meta models

Price per 1,000 input tokens

Price per 1,000 output tokens

Price per 1,000 input tokens (batch)

Price per 1,000 output tokens (batch)

Llama 3.2 Instruct (1B)

$0.0001

$0.0001

N/A

N/A

Llama 3.2 Instruct (3B)

$0.00015

$0.00015

N/A

N/A

Llama 3.2 Instruct (11B)

$0.00035

$0.00035

N/A

N/A

Llama 3.2 Instruct (90B)

$0.002

$0.002

N/A

N/A

Pricing for provisioned throughput in the US West (Oregon) region

AI Mistral

Modern language models tailored for a range of NLP activities are available from Mistral AI, which provides excellent accuracy and efficiency in data processing and understanding.

US East (N. Virginia) is the region for on-demand and batch pricing.

Mistral models

Price per 1,000 input tokens

Price per 1,000 output tokens

Price per 1,000 input tokens (batch)

Price per 1,000 output tokens (batch)

Mistral 7B

$0.00015

$0.0002

N/A

N/A

Mixtral 8*7B

$0.00045

$0.0007

N/A

N/A

Mistral Small (24.02)

$0.001

$0.003

$0.0005

$0.0015

Mistral Large (24.02)

$0.004

$0.012

N/A

N/A