NVIDIA: Llama 3.1 Nemotron Ultra 253B v1

Nvidia

Llama-3.1-Nemotron-Ultra-253B-v1 is a large language model (LLM) optimized for advanced reasoning, human-interactive chat, retrieval-augmented generation (RAG), and tool-calling tasks. Derived from Meta’s Llama-3.1-405B-Instruct, it has been significantly customized using Neural Architecture Search (NAS), resulting in enhanced efficiency, reduced memory usage, and improved inference latency. The model supports a context length of up to 128K tokens and can operate efficiently on an 8x NVIDIA H100 node. Note: you must include `detailed thinking on` in the system prompt to enable reasoning. Please see [Usage Recommendations](https://huggingface.co/nvidia/Llama-3_1-Nemotron-Ultra-253B-v1#quick-start-and-usage-recommendations) for more.

料金情報

入力料金$0.600/100万トークン

出力料金$1.80/100万トークン

入力料金$0.600

出力料金$1.80

スペック

コンテキスト長: 131,072 トークン
最大出力トークン: -
モダリティ: text→text
トークナイザー: Llama3
モデレーション: なし

スペック

✓Text Generation

✗Image Input

✗Audio Input

✗Image Output

✓Function Calling

✓Streaming

Quick Comparison

Speed70%

Cost Efficiency98%

Context Size7%

このモデルを使ってみる (OpenRouter)

APIコードサンプル

from openai import OpenAI

client = OpenAI(
    base_url="https://openrouter.ai/api/v1",
    api_key="YOUR_API_KEY"
)

response = client.chat.completions.create(
    model="nvidia/llama-3.1-nemotron-ultra-253b-v1",
    messages=[
        {"role": "user", "content": "Hello!"}
    ]
)

print(response.choices[0].message.content)

料金アラート:

ユーザーレビュー

このページは役に立ちましたか？

Quick Compare

類似モデル

NVIDIA: Llama 3.1 Nemotron Ultra 253B v1Try on OpenRouter

NVIDIA: Llama 3.1 Nemotron Ultra 253B v1

料金情報

スペック

スペック

Quick Comparison

APIコードサンプル

ユーザーレビュー

レビューを書く

Quick Compare

類似モデル