← All Models

NVIDIA: Llama 3.1 Nemotron 70B Instruct

🇺🇸 NVIDIA · Llama 3.1

Input Price $1.20 per million tokens NT$38.4
Output Price $1.20 per million tokens NT$38.4
Context Window 131K tokens Output limit: 16K
OpenRouter Route Price Please verify with official pricing pages
Use this model via OpenRouter →

Overview

NVIDIA: Llama 3.1 Nemotron 70B Instruct is a large language model API from NVIDIA, part of its Llama 3.1 model family. Priced at $1.20 per million input tokens and $1.20 per million output tokens, it occupies the mid-range, balancing capability against running cost. The 131K-token context window — around 197 pages of text — comfortably handles long documents, multi-file code, or extended conversations. On Artificial Analysis's Intelligence Index it scores 7 (F grade), a useful proxy for its general reasoning strength relative to the other models tracked here. All prices on this page reflect OpenRouter's routed rates and are re-synced automatically every day; confirm against the provider's official pricing before committing to production.

Dimension Unit Price (USD) Price (TWD) Effective From
Input per 1M tokens $1.20 NT$38.4
Output per 1M tokens $1.20 NT$38.4

Provider
NVIDIA
Model Family
Llama 3.1
Version String
nvidia/llama-3.1-nemotron-70b-instruct
Status
Active
Modality
Text
Context Window
131,072 tokens
Output Limit
16,384 tokens

Index Metrics

Cross-domain capability indexes evaluated by Artificial Analysis — Artificial Analysis

Intelligence Index 7 F Measured: 2026-08-26

Benchmark Scores

Data source: Artificial Analysis

AA-LCR 7.3% F Measured: 2026-08-26
GPQA Diamond 46.5% C Measured: 2026-08-26
HLE 4.2% D Measured: 2026-08-26
IFBench 30.8% D Measured: 2026-08-26
Non-Hallucination 28.7% Measured: 2026-08-26
Omniscience Accuracy 17.8% Measured: 2026-08-26
SciCode 23.3% C Measured: 2026-08-26
Tau2 23.1% Measured: 2026-08-26
TerminalBench 4.5% Measured: 2026-08-26

Performance Metrics

Real-world benchmarks, updated every 72 hours by Artificial Analysis — Artificial Analysis

First Token Latency 7.0s Measured: 2026-08-26
Output Speed 98 t/s Measured: 2026-08-26
Response Time 12.1s Measured: 2026-08-26

Key Insights

Key data points from this page for quick reference and citation.

  • NVIDIA: Llama 3.1 Nemotron 70B Instruct Input price: $1.2/M tokens
  • NVIDIA: Llama 3.1 Nemotron 70B Instruct Output price: $1.2/M tokens
  • Context window: 131,072 tokens
  • Provider: NVIDIA
  • Model family: Llama 3.1
  • Modalities: Text
  • Data source: OpenRouter, updated daily