NVIDIA NIM

Using NVIDIA NIM hosted models in Aimogen Pro, the available model catalogue, and how NIM fits alongside other providers.

NVIDIA NIM hosts a broad catalogue of open and NVIDIA-developed models behind an OpenAI-compatible API.

Get a key#

  1. Go to build.nvidia.com.
  2. Create an account and generate an API key.

NVIDIA offers a free credit allocation for evaluation.

Configure#

Aimogen Pro › Settings › API Keys › Nvidia AI API Keys

One key per line, Test API key, save.

Models#

Aimogen Pro 2.8.8 lists 32 NVIDIA models. Highlights:

ModelNotes
nvidia/nemotron-3-ultra-550b-a55bLargest Nemotron, added in 2.8.2
nvidia/nemotron-3-super-120b-a12bMid tier
nvidia/nemotron-3-nano-omni-30b-a3b-reasoningSmall reasoning model
deepseek-ai/deepseek-v4-pro, deepseek-ai/deepseek-v4-flashDeepSeek v4
qwen/qwen3.5-397b-a17b, qwen/qwen3.5-122b-a10b, qwen/qwen3-coder-480b-a35b-instructQwen 3.5 and a coding model
moonshotai/kimi-k2.6Long-context
mistralai/mistral-medium-3.5-128b, mistralai/mistral-small-4-119b-2603, mistralai/mistral-nemotronMistral
z-ai/glm-5.1, stepfun-ai/step-3.7-flash, bytedance/seed-oss-36b-instructOther open models
google/gemma-4-31b-it, google/gemma-3n-e2b-it, google/gemma-2-2b-itGemma
meta/llama-3.1-70b-instruct, meta/llama-3.1-8b-instruct, meta/llama-3.2-3b-instructLlama 3.x
nvidia/llama-3.1-nemotron-nano-vl-8b-v1, nvidia/nemotron-nano-12b-v2-vlVision-capable variants

When to use NVIDIA NIM#

Good for

  • Access to a wide open-model catalogue through one credential
  • Trying a specific open model without self-hosting it
  • Models not offered by the other configured providers

Less good for

  • Guaranteed availability. Catalogue entries change as NVIDIA rotates models
  • Feature depth. Assistants, Batch and fine-tuning are OpenAI-only in the plugin

Common problems#

A listed model returns 404 NVIDIA rotated it out of the catalogue. Pick another. The plugin list is compiled per release, so it can drift from the live catalogue between updates.

Slow first response Cold-start on a rarely used model. Subsequent calls are faster.

401 Key error, or the key was generated under a different NVIDIA account.

Endpoint used#

https://integrate.api.nvidia.com/v1
  • Ollama — run the same class of open models yourself
  • OpenRouter — another aggregator

Still stuck? Open a support ticket and include the diagnostics from Aimogen Pro › System & Logs › System Info.