NVIDIA NIM
Using NVIDIA NIM hosted models in Aimogen Pro, the available model catalogue, and how NIM fits alongside other providers.
NVIDIA NIM hosts a broad catalogue of open and NVIDIA-developed models behind an OpenAI-compatible API.
Get a key#
- Go to build.nvidia.com.
- Create an account and generate an API key.
NVIDIA offers a free credit allocation for evaluation.
Configure#
Aimogen Pro › Settings › API Keys › Nvidia AI API Keys
One key per line, Test API key, save.
Models#
Aimogen Pro 2.8.8 lists 32 NVIDIA models. Highlights:
| Model | Notes |
|---|---|
nvidia/nemotron-3-ultra-550b-a55b | Largest Nemotron, added in 2.8.2 |
nvidia/nemotron-3-super-120b-a12b | Mid tier |
nvidia/nemotron-3-nano-omni-30b-a3b-reasoning | Small reasoning model |
deepseek-ai/deepseek-v4-pro, deepseek-ai/deepseek-v4-flash | DeepSeek v4 |
qwen/qwen3.5-397b-a17b, qwen/qwen3.5-122b-a10b, qwen/qwen3-coder-480b-a35b-instruct | Qwen 3.5 and a coding model |
moonshotai/kimi-k2.6 | Long-context |
mistralai/mistral-medium-3.5-128b, mistralai/mistral-small-4-119b-2603, mistralai/mistral-nemotron | Mistral |
z-ai/glm-5.1, stepfun-ai/step-3.7-flash, bytedance/seed-oss-36b-instruct | Other open models |
google/gemma-4-31b-it, google/gemma-3n-e2b-it, google/gemma-2-2b-it | Gemma |
meta/llama-3.1-70b-instruct, meta/llama-3.1-8b-instruct, meta/llama-3.2-3b-instruct | Llama 3.x |
nvidia/llama-3.1-nemotron-nano-vl-8b-v1, nvidia/nemotron-nano-12b-v2-vl | Vision-capable variants |
When to use NVIDIA NIM#
Good for
- Access to a wide open-model catalogue through one credential
- Trying a specific open model without self-hosting it
- Models not offered by the other configured providers
Less good for
- Guaranteed availability. Catalogue entries change as NVIDIA rotates models
- Feature depth. Assistants, Batch and fine-tuning are OpenAI-only in the plugin
Common problems#
A listed model returns 404 NVIDIA rotated it out of the catalogue. Pick another. The plugin list is compiled per release, so it can drift from the live catalogue between updates.
Slow first response Cold-start on a rarely used model. Subsequent calls are faster.
401 Key error, or the key was generated under a different NVIDIA account.
Endpoint used#
https://integrate.api.nvidia.com/v1Related#
- Ollama — run the same class of open models yourself
- OpenRouter — another aggregator
Still stuck? Open a support ticket and include the diagnostics from Aimogen Pro › System & Logs › System Info.