FLUX 3 Action
VLMBlack Forest Labs' world action model for robotics, with a guide to generating action predictions on Jetson Thor using LeRobot.
Run the latest generative AI models on your Jetson device
Latest releases with day-0 support on Jetson
Browse by family cards or switch to a sortable table.
Quick Start Runner
Loading command... Commands are auto-generated based on your configuration settings.
This model requires a Hugging Face access token. The token is inserted into the command and never stored.
Configure vLLM server parameters. Leave empty to use defaults.
Maximum context length the model can handle
Fraction of GPU memory to use (0.1 - 1.0)
Quick Start Runner
Loading command... Commands are auto-generated based on your configuration settings.
This model requires a Hugging Face access token. The token is inserted into the command and never stored.
Configure vLLM server parameters. Leave empty to use defaults.
Maximum context length the model can handle
Fraction of GPU memory to use (0.1 - 1.0)
Quick Start Runner
Loading command... Commands are auto-generated based on your configuration settings.
This model requires a Hugging Face access token. The token is inserted into the command and never stored.
Configure vLLM server parameters. Leave empty to use defaults.
Maximum context length the model can handle
Fraction of GPU memory to use (0.1 - 1.0)
Quick Start Runner
Loading command... Commands are auto-generated based on your configuration settings.
This model requires a Hugging Face access token. The token is inserted into the command and never stored.
Configure vLLM server parameters. Leave empty to use defaults.
Maximum context length the model can handle
Fraction of GPU memory to use (0.1 - 1.0)
Quick Start Runner
Loading command... Commands are auto-generated based on your configuration settings.
This model requires a Hugging Face access token. The token is inserted into the command and never stored.
Configure vLLM server parameters. Leave empty to use defaults.
Maximum context length the model can handle
Fraction of GPU memory to use (0.1 - 1.0)
Quick Start Runner
Loading command... Commands are auto-generated based on your configuration settings.
This model requires a Hugging Face access token. The token is inserted into the command and never stored.
Configure vLLM server parameters. Leave empty to use defaults.
Maximum context length the model can handle
Fraction of GPU memory to use (0.1 - 1.0)
Quick Start Runner
Loading command... Commands are auto-generated based on your configuration settings.
This model requires a Hugging Face access token. The token is inserted into the command and never stored.
Configure vLLM server parameters. Leave empty to use defaults.
Maximum context length the model can handle
Fraction of GPU memory to use (0.1 - 1.0)
Quick Start Runner
Loading command... Commands are auto-generated based on your configuration settings.
This model requires a Hugging Face access token. The token is inserted into the command and never stored.
Configure vLLM server parameters. Leave empty to use defaults.
Maximum context length the model can handle
Fraction of GPU memory to use (0.1 - 1.0)
Quick Start Runner
Loading command... Commands are auto-generated based on your configuration settings.
This model requires a Hugging Face access token. The token is inserted into the command and never stored.
Configure vLLM server parameters. Leave empty to use defaults.
Maximum context length the model can handle
Fraction of GPU memory to use (0.1 - 1.0)
Quick Start Runner
Loading command... Commands are auto-generated based on your configuration settings.
This model requires a Hugging Face access token. The token is inserted into the command and never stored.
Configure vLLM server parameters. Leave empty to use defaults.
Maximum context length the model can handle
Fraction of GPU memory to use (0.1 - 1.0)
Quick Start Runner
Loading command... Commands are auto-generated based on your configuration settings.
This model requires a Hugging Face access token. The token is inserted into the command and never stored.
Configure vLLM server parameters. Leave empty to use defaults.
Maximum context length the model can handle
Fraction of GPU memory to use (0.1 - 1.0)
Quick Start Runner
Loading command... Commands are auto-generated based on your configuration settings.
This model requires a Hugging Face access token. The token is inserted into the command and never stored.
Configure vLLM server parameters. Leave empty to use defaults.
Maximum context length the model can handle
Fraction of GPU memory to use (0.1 - 1.0)
Quick Start Runner
Loading command... Commands are auto-generated based on your configuration settings.
This model requires a Hugging Face access token. The token is inserted into the command and never stored.
Configure vLLM server parameters. Leave empty to use defaults.
Maximum context length the model can handle
Fraction of GPU memory to use (0.1 - 1.0)
Quick Start Runner
Loading command... Commands are auto-generated based on your configuration settings.
This model requires a Hugging Face access token. The token is inserted into the command and never stored.
Configure vLLM server parameters. Leave empty to use defaults.
Maximum context length the model can handle
Fraction of GPU memory to use (0.1 - 1.0)
Quick Start Runner
Loading command... Commands are auto-generated based on your configuration settings.
This model requires a Hugging Face access token. The token is inserted into the command and never stored.
Configure vLLM server parameters. Leave empty to use defaults.
Maximum context length the model can handle
Fraction of GPU memory to use (0.1 - 1.0)
Quick Start Runner
Loading command... Commands are auto-generated based on your configuration settings.
This model requires a Hugging Face access token. The token is inserted into the command and never stored.
Configure vLLM server parameters. Leave empty to use defaults.
Maximum context length the model can handle
Fraction of GPU memory to use (0.1 - 1.0)
Quick Start Runner
Loading command... Commands are auto-generated based on your configuration settings.
This model requires a Hugging Face access token. The token is inserted into the command and never stored.
Configure vLLM server parameters. Leave empty to use defaults.
Maximum context length the model can handle
Fraction of GPU memory to use (0.1 - 1.0)
Quick Start Runner
Loading command... Commands are auto-generated based on your configuration settings.
This model requires a Hugging Face access token. The token is inserted into the command and never stored.
Configure vLLM server parameters. Leave empty to use defaults.
Maximum context length the model can handle
Fraction of GPU memory to use (0.1 - 1.0)
Quick Start Runner
Loading command... Commands are auto-generated based on your configuration settings.
This model requires a Hugging Face access token. The token is inserted into the command and never stored.
Configure vLLM server parameters. Leave empty to use defaults.
Maximum context length the model can handle
Fraction of GPU memory to use (0.1 - 1.0)
Quick Start Runner
Loading command... Commands are auto-generated based on your configuration settings.
This model requires a Hugging Face access token. The token is inserted into the command and never stored.
Configure vLLM server parameters. Leave empty to use defaults.
Maximum context length the model can handle
Fraction of GPU memory to use (0.1 - 1.0)
Quick Start Runner
Loading command... Commands are auto-generated based on your configuration settings.
This model requires a Hugging Face access token. The token is inserted into the command and never stored.
Configure vLLM server parameters. Leave empty to use defaults.
Maximum context length the model can handle
Fraction of GPU memory to use (0.1 - 1.0)
Quick Start Runner
Loading command... Commands are auto-generated based on your configuration settings.
This model requires a Hugging Face access token. The token is inserted into the command and never stored.
Configure vLLM server parameters. Leave empty to use defaults.
Maximum context length the model can handle
Fraction of GPU memory to use (0.1 - 1.0)
Quick Start Runner
Loading command... Commands are auto-generated based on your configuration settings.
This model requires a Hugging Face access token. The token is inserted into the command and never stored.
Configure vLLM server parameters. Leave empty to use defaults.
Maximum context length the model can handle
Fraction of GPU memory to use (0.1 - 1.0)
Quick Start Runner
Loading command... Commands are auto-generated based on your configuration settings.
This model requires a Hugging Face access token. The token is inserted into the command and never stored.
Configure vLLM server parameters. Leave empty to use defaults.
Maximum context length the model can handle
Fraction of GPU memory to use (0.1 - 1.0)
Quick Start Runner
Loading command... Commands are auto-generated based on your configuration settings.
This model requires a Hugging Face access token. The token is inserted into the command and never stored.
Configure vLLM server parameters. Leave empty to use defaults.
Maximum context length the model can handle
Fraction of GPU memory to use (0.1 - 1.0)
Quick Start Runner
Loading command... Commands are auto-generated based on your configuration settings.
This model requires a Hugging Face access token. The token is inserted into the command and never stored.
Configure vLLM server parameters. Leave empty to use defaults.
Maximum context length the model can handle
Fraction of GPU memory to use (0.1 - 1.0)
Quick Start Runner
Loading command... Commands are auto-generated based on your configuration settings.
This model requires a Hugging Face access token. The token is inserted into the command and never stored.
Configure vLLM server parameters. Leave empty to use defaults.
Maximum context length the model can handle
Fraction of GPU memory to use (0.1 - 1.0)
Quick Start Runner
Loading command... Commands are auto-generated based on your configuration settings.
This model requires a Hugging Face access token. The token is inserted into the command and never stored.
Configure vLLM server parameters. Leave empty to use defaults.
Maximum context length the model can handle
Fraction of GPU memory to use (0.1 - 1.0)
Quick Start Runner
Loading command... Commands are auto-generated based on your configuration settings.
This model requires a Hugging Face access token. The token is inserted into the command and never stored.
Configure vLLM server parameters. Leave empty to use defaults.
Maximum context length the model can handle
Fraction of GPU memory to use (0.1 - 1.0)
Quick Start Runner
Loading command... Commands are auto-generated based on your configuration settings.
This model requires a Hugging Face access token. The token is inserted into the command and never stored.
Configure vLLM server parameters. Leave empty to use defaults.
Maximum context length the model can handle
Fraction of GPU memory to use (0.1 - 1.0)
Quick Start Runner
Loading command... Commands are auto-generated based on your configuration settings.
This model requires a Hugging Face access token. The token is inserted into the command and never stored.
Configure vLLM server parameters. Leave empty to use defaults.
Maximum context length the model can handle
Fraction of GPU memory to use (0.1 - 1.0)
Quick Start Runner
Loading command... Commands are auto-generated based on your configuration settings.
This model requires a Hugging Face access token. The token is inserted into the command and never stored.
Configure vLLM server parameters. Leave empty to use defaults.
Maximum context length the model can handle
Fraction of GPU memory to use (0.1 - 1.0)
Quick Start Runner
Loading command... Commands are auto-generated based on your configuration settings.
This model requires a Hugging Face access token. The token is inserted into the command and never stored.
Configure vLLM server parameters. Leave empty to use defaults.
Maximum context length the model can handle
Fraction of GPU memory to use (0.1 - 1.0)
Quick Start Runner
Loading command... Commands are auto-generated based on your configuration settings.
This model requires a Hugging Face access token. The token is inserted into the command and never stored.
Configure vLLM server parameters. Leave empty to use defaults.
Maximum context length the model can handle
Fraction of GPU memory to use (0.1 - 1.0)
Quick Start Runner
Loading command... Commands are auto-generated based on your configuration settings.
This model requires a Hugging Face access token. The token is inserted into the command and never stored.
Configure vLLM server parameters. Leave empty to use defaults.
Maximum context length the model can handle
Fraction of GPU memory to use (0.1 - 1.0)
Quick Start Runner
Loading command... Commands are auto-generated based on your configuration settings.
This model requires a Hugging Face access token. The token is inserted into the command and never stored.
Configure vLLM server parameters. Leave empty to use defaults.
Maximum context length the model can handle
Fraction of GPU memory to use (0.1 - 1.0)
Quick Start Runner
Loading command... Commands are auto-generated based on your configuration settings.
This model requires a Hugging Face access token. The token is inserted into the command and never stored.
Configure vLLM server parameters. Leave empty to use defaults.
Maximum context length the model can handle
Fraction of GPU memory to use (0.1 - 1.0)
Quick Start Runner
Loading command... Commands are auto-generated based on your configuration settings.
This model requires a Hugging Face access token. The token is inserted into the command and never stored.
Configure vLLM server parameters. Leave empty to use defaults.
Maximum context length the model can handle
Fraction of GPU memory to use (0.1 - 1.0)
Quick Start Runner
Loading command... Commands are auto-generated based on your configuration settings.
This model requires a Hugging Face access token. The token is inserted into the command and never stored.
Configure vLLM server parameters. Leave empty to use defaults.
Maximum context length the model can handle
Fraction of GPU memory to use (0.1 - 1.0)
Quick Start Runner
Loading command... Commands are auto-generated based on your configuration settings.
This model requires a Hugging Face access token. The token is inserted into the command and never stored.
Configure vLLM server parameters. Leave empty to use defaults.
Maximum context length the model can handle
Fraction of GPU memory to use (0.1 - 1.0)
Quick Start Runner
Loading command... Commands are auto-generated based on your configuration settings.
This model requires a Hugging Face access token. The token is inserted into the command and never stored.
Configure vLLM server parameters. Leave empty to use defaults.
Maximum context length the model can handle
Fraction of GPU memory to use (0.1 - 1.0)
Quick Start Runner
Loading command... Commands are auto-generated based on your configuration settings.
This model requires a Hugging Face access token. The token is inserted into the command and never stored.
Configure vLLM server parameters. Leave empty to use defaults.
Maximum context length the model can handle
Fraction of GPU memory to use (0.1 - 1.0)
Quick Start Runner
Loading command... Commands are auto-generated based on your configuration settings.
This model requires a Hugging Face access token. The token is inserted into the command and never stored.
Configure vLLM server parameters. Leave empty to use defaults.
Maximum context length the model can handle
Fraction of GPU memory to use (0.1 - 1.0)
Quick Start Runner
Loading command... Commands are auto-generated based on your configuration settings.
This model requires a Hugging Face access token. The token is inserted into the command and never stored.
Configure vLLM server parameters. Leave empty to use defaults.
Maximum context length the model can handle
Fraction of GPU memory to use (0.1 - 1.0)
Quick Start Runner
Loading command... Commands are auto-generated based on your configuration settings.
This model requires a Hugging Face access token. The token is inserted into the command and never stored.
Configure vLLM server parameters. Leave empty to use defaults.
Maximum context length the model can handle
Fraction of GPU memory to use (0.1 - 1.0)
Quick Start Runner
Loading command... Commands are auto-generated based on your configuration settings.
This model requires a Hugging Face access token. The token is inserted into the command and never stored.
Configure vLLM server parameters. Leave empty to use defaults.
Maximum context length the model can handle
Fraction of GPU memory to use (0.1 - 1.0)
Quick Start Runner
Loading command... Commands are auto-generated based on your configuration settings.
This model requires a Hugging Face access token. The token is inserted into the command and never stored.
Configure vLLM server parameters. Leave empty to use defaults.
Maximum context length the model can handle
Fraction of GPU memory to use (0.1 - 1.0)
Quick Start Runner
Loading command... Commands are auto-generated based on your configuration settings.
This model requires a Hugging Face access token. The token is inserted into the command and never stored.
Configure vLLM server parameters. Leave empty to use defaults.
Maximum context length the model can handle
Fraction of GPU memory to use (0.1 - 1.0)
Quick Start Runner
Loading command... Commands are auto-generated based on your configuration settings.
This model requires a Hugging Face access token. The token is inserted into the command and never stored.
Configure vLLM server parameters. Leave empty to use defaults.
Maximum context length the model can handle
Fraction of GPU memory to use (0.1 - 1.0)
| | Modules | Inference engines | Actions | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| FLUX 3 Action New | Black Forest Labs | VLM | — | — | — | — | — | — | — | — | — | |
| Qwen3.8 Flash Next New | Alibaba Qwen3.8 | VLM | ✓ | — | — | — | — | ✓ | — | — | — | Qwen3.8 Flash NextQuick Start Runner Inference Engine Commands are auto-generated based on your configuration settings. Advanced configurationAuthenticationThis model requires a Hugging Face access token. The token is inserted into the command and never stored. vLLM ConfigurationConfigure vLLM server parameters. Leave empty to use defaults. Maximum context length the model can handle Fraction of GPU memory to use (0.1 - 1.0) |
| Qwen3.8 27B New | Alibaba Qwen3.8 | VLM | ✓ | ✓ | ✓ | — | — | — | — | ✓ | — | Qwen3.8 27BQuick Start Runner Inference Engine Commands are auto-generated based on your configuration settings. Advanced configurationAuthenticationThis model requires a Hugging Face access token. The token is inserted into the command and never stored. vLLM ConfigurationConfigure vLLM server parameters. Leave empty to use defaults. Maximum context length the model can handle Fraction of GPU memory to use (0.1 - 1.0) |
| Nemotron 3.5 Lightning | NVIDIA Nemotron | — | ✓ | ✓ | ✓ | — | — | ✓ | — | ✓ | ✓ | Nemotron 3.5 LightningQuick Start Runner Inference Engine Commands are auto-generated based on your configuration settings. Advanced configurationAuthenticationThis model requires a Hugging Face access token. The token is inserted into the command and never stored. vLLM ConfigurationConfigure vLLM server parameters. Leave empty to use defaults. Maximum context length the model can handle Fraction of GPU memory to use (0.1 - 1.0) |
| Muse Glimmer 30B | Meta Muse | VLM | ✓ | ✓ | ✓ | — | — | — | — | ✓ | — | Muse Glimmer 30BQuick Start Runner Inference Engine Commands are auto-generated based on your configuration settings. Advanced configurationAuthenticationThis model requires a Hugging Face access token. The token is inserted into the command and never stored. vLLM ConfigurationConfigure vLLM server parameters. Leave empty to use defaults. Maximum context length the model can handle Fraction of GPU memory to use (0.1 - 1.0) |
| Nemotron3 Nano 4B | NVIDIA Nemotron | — | ✓ | ✓ | ✓ | ✓ | ✓ | — | — | ✓ | — | Nemotron3 Nano 4BQuick Start Runner Inference Engine Commands are auto-generated based on your configuration settings. Advanced configurationAuthenticationThis model requires a Hugging Face access token. The token is inserted into the command and never stored. vLLM ConfigurationConfigure vLLM server parameters. Leave empty to use defaults. Maximum context length the model can handle Fraction of GPU memory to use (0.1 - 1.0) |
| FunctionGemma | Google Gemma3 | — | ✓ | ✓ | ✓ | ✓ | ✓ | — | — | ✓ | — | FunctionGemmaQuick Start Runner Inference Engine Commands are auto-generated based on your configuration settings. Advanced configurationAuthenticationThis model requires a Hugging Face access token. The token is inserted into the command and never stored. vLLM ConfigurationConfigure vLLM server parameters. Leave empty to use defaults. Maximum context length the model can handle Fraction of GPU memory to use (0.1 - 1.0) |
| Cosmos Reason 1 7B | NVIDIA Cosmos Reason | VLM | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | — | — | Cosmos Reason 1 7BQuick Start Runner Inference Engine Commands are auto-generated based on your configuration settings. Advanced configurationAuthenticationThis model requires a Hugging Face access token. The token is inserted into the command and never stored. vLLM ConfigurationConfigure vLLM server parameters. Leave empty to use defaults. Maximum context length the model can handle Fraction of GPU memory to use (0.1 - 1.0) |
| Gemma 3 270M | Google Gemma3 | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | — | Gemma 3 270MQuick Start Runner Inference Engine Commands are auto-generated based on your configuration settings. Advanced configurationAuthenticationThis model requires a Hugging Face access token. The token is inserted into the command and never stored. vLLM ConfigurationConfigure vLLM server parameters. Leave empty to use defaults. Maximum context length the model can handle Fraction of GPU memory to use (0.1 - 1.0) |
| Gemma 4 E2B | Google Gemma4 | VLM | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | ✓ | — | Gemma 4 E2BQuick Start Runner Inference Engine Commands are auto-generated based on your configuration settings. Advanced configurationAuthenticationThis model requires a Hugging Face access token. The token is inserted into the command and never stored. vLLM ConfigurationConfigure vLLM server parameters. Leave empty to use defaults. Maximum context length the model can handle Fraction of GPU memory to use (0.1 - 1.0) |
| GPT OSS 20B | OpenAI GPT OSS | — | ✓ | ✓ | ✓ | — | — | ✓ | — | — | — | GPT OSS 20BQuick Start Runner Inference Engine Commands are auto-generated based on your configuration settings. Advanced configurationAuthenticationThis model requires a Hugging Face access token. The token is inserted into the command and never stored. vLLM ConfigurationConfigure vLLM server parameters. Leave empty to use defaults. Maximum context length the model can handle Fraction of GPU memory to use (0.1 - 1.0) |
| Llama 3.2 3B | Meta Llama 3 | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | — | Llama 3.2 3BQuick Start Runner Inference Engine Commands are auto-generated based on your configuration settings. Advanced configurationAuthenticationThis model requires a Hugging Face access token. The token is inserted into the command and never stored. vLLM ConfigurationConfigure vLLM server parameters. Leave empty to use defaults. Maximum context length the model can handle Fraction of GPU memory to use (0.1 - 1.0) |
| MiniMax M2.7 | MiniMax M2.7 | — | ✓ | — | — | — | — | — | — | ✓ | — | MiniMax M2.7Quick Start Runner Inference Engine Commands are auto-generated based on your configuration settings. Advanced configurationAuthenticationThis model requires a Hugging Face access token. The token is inserted into the command and never stored. vLLM ConfigurationConfigure vLLM server parameters. Leave empty to use defaults. Maximum context length the model can handle Fraction of GPU memory to use (0.1 - 1.0) |
| Ministral 3 3B Instruct | Mistral AI Ministral 3 | VLM | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | — | Ministral 3 3B InstructQuick Start Runner Inference Engine Commands are auto-generated based on your configuration settings. Advanced configurationAuthenticationThis model requires a Hugging Face access token. The token is inserted into the command and never stored. vLLM ConfigurationConfigure vLLM server parameters. Leave empty to use defaults. Maximum context length the model can handle Fraction of GPU memory to use (0.1 - 1.0) |
| Nemotron3 Nano 30B-A3B | NVIDIA Nemotron | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | — | Nemotron3 Nano 30B-A3BQuick Start Runner Inference Engine Commands are auto-generated based on your configuration settings. Advanced configurationAuthenticationThis model requires a Hugging Face access token. The token is inserted into the command and never stored. vLLM ConfigurationConfigure vLLM server parameters. Leave empty to use defaults. Maximum context length the model can handle Fraction of GPU memory to use (0.1 - 1.0) |
| Qwen3 4B | Alibaba Qwen3 | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | — | — | Qwen3 4BQuick Start Runner Inference Engine Commands are auto-generated based on your configuration settings. Advanced configurationAuthenticationThis model requires a Hugging Face access token. The token is inserted into the command and never stored. vLLM ConfigurationConfigure vLLM server parameters. Leave empty to use defaults. Maximum context length the model can handle Fraction of GPU memory to use (0.1 - 1.0) |
| Qwen3.5 35B-A3B (MoE) | Alibaba Qwen3.5 | — | ✓ | ✓ | ✓ | — | — | ✓ | — | — | — | Qwen3.5 35B-A3B (MoE)Quick Start Runner Inference Engine Commands are auto-generated based on your configuration settings. Advanced configurationAuthenticationThis model requires a Hugging Face access token. The token is inserted into the command and never stored. vLLM ConfigurationConfigure vLLM server parameters. Leave empty to use defaults. Maximum context length the model can handle Fraction of GPU memory to use (0.1 - 1.0) |
| Qwen3.6 35B-A3B (MoE) | Alibaba Qwen3.6 | — | ✓ | ✓ | ✓ | — | — | ✓ | — | — | — | Qwen3.6 35B-A3B (MoE)Quick Start Runner Inference Engine Commands are auto-generated based on your configuration settings. Advanced configurationAuthenticationThis model requires a Hugging Face access token. The token is inserted into the command and never stored. vLLM ConfigurationConfigure vLLM server parameters. Leave empty to use defaults. Maximum context length the model can handle Fraction of GPU memory to use (0.1 - 1.0) |
| Gemma 3 1B | Google Gemma3 | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | — | Gemma 3 1BQuick Start Runner Inference Engine Commands are auto-generated based on your configuration settings. Advanced configurationAuthenticationThis model requires a Hugging Face access token. The token is inserted into the command and never stored. vLLM ConfigurationConfigure vLLM server parameters. Leave empty to use defaults. Maximum context length the model can handle Fraction of GPU memory to use (0.1 - 1.0) |
| Gemma 4 12B | Google Gemma4 | VLM | ✓ | ✓ | ✓ | — | — | ✓ | — | ✓ | — | Gemma 4 12BQuick Start Runner Inference Engine Commands are auto-generated based on your configuration settings. Advanced configurationAuthenticationThis model requires a Hugging Face access token. The token is inserted into the command and never stored. vLLM ConfigurationConfigure vLLM server parameters. Leave empty to use defaults. Maximum context length the model can handle Fraction of GPU memory to use (0.1 - 1.0) |
| Gemma 4 E4B | Google Gemma4 | VLM | ✓ | ✓ | ✓ | ✓ | — | ✓ | — | ✓ | — | Gemma 4 E4BQuick Start Runner Inference Engine Commands are auto-generated based on your configuration settings. Advanced configurationAuthenticationThis model requires a Hugging Face access token. The token is inserted into the command and never stored. vLLM ConfigurationConfigure vLLM server parameters. Leave empty to use defaults. Maximum context length the model can handle Fraction of GPU memory to use (0.1 - 1.0) |
| GPT OSS 120B | OpenAI GPT OSS | — | ✓ | ✓ | — | — | — | ✓ | — | — | — | GPT OSS 120BQuick Start Runner Inference Engine Commands are auto-generated based on your configuration settings. Advanced configurationAuthenticationThis model requires a Hugging Face access token. The token is inserted into the command and never stored. vLLM ConfigurationConfigure vLLM server parameters. Leave empty to use defaults. Maximum context length the model can handle Fraction of GPU memory to use (0.1 - 1.0) |
| Llama 3.1 8B | Meta Llama 3 | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | — | Llama 3.1 8BQuick Start Runner Inference Engine Commands are auto-generated based on your configuration settings. Advanced configurationAuthenticationThis model requires a Hugging Face access token. The token is inserted into the command and never stored. vLLM ConfigurationConfigure vLLM server parameters. Leave empty to use defaults. Maximum context length the model can handle Fraction of GPU memory to use (0.1 - 1.0) |
| Ministral 3 8B Instruct | Mistral AI Ministral 3 | VLM | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | — | Ministral 3 8B InstructQuick Start Runner Inference Engine Commands are auto-generated based on your configuration settings. Advanced configurationAuthenticationThis model requires a Hugging Face access token. The token is inserted into the command and never stored. vLLM ConfigurationConfigure vLLM server parameters. Leave empty to use defaults. Maximum context length the model can handle Fraction of GPU memory to use (0.1 - 1.0) |
| Nemotron Nano 9B v2 | NVIDIA Nemotron | — | ✓ | ✓ | — | — | — | ✓ | — | — | — | Nemotron Nano 9B v2Quick Start Runner Inference Engine Commands are auto-generated based on your configuration settings. Advanced configurationAuthenticationThis model requires a Hugging Face access token. The token is inserted into the command and never stored. vLLM ConfigurationConfigure vLLM server parameters. Leave empty to use defaults. Maximum context length the model can handle Fraction of GPU memory to use (0.1 - 1.0) |
| Qwen3.5 27B | Alibaba Qwen3.5 | — | ✓ | ✓ | ✓ | — | — | ✓ | — | — | — | Qwen3.5 27BQuick Start Runner Inference Engine Commands are auto-generated based on your configuration settings. Advanced configurationAuthenticationThis model requires a Hugging Face access token. The token is inserted into the command and never stored. vLLM ConfigurationConfigure vLLM server parameters. Leave empty to use defaults. Maximum context length the model can handle Fraction of GPU memory to use (0.1 - 1.0) |
| Qwen3.6 27B | Alibaba Qwen3.6 | — | ✓ | ✓ | ✓ | — | — | ✓ | — | — | — | Qwen3.6 27BQuick Start Runner Inference Engine Commands are auto-generated based on your configuration settings. Advanced configurationAuthenticationThis model requires a Hugging Face access token. The token is inserted into the command and never stored. vLLM ConfigurationConfigure vLLM server parameters. Leave empty to use defaults. Maximum context length the model can handle Fraction of GPU memory to use (0.1 - 1.0) |
| Qwen3 8B | Alibaba Qwen3 | — | ✓ | ✓ | ✓ | ✓ | — | ✓ | — | — | — | Qwen3 8BQuick Start Runner Inference Engine Commands are auto-generated based on your configuration settings. Advanced configurationAuthenticationThis model requires a Hugging Face access token. The token is inserted into the command and never stored. vLLM ConfigurationConfigure vLLM server parameters. Leave empty to use defaults. Maximum context length the model can handle Fraction of GPU memory to use (0.1 - 1.0) |
| Cosmos Reason 2 2B | NVIDIA Cosmos Reason | VLM | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | ✓ | — | |
| Gemma 3 4B | Google Gemma3 | VLM | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | — | Gemma 3 4BQuick Start Runner Inference Engine Commands are auto-generated based on your configuration settings. Advanced configurationAuthenticationThis model requires a Hugging Face access token. The token is inserted into the command and never stored. vLLM ConfigurationConfigure vLLM server parameters. Leave empty to use defaults. Maximum context length the model can handle Fraction of GPU memory to use (0.1 - 1.0) |
| Gemma 4 26B-A4B | Google Gemma4 | VLM | ✓ | ✓ | ✓ | — | — | ✓ | — | ✓ | — | Gemma 4 26B-A4BQuick Start Runner Inference Engine Commands are auto-generated based on your configuration settings. Advanced configurationAuthenticationThis model requires a Hugging Face access token. The token is inserted into the command and never stored. vLLM ConfigurationConfigure vLLM server parameters. Leave empty to use defaults. Maximum context length the model can handle Fraction of GPU memory to use (0.1 - 1.0) |
| Llama 3.1 70B | Meta Llama 3 | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | — | Llama 3.1 70BQuick Start Runner Inference Engine Commands are auto-generated based on your configuration settings. Advanced configurationAuthenticationThis model requires a Hugging Face access token. The token is inserted into the command and never stored. vLLM ConfigurationConfigure vLLM server parameters. Leave empty to use defaults. Maximum context length the model can handle Fraction of GPU memory to use (0.1 - 1.0) |
| Ministral 3 14B Instruct | Mistral AI Ministral 3 | VLM | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | — | Ministral 3 14B InstructQuick Start Runner Inference Engine Commands are auto-generated based on your configuration settings. Advanced configurationAuthenticationThis model requires a Hugging Face access token. The token is inserted into the command and never stored. vLLM ConfigurationConfigure vLLM server parameters. Leave empty to use defaults. Maximum context length the model can handle Fraction of GPU memory to use (0.1 - 1.0) |
| Nemotron 3 Super 120B-A12B | NVIDIA Nemotron | — | ✓ | — | — | — | — | ✓ | — | ✓ | — | Nemotron 3 Super 120B-A12BQuick Start Runner Inference Engine Commands are auto-generated based on your configuration settings. Advanced configurationAuthenticationThis model requires a Hugging Face access token. The token is inserted into the command and never stored. vLLM ConfigurationConfigure vLLM server parameters. Leave empty to use defaults. Maximum context length the model can handle Fraction of GPU memory to use (0.1 - 1.0) |
| Nemotron Nano 12B VL | NVIDIA Nemotron | VLM | ✓ | ✓ | — | — | — | ✓ | — | — | — | Nemotron Nano 12B VLQuick Start Runner Inference Engine Commands are auto-generated based on your configuration settings. Advanced configurationAuthenticationThis model requires a Hugging Face access token. The token is inserted into the command and never stored. vLLM ConfigurationConfigure vLLM server parameters. Leave empty to use defaults. Maximum context length the model can handle Fraction of GPU memory to use (0.1 - 1.0) |
| Qwen3 30B-A3B (MoE) | Alibaba Qwen3 | — | ✓ | ✓ | ✓ | — | — | ✓ | — | — | — | Qwen3 30B-A3B (MoE)Quick Start Runner Inference Engine Commands are auto-generated based on your configuration settings. Advanced configurationAuthenticationThis model requires a Hugging Face access token. The token is inserted into the command and never stored. vLLM ConfigurationConfigure vLLM server parameters. Leave empty to use defaults. Maximum context length the model can handle Fraction of GPU memory to use (0.1 - 1.0) |
| Qwen3.5 9B | Alibaba Qwen3.5 | VLM | ✓ | ✓ | ✓ | ✓ | — | ✓ | — | — | — | Qwen3.5 9BQuick Start Runner Inference Engine Commands are auto-generated based on your configuration settings. Advanced configurationAuthenticationThis model requires a Hugging Face access token. The token is inserted into the command and never stored. vLLM ConfigurationConfigure vLLM server parameters. Leave empty to use defaults. Maximum context length the model can handle Fraction of GPU memory to use (0.1 - 1.0) |
| Cosmos3 Edge | NVIDIA Cosmos | VLM | ✓ | ✓ | ✓ | — | — | ✓ | — | — | — | Cosmos3 EdgeQuick Start Runner Inference Engine Commands are auto-generated based on your configuration settings. Advanced configurationAuthenticationThis model requires a Hugging Face access token. The token is inserted into the command and never stored. vLLM ConfigurationConfigure vLLM server parameters. Leave empty to use defaults. Maximum context length the model can handle Fraction of GPU memory to use (0.1 - 1.0) |
| Gemma 3 12B | Google Gemma3 | VLM | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | — | Gemma 3 12BQuick Start Runner Inference Engine Commands are auto-generated based on your configuration settings. Advanced configurationAuthenticationThis model requires a Hugging Face access token. The token is inserted into the command and never stored. vLLM ConfigurationConfigure vLLM server parameters. Leave empty to use defaults. Maximum context length the model can handle Fraction of GPU memory to use (0.1 - 1.0) |
| Cosmos Reason 2 8B | NVIDIA Cosmos Reason | VLM | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | ✓ | — | |
| Gemma 4 31B | Google Gemma4 | VLM | ✓ | ✓ | ✓ | — | — | ✓ | — | ✓ | — | Gemma 4 31BQuick Start Runner Inference Engine Commands are auto-generated based on your configuration settings. Advanced configurationAuthenticationThis model requires a Hugging Face access token. The token is inserted into the command and never stored. vLLM ConfigurationConfigure vLLM server parameters. Leave empty to use defaults. Maximum context length the model can handle Fraction of GPU memory to use (0.1 - 1.0) |
| Ministral 3 3B Reasoning | Mistral AI Ministral 3 | VLM | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | — | — | Ministral 3 3B ReasoningQuick Start Runner Inference Engine Commands are auto-generated based on your configuration settings. Advanced configurationAuthenticationThis model requires a Hugging Face access token. The token is inserted into the command and never stored. vLLM ConfigurationConfigure vLLM server parameters. Leave empty to use defaults. Maximum context length the model can handle Fraction of GPU memory to use (0.1 - 1.0) |
| Qwen3 32B | Alibaba Qwen3 | — | ✓ | ✓ | — | — | — | ✓ | — | — | — | Qwen3 32BQuick Start Runner Inference Engine Commands are auto-generated based on your configuration settings. Advanced configurationAuthenticationThis model requires a Hugging Face access token. The token is inserted into the command and never stored. vLLM ConfigurationConfigure vLLM server parameters. Leave empty to use defaults. Maximum context length the model can handle Fraction of GPU memory to use (0.1 - 1.0) |
| Nemotron 3 Nano Omni | NVIDIA Nemotron | VLM | ✓ | ✓ | ✓ | — | — | ✓ | ✓ | ✓ | — | Nemotron 3 Nano OmniQuick Start Runner Inference Engine Commands are auto-generated based on your configuration settings. Advanced configurationAuthenticationThis model requires a Hugging Face access token. The token is inserted into the command and never stored. vLLM ConfigurationConfigure vLLM server parameters. Leave empty to use defaults. Maximum context length the model can handle Fraction of GPU memory to use (0.1 - 1.0) |
| Qwen3.5 4B | Alibaba Qwen3.5 | VLM | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | — | — | Qwen3.5 4BQuick Start Runner Inference Engine Commands are auto-generated based on your configuration settings. Advanced configurationAuthenticationThis model requires a Hugging Face access token. The token is inserted into the command and never stored. vLLM ConfigurationConfigure vLLM server parameters. Leave empty to use defaults. Maximum context length the model can handle Fraction of GPU memory to use (0.1 - 1.0) |
| DiffusionGemma 26B-A4B | Google Gemma4 | — | ✓ | ✓ | ✓ | — | — | ✓ | — | — | — | DiffusionGemma 26B-A4BQuick Start Runner Inference Engine Commands are auto-generated based on your configuration settings. Advanced configurationAuthenticationThis model requires a Hugging Face access token. The token is inserted into the command and never stored. vLLM ConfigurationConfigure vLLM server parameters. Leave empty to use defaults. Maximum context length the model can handle Fraction of GPU memory to use (0.1 - 1.0) |
| Cosmos3 Nano | NVIDIA Cosmos | VLM | ✓ | ✓ | — | — | — | ✓ | — | — | — | |
| Gemma 3 27B | Google Gemma3 | VLM | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | — | Gemma 3 27BQuick Start Runner Inference Engine Commands are auto-generated based on your configuration settings. Advanced configurationAuthenticationThis model requires a Hugging Face access token. The token is inserted into the command and never stored. vLLM ConfigurationConfigure vLLM server parameters. Leave empty to use defaults. Maximum context length the model can handle Fraction of GPU memory to use (0.1 - 1.0) |
| Ministral 3 8B Reasoning | Mistral AI Ministral 3 | VLM | ✓ | ✓ | ✓ | ✓ | — | ✓ | — | — | — | Ministral 3 8B ReasoningQuick Start Runner Inference Engine Commands are auto-generated based on your configuration settings. Advanced configurationAuthenticationThis model requires a Hugging Face access token. The token is inserted into the command and never stored. vLLM ConfigurationConfigure vLLM server parameters. Leave empty to use defaults. Maximum context length the model can handle Fraction of GPU memory to use (0.1 - 1.0) |
| Qwen3.5 0.8B | Alibaba Qwen3.5 | VLM | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | — | — | Qwen3.5 0.8BQuick Start Runner Inference Engine Commands are auto-generated based on your configuration settings. Advanced configurationAuthenticationThis model requires a Hugging Face access token. The token is inserted into the command and never stored. vLLM ConfigurationConfigure vLLM server parameters. Leave empty to use defaults. Maximum context length the model can handle Fraction of GPU memory to use (0.1 - 1.0) |
| Qwen3 VL 4B | Alibaba Qwen3 | VLM | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | — | — | — | Qwen3 VL 4BQuick Start Runner Inference Engine Commands are auto-generated based on your configuration settings. Advanced configurationAuthenticationThis model requires a Hugging Face access token. The token is inserted into the command and never stored. vLLM ConfigurationConfigure vLLM server parameters. Leave empty to use defaults. Maximum context length the model can handle Fraction of GPU memory to use (0.1 - 1.0) |
| Ministral 3 14B Reasoning | Mistral AI Ministral 3 | VLM | ✓ | ✓ | ✓ | — | — | ✓ | — | — | — | Ministral 3 14B ReasoningQuick Start Runner Inference Engine Commands are auto-generated based on your configuration settings. Advanced configurationAuthenticationThis model requires a Hugging Face access token. The token is inserted into the command and never stored. vLLM ConfigurationConfigure vLLM server parameters. Leave empty to use defaults. Maximum context length the model can handle Fraction of GPU memory to use (0.1 - 1.0) |
| Qwen3 VL 8B | Alibaba Qwen3 | VLM | ✓ | ✓ | ✓ | ✓ | — | ✓ | — | — | — | Qwen3 VL 8BQuick Start Runner Inference Engine Commands are auto-generated based on your configuration settings. Advanced configurationAuthenticationThis model requires a Hugging Face access token. The token is inserted into the command and never stored. vLLM ConfigurationConfigure vLLM server parameters. Leave empty to use defaults. Maximum context length the model can handle Fraction of GPU memory to use (0.1 - 1.0) |
Try a different model name filter or clear column checkboxes in the table header
Benchmarks on Jetson across supported inference runtimes
* ISL/OSL for all benchmarks: 2048/128
* Unless otherwise specified, all models utilize W4A16 quantization for Orin and NVFP4 for Thor.
* NVFP4 and MXFP4 require Blackwell FP4 tensor cores and are not available on Orin (Ampere).
* Edge-LLM benchmark values are generation throughput from the official TensorRT Edge-LLM performance benchmarks on Jetson AGX Thor Developer Kit.
* Cosmos3 Edge (Reasoner) (BF16): streaming chat harness, ISL 1705 / OSL 128, per-request prompt-cache salting, median of 3 - not the aiperf ISL-2048 protocol used for other rows.