Available Models¶
Models are organized by cluster and marked with the following capabilities:
- B - Batch Processing Enabled
- T - Tool Calling Enabled
- R - Reasoning Enabled
- H - Always Hot Model
Sophia Cluster¶
Chat Language Models (vLLM)
Meta Llama Family
- meta-llama/Meta-Llama-3.1-8B-InstructBTH
- meta-llama/Meta-Llama-3.1-70B-InstructBTH
- meta-llama/Meta-Llama-3.1-405B-InstructBT
- meta-llama/Llama-3.3-70B-InstructBT
- meta-llama/Llama-4-Scout-17B-16E-InstructBT
- meta-llama/Llama-4-Maverick-17B-128E-InstructT
Mistral Family
- mistralai/Mistral-Large-Instruct-2407
- mistralai/Mixtral-8x22B-Instruct-v0.1
- mistralai/Devstral-2-123B-Instruct-2512
OpenAI Family
- openai/gpt-oss-20bBRTH
- openai/gpt-oss-120bBRTH
Aurora GPT Family
- argonne/AuroraGPT-IT-v4-0125B
- argonne/AuroraGPT-Tulu3-SFT-0125B
- argonne/AuroraGPT-DPO-UFB-0225B
- argonne/AuroraGPT-KTO-UFB-0325B
Google Family
- google/gemma-3-27b-itBT
- google/gemma-4-26B-A4B-itRT
- google/gemma-4-31B-itRTH
- google/gemma-4-E4B-itRTH
Other Models
- arcee-ai/Trinity-Large-Thinking-W4A16RT
- nvidia/nemotron-3-super-120bRT
- AstroMLab/AstroSage-70B-20251009
Vision Language Models (vLLM)
- meta-llama/Llama-3.2-90B-Vision-Instruct
Embedding Models (vLLM)
- mistralai/Mistral-7B-Instruct-v0.3-embed
- google/embeddinggemma-300m
- Salesforce/SFR-Embedding-Mistral
- genslm-test/genslm-esmc-300M-aminoacid
- genslm-test/genslm-esmc-300M-codon
- genslm-test/genslm-esmc-300M-contrastive-aminoacid
- genslm-test/genslm-esmc-300M-contrastive-codon
- genslm-test/genslm-esmc-300M-joint-aminoacid
- genslm-test/genslm-esmc-300M-joint-codon
- genslm-test/genslm-esmc-600M-aminoacid
- genslm-test/genslm-esmc-600M-codon
- genslm-test/genslm-esmc-600M-contrastive-aminoacid
- genslm-test/genslm-esmc-600M-contrastive-codon
- genslm-test/genslm-esmc-600M-joint-aminoacid
- genslm-test/genslm-esmc-600M-joint-codon
Image Segmentation
- sam3
- dinov3
Promptable Image Segmentation Models
The SAM 3 and DINOv3 models are deployed on Sophia. Install the alcf-ai package for the command line and Python toolkit. See alcf-ai CLI and SDK for worked examples, including batch and DINOv3 segmentation.
If this is your first time using alcf-ai, run alcf-tokens login inference first. Without a valid token, alcf-ai exits with an authentication error naming that command.
Metis Cluster (SambaNova)¶
Endpoint and model status is on the Metis status page. See SambaNova's OpenAI compatible API documentation for API details.
Chat Language Models
- gpt-oss-120bH
- Mistral-Large-3-675B-Instruct-2512H
- gemma-4-31B-itH
Metis Limitations
- Batch processing is not currently supported on the Metis cluster.
- Tool calling is advertised by the API, but SambaNova's sanitization of tool calls has a known issue. See the Metis Tool Calling warning under Shims/Proxies.
Minerva Cluster (NVIDIA)¶
Chat Language Models
- nemotron-3-ultraH
- inkling-bf16H
- gpt-oss-120b
Model Serving Configuration¶
When available, model serving configuration details can be viewed for each cluster.
#!/bin/bash
# Get your access token
access_token=$(alcf-tokens get-token inference)
# Check serving configuration for all Sophia models
curl -X GET "https://inference-api.alcf.anl.gov/resource_server/sophia/models" \
-H "Authorization: Bearer ${access_token}"
# Check serving configuration for a specific model (e.g., openai/gpt-oss-120b)
curl -X GET "https://inference-api.alcf.anl.gov/resource_server/sophia/models?model_id=openai/gpt-oss-120b" \
-H "Authorization: Bearer ${access_token}"
Model Capabilities¶
Each model includes a versioned capabilities object describing what the deployment supports. The current schema is version 1:
| Field | Description |
|---|---|
schema_version | Version of the capability object. |
api_protocols | API protocols the model accepts. |
context_window_tokens | Maximum context window, in tokens. |
input_modalities | Accepted input types. |
streaming | Whether streaming responses are supported. |
reasoning.supported | Whether the model supports reasoning. |
reasoning.effort_levels | Accepted reasoning effort values. |
reasoning.default_effort | Effort applied when a request omits one. |
reasoning.separate_output | Whether reasoning is returned separately as reasoning_content instead of inline. |
tool_calling.supported | Whether the model supports tool calling. |
Optional fields are omitted when they do not apply, so clients should fall back to conservative defaults. Models served by vLLM also report max_model_len, max_num_seqs, tool_call_parser, and enable_auto_tool_choice alongside capabilities.
Possible Values¶
| Field | Values |
|---|---|
schema_version | 1 |
api_protocols | chat_completions, responses, messages |
context_window_tokens | Any positive integer |
input_modalities | text, image, video |
streaming | true, false |
reasoning.supported | true, false |
reasoning.effort_levels | none, minimal, low, medium, high, xhigh, max |
reasoning.default_effort | One of the model's reasoning.effort_levels |
reasoning.separate_output | true, false |
tool_calling.supported | true, false |
alcf-ai ls-models <cluster> returns the same objects and uses them to configure agent harnesses. See alcf-ai CLI and SDK.
Want to add a model?
To request a new model, please contact ALCF Support.