The frameworks module¶
In this module we provide pre-installed packages for various AI/ML frameworks like pytorch and vllm through a conda environment as a part of the compute image on Aurora.
help([[The Frameworks Intelpython environment.
Includes installations of PyTorch with extensions from Intel
torch 2.13.0a0+gitcf30153
torchao 0.17.0+git02105d46c
torchcodec 0.15.0
torchcomms 0.3.1
torchdata 0.11.0+377e64c
torchvision 0.28.0+8fb8771
triton-xpu 3.7.2
mpi4py 4.1.2
vllm 0.26.1.dev0+g568afb3a1.d20260803.xpu
vllm-xpu-kernels 0.1.11.2.dev0+ga692986.d20260803
deepspeed 0.19.3
dpctl 0.23.0.dev0+205.gb24f931fde
dpnp 0.21.0.dev3+8.g987f2992697
scikit-learn 1.9.0
scikit-learn-intelex 20260728.214749
--##
You can modify this environment as follows:
- Extend this environment locally
$ pip install --user [package]
- Create a new one of your own
$ conda create -n [environment_name] [package]
https://docs.conda.io/projects/conda/en/latest/user-guide/getting-started.html
]])
whatis("Name: frameworks")
whatis("Version: 2026.1.0")
whatis("Category: oneapi frameworks")
whatis("Keywords: oneapi frameworks")
whatis("Description: Aurora frameworks python environment")
whatis("URL: https://docs.conda.io/projects/conda/en/latest/user-guide/getting-started.html")
depends_on("oneapi/release/2026.1.0")
depends_on("intel_gpu_umd_aicoe")
depends_on("hdf5")
depends_on("pti-gpu")
depends_on("miniforge3")
setenv("ENV_NAME","frameworks/2026.1.0")
setenv("PYTHONUSERBASE","$HOME/.local/aurora/frameworks/2026.1.0")
unsetenv("PYTHONSTARTUP")
setenv("ZE_FLAT_DEVICE_HIERARCHY","FLAT")
setenv("CCL_PROCESS_LAUNCHER","pmix")
setenv("TORCH_CPP_LOG_LEVEL","ERROR")
execute{cmd="conda activate /opt/aurora/26.181.0/frameworks/aurora_frameworks-2026.1.0;", modeA={"load"}}
family("frameworks")
Global Changes in frameworks/2026.1.0¶
The following are the global changes that we have introduced in this iteration
CCL_OP_SYNC=0. Frameworks module does not setCCL_OP_SYNCto1any more (see Hangs withCCL_OP_SYNC=0).ONEAPI_DEVICE_SELECTOR="level_zero:gpu"as set by theoneapi/release/2026.1.0. The module does not set it any more.
Known issues¶
MPI_Init error with CCL_KVS_MODE=mpi¶
To scale out beyond 1024 nodes on Aurora, you may need to set export CCL_KVS_MODE=mpi. Because of a change in how oneCCL interacts with MPI, a distributed application may then fail with an error about using mpi before it is initialized. We are investigating the issue further.
Workaround¶
Initialize MPI manually. From a python/PyTorch standpoint, import mpi4py performs the MPI_Init. The application does not need to use mpi4py; only the import is needed.
Hangs with CCL_OP_SYNC=0¶
Historically, the frameworks module set CCL_OP_SYNC=1, so collectives ran in a synchronized fashion. It no longer does, because of an issue with XPUGraph capturing, a new feature in this iteration. As a side effect, multi-node distributed jobs may hang.
Workaround¶
If a job hangs, restore the legacy behavior with export CCL_OP_SYNC=1. You may also set export CCL_ATL_SYNC_COLL=1.
Tracking changes¶
This section is an attempt to keep track of high-level changes to the module
Major changes in the frameworks/2025.3.1¶
- The
torch_cclmodule has been removed.import oneccl_bindings_for_pytorch as torch_cclis no longer needed. - When initializing
torch.distributed, thebackendmust be changed toxcclfromccl. import intel_extension_for_pytorch as ipexis now deprecated. The vendor is upstreaming all of the functionality from IPEX to the mainline PyTorch distribution. If you experience performance variations after removing the import, please switch back to importing it.horovodsupport for PyTorch has been removed.ONEAPI_DEVICE_SELECTORhas been set to"opencl:gpu;level_zero:gpu", if this causes any issues, please revert to Level Zero only withexport ONEAPI_DEVICE_SELECTOR="level_zero:gpu"