Skip to main content

framework-rocm-containers

framework-rocm is a set of reproducible ROCm containers for the Framework Desktop (AMD Ryzen AI MAX / Strix Halo / Radeon 8060S, gfx1151). Developed by community member geoff-davis, it provides pinned PyTorch and JAX Docker images with GPU smoke tests and measured attention/training performance findings1.

The project exists because getting a recent ROCm version and a matching PyTorch or JAX talking to the gfx1151 GPU is fiddly, and re-deriving the right versions every time is tedious. The containers are built on AMD-published base images so ROCm and the framework arrive pre-pinned and mutually consistent1.

Images

Two sibling images, each built on the matching AMD-published base image1:

ImageBaseDefault Versions
pytorchrocm/pytorchROCm 7.2.4, PyTorch 2.10.0
jaxrocm/jaxROCm 7.2.4, JAX 0.8.2

Both images were verified on an actual Framework Desktop (last verified 2026-07-05): the bases see gfx1151 (Radeon 8060S) and run GPU compute natively — no HSA_OVERRIDE_GFX_VERSION and no gfx1151 fallback wheels required1.

gfx1151 Support

AMD's Radeon/Ryzen Linux matrix lists the Ryzen AI MAX+ 395 / Radeon 8060S (gfx1151) with production support on ROCm 7.2.1 + PyTorch 2.9.1 (FP16 validated). This repo defaults to AMD's newer published images (ROCm 7.2.4 base; PyTorch 2.10.0 / JAX 0.8.2) because they measure faster and cleaner on this hardware. To stay strictly on the AMD-validated combo, pin a rocm7.2.1/torch-2.9.1 tag instead1.

The gfx1151 Attention Gotcha

Enabling AOTriton mem-efficient SDPA (which these containers do by default) takes bf16 attention from 92 → 8.4 ms/iter (~11x) at BERT-base shape, and a real 110M-param sentence-encoder fine-tune from 9.8 → 1.10 s/step (~9x) once combined with the other documented levers (no gradient checkpointing, TunableOp, seq-length cap)1.

Two levers, in order1:

  1. bf16 — fp32 attention on gfx1151 is memory-bandwidth-bound
  2. TORCH_ROCM_AOTRITON_ENABLE_EXPERIMENTAL=1 — unlocks AOTriton mem-efficient SDPA (~11x faster bf16 attention)

PyTorch TunableOp (PYTORCH_TUNABLEOP_ENABLED=1) finds better GEMM kernels than the untuned gfx1151 defaults (~1.26x measured on a BERT fine-tune)1.

Quick Start

./run.sh pytorch build
./run.sh pytorch check # quick deterministic GPU correctness check
./run.sh jax check # quick deterministic GPU correctness check
./run.sh pytorch bench # correctness check + attention benchmark
./run.sh jax bench # same for JAX (includes XLA compilation)
./run.sh pytorch shell # interactive shell
./run.sh jax python your_script.py

A successful PyTorch correctness check output:

torch version : 2.10.0+rocm7.2.4.git3d3aa833
ROCm/HIP ver : 7.2.53211
device count : 1
[0] Radeon 8060S Graphics
gpu arch : gfx1151
matmul OK : sum=1073741824.000 on Radeon 8060S Graphics

Runtime Behaviour

  • Runs as your host user so files written to the mounted workspace aren't root-owned
  • HOME is /workspace inside the container
  • Caches persist across runs~/.cache/framework-rocm mounted at $HOME/.cache (pip downloads, MIOpen compiled-kernel cache)
  • Hugging Face models use the host's standard cache~/.cache/huggingface shared with native tools
  • Ports — Jupyter, TensorBoard via ROCM_PORTS="8888:8888"1

Requirements

  • A Framework Desktop (Ryzen AI Max / Strix Halo) — or any gfx1151 machine
  • A recent kernel with amdgpu loaded (/dev/kfd and /dev/dri/renderD* present)
  • Docker with the Compose plugin1
Not a Security Boundary

These containers use seccomp=unconfined, ipc=host, direct GPU device access, and host cache/workspace mounts. They are not a security boundary — treat code you run in them like code you'd run on the host, and don't point them at untrusted models or notebooks you wouldn't run natively1.

Technical Details

DetailValue
LanguagePython
LicenseMIT
Stars0
Commits21
Container formatDocker (with Compose plugin)
Default PyTorchROCm 7.2.4 / torch 2.10.0
Default JAXROCm 7.2.4 / jax 0.8.2
GPU verifiedRadeon 8060S (gfx1151), last verified 2026-07-05
AMD base images17–23 GB (shared via Docker layer storage)
Last updatedJuly 9, 2026
ProjectDescription
Strix Halo GuideComprehensive local LLM setup and benchmark guide for Strix Halo
strix-llmKnown-good local LLM inference configuration for Framework Desktop

Footnotes

  1. geoff-davis/framework-rocm — GitHub 2 3 4 5 6 7 8 9 10 11