❌

Normale Ansicht

Received before yesterday

AMD begins to add GDDR7 support to its Linux GPU drivers

AMD has started to add support for GDDR7 memory and several new graphics IP blocks to its open-source Linux kernel driver, which indicates that software enablement for the company's next-generation standalone GPUs is underway, reports Phoronix. While AMD does not identify the upcoming architecture, the changes are likely tied to its RDNA 5 project — but they don't necessarily herald an imminent launch.

AMD's existing Radeon RX 9000-series graphics processors based on the RDNA 4 architecture use GDDR6 memory, so the addition of GDDR7 identification to drivers is arguably the most explicit confirmation that AMD is setting the stage for its next-generation discrete Radeon graphics processors based on the RDNA 5 architecture.

AMD also submitted patches that enable IH 8.0, a new version of its Interrupt Handler IP block, as well as NBIF 7.10, the latest revision of the company's New Bus Interface. Several smaller patches make additional preparations for the new hardware. These changes follow earlier Linux driver work involving Display Core Next 6 (DCN6) as well as GFX 13.0.x, suggesting that AMD is gradually upstreaming support for multiple components of an upcoming GPU architecture. Yet, while enablement of IH8, NBIF 7.10, DCN6, and GFX 13.0.x clearly point to new graphics hardware, it does not necessarily point to new discrete GPUs, unlike the mention of GDDR7.

As revealed in August, 2025, Laks Pappu, Senior Fellow at AMD, is the lead architect for AMD's next generation datacenter GPU and discrete graphics platforms. Pappu was building next-generation "competitive 2.5D/3.5D chiplet-based and monolithic graphics SoCs on various packaging technologies," according to his LinkedIn profile before it was edited to remove those disclosures.

Essentially, Pappu's profile indicated that AMD's next-generation GPU architecture could support multiple physical implementations — including multi-chiplet and monolithic — depending on performance, cost, market requirements, and AMD's willingness to compete in certain market segments

AMD has already used a multi-chiplet design with its Navi 31 GPU, an implementation that kept graphics processing hardware on a large Graphics Compute Die (GCD), but disaggregated memory interfaces and caches into smaller Memory Cache Dies (MCDs). A more ambitious implementation could potentially distribute graphics processing resources between multiple dies, although Pappu's profile did not disclose how AMD intended to partition its future GPUs.

Such an approach would be considerably more complicated than separating memory and cache functionality. Multiple compute dies would require high-bandwidth, low-latency connections as well as mechanisms for synchronization and coherency, while software would ideally continue to see the set of chiplets as a single graphics processor.

Based on conventional GPU development cycles of roughly 2.5 to 3.5 years, RDNA 5 could already have been approaching tape-out or early post-tape-out stages when the information surfaced in August 2025, which means that the company is already likely to be testing the new GPUs internally.

Officially, AMD hasn't disclosed any details about its RDNA 5 architecture nor launch schedule, so take these developments with a grain of salt. For now, all we know for sure is that AMD is working on standalone RDNA 5-based GPUs, and they're likely to use GDDR7 in at least some configurations, assuming the AI-driven memory crunch doesn't choke supply of those chips even further for consumer applications.

PlayStation 3 emulation devs work around an Nvidia bug for up to 37% faster performance

21. September 2026 um 15:00

The developers of the open-source PlayStation 3 emulator RPCS3 are celebrating finding a workaround for an Nvidia driver bug. The social media celebration stems from their testing showing that some PC configurations with Nvidia GPUs can enjoy up to 37% faster gaming performance. That’s more than a small wrinkle that’s been ironed out.

More RPCS3 performance on NVIDIA GPUs!Yahfz found a workaround for an @NVIDIA driver bug that was bottlenecking performance.In this boat test scenario:13900K + RTX 3080: +20% FPS9800X3D + RTX 5090: +25% FPSAMD GPU drivers do not have this bug, thus seeing no change. pic.twitter.com/BRad9AyesfSeptember 16, 2026

We don’t know the exact nature of the Nvidia driver bug here, but the RPCS3 devs say that it bottlenecks the performance of their PS3 emulator. The first example cited provides two PC configurations running Red Dead Redemption in a scene featuring a paddle steamer. This is a popular ‘benchmarking’ scene focused on the Blackwater docks in the game, featuring relatively consistent camera paths and water effects.

RDR boat test

Configuration

Performance uplift

System 1

13900K + RTX 3080

20% FPS

System 2

9800X3D + RTX 5090

25% FPS

Those are very gratifying performance uplifts resulting from a driver bug workaround. Plenty of readers are probably more interested if the purported bug might also be bottlenecking native PC gaming performance. If it were, that would be a hugely significant find/workaround from the RPCS3 team, specifically ‘Yahfz.’ Please note that “AMD GPU drivers do not have this bug, thus seeing no change,” say the emulator developers.

Gran Turismo 5 ran 37% faster, used far less VRAM

Follow-up social media postings underline that RDR isn’t an outlier, and “most games are impacted” by the bottleneck in the Nvidia driver. Great frame rate improvements can also be seen when using lower-end GPUs, too. The open sourcerers tested Gran Turismo 5 and saw frame rates climb by an incredible 37%, while VRAM usage dropped from 7.1 to 4.8GB. Meanwhile, playing Saints Row IV on a more mainstream GTX 1060 saw a 20% improvement in performance.

Sadly, RPCS3 dev reach-outs to Nvidia are falling on deaf ears. Another follow-up tweet states that “We tried using the developer forums in the past to no avail and gave up.” It then asks @Nvidia to establish a direct communication channel to receive driver bug reports. Nvidia Software QA boss Manuel Guzman has now replied directly, so we will see what happens in due course.

Solo dev enables running CUDA on AMD hardware in Windows, getting multiple CUDA libraries running on a gaming Radeon RX 9060 XT GPU in Windows — CUDA-exclusive workloads on AMD hardware in Windows possible without virtualization or dual-booting

14. September 2026 um 16:23

With the latest ROCm updates, AMD finally brought robust, official PyTorch and HIP SDK support to Windows for consumer GPUs, fully supporting the Radeon RX 7000 and the RX 9000 series. For native, supported frameworks, AMD on Windows is finally a viable reality, but what happens when you want to run a proprietary application, an older repository, or a specialized AI tool that absolutely refuses to support anything but NVIDIA's CUDA? That's where a new project, Speedstu's "CUDA-for-AMD-Windows," could save the day. It proves that running rigidly CUDA-exclusive workloads on AMD hardware in Windows is possible without virtualization or dual-booting.

To be clear, this project is not a brand-new runtime. Instead, it is a highly automated and reproducible PowerShell setup that bridges the gap between ZLUDA, the well-known, formerly AMD-funded translation layer, and AMD's native HIP/ROCm SDK for Windows. Through a series of clever scripts, the toolkit automatically detects the user's GPU architecture, grabs a specifically pinned version of ZLUDA (v6-preview.69), and, at least in theory, seamlessly maps it to the ROCm math libraries already present in Windows.

A screenshot of the CUDA for AMD Windows GitHub documentation.

Several important CUDA libraries link up, but the important cuDNN doesn't work yet. (Image credit: Speedstu/GitHub)

The result is that the developer successfully intercepted and mapped the CUDA driver API as well as the cuBLAS, cuSPARSE, and cuFFT libraries directly over to their AMD equivalents. As a proof-of-concept, the author even trained a 2.2-million-parameter PPO reinforcement-learning network end-to-end using unmodified CUDA libraries on an AMD Radeon RX 9060 XT, which happens to be the only officially supported GPU at this time.

Even with AMD's official ROCm support on Windows, the local developer community frequently runs into dependency hell when trying out experimental GitHub repos or specialized AI tools that hardcode CUDA as a requirement. For developers who want to experiment with these CUDA-only tools natively on their Windows daily driver machines without dealing with WSL2 passthrough issues or waiting for the original author to write a HIP port, this project offers a highly desirable translation pipeline. It acts as a sort of hacky adapter for software that stubbornly demands an NVIDIA card.

Benchmark testing included in the project's documentation offers some interesting findings. In a controlled A/B test running a 2.2M-parameter reinforcement learning workload on a Radeon RX 9060 XT, the "public upstream path," which relies purely on official ZLUDA releases and AMD's stock HIP SDK 6.4, achieved a median throughput of 13,278 steps per second (SPS). By contrast, an optional "recovered custom overlay" apparently built from salvaged legacy ZLUDA binaries ran slightly worse at 12,876 SPS, making it roughly 3% slower. While the clean official setup is faster, the author notes that "a later rewrite removed LibTorch/ZLUDA from PPO and achieved substantially higher throughput," indicating that there is still a performance hit for this stack of translators.

CUDA for AMD Windows Official benchmarks (BENCHMARKS.md)

Metric

Public upstream

Recovered custom

Custom delta

Overall SPS, median

13,278.46

12,875.80

-3.03%

Overall SPS, mean

13,172.49

12,649.83

-3.97%

Collection SPS, median

63,306.00

59,360.67

-6.23%

Consumption SPS, median

16,806.36

16,445.66

-2.15%

Inference time, median

0.5863 s

0.6293 s

+7.33%

PPO learn time, median

3.2076 s

3.2958 s

+2.75%

Now, before anyone declares the CUDA moat officially drained, it is crucial to set realistic expectations. First, this is a solo open-source project, not an enterprise-grade solution. The author is extremely transparent about its narrow scope, as crucial machine learning libraries like cuDNN, TensorRT, and NCCL do not resolve yet. This means compatibility is strictly workload-dependent; if your specific AI tool relies heavily on cuDNN, this setup will fail. Furthermore, it's worth pointing out that ZLUDA itself is currently being maintained as a "weekend hobby project" after losing its commercial backing a second time. Relying on this pipeline for production-level work remains a massive risk. This repository is a tinkerer's tool, not a corporate IT deployment strategy.

Despite these limitations, "CUDA-for-AMD-Windows" is pretty exciting, as it proves that the barrier to entry for running CUDA-exclusive software on AMD GPUs isn't an insurmountable hardware flaw, but a relatively tractable translation tooling problem. Because the project is entirely open-source, its potential extends far beyond this initial proof-of-concept. With community contributions, we could see expanded hardware detection and clever patches to get more stubborn CUDA libraries resolving properly.

❌