Daftar Isi
1. Fix SageAttention in Linux and WSL (Part 1)
This is part 1 of 2. This video is for SageAttention on a 4090. See the next video for Conda and SageAttention 3 on a 5090.
2. CUDA vs ROCm: Why ROCm Sucks at Every Step
Why does AMD's flagship MI355X carry 288GB of VRAM and match Nvidia's B200 on paper, only to lose by 33% on real ...
3. This TPU is 4500% Faster Than A GPU
Google's TPU was built specifically for machine learning, but does that actually make it a GPU killer? In this video, we break down ...
4. Colibrì vs llama.cpp: Running DeepSeek V4 284B on CPU
Run 360GB models on 8GB RAM using Colibri technology. This approach challenges NVIDIA dominance in local LLM inference.
5. Running Gemma4 31B FP8 and NVFP4 with MTP locally on an RTX PRO 6000 Blackwell
Additional SGLang version of this benchmark: youtube.com/watch?v=YlysBl5Qg34 The SGLang run produced ...
6. Nemotron 3 Super 120B NVFP4 on RTX PRO 6000: 150 tok/s with vLLM MTP
Running NVIDIA Nemotron 3 Super 120B A12B NVFP4 locally on a single RTX PRO 6000 Blackwell 96GB GPU with vLLM 0.21.0 ...
7. GPU Coding: PyTorch, torch.compile, CUDA, Triton | Build Your Own LLM Workshop #5 [Refreshed]
GPU Performance for LLMs: PyTorch, torch.compile(), Fused Kernels, CUDA & Triton, and benchmarking. Part of a Build your own ...
8. Spectral Compute: Compile CUDA everywhere
Chris Kitching, CTO at Spectral Compute, explains how SCALE recompiles unmodified CUDA source to native AMD and NVIDIA ...
9. This AI Agent Runs My Supercomputer While I Go Running
Turn your GitHub repo into a research robot! In this Vibe Coding 101 video, I show you how to run Claude Code agents inside ...
10. Long Context Training and Inference on AMD GPUs
Mehdi Rezagholizadeh, Principal Research Scientist, AMD About the Speaker: Mehdi Rezagholizadeh is a Principal Member of ...
11. How to Fine-Tune LLMs for AI Agents: SFT and QLoRA Guide using Unsloth
I fine-tuned a small open-source LLM (Qwen2.5-1.5B) into a custom AI agent using SFT and QLoRA — then ran it fully offline on ...
12. Qdrant Edge | Edge Shard from Collection | Rust App | iced-rs
A small iced (0.14) desktop GUI that does what the Qdrant Edge docs describe by hand: Download a shard snapshot from a ...
13. 5 PM PKT | Fundamentals of System Design | Week 4 | Day 2
5 PM PKT | Fundamentals of System Design | Week 4 | Day 2.
14. Qwen3.6-27B NVFP4+MTP vLLM Benchmark TG 190tok/s — RTX PRO 6000 Blackwell Max-Q x 2
Running Qwen3.6-27B-Text-NVFP4-MTP on vLLM v0.19.2rc1 with MTP speculative decoding (num_speculative_tokens=3) on 2x ...
15. [PLDI'26] [TOPLAS] StreamAlloc: A Framework for Analyzing and Transforming CUDA Code to Enable(…)
[TOPLAS] StreamAlloc: A Framework for Analyzing and Transforming CUDA Code to Enable Asynchronous Execution (Video, ...
Torch_cuda_arch_list Information Guide
About of Torch_cuda_arch_list

Main Features

Recent Updates

Detailed Analysis
Data is compiled from public records and verified media reports.
Last Updated: August 11, 2026
Final Thoughts

Disclaimer: Disclaimer: All information is compiled from publicly available data, media reports, and analysis. Actual details may vary.


![GPU Coding: PyTorch, torch.compile, CUDA, Triton | Build Your Own LLM Workshop #5 [Refreshed]](https://i.ytimg.com/vi/W0mCz4bTHSU/mqdefault.jpg)







![[PLDI'26] [TOPLAS] StreamAlloc: A Framework for Analyzing and Transforming CUDA Code to Enable(…)](https://i.ytimg.com/vi/aeF1q_VmtxE/mqdefault.jpg)