Seminar Schedule
All Events
| Date / time | Event | Summary | Speakers | Materials / assignments |
|---|---|---|---|---|
Ended18:00–20:30 | Topic 1: Kernel & ML Compiler — First SeminarSession 1.0 replay · Session 1.1 replay · Session 1.2 replay | Session 1.0 From HPC to AI Infra: Parallel Computing and Parallel Programming Session 1.1 CUDA Programming Model Session…Collapse descriptionSession 1.0 From HPC to AI Infra: Parallel Computing and Parallel Programming Session 1.1 CUDA Programming Model Session 1.2 Triton/TileLang Tile Level Programming Seminar overview, curriculum design, evaluation, and computing resources. | Jiajun ChenYi ZhengYifei Wang | |
Ended19:30–21:30 | Topic 1, Session 2: Memory Abstraction & HierarchyTencent Meeting · Bilibili Live · LCPU Live | 本讲以 CUDA Core FP32 GEMM 为主线,从计算语义出发理解 data reuse 与 memory hierarchy 的抽象;从 strawman GEMM 演进到 tiling,并使用 Roofline Model 分析…Collapse description本讲以 CUDA Core FP32 GEMM 为主线,从计算语义出发理解 data reuse 与 memory hierarchy 的抽象;从 strawman GEMM 演进到 tiling,并使用 Roofline Model 分析数据复用;再从硬件视角理解 SM、warp 执行、 coalescing、latency hiding、shared memory bank conflict、padding、 swizzle 与 per-thread microtile,最后使用 Nsight Compute 完成一个 Kernel 的编写与调优。 | Yuxuan ZhouRuoyu Lin | — |
Upcoming19:00–22:00 | Topic 1, Session 3: Tensor CoreTencent Meeting · Bilibili Live · LCPU Live | 本讲以 Tensor Core 的硬件与编程模型为主线,贯穿 mma.sync(Ampere)、 wgmma(Hopper)和 tcgen05(Blackwell)三代指令的演进,从数据搬运与 发射带宽推导 Tensor Core 的设计逻…Collapse description本讲以 Tensor Core 的硬件与编程模型为主线,贯穿 mma.sync(Ampere)、 wgmma(Hopper)和 tcgen05(Blackwell)三代指令的演进,从数据搬运与 发射带宽推导 Tensor Core 的设计逻辑,理解 fragment 布局以及使用 layout 与 swizzle 解决 bank conflict 的方法,最后以 FP8 fine-grained scaling 介绍低精度 Tensor Core 的使用。 | Yuanhang Sun | — |
Calendar
The event list above will remain available.
