Skip to content

Seminar Schedule

Download course calendar (.ics)

All Events

Date / timeEventSummarySpeakersMaterials / assignments
Ended18:00–20:30
Topic 1: Kernel & ML Compiler — First SeminarSession 1.0 replay · Session 1.1 replay · Session 1.2 replay
Session 1.0 From HPC to AI Infra: Parallel Computing and Parallel Programming Session 1.1 CUDA Programming Model Session…Show full descriptionCollapse description

Session 1.0 From HPC to AI Infra: Parallel Computing and Parallel Programming Session 1.1 CUDA Programming Model Session 1.2 Triton/TileLang Tile Level Programming Seminar overview, curriculum design, evaluation, and computing resources.

Jiajun ChenYi ZhengYifei Wang
Ended19:30–21:30
Topic 1, Session 2: Memory Abstraction & HierarchyTencent Meeting · Bilibili Live · LCPU Live
本讲以 CUDA Core FP32 GEMM 为主线,从计算语义出发理解 data reuse 与 memory hierarchy 的抽象;从 strawman GEMM 演进到 tiling,并使用 Roofline Model 分析…Show full descriptionCollapse description

本讲以 CUDA Core FP32 GEMM 为主线,从计算语义出发理解 data reuse 与 memory hierarchy 的抽象;从 strawman GEMM 演进到 tiling,并使用 Roofline Model 分析数据复用;再从硬件视角理解 SM、warp 执行、 coalescing、latency hiding、shared memory bank conflict、padding、 swizzle 与 per-thread microtile,最后使用 Nsight Compute 完成一个 Kernel 的编写与调优。

Yuxuan ZhouRuoyu Lin
Upcoming19:00–22:00
Topic 1, Session 3: Tensor CoreTencent Meeting · Bilibili Live · LCPU Live
本讲以 Tensor Core 的硬件与编程模型为主线,贯穿 mma.sync(Ampere)、 wgmma(Hopper)和 tcgen05(Blackwell)三代指令的演进,从数据搬运与 发射带宽推导 Tensor Core 的设计逻…Show full descriptionCollapse description

本讲以 Tensor Core 的硬件与编程模型为主线,贯穿 mma.sync(Ampere)、 wgmma(Hopper)和 tcgen05(Blackwell)三代指令的演进,从数据搬运与 发射带宽推导 Tensor Core 的设计逻辑,理解 fragment 布局以及使用 layout 与 swizzle 解决 bank conflict 的方法,最后以 FP8 fine-grained scaling 介绍低精度 Tensor Core 的使用。

Yuanhang Sun

Calendar

LectureGuest LectureAssignment deadline
Loading the interactive calendar…

The event list above will remain available.