Keyboard shortcuts

Press or to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Per-SM Scheduling Model — Sample Tables

Addresses apply to ptxas v13.0.88 (CUDA 13.0). VA base 0x400000 (non-PIE).

Representative slices of the three per-SM scheduling table families. The full per-SM TSVs (11 dependency-rule sets, 3 latency families, 7 scoreboard configs) live in the repo at decoded/ptxas-sched-full/. See Latency Model and Scoreboards for the field semantics.

Dependency rule (40 B) — sm_90 vs sm_90a (first 12 classes + the WGMMA split)

rule_type 4 = disabled/unit-absent (always paired with latency=255). The six classes 41, 561, 562, 563, 566, 567 are disabled on sm_90 but active on sm_90a — the WGMMA / async-MMA tensor classes. This is the entire sm_90 vs sm_90a divergence.

smidxunitrulelattput_invbar_latbar_tputrd_latwr_latstallslots
sm_90021170568-1-101
sm_90131170568-1-1332
sm_902414215564-1-111
sm_903514215564-1-111
sm_904614215564-1-111
sm_905714215564-1-123
sm_906814215564-1-123
sm_907914215564-1-111
sm_9081014215564-1-141
sm_909111170568-1611
sm_9010121170568-1-111
sm_9011150222562-1-1394
sm_90334142553556-1-1-1394
sm_9017156142553556-1-1-1394
sm_9017256242553556-1-1-1394
sm_9017356342553556-1-1-1394
sm_9017456642553556-1-1-1394
sm_9017556742553556-1-1-1394
sm_90a021170568-1-101
sm_90a131170568-1-1332
sm_90a2414215564-1-111
sm_90a3514215564-1-111
sm_90a4614215564-1-111
sm_90a5714215564-1-123
sm_90a6814215564-1-123
sm_90a7914215564-1-111
sm_90a81014215564-1-141
sm_90a9111170568-1611
sm_90a10121170568-1-111
sm_90a11150222562-1-1394
sm_90a334114215564-1-1332
sm_90a17156101519178-1-1394
sm_90a17256201519178-1-1394
sm_90a173563014191816-1-1394
sm_90a17456601519178-1-1394
sm_90a175567014191816-1-1394

Latency / sched-class descriptor (72 B) — sm_8x family (first 16 classes)

p7_self equals class_id in every record (self-reference). p1_tput = throughput class, p5_maxstall = max-stall cycles, p11 always 0. pipeA/pipeB are per-pipe eligibility byte vectors (0xFF byte = pipe N/A).

idxclass_idpipeA_hexpipeB_hexp0_flagsp1_tputp2p3p4p5_maxstallp6p7_selfp8p9p10p11
020303ffffffffffff0000ffffffff0000040107321230
130303ffffffffffff0000ffffffff00000433207331230
2404040202ffffffff0000ffffffff0000041137341230
3504040202ffffffff0000ffffffff000026843545641137351230
4604040202ffffffff0000ffffffff000026843545641137361230
5704040202ffff01010000ffffffff0000042337371230
6804040202ffff01010000ffffffff0000042337381230
7904040202ffffffff0000ffffffff0000041137391230
8100404ffffffff01010000ffffffff00000441373101210
9110404ffffffff01010000ffffffff0000276824064411373111210
10120404ffffffff01010000ffffffff0000268439552411373121210
11150000ffffffffffff0000ffffffff0000003941932150110
12161212ffffffffffff00000303ffff0000041311461163110
13171212ffffffffffff00000303ffff0000041311461173110
14181313ffffffffffff00000303ffff0000041311461183110
15191111ffffffffffff00000303ffff0000041311461193110

Scoreboard config — sm_100 (≤6 triplets) vs sm_90 (single triplet)

Each config is up to 6 {sb_id, threshold, mask} triplets. mask = -1 = unconditional wait; small masks (2/4/8/32) = pipe-specific. threshold 56 dominates. sm_100 (Blackwell) uses up to 6 scoreboards per config (async dependency-barrier model); sm_90 uses one triplet with a pipe mask.

sm_100 (first 5 configs)

idxtriplet_countsb_idthresholdmask
02056-1
02256-1
13056-1
13256-1
13556-1
24056-1
24256-1
24556-1
241556-1
36056-1
36256-1
361256-1
361556-1
362356-1
363456-1
43056-1
43256-1
431556-1

sm_90 (first 10 configs)

idxtriplet_countsb_idthresholdmask
010568
112562
215564
316568
411282
511291
6112124
7112132
8113568
9115564