Per-category breakdown with percentages — identifies compilation instruction bottlenecks
Category
LLVM
LLVM %
ScratchV
SV %
SV/LLVM
ALU R-type
131,372,739
7.1%
929,645,352
32.0%
7.08×
ALU I-type
524,713,477
28.4%
922,090,821
31.7%
1.76×
FP
524,722,184
28.4%
—
0.0%
0.00×
Shift
—
0.0%
255,137,123
8.8%
—
Load
527,439,523
28.6%
528,104,451
18.2%
1.00×
Store
3,053,618
0.2%
6,830,882
0.2%
2.24×
Branch
80,766,507
4.4%
265,412,804
9.1%
3.29×
Jump
25,779,684
1.4%
—
0.0%
0.00×
Upper immediate
27,157,604
1.5%
—
0.0%
0.00×
Total
1,845,005,337
100%
2,907,221,433
100%
1.58×
Biggest bottleneck: ALU R-type (7.08x vs LLVM). "
Store ratio 2.24×: LLVM keeps accumulators in FP registers (few stores). ScratchV spills to stack every MAC due to limited registers.
3. Operator Comparison · 算子粒度
Per-operator-type dynamic instruction ratio (ScratchV / LLVM) — identifies which operator types have the most optimization headroom
Per-Operator Instruction Type Breakdown
Category distribution (% of total dynamic instructions) per operator type — compare compiler instruction mix patterns
conv (3 ops, ratio: 1.14x)
Category
SV %
SV bar
LLVM %
LLVM bar
ALU I
31.3%
28.6%
FP
0.0%
28.6%
Load
25.0%
28.6%
ALU R
25.0%
7.1%
Shift
12.5%
0.0%
Branch
3.7%
4.3%
Jump
1.2%
1.4%
Upper
1.2%
1.4%
gemm (2 ops, ratio: 1.38x)
Category
SV %
SV bar
LLVM %
LLVM bar
ALU I
33.3%
23.1%
FP
0.0%
30.8%
Load
22.2%
30.8%
ALU R
22.2%
7.7%
Shift
11.1%
0.0%
Branch
5.6%
4.6%
Upper
3.3%
3.1%
Jump
2.2%
0.0%
maxpool (3 ops, ratio: 0.96x)
Category
SV %
SV bar
LLVM %
LLVM bar
ALU I
34.8%
33.3%
Load
26.1%
25.0%
Branch
13.0%
12.5%
ALU R
8.7%
8.3%
Store
8.7%
4.2%
FP
0.0%
8.3%
Jump
4.3%
4.2%
Upper
4.3%
4.2%
relu (4 ops, ratio: 1.88x)
Category
SV %
SV bar
LLVM %
LLVM bar
ALU I
26.7%
25.0%
Load
26.7%
25.0%
Store
13.3%
25.0%
Branch
13.3%
12.5%
FP
0.0%
12.5%
ALU R
6.7%
0.0%
Jump
6.7%
0.0%
Upper
6.7%
0.0%
sigmoid (1 ops, ratio: 0.72x)
Category
SV %
SV bar
LLVM %
LLVM bar
FP
0.0%
32.0%
ALU I
27.8%
20.0%
ALU R
16.7%
12.0%
Branch
16.7%
8.0%
Load
11.1%
12.0%
Shift
11.1%
0.0%
Store
5.6%
8.0%
Jump
5.6%
4.0%
Upper
5.6%
4.0%
Orange = ScratchV dominant · Green = SV lower · Blue = LLVM. Longer bar = higher % of total instructions.
Model
Op Type
SV Static
SV Dynamic
LLVM Dynamic
Ratio
cnn_fc1_Gemm
gemm
87
62,005,247
44,781,567
1.38x
cnn_fc2_Gemm
gemm
80
1,151
831
1.39x
cnn_layer1.0_Conv
conv
165
425,115,646
371,976,190
1.14x
cnn_layer1.1_Relu
relu
56
14,760,960
7,872,512
1.88x
cnn_layer1.2_MaxPool
maxpool
97
5,658,368
5,904,384
0.96x
cnn_layer2.0_Conv
conv
165
1,097,367,551
960,196,607
1.14x
cnn_layer2.1_Relu
relu
56
3,572,160
1,905,152
1.88x
cnn_layer2.2_MaxPool
maxpool
97
1,369,328
1,428,864
0.96x
cnn_layer3.0_Conv
conv
165
513,294,335
449,132,543
1.14x
cnn_layer3.1_Relu
relu
56
1,670,880
891,136
1.88x
cnn_layer3.2_MaxPool
maxpool
96
618,976
645,888
0.96x
cnn_relu1_Relu
relu
54
960
512
1.88x
cnn_sigmoid1_Sigmoid
sigmoid
69
18
25
0.72x
4. Optimization Timeline · 优化时间线
Click to expand each version — instruction category breakdown with percentages and per-operator comparison