ARASHI

Architecture Reconfiguration for Accelerating Sparse-tensor Hardware Implementations

Keep the contraction. Specialize the hardware.

ARASHI investigates sparse-aware hardware compilation. Our 2026 prototype uses sparse profiles to specialize generated HLS from MLIR-based toolchains, preserving the contraction while changing the hardware architecture. Compiler-native analysis and automatic selection are the next research step.

From 2025 feasibility to 2026 optimization

From algorithm-to-layout feasibility to compiler-native sparse architecture selection.

2025

Algorithm-to-layout feasibility

Investigate model-to-layout feasibility, document ASIC-flow barriers and FPGA evaluation, and establish the next toolchain direction.

2026

Sparse optimization experiments

Check sparse structural preconditions, compare generated-HLS transformations, and measure architecture trade-offs across two lowering paths.

2027

Generalization and scale

Extend automatic selection across contraction families, sparse formats, toolchains, and parallel memory systems.

The path

1Tensor contraction
2LLVM / MLIR
3Backend lowering
4Generated HLS
5Sparse specialization
6Vitis / RTL

The experimental pipeline preserves the contraction and specializes generated HLS using checked sparse structure. Bringing these transformations into common MLIR passes is the next integration step.

Cross-toolchain validation

HIDA and STAR Lab's independent implementation of StreamTensor's methodology provide two MLIR-based lowering paths. ARASHI applies profile-guided transformations around their generated HLS. A common AXI boundary and Vitis backend support a controlled transfer study; native MLIR integration remains future work.

2026 · current work

Does sparse-aware compiler optimization improve synthesized hardware?

Across two independent lowering paths and two distinct bottlenecks, selected designs reduce HLS cycle estimates by 10.01–11.31× while preserving every FP32 output.

Physical implementation5 ns

baseline and selected E2/S1 IP meets the U280 out-of-context timing target in both lowering paths.

C++ oracle checks and OOC routing; board execution remains open

Controlled comparison

Minimally enabled baseline versus ARASHI selection

Identical contraction, inputs, AXI interface, target device, and full-output oracle; the experimental hardware transformations change.

HLS latencyAverage cycle estimate from Vitis HLS synthesis; lower is better.
TIDES E2 fbag,bhai→fghi10.01× cycle reduction
Baseline572,136,789
Optimized57,154,011
TIDES S1 dae,fga→defg11.31× cycle reduction
Baseline144,140,318
Optimized12,748,931
Initiation intervalCycles between successive loop-iteration launches; lower is better.
TIDES E2 fbag,bhai→fghi10× II reduction
BaselineII 20
OptimizedII 2
TIDES S1 dae,fga→defg14× II reduction
BaselineII 56
OptimizedII 4

Each workload is normalized to its own minimally enabled, correct, synthesizable baseline. Contraction, inputs, AXI boundary, U280 target, and full-output oracle are identical; values are Vitis HLS estimates, not FPGA board measurements.

Presentation archive

One research story, preserved by year.