concept-collection / matmul-bench
matmul-bench
Go to fileHistoryFork
.githubBuild libFLAME/BLIS from source in CI instead of vendoring binaries
blisBuild libFLAME/BLIS from source in CI instead of vendoring binaries
nativematmul-bench: browser GEMM benchmark (JS, WebGPU, custom C WASM, libFLAME/BLIS WASM)
publicBuild libFLAME/BLIS from source in CI instead of vendoring binaries
srcBuild libFLAME/BLIS from source in CI instead of vendoring binaries
wasmmatmul-bench: browser GEMM benchmark (JS, WebGPU, custom C WASM, libFLAME/BLIS WASM)
.gitignoreBuild libFLAME/BLIS from source in CI instead of vendoring binaries
.oxlintrc.jsonmatmul-bench: browser GEMM benchmark (JS, WebGPU, custom C WASM, libFLAME/BLIS WASM)
CLAUDE.mdBuild libFLAME/BLIS from source in CI instead of vendoring binaries
index.htmlmatmul-bench: browser GEMM benchmark (JS, WebGPU, custom C WASM, libFLAME/BLIS WASM)
package-lock.jsonmatmul-bench: browser GEMM benchmark (JS, WebGPU, custom C WASM, libFLAME/BLIS WASM)
package.jsonmatmul-bench: browser GEMM benchmark (JS, WebGPU, custom C WASM, libFLAME/BLIS WASM)
README.mdBuild libFLAME/BLIS from source in CI instead of vendoring binaries
tsconfig.app.jsonmatmul-bench: browser GEMM benchmark (JS, WebGPU, custom C WASM, libFLAME/BLIS WASM)
tsconfig.jsonmatmul-bench: browser GEMM benchmark (JS, WebGPU, custom C WASM, libFLAME/BLIS WASM)
tsconfig.node.jsonmatmul-bench: browser GEMM benchmark (JS, WebGPU, custom C WASM, libFLAME/BLIS WASM)
vite.config.tsmatmul-bench: browser GEMM benchmark (JS, WebGPU, custom C WASM, libFLAME/BLIS WASM)

matmul-bench#

Compares matrix-matrix multiply (GEMM) performance across implementations running in the browser, plus a native reference measured outside it.

Live: https://concept-collection.github.io/matmul-bench/

Methods:

  • JavaScript/TypeScript — plain triple loop over Float64Array (f64), no special-casing. The "no special effort" baseline.
  • WebGPU — a tiled compute shader (f32; WGSL has no double precision).
  • WASM from custom C — naive, blocked+SIMD (f64x2, -msimd128), and a threaded (WASM pthreads) version of the blocked+SIMD kernel, built with emscripten (f64). The threaded one needs cross-origin isolation.
  • libFLAME/BLIS in WASM — a real BLAS (dgemm via libFLAME + BLIS) compiled to WebAssembly, single-threaded and multi-threaded (WASM pthreads), from libflame2wasm (f64). The threaded build needs cross-origin isolation (SharedArrayBuffer), supplied on GitHub Pages by a vendored service worker.
  • Native LAPACK/OpenBLAS — a fixed reference table from dgemm run outside the browser (native/).

Develop#

npm install
npm run dev            # local dev server (sends COOP/COEP for threaded methods)
wasm/build-wasm.sh      # rebuild wasm/dist/matmul.* + public/matmul/matmul_mt.*
                        #   (needs emsdk)

# One-time: build the WASM libflame/BLIS libraries in a sibling checkout
git clone https://github.com/concept-collection/libflame2wasm ../libflame2wasm
(cd ../libflame2wasm && ./build-all-wasm.sh)        # slow; builds libflame twice

blis/build-blis-wasm.sh # rebuild public/matmul/matmul_blis_{st,mt}.*
                        #   (needs emsdk + the sibling libflame2wasm above;
                        #    override its path with LIBFLAME=/path/to/libflame2wasm)
npm run build           # tsc -b && vite build

Nothing prebuilt is committed. The WASM modules under public/matmul/ (except worker.js) are built by CI: the custom kernels from wasm/matmul.c, and the libFLAME/BLIS modules by cloning libflame2wasm as a sibling, running its build-all-wasm.sh (cached, keyed on its commit SHA), and linking against the resulting .a libraries.

Native reference#

native/build.sh
OPENBLAS_NUM_THREADS=1  native/bench_native 128 256 512 1024 2048
OPENBLAS_NUM_THREADS=12 native/bench_native 128 256 512 1024 2048

Hand-copy the results into src/data/nativeReference.ts.