Skip to content

WebGPU Micro-Benchmark Result

Per-operation GPU timings for the profiled device. Benchmark details available in the webgpu-bench open source project.

Summary

Average Bandwidth
—
FP32 FLOPS
—
FP16 FLOPS
—
INT8 OPS
—

Device

Recorded 9/8/2026, 3:52:36 PM

WebGPU device
Vendorapple
Device—
Architecturemetal-3
Description—
WebGL device
VendorGoogle Inc. (Apple)
RendererANGLE (Apple, ANGLE Metal Renderer: Apple M3, Unspecified Version)
Browser
NameChrome
Version152.0.0.0
Operating system
NameMac OS
Version10.15.7

Benchmark Results

Throughput from this run, in bytes or operations per second. Higher is faster.

BenchmarkThroughput
atomic direct
—
atomic sharded
—
atomic workgroup
—
branch coherent
—
branch divergent
—
branch none
—
branch uniform
—
branch vec4 if
—
branch vec4 select
—
f32<->f16 convert
—
fp16 div
—
fp16 ln
—
fp16 mat4 FMA
—
fp16 matvec FMA
—
fp16 pow
—
fp16 rsqrt
—
fp16 scalar FMA
—
fp16 sin/cos
—
fp16 sqrt
—
fp16 vec4 FMA
—
fp32 clamp
—
fp32 div
—
fp32 fma() builtin
—
fp32 ln
—
fp32 mat4 FMA
—
fp32 matvec FMA
—
fp32 min/max
—
fp32 pow
—
fp32 rsqrt
—
fp32 scalar FMA
—
fp32 select
—
fp32 sin/cos
—
fp32 sqrt
—
fp32 vec4 FMA
—
i32 div
—
i32 mat4 multiply-add
—
i32 matvec multiply-add
—
i32 scalar multiply-add
—
i32 vec4 multiply-add
—
i32<->f32 convert
—
int8 dp4a
—
int8 dp4a matvec
—
layout aos
—
layout aosoa
—
layout soa
—
read dependent chain
—
read gather 16kb
—
read gather 4mb
—
read gather 64mb
—
read linear
—
reduction serial
—
reduction subgroup
—
reduction workgroup
—
texture interp (built-in)
—
texture interp (manual)
—
tile direct
—
tile shared
—
tile shared padded
—
u32 byte pack/unpack
—
u32 countOneBits
—
u32 firstLeadingBit
—
u32 reverseBits
—
u32 variable shift
—
uint8 dp4a
—
workgroup 128
—
workgroup 256
—
workgroup 64
—
write linear
—
write scatter 16kb
—
write scatter 4mb
—
write scatter 64mb
—

Recent Results

Loading Recent Results