WebGPU Micro-Benchmark Result
Per-operation GPU timings for the profiled device. Benchmark details available in the webgpu-bench open source project.
Device
Recorded 9/14/2026, 2:02:27 PM
WebGPU device
Vendorapple
Deviceapple
Architectureapple
Descriptionapple
WebGL device
VendorApple Inc.
RendererApple GPU
Browser
NameSafari
Version27.0
Operating system
NameMac OS
Version10.15.7
Benchmark Results
Throughput from this run, in bytes or operations per second. Higher is faster.
| Benchmark | Throughput |
|---|---|
atomic direct | 1 Gupdate/s |
atomic sharded | 3.44 Gupdate/s |
atomic workgroup | 6.1 Gupdate/s |
branch coherent | 1.69 TFLOP/s |
branch divergent | 954 GFLOP/s |
branch none | 2.7 TFLOP/s |
branch uniform | 1.76 TFLOP/s |
branch vec4 if | 300 Gchoice/s |
branch vec4 select | 273 Gchoice/s |
f32<->f16 convert | 2.45 TOP/s |
fp16 div | 828 GFLOP/s |
fp16 ln | 818 GFLOP/s |
fp16 mat4 FMA | 1.62 TFLOP/s |
fp16 matvec FMA | 2.01 TFLOP/s |
fp16 pow | 412 GFLOP/s |
fp16 rsqrt | 838 GFLOP/s |
fp16 scalar FMA | 2.78 TFLOP/s |
fp16 sin/cos | 220 GFLOP/s |
fp16 sqrt | 828 GFLOP/s |
fp16 vec4 FMA | 3.09 TFLOP/s |
fp32 clamp | 1.34 TOP/s |
fp32 div | 828 GFLOP/s |
fp32 fma() builtin | 2.7 TFLOP/s |
fp32 ln | 838 GFLOP/s |
fp32 mat4 FMA | 1.46 TFLOP/s |
fp32 matvec FMA | 1.61 TFLOP/s |
fp32 min/max | 2.65 TOP/s |
fp32 pow | 412 GFLOP/s |
fp32 rsqrt | 838 GFLOP/s |
fp32 scalar FMA | 2.7 TFLOP/s |
fp32 select | 962 GOP/s |
fp32 sin/cos | 177 GFLOP/s |
fp32 sqrt | 828 GFLOP/s |
fp32 vec4 FMA | 2.3 TFLOP/s |
i32 div | 153 GOP/s |
i32 mat4 multiply-add | 819 GOP/s |
i32 matvec multiply-add | 799 GOP/s |
i32 scalar multiply-add | 838 GOP/s |
i32 vec4 multiply-add | 842 GOP/s |
i32<->f32 convert | 419 GOP/s |
int8 dp4a | 330 GOP/s |
int8 dp4a matvec | 505 GOP/s |
layout aos | 1.61 Gparticle/s |
layout aosoa | 4.88 Gparticle/s |
layout soa | 4.71 Gparticle/s |
read dependent chain | 6.39 Mhop/s |
read gather 16kb | 701 GB/s |
read gather 4mb | 64.1 GB/s |
read gather 64mb | 23.8 GB/s |
read linear | 90.9 GB/s |
reduction serial | 1.92 Gelement/s |
reduction subgroup | — |
reduction workgroup | 4.55 Gelement/s |
texture interp (built-in) | 471 GB/s |
texture interp (manual) | 537 GB/s |
tile direct | 4.97 Goutput/s |
tile shared | 4.19 Goutput/s |
tile shared padded | 4.13 Goutput/s |
u32 byte pack/unpack | 609 GOP/s |
u32 countOneBits | 414 GOP/s |
u32 firstLeadingBit | 300 GOP/s |
u32 reverseBits | 419 GOP/s |
u32 variable shift | 414 GOP/s |
uint8 dp4a | 790 GOP/s |
workgroup 128 | 4.88 Gelement/s |
workgroup 256 | 4.97 Gelement/s |
workgroup 64 | 4.97 Gelement/s |
write linear | 81 GB/s |
write scatter 16kb | 8.08 GB/s |
write scatter 4mb | 40.5 GB/s |
write scatter 64mb | 18.2 GB/s |