Skip to content

WebGPU Micro-Benchmark Result

Per-operation GPU timings for the profiled device. Benchmark details available in the webgpu-bench open source project.

Summary

Average Bandwidth
86 GB/s
FP32 FLOPS
2.7 TFLOP/s
FP16 FLOPS
3.09 TFLOP/s
INT8 OPS
842 GOP/s

Device

Recorded 9/14/2026, 2:02:27 PM

WebGPU device
Vendorapple
Deviceapple
Architectureapple
Descriptionapple
WebGL device
VendorApple Inc.
RendererApple GPU
Browser
NameSafari
Version27.0
Operating system
NameMac OS
Version10.15.7

Benchmark Results

Throughput from this run, in bytes or operations per second. Higher is faster.

BenchmarkThroughput
atomic direct
1 Gupdate/s
atomic sharded
3.44 Gupdate/s
atomic workgroup
6.1 Gupdate/s
branch coherent
1.69 TFLOP/s
branch divergent
954 GFLOP/s
branch none
2.7 TFLOP/s
branch uniform
1.76 TFLOP/s
branch vec4 if
300 Gchoice/s
branch vec4 select
273 Gchoice/s
f32<->f16 convert
2.45 TOP/s
fp16 div
828 GFLOP/s
fp16 ln
818 GFLOP/s
fp16 mat4 FMA
1.62 TFLOP/s
fp16 matvec FMA
2.01 TFLOP/s
fp16 pow
412 GFLOP/s
fp16 rsqrt
838 GFLOP/s
fp16 scalar FMA
2.78 TFLOP/s
fp16 sin/cos
220 GFLOP/s
fp16 sqrt
828 GFLOP/s
fp16 vec4 FMA
3.09 TFLOP/s
fp32 clamp
1.34 TOP/s
fp32 div
828 GFLOP/s
fp32 fma() builtin
2.7 TFLOP/s
fp32 ln
838 GFLOP/s
fp32 mat4 FMA
1.46 TFLOP/s
fp32 matvec FMA
1.61 TFLOP/s
fp32 min/max
2.65 TOP/s
fp32 pow
412 GFLOP/s
fp32 rsqrt
838 GFLOP/s
fp32 scalar FMA
2.7 TFLOP/s
fp32 select
962 GOP/s
fp32 sin/cos
177 GFLOP/s
fp32 sqrt
828 GFLOP/s
fp32 vec4 FMA
2.3 TFLOP/s
i32 div
153 GOP/s
i32 mat4 multiply-add
819 GOP/s
i32 matvec multiply-add
799 GOP/s
i32 scalar multiply-add
838 GOP/s
i32 vec4 multiply-add
842 GOP/s
i32<->f32 convert
419 GOP/s
int8 dp4a
330 GOP/s
int8 dp4a matvec
505 GOP/s
layout aos
1.61 Gparticle/s
layout aosoa
4.88 Gparticle/s
layout soa
4.71 Gparticle/s
read dependent chain
6.39 Mhop/s
read gather 16kb
701 GB/s
read gather 4mb
64.1 GB/s
read gather 64mb
23.8 GB/s
read linear
90.9 GB/s
reduction serial
1.92 Gelement/s
reduction subgroup
reduction workgroup
4.55 Gelement/s
texture interp (built-in)
471 GB/s
texture interp (manual)
537 GB/s
tile direct
4.97 Goutput/s
tile shared
4.19 Goutput/s
tile shared padded
4.13 Goutput/s
u32 byte pack/unpack
609 GOP/s
u32 countOneBits
414 GOP/s
u32 firstLeadingBit
300 GOP/s
u32 reverseBits
419 GOP/s
u32 variable shift
414 GOP/s
uint8 dp4a
790 GOP/s
workgroup 128
4.88 Gelement/s
workgroup 256
4.97 Gelement/s
workgroup 64
4.97 Gelement/s
write linear
81 GB/s
write scatter 16kb
8.08 GB/s
write scatter 4mb
40.5 GB/s
write scatter 64mb
18.2 GB/s

Recent Results

Loading Recent Results