perf(importer): decode glTF images in parallel #834

Merged
SakulFlee merged 1 commit from perf/gltf-import-parallel-decode into main 2026-09-29 17:01:44 +00:00
Owner

Image decoding dominates glTF import time. Measured on DamagedHelmet.glb
(5 x 2048x2048 textures), decoding accounts for ~99% of the import: ~97 ms
of decode versus ~0.36 ms for the document parse and ~0.25 ms for all of
parse_texture/parse_dual_texture combined. The decode loop was strictly
sequential, so multi-core machines left most of their capacity idle during
imports.

Split the loop in two. The first stage gathers each image's encoded bytes
and stays sequential: it is I/O and bookkeeping, and keeping buffer-view
images borrowed via Cow avoids copying the encoded payload. The second stage
runs image::load_from_memory over rayon::par_iter, which needs no
synchronization because each task reads only its own slice and returns a
value.

Measured end-to-end on a 6-core machine: ~1.05-1.14x faster (min of 40 runs,
109.5 ms -> 96.0 ms). The decode stage in isolation is ~1.27-1.38x faster.
The end-to-end figure is lower than the isolated one because Amdahl's law
caps the win: on this asset a single 2048x2048 texture takes ~66 ms of the
~97 ms decode, so that one image sets the critical path no matter how many
threads are available. Assets with more evenly sized textures should scale
closer to the core count.

Adds an import_bench example for measuring this. Note that import timings are
very sensitive to machine load - on a busy host the same import varies between
~80 ms and ~200 ms - so the benchmark reports the minimum across rounds, which
is the most reliable estimator of uncontended cost.

Co-Authored-By: Claude Opus 4.8 (1M context) noreply@anthropic.com

Image decoding dominates glTF import time. Measured on DamagedHelmet.glb (5 x 2048x2048 textures), decoding accounts for ~99% of the import: ~97 ms of decode versus ~0.36 ms for the document parse and ~0.25 ms for all of parse_texture/parse_dual_texture combined. The decode loop was strictly sequential, so multi-core machines left most of their capacity idle during imports. Split the loop in two. The first stage gathers each image's encoded bytes and stays sequential: it is I/O and bookkeeping, and keeping buffer-view images borrowed via Cow avoids copying the encoded payload. The second stage runs image::load_from_memory over rayon::par_iter, which needs no synchronization because each task reads only its own slice and returns a value. Measured end-to-end on a 6-core machine: ~1.05-1.14x faster (min of 40 runs, 109.5 ms -> 96.0 ms). The decode stage in isolation is ~1.27-1.38x faster. The end-to-end figure is lower than the isolated one because Amdahl's law caps the win: on this asset a single 2048x2048 texture takes ~66 ms of the ~97 ms decode, so that one image sets the critical path no matter how many threads are available. Assets with more evenly sized textures should scale closer to the core count. Adds an import_bench example for measuring this. Note that import timings are very sensitive to machine load - on a busy host the same import varies between ~80 ms and ~200 ms - so the benchmark reports the minimum across rounds, which is the most reliable estimator of uncontended cost. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
perf(importer): decode glTF images in parallel
All checks were successful
Lint / lint (pull_request) Successful in 5m4s
25b56855d8
Image decoding dominates glTF import time. Measured on DamagedHelmet.glb
(5 x 2048x2048 textures), decoding accounts for ~99% of the import: ~97 ms
of decode versus ~0.36 ms for the document parse and ~0.25 ms for all of
parse_texture/parse_dual_texture combined. The decode loop was strictly
sequential, so multi-core machines left most of their capacity idle during
imports.

Split the loop in two. The first stage gathers each image's encoded bytes
and stays sequential: it is I/O and bookkeeping, and keeping buffer-view
images borrowed via Cow avoids copying the encoded payload. The second stage
runs image::load_from_memory over rayon::par_iter, which needs no
synchronization because each task reads only its own slice and returns a
value.

Measured end-to-end on a 6-core machine: ~1.05-1.14x faster (min of 40 runs,
109.5 ms -> 96.0 ms). The decode stage in isolation is ~1.27-1.38x faster.
The end-to-end figure is lower than the isolated one because Amdahl's law
caps the win: on this asset a single 2048x2048 texture takes ~66 ms of the
~97 ms decode, so that one image sets the critical path no matter how many
threads are available. Assets with more evenly sized textures should scale
closer to the core count.

Adds an import_bench example for measuring this. Note that import timings are
very sensitive to machine load - on a busy host the same import varies between
~80 ms and ~200 ms - so the benchmark reports the minimum across rounds, which
is the most reliable estimator of uncontended cost.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
SakulFlee force-pushed perf/gltf-import-parallel-decode from 25b56855d8
All checks were successful
Lint / lint (pull_request) Successful in 5m4s
to ecef6278ce
All checks were successful
Lint / lint (pull_request) Successful in 4m27s
2026-09-29 16:56:28 +00:00
Compare
SakulFlee deleted branch perf/gltf-import-parallel-decode 2026-09-29 17:01:44 +00:00
Sign in to join this conversation.
No description provided.