perf(importer): decode glTF images in parallel #834
No reviewers
Labels
No labels
Context: Bug
Context: Enhancements
Platform: Android
Platform: Linux
Platform: Web
Platform: Windows
Platform: iOS
Platform: macOS
Target: CI
Target: CLI
Target: Dependency
Target: Engine
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
SakulFlee/Orbital!834
Loading…
Reference in a new issue
No description provided.
Delete branch "perf/gltf-import-parallel-decode"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Image decoding dominates glTF import time. Measured on DamagedHelmet.glb
(5 x 2048x2048 textures), decoding accounts for ~99% of the import: ~97 ms
of decode versus ~0.36 ms for the document parse and ~0.25 ms for all of
parse_texture/parse_dual_texture combined. The decode loop was strictly
sequential, so multi-core machines left most of their capacity idle during
imports.
Split the loop in two. The first stage gathers each image's encoded bytes
and stays sequential: it is I/O and bookkeeping, and keeping buffer-view
images borrowed via Cow avoids copying the encoded payload. The second stage
runs image::load_from_memory over rayon::par_iter, which needs no
synchronization because each task reads only its own slice and returns a
value.
Measured end-to-end on a 6-core machine: ~1.05-1.14x faster (min of 40 runs,
109.5 ms -> 96.0 ms). The decode stage in isolation is ~1.27-1.38x faster.
The end-to-end figure is lower than the isolated one because Amdahl's law
caps the win: on this asset a single 2048x2048 texture takes ~66 ms of the
~97 ms decode, so that one image sets the critical path no matter how many
threads are available. Assets with more evenly sized textures should scale
closer to the core count.
Adds an import_bench example for measuring this. Note that import timings are
very sensitive to machine load - on a busy host the same import varies between
~80 ms and ~200 ms - so the benchmark reports the minimum across rounds, which
is the most reliable estimator of uncontended cost.
Co-Authored-By: Claude Opus 4.8 (1M context) noreply@anthropic.com
25b56855d8ecef6278ce