Last Week in Unreal: The Build Farm Stops Blaming the Network (August 10 to 16, 2026)

Unreal Engine Weekly: The Build Farm Stops Blaming the Network
Week of August 10 – 16, 2026 | 847 commits on ue6-main, 55 on ue5-main | Branch: ue6-main
Overview
847 commits, 198 worth reading, and the week belongs almost entirely to build infrastructure. Build is the top subsystem at 121, and a single engineer accounts for 47 of those 198.
The through-line is diagnosis. Several long-running build failures got root-caused this week, and in most of them the system had been pointing at the wrong culprit. Farm hosts were freezing whole builds on single filesystem calls that hung for up to 28 minutes, and the resulting timeouts looked like network bugs. Horde agents on streams with moving Perforce pins were producing builds from a mix of old and new files, and the detector written to catch exactly that stayed dormant, because it read pins from a field Perforce never writes @CL revspecs into. A material instance cook fix landed with a six-of-six versus zero-of-six A/B instead of a claim.
The last two weeks were about caches returning wrong answers. This week is about the instruments that name which thing went wrong, and how long a build farm can run without them.
Development Trends
Build has been the top or second subsystem for eight consecutive full-volume weeks, and the work has moved through three distinct phases. July was modernization: C++23 by default, circular dependencies eliminated, EpicClang versioned, the C# BuildGraph VM deleted. Early August was cache correctness: seven independent defects where a cache key was not a function of its inputs. This week is operations. The UBA cache server got sweep-budget scaling, automatic bucket retirement, retention telemetry, per-bucket metadata, and a web console that scales, because the bucket key gained target type plus branch name for the one heavily used target, taking the count from roughly 500 to roughly 5,000 (6e5e1585, 98dccb91). Retirement policy is service work, and Epic is now doing service work on what used to be a build feature.
The UBA cache server has now been hardened under live production failure three weeks running, and the failure class changes each time. Jul 20-26 was throughput and corruption. Aug 3-9 was process death, with SIGBUS on a full filesystem. This week is scale and misattribution: a worker pool that could enter an absorbing state and freeze a build for roughly 26 minutes (26e9b3dd), a torn read on Apple Silicon that segfaulted inside blake3 (5fb5c544), and a stall detector that reclassifies timeouts caused by local disk (b8a0758b). Three consecutive weeks on the same component means the thing is being operated past the design it shipped with.
The cook-determinism campaign is now in its fifth week and produced its most concrete result yet, though the fixes carry three different reviewers and one jira between them, so they read as unrelated one-offs unless you line them up. A material instance parameter strip that mutates the live object mid-save (68070a5d), Blueprint RNG nodes drawing from the process-global stream instead of the deterministic cook stream (3bc9b39d), and a shader AssetInfo sort truncating 8-byte hashes to 4 bytes at an estimated 99% collision probability across 200,000 elements (8f43677d). Two of the three present as incremental-cook validation failures, the mechanism that pushes studios into full recooks; the shader hash truncation surfaces as raw cook indeterminism on large projects.
Backout volume stayed high and clustered on the same two areas as last week: 27 commits tagged [Backout], and five land-backout-reland cycles inside seven days. Material instance parameter stripping alone burned three commits between two engineers on a single day (364f4a0b, e5601e1b, 68eaf5d3), on top of two diagnostic backouts the day before. Cook determinism, dynamic subsystems, an input debug overlay, and the crash reporter statics each completed a full cycle. The input thread default landed and was reverted with no reland. If you track ue6-main closely, mid-week changelists in the material and cook areas are not safe sync points.
TL;DR
- Horde agents on streams with moving import CL pins have been intermittently producing builds from a mix of old and new files. The existing view-change detection stayed dormant for pin moves, because Perforce never emits
@CLrevspecs in the generated view ofp4 stream -o -v, so pins surface only in the computedChangeViewfield. That field is now captured, hashed, persisted, and diffed on client recreation (61e34d05). - The UBA cache bucket key now includes target type and, for one target, branch name, taking bucket count from roughly 500 to roughly 5,000. A fixed sweep budget of one bucket per poll saturates around 3,600 buckets and the backlog then grows forever, so the budget now scales at one per 500 per poll and idle empty buckets auto-retire (
6e5e1585). - A passive disk stall detector landed in UBA after Ntfs event-log history showed farm hosts holding single filesystem calls for 30 seconds to 28 minutes, near daily per affected machine. Timeout warnings whose wait overlapped a detected stall are now downgraded to info with an explanation (
b8a0758b). - The minimum Windows SDK is now 10.0.26100.0, with the commit sweeping in-tree references to 22621 (
05d13aec). Build images that only carry the older SDK are now below the floor. - Material instance cook non-determinism was traced to
CleanUpUnusedParameters()mutating the live object during cook-save, validated with an A/B that reproduced the failure on six of six stripping instances with the fix off and zero of six with it on (68070a5d). FAssetRegistryGenerator::FinalizeChunkIdswent from 187s to 1s, taking a 300s stage down to 114s, by fixing pathologicalTSetreallocations in the package graph (32e19f61).- Restricting ICF defeating to only where AutoRTFM needs it cut a large server binary by 10.5% (
1f14b550). - The input thread was enabled by default on August 14 and backed out three hours later. As of this writing
input.UseInputThreadstill defaults to 0 onue6-main(fbd5d79a,3ba05293).
Highlights
Horde was building from two changelists at once, and nothing reported it
What changed. Horde agents on streams with moving import CL pins were intermittently producing builds from a mix of old and new files. When a partitioned client is recreated after an edge-server change, the have table is rebuilt with a flush under today’s stream view. If a pin moved since the workspace was last synced, that flush claims files at the new pin that were never written to disk, and incremental syncs never repair them.
Worth reading twice: why the existing safety net never caught it. A previous changelist added view-change detection to reconcile exactly this case, but it read pins from the generated View of p4 stream -o -v, and the server never emits @CL revspecs there. Pins, including ones inherited from parent stream specs, surface only in the computed ChangeView field. The view hash therefore never changed when a pin moved, and the reconcile stayed dormant. ChangeView is now captured, hashed, persisted in workspace state, and diffed on client recreation, reconciled against what can actually be on disk and capped by both the pin and the last-synced CL: forward moves and pin removal flush to the old snapshot and sync only the delta, while backward moves delete the exact set of formerly-versioned stale files and force-sync the new snapshot, so untracked build artifacts survive (61e34d05). A separate fix stops PerforceConfigSource login failures from stranding one native connection per minute-poll with no finalizer (8cec04cf).
Why this is important. A build-correctness failure with no error attached to it. The job succeeds, the artifact is wrong, and the wrongness is a mix of two changelists rather than an obviously old build. The commit describes Epic’s own agents, but the shape generalizes to anyone pinning imports across streams and running edge servers. The detection gap is the transferable lesson: a check that reads the wrong field is worse than no check, because it looks like coverage.
Who should care. Anyone running Horde with pinned stream imports and Perforce edge servers. Release engineers who have ever chased a build that behaved like a stale sync and could not reproduce it.
Urgency: Act now.
The UBA cache server became a multi-tenant service, in one week
What changed. The cache bucket key changed. Target type (client, server, editor, program, game) is now part of the bucket id, and branch name is added for the win64 modular development editor target, with the commit explicitly avoiding wider use to prevent “a permutation explosion (Epic has lots of branches)” (98dccb91, 8af204f5, 11fe8b2d). Bucket count goes from roughly 500 to roughly 5,000. Worth noting the order: the server-side scaling work landed earlier in the week, and the changelist that actually put branch name into the bucket id came last, on August 14. Epic built the headroom first.
Sweep budget now scales at one bucket per 500 per poll, eight times that while a force-deleted wave is pending. The commit shows the math: a fixed budget of one saturates around 3,600 buckets, after which the backlog grows forever, and force-deleted waves took bucketCount/8 seconds to consume while serving failed fetches (6e5e1585). Empty idle buckets retire automatically, since per-branch buckets mean dead branches would otherwise accumulate zombie buckets forever, and the idle grace period doubles as the use-after-free guard for erasing a bucket outside the forced path. Bucket indices now come from a counter rather than map size, so retired indices do not alias logs or round-robin.
Retention became measurable rather than assumed. A Retention column reports the longest-unused entry age, an orange trimmed N marks entries that the last full maintenance pass sacrificed which configured expiration would have kept, since only runs holding the bucket exclusively can compact and enforce, the squeeze step moved from a flat one-hour cut to a proportional 10% per iteration, and a starvation bug was fixed where a bucket left dirty by a non-exclusive run never requeued for maintenance (871534b8, 13adc404, d05288a0). Two metadata paths landed alongside: per-bucket metadata with dedup and eviction past a default of five entries (0ff4cf28), and server-wide key/value metadata (604510d9), with UBT sending the unhashed cache key string as bucket metadata on first use (11fe8b2d), which should make it possible to tell which bucket belongs to which build configuration. The web console moved sorting, filtering, and truncation server-side, because /api/buckets was shipping roughly 1MB per poll at 5,000 buckets and now ships roughly 15KB for the top 50 (ea085af7). Zen server also moved to v5.8.19 (17032ade), though that commit resolves only to Commit.gitdeps.xml, so only the changelog text is verifiable.
Why this is important. Operating a UBA cache is now a different job. Cache buckets used to be a coarse partition you could ignore. Now they are per branch and per target type, they retire on their own, and they report when trimming cost you real cache hits. Learn the orange trimmed N marker: it means the expiration you configured was not the expiration you got, and the recommended response is to split the bucket.
Who should care. Build and CI engineers self-hosting a UBA cache server. Anyone whose cache hit rate has been unexplainably mediocre across many branches.
Urgency: Act now.
Your build farm has been blaming the network
What changed. A passive local disk stall detector landed in UBA, and the commit’s justification is the interesting part. Farm hosts on AWS with EBS storage can hold a single filesystem call for anything from 30 seconds to many minutes while the device otherwise looks healthy, confirmed against Ntfs Operational event 147 history showing Create, Write, and Cleanup calls stuck between 30 seconds and 28 minutes, near daily per affected machine. Those stalls freeze the whole build, and the timeouts they cause downstream looked like network or UBA bugs.
The detector costs two relaxed stores per filesystem call. Wrappers publish the in-flight call in a per-thread slot, and a watchdog scans for calls stuck longer than 15 seconds, logging one line when the stall resolves. The load-bearing part is the reclassification: timeout warnings and errors whose wait overlapped a detected stall are downgraded to info with a suffix naming the stall, covering FetchBegin hangs, WaitForWritten, SendBatchMessages, FetchCompactTable, session ping, and TCP WSAPoll timeouts (b8a0758b).
Two other freezes got root-caused the same way. NetworkServer receive threads could park forever waiting for a worker, and receive threads are where disconnect detection happens, so worker-pool exhaustion was an absorbing state that produced a roughly 26 minute whole-build freeze. PopWorker was changed to wait 10 seconds and then spawn an emergency worker beyond maxWorkerCount, capped at twice the limit, with a warning (26e9b3dd). Do not go looking for that behavior: the same author reverted it hours later, having concluded it would trigger constantly (d97e9372), and shipped a callstack dump for a PopWorker stalled past 15 minutes instead (2ab17b2e). The freeze is diagnosed, not yet fixed. And on Apple Silicon build machines, m_memorySize and m_overlaySize were plain u32 published across threads with no barrier, which is safe on x86 total store ordering and not on weakly ordered ARM, so a reader could see the new size before the bytes were visible and hand a torn filename length to a routine that crashed inside blake3_hasher_update. Both fields are now Atomic<u32> with explicit acquire and release, plus a bounds assert so a future torn write fails loudly instead of segfaulting deep in a hash function (5fb5c544).
A related Mac farm link collapse traced to file-mapping handles carrying the creator’s raw file descriptor, so a close-then-recycle served the wrong file to a live fetch. Handles now carry a monotonic mapping id resolved through a live-mapping registry, the descriptor server thread no longer exits on transient accept errors, and descriptor fetches get 10 second socket timeouts (3d4eb7a3). Throughput work stayed measurement-driven: stall sampling showed roughly 380 of 400 workers parked in FetchSegments.Send, so receive-stall thresholds now reserve the upper half of each connection’s send budget as headroom and segment fetch batching auto-sizes to two batches per connection instead of a fixed eight (4cce4d3b, 75c07a40, bfafea3b). Cold-page faults inside fetch memcpys got an async readahead, with a POSIX madvise(MADV_WILLNEED) implementation added because, as the commit puts it, the production cache servers run Linux (ba435bd5). Trace buffer exhaustion now logs and saves a truncated trace instead of failing silently, with a configurable reserve size (7d10a313).
Why this is important. Three separate classes of report that read as flaky networking were local resource starvation, and the fix in each case was a detector that renames the symptom rather than a retry that hides it. If you operate a build farm on cloud block storage, the specific finding is worth checking on your own hardware before you assume it is an Epic-scale problem: multi-minute filesystem stalls on healthy-looking instances, near daily.
Who should care. Build engineers running distributed builds on AWS or equivalent, anyone with Apple Silicon build machines, anyone who has written off recurring build freezes as network flakiness.
Urgency: Act now.
The cook was baking fresh random numbers into your packages
What changed. Three cook-determinism fixes landed.
UMaterialInstance::Serialize calls CleanUpUnusedParameters() while cooking, which destructively removes unused entries from the instance’s parameter arrays before Super::Serialize writes them. SavePackage serializes each export more than once, a reference-harvest pass and then an export-write pass, and the cooker can save the same instance more than once in a process. Because the strip mutates the live object and is skipped during the reference-collector pass, a later save harvests a smaller parameter set than the first, and the harvested name table then differs between the two saves. IncrementalValidate reported that as a spurious DeclaredUnmodified_FoundModified false positive, described in the commit as the dominant MaterialInstanceConstant spike. The fix snapshots the seven mutated arrays and restores them after serialization, behind r.Material.RestoreStrippedParametersAfterCookSave, default enabled. It shipped with an A/B: with the restore off, six of six stripping instances emit NameMap-is-different and Indeterministic-header-size; with it on, zero of six (68070a5d).
Blueprint RNG nodes in construction scripts were baking a fresh random value into cooked packages on every cook. AActor::ProcessUserConstructionScript already establishes a deterministic cook-seeded scope, but only KismetGuidLibrary::NewGuid consulted it, so RandomIntegerInRange, RandomFloat, RandomRotator, and array Random and Shuffle still drew from the process-global stream. Those nodes now consult the in-scope cook stream when one is active, gated to editor builds, so shipping behavior is unchanged (3bc9b39d, a resubmit of the first attempt 9a961ed9 after it was backed out in af23eb62 for a Linux and Mac compile error, where RAND_MAX + 1 overflowed int because RAND_MAX is MAX_int32 outside the MSVC runtime). And the shader AssetInfo sort was clamping 8-byte FShaderHash values to 4 bytes via GetTypeHash, which the commit puts at an estimated 99% collision probability at 200,000 elements, so large projects specifically got indeterminate output (8f43677d).
Not a determinism fix but from the same author and the same cook path: FAssetRegistryGenerator::FinalizeChunkIds had pathological TSet reallocations in the package graph, and the commit reports 187s dropping to 1s, taking the enclosing stage from 300s to 114s, measured on a NullCook performance run rather than a production cook (32e19f61).
The material instance area also produced the week’s worst thrash. A commit disabling MIC unused-parameter stripping outright, because it was causing issues with texture parameters used in runtime virtual texture rendering, was submitted, backed out, and then restored by backing out the backout, all on August 12 across two engineers (364f4a0b, e5601e1b, 68eaf5d3).
Why this is important. Every one of these presents as an incremental cook validation failure rather than as an error, which is the mechanism that converts incremental cooks into full ones. If your team has concluded that incremental cook validation is noisy and started ignoring it, the material instance finding is a concrete reason it was noisy, with a measurement attached.
Who should care. Anyone running incremental cooks at scale, anyone with a content-addressed artifact cache, technical directors who have stopped trusting cook validation output.
Urgency: Act now if you cook material instances or use Blueprint RNG in construction scripts.
Version floors moved, and one of them will stop your build
What changed. The minimum Windows SDK is now 10.0.26100.0, with the commit noting it fixes all in-tree instances of 22621 found through internal code search (05d13aec). The commit does not spell out the consequence for build machines, but a floor is a floor: an image carrying only the older SDK is now below it. Agility SDK 1.619.5 was added and made the default (90ed17ad, 4875ceb2), 1.619.3 was removed (8811b76a), and two debug-layer suppressions were dropped because the validation false positives they covered are fixed in that SDK (ab6243fe). AMD AGS went from 6.0.1 to 6.3.1, adding RDNA3 and RDNA4 asic detection and DX12 shader clock intrinsics, and replacing VS2017 static libraries with VS2026 (cd727840). New DXC binaries shipped alongside a VulkanRHI fix for an invalid OpStore to OpTypeImage from Interlocked* on copied resources, which the commit says was causing crashes on some drivers (9bbbbb5c, 90a9fcf7).
One toolchain change is worth stealing. Restricting ICF defeating to only the cases AutoRTFM actually needs, rather than applying it blanket, reduced a large server binary by 10.5%. The commit also documents two behaviors the blanket approach was providing by accident, including AddressSanitizer disabling its fake stack for functions containing inline assembly, which the AutoRTFM runtime depended on (1f14b550).
Why this is important. SDK floor changes are the least interesting and most disruptive category of build change. They do not appear in release notes as features, and they turn into a morning of red CI for whoever syncs first.
Who should care. Build engineers, anyone maintaining pinned toolchain images, console and Windows platform teams.
Urgency: Act now.
BuildGraph’s C# authoring API finished collapsing
What changed. Two weeks after the bytecode VM was deleted, the rest of that layer went with it. BgContext, BgThunkDef, and the remaining expression and VM plumbing are gone, the Def-suffixed hierarchy collapses into direct BgGraph, BgNode, and BgAgent types, graph authoring methods moved from BgGraphDefBuilder onto BgAgent and BgNode, authored node delegates now execute directly instead of through reflection-based thunk binding, and unsupported Task<T> node functions are now rejected outright instead of running with their results silently discarded (36037439). All CustomTask implementations migrated to BgTaskImpl and the deprecated base class was removed (61b23381, cf13ef3e).
Horde also spent the week cutting the cost of its own CI. Analyzers moved inside the Docker compile that produces the shipped binary, so only one build runs per change instead of the entire server graph twice (f5da8afd). Interprocedural security rules moved to an opt-in configuration after costing 33 seconds per change while producing zero diagnostics (86e09fb5). Integration tests moved from creating and dropping a MongoDB database per test to a pooled model (775365c6), and test projects now run in parallel (210d5e5d, c4ca5172, 3c1c8e40). Tool deployments now precompute a zstd-compressed tar archive stored as a flat blob, served by presigned redirect straight from the storage backend, so the server spends neither CPU nor bandwidth on it and the artifact extracts with plain tar --zstd -xf while still downloading (db4fe1b3, 9c50776b).
Why this is important. The XML authoring surface looks untouched, which is what most studios actually use, though the commit rewrites XML graph reading and execution onto the consolidated model and CustomTask subclasses have to move to BgTaskImpl. If you invested in C# graph authoring, this is the second breaking pass in three weeks, and the API you migrate to has a different shape as well as a smaller surface. The Horde CI cuts are worth reading as a pattern even if you do not run Horde: two of the three were removing work that produced no output.
Who should care. Anyone authoring BuildGraph in C#, Horde operators, CI engineers looking for cheap wins in their own pipelines.
Urgency: Act now if you author C# BuildGraph. Track otherwise.
The renderer decomposition took Nanite material bins and distance fields
What changed. Two more large migrations moved subsystems off explicit FScene calls and onto dirty-mask tracking resolved inside a defined update window. Nanite material bin lifecycle left FScene and FPrimitiveSceneInfo: PrimitivesNeedingStaticMeshUpdate went from a packed-index TBitArray on FScene to an FDataArrayDirty on FScenePrimitiveDataRenderer, a PostStaticMeshUpdate hook collects the rebuild set, dirty marks can be locked for the remainder of an update, and the mutexes were removed because the update task and the render thread no longer touch the same data (7e6eb874, 374879a1).
FDistanceFieldSceneData became FDistanceFieldSceneExtension, and that migration carried several fixes that have nothing to do with refactoring. Three worth calling out. ShouldPrepareDistanceFieldScene now mirrors ShouldCompileDistanceFieldShaders, so render prep cannot request shaders that are absent from the global shader map when distance fields are disabled. GetDistanceToNearestSurfaceGlobal returns no-surface-in-range instead of distance 0 when no global signed distance field clipmaps exist, which had been pinning Niagara particles at the emitter. And hidden shadow or GI casters keep their distance field while hidden (39fc371d).
The fallout confirms how load-bearing this work is. A use-after-free race had scene extension task handles stored on the scene data object awaited on a worker thread under RDG parallel destruction, racing the render thread replacing handles for the next frame (fa537262). Out-of-band dirty marking from PSO compilation completion asserted on the locked update array (374879a1). A TAtomic was swapped for a transactionally safe mutex to support in-transaction render proxy initialization (37f10060). Separately, occlusion culling got an element-id cache for the occlusion history set, where the lookup had been 43% of sampled cycles on a gen8 mobile console, plus opt-in visible re-test throttling that is off by default, the two together measured at 0.4ms render thread, 0.3ms game thread, and 0.1 RHI (fede2869), and occlusion bounds moved to sparse custom bounds to cut memory (00cdc259).
Why this is important. Seven consecutive weeks of the same campaign, and the pattern has not changed: explicit add and remove calls become dirty flags resolved in a window, which is what makes parallel scene update safe. If you carry renderer patches, these are real conflicts, and the distance field fixes are worth pulling on their own merits regardless of the refactor.
Who should care. Rendering engineers, anyone maintaining an engine fork with renderer changes, anyone using global distance fields with Niagara.
Urgency: Track. Act now if you maintain a renderer fork.
Lumen gained a reference path, and Nanite ray tracing stopped holding everything resident
What changed. A new LumenRef subsystem appeared inside the renderer, with its own world cache, initial sampling, light and material sampling, path resolve, and visualization passes, wired into the deferred shading renderer and scene view state (5e889c76), followed by a world cache freeze visualization fix the same week (209c4705). The commit message is two words, “Initial prototype,” and says nothing more.
The rest is measured performance work. Lumen irradiance field gather now reuses downsampled depth and normal from the stochastic lighting tile classification pass, which itself gains Substrate tile size 16 support and caches translated world position in shared memory. The change also swaps some shared arrays from float3 to float2 plus float because, per the commit, some platforms pad float3 to float4, which increases shared memory usage and reduces occupancy. Measured at 0.51ms saved on console at 900p with the Substrate blendable gbuffer, and 0.88ms with MegaLights enabled (ca72133c). Nanite ray tracing now builds CLAS on demand and reclaims CLAS for unreferenced pages instead of holding everything resident, with streaming requests originating from ray tracing flagged through a previously unused bit and LRU ordering maintained for ray-tracing-referenced pages (e8aa742e), and automatic quality scaling now keys off CLAS pool usage (cbeb9852). Lumen reflections gained optional translucency depth during screen tracing so screen traces hit the frontmost depth-writing translucent layer, at roughly 0.15ms fixed cost at 1080p on console (c7740741). CSM shadow depth reuse gained validity-based invalidation and frustum expansion under r.Shadow.CSMShadowDepthReuse, reusing static and far cascade depth until the light rotates or the cascade footprint leaves the frozen, expanded footprint (50747fde). ACES 2.0 transforms migrated to match OpenColorIO’s adaptive hue and cusp construction, packed into a single 363×1 RGBA32F texture (521db891).
Why this is important. The CLAS change is the practical one: hardware ray tracing memory pressure on Nanite content stops being a function of everything you loaded and becomes a function of what ray tracing actually referenced. The Lumen gather saving is free if you are on Substrate. LumenRef is worth exactly one line in your notes until it does something.
Who should care. Rendering engineers, technical artists on GI-heavy or hardware ray tracing projects, console teams.
Urgency: Track.
A week of crashes that all happen after the useful work is done
What changed. An unusual share of this week’s crash fixes are teardown crashes. The clearest one: a fatal error raised from an atexit destructor died on already-destroyed function-local statics before it could be reported, exiting with a crash-reporter-crashed code and producing no report at all. The fix makes the thread manager singleton, the log output device, the log-filename lock, the error and feedback context devices, and the ApplicationWillTerminate and memory-pressure delegate accessors immortal, and with the singleton immortal the runnable thread destructor always unregisters, so shutdown crash reports keep every thread’s name and callstack. Note the landing sequence: it landed, was backed out, and was restored with the note that the backout did not fix the continuous-integration failure it was meant to fix (a6d3dc77, a3eda6f3, cfc74eef).
A sweep fixed existing use-after-free instances across ContentBrowser, the Niagara editor, Slate accessibility, the virtual shadow map cache manager, and the console, and then FText::ToString was marked lifetime-bound so static analysis catches the same bug going forward (3a42e0b6, then b93c6cea). Sweep plus annotation is worth copying into your own codebase for this class of bug.
Elsewhere: Control Rig had two concurrency crashes, including InitializeFromCDO reading the shared source hierarchy while also writing to it with nothing serializing it, so instances initializing on anim worker threads corrupted the listening hierarchies array (85d0e5b2, 0e9cee80). Async loading’s state dump moved into a commandlet with its use-after-free fixed by taking references rather than relying on package lifetime (d08fb4ca). IoStore container mapped file handles are now kept alive while mapped regions are outstanding, fixing a fatal assert when a game feature plugin unmounts while memory-mapped animation or sound bulk data is still resident (01bc8239). Audio wave instances outliving their active sound left a dangling back-pointer that the staleness check could not see (f9aa5086). And two that are easy to miss: an out-of-bounds read of the hardcoded-name table on Iris name deserialize (841d1f05), and a data race on a bit-packed Visual Logger flag where a concurrent load-mask-store could tear neighboring flags (1b008a1a).
Why this is important. Shutdown crashes get deprioritized because the work is already done, and the crash reporter fix explains part of why: some of them were never producing a report to prioritize. If your crash telemetry shows a suspiciously clean shutdown path, that may be the reason.
Who should care. Engine and tools programmers, anyone shipping plugins that unmount at runtime, anyone triaging shutdown crashes or missing crash reports.
Urgency: Track. The Iris deserialize fix is worth a look if you ship networked builds.
Also worth knowing
- All subsystems became dynamic subsystems, and
UDynamicSubsystemis deprecated. The driver is merged modules, where dynamic libraries may carry subsystem definitions without game code opting in everywhere. Subsystems in automatically loaded plugins keep their old behavior, explicitly loaded plugins get dynamic behavior, and a new opt-in restores per-module lifetime. It landed, was backed out, and relanded (c11e8cb7,783fba87,9a17b956). Treat it as a source-compatibility item on merge. - Plugin unmount got several minutes faster. Package unloading no longer loads data unless a rename is actually being performed, tagged for release notes by Epic (
5ccdf834). A related fix stops evicted thumbnail scenes from outliving a plugin unload and keeping its packages referenced (ff42b7c2). - D3D12 enhanced barriers got a correctness pass while the default stays legacy. Layout selection had been picking queue-type-specific layouts from whichever pipe a transition happened to be recorded on, which the commit calls wrong twice over, and shader-resource, unordered-access, and generic-read layouts are now derived from access alone (
dd798781). Automatic selection narrowed so it only picks enhanced barriers when the SDK provides the required layout constant (7e0511ae), and a global was exposed specifically so projects can report barrier adoption through their own telemetry (b84bdb47). Correctness fixes plus adoption instrumentation usually come before a default flip. - CTR encryption landed for plugin and collection encryption, with a deterministic Blake3-derived initialization vector per block so delta patching still works, and the commit says the foundation should later allow a full-project switchover at a one-time patch size cost (
aa47a726). - A black screen in UE Remote is fixed. Signalling server CORS filtering added earlier had a loopback-only allowed-origins default that rejected every UE Remote client, so no codec was ever negotiated (
928808b3). - The CSV profiler changed what it measures by default.
csv.UseLegacyFrameTimewas renamed tocsv.MeasureFrameTimeAtEndFrameand defaulted on, with Windows opting out via ini. Any cross-platform frame-time capture comparison spanning this change is not comparing the same measurement (c3aac366).
Worth Tracking
The input thread default flipped on and back off inside three hours. It was enabled by default on August 14 with input.UseInputThread and -NoInputThread as escape hatches (fbd5d79a), and a backout with the identical description landed the same afternoon (3ba05293). We checked the branch: input.UseInputThread still defaults to 0 on ue6-main, so nothing has changed for you yet. Worth watching, because a default-on input thread changes threading assumptions for anything reading input state off the game thread, and Epic has now tried once.
A new experimental geometry plugin called JOULE appeared at Engine/Plugins/Experimental/JOULE, with core mesh, grid, atlas packing, and cut-face headers. The author describes it as an initial assisted port and a proof of concept to be matured as it is used, and it landed with no reviewer (f7987d2d). The name is the technique: jointly optimized u(v) lattice embeddings, an approach to continuous parameterization with a jointly optimized volumetric step. Two follow-ups the same week are a static analysis pass and a build fix, both housekeeping. The next few weeks of volume will tell you whether this is the seed of a real parameterization pipeline or something that disappears.
An engine-wide logging migration is underway under one ticket. Roughly eight commits this week move UE_LOG and UE_CLOG onto narrow-format UE_LOGF and UE_CLOGF, including the public wrapper macros and the Visual Logger wrappers, with the wide UE_VLOG_UELOG and UE_CVLOG_UELOG variants deprecated and several per-system logging macros migrated onto the narrow-format replacements (a42421cd, 4f1a1165, 28e1dbae, 22070966). Broad mechanical churn across public headers plus a deprecation means a removal step is coming and merge pain with it.
The modular test framework continued at roughly 20 commits for a third consecutive week, adding a standalone hub program that mounts a test plugin with its dependency closure (9a258893), a runner UI with warm-pool visibility and per-worker kill (e37999f9), elision of the reset and travel cycle between suites sharing a resolved environment (4b3252da), an external-session worker mode (d5960e48), and a Verse-facing plugin (b6d2c48f). Three straight weeks at that pace makes this a program.
For teams shipping on 5.8
ue5-main ran 55 commits this week, down from 122, and the mix is narrow: animation work dominates once you count UAF, Control Rig, and PoseSearch together, with the rest split between a procedural vegetation export cluster, localization automation, and isolated fixes.
The one to pull is a heap corruption. UControlRig::InitializeFromCDO reads the shared source hierarchy while also writing to it, registering a listening hierarchy and toggling a cache validity flag, with nothing serializing access, so multiple Control Rig instances of the same class initializing concurrently on anim worker threads corrupted the listening hierarchies array (2da7adde, b3af43db). If you run Control Rig on characters and have ever seen unexplained anim-thread instability at scale, this is a candidate.
Also on the branch: UAF variable refactor phase 1 replaces the binary public and private variable flag with input, internal, and local scopes plus access control, where external callers reach input-scope variables only and inputs are read-only inside UAF (5c3106ac). A PoseSearch fatal check on a database and schema race in completion handling was fixed (d6d94482). Lumen hit-lighting picked up inline ray tracing from ray generation shaders and disabled unused permutations (e3cc02de). Movie Render Queue PIE lifecycle issues were fixed, including a Quick Render crash when a render is aborted because a Blueprint does not compile (420759f4). And a use-after-free from binding a const reference to a temporary returned by FText::ToString() was fixed on the platform text entry path, the same class of bug that got the annotation and sweep treatment on ue6-main (de14038d).
No 5.8.x point release is named in this week’s commits on the branch.
What We Ignored
We passed over roughly 480 commits this week: static analysis and warning cleanup passes, localization automation, formatting, test plumbing, and content asset updates that resolve only to binary manifest blobs. That figure includes the 27 commits tagged [Backout], whose net effect shows up in the changes they landed on or removed.
Closing
A quiet week for anyone shipping content and a loud one for anyone running the machines that build it. Three separate multi-week failures ended the same way: someone added an instrument that names the cause instead of another retry that hides it.
Commit SHAs are cited inline, in parentheses, next to each claim. They reference the ue6-main branch of EpicGames/UnrealEngine unless the claim is in the 5.8 section, which references ue5-main. Some of the highest-interest entries resolve to Commit.gitdeps.xml binary-manifest blobs: the commit message is real, but the code lives in Epic’s internal binaries, which we cannot see.
Never miss a ue5-main update.
Get them delivered straight to your inbox.


