-
b9160962
by Simon Jakobi at 2026-08-20T14:57:52-04:00
testsuite: Show baseline sample count and range in perf failures
A perf baseline is the mean of all samples recorded for a commit, and
it prints as a single number, hiding how far the samples spread. When
the spread is wide, this can indicate an unstable metric that isn't
actually useful as a signal for the perf tests.
For example, in #27602, T27336's peak_megabytes_allocated baseline
showed as 757 when the underlying samples were 605 and 909.
When the baseline is averaged from more than one sample, say so in the
failure output: the one-line stat-failure reason shows the sample
range, and the detail block lists the raw samples. Single-sample
baselines print exactly as before.
Context: #27602
Assisted-by: Claude Fable 5
-
a4979877
by Simon Jakobi at 2026-08-20T14:57:52-04:00
testsuite: Fold Baseline into CommitMetric
A Baseline was just a CommitMetric plus the commit it came from, built
by copying fields across. Since get_commit_metric already knows that
commit, record it on CommitMetric itself and drop Baseline. This also
collapses both branches of find_baseline into plain returns.
Assisted-by: Claude Fable 5
-
99fb8d68
by Simon Jakobi at 2026-08-20T14:57:52-04:00
ci: Clarify comment on pushing perf notes after failures
Context: #27602
Assisted-by: Claude Fable 5
-
2ca87972
by Alan Zimmerman at 2026-08-20T14:58:36-04:00
EPA: Remove LocatedBC / SrcSpanBF
The custom annotations are now in the BooleanFormula TTG extension
points, so LBooleanFormula can now use the standard LocatedA.
-
d2bc32aa
by Simon Peyton Jones at 2026-08-21T12:59:26-04:00
Better handling of serialisation of wired-in names
Fixes #27501
-
d2795ffc
by Alan Zimmerman at 2026-08-21T13:00:05-04:00
EPA: Remove NoEpTok/NoEpUniTok, using an unhelpful SrcSpan instead
Also introduce helper functions noEpTok and noEpUniTok to serve
as simple replacements in code inserting an token annotation without
location information.
-
b5d29ab8
by Brandon Chinn at 2026-08-25T18:42:08-04:00
Add Data.RealFloat and Infinity/NegInfinity/NaN pattern synonyms (#26961)
-
e60eb3bc
by Andreas Klebinger at 2026-08-25T18:42:59-04:00
rts linker: Fix pointer arithmetic issue in flushInstructionCacheRISCV64
We accidentally operated over `uint64_t*` when we should use `uint8_t`.
Fixes #27569
-
e9bbe8f9
by Andreas Klebinger at 2026-08-26T15:09:23-04:00
cmm dumps: Add machop width info with -dppr-debug for infix ops.
-
86e3a9d8
by Andreas Klebinger at 2026-08-26T15:09:24-04:00
CmmLint: Check for unsupported MachOp widths
machOpArgReps now maps MachOp + Width to a list of supported
argument widths or Nothing if the given operation is not supported
at the given width.
This allows us to check for nonsensical combinations like FloatToInt
at Word16.
Similarly we now check that every address is actually wordwidth.
-
13781cca
by Andreas Klebinger at 2026-08-26T15:09:24-04:00
arm64 ncg: The big subword truncation fix.
A set of slightly related fixes to arm subword handling:
Bitmask immediates:
Don't produce overflowing assembly literals.
There is still another bug here that causes us to miss some valid
literals but we will fix that later.
Improve subword truncation handling:
We now use a small set of helpers to truncate `Register` values rather
than truncating immediate `Reg` values which greatly simplifies the code
structure. This fixes a great many bugs to do with sign/zero extending subwords
or the lack thereof.
We now establish the invariant that subword values are zero-extended at
every site at which they come into "scope" of the ncg, and rely on the
invariant throughout rather than pessimistically inserting redundant
extensions in a hodgepodge manner at the use sites of these values.
This fixes at least the bugs described in issues #27533, #27430
#27537, #27538, #27539, and #27550. But likely more bugs yet not
found.
Subword ffi results:
Apply truncations when calling functions returning
subword values.
genCondJump:
Don't sign extend signed values in the input register as
it might map to a local variable, corrupting the value stored within.
Fix subword store/load instructions.:
We used to read those at 32bit width even for smaller values possibly
resulting in invalid memory access. Now we construct the suffix for
subword variants based on the instruction format for these.
-
d8fa5d7c
by Andreas Klebinger at 2026-08-26T15:09:24-04:00
arm64 ncg: Fix MO_V_Broadcast for non-literals.
We now use OpReg instead of OpScalarAsVec as required since we broadcast a gp register.
Also adds a test. Fixes #27565.
-
94822c95
by Andreas Klebinger at 2026-08-26T15:09:24-04:00
Add some test cases covering bugs in the arm ncg.
* Test for #27430 (subword ffi results)
* #27537 - subword conversions
* #27538 - subwords used in conditional
* #27533 - single byte read
-
dd1ba88a
by Andreas Klebinger at 2026-08-26T15:09:24-04:00
cmmLint: Lint against MO_FS_Truncate subword use.
-
fd22f71e
by Zubin Duggal at 2026-08-26T15:10:20-04:00
ghc-internal: annotateSTM should use catchSTM# rather than catch#
A catch# frame inside a transaction breaks retry and async exception
delivery.
Fixes #27657
-
bb324171
by Rodrigo Mesquita at 2026-08-26T15:10:59-04:00
rts: refactor to reduce THREADED_RTS in MSG_UPD_TSO_FLAGS
- No behavior change in this commit (well, a small optimization here
makes us do less work if the target TSO owned by the curr. capability)
- Move all THREADED_RTS CPP needed into `updThreadFlag`
- Merge MSG_SET_TSO_FLAGS and MSG_UNSET_TSO_FLAGS into MSG_UPD_TSO_FLAGS
plus a `set` bool field in the MessageUpdTSOFlag struct
Towards #27729
-
ed99b7b7
by Rodrigo Mesquita at 2026-08-26T15:10:59-04:00
rts: Fix race condition in MSG_UPD_TSO_FLAGS execution
The code for processing the MSG_UPD_TSO_FLAGS message was not taking
into consideration that the TSO's owner might have moved in between that
capability receiving the message (since it was its previous owner) and
starting to process its inbox (a point at which it was no longer the
owner)
Added Note [TSO owner may change in between Msg being sent and received]
to explain this race and the pattern used to fix this, where we just
forward the message to the new owner.
Fixes #27729
-
cd653714
by Alan Zimmerman at 2026-08-26T15:11:49-04:00
EPA: Uses Parsers.parseModule for exactprint tests
Parsers.parseModule is the advertised way to parse for use for exact
printing in the ghc-exactprint library. This commit updates the GHC
exact print testing to use it.
This requires moving the comment balancing that was occurring
only in the test path into the advertising parser path, so it moves
from Transforms.hs to Utils.hs.
Also update the comment adding to honour trailing annotations
-
d1d01fa5
by Wolfgang Jeltsch at 2026-08-27T13:17:59+03:00
Add `rethrowSTM` and improve STM-related documentation
Adding `rethrowSTM` resolves #26758.
The implementation of `rethrowSTM` is completely analogous to the one of
`rethrowIO`.
The following is established for the documentation of `throwSTM` and
`catchSTM`:
* Both operations are directly described as analogs of their `IO`
counterparts.
* There is no reference to `throw` in the documentation of `throwSTM`,
because, although such a reference is great in the documentation of
`throwIO`, it is somewhat out of place in the documentation of
`throwSTM`.
* Instead of repeating part of `throwIO`’s documentation, the
documentation of `throwSTM` just recommends using `throwSTM` instead
of `throw` and references the corresponding arguments in the
documentation of `throwIO`.
-
06fde293
by fendor at 2026-08-28T06:06:44-04:00
GHCi: Fix order of `PackageDBFlag`s for interactive home unit
`PackageDBFlag`s are stored in reverse order of cli specification.
When sorting the `PackageDBFlag`s by longest common prefix, we need thus
to reverse the package db stacks before calculating the prefix.
We make sure to reverse the package db stack for the interactive home
unit to uphold that later specified package dbs overwrite earlier ones.
Resolved and adds regression test for #27640
-
024c4d04
by fendor at 2026-08-28T06:07:23-04:00
Reuse the UnitIndexCache after initialising multiple home units
-
55326fa0
by Alan Zimmerman at 2026-08-28T06:08:03-04:00
EPA: Some Haddock processing tweaks
These changes to the Haddock postprocessing should not change
behaviour, but just bring it more closely in line with the
original, changed at 44309cd377f
And add some haddock exactprint tests to show they work.
-
b3ddee95
by Andreas Klebinger at 2026-08-28T13:57:46-04:00
hadrian: Deprecate quickest flavour.
It was more of a trap for new users than actually beneficial so we
deprecate it and suggest quick+no_dynamic_libs to users instead.
-
a1d81390
by Andreas Klebinger at 2026-08-28T13:58:37-04:00
cmm: Always favour entry block during block deduplication.
We now always keep the first block in the CmmGraph. This way we avoid
the need to update the entry info table.
Failing to do so caused #27722.
Fixes #27722.
-
5bd65f00
by Andreas Klebinger at 2026-08-28T13:59:16-04:00
test: FamAppCachePerf - Only collect bytes allocated. Fixes 27747
-
ced53ce6
by mangoiv at 2026-08-29T07:15:24-04:00
nightlies: output yaml to file only
Previously we would just output the metadata to stdout
which risks that it's clobbered by incidental debugt output.
We now output to file only.
Fixes #27511
-
578bd185
by Andreas Klebinger at 2026-08-29T07:16:05-04:00
Specialise: Stop looping on recursive dictionaries in interestingDict
interestingDict now doesn't look through loopbreaker unfoldings.
Doing so would cause infinite loops on certain dictionaries.
Fixes #27705.
-
9f945c90
by Duncan Coutts at 2026-08-29T14:57:04+01:00
Minor doc & comment improvements to releaseCapability_
-
29cd4bfa
by Duncan Coutts at 2026-08-29T14:57:04+01:00
Remove redundant USED_IF_THREADS attribute on Capability utilities
-
414c6c64
by Duncan Coutts at 2026-08-29T14:57:04+01:00
Move several Capability utils from Schedule.{c,h} to Capability.{c,h}
They probably should have been there all along. This means all the
pending_sync functionality is within Capability.{c,h}. We only expose
pending_sync for the purpose of inline header functions.
-
aae57993
by Duncan Coutts at 2026-08-29T14:57:05+01:00
Shuffle the pending sync type declarations for better readability
Move them together into the section with the related functions that use
them.
Also drop the legacy use of the C 'volatile' modifier on the
pending_sync variable. We use C atomics for such access, not volatile.
-
c1586415
by Duncan Coutts at 2026-08-29T14:57:05+01:00
Rename returning task queue helpers
Follows a naming convention elsewhere. It also gives us suitable names
to distinguish appending vs prepending to the queue, and we're about to
add a prepend operation.
-
73798761
by Duncan Coutts at 2026-08-29T14:57:05+01:00
Add a prepend operation for the returning task queue
With a pending sync (e.g. for GC), we really want to be able to
prioritise the task waiting on the sync over all other returning tasks.
To do that we will need to prepend to the queue rather than append.
-
c0d0b784
by Duncan Coutts at 2026-08-29T14:57:05+01:00
Split waitForSomeCapability out of waitForCapability and adjust callers
Previously waitForCapability had a general interface covering serveral
situations, but it's more useful and easier to understand with two more
specialised functions.
Previously waitForCapability had an in/out Capability parameter: if the
cap was non-null then it would wait to acquire that specific capability
(though in this case it had to be the capability the task was already
associated with), or if it was null then it would pick a suitable
capability, associate the task with it and wait for that capability.
So overall, this required and in/out parameter and suggested that the
capability returned could be different from the one passed in. That was
true if the task was associated with no capability, but false if it was.
This complicated the call sites where we know that the task does have an
associated capability, as it implied the capability could change when in
fact it cannot.
So we now split it: waitForSomeCapability handles this general case
where the task may or may not have an associated capability, and it
returns that new capability. And the waitForCapability now assumes that
the task is associated with a capability and does not need to return
anything. Neither function needs a Capability *cap argument any more.
This simplifies several call sites.
-
4aa0363d
by Duncan Coutts at 2026-08-29T14:57:05+01:00
Update stale comments related to waitFor[Some]Capability
In particular waitForCapability cannot change the capability.
Update a comment about grabbing a new Task in performGC_. This comment was
accurate in 2006, but it is no longer accurate since newBoundTask does not
in fact produce a new Task (typically). The right thing now is to talk about
InCalls, not tasks. That said, the code is still correct since
newBoundTask will push a new incall, but also deal with the general case
of an OS thread that does not yet have an associated Task.
-
84ec7bba
by Duncan Coutts at 2026-08-29T14:57:05+01:00
Introduce waitForCapability_ with additional priority arg
Split waitForCapability into a wrapper with the existing type and a
worker with an extra argument.
The new high_priority argument controls whether the task is appended or
prepended to the returing task queue. The default, used by the
waitForCapability wrapper, is false, meaning append to the end of the
queue. This gives fairness.
The high_priority==true case will be used in the subsequent commit.
-
2463167d
by Duncan Coutts at 2026-08-29T14:57:05+01:00
Make acquireAllCapabilities use high_priority on waitForCapability_
As discussed in issue #27473, a sync of all capabilities is something
that needs to happen promptly (but often doesn't).
One source of delay is that acquireAllCapabilities using
waitForCapability would put the task trying to acquire each capability
at the _end_ of the returning task queue. This gave every other returing
task a full timeslice to run. Meanwhile, several other capabilities are
blocked waiting for the sync to complete, leading to a loss of
throughput.
We use the new high_priority arg to waitForCapability_ to ensure that
the requesting task is put on the front of the returing task queue. This
will ensure that releaseCapability_ will prioritise giving the
capability to the task requesting the sync.
-
be6c45e8
by Duncan Coutts at 2026-08-29T14:57:05+01:00
In releaseCapability_ make the pending_sync case self-contained
Previously the pending_sync case had to be checked _after_ the returning
tasks case, since one of the possibilities (indeed the more likely
possibility) is that the task calling waitForCapability will have
enqueued itself as a returning task.
Now we make the pending_sync case self-contained. We note in a comment
the two possibilities: either the task calling waitForCapability has
enqueued itself already and is waiting, or it's not got there yet. We
can handle the first case by giving the capability to the task at the
head of the returning tasks queue, and the second case by leaving the
capability free.
Another way to look at this, is that we move a special case of handling
of the returning task case into the pending_sync case. That special case
being a returing task during a pending sync.
This makes the order of handling returning tasks vs pending sync
independent. This is good, because really they're in the wrong priority
order and we want to flip them around.
-
6bb34033
by Duncan Coutts at 2026-08-29T14:57:05+01:00
In releaseCapability_ prioritise pending sync over returning tasks
Fixes issue #27460
As explained in the issue, a pending sync (e.g. for GC) should be dealt
with promptly. Returning tasks are a lower priority.
Historically however we had to check returning tasks first, because the
synchronisation mechanism mixed up the task doing a sync with the tasks
returning from safe FFI calls. The task performing the sync was very
likely to be queued on the returning task list (and historically it was
at the _end_ of this list!).
We have now arranged that the task performing the sync is at the front
of the returning task list, and in the pending sync case we now check
the returning task list and run the first task from there if it there is
one.
This is by no means perfect, but it is better. See issue #27473 for a
more general issue of cleaning up the design of the pending sync.
-
74f2b836
by Duncan Coutts at 2026-08-29T14:57:05+01:00
Eliminate a use of releaseCapability_ with always_wakeup
This one was purely artificial, just due to the unnecessarily strong
pre-condition. We can just weaken the precondition. The capability inbox
is non-empty so releaseCapability_ will certainly wake up a task for the
capability anyway.
We are trying to eliminate the always_wakeup parameter entirely since it
is a bit of a design wart.
-
6d52c0f4
by Duncan Coutts at 2026-08-29T14:57:05+01:00
Eliminate another use of releaseCapability_ with always_wakeup
Previusly in schedulePushWork, it checked if there are sparks for the
capability and called releaseAndWakeupCapability if there were and
releaseCapability if there were none. This is unnecessary:
releaseCapability already ensures that a task will be worken if there
are sparks available for the capability.
This also lets us remove the now unused releaseAndWakeupCapability,
eliminating another use of always_wakeup==true.
-
69f6f66b
by Duncan Coutts at 2026-08-29T14:57:05+01:00
Make prodCapability reliable, fix race condition
Also eliminate the last use of releaseCapability_ using the
always_wakeup param.
Add a Note that describes the problem and solution.
Now that prodCapability also does an interruptCapability (if the
capability is active) then we don't need to use interruptCapability as
well at call sites of prodCapability.
-
d60466e6
by Duncan Coutts at 2026-08-29T14:57:05+01:00
Eliminate the now-unused always_wakeup param from releaseCapability_
releaseCapability_ had an extra bool param: always_wakeup. This was
rather a design wart. We have now eliminated all uses of it so we can
remove the param entirely.
This will also reduce churn at call site when we add a new (rarely
used) parameter in the subsequent commit.
-
ae30e363
by Duncan Coutts at 2026-08-29T14:57:05+01:00
Add releaseCapability_ worker with a wakeup_worker modifier
Split releaseCapability_ into a worker and wrapper. The worker gains the
extra wakeup_worker parameter.
Document within releaseCapability__ the basic approach of looking for a
series of conditions in priority order and acting on them.
Then add a modifier, wakeup_worker and explain it in similar terms.
What it does is skip two of the conditions in the priority list, with
the effect that we prioritise waking up a worker task over a returning
task or bound task.
This feature is not yet used in this commit, but it will be used as
part of a scheme to allow in-RTS I/O managers in the threaded RTS. This
scheme will make use of being able to start a background worker thread,
and that will use this feature to start it promptly.
Also correct the yieldCapability docs to cover all the conditions, and
in priority order for consistency.
-
2e11787d
by Duncan Coutts at 2026-08-29T14:57:05+01:00
Move enqueueWorker next to where it is used.
It's not general purpose at all. It's very specifically crafted to work
with it's only caller: yieldCapability. It does very suprising things
like releaseCapability_, release locks and terminate threads. This logic
would be much clearer if done within yieldCapability.
-
f74519cd
by Duncan Coutts at 2026-08-29T14:57:05+01:00
Move code out of enqueueWorker and into releaseCapability_
Instead of directly releasing locks and terminating tasks, have it
return whether the enqueue was successful or not. In the latter case,
releaseCapability_ itself will release locks and terminate the task.
This makes the logic of releaseCapability_ a lot clearer. Fiddling with
tasks is what releaseCapability_ does, so it's better not to try and
encapsulate this within a helper function.
-
7ae21ca9
by Duncan Coutts at 2026-08-29T14:57:05+01:00
Clarify the logic and control flow in yieldCapability
yieldCapability is unfortunately a bit complicated. This change
restructures things slightly but should keep the behaviour the same.
Previously after calling releaseCapability_ we had a bunch of
alternatives, where in each branch we would use RELEASE_LOCK(cap->lock)
and do various things before/after the lock is released. This was a bit
hard to follow, or to extend (which we need to do).
So now we have unconditional acquire and release of the cap->lock, so
it's clear where that happens, with releaseCapability_ in between. Then
in between these steps we have the various other pre/post actions. Some
before releaseCapability_, some after while holing the lock, and some
after having released the lock.
We explain this structure in a longer comment, and refer back to the
structure from the code.