-
7bf546fc
by Simon Peyton Jones at 2026-08-31T23:48:53-04:00
Never make an absent filler for a constraint type
mkAbsentFiller used isTerminatingType to decide, but that is not enough.
Consider
class Eq a => UC a where {}
let u :: UC Int -- UC Int is a "non-terminating type"
u = error "Absent"
let e :: Eq Int -- Eq Int is a "terminating type"
e = $p1UC u
We clearly must not make a filler for `e`, because we speculatively
evaluate it. But speculatively evaluating `e` forces `u`, so we must not
make one for `u` either.
Asking isDictTy instead is not enough either, because it does not catch a
constraint hidden behind an unreduced type family application:
type family F a :: Constraint
type instance F W = TC W
a :: F W => Int -> Int -- (F W) argument is absent
Oops! Entered absent arg Arg: irred
Type: F W
So play safe and use isPredTy: never make an absent filler for any
constraint-kinded type.
Fixes #27627
-
5f474953
by Zubin Duggal at 2026-08-31T23:48:53-04:00
Add tests for absent fillers at dictionary types
T27627 a unary class whose superclass is a non-unary class
T27627a ...whose superclass is a Constraint-kinded type family
T27627b ...whose superclass is a quantified constraint
T27627c a unary class applied to itself, (UC (UC (TC a)))
T27627e a (forall b. P b) dictionary that loops
-
cd5c6bcc
by Zubin Duggal at 2026-08-31T23:48:53-04:00
An abstract TyCon may hide a unary class
A class declared in an hs-boot file is an AbstractTyCon inside the
module loop, and compiling the real declaration may reveal it to be a
UnaryClassTyCon.
- isTerminatingType returned True for such AbstractTyCons
- IfaceToCore set the unary flag to False in the DFunId
So we could end up speculating bottom dictionaries because inside a module
loop we see an UnaryClassTyCon as an AbstractTyCon
Use isTerminatingTyCon, which returns False for an abstract TyCon.
The Bool in DFunId is now a cache for isTerminatingTyCon, set in
mkDFunIdDetails.
Fixes #27704
-
abfc224a
by Zubin Duggal at 2026-08-31T23:48:53-04:00
Specialise: don't replace dead args with absent fillers
specHeader decides an argument is dead by calling isDeadBinder on a binder of
the /optimised RHS/, then applies the filler to the /stable unfolding/
template instead. The two may differ, so the argument can be dead in
the RHS and not in the template.
The specialised function's unfolding then has an absent filler, and any call
site that inlines it evaluates the error thunk.
Dropping dead args in the specialiser is rarely worth it, to quote Simon,
"The later worker/wrapper pass will pick up the dead arg later if it is really dead. Keeps the specialiser simpler."
So instead of trying to check if the arg really is dead in the stable unfolding,
just drop the logic for dropping dead args in the specialiser altogeher.
Fixes #27703
-
1557fd1c
by Zubin Duggal at 2026-08-31T23:48:53-04:00
CorePrep: don't speculate a call across an hs-boot edge
We take care not to evaluate things that might be bottom, like a
looping dictionary group, but our analysis is defeated by boot files.
We only track recursion within a module, so two dictionaries that
depend on each other across a module loop each look non-recursive, and
we might speculate them.
Any recursion we cannot see must cross an hs-boot edge, so refuse to
speculate calls that cross one.
Fixes #27717
-
4117e5ae
by Wolfgang Jeltsch at 2026-08-31T23:49:35-04:00
Incorporate the `rethrowSTM` reexport into the `stm` submodule
-
4bfbf5c8
by ARATA Mizuki at 2026-09-01T18:55:08-04:00
testsuite: Fix out-of-bounds access in T3586
unsafeRead and unsafeWrite use 0-based index.
Looking at #3586, the expected output seems to be 2.8e8.
Addresses #27596
-
3f9db6d4
by ARATA Mizuki at 2026-09-01T18:55:08-04:00
testsuite: Fix out-of-bounds access in T21305
writeInt64Array# takes an index measured in units of Int64 elements.
Fixes #27596
-
44d7788f
by Simon Peyton Jones at 2026-09-01T18:55:53-04:00
Fix buglet in INLINE-arity calculation for pattern synonyms
This fixes #27744.
The buglet was accidentally introduced by
commit 3a0f9a51c1dacc474c7fd128082edd8bf4081256
Author: Simon Peyton Jones <simon.peytonjones@gmail.com>
Date: Sat Aug 1 00:13:02 2026 +0100
Fix three bugs related to required type args and INLINE pragmas
I failed to find all the calls to `addInlinePragArity`!
-
df058f1d
by Simon Jakobi at 2026-09-02T07:18:06-04:00
testsuite: Migrate perf tests off collect_compiler_stats('all')
The 'all' metric argument applies a single tolerance to bytes
allocated, max_bytes_used and peak_megabytes_allocated, although their
noise profiles are incompatible (#27653): allocations are nearly
deterministic, residency needs 10-20%, and peak is quantized to 1 MB.
Any single tolerance is too tight for one metric or too slack for
another. This migrates the remaining users of 'all' (and of the 'all'
default) to explicit per-metric collection, ahead of removing 'all'
from the driver.
peak_megabytes_allocated is dropped everywhere: its 1 MB granularity
makes tight relative windows meaningless (#27613), and it is sensitive
to GC timing. In #27489 it drifted by -5.3% while max_bytes_used moved
by less than 0.1%. Where a test guards a memory property,
max_bytes_used covers it at byte granularity.
Where the motivating ticket was about compile-time memory
(T11545, T15304, T26425), residency remains gated via max_bytes_used,
now with a residency-appropriate tolerance.
max_bytes_used is dropped where residency was only ever an accident of
'all':
* T15630, T15630a, T20261: the underlying tickets (#15630, #20261)
contain no memory data at all. One is a simplifier-ticks blowup and
the other is stated entirely in allocation numbers, so the 20%
window never had teeth.
* T21839c: #21839's measurements show residency essentially flat
(+0.16%) while allocations moved +7%, so allocations are the
discriminating metric. They are already gated at 1% via
collect_compiler_runtime. The ghc/max gate had previously broken CI
spuriously (9fd11585eb widened it from 1% to 10% for that reason).
Allocation tolerances are tightened to the testsuite's conventional 2%
where 'all' previously left them at 10-20%.
Assisted-by: Claude Fable 5
-
72dd2432
by Simon Jakobi at 2026-09-02T07:18:06-04:00
testsuite: Remove the 'all' metric argument of collect_stats
'all' gated bytes allocated, max_bytes_used and peak_megabytes_allocated
at a single tolerance, although their noise profiles are incompatible,
making such tests either flaky or toothless (#27653).
Closes #27653.
Assisted-by: Claude Fable 5
-
2228cb30
by Simon Jakobi at 2026-09-02T07:18:06-04:00
testsuite: Make the deviation argument of collect_stats mandatory
Almost every caller passes an explicit tolerance matched to the
metric's noise profile, and the silent 20% default is far slacker than
'bytes allocated' merits. Only two tests relied on it. They now state
their tolerance explicitly:
large-project gets 10%, in line with other large compile-time tests.
T9848 gets 2%: its metric is byte-for-byte deterministic across CI
jobs and platforms of a given test_env, has drifted only about 2.5%
since 2015, and the fusion failure it guards against would show up as
a roughly +30000% jump.
Assisted-by: Claude Fable 5
-
56291fc5
by Cheng Shao at 2026-09-02T07:18:45-04:00
autoconf/ghc-toolchain: bump llvm upper bound to support llvm 23
This commit bumps llvm upper bound to support llvm 23.
-
20eb3f41
by Cheng Shao at 2026-09-02T07:18:45-04:00
rts: fix compilation issues with clang 23
clang 23 has broadened `-Wall`/`-Wextra` ranges, exposing some minor
issues in the rts when building with validate flavours:
- Unused locals
- `#pragma GCC diagnostic pop` mismatch
This commit fixes those.
-
9e7739ab
by Simon Jakobi at 2026-09-03T13:45:43-04:00
Add -Wimplicit-field-strictness (#16836)
This opt-in warning fires when a data constructor field lacks an
explicit strictness annotation (`!` or `~`). It complements the
LazyFieldAnnotations extension (4762a8bf30f) from GHC proposal 752,
which makes `~` annotations available for this purpose.
Deciding which fields to report requires their levity, so the check
runs after typechecking. To keep the noise down, the diagnostic is
emitted once per data declaration, grouped by constructor.
Closes #16836.
Assisted-by: Claude Fable 5
-
6a44d591
by mangoiv at 2026-09-03T13:46:22-04:00
simplifier: remove a bogus `assert` in `rebuild_app'` in `GHC.CoreToStg.Prep.cpeApp`.
Prior to this commit
commit 08bc245be70d95801bc1138804ed1de9474fbdc0
Author: sheaf <sam.derbyshire@gmail.com>
Date: Sat Feb 28 16:30:43 2026 +0100
Clean up join points, casts & ticks
This commit shores up the logic dealing with casts and ticks occurring
in between a join point binding and a jump
any `PlaceRuntime` ticks we observed were profiling ticks, even though that
isn't necessary.
The more liberal rules in the above commit allow e.g. breakpoint ticks
(which are valid PlaceRuntime ticks) to legitimately appear in an argument
position.
Fixes #27556
-
eaa95a6f
by Andrei Borzenkov at 2026-09-04T08:08:40-04:00
Parentheses in prefix GADT constructors (#27423)
Updated `splitLHsGadtTy` to allow looking
through the parentheses for inner binders. General example
of a code pattern that's allowed now:
data S a where
MkS :: (forall a. S a)
That should work now with any combination of nested
foralls and parentheses.
We don't perform parenthesis unwrapping for record
GADT constructors in accordance with GHC Proposal #402.
To this end `con_inner_bndrs` no longer stores plain forall
telescopes: `[HsForAllTelescope pass]` is replaced with
`[LHsGadtTelescope pass]`, a new `HsArg`-style type whose
`HsGadtForAll` holds an inner telescope and whose `HsGadtPar`
holds a pair of parentheses. The parentheses carry no meaning
for renaming or type checking; the only reason to record them
is exact-printing.
Updated `pprConDecl` to improve the `parse == parse . ppr . parse`
property of GADT pretty-printing.
The pretty printer can now output code that's similar to this:
data T a where
MkT1 :: (forall a. T a)
MkT2 :: forall . forall a. T a
These are special cases of inner forall binders
for prefix GADT constructors, when we have either
implicit or zero explicit outer binders.
-
639456a8
by Evgeny Malyshev at 2026-09-04T08:09:25-04:00
HsPat: Add spacing to unary tuple patterns
Use hsep when printing unary boxed TuplePat applications so that the
MkSolo constructor is separated from its argument. The old hcat rendered
wildcard and parenthesized arguments as MkSolo_ and MkSolo().
Add T27034 to exercise the HsPat splice-dump path with wildcard and
variable arguments, and update existing affected golden output.
Fixes #27034
Assisted-By: OpenAI Codex
-
2908ab36
by Simon Jakobi at 2026-09-05T07:22:17-04:00
X86 NCG: use btr/bts/btc for single-bit operations
Previously the Cmm patterns
x & ~(1 << i)
x | (1 << i)
x ^ (1 << i)
compiled to mov/shl/not/and-style sequences of 3-4 instructions. Now
they compile to a single btr, bts or btc, matching what C compilers
produce.
When the bit index is a literal, constant folding has already collapsed
these patterns into ones with a literal mask, such as
x & 0xfffffeffffffffff for x & ~(1 << 40). Such masks are now also
compiled to a bit-test instruction when they don't fit in an imm32 and
would otherwise have to be loaded into a register first.
For a variable bit index, this applies only when the shift is unchecked
(uncheckedShiftL#, Data.Bits.unsafeShiftL): the bounds-checked shiftL
used by e.g. the default clearBit/setBit/complementBit implementations
wraps the shift in a bounds mask that this optimisation does not see
through. With a literal index, the bounds mask is constant-folded away,
so the checked operations benefit too.
See Note [Bit-test instructions] in GHC.CmmToAsm.X86.CodeGen.
Fixes #25233.
Assisted-by: Claude Fable 5
-
c673ecf0
by Simon Peyton Jones at 2026-09-05T07:22:59-04:00
Move HsStatic free-var test to typechecker
A `static` form should have no free *term* variables, but it
can have free *type* variables. Alas, the renamer does not really know what
is a term variable and what is a type variable, because of required type
arguments. This patch moves the test to the typechecker, which does know.
Addresses #27664
-
5939ceaf
by sheaf at 2026-09-05T22:06:18-04:00
Windows: enforce path convention in ./configure
As detailed in Note [MSYS paths] in Hadrian.Utilities, the standing
convention (using Windows-style paths with forward slashes) is now
enforced in ./configure instead of within Hadrian, removing the need for
'cygpath' calls within Hadrian.
Fixes #26683
-
a6061455
by sheaf at 2026-09-05T22:06:18-04:00
Hadrian: introduce ExeSpawnPath
Specific details about the filepath used to specify the executable to
spawn with CreateProcess matters on Windows: whether we use forward or
backward slashes, a leading ./, or an absolute path changes how the
executable is found.
This commit introduces 'ExeSpawnPath' which is a path that is guaranteed
to be found when spawning a process. All command invocations now go
through this type to ensure the path has been properly sanitised.
See Note [NeedCurrentDirectoryForExePath] in Hadrian.Utilities.
The same treatment is applied to hsc2hs. Updates hsc2hs submodule.
-
e28313e3
by sheaf at 2026-09-05T22:07:06-04:00
Preserve tick ordering in 'tickTickedExpr'
'GHC.Core.Utils.tickTickedExpr' tries to combine a tick 't1' into an
existing stack of ticks 't2s'. There are two situations:
1. 't1' is subsumed by a tick in 't2s': drop it.
2. A tick in 't2s' is subsumed by 't1', say 't2'.
This commit ensures that in case (2) we keep 't1' on the outside instead
of replacing 't2' at its position in the stack. This avoids re-ordering
source notes (which was the cause of #27749).
This fixes a regression introduced in 2dadf3b0d05.
Fixes #27749
-
3172f557
by sheaf at 2026-09-05T22:07:06-04:00
Consistently prefer local source note ticks
GHC.Cmm.DebugBlock.cmmDebugGen (DWARF annotations) and
GHC.Stg.Debug.quickSourcePos (-finfo-table-map) both contained logic to
prioritise source note ticks from the current module.
This commit commons up this logic and propagates it to a third consumer:
IPE stack frames, in GHC.Driver.GenerateCgIPEStub.
See the new function GHC.Types.Tickish.bestSourceNote.
-
192be0b6
by Luite Stegeman at 2026-09-07T19:43:36-04:00
rts: fix ctoi_tuple_spill_words getting out of sync
Fix a few places that were not updating ctoi_tuple_spill_words
correctly, leading to corruption/crashes when dealing with large
unboxed tuples in bytecode:
- captureContinuationAndAbort
- findRetryFrameHelper/findAtomicallyFrameHelper
- interpretBCO bci_BRK_FUN
fixes #27633
-
604eb43c
by Simon Jakobi at 2026-09-07T19:44:15-04:00
Hadrian: don't capture the testsuite driver's output (#27780)
a6061455d54 switched the Testsuite RunTest case from Shake's cmd to the
cmd' wrapper. cmd' always captures stdout and stderr, and since the
caller asks for Exit, it returns without dumping what it captured. As a
result the testsuite output no longer appears in CI job logs: failures,
performance metrics and the summary were lost with the job.
Use cmdExe, the uncaptured cmd, as the other plain-cmd sites in that
commit do.
Fixes #27780.
Assisted-by: Claude Fable 5.1
-
06eee015
by Luite Stegeman at 2026-09-08T06:03:06+02:00
rts: Fix missing memory barrier in eval_thunk_selector (#27477)
unchain_thunk_selectors() was missing an ACQUIRE_LOAD for the
indirectee, leading to segfaults and corruption during GC on
weakly-ordered architectures.
Fixes #27477
-
f8b2bd8f
by Alan Zimmerman at 2026-09-08T09:39:12-04:00
EPA Fix HsCmdDo exact print with comments
Exact printing of HsCmdDo was ignoring the location for the do
statements, and this is an annotation that can have comments in it.
Update it so we print the statements as a unit, including any
comments.
Also add the result of auditing that we capture comments in all needed
places, noting that the remaining Anno SrcSpan instances are benign.
-
430967ab
by Duncan Coutts at 2026-09-09T11:07:01+01:00
Reorder cmm decls in HeapStackCheck for a better logical grouping
And put more section headers in to deliniate the groups.
We're about to add more here, so better to organise it first.
-
344080fd
by Duncan Coutts at 2026-09-09T11:07:01+01:00
Add raisePrimIOException and add it to RTS<->ghc-internal API
The raisePrimIOException is a new helper function that I/O primops will
use to help them report I/O errors. This is implemented in Haskell
(since that's a lot easier), but has a calling convention that is
easy(ish) to use from Cmm in the I/O primops.
So we add it to the RTS API struct, and since we'll use it from Cmm we
also need a field accessor macro for cmm (in deriveConstants).
See the Note about how we cannot have nice things due to async
exceptions and thunks preventing us from using catch.
-
cfdb3399
by Duncan Coutts at 2026-09-09T11:07:01+01:00
Add new blocking functions for I/O primops
See the Note [Thread blocking for new I/O primops],
and the Note [Calling convention for raisePrimIOException].
The point is, it will allow us to report synchronous exceptions from I/O
primops, and do so much more flexibly.
Previously the I/O managers could only report async exceptions and only
nullary exceptions. This was OK historically, but no good as we add
more I/O managers and expand the range of I/O operations we support.
-
f09c09d8
by Duncan Coutts at 2026-09-09T11:07:01+01:00
Change the encoding of results from the I/O manager to I/O primops
Previously we just had async continue or heap overflow.
We now extend what we can report with synchronous success, and
synchronous failure with an errno.
See Note [Encoding of result of I/O manager operations]
We don't use these two new cases yet, but we will. In particular an
epoll I/O manager needs to be able to report synchronous success or
failure for waitRead#/waitWrite#.
-
30f074d9
by Duncan Coutts at 2026-09-09T11:07:01+01:00
Switch waitRead/Write# to use new blocking return frames
and update the I/O managers to set the result before resuming the
blocked threads.
This makes it possible for I/O managers to report synchronous exceptions
from the I/O primops, but that will be done in a subsequent commit.
-
a8786ced
by Duncan Coutts at 2026-09-09T11:07:01+01:00
Switch Poll and Select I/O managers to report sync exceptions
rather than using raiseAsync with blockedOnBadFD_closure.
This uses the new mechanism in the blocking frame return code to report
synchronous exceptions.
-
7a9df79a
by Duncan Coutts at 2026-09-09T11:07:01+01:00
Remove now-unused blockedOnBadFD
It was previously thrown by the select and poll I/O managers, but now
they use raisePrimIOException (with an EBADF errno).
-
6b0424e3
by Duncan Coutts at 2026-09-09T11:07:01+01:00
Improve the docs for delay# waitRead# and waitWrite#
Document that the waitRead/Write# can throw exceptions (this was true
before too), and that all of them are async exception cancellation
points.
-
afeb2c30
by Duncan Coutts at 2026-09-09T20:08:58-04:00
Enable printf warnings for trace functions and fix resulting warnings
Most of the existing printf-style functions are annotated with
attributes to enable gcc/clang warnings for the printf format string,
but several trace functions in Trace.h were missing this annotation.
Enable them, and fix the resulting warnings.
-
9a442c93
by Alan Zimmerman at 2026-09-09T20:09:36-04:00
EPA: Exact print ConDeclGADT without custom enterAnn
!16321 brought in explicit capture of parens in a ConDeclGADT.
The ExactPrint update introduced a modification of the fundamental
function in exact printing, `enterAnn`, by splitting it into a version
allowing injection of functionality normally handled by the
ExactPrint class methods.
This commit refactors that code, to restore the prior `enterAnn`
version, by following the convention in ExactPrint of introducing a
helper data structure with its own `ExactPrint` instance to achieve
the same effect.
-
a01d8943
by Duncan Coutts at 2026-09-13T22:55:59+01:00
Make signal handling be a responsibility of the I/O manager(s)
Previously it was scattered between I/O managers and the scheduler, and
especially the scheduler's deadlock detection.
Previously the scheduler would poll for pending signals each iteration
of the scheduler loop. The scheduler also had some hairy signal
functionality in the deadlock detection: in the non-threaded RTS (only)
if there were still no threads running after deadlock detection then it
would block waiting for signals.
But signals can and (in my opinion) should be thought of as just a funny
kind of I/O, and thus should be a responsibility of the I/O manager.
So now we have the I/O managers poll for signals when they are polling
for I/O completion (and removing the separate poll in the scheduler).
And when I/O managers block waiting for I/O then they now also start
signal handlers if they get interrupted by a signal. Crucially, if there
is no pending I/O or timers, the awaitCompletedTimeoutsOrIO will still
block waiting for signals.
This patch puts us into an intermediate state: it temporarily breaks
deadlock detection in the non-threaded RTS. The waiting on I/O currently
happens before deadlock detection. This means we'll now wait forever on
signals before doing deadlock detection. We need to move waiting after
deadlock detection. We'll do that in a later patch.
-
42b765b0
by Duncan Coutts at 2026-09-13T22:55:59+01:00
Clean up the RTS internal signal handling API
Now that the I/O manager is responsible for signals, we can simplify the
API we present for signal handling.
We now just need startPendingSignalHandlers, which is called from the
I/O managers. We can get rid of awaitUserSignals. We also don't need
RtsSignals.h to re-export the platform-specific posix/Signals.h or
win32/ConsoleHandler.h
We can also hide more of the implementation of signals. Less has to be
exposed in posix/Signals.h or win32/ConsoleHandler.h. Indeed,
posix/Signals.h becomes empty and we remove it. Partly this is because
we don't need inline functions (or macros) in the interface.
Also remove signal_handlers from RTS ABI exported symbols list. It does
not appear to have any users in the core libs, and its really an
internal implementation detail. It should not be exposed unless it's
really necessary.
-
c39b3472
by Duncan Coutts at 2026-09-13T22:56:00+01:00
In the scheduler, move I/O blocking after deadlock detection
To make deadlock detection effective in the non-threaded RTS when there
are deadlocked threads and other unrelated threads waiting on I/O, we
need to arrange to do deadlock detection before we block in scheduler
to wait on I/O.
The solution is to:
1. adjust scheduleFindWork, which runs before deadlock detection, to
only poll for I/O and not block; and
2. add a step after deadlock detection to wait on I/O if there are
still no threads to run (and there's any I/O or timeouts outstanding)
The scheduleCheckBlockedThreads is now so simple that it made more sense
to inline it into scheduleFindWork.
-
08ab9618
by Duncan Coutts at 2026-09-13T22:56:00+01:00
Remove bogus anyPendingTimeoutsOrIO guard from scheduleDetectDeadlock
The deadlock detection was only invoked if both of these conditions
hold:
1. the run queue is empty
2. there is no pending I/O or timeouts
The second condition is unnecessary. The deadlock detection mechanism
can find deadlocks even if there are other threads waiting on I/O or
timers. Having this extra condition means that we fail to detect
blocked threads if there are any threads waiting on I/O or timers.
Part of fixing issue #26408
-
c53825b0
by Duncan Coutts at 2026-09-13T22:56:00+01:00
Don't consider pending I/O for early context switch optimisation
Context switches are normally initiated by the timer signal. If however
the user specifies "context switch as often as possible", with +RTS -C0
then the scheduler arranges for an early context switch (when it's just
about to run a Haskell thread).
Context switching very often is expensive, so as an optimisation there
cases where we do not arrange an early context switch:
1. if there's no other threads to run
2. if there is no pending I/O or timers
This patch eliminates case 2, leaving only case 1.
The rationale is as follows. The use of this was inconsistent across
platforms and threaded/non-threaded RTS ways. It only worked on the
non-threaded RTS and on Windows only worked for the win32-legacy I/O
manager. On all other combinations anyPendingTimeoutsOrIO would always
return false. The fact that nobody noticed and complained about this
inconsistency suggests that the feature is not relied upon.
If however it turns out that applications do rely on this, then the
proper thing to do is not to restore this check, but to add a new I/O
manager hint function that returns if there is any pending events that
are likely to happen *soon*: for example timeouts expiring within one
timeslice, or I/O waits on things likely to complete soon like disk I/O,
but not for example socket/pipe I/O.
The motivation to avoid this use of anyPendingTimeoutsOrIO is to
allow us to eliminate anyPendingTimeoutsOrIO entirely. All other uses
of this are just guards on {await,poll}CompletedTimeoutsOrIO and
the guards can safely be folded into those functions. This will better
cope with some I/O managers having no proper implementation of
anyPendingTimeoutsOrIO.
Ultimately this will let us simplify the scheduler which currently has
to have special #ifdef mingw32_HOST_OS cases to cope with the lack of a
working anyPendingTimeoutsOrIO for some Windows I/O managers
-
c194dbdc
by Duncan Coutts at 2026-09-13T22:56:00+01:00
Remove anyPendingTimeoutsOrIO guarding {poll,await}CompletedTimeoutsOrIO
Previously the API of the I/O manager used a two step process: check
anyPendingTimeoutsOrIO and then call {poll,await}CompletedTimeoutsOrIO.
This was primarily there as a performance thing, to cheaply check if we
need to do anything.
And then because anyPendingTimeoutsOrIO existed, it was used for other
things too. We have now eliminated the other uses, and are just left
with the performance pattern.
But this was problematic because not all I/O managers correctly
implement anyPendingTimeoutsOrIO (specifically the win32 ones), and now
that we also make I/O managers responsible for signals then we need to
poll/await even if there is no pending I/O or timeouts. If there is no
pending I/O or timeouts then await needs to degenerate to just waiting
forever for any signals.
-
8a2963c0
by Duncan Coutts at 2026-09-13T22:56:00+01:00
Remove anyPendingTimeoutsOrIO, it is no longer used
And this avoids the problems arising from the win32 I/O managers having
had a bogus implementation.
-
d3b6a429
by Duncan Coutts at 2026-09-13T22:56:00+01:00
Remove second scheduler call to awaitCompletedTimeoutsOrIO
Previously awaitCompletedTimeoutsOrIO was called both before and after
deadlock detection in the scheduler. The reason for that was that the
win32 I/O managers had a bogus implementation of anyPendingTimeoutsOrIO
and this was used to guard the call of awaitCompletedTimeoutsOrIO prior
to deadlock detection. This meant the first call site was never actually
called when using the win32 I/O managers. This was the reason for the
second call: the first one was never used. What a mess.
So now we have a simple design in the scheduler:
1. poll for completed I/O, timers or signals
2. if no runnable threads: do deadlock detection
3. if still no runnable threads: block waiting for I/O, timers or
signals.
-
c3afe473
by Duncan Coutts at 2026-09-13T22:56:00+01:00
Lift emptyRunQueue guard out of scheduleDetectDeadlock
this improved the clarity of the logic when reading the scheduler code.
-
498ada58
by Duncan Coutts at 2026-09-13T22:56:00+01:00
Make non-threaded deadlock detection also rely on idle GC
Only do deadlock detection GC when idle GC kicks in. This also relies on
using wakeUpRts, so now do this unconditionally. Previously wakeUpRts
was for the threaded rts only.
-
df7b9ead
by Duncan Coutts at 2026-09-13T22:56:00+01:00
Enable idle GC by default on non-threaded RTS
The behaviour is now uniform between the threaded and non-threaded RTS
ways. The deadlock detection now relies on idle GC for both threaded
and non-threaded ways. Previously deadlock detection did not rely on
idle GC for the non-threaded way.
Also tweak test T7275 to account for idle GC. This test's output is
sensitive to the number of major GCs run. Since this commit enables idle
GC for the non-threaded RTS, for this test that increases the number of
major GCs, since the test program is frequently idle for more than 300ms.
-
d41dd174
by Duncan Coutts at 2026-09-13T22:56:00+01:00
Fix state of idle GC control vars with +RTS -V0
Currently when the user uses +RTS -I0, then doIdleGC is set to false.
But if the master tick interval -V is set to 0 then the idleGCDelayTime
was being set to 0 but doIdleGC was not being set to false, which is
inconsistent, and almost certainly buggy.
-
596ddecb
by Duncan Coutts at 2026-09-13T22:56:00+01:00
Add a long Note [Deadlock detection]
It describes the historical and modern designs and their trade-offs.
The point is we've now unified the code for deadlock detection between
the threaded and non-threaded ways, by changing the non-threaded to
follow the same design as the threaded.
-
21699497
by Duncan Coutts at 2026-09-13T22:56:00+01:00
Add a test for deadlock detection, issue #26408
-
4423cc6e
by Duncan Coutts at 2026-09-13T22:56:00+01:00
Update the user guide with the revised idle GC behaviour
i.e. it's now not just for the threaded RTS, but general.
Also document the fact that disabling idle GC also disables deadlock
detection.
And add a changelog entry.
-
47bf6704
by Duncan Coutts at 2026-09-13T22:56:00+01:00
Move idle GC tracking to its own file, step 1 of 2
Add a new IdleGC.{c,h} module. Move the RecentActivity type,
recent_activity variable and {get,set}RecentActivity wrappers.
-
c7ae301d
by Duncan Coutts at 2026-09-13T22:56:00+01:00
Move idle GC functionality to its own file, step 2 of 2
We can now fully encapsulate the recent_activity state, exposing just
three hooks, and one query used in the scheduler and timer tick.
This is a prelude to making changes to the implementation of tracking
when to perform an idle gc.
-
74c8aefc
by Duncan Coutts at 2026-09-14T16:18:17+01:00
Add idle GC mode for !HAVE_PREEMPTION
We have to do idle GC (and thus deadlock detection) differently when we
do not have the ticker and the ability to interrupt the I/O manager
when it's blocking.