I think Simon is offering a free service here. If you don't want to see these reviews for your MRs, then don't request them. I'd surely find them helpful, although I am fortunate enough to have a Claude subscription of my own. I don't think discarding the idea on the grounds that "LLM reviews are too noisy" is sensible. Firstly, linters and warnings are noisy as well, yet they bring our validate pipelines to a halt. This is *much* noisier than an LLM review you requested but now don't want to read for some reason. Secondly, if a contributor requests an LLM "review" that has a false positive (which happens, but far less frequently than true positives!), then either the contributor is able to label it as such or will discuss it with the maintainers. At the end of such a discussion, often, both contributor and maintainer will have learned something. Am Mo., 27. Juli 2026 um 22:54 Uhr schrieb Cheng Shao via ghc-devs < ghc-devs@haskell.org>:
<detail> tag with a big "oh fable found this oh sol found that here's full diagnosis and fix" caption is not a visual noise, it's a cognitive one. a human reviewer, no matter pro or against llms, would have this temptation with one click away, and this is a constant distraction when they spend their own personal time looking at that patch. and regardless of if the reviewer is organic or not, a pre-existing blob of text there will diverge their attention towards what's already been spotted, away from what might not have been spotted yet (deeper architectural issues, something shouldn't be done in the first place, etc etc).
again, i don't think the gitlab instance is the right data plane for llm generated reviews, or for "experiments" that you propose.
On Mon, Jul 27, 2026 at 10:41 PM Simon Jakobi <simon.jakobi@gmail.com> wrote:
Hi Cheng,
Thanks for your kind reply! :)
Am Mo., 27. Juli 2026 um 22:18 Uhr schrieb Cheng Shao <terrorjack@type.dance>:
i think simon jakobi's proposal is not installing some label-triggered
auto-review bot on our gitlab instance, it's about him using his own account to paste llm-generated reviews to patches where authors explicitly request so.
thanks simon, it's indeed a generous offer, but i don't think the
gitlab instance is the right place for pasting large blobs of llm-generated reviews. imho it's better to move these to e.g. a self-hosted gerrit instance, or a discord/matrix channel etc, where people who want this service can ping you, get a piece of review to take time to digest and improve their patch accordingly.
If the problem is with visual noise, I don't mind putting the blobs into `<details>` tags so they don't take up as much space.
However if the goal is to reduce maintainers' workload, I think the MR is still the best place, so maintainers can use the presence of the LLM-generated review as a signal.
A decent amount of visibility would probably also help to improve the LLM review system more quickly.
i personally find llm-generated reviews very useful in my own work, but the signal-to-noise ratio of review outputs by even the best models/harnesses today is simply low, the insight needs to be mined from the text, and this can be a huge distraction and prevent a human reviewer from entering the flow when looking at patches.
If the signal-to-noise ratio is indeed so bad, I'm ready to improve the system until the ratio is more acceptable. It would be great to have some collaborators in this, but I'll also give a try on my own.
Cheers, Simon
cheers, cheng
On Mon, Jul 27, 2026 at 9:40 PM Zubin Duggal via ghc-devs <
ghc-devs@haskell.org> wrote:
I am strongly opposed to this. I also don't see the value in it. If you want to ask an LLM to review your code, you can probably do this much quicker by just asking one. Today you are never more than a few clicks away from a text box that will put you in direct contact with one.
If you want to maintain a shared repository of common prompts, skills, guidance etc for LLMs regarding work on GHC, I have no objection to this (I would prefer to only have documentation and prose targetting humans in the main GHC repository though).
In fact I would like to propose a policy in the exact opposite direction:
Disallow _purely autonomous_ systems from posting on the GHC Gitlab, unless such a system is cleared in advance by the administrators of the instance and/or contributors to GHC.
What this won't disallow:
* Getting an LLM to isolate and create a minimal reproducer or bug description and posting this.
* Getting an LLM to propose a patch or fix for a particular issue
* Asking an LLM about the details of how some particular external system works, and then posting the results.
* Asking an LLM to catalogue a bunch of related issues and extract a common pattern
* etc.
What this will disallow:
Autonomous systems without humans in the loop writing and interacting with the gitlab.
My reasons
1. Code review is a tool for building a shared understanding and a common mental model of the codebase and problem domain. An LLM model cannot do that, or even if you believe it can, this internal model is lost as soon as that particular session is garbage collected. At best you can try to reconstruct some approximation of this model later by feeding the same inputs into a fresh session on a probably different model.
2. I would find it a very sad if the GHC Gitlab became a place where bots talk to other bots who act upon and those comments to make addtional changes in response to this in an uncontrolled feedback loop. I predict that in the limit any such feedback loop becomes quickly detached from the original problem and devolves into a lot of noise totally divorced from any real human concerns.
We can already see examples of how this goes wrong **today** Take a look at a prominent example of a codebase that went "all in" on mostly unsupervised autonomous LLM development, the javascript runtime Bun: https://github.com/oven-sh/bun/pulls
It would kill any motivation I would have to contribute to GHC and collaborate with others if the merge request queue looked anything approximating that. I use the gitlab to talk to and collaborate with actual humans, who I hope will actually engage their brains and help get us to a better and more corret mental model of the space.
If I wanted to talk to an LLM, I have countless other options for ways in which to do that.
3. Prompt injection. Gitlab comments are an untrusted input source and LLMs fundamentally cannot discriminate between instructions and data. You can imagine prompt injection attacks where hidden text in a gitlab comment, attached bug reproducer, link etc. gets the model to do various destructive things, like spamming the gitlab, editing/deleting your prior comments, mining bitcoin on the CI runners, attempting to inject malicious code into the compiler, testcases, reproducers etc.
That doesn't mean that there isn't value for LLM assisted autonomous processes on the gitlab, given that they are desiged properly keeping in mind the concerns about human centeredness, resource usage and prompt injection.
Here are some examples of LLM assisted system/bots that could actually be useful if audited and pre-approved with the community:
* A bot that automatically bisects regressions and identifies the change responsible for it. In principle this could be done with a purely deterministic script, but often you run into problems with the toolchain or commits not being individually buildable etc that require some adjustements on the fly.
* A bot that automatically minimises reproducers and posts the results after the testcase can been verified on a deterministic system
On 26/07/27 19:38, Simon Jakobi via ghc-devs wrote:
Hi devs!
Inspired by the recent discussion on the LLM policy, I'd like to try introducing LLM-generated code reviews in GHC, as an opt-in service.
My motivation:
* Catch issues before human reviewers spend time on them. * Reduce the number of bugs merged.
What this is — and what it isn't --------------------------------
This is tool output: requested on the MR, produced by a model, posted
verbatim
and clearly marked as machine-generated. This is not a review in the sense of draft LLM policy.
I'm not vouching for any of it. Think of it as a linter the author could have run themselves, except on my LLM budget. Nobody is required to act on it even read it.
I mention this explicitly because the policy asks reviewers to take full responsibility for each line of their review. I can't do that here — it would mean checking every finding myself, which defeats the purpose and would restrict the service to areas of the compiler I already know well.
How to request one ------------------
* Ping me on the MR with a comment like "@sjakobi llm-review please". * Feel free to request particular aspects the review should cover. * Authors can request a review for their own MRs. Maintainers can request one on any MR. * If I have enough usage left on my plan, I'll confirm that I'm on it. * I'll then prompt the model, and post the result on the MR. * Feedback on the quality of the review is of course very welcome.
What I promise --------------
* Every review is labelled as machine-generated, with model, effort level and the exact prompt used. * Reviews are posted verbatim. I don't edit or curate them, so what you see is exactly what the model produced. * Only the public MR diff and public repository context go to the model. Nothing else. * I'll keep working on the prompts and the choice of model to improve the signal-to-noise ratio. * If the signal-to-noise ratio stays inacceptably bad, we can simply stop this service.
Review coverage ---------------
I'll probably start with a fairly basic review prompt, primarily aimed at correctness issues.
Possible extensions:
* Documentation consistency: check that documentation stays consistent, including Notes elsewhere that reference the changed code. * Performance: check for performance issues — potentially including checking the -ddump-simpl output for the changed code for unnecessary allocation etc. * Feel free to suggest anything else!
Format ------
Initially a single comment per review. Once we reach a good signal-to-noise ratio, we can consider inline comments on the diff.
Caveats -------
This is a volunteer service I intend to provide in my spare time. If I'm AFK or on vacation, reviews will take longer. If I can't keep up with requests, I will prioritise which MRs I run. I might take a break entirely.
Getting involved ----------------
Is anyone interested in joining this effort? If so we could form a team (@llm-reviewers?), and share prompting techniques, model choices, etc.
If the service turns out to be genuinely useful, we can look into automating it and – if necessary – possible funding.
Cheers, Simon _______________________________________________ ghc-devs mailing list -- ghc-devs@haskell.org To unsubscribe send an email to ghc-devs-leave@haskell.org
ghc-devs mailing list -- ghc-devs@haskell.org To unsubscribe send an email to ghc-devs-leave@haskell.org
_______________________________________________ ghc-devs mailing list -- ghc-devs@haskell.org To unsubscribe send an email to ghc-devs-leave@haskell.org