critical
#29454
Update dependency org.jsoup:jsoup to v1.23.1 [SECURITY]
This PR contains the following updates:
| Package | Type | Update | Change |
| --- | --- | --- | --- |
| org.jsoup:jsoup (source) | compile | minor | 1.22.2 → 1.23.1 |
---
⚠️ Warning
Some dependencies could not be looked up. Check the warning logs for more information.
---
jsoup: Cleaner may expose markup with custom raw-text elements
CVE-2026-71497 / GHSA-pmhh-3w7g-xqp8
Details
When a custom Safelist permits certain raw-text elements, jsoup may incorrectly sanitize malformed HTML containing a tag name that ends in a control character. The tag may acquire the parsing behavior of a different element, causing content that should remain text to be emitted as active markup after serialization and potentially allowing XSS.
jsoup’s built-in Safelists are unaffected.
##### Patches
Upgrade to jsoup 1.23.1.
##### Workarounds
Until upgrading, do not permit raw-text elements in custom Safelists used to clean untrusted HTML.
##### Additional security considerations
This fix addresses malformed tag-name handling only.
Permitting raw-text elements in a custom Safelist does not make their contents inherently safe. For example, applications that permit style must apply appropriate CSS safeguards separately, because jsoup does not parse or sanitize CSS.
Severity
- CVSS Score: 4.7 / 10 (Medium)
- Vector String: CVSS:3.1/AV:N/AC:H/PR:N/UI:R/S:C/C:L/I:L/A:N
References
- https://github.com/jhy/jsoup/security/advisories/GHSA-pmhh-3w7g-xqp8
- https://github.com/jhy/jsoup/issues/2538
- https://github.com/jhy/jsoup/commit/92f1aca552548b484bc7d4b94c51e48b8e6eca70
- https://github.com/jhy/jsoup
- https://github.com/jhy/jsoup/releases/tag/jsoup-1.23.1
This data is provided by OSV and the GitHub Advisory Database (CC-BY 4.0).
---
Release Notes
[https://github.com/jhy/jsoup/blob/HEAD/CHANGES.md#1231-2026-Jul-30 v1.23.1 ]
##### Improvements
- Reduced retained memory when parsing with source position tracking enabled ( Parser#setTrackPosition(true) ). Source ranges are now stored in compact parser-owned span records instead of node and attribute user data, and Position objects are created lazily when source ranges are read. This cuts tracked DOM retained size by about 50-60% on representative benchmark documents, while keeping Node#sourceRange() , Element#endSourceRange() , and Attribute#sourceRange() behavior intact. #2498
- Added Element#classList() , an immutable snapshot of an element's class names in attribute order. Use hasClass() when you just need to test for one class, classList() when you want to read or iterate classes without needing a mutable result, and classNames() when you want the existing mutable, deduplicated set that can be written back with classNames(Set) . The class APIs now share an HTML-whitespace scanner, which also makes classNames() faster and lighter on allocation, especially when walking many elements without class names. #2500
- Aligned HTML parser scope classification with the current HTML spec for select , foreignObject , and template . #2501
- Simplified the HTML tree builder's scope, implied-end-tag, and special-element checks by caching parser-only options on Tag. That improves HTML parser throughput by about 10% on small inputs and up to about 30% on larger inputs in the benchmark fixtures. #2502
- Improved HTML parser throughput stability by making hot tokeniser scan paths compile more predictably. #2507
- <noscript> fallback markup is now parsed into an inspectable DOM subtree in both the document head and body. The fallback acts as a contained parsing island, so malformed markup cannot disrupt the surrounding document structure, while normal HTML tokenization still applies within it. This also improves round-trip serialization. #2537
- Improved redirect credential handling as a defense-in-depth measure: explicit authorization headers and request cookies are no longer forwarded across origins, reducing exposure through open redirects and aligning with HTTP guidance. Cookies managed by a CookieStore continue to follow their configured scope. #2540
- Elements can now append their outer HTML, including their own tags, directly to an Appendable with Node#outerHtml(Appendable) , without first creating a String . This complements Element#html(Appendable) , which appends inner HTML only. #2532
- Aligned CDATA tokenization with the HTML spec: CDATA syntax in HTML content is parsed as a bogus comment, while it remains supported in SVG, MathML, and XML. Also improved namespace-aware fragment parsing so SVG and MathML contexts, HTML integration points, and context-sensitive tokenizer states are handled correctly. #2542
- When using the optional re2j regular expression engine, stack overflows caused by complex selector patterns are now normalized to a ValidationException with a Pattern complexity error message. #2548
##### Bug Fixes
- Fixed HTML parsing of mixed-case RCDATA end tags after tag-shaped text. For example, <title><p>Foo</TiTLE> and <textarea><img src=x></TeXtArEa> now keep the tag-shaped content as text instead of promoting it to markup. #2503
- Fixed W3CDom XML conversion so plain XML elements don't serialize with the reserved XML namespace as the default namespace. Explicit XML namespaces and xml:* attributes are still preserved. #2504
- Preserve control characters in parsed tag names #2538
- Updated HTTP redirects to follow the specification: 307 and 308 preserve the request method and content, 301 and 302 only change POST to GET, and Location is followed only for 301, 302, 303, 307, and 308 responses. Streamed request bodies are not buffered; if an automatic redirect requires replaying one, execution fails, so the caller can resend with a fresh stream. #2540
- Corrected the Cleaner's same-site link detection to compare hostnames rather than URL prefixes when applying rel=nofollow . #2543
##### Build Changes
- Cleaned up the Maven build for the multi-release JAR so Java 8 and Java 11+ sources compile as separate source sets. This avoids spurious Java 8 compiler warnings from newer-language overlay sources, keeps long-running parser checks behind an explicit profile, and preserves the same published artifacts and runtime behavior.
- Improved parallelism and tuned timing in our integration tests, so that a full mvn clean verify drops from \~ 1m18s to \~ 21 seconds.
---