← board

xml.etree.ElementTree: a thin tree model, or a real XML library?

Why this is the question now

After the 2026-08-18 shim batch, the corpus has no thin-stdlib-shim work left except this row. The remaining missing-module rows on the ladder are:

row files what it actually is
xml_etree_elementtree 4 this question
genshi_core 2 a third-party package, not stdlib
lxml 1 a third-party package (a C library binding)
weakref 1 a runtime lifetime facility, not a shim

So this row is the entire remaining shim lever, and it is blocked on a naming/scope decision rather than on effort.

What the corpus ACTUALLY uses — measured, not assumed

Every ElementTree.<name> reference in html5lib's library code:

5  ElementTree.Element
3  ElementTree.Comment
1  ElementTree.ElementTree

And every member touched on the elements it builds:

.tag  .text  .tail  .attrib  .get  .set  .append  .insert  .remove  .find

plus len(elem), elem[i], and iteration.

What it does NOT use: parse, fromstring, iterparse, SubElement, findall, itertext, namespaces beyond plain string tags, or any serialisation — html5lib/treebuilders/etree.py:262 defines its own tostring. There is no XML text on either side of this interface: html5lib parses HTML itself and only wants somewhere to hang the resulting tree.

One quirk that must be reproduced exactly: ElementTree.Comment("x").tag is the Comment function itself (CPython uses the factory as a sentinel tag), and html5lib relies on that identity — ElementTreeCommentType = ElementTree.Comment("asd").tag, then node.tag == ElementTreeCommentType.

The fork

It is not "how much work". A tree model with those members is roughly 60 lines and is exactly the kind of closed, testable thing the other mimic_* shims are. The question is what we are willing to put an upstream name on:

Option A — thin tree model (recommended). Ship mimic_xml_etree_elementtree.py with Element, Comment, ElementTree and the ten members above. parse, fromstring and iterparse are absent, so a caller reaching for them gets a loud unresolved-name error naming this decision.

Option B — a real XML implementation. Tokeniser, well-formedness, entity and namespace handling, serialisation, and enough XPath for find/findall.

Option C — ship A under a non-upstream name (e.g. pxxtree) and leave the upstream name unclaimed.

Recommendation

A, with the absences explicit and loud, and a docstring that says the module is a tree model rather than an XML implementation. If B is ever wanted it is a separate, honestly-ranked project — the same call already made for minidom in [[feature-b-a-real-minidom-is-an-implementation-not-a-shim]], and A does not foreclose it.

If A is chosen, the work is ready to start immediately and is measured: ~60 lines plus a differential test against CPython's real xml.etree.ElementTree, in the shape of the eight mimic_* shims already gated by make lib-test.


DECIDED 2026-08-19 by the user — Option A: the minimal shim

"For now the minimal shim; if we want to extend it we can write the XML importer later."

Ship mimic_xml_etree_elementtree.py as the tree model only: Element, Comment, ElementTree, and the ten members the corpus actually touches. The XML reader is a separate, later, optional piece of work — chosen on its own merits if it is ever wanted, not smuggled in as a side quest attached to a corpus goal.

The sub-question was NOT decided, and does not need to be

How should a shim declare its own incompleteness — omit the missing entry point (compile error), or include it and refuse (runtime error)? The repo has one of each today: mimic_urllib_request includes urlopen and raises; this ticket proposed omitting parse. Asked, and the user did not settle it — so do not write a general rule.

Turn it into a measurement instead of a decision. The distinction that matters is whether anything imports the missing entry point without calling it — that is the only reason present-and-refusing exists (mimic_urllib_request is a refusing stub precisely so importing code still compiles). So:

If a second case ever needs a judgement call rather than a measurement, then file the general rule as its own decision.

Must be exact, or comments silently stop working

ElementTree.Comment("x").tag is the Comment function itself — CPython uses the factory as its own sentinel tag, and html5lib depends on the identity: ElementTreeCommentType = ElementTree.Comment("asd").tag, then node.tag == ElementTreeCommentType.

Re-filed as work

See feature-b-mimic-xml-etree-elementtree-tree-model. A decided ticket that is never re-filed is invisible to ready/next.