← board

Every compile-time intrinsic hand-rolls its own operand parser

RESOLVED 2026-08-26. All four rows landed; see Outcome below.

Umbrella, opened 2026-08-26. Four tickets filed between 2026-08-20 and 2026-08-22 against SizeOf, Low and High. Each was filed as a missing case; together they are one design.

The design, stated once

SizeOf, Low and High do not share an operand parser or a result-typing rule. Each is a private re-implementation of the selector chain, with its own set of accepted shapes and its own size/bound formula:

Why this is one ticket and not four

The SizeOf(p^.A) ticket already reached this conclusion on its own and recorded it: it was filed at low prio specifically because "the cheap version means adding a fourth shape to SizeOf's hand-rolled operand walk, which is the structure that produced the parent bug in the first place." Its parent (bug-p-sizeof-an-array-field-returns-the-element-size) was that structure failing once already.

So the ordering is: one operand parser and one result-typing rule for the whole intrinsic family first, then the four rows fall out. Doing them in the filed order adds a fifth, sixth and seventh hand-rolled branch and makes the eventual consolidation harder. devdocs/dev/normalise-dont-special-case.md is the reference; this is a textbook instance of it.

What NOT to normalise away

Two rules are genuinely per-intrinsic and must survive the consolidation:

  1. a literal is typed by its value for SizeOf (the table in the folded ticket below is the spec, measured against fpc 3.2.2);
  2. Low/High over a string CONSTANT are 0-based over its own length, over a managed-string EXPRESSION are 1-based, and over a shortstring expression are 0-based over the DEFAULT capacity — three different bases, so "it starts with a quote" tells you nothing. That table is in the folded ticket too.

A consolidation that flattens either of those is a regression that compiles.

Gate

make compiler/pascal26 + every row of both tables diffed against fpc 3.2.2 + tools/gate.sh quick.

Outcome

All four rows landed, each gated against fpc 3.2.2 and each leaving its measured table pinned as a test rather than as prose.

What the last row cost that the ticket did not predict

Two things, both found by measurement rather than by reading:

  1. The Low arm clobbered its own fix. Its chain ended in a blanket ASTTk[CurASTNode] := Ord(tyInteger) AFTER the arms, so High(c) answered 'e' while Low(c) still answered 97 from what looked like symmetric code. The default now goes in before the chain, where an arm can override it.
  2. Low(TE) over an enum TYPE was wrong too, and was not in any of the four tickets: WriteLn(Low(TE)) printed 0 where fpc prints eA. The ordinal was right and the node had simply forgotten which enum it came from — one line, the same missing fact, found only because the test happened to include it.

What it deliberately did NOT fix

const CLO = Low(TC); WriteLn(CLO) still prints 97. That is not this ticket: pxx's constant evaluator represents every ordinal as a bare Int64 by design, so const X = eB already prints 1 rather than eB with no Low/High involved. TryConstHighLowValue is handed the index type and discards it on purpose, with a comment saying so. Filed as [[bug-p-the-constant-evaluator-erases-an-ordinals-type]], which is where the held-out rows of test/test_low_high_index_type.pas go when it is fixed.

What is left of the design complaint

The consolidation this ticket actually asked for — ONE operand parser for the whole family — is not done, and is smaller now than it was. Low/High share HighLowOperandIsExpr for shape and StampArrIdxType for result typing; SizeOf still has its own walk. Nobody should reopen this to finish that: the four symptoms are gone and there is a pinned table behind each, so the next person to touch the family has a spec to move against instead of a description of one. If a fifth symptom appears, THAT is the moment the shared parser pays.


The folded tickets, verbatim

Each section below is a ticket that was filed separately and is now part of this one. Nothing is summarised away: the repro tables, the measured oracle output and the located source lines are the reason these are worth keeping, and they are reproduced unchanged.

SizeOf rejects a pointer deref in its operand

(was bug-p-sizeof-rejects-a-pointer-deref-in-its-operand, prio 55)

SizeOf rejects a pointer deref in its operand

Repro

type TR = record A: array[0..2] of Integer; end;
     PR = ^TR;
var r: TR; p: PR;
begin
  p := @r;
  WriteLn(SizeOf(p^.A));    { pxx: Expected: ), but got: ^   FPC 3.2.2: 12 }
end.

SizeOf(r.A) and SizeOf(p^) are fine; it is the deref inside a selector chain that has no case.

Why it is filed separately, and low

It is a loud failure — a parse error at the exact token, not a wrong value — which is the opposite of its parent ticket's failure mode, and nothing silently miscomputes. It is filed at 35 rather than folded into that fix because the cheap version means adding a fourth shape to SizeOf's hand-rolled operand walk, which is the structure that produced the parent bug in the first place.

The shape worth considering first

SizeOf's operand parser is a private re-implementation of the selector chain: separate branches for a bare var, v.f.g, v[i] on a 1-D array, and v[i,j] on an N-D one, none of which know about ^, and each carrying its own size formula. The parent ticket already collapsed the field formula into RecFieldByteSize. The real fix is probably to stop hand-rolling: parse the operand with the ORDINARY lvalue parser (which handles ^, indexing, chains and casts already) in a non-evaluating mode, and ask the resulting node for its type + extent. That is a bigger change than this symptom justifies on its own, which is why this sits in backlog rather than being squeezed in — but it is the version that would also retire the wrong number of array subscripts and SizeOf: unknown field special cases. See devdocs/dev/root-cause-over-microfix.md.

Low/High of a char-indexed array answer the ordinal, not the char

(was bug-a-low-high-of-a-char-indexed-array-answer-the-ordinal, prio 48)

Low/High of a char-indexed array answer the ordinal, not the char

Found 2026-08-22 alongside [[bug-a-low-high-of-an-ordinal-variable-answer-0-and-minus-1]] and split out of it, because it is a different missing fact rather than a missing arm.

The measurement

fpc -Mobjfpc -O1 3.2.2 vs pxx 0b77e2bea.

type TC = array['a'..'e'] of Integer;
var c: TC;
expression fpc pxx
WriteLn(Low(c)) a 97
WriteLn(High(c)) e 101
WriteLn(Low(TC)) a 97
WriteLn(High(TC)) e 101

The VALUES are right; only the type is. The consequence is that the natural loop does not compile:

var ch: Char;
for ch := Low(c) to High(c) do ...    { pxx: bound is Integer, not Char }

Note this is NOT the case of a named char SUBRANGE, which is already correct: TL = 'a'..'e' gives 'a'/'e' for both the type name and a variable, because SymIsSub/AliasIsSub carry the base type along with the bounds.

Root cause

An array's INDEX type is not recorded. ArrTypeDimLo/ArrTypeDimSpan (and SymArrDimLo/SymArrDimSpan for a variable) store the bounds as plain Integers, so by the time Low/High folds there is nothing left saying the index was a Char, a Boolean or an enum. Both fold sites therefore stamp Ord(tyInteger) on the literal.

Boolean- and enum-indexed arrays have the same shape and are worth checking in the same pass (array[Boolean] of T, array[TEnum] of T) — the ordinal values 0/1 and 0..n happen to be indistinguishable from the right answer when printed through Ord, so they may be silently wrong in the same way.

The fix

Record the index type kind next to the bounds — an ArrTypeIdxTk parallel to ArrTypeDimLo, and the matching SymArrIdxTk — then have both fold sites use it instead of tyInteger. TryArrayTypeBound already returns through a var parameter and the variable arms already build the literal by hand, so both take a type the same way TryOrdinalVarBound already does.

Fix both arms together. The type-name arm currently answers 97 on purpose, to agree with the variable arm; changing one without the other would make Low(TC) and Low(c) disagree, which is worse than the present state.

Gate

The four rows above matching fpc -O1, plus for ch := Low(c) to High(c) compiling with a Char loop variable, plus the Boolean- and enum-indexed cases, and self-host byte-identical.

High/Low refuse every non-identifier operand

(was bug-p-high-and-low-refuse-every-non-identifier-operand, prio 30)

Refused today, legal in fpc 3.2.2

var s: AnsiString;
begin
  s := 'qxy';
  WriteLn(High('abc'));      { fpc: 2   pxx: High: expected array variable or ordinal type }
  WriteLn(High('ab' + s));   { fpc: 5   pxx: same error }
  WriteLn(High(('ab')));     { fpc: 1   pxx: same error }
  WriteLn(Low('abc'));       { fpc: 0   pxx: Low: expected array variable or ordinal type }
  WriteLn(Low('ab' + s));    { fpc: 1   pxx: same error }
end.

compiler/pasparser_expr.inc, both arms:

if CurTok.Kind <> tkIdent then Error('High: expected array variable or ordinal type');

A proc NAME already escapes this (it starts with an identifier — fixed by compat-pascal-index-a-function-call-result), so High(F) works. A literal or a ( does not.

Why this is not a one-line widening — THREE bases

Measured, and this is the whole reason the ticket exists:

operand fpc Low fpc High why
'abc' 0 2 a string CONSTANT is an array-of-CHAR constant, 0-based over its own length
('ab') 0 1 …and the parens do NOT change that — so the test must ask the NODE, not the token
'ab' + s 1 5 a managed-string EXPRESSION, 1-based over its length
'ab' + 'cd' 1 4 two literals concatenated are an ANSISTRING expression, not a char array
sh + 'x' (sh: string[10]) 0 255 a shortstring EXPRESSION, 0-based over the DEFAULT capacity

So "it starts with a quote" tells you nothing; the answer depends on the value's type and on whether it is a constant. The bases themselves already exist in the compiler after [[bug-p-high-and-low-of-a-string-are-off-by-one]] — managed = 1 .. Length, frozen = 0 .. capacity, array = own bounds. What this ticket adds is (a) letting the operand parse, and (b) an ASTKind[valNode] = AN_STR_LIT arm for the constant case, which folds to ASTSLen[valNode] - 1 / 0 and must sit ABOVE the frozen arm (a literal's LastExprTk is tyString, so the frozen arm would otherwise claim it and answer 255).

Do not widen the guard to "anything": High(3) would then reach the runtime Length tail and produce garbage where it is currently a clear error. Admit tkString and tkLParen and keep the diagnostic for the rest.

Sibling gap in the same arm — a frozen-string RESULT has no capacity

High(G) where G: TSA returns string[6] answers 1 (Length - 1); fpc says 6. There is no ProcRetStrCap row for the frozen arm to read, and defaulting to 255 would turn a too-SMALL bound into a too-LARGE one — a loop reading past the end rather than stopping early — so the arm currently declines the operand rather than guessing. Adding the capacity to the proc row is a small Track A metadata change and would close this row; do it in the same pass if the row is easy to add, and keep the decline if it is not.

Gate

Track P's. Every row of both tables above in a test wired into test-core, each diffed against fpc 3.2.2 — including the rows that already work, since the fix moves them onto a different path. Grep Length's arm before closing: it took the same class of fix on 2026-08-25 ([[bug-p-length-of-a-string-literal-plus-anything-does-not-parse]]) and its literal fold is the model for the AN_STR_LIT arm here.

SizeOf(<literal>) is refused

(was feature-p-sizeof-of-a-literal, prio 20)

SizeOf(<literal>) is refused

Split out of [[feature-p-sizeof-of-an-expression]] when that landed on 2026-08-22. Expression operands now work; literal operands still raise SizeOf: unknown type or variable.

They were left out deliberately: fpc types a literal by its VALUE, not by the type the expression parser would give it, so routing them through the new expression path would answer 8 for most of this table — a wrong size, silently, that GetMem and Move would carry straight into the allocator.

The rule to implement (measured, fpc -Mobjfpc -O1 3.2.2)

operand fpc why
1 1 smallest type holding the value
127 1
128 1 unsigned range, so Byte
255 1
256 2
32767 2
32768 2 still Word
65535 2
65536 4
100000 4
5000000000 8
-1 1 ShortInt
-129 2 SmallInt
3.5 4 a real literal is Single, not Double
'a' 1 Char
'' 1
'abc' 3 its LENGTH, not a string handle
nil 8 pointer width
[1, 2] 2 the set's storage size

Two of these are traps worth calling out: SizeOf(3.5) is 4 because an untyped real constant is Single-typed for this purpose even though it would promote to Double in arithmetic; and SizeOf('abc') is 3 because a string literal is typed as its own array[1..3] of Char. Both look harmless to get wrong and are wrong everywhere, so both belong in the test.

[1, 2] also depends on compat-pascal-set-storage-size-is-always-32-bytes — our sets are 32 bytes, so that row cannot match until that ticket does. Skip it or assert the pxx value with a comment; do not "fix" set sizing from here.

Where the code is

compiler/pasparser_expr.inc, the szIsExpr dispatch block at the top of the sizeof intrinsic. Today the first token being a literal falls through to the name path and its error. Add a literal arm ahead of that, keyed on the token kind, implementing the table above.

Prio

  1. Nobody writes SizeOf(1) in earnest — the value of the parent ticket was type-probing an expression, and that half has landed. This is conformance tidiness, and the wrong answer is currently a loud compile error rather than a silent number, which is the right failure mode to wait in.

Gate

Every row above matching fpc -O1 (except the set row, see above), the expression and name paths unchanged, and self-host byte-identical.