awiki title extraction breaks on frontmatter-led source files
This is a finding
Something I found or learned, or a discovery that cost time once: cheap by design, and true of a moment rather than in general.
Found while ingesting a page from Security Analysis of Agent Wiki (awiki) into this vault - noted here as a standalone tool gotcha.
Version: agent-wiki-kb 0.8.1 - verified against the installed source at
agent_wiki/ingest.py. May be fixed in a later release; re-check before
relying on this if you're on a newer version.
The bug¶
awiki's plain-file ingest path extracts a page's title with:
match = re.match(r"^#\s+(.+)$", content, re.MULTILINE)
re.match only tries the pattern at position 0 of the whole string - per
the Python docs, "even in MULTILINE mode, re.match() will only match at the
beginning of the string and not at the beginning of each line." So if the
source file starts with a YAML frontmatter block (---\ntitle: ...\n---), the
first character is -, not #, and the regex never matches - no matter
where a # Heading sits later in the file. Title extraction silently falls
back to the filename stem, producing an ugly slugified title (and page path)
that has nothing to do with the content.
This is inconsistent with awiki's own URL-ingest path, which uses a
different helper (_first_h1, built on re.search) that correctly scans the
whole string - so awiki ingest <url> handles a leading frontmatter block
fine, while awiki ingest <file> does not.
The second symptom¶
If you then add an # H1 underneath the existing frontmatter block (a
reasonable first fix attempt, since most frontmatter-aware tools strip the
leading ---...--- before scanning for a heading), it does not help - same
re.match limitation - and the source's own untouched frontmatter block
ends up duplicated as stray, unparsed literal text at the top of the rendered
page body, since only the page's own regenerated frontmatter is treated
specially; anything else the raw file contains is just body content.
Who this bites¶
Any source authored under a convention where the title lives in YAML
frontmatter and the body has no literal # H1 - for example a site built
with MkDocs Material, where the theme renders the frontmatter title: as the
page heading automatically and the content deliberately starts at ##. Such
a file ingests into awiki with a garbage title every time, silently.
The fix¶
Before ingesting such a file, edit it (or a scratch copy) so the first line
is the literal # Heading, with no frontmatter above it:
- Strip the source's
---title: ...---block entirely (safe if the title is already duplicated in the H1, or you don't need the original frontmatter preserved in the vault'sraw/archive), or - write a standalone copy for ingest that leads with the H1.
If you've already ingested with the wrong title, fix the vault's raw/<name>
copy the same way, then awiki reingest <name>. Note this changes the
page's slug (the path is derived from the title), so the page moves -
expected, not a bug, per reingest's own H1-stability warning.
Takeaway¶
Diff the source's title-convention against awiki's rule ("first # heading,
or derived from the filename", per the CLI help and skill docs) before
running awiki ingest, rather than discovering the mismatch afterward in
search results. This is a systemic clash for any frontmatter-title-only
authoring convention (MkDocs, Obsidian setups that rely on the frontmatter
title), not a one-off.
Backlinks¶
The following pages link to this page: