RedlineKnowledge base

B1 Build-time index

Make page history a constant-cost build step. This is the on-ramp to the whole experiment: the consuming site has history switched off because of this cost.


T1.1 Benchmark harness

  • Wave: 1
  • Depends on: none
  • Size: M
  • Design: build-index.md (observability)
  • Touches: scripts/bench-history.js (new), git.js (stats only), test/bench.test.js (new), package.json (script)
  • Why: the gate is a bound on git invocations. The harness measures it the same way before and after T1.2.

Steps

  1. git.js: add a module-level _stats = { gitCalls: 0, indexSource: 'none', indexPaths: 0, indexMs: 0 }. Increment gitCalls inside the git() helper. Export _getGitStats() returning a copy and _resetGitState() that resets stats and forces the next call to re-run init (set _ready = false). Do not change behaviour.
  2. scripts/bench-history.js with flags --pages N (default 300), --commits C (default 1200), --seed S (default 1), --no-index (passes index: false, meaningful after T1.2), --keep (do not delete the temp repo; print its path).
    • Create a temp git repository with docs/page-0001.md … docs/page-N.md.
    • Make C commits; each touches between 1 and 5 pages chosen by a seeded PRNG (implement a 32-bit LCG; do not add a dependency). Every 200th commit renames one page with git mv. Commit with -c commit.gpgsign=false, fixed author, deterministic dates (GIT_AUTHOR_DATE/GIT_COMMITTER_DATE stepping one hour per commit from 2026-01-01T00:00:00Z).
    • Call _resetGitState(), then getGitMeta(path, { cwd, cacheFile, silent: true, contentRoots: ['docs'] }) for every current page path, timing the loop.
    • Print exactly one JSON line: { "pages", "commits", "gitCalls", "ms", "indexSource", "indexed": indexSource !== 'none', "cacheFile" }.
    • --warm runs the loop a second time after _resetGitState() with the same cacheFile and prints the second run’s JSON.
  3. package.json script "bench:history": "node scripts/bench-history.js".

Acceptance

test/bench.test.js:

  • benchmark harness produces a deterministic repository (run twice with --pages 5 --commits 12 --seed 7 --keep, assert the two git log --format=%H outputs are equal)
  • benchmark harness prints one json line with the expected keys (--pages 20 --commits 30, parse stdout, assert keys and pages === 20)
npm run bench:history -- --pages 50 --commits 100   # one JSON line; gitCalls will be > 150 before T1.2

T1.2 Build-time history index

  • Wave: 2
  • Depends on: T1.1
  • Size: L
  • Design: build-index.md
  • Touches: git.js, index.js (pass index and contentRoots through; nothing else), index.d.ts (options and localCommits), test/git-index.test.js (new), README.md (options table under “Consumer build requirements”)
  • Why: the core of B1.

Steps

  1. Keep the existing per-path implementation; rename the body of getGitMeta to getGitMetaByWalk(repoPath, candidates, opts) and leave it intact.
  2. In init(opts): after the existing rev-parse calls, when opts.index !== false, build the index as the design specifies: one git log, one git status --porcelain, one git ls-files, with contentRoots (default ['.']) as pathspecs. Store _index = { rows: Map<path, Row[]>, tracked: Set, dirty: Map<path, status>, roots }. Record indexSource = 'walk', indexPaths, indexMs, and print the summary line unless silent.
  3. New getGitMeta: resolve candidates as today; if the index exists and the first candidate that is in tracked or rows is under contentRoots, derive the result from rows as the design describes and attach localCommits. Otherwise call getGitMetaByWalk.
  4. index.js: serializableIntegrationOptions already serialises arbitrary options; nothing to change except the JSDoc. Update index.d.ts.

Acceptance

test/git-index.test.js (build temp repositories as test/git.test.js does):

  • index returns identical metadata to the per-path walk (6 files, 8 commits including one git mv followed by an edit; for every path compare created, updated, revisions, contributors, authors, signed, state between { index: true } and { index: false })
  • index git call count is independent of page count (repo A with 3 files, repo B with 60 files; query every file; gitCalls equal for A and B and <= 8)
  • index attaches local commits newest first (localCommits[0].sha equals git rev-parse HEAD for a file touched by the last commit)
  • paths outside content roots fall back to the per-path walk (contentRoots: ['docs'], ask for README.md; result matches the walk; indexSource stays walk)
  • dirty and untracked states survive indexing (modify a tracked file → dirty: true; add a new file → state: 'untracked'; git add another → state: 'staged')
npm test 2>&1 | grep -c "^✔ index \|^✔ paths outside\|^✔ dirty and untracked"   # prints 5
npm run bench:history -- --pages 300 --commits 1200   # gitCalls <= 8, indexed true
npm run bench:history -- --pages 3 --commits 1200     # gitCalls equal to the line above

T1.3 Warm index cache

  • Wave: 3
  • Depends on: T1.2
  • Size: S
  • Design: build-index.md (cache)
  • Touches: git.js, test/git-index.test.js
  • Why: rebuilds on the same commit should not walk history at all.

Steps

  1. On a walk, store __index: { head, roots, rows } in the cache object and mark it dirty so flush() writes it.
  2. On init, if the cache has __index with head === git rev-parse HEAD and equal roots, load rows from it and set indexSource = 'cache'; skip the git log.
  3. Rows must round-trip through JSON (plain objects, no Map in the file).

Acceptance

Add to test/git-index.test.js:

  • warm index skips the history walk (second _resetGitState() + query with the same cacheFile: indexSource === 'cache', gitCalls <= 7, results deep-equal the cold run)
  • index cache is invalidated by a new commit (commit again; indexSource === 'walk')
npm run bench:history -- --pages 300 --commits 1200 --warm   # second line: indexSource "cache", gitCalls <= 7

T1.4 GitLab enrichment from the index

  • Wave: 3
  • Depends on: T1.2
  • Size: L
  • Design: build-index.md (localCommits)
  • Touches: gitlab.js, index.d.ts, test/gitlab-index.test.js (new), README.md (one paragraph under “Limits, caching, and retries”)
  • Why: after T1.2 the only per-page remote call left is repository/commits?path=…. With local commits known, the remote work becomes one signature lookup per unique commit and one user lookup per unique email, both already cached.

Steps

  1. In getGitMetaEnhanced, when local.localCommits exists, local.state === 'committed' and !local.historyIncomplete: do not call gitlabPaged(... repository/commits ...). Build rawCommits from localCommits (first maxCommits) as objects with the fields mapCommit reads: id, short_id (first 8), title, message (title), author_name, author_email, authored_date, committed_date (both the row date), web_url (${buildProjectUrl(client)}/-/commit/${sha} when projectPath is known, else undefined), stats: null. complete = true.
  2. Everything downstream (mapCommit, signatures, user resolution, tags, in-flight) stays as it is.
  3. Signature responses are immutable; give them a cache TTL of 30 days in the response cache regardless of cacheTtlMs. Keep the user lookup TTL.
  4. trackingSource for this path is 'mixed'.

Acceptance

test/gitlab-index.test.js (temp repo with 3 pages sharing 2 commits by 1 author; counting fake fetch built from test/git.test.js’s createGitLabFetch pattern):

  • enriched history uses local commits and makes no per-path commits request (zero requests whose path ends in /repository/commits and carry a path query)
  • enriched history requests one signature per unique commit and one user per unique email (signature requests === 2, users?search requests === 1 across the three pages)
  • enriched history links commits to the project when the project path is known (meta.gitlab.commits[0].url === 'https://gitlab.example.test/group/project/-/commit/<sha>')
  • enriched history falls back to the commits request for a shallow clone (historyIncomplete: true input → one /repository/commits request with path)
npm test 2>&1 | grep -c "^✔ enriched history"   # prints 4
Git history

Loading the page's history…