Repo Listing Auditor
Paste a list of repository names and see how the collection reads at listing depth: which first hyphen tokens are concentrated, and which sorted neighbours share so many leading characters that they render as one unreadable block. A slug echo check runs alongside and currently finds nothing on the author's own 373 repos.
Everything runs in this page. Nothing is uploaded, stored or transmitted. This is a collection level verdict, not a per repo linter, and it never fetches a repository.
Input
Four accepted shapes: a JSON array of objects, tab separated name and description, comma separated name and description with double quoted fields, or a bare newline separated list of names. A bare list of names with no descriptions at all is enough.
The 373 row example is the real reference corpus: the author's own public repository list, name, tab, description, as it stood in the snapshot of 2026-09-02, shipped in full and never trimmed, so every figure in this page's README is reproducible here in one click. It is a snapshot, not a feed, so a list pulled today is a different set of rows. The corpus is stored in this file as literal text with no substitution of any kind, and the self test checks it row by row. The 40 row detector sample is neutral placeholder data that exercises every panel, which the reference corpus does not.
Panel A. Prefix concentration
Will count the first hyphen token of every name and list every token appearing 3 or more times, most frequent first.
Panel B. Adjacent collisions
Will sort a copy of the list, find every maximal contiguous run of 3 or more entries sharing a first hyphen token, and every adjacent pair sharing 8 or more leading characters.
Panel C. Mock listing
Will render the sorted names as a plain listing so the collision blocks are visible rather than merely counted.
Panel D. Description checks
Secondary. Never the headline.
Will check whether a description merely repeats its own name, and will report a separate WARN tier for one string match rule.
Self test
Runs the shipped functions against known correct values, including two positive controls: a deliberately wrong assertion that must be reported as correctly rejected, and a malformed input that must raise a named refusal.
Method, and what each number means
Sorting
The input is never reordered in place. A copy is sorted. Codepoint order is a plain a < b comparison, not localeCompare, because the default locale comparison is punctuation insensitive on many engines and would silently change the default output. The alternative setting uses the browser's own collator with base sensitivity and punctuation ignored, plus a codepoint tiebreak so the sort stays total and stable.
Which order any given listing page uses is a property of that page, not of this tool. No document on hand states it, so the choice is exposed rather than assumed. The author's own catalog notes recorded their prefix runs under the punctuation insensitive order while calling that order lexicographic, which is where two conflicting sets of run lengths in that file came from.
Prefix histogram
Everything before the first hyphen, lowercased. No stemming, no lemmatising, no merging of singular and plural. embedding and embeddings are two different tokens, and merging them would destroy the finding this check exists to surface.
Maximal contiguous runs
A run is a block of consecutive entries in the sorted copy whose first hyphen token is identical. It is defined by adjacency, not by how many times the token appears anywhere in the list. That distinction is the point: a token can appear nine times and still not form a run of nine if another name sorts into the middle of the block. The sorted index of the entry immediately after each run is printed, so you can see for yourself what does or does not split a block.
Longest common prefix
Matching leading characters between two adjacent sorted names, counted until the first difference or the end of the shorter name. Computed for every adjacent pair, which is one comparison fewer than the number of names. The threshold is yours to change.
Panel D, and the WARN tier
Slug echo is exact equality of the sorted token multisets of name and description, plus a separate exact title cased slug string test. Token set Jaccard, intersection over union, is reported at 0.45 or above as a NEAR tier. An empty description is counted separately and is never an echo.
WARN Copula opener rule. This is a string match, not a grammatical judgement. Clause parsing cannot ship in one self contained page with no network and no library, so it is not attempted. Measured on the author's 373 row reference corpus, this rule scores 1 hit, 0 true positives, 1 false positive (conventional-commits-reference). It is never counted in any headline total.
A broader variant was written and rejected. Looking for any form of "to be" anywhere before the first subordinating conjunction fires on 40 of the same 373 rows with visible false positives, including eagles-eye ("A spy satellite simulator in your browser, except the data is real.") and text-encryptor ("Encrypt and decrypt text in your browser with a passphrase using the Web Crypto API."), neither of which opens with a copula. The narrow rule ships, labelled WARN, with its own failure rate printed above.
What is deliberately not here
No near duplicate name detector by edit distance. It was implemented and measured: at distance 3 or less with a first token bucket prefilter it returns zero pairs on the 373 row corpus. It is dead weight, so it is not shipped in any form.