What Makes for Successful Contributions to Perseus canonical-greekLit

A tiered analysis of the commit record, 2013–2026

Author

LLM-generated, prompted and supervised by Mark G. Bilby

Published

July 29, 2026

Companion document. This report is the evidence base for the checklist in the PR template (“A Template for Successful Pull Requests”), kept in the repository as perseus-PR-template.md. Each success factor in §6 maps to a section of that guide.

1 Executive summary

The PerseusDL/canonical-greekLit history contains roughly 3,900 non-merge commits and 1,155 merged pull requests, spanning mid-2013 to mid-2026. Contributions divide into four tiers of contributor and three modes of work. Only the fourth tier is unpaid; the first three are compensated — long-term staff, contract engineers, and funded interns.

Tiers, by who the contributor is:

  • Tufts PDL long-term staff (Cerrato, Babeu, Crane, Almas, Buckingham): most of the volume. Expected, and not the subject of this report.
  • Leipzig / Open Greek and Latin engineers on fixed-term contracts (Clérice, Munson, Dee, Stoyanova): funded engineers, not volunteers, responsible for most structural change in short, concentrated windows.
  • CHS / Perseids interns and fellows (Konieczny, Hanhardt, shared intern accounts): funded, and the principal source of new editions.
  • Outside volunteers and independent scholars (Fifield, McCallum, Dik, Beine, Drymon, Driscoll, Berra, and ~20 others): the only unpaid tier, and almost exclusively correctors of existing editions.

Modes, by what a contribution does:

  1. Micro-corrections — single-line philological, OCR, or diacritic fixes.
  2. Edition-level work — converting or authoring a whole EpiDoc/CTS edition.
  3. Batch normalization — one systematic markup or metadata fix across files.

Principal finding. Substantial, durable contribution does not require staff status. The most prolific outside contributor, David Fifield, corrected ~59 editions over five years without creating a new file; one independent scholar, Helma Dik, earned merge rights; and the interns who converted the Galen corpus each created 65–70 editions in a single term. The shared method is set out in §6.

2 Data and method

Four TSVs were exported from a full clone (git log, default branch, merges excluded except where noted):

File Contents Used for
perseus-editlog.tsv A/M/D status per file, scoped to data/**/*.xml new-edition vs. correction counts
perseus-authors.tsv one line per commit: hash, name, email, date, subject totals, cadence, commit style
perseus-numstat.tsv added/deleted lines per file, scoped to data/**/*.xml magnitude of change
perseus-merges.tsv merge commits with PR number and source branch PR workflow, cadence

Caveats.

  1. The log records what landed; rejected or abandoned PRs are absent, so “what gets rejected” is inferred, not measured.
  2. A/M separates a new file from an edited one, but an edition-scale conversion of a pre-existing stub registers as M with a large +lines; line volume is therefore reported alongside A/M (see Berra, below).
  3. Identity is resolved to a person across aliases and emails, and each person is assigned a tier deterministically. The mapping is printed below so it can be audited.
Identity resolution, tier classification, and TSV parsers
import re
from collections import defaultdict, Counter

# --- alias -> canonical person ---
CANON = {
 'lcerrato':'Lisa Cerrato','Lisa Cerrato':'Lisa Cerrato','cerrato':'Lisa Cerrato',
 'Alison Babeu':'Alison Babeu','AlisonBabeu':'Alison Babeu',
 'gregorycrane':'Gregory Crane','Gregory Crane':'Gregory Crane','Crane':'Gregory Crane',
 'Bridget Almas':'Bridget Almas','balmas':'Bridget Almas','balmas01':'Bridget Almas',
 'TDBuck':'Tim Buckingham','Tim Buckingham':'Tim Buckingham',
 'Thibault Clérice':'Thibault Clérice','Thibault Clerice':'Thibault Clérice',
 'sonofmun':'Matthew Munson','Matthew Munson':'Matthew Munson',
 'Stella Dee':'Stella Dee','srdee':'Stella Dee',
 'simonastoyanova':'Simona Stoyanova','stoyanova':'Simona Stoyanova',
 'Michael Konieczny':'Michael Konieczny','Angelia Hanhardt':'Angelia Hanhardt',
 'helmadik':'Helma Dik','Helma Dik':'Helma Dik',
 'David Fifield':'David Fifield','Nathaniel McCallum':'Nathaniel McCallum',
 'Julia Jennifer Beine':'Julia Jennifer Beine','juliajbeine':'Julia Jennifer Beine',
 'Charles Pletcher':'Charles Pletcher','Jacob Wegner':'Jacob Wegner',
 'Isaac Bennett-Smith':'Isaac Bennett-Smith','Stephen Scott':'Stephen Scott',
 'Chris Drymon':'Chris Drymon','chrisdrymon':'Chris Drymon',
 'David F. Driscoll':'David F. Driscoll','David Smith':'David Smith (dasmiq)',
 'Jeroen Hellingman':'Jeroen Hellingman','Aurélien Berra':'Aurélien Berra',
 'ChiaraPalladino':'Chiara Palladino','Chiara Palladino':'Chiara Palladino',
 'Adiel Mittmann':'Adiel Mittmann','Joel Kalvesmaki':'Joel Kalvesmaki',
 'Scott Fleischman':'Scott Fleischman','Eric Sowell':'Eric Sowell',
 'emNeven Jovanovic/em':'Neven Jovanović','KATEBHN':'KATEBHN',
 'Kevin Krahn':'Kevin Krahn','kevinkrahn':'Kevin Krahn','TickleForce':'Kevin Krahn',
 'achitu':'Andrei Chitu',
 'Intern4':'CHS interns (shared)','Intern3':'CHS interns (shared)',
 'Intern2':'CHS interns (shared)','Intern 1':'CHS interns (shared)',
 'pubintern1':'Perseids pubinterns','pubintern2':'Perseids pubinterns',
 'pubintern5':'Perseids pubinterns',
}
def canon(name):
    return CANON.get(name, name)

# --- person -> tier (assigned by identity, NOT email domain) ---
# TUFTS: long-term paid staff.  OGL: funded fixed-term Leipzig / Open Greek
# & Latin engineers.  CHS: funded interns / fellows.  EXT: outside volunteers.
TUFTS = {'Lisa Cerrato','Alison Babeu','Gregory Crane','Bridget Almas',
         'Tim Buckingham','Rashmi Singhal','Anne Mahoney','William Merrill',
         'Elli Mylonas','David Mimno'}
OGL   = {'Thibault Clérice','Matthew Munson','Stella Dee','Simona Stoyanova',
         'Monica Berti','Greta Franzini','Gabriel Weaver'}
CHS   = {'Michael Konieczny','Angelia Hanhardt','CHS interns (shared)',
         'Perseids pubinterns'}

def tier(name, email):
    if ('github-actions' in name or name.lower().startswith('travis')
            or 'dependabot' in name):
        return 'BOT'
    c = canon(name)
    if c in TUFTS: return 'TUFTS'   # long-term paid staff
    if c in OGL:   return 'OGL'     # funded fixed-term engineers
    if c in CHS:   return 'CHS'     # funded interns / fellows
    return 'EXT'                    # outside volunteers

def is_edition(p):
    return (p.startswith('data/') and p.endswith('.xml')
            and not p.endswith('__cts__.xml'))

HASH = re.compile(r'^[0-9a-f]{40}$')

def parse_editlog(path='data/perseus-editlog.tsv'):
    d = defaultdict(lambda: dict(A=set(), M=set(), D=set()))
    cat = {}; cur = None
    for line in open(path, encoding='utf-8', errors='replace'):
        f = line.rstrip('\n').split('\t')
        if f and f[0] == 'C' and len(f) >= 6 and HASH.match(f[1]):
            cur = canon(f[2]); cat[cur] = tier(f[2], f[3])
        elif cur and len(f) >= 2 and f[0][:1] in 'AMD':
            if is_edition(f[-1]):
                d[cur][f[0][0]].add(f[-1])
    return d, cat

def parse_authors(path='data/perseus-authors.tsv'):
    tot = Counter(); cat = {}; dates = defaultdict(list)
    for line in open(path, encoding='utf-8', errors='replace'):
        f = line.rstrip('\n').split('\t')
        if len(f) < 5:
            continue
        c = canon(f[1]); t = tier(f[1], f[2])
        if t == 'BOT':
            continue
        tot[c] += 1; cat[c] = t; dates[c].append(f[3])
    return tot, cat, dates

def parse_numstat(path='data/perseus-numstat.tsv'):
    add = Counter(); dele = Counter(); cur = None
    for line in open(path, encoding='utf-8', errors='replace'):
        f = line.rstrip('\n').split('\t')
        if f and f[0] == 'C' and len(f) >= 6 and HASH.match(f[1]):
            cur = (canon(f[2]), tier(f[2], f[3]))
        elif cur and len(f) >= 3 and cur[1] != 'BOT' and is_edition(f[2]):
            add[cur[0]]  += int(f[0]) if f[0].isdigit() else 0
            dele[cur[0]] += int(f[1]) if f[1].isdigit() else 0
    return add, dele

ed, catE = parse_editlog()
tot, catA, dates = parse_authors()
add, dele = parse_numstat()
cat = {**catE, **catA}

def profile(c):
    e = ed.get(c, {'A': set(), 'M': set(), 'D': set()})
    ds = sorted(dates.get(c, []))
    span = f"{ds[0][:7]}..{ds[-1][:7]}" if ds else ""
    return dict(name=c, tier=cat.get(c, 'EXT'), commits=tot.get(c, 0),
                new=len(e['A']), corr=len(e['M']),
                ladd=add.get(c, 0), ldel=dele.get(c, 0),
                months=len({d[:7] for d in ds}), span=span)

rows = [profile(c) for c in set(list(ed) + list(tot))]
print(f"Contributors (excl. bots): {len(rows)}")
print(f"Window: {min(min(dates[c]) for c in dates)[:7]}"
      f" .. {max(max(dates[c]) for c in dates)[:7]}")
Contributors (excl. bots): 39
Window: 2013-07 .. 2026-07

3 The landscape: who contributes, and how much

Category rollup
roll = defaultdict(lambda: [0, 0, 0, 0])
for r in rows:
    roll[r['tier']][0] += 1
    roll[r['tier']][1] += r['commits']
    roll[r['tier']][2] += r['new']
    roll[r['tier']][3] += r['corr']
print(f"{'tier':6s}{'people':>7}{'commits':>9}{'new':>7}{'corrected':>11}")
for k in ['TUFTS', 'OGL', 'CHS', 'EXT']:
    v = roll[k]
    print(f"{k:6s}{v[0]:7d}{v[1]:9d}{v[2]:7d}{v[3]:11d}")
tier   people  commits    new  corrected
TUFTS       5     3183   1456       3763
OGL         4      119   1544       1348
CHS         4      357    210         44
EXT        26      253      3        251

Long-term Tufts staff account for most commits and most corrections, since maintaining the collection is their job and the managing editor also commits others’ reviewed work. The four OGL engineers, on ~119 commits between them, account for the largest share of new editions: one of them ran the CTS migration that created the file tree. Interns are the next-largest source of new editions. The outside (unpaid) tier is numerous (~26 people) but corrects rather than creates — three new editions among them, against ~250 corrected editions.

The rest of this report sets long-term staff aside and examines the three groups an entrant can join: funded engineers, interns, and outside volunteers.

4 Three modes of successful contribution

4.1 Mode 1 — Micro-corrections (outside volunteers)

Top correctors among outside volunteers (EXT), by editions modified
outside = [r for r in rows if r['tier'] == 'EXT']
print(f"{'contributor':22s}{'tier':4s}{'corr':>8}{'commits':>8}{'mo':>4}  span")
for r in sorted(outside, key=lambda r: -r['corr'])[:12]:
    print(f"{r['name'][:22]:22s}{r['tier']:4s}{r['corr']:8d}"
          f"{r['commits']:8d}{r['months']:4d}  {r['span']}")
contributor           tier    corr commits  mo  span
Nathaniel McCallum    EXT       84      26   2  2023-10..2023-11
David Fifield         EXT       59      58  16  2021-05..2026-07
Helma Dik             EXT       29      38   6  2021-09..2026-06
Chris Drymon          EXT       18       5   3  2021-01..2026-05
Julia Jennifer Beine  EXT       17      14   2  2025-05..2026-07
Scott Fleischman      EXT       14      30   2  2016-09..2016-11
Stephen Scott         EXT        5       6   5  2021-03..2021-07
Isaac Bennett-Smith   EXT        4       6   1  2025-06..2025-06
Chiara Palladino      EXT        2       4   2  2016-03..2016-04
Jeroen Hellingman     EXT        2       2   1  2020-10..2020-10
Aurélien Berra        EXT        2       9   3  2019-03..2019-07
Kevin Krahn           EXT        2       2   2  2023-02..2024-02

This is the most accessible path. David Fifield (a software engineer, not a classicist) is representative: ~59 editions corrected across 16 active months over 2021–2026, each a single-purpose commit naming the URN and the exact change — e.g. add missing first letter of ὡς in 14.22; change Latin o to Greek omicron; use <q> for nested quotations. Helma Dik (University of Chicago; Logeion / Perseus-under-PhiloLogic) submits OCR and diacritic corrections in many small PRs and is the only outside contributor with merge rights (§6). Julia Jennifer Beine and Chris Drymon follow the same single-fix pattern. A sustained stream of small, verifiable corrections is itself a first-class contribution.

4.2 Mode 2 — Edition-level work (interns and funded engineers)

Top creators of new editions (new file added), staff excluded
makers = [r for r in rows if r['tier'] in ('EXT', 'OGL', 'CHS') and r['new'] > 0]
print(f"{'contributor':22s}{'tier':4s}{'new':>6}{'+lines':>10}{'mo':>4}  span")
for r in sorted(makers, key=lambda r: -r['new'])[:10]:
    print(f"{r['name'][:22]:22s}{r['tier']:4s}{r['new']:6d}"
          f"{r['ladd']:10d}{r['months']:4d}  {r['span']}")
contributor           tier   new    +lines  mo  span
Thibault Clérice      OGL   1529   3996745  13  2014-10..2017-05
CHS interns (shared)  CHS     70    116808   2  2019-06..2019-07
Michael Konieczny     CHS     69    106186   7  2019-11..2020-09
Angelia Hanhardt      CHS     66     71434   9  2018-07..2020-04
Stella Dee            OGL     14    113037   9  2013-09..2015-08
Perseids pubinterns   CHS      5      2812   1  2018-07..2018-07
Chiara Palladino      EXT      2      1089   2  2016-03..2016-04
Charles Pletcher      EXT      1      4799   1  2023-07..2023-07
Simona Stoyanova      OGL      1      8786   2  2014-08..2015-07

New editions come overwhelmingly from interns and funded engineers. The CHS interns who converted the Galen corpus each produced 65–70 new EpiDoc/CTS editions in one term (Konieczny 69, Hanhardt 66, the shared intern account 70). Thibault Clérice’s large new/+lines figure is the 2014–15 CTS migration that created the canonical data/ tree — an infrastructural act, not 1,500 hand-edited editions, and the reason line volume must be read in context.

Edition-level work can also register as M. Aurélien Berra’s conversion of Athenaeus (tlg0008) and Aristophanes (tlg0019) shows only ~2 modified files but +34,000 lines: he rewrote existing stubs into full EpiDoc editions and validated them with HookTest before submitting.

Edition-level work that registers as ‘M’ — line volume, EXT tier
big = [r for r in rows if r['tier'] == 'EXT' and r['ladd'] > 2000]
print(f"{'contributor':22s}{'new':>4}{'corr':>5}{'+lines':>9}{'-lines':>9}  span")
for r in sorted(big, key=lambda r: -r['ladd']):
    print(f"{r['name'][:22]:22s}{r['new']:4d}{r['corr']:5d}"
          f"{r['ladd']:9d}{r['ldel']:9d}  {r['span']}")
contributor            new corr   +lines   -lines  span
Aurélien Berra           0    2    34351    27727  2019-03..2019-07
Nathaniel McCallum       0   84    14895    14895  2023-10..2023-11
Charles Pletcher         1    1     4799     4788  2023-07..2023-07
David Fifield            0   59     2226     2225  2021-05..2026-07

4.3 Mode 3 — Batch normalization (one operation, many files)

Batch-style contributions: many editions touched, low lines-per-file
cand = [r for r in rows if r['tier'] in ('EXT', 'OGL') and r['corr'] >= 10]
print(f"{'contributor':22s}{'tier':4s}{'editions':>9}{'+lines':>9}{'lines/file':>11}")
for r in sorted(cand, key=lambda r: -r['corr']):
    lpf = r['ladd'] / r['corr'] if r['corr'] else 0
    print(f"{r['name'][:22]:22s}{r['tier']:4s}{r['corr']:9d}"
          f"{r['ladd']:9d}{lpf:11.1f}")
contributor           tier editions   +lines lines/file
Matthew Munson        OGL      1175     1429        1.2
Thibault Clérice      OGL       161  3996745    24824.5
Nathaniel McCallum    EXT        84    14895      177.3
David Fifield         EXT        59     2226       37.7
Helma Dik             EXT        29     1933       66.7
Chris Drymon          EXT        18       85        4.7
Julia Jennifer Beine  EXT        17       57        3.4
Scott Fleischman      EXT        14      179       12.8
Stella Dee            OGL        12   113037     9419.8

A single well-scoped normalization can touch dozens or hundreds of files. Matthew Munson (OGL, funded staff) normalized ~1,175 editions in short bursts. Among outside volunteers, Nathaniel McCallum (a software engineer) is the clearest case: 84 editions in about one month, with commit messages that name the operation — Normalize lang="greek" => lang="grc", Correctly identify translations, Fix incorrect div type assignment — and near-equal added and deleted counts, consistent with a line-for-line automated rewrite that was then spot-checked. Scott Fleischman, also an outside volunteer, did the same at smaller scale. This mode requires no classical training, only care and one clear objective per PR.

5 Profiles of note

Funded engineers (Leipzig / Open Greek and Latin). A small group with concentrated structural impact. Clérice built the CTS/EpiDoc plumbing and ran the migration; Munson ran corpus-wide normalizations; Dee and Stoyanova did migration and conversion. The work is bursty — a few active months — and infrastructural, the opposite of the outside correctors’ long, thin streams.

Interns and fellows (CHS / Perseids). The main source of new editions. They work one text-group at a time (Konieczny and Hanhardt mostly within Galen, tlg0007), one file per commit, converting to EpiDoc/CTS. Hanhardt links 96% of her commits to a tracked issue, the strongest issue-discipline in the dataset, which is how a supervised intern pipeline is expected to look.

Outside volunteers and independent scholars. The long tail (~26 people) is dominated by correctors, in two recurring archetypes: the domain scholar (Dik, Beine, Driscoll, Berra) fixing philology in texts they study, and the software engineer (Fifield, McCallum, Fleischman, Krahn) fixing markup and encoding at scale. Both succeed by keeping to a single, well-defined scope.

6 What made them successful (and how it maps to the guide)

PR workflow and commit-style metrics
# PR cadence and branch style from the merges export
prs = 0; branch = Counter(); merger = Counter()
patch = re.compile(r'-patch-|patch-\d+', re.I)
for line in open('data/perseus-merges.tsv', encoding='utf-8', errors='replace'):
    f = line.rstrip('\n').split('\t')
    if len(f) < 4:
        continue
    m = re.search(r'Merge pull request #(\d+) from (\S+)', f[3])
    if not m:
        continue
    prs += 1; merger[f[1]] += 1
    br = m.group(2)
    branch['single-fix patch branch' if patch.search(br)
           else 'topic branch (tlgXXXX)' if '/tlg' in br.lower()
           else 'other named branch'] += 1
print(f"Total merged PRs: {prs}")
print("Merge rights (PRs merged, top 6):")
for who, n in merger.most_common(6):
    print(f"  {who:16s}{n:5d}")
print("Branch style:")
for style, n in branch.most_common():
    print(f"  {style:26s}{n:5d}")

# commit granularity + issue-linking for exemplar contributors
issue = re.compile(r'#\d+')
vague = re.compile(r'^(edits?|text fixes?|update|changes?|fixes?)\.?$', re.I)
FOCUS = {'David Fifield', 'Nathaniel McCallum', 'helmadik',
         'Julia Jennifer Beine', 'Chris Drymon', 'Michael Konieczny',
         'Angelia Hanhardt'}
agg = defaultdict(lambda: dict(n=0, iss=0, vg=0, ml=0))
for line in open('data/perseus-authors.tsv', encoding='utf-8', errors='replace'):
    f = line.rstrip('\n').split('\t')
    if len(f) < 5 or f[1] not in FOCUS:
        continue
    c = canon(f[1]); a = agg[c]; a['n'] += 1; a['ml'] += len(f[4])
    if issue.search(f[4]): a['iss'] += 1
    if vague.match(f[4].strip()): a['vg'] += 1
print(f"\n{'contributor':22s}{'commits':>8}{'%iss':>6}{'%vague':>8}{'avg_len':>9}")
for c, a in sorted(agg.items(), key=lambda x: -x[1]['n']):
    n = max(a['n'], 1)
    print(f"{c[:22]:22s}{a['n']:8d}{100*a['iss']//n:5d}%"
          f"{100*a['vg']//n:7d}%{a['ml']//n:9d}")
Total merged PRs: 1155
Merge rights (PRs merged, top 6):
  Lisa Cerrato      964
  Alison Babeu       86
  helmadik           25
  Bridget Almas      22
  gregorycrane       20
  srdee              19
Branch style:
  other named branch          611
  topic branch (tlgXXXX)      378
  single-fix patch branch     166

contributor            commits  %iss  %vague  avg_len
Michael Konieczny           99    3%      0%       17
Angelia Hanhardt            99   96%      0%       36
David Fifield               58    8%      0%       60
Helma Dik                   35    0%      0%       19
Nathaniel McCallum          26    0%      0%       28
Julia Jennifer Beine        13    0%      0%       17
Chris Drymon                 3    0%      0%       16

Six behaviors recur among successfully merged contributions; each maps to a section of the companion guide perseus-PR-template.md:

  1. Atomic commits — one fix, one file. Every successful corrector averages about one file per commit. → Guide §5, “One work per PR.”
  2. Self-describing messages. Successful contributors name the URN and the exact change and never ship “edits” as the sole message (0% vague among the exemplars); Fifield’s average message is ~60 characters of specifics. → Guide §3, the <change> note, and §5, the PR description.
  3. Issue linking. Tying a change to a tracked issue (#1234) — near-universal in the intern pipeline — identifies the report a fix resolves. → Guide §0 (open an issue first) and §5.
  4. Validate before submitting. The CI gate is HookTest against the EpiDoc/CapiTainS-CTS scheme; contributors who run it locally first (Berra documents doing so) are not bounced. → Guide §1 and §4.
  5. Keep to one scope. Scholars correct philology; engineers normalize markup or convert editions. → Guide §6.
  6. Contribute consistently. The most prolific outside contributor spread ~58 commits across five years; sustained small contributions build the reviewer trust that, for Helma Dik, culminated in merge rights (she merged 25 of her own patch-branch PRs). → Guide §6, durable-success habits.

The PR mechanics agree: of 1,155 merged PRs, a large share arrive through single-fix patch branches (the you/you-patch-N workflow) and per-text-group topic branches — both small, reviewable units. There is no evidence of large omnibus PRs succeeding.

7 Reproducibility

Every figure above is produced by the embedded chunks running against the four TSVs. To regenerate the inputs from a full clone:

git log --no-merges --name-status --no-renames --date=short \
  --pretty=format:'C%x09%H%x09%an%x09%ae%x09%ad%x09%s' \
  -- 'data/**/*.xml' > data/perseus-editlog.tsv

git log --no-merges --date=short \
  --pretty=format:'%H%x09%an%x09%ae%x09%ad%x09%s' > data/perseus-authors.tsv

git log --no-merges --numstat --no-renames --date=short \
  --pretty=format:'C%x09%H%x09%an%x09%ae%x09%ad%x09%s' \
  -- 'data/**/*.xml' > data/perseus-numstat.tsv

git log --merges --date=short \
  --pretty=format:'%H%x09%an%x09%ad%x09%s' > data/perseus-merges.tsv

Then quarto render (or quarto preview while editing). The tier map in the setup chunk is the one place to audit: it is deterministic and identity-based, so adding a newly identified person, or moving someone between tiers, is a one-line edit that flows through every table.

Known limits.

  1. Only merged history is visible; rejected PRs are not counted.
  2. A/M is exact for new-vs-edited, but edition-scale conversions of existing stubs appear as high-line M (handled by reading line volume alongside).
  3. Tier boundaries involve judgment: the OGL engineers were paid by the broader project but not by Tufts, and CHS interns are program participants rather than volunteers. Both are reported separately so the reader can draw the line where they prefer.