---
title: "What Makes for Successful Contributions to Perseus `canonical-greekLit`"
subtitle: "A tiered analysis of the commit record, 2013–2026"
author: "LLM-generated, prompted and supervised by Mark G. Bilby"
date: late updated, July 29, 2026
jupyter: python3
execute:
echo: true
warning: false
---
> **Companion document.** This report is the evidence base for the checklist in
> the [PR template](template.qmd) ("A Template for Successful Pull Requests"),
> kept in the repository as `perseus-PR-template.md`. Each success factor in §6
> maps to a section of that guide.
# Executive summary
The `PerseusDL/canonical-greekLit` history contains roughly 3,900 non-merge
commits and 1,155 merged pull requests, spanning mid-2013 to mid-2026.
Contributions divide into four tiers of contributor and three modes of work.
Only the fourth tier is unpaid; the first three are compensated — long-term
staff, contract engineers, and funded interns.
Tiers, by who the contributor is:
* **Tufts PDL long-term staff** (Cerrato, Babeu, Crane, Almas, Buckingham):
most of the volume. Expected, and not the subject of this report.
* **Leipzig / Open Greek and Latin engineers** on fixed-term contracts (Clérice,
Munson, Dee, Stoyanova): funded engineers, not volunteers, responsible for most
structural change in short, concentrated windows.
* **CHS / Perseids interns and fellows** (Konieczny, Hanhardt, shared intern
accounts): funded, and the principal source of new editions.
* **Outside volunteers and independent scholars** (Fifield, McCallum, Dik, Beine,
Drymon, Driscoll, Berra, and ~20 others): the only unpaid tier, and almost
exclusively correctors of existing editions.
Modes, by what a contribution does:
1. **Micro-corrections** — single-line philological, OCR, or diacritic fixes.
2. **Edition-level work** — converting or authoring a whole EpiDoc/CTS edition.
3. **Batch normalization** — one systematic markup or metadata fix across files.
**Principal finding.** Substantial, durable contribution does not require staff
status. The most prolific outside contributor, David Fifield, corrected ~59
editions over five years without creating a new file; one independent scholar,
Helma Dik, earned merge rights; and the interns who converted the Galen corpus
each created 65–70 editions in a single term. The shared method is set out in §6.
# Data and method
Four TSVs were exported from a full clone (`git log`, default branch, merges
excluded except where noted):
| File | Contents | Used for |
|---|---|---|
| `perseus-editlog.tsv` | `A`/`M`/`D` status per file, scoped to `data/**/*.xml` | new-edition vs. correction counts |
| `perseus-authors.tsv` | one line per commit: hash, name, email, date, subject | totals, cadence, commit style |
| `perseus-numstat.tsv` | added/deleted lines per file, scoped to `data/**/*.xml` | magnitude of change |
| `perseus-merges.tsv` | merge commits with PR number and source branch | PR workflow, cadence |
**Caveats.**
1. The log records what landed; rejected or abandoned PRs are absent, so "what
gets rejected" is inferred, not measured.
2. `A`/`M` separates a new file from an edited one, but an edition-scale
conversion of a pre-existing stub registers as `M` with a large `+lines`; line
volume is therefore reported alongside `A`/`M` (see Berra, below).
3. Identity is resolved to a person across aliases and emails, and each person is
assigned a tier deterministically. The mapping is printed below so it can be
audited.
```{python}
#| label: setup
#| code-summary: "Identity resolution, tier classification, and TSV parsers"
import re
from collections import defaultdict, Counter
# --- alias -> canonical person ---
CANON = {
'lcerrato':'Lisa Cerrato','Lisa Cerrato':'Lisa Cerrato','cerrato':'Lisa Cerrato',
'Alison Babeu':'Alison Babeu','AlisonBabeu':'Alison Babeu',
'gregorycrane':'Gregory Crane','Gregory Crane':'Gregory Crane','Crane':'Gregory Crane',
'Bridget Almas':'Bridget Almas','balmas':'Bridget Almas','balmas01':'Bridget Almas',
'TDBuck':'Tim Buckingham','Tim Buckingham':'Tim Buckingham',
'Thibault Clérice':'Thibault Clérice','Thibault Clerice':'Thibault Clérice',
'sonofmun':'Matthew Munson','Matthew Munson':'Matthew Munson',
'Stella Dee':'Stella Dee','srdee':'Stella Dee',
'simonastoyanova':'Simona Stoyanova','stoyanova':'Simona Stoyanova',
'Michael Konieczny':'Michael Konieczny','Angelia Hanhardt':'Angelia Hanhardt',
'helmadik':'Helma Dik','Helma Dik':'Helma Dik',
'David Fifield':'David Fifield','Nathaniel McCallum':'Nathaniel McCallum',
'Julia Jennifer Beine':'Julia Jennifer Beine','juliajbeine':'Julia Jennifer Beine',
'Charles Pletcher':'Charles Pletcher','Jacob Wegner':'Jacob Wegner',
'Isaac Bennett-Smith':'Isaac Bennett-Smith','Stephen Scott':'Stephen Scott',
'Chris Drymon':'Chris Drymon','chrisdrymon':'Chris Drymon',
'David F. Driscoll':'David F. Driscoll','David Smith':'David Smith (dasmiq)',
'Jeroen Hellingman':'Jeroen Hellingman','Aurélien Berra':'Aurélien Berra',
'ChiaraPalladino':'Chiara Palladino','Chiara Palladino':'Chiara Palladino',
'Adiel Mittmann':'Adiel Mittmann','Joel Kalvesmaki':'Joel Kalvesmaki',
'Scott Fleischman':'Scott Fleischman','Eric Sowell':'Eric Sowell',
'emNeven Jovanovic/em':'Neven Jovanović','KATEBHN':'KATEBHN',
'Kevin Krahn':'Kevin Krahn','kevinkrahn':'Kevin Krahn','TickleForce':'Kevin Krahn',
'achitu':'Andrei Chitu',
'Intern4':'CHS interns (shared)','Intern3':'CHS interns (shared)',
'Intern2':'CHS interns (shared)','Intern 1':'CHS interns (shared)',
'pubintern1':'Perseids pubinterns','pubintern2':'Perseids pubinterns',
'pubintern5':'Perseids pubinterns',
}
def canon(name):
return CANON.get(name, name)
# --- person -> tier (assigned by identity, NOT email domain) ---
# TUFTS: long-term paid staff. OGL: funded fixed-term Leipzig / Open Greek
# & Latin engineers. CHS: funded interns / fellows. EXT: outside volunteers.
TUFTS = {'Lisa Cerrato','Alison Babeu','Gregory Crane','Bridget Almas',
'Tim Buckingham','Rashmi Singhal','Anne Mahoney','William Merrill',
'Elli Mylonas','David Mimno'}
OGL = {'Thibault Clérice','Matthew Munson','Stella Dee','Simona Stoyanova',
'Monica Berti','Greta Franzini','Gabriel Weaver'}
CHS = {'Michael Konieczny','Angelia Hanhardt','CHS interns (shared)',
'Perseids pubinterns'}
def tier(name, email):
if ('github-actions' in name or name.lower().startswith('travis')
or 'dependabot' in name):
return 'BOT'
c = canon(name)
if c in TUFTS: return 'TUFTS' # long-term paid staff
if c in OGL: return 'OGL' # funded fixed-term engineers
if c in CHS: return 'CHS' # funded interns / fellows
return 'EXT' # outside volunteers
def is_edition(p):
return (p.startswith('data/') and p.endswith('.xml')
and not p.endswith('__cts__.xml'))
HASH = re.compile(r'^[0-9a-f]{40}$')
def parse_editlog(path='data/perseus-editlog.tsv'):
d = defaultdict(lambda: dict(A=set(), M=set(), D=set()))
cat = {}; cur = None
for line in open(path, encoding='utf-8', errors='replace'):
f = line.rstrip('\n').split('\t')
if f and f[0] == 'C' and len(f) >= 6 and HASH.match(f[1]):
cur = canon(f[2]); cat[cur] = tier(f[2], f[3])
elif cur and len(f) >= 2 and f[0][:1] in 'AMD':
if is_edition(f[-1]):
d[cur][f[0][0]].add(f[-1])
return d, cat
def parse_authors(path='data/perseus-authors.tsv'):
tot = Counter(); cat = {}; dates = defaultdict(list)
for line in open(path, encoding='utf-8', errors='replace'):
f = line.rstrip('\n').split('\t')
if len(f) < 5:
continue
c = canon(f[1]); t = tier(f[1], f[2])
if t == 'BOT':
continue
tot[c] += 1; cat[c] = t; dates[c].append(f[3])
return tot, cat, dates
def parse_numstat(path='data/perseus-numstat.tsv'):
add = Counter(); dele = Counter(); cur = None
for line in open(path, encoding='utf-8', errors='replace'):
f = line.rstrip('\n').split('\t')
if f and f[0] == 'C' and len(f) >= 6 and HASH.match(f[1]):
cur = (canon(f[2]), tier(f[2], f[3]))
elif cur and len(f) >= 3 and cur[1] != 'BOT' and is_edition(f[2]):
add[cur[0]] += int(f[0]) if f[0].isdigit() else 0
dele[cur[0]] += int(f[1]) if f[1].isdigit() else 0
return add, dele
ed, catE = parse_editlog()
tot, catA, dates = parse_authors()
add, dele = parse_numstat()
cat = {**catE, **catA}
def profile(c):
e = ed.get(c, {'A': set(), 'M': set(), 'D': set()})
ds = sorted(dates.get(c, []))
span = f"{ds[0][:7]}..{ds[-1][:7]}" if ds else ""
return dict(name=c, tier=cat.get(c, 'EXT'), commits=tot.get(c, 0),
new=len(e['A']), corr=len(e['M']),
ladd=add.get(c, 0), ldel=dele.get(c, 0),
months=len({d[:7] for d in ds}), span=span)
rows = [profile(c) for c in set(list(ed) + list(tot))]
print(f"Contributors (excl. bots): {len(rows)}")
print(f"Window: {min(min(dates[c]) for c in dates)[:7]}"
f" .. {max(max(dates[c]) for c in dates)[:7]}")
```
# The landscape: who contributes, and how much
```{python}
#| label: rollup
#| code-summary: "Category rollup"
roll = defaultdict(lambda: [0, 0, 0, 0])
for r in rows:
roll[r['tier']][0] += 1
roll[r['tier']][1] += r['commits']
roll[r['tier']][2] += r['new']
roll[r['tier']][3] += r['corr']
print(f"{'tier':6s}{'people':>7}{'commits':>9}{'new':>7}{'corrected':>11}")
for k in ['TUFTS', 'OGL', 'CHS', 'EXT']:
v = roll[k]
print(f"{k:6s}{v[0]:7d}{v[1]:9d}{v[2]:7d}{v[3]:11d}")
```
Long-term Tufts staff account for most commits and most corrections, since
maintaining the collection is their job and the managing editor also commits
others' reviewed work. The four OGL engineers, on ~119 commits between them,
account for the largest share of new editions: one of them ran the CTS migration
that created the file tree. Interns are the next-largest source of new editions.
The outside (unpaid) tier is numerous (~26 people) but corrects rather than
creates — three new editions among them, against ~250 corrected editions.
The rest of this report sets long-term staff aside and examines the three groups
an entrant can join: funded engineers, interns, and outside volunteers.
# Three modes of successful contribution
## Mode 1 — Micro-corrections (outside volunteers)
```{python}
#| label: correctors
#| code-summary: "Top correctors among outside volunteers (EXT), by editions modified"
outside = [r for r in rows if r['tier'] == 'EXT']
print(f"{'contributor':22s}{'tier':4s}{'corr':>8}{'commits':>8}{'mo':>4} span")
for r in sorted(outside, key=lambda r: -r['corr'])[:12]:
print(f"{r['name'][:22]:22s}{r['tier']:4s}{r['corr']:8d}"
f"{r['commits']:8d}{r['months']:4d} {r['span']}")
```
This is the most accessible path. **David Fifield** (a software engineer, not a
classicist) is representative: ~59 editions corrected across 16 active months
over 2021–2026, each a single-purpose commit naming the URN and the exact change
— e.g. *add missing first letter of ὡς in 14.22*; *change Latin o to Greek
omicron*; *use `<q>` for nested quotations*. **Helma Dik** (University of Chicago;
Logeion / Perseus-under-PhiloLogic) submits OCR and diacritic corrections in many
small PRs and is the only outside contributor with merge rights (§6). **Julia
Jennifer Beine** and **Chris Drymon** follow the same single-fix pattern. A
sustained stream of small, verifiable corrections is itself a first-class
contribution.
## Mode 2 — Edition-level work (interns and funded engineers)
```{python}
#| label: neweditions
#| code-summary: "Top creators of new editions (new file added), staff excluded"
makers = [r for r in rows if r['tier'] in ('EXT', 'OGL', 'CHS') and r['new'] > 0]
print(f"{'contributor':22s}{'tier':4s}{'new':>6}{'+lines':>10}{'mo':>4} span")
for r in sorted(makers, key=lambda r: -r['new'])[:10]:
print(f"{r['name'][:22]:22s}{r['tier']:4s}{r['new']:6d}"
f"{r['ladd']:10d}{r['months']:4d} {r['span']}")
```
New editions come overwhelmingly from interns and funded engineers. The CHS
interns who converted the Galen corpus each produced 65–70 new EpiDoc/CTS
editions in one term (Konieczny 69, Hanhardt 66, the shared intern account 70).
**Thibault Clérice**'s large `new`/`+lines` figure is the 2014–15 CTS migration
that created the canonical `data/` tree — an infrastructural act, not 1,500
hand-edited editions, and the reason line volume must be read in context.
Edition-level work can also register as `M`. **Aurélien Berra**'s conversion of
Athenaeus (`tlg0008`) and Aristophanes (`tlg0019`) shows only ~2 modified files
but **+34,000 lines**: he rewrote existing stubs into full EpiDoc editions and
validated them with HookTest before submitting.
```{python}
#| label: berra
#| code-summary: "Edition-level work that registers as 'M' — line volume, EXT tier"
big = [r for r in rows if r['tier'] == 'EXT' and r['ladd'] > 2000]
print(f"{'contributor':22s}{'new':>4}{'corr':>5}{'+lines':>9}{'-lines':>9} span")
for r in sorted(big, key=lambda r: -r['ladd']):
print(f"{r['name'][:22]:22s}{r['new']:4d}{r['corr']:5d}"
f"{r['ladd']:9d}{r['ldel']:9d} {r['span']}")
```
## Mode 3 — Batch normalization (one operation, many files)
```{python}
#| label: batch
#| code-summary: "Batch-style contributions: many editions touched, low lines-per-file"
cand = [r for r in rows if r['tier'] in ('EXT', 'OGL') and r['corr'] >= 10]
print(f"{'contributor':22s}{'tier':4s}{'editions':>9}{'+lines':>9}{'lines/file':>11}")
for r in sorted(cand, key=lambda r: -r['corr']):
lpf = r['ladd'] / r['corr'] if r['corr'] else 0
print(f"{r['name'][:22]:22s}{r['tier']:4s}{r['corr']:9d}"
f"{r['ladd']:9d}{lpf:11.1f}")
```
A single well-scoped normalization can touch dozens or hundreds of files.
**Matthew Munson** (OGL, funded staff) normalized ~1,175 editions in short
bursts. Among outside volunteers, **Nathaniel McCallum** (a software engineer) is
the clearest case: 84 editions in about one month, with commit messages that name
the operation — *Normalize `lang="greek"` => `lang="grc"`*, *Correctly identify
translations*, *Fix incorrect div type assignment* — and near-equal added and
deleted counts, consistent with a line-for-line automated rewrite that was then
spot-checked. **Scott Fleischman**, also an outside volunteer, did the same at
smaller scale. This mode requires no classical training, only care and one clear
objective per PR.
# Profiles of note
**Funded engineers (Leipzig / Open Greek and Latin).** A small group with
concentrated structural impact. Clérice built the CTS/EpiDoc plumbing and ran the
migration; Munson ran corpus-wide normalizations; Dee and Stoyanova did migration
and conversion. The work is bursty — a few active months — and infrastructural,
the opposite of the outside correctors' long, thin streams.
**Interns and fellows (CHS / Perseids).** The main source of new editions. They
work one text-group at a time (Konieczny and Hanhardt mostly within Galen,
`tlg0007`), one file per commit, converting to EpiDoc/CTS. Hanhardt links 96% of
her commits to a tracked issue, the strongest issue-discipline in the dataset,
which is how a supervised intern pipeline is expected to look.
**Outside volunteers and independent scholars.** The long tail (~26 people) is
dominated by correctors, in two recurring archetypes: the domain scholar (Dik,
Beine, Driscoll, Berra) fixing philology in texts they study, and the software
engineer (Fifield, McCallum, Fleischman, Krahn) fixing markup and encoding at
scale. Both succeed by keeping to a single, well-defined scope.
# What made them successful (and how it maps to the guide)
```{python}
#| label: workflow
#| code-summary: "PR workflow and commit-style metrics"
# PR cadence and branch style from the merges export
prs = 0; branch = Counter(); merger = Counter()
patch = re.compile(r'-patch-|patch-\d+', re.I)
for line in open('data/perseus-merges.tsv', encoding='utf-8', errors='replace'):
f = line.rstrip('\n').split('\t')
if len(f) < 4:
continue
m = re.search(r'Merge pull request #(\d+) from (\S+)', f[3])
if not m:
continue
prs += 1; merger[f[1]] += 1
br = m.group(2)
branch['single-fix patch branch' if patch.search(br)
else 'topic branch (tlgXXXX)' if '/tlg' in br.lower()
else 'other named branch'] += 1
print(f"Total merged PRs: {prs}")
print("Merge rights (PRs merged, top 6):")
for who, n in merger.most_common(6):
print(f" {who:16s}{n:5d}")
print("Branch style:")
for style, n in branch.most_common():
print(f" {style:26s}{n:5d}")
# commit granularity + issue-linking for exemplar contributors
issue = re.compile(r'#\d+')
vague = re.compile(r'^(edits?|text fixes?|update|changes?|fixes?)\.?$', re.I)
FOCUS = {'David Fifield', 'Nathaniel McCallum', 'helmadik',
'Julia Jennifer Beine', 'Chris Drymon', 'Michael Konieczny',
'Angelia Hanhardt'}
agg = defaultdict(lambda: dict(n=0, iss=0, vg=0, ml=0))
for line in open('data/perseus-authors.tsv', encoding='utf-8', errors='replace'):
f = line.rstrip('\n').split('\t')
if len(f) < 5 or f[1] not in FOCUS:
continue
c = canon(f[1]); a = agg[c]; a['n'] += 1; a['ml'] += len(f[4])
if issue.search(f[4]): a['iss'] += 1
if vague.match(f[4].strip()): a['vg'] += 1
print(f"\n{'contributor':22s}{'commits':>8}{'%iss':>6}{'%vague':>8}{'avg_len':>9}")
for c, a in sorted(agg.items(), key=lambda x: -x[1]['n']):
n = max(a['n'], 1)
print(f"{c[:22]:22s}{a['n']:8d}{100*a['iss']//n:5d}%"
f"{100*a['vg']//n:7d}%{a['ml']//n:9d}")
```
Six behaviors recur among successfully merged contributions; each maps to a
section of the companion guide `perseus-PR-template.md`:
1. **Atomic commits — one fix, one file.** Every successful corrector averages
about one file per commit. → *Guide §5, "One work per PR."*
2. **Self-describing messages.** Successful contributors name the URN and the
exact change and never ship "edits" as the sole message (0% vague among the
exemplars); Fifield's average message is ~60 characters of specifics.
→ *Guide §3, the `<change>` note, and §5, the PR description.*
3. **Issue linking.** Tying a change to a tracked issue (`#1234`) — near-universal
in the intern pipeline — identifies the report a fix resolves.
→ *Guide §0 (open an issue first) and §5.*
4. **Validate before submitting.** The CI gate is HookTest against the
EpiDoc/CapiTainS-CTS scheme; contributors who run it locally first (Berra
documents doing so) are not bounced. → *Guide §1 and §4.*
5. **Keep to one scope.** Scholars correct philology; engineers normalize markup
or convert editions. → *Guide §6.*
6. **Contribute consistently.** The most prolific outside contributor spread ~58
commits across five years; sustained small contributions build the reviewer
trust that, for Helma Dik, culminated in merge rights (she merged 25 of her own
patch-branch PRs). → *Guide §6, durable-success habits.*
The PR mechanics agree: of 1,155 merged PRs, a large share arrive through
single-fix patch branches (the `you/you-patch-N` workflow) and per-text-group
topic branches — both small, reviewable units. There is no evidence of large
omnibus PRs succeeding.
# Reproducibility
Every figure above is produced by the embedded chunks running against the four
TSVs. To regenerate the inputs from a full clone:
```bash
git log --no-merges --name-status --no-renames --date=short \
--pretty=format:'C%x09%H%x09%an%x09%ae%x09%ad%x09%s' \
-- 'data/**/*.xml' > data/perseus-editlog.tsv
git log --no-merges --date=short \
--pretty=format:'%H%x09%an%x09%ae%x09%ad%x09%s' > data/perseus-authors.tsv
git log --no-merges --numstat --no-renames --date=short \
--pretty=format:'C%x09%H%x09%an%x09%ae%x09%ad%x09%s' \
-- 'data/**/*.xml' > data/perseus-numstat.tsv
git log --merges --date=short \
--pretty=format:'%H%x09%an%x09%ad%x09%s' > data/perseus-merges.tsv
```
Then `quarto render` (or `quarto preview` while editing). The tier map in the setup
chunk is the one place to audit: it is deterministic and identity-based, so
adding a newly identified person, or moving someone between tiers, is a one-line
edit that flows through every table.
**Known limits.**
1. Only merged history is visible; rejected PRs are not counted.
2. `A`/`M` is exact for new-vs-edited, but edition-scale conversions of existing
stubs appear as high-line `M` (handled by reading line volume alongside).
3. Tier boundaries involve judgment: the OGL engineers were paid by the broader
project but not by Tufts, and CHS interns are program participants rather than
volunteers. Both are reported separately so the reader can draw the line where
they prefer.