Coding Board Composition Constructs in the Governance Literature

2026-08-11
STATUS: three papers LLM-coded, waiting for human decision on disagreements (see "Disagreements to Check" tab) before LLM-coding all papers.

What this is

Workflow and status

Fields coded (per variable)

Exact instructions and definitions provided to the LLM are stated in the “LLM Prompt” section.

FieldWhat it records
name_in_paperThe variable's label as the authors write it — the label from the variable-definitions table, or the row label in a regression table — with nothing appended by the coder. Table locations go in appears_in_tables, descriptions in operationalization.
roleThe part the variable plays in the paper's main analysis: dv, iv, moderator, mediator, control, instrument. Other parts in other analyses go in notes.
scopemain (main specifications), robustness (robustness checks only), dropped (considered and excluded). Every variable in the empirical analyses gets a record — none skipped for seeming minor.
contentWhich construct the authors intend the variable to measure, from the Content Vocabulary below. Identified from the paper's own framing — the variable's name, the hypothesis or argument it serves, the literature cited when introducing it — never from the measurement formula alone. Null when the paper provides no framing anywhere (no guessing). not_applicable for variables measuring no person or board trait (firm outcomes, industry, macro). Multi-valued (an ordered list) when one measure's definition spans more than one construct — a multi-trait faultline, or a disjunctive category that fires on conditions from different trait families. The primary construct is listed first, and a paper counts toward every construct in the list, not only the first. Multi-valued content applies only to variables with an empty derived_from: a composite takes the content its own framing supports, never the union of its constituents', which stays recoverable from the constituents' own records.
content_evidenceThe framing evidence for the content choice: a short verbatim phrase plus location. Empty when content is null or not_applicable.
measurement_formHow the variable is measured over people, from the Form Vocabulary below. Decided by the unit the value varies over: differs across directors → individual; one value per board → a board_* form; a board feature that is not an aggregate of member traits → structural. Derived variables take a form only if one genuinely fits.
relative_toThe benchmark a relational measure is defined against: none, ceo, chair, other_directors, firm. The benchmark never changes the form, and benchmarking against the CEO or firm does not make the variable a CEO or firm variable.
levelWhose attribute the value is: director, board, ceo, firm, other. Person-trait contents apply at any person level — a CEO's age is (content: age, level: ceo). Board tables later filter to director/board.
derived_fromFor variables built from other named measures (composite index, interaction term, ratio): the constituent variable names. Each constituent gets its own record; the derived record is coded in its own right (unframed derived variables take content = null like anything else).
operationalizationThe precise definition of how the variable is constructed, as the authors define it.
transformationsLog, +1 offset, winsorization, reverse-coding, threshold dummy, standardization.
data_sourceWhere the variable's data comes from, as named by the authors.
measure_labelA short label: the operationalization phrase, plus the data source in parentheses where the paper names one. Full definitions stay in operationalization.
inclusion_quote (+ location)Verbatim quote where the authors justify why this variable is in the analysis. Empty if none exists — never a paraphrase.
measurement_quote (+ location)Verbatim quote justifying why the variable is measured this way. Empty if none. Quote policy for both: verbatim wording and word order; OCR spacing may be repaired; no word changed, added, or deleted.
appears_in_tablesThe tables/panels where the variable appears.
notesAmbiguities, secondary roles, suspected conversion artifacts, judgment calls.

Content field vocabulary list

IDDefinitionBoundary notes
ageA person's age: years, bins, retirement-age dummies.
genderGender; female indicators; gender in matching or tie definitions.Gender-based ties also code networks.
race_ethnicityRace or ethnicity; minority indicators.
nationalityThe person's own nationality or foreign status.International experience is expertise.
political_ideologyPolitical leanings, usually donation-based; ideological distance.Office-holding is political_experience.
risk_preferencesRevealed risk preferences (e.g., option-exercise behavior).
personality_psychologyPersonality and cognition: narcissism, overconfidence, cognitive style.Risk attitudes are risk_preferences.
experience_tenureExperience as a director: tenure on the focal board, count of current/past board seats, first-time-director indicators, years served with the current CEO.Seat counts framed as attention constraints are busyness. Executive work experience is expertise.
educationDegree level, institution prestige/elite lists, field of study, exam scores.Education-based ties (same school) are networks.
expertiseSkills, functional background (finance, legal, accounting, media, medicine), industry experience, executive experience (including being a CEO elsewhere), country-specific/foreign experience.Political office is political_experience; board seats as experience are experience_tenure.
political_experienceHeld political office or worked in government; years in government.Donation-based ideology is political_ideology.
busynessOvercommitment: directorship-count thresholds, seat counts framed as attention constraints, outside employment.Temporary shocks are distraction.
distractionTemporary attention shocks (events at the person's other firms).Chronic overload is busyness.
attendanceMeeting attendance records.Count of meetings held is board_structure.
networksTies and network positions: interlocks, employment/education/social ties (including ties to the CEO or chair), centrality at any level.
statusStanding in a prestige hierarchy, usually composite; directorship counts explicitly framed as status.Negative standing is stigma_reputation.
stigma_reputationNegative or at-risk public standing: fraud taint, lawsuit association, media reputation.
independenceIndependence classification (exchange/ISS definitions; inside/outside/affiliated).
pay_ownershipCompensation; equity ownership; incentive structure.Blockholder-scale stakes are blockholder.
blockholderThe person is, or represents, a large shareholder.
founder_familyFounder or founding/controlling-family member.
ceo_dualityThe CEO also holds the board chair role, or the two roles are separated.The chair role on its own, held by anyone, is board_chair.
board_chairHolds the chair (or presiding chair) role of the board itself, as an attribute of a person.Whether the CEO is the one holding it is ceo_duality; chairing a board committee is committee_chair.
committee_chairChairs a board committee.
committee_memberSits on a board committee; assignment used as a variable.A paper using chair and membership variables codes both.
ceo_appointmentDirector joined the board during the sitting CEO's tenure (co-opted directors, CEO allies).“Currently a CEO elsewhere” is expertise.
election_outcomeShareholder votes on the director: election/withhold results, vote shares.
investor_appointedAppointed by or representing an investor, typically an activist.
lead_independent_directorLead/presiding independent director role.
board_structureStructural board features that are not aggregates of member traits: staggered/classified board, board size, meeting frequency.
not_applicableThe variable measures no person or board trait (firm characteristics and outcomes, industry and macro conditions).

Measurement form field vocabulary list

IDDefinition
individualMeasured per person, including role dummies and dyadic measures that vary across directors.
board_mean_levelBoard-level mean or level of a member trait.
board_share_countShare or count of members with a trait; totals of member traits.
board_dispersionHeterogeneity: SD, Blau/HHI, similarity indices, heterogeneity measures.
board_faultlineFaultline measures, including multi-trait ones (content is then multi-valued).
board_overlapPairwise overlap/similarity among members; uniqueness relative to the rest of the board (with relative_to = other_directors).
structuralBoard features that are not aggregates of member traits.

A variable can also have a benchmark, recorded separately in relative_to. The benchmark does not change the measurement form: average board tenure divided by CEO tenure is still a board mean — it is simply a board mean with relative_to = CEO.

Paper-level and coding-notes fields

BlockFields
papertitle, authors, year, journal; research_question (one sentence); theory_framing; samples (one entry per estimation sample: description, n, period); data_sources; empirical_methods; identification_strategy (with the authors' own caveats); measurement_innovation (reusable dictionaries, indices, coding protocols).
coding_notesconversion_quality; ambiguities (judgment calls); possible_missing (content suspected lost in PDF conversion).

The prompt the LLM coder sees

Verbatim copy of the coding prompt (prompt body only; the coder sees raw markdown). {PAPER_MD_PATH} is filled in with the paper being coded. Every returned file is checked against the schema before it is accepted.

You are coding one academic paper for a structured literature database on board composition research.
Your final message must be a single JSON object and nothing else; it is parsed by a program. Do not
summarize the paper, evaluate its quality, or address a human reader.

The paper (a markdown conversion of the PDF) is at: `{PAPER_MD_PATH}`

## How to read the paper

- Read the ENTIRE file. Page through it with the Read tool (offset/limit) until you reach the end.
- If the paper has a variable-definitions appendix or table, read it FIRST and use it as the spine of
  your inventory; then walk every regression table to catch variables the appendix omits; then read the
  text for framing and justification quotes.
- Expect PDF-conversion artifacts: minus signs rendered as "2" (a table cell "20.30" may be −0.30), "×"
  rendered as "3", "<" rendered as ",", dropped ligatures ("frm" for "firm"), collapsed whitespace, and
  unreadable figures (captions survive). When you suspect an artifact, record the suspicion in
  `coding_notes` instead of silently guessing.
- Never invent a variable, a number, or a quote. When you are unsure, say so in the `notes` field of the
  affected record; uncertainty belongs in the data, not resolved by guessing.

## Which variables get a record

Create one record for every variable that appears in the paper's empirical analyses: dependent
variables, focal predictors, moderators, mediators, instruments, and controls. Set `role` to the part
the variable plays in the paper's main analysis; if it plays different parts in different analyses, use
the main analysis and note the others in `notes`. Set `scope` to `main` for variables in the main
specifications, `robustness` for variables appearing only in robustness checks, and `dropped` for
variables the authors considered but excluded. Do not skip a variable because it seems minor.

### What counts as one record

These two rules fix the boundaries of the inventory. Apply them literally.

- **Do not record estimation apparatus.** Fixed effects of any kind, intercepts, time dummies, offsets,
  exposure terms, and weights are not variables. They measure no person or board trait and get no
  record, however prominently the tables list them.
- **A variable and its transformation are one record.** Where variable X enters the analyses only in
  transformed form — logged, winsorized, standardized, reverse-coded, or cut into a threshold dummy —
  create one record, named as the authors label it, and put the transformation in `transformations`.
  Create a second record only where the untransformed and transformed forms both enter as distinct
  regressors, and say so in `notes`.

### Reconcile before you emit

Check your inventory against the paper, and fix any gap you find:

1. Every row of the variable-definitions appendix or table has a record.
2. Every row of every regression table has a record, or is estimation apparatus excluded above.

This check leaves no trace in your output; just complete it.

## Field guide

Every entry below is the rule for that field. Vocabulary values are defined once, in the tables after
this guide; nothing here redefines them.

**`name_in_paper`** — the variable's label as the authors write it: the label from the
variable-definitions table, or the row label in a regression table. Add nothing of your own devising —
not a description, not a table reference, not a clarifying parenthetical. Table locations belong in
`appears_in_tables` and descriptions in `operationalization`.

**`role`** — one of: `dv`, `iv`, `moderator`, `mediator`, `control`, `instrument`.

**`scope`** — one of: `main`, `robustness`, `dropped` (see "Which variables get a record").

**`content`** — which construct the authors intend the variable to measure, as a Content Vocabulary ID.
Identify the construct from the paper's own framing — the variable's name, the hypothesis or argument it
serves, and the literature cited when introducing it — never from the measurement formula alone. If the
paper provides no framing anywhere (typically a control that is listed but never discussed), set
`content` to null rather than guessing. Variables that measure no person or board trait — firm
characteristics and outcomes, industry and macro conditions — take `content` = `not_applicable`.

Some variables are defined DISJUNCTIVELY: one measure fires on any of several conditions that belong
to different constructs, and the data cannot show which condition fired for a given person.
Disclosure-based categories are the common case — a single dummy defined so that it fires when a
person satisfies either of two conditions drawn from different trait families, sometimes operationalized
with a keyword list spanning both. For these, `content` is a LIST naming every construct the definition
can fire on, with the PRIMARY construct FIRST. The primary construct is the one the authors' own framing
leads with; where the framing does not settle it, the one covering most of the definition. Do not
silently pick one condition and drop the others, and do not reorder the list alphabetically or by
vocabulary order; the order is meaningful, not cosmetic. A single-construct variable stays a plain
string, not a one-element list.

This applies ONLY to variables whose `derived_from` is empty. A variable built from other named
measures takes only the content its OWN framing supports, or null — never the union of its
constituents' contents; the constituents carry those in their own records. If you find yourself
listing three or more contents for one variable, check whether it is in fact built from constituents;
if it is, name them in `derived_from` and leave the extra contents off this record.

**`content_evidence`** — the framing evidence for your `content` choice: a short verbatim phrase or
sentence plus its location. Empty string when `content` is null or `not_applicable`.

**`measurement_form`** — how the variable is measured over people, as a Measurement Form ID. Decide by
the unit the value varies over: it differs across individual directors → `individual`; one value per
board → the matching `board_*` form; a board feature that is not an aggregate of member traits →
`structural`. For a derived variable (see `derived_from`), assign a form only if the derived variable
itself matches one; otherwise null, with a note.

**`relative_to`** — one of: `none`, `ceo`, `chair`, `other_directors`, `firm`. The benchmark a
relational measure is defined against. The benchmark never changes `measurement_form`, and using the CEO
or the firm as a benchmark does not make the variable a CEO or firm variable.

**`level`** — one of: `director`, `board`, `ceo`, `firm`, `other`. Content values describe traits of
people and boards; code every person-trait variable with its content regardless of whose trait it is —
`level` records whether the person is a director or the CEO.

**`derived_from`** — when the variable is built from other named measures (a composite index, an
interaction term, a ratio of two variables), list the constituent variable names here, and create a
record for each constituent as well. Code every record — including the derived one — in its own right
under this field guide; a derived variable with no framing of its own takes `content` = null like any
other unframed variable. Empty list otherwise.

**`operationalization`** — the precise definition of how the variable is constructed, as the authors
define it (include sample statistics only if the authors state them).

**`transformations`** — measurement transformations applied: log, +1 offset, winsorization,
reverse-coding, threshold dummy, standardization. Empty string if none.

**`data_source`** — where the variable's data comes from, as named by the authors.

**`measure_label`** — a short label of the form "[operationalization phrase] ([data source])", or just
the phrase where the paper names no source. Full definitions belong in `operationalization`, never in
the label.

**`inclusion_quote`, `inclusion_quote_location`** — verbatim quote where the authors justify WHY this
variable is in the analysis, and where it appears. Empty strings if no such justification exists — never
substitute a paraphrase.

**`measurement_quote`, `measurement_quote_location`** — verbatim quote where the authors justify why
the variable is measured THIS WAY, and where it appears. Empty strings if none.

Quote policy for both quote fields: verbatim in wording and word order; you may repair spacing lost in
PDF conversion; you may not change, add, or delete a word.

**`appears_in_tables`** — the tables/panels where the variable appears (e.g., "Table 2; Table 5 Panel A").

**`notes`** — ambiguities, secondary roles, suspected conversion artifacts, judgment calls.

### Paper block fields

`title`, `authors`, `year`, `journal`; `research_question` (one sentence); `theory_framing` (the
literature/theory motivating the paper, brief); `samples` (a list — one entry per estimation sample:
description, n, period); `data_sources`; `empirical_methods`; `identification_strategy` (the causal
strategy, or "descriptive/associational", with the authors' own caveats if any); `measurement_innovation`
(any reusable measurement contribution — a new dictionary, index, or coding protocol — or empty string).

Types in this block: every field above is a STRING except `samples`, which is a list of OBJECTS, one
per estimation sample, each with exactly the keys `description`, `n` and `period`. Do not write
`samples` as a list of sentences.

### Coding-notes block fields

`conversion_quality` (were tables and definitions readable?); `ambiguities` (judgment calls you made);
`possible_missing` (content you suspect exists in the PDF but was lost in conversion). All three are
strings.

## Content Vocabulary

Assign the ID whose definition matches the construct the paper's framing invokes. Boundary notes resolve
neighboring pairs. Multi-valued where one measure genuinely spans trait families — a faultline combining
demographic and skill traits, or a disjunctive category that fires on conditions from different families.
List the primary construct first.

| ID | Definition | Boundary notes |
|---|---|---|
| `age` | A person's age: years, bins, retirement-age dummies. | |
| `gender` | Gender; female indicators; gender in matching or tie definitions. | Gender-based ties also code `networks`. |
| `race_ethnicity` | Race or ethnicity; minority indicators. | |
| `nationality` | The person's own nationality or foreign status. | International *experience* is `expertise`. |
| `political_ideology` | Political leanings, usually donation-based; ideological distance. | Office-holding is `political_experience`. |
| `risk_preferences` | Revealed risk preferences (e.g., option-exercise behavior). | |
| `personality_psychology` | Personality and cognition: narcissism, overconfidence, cognitive style. | Risk attitudes are `risk_preferences`. |
| `experience_tenure` | Experience as a director: tenure on the focal board, count of current/past board seats, first-time-director indicators, years served with the current CEO. | Seat counts framed as attention constraints are `busyness`. Executive work experience is `expertise`. |
| `education` | Degree level, institution prestige/elite lists, field of study, exam scores. | Education-based *ties* (same school) are `networks`. |
| `expertise` | Skills, functional background (finance, legal, accounting, media, medicine), industry experience, executive experience (including being a CEO elsewhere), country-specific/foreign experience. | Political office is `political_experience`; board seats as experience are `experience_tenure`. |
| `political_experience` | Held political office or worked in government; years in government. | Donation-based ideology is `political_ideology`. |
| `busyness` | Overcommitment: directorship-count thresholds, seat counts framed as attention constraints, outside employment. | Temporary shocks are `distraction`. |
| `distraction` | Temporary attention shocks (events at the person's other firms). | Chronic overload is `busyness`. |
| `attendance` | Meeting attendance records. | Count of meetings held is `board_structure`. |
| `networks` | Ties and network positions: interlocks, employment/education/social ties (including ties to the CEO or chair), centrality at any level. | |
| `status` | Standing in a prestige hierarchy, usually composite; directorship counts explicitly framed as status. | Negative standing is `stigma_reputation`. |
| `stigma_reputation` | Negative or at-risk public standing: fraud taint, lawsuit association, media reputation. | |
| `independence` | Independence classification (exchange/ISS definitions; inside/outside/affiliated). | |
| `pay_ownership` | Compensation; equity ownership; incentive structure. | Blockholder-scale stakes are `blockholder`. |
| `blockholder` | The person is, or represents, a large shareholder. | |
| `founder_family` | Founder or founding/controlling-family member. | |
| `ceo_duality` | The CEO also holds the board chair role, or the two roles are separated. | The chair role on its own, held by anyone, is `board_chair`. |
| `board_chair` | Holds the chair (or presiding chair) role of the board itself, as an attribute of a person. | Whether the CEO is the one holding it is `ceo_duality`; chairing a board committee is `committee_chair`. |
| `committee_chair` | Chairs a board committee. | |
| `committee_member` | Sits on a board committee; assignment used as a variable. | A paper using chair and membership variables codes both. |
| `ceo_appointment` | Director joined the board during the sitting CEO's tenure (co-opted directors, CEO allies). | "Currently a CEO elsewhere" is `expertise`. |
| `election_outcome` | Shareholder votes on the director: election/withhold results, vote shares. | |
| `investor_appointed` | Appointed by or representing an investor, typically an activist. | |
| `lead_independent_director` | Lead/presiding independent director role. | |
| `board_structure` | Structural board features that are not aggregates of member traits: staggered/classified board, board size, meeting frequency. | |
| `not_applicable` | The variable measures no person or board trait (firm characteristics and outcomes, industry and macro conditions). | |

## Measurement Form Vocabulary

| ID | Definition |
|---|---|
| `individual` | Measured per person, including role dummies and dyadic measures that vary across directors. |
| `board_mean_level` | Board-level mean or level of a member trait. |
| `board_share_count` | Share or count of members with a trait; totals of member traits. |
| `board_dispersion` | Heterogeneity: SD, Blau/HHI, similarity indices, heterogeneity measures. |
| `board_faultline` | Faultline measures, including multi-trait ones (content is then multi-valued). |
| `board_overlap` | Pairwise overlap/similarity among members; uniqueness relative to the rest of the board (with `relative_to` = `other_directors`). |
| `structural` | Board features that are not aggregates of member traits. |

## Worked examples (schematic)

These examples use placeholder names (X, A, Z, D) and placeholder content tokens (`<CONTENT_C>` stands
for whichever Content Vocabulary ID the paper's framing supports). They demonstrate mechanics and format
only; they carry no information about which contents are common or likely.

**Example 1 — a framed focal variable.** The paper says: *"Our main independent variable is X, measured
as [definition] following Author (Year). We argue that X captures [construct C], and expect it to
increase Y because [mechanism]."*

```json
{
  "name_in_paper": "X",
  "role": "iv",
  "scope": "main",
  "content": "<CONTENT_C>",
  "content_evidence": "\"We argue that X captures [construct C], and expect it to increase Y because [mechanism].\" (Hypotheses section, p. 7)",
  "measurement_form": "individual",
  "relative_to": "none",
  "level": "director",
  "derived_from": [],
  "operationalization": "[definition as the authors state it]",
  "transformations": "",
  "data_source": "[source named by the authors]",
  "measure_label": "[short phrase] ([source])",
  "inclusion_quote": "\"We argue that X captures [construct C]...\"",
  "inclusion_quote_location": "Hypotheses section, p. 7",
  "measurement_quote": "\"measured as [definition] following Author (Year)\"",
  "measurement_quote_location": "Data section, p. 12",
  "appears_in_tables": "Table 3, Columns 1-4",
  "notes": ""
}
```

**Example 2 — an unframed control.** The paper says only: *"Controls include A, B, and C."* and A is
never discussed again.

```json
{
  "name_in_paper": "A",
  "role": "control",
  "scope": "main",
  "content": null,
  "content_evidence": "",
  "measurement_form": "individual",
  "relative_to": "none",
  "level": "director",
  "derived_from": [],
  "operationalization": "[definition if given anywhere, e.g., a definitions appendix]",
  "transformations": "",
  "data_source": "[source if named]",
  "measure_label": "[short phrase]",
  "inclusion_quote": "",
  "inclusion_quote_location": "",
  "measurement_quote": "",
  "measurement_quote_location": "",
  "appears_in_tables": "Table 3, Columns 1-4",
  "notes": "Listed among controls only; no framing anywhere in the paper."
}
```

**Example 3 — a derived variable and its constituents.** The paper interacts X with Z in its moderation
test. X and Z each get their own full record (as in Examples 1-2); the interaction gets:

```json
{
  "name_in_paper": "X × Z",
  "role": "moderator",
  "scope": "main",
  "content": null,
  "content_evidence": "",
  "measurement_form": "individual",
  "relative_to": "none",
  "level": "director",
  "derived_from": ["X", "Z"],
  "operationalization": "Product of X and Z.",
  "transformations": "",
  "data_source": "",
  "measure_label": "X × Z",
  "inclusion_quote": "\"Hypothesis 2. The relationship between X and Y is weaker when Z is high.\"",
  "inclusion_quote_location": "Hypotheses section, p. 8",
  "measurement_quote": "",
  "measurement_quote_location": "",
  "appears_in_tables": "Table 4",
  "notes": "Interaction term; constituents X and Z recorded separately."
}
```

**Example 4 — a disjunctive category.** The paper defines a single dummy D as: *"D equals one if the
person satisfies condition P or condition Q"*, where P belongs to one trait family and Q to another,
and introduces D as a measure of P.

```json
{
  "name_in_paper": "D",
  "role": "iv",
  "scope": "main",
  "content": ["<CONTENT_P>", "<CONTENT_Q>"],
  "content_evidence": "\"D equals one if the person satisfies condition P or condition Q.\" (Data section, p. 9)",
  "measurement_form": "individual",
  "relative_to": "none",
  "level": "director",
  "derived_from": [],
  "operationalization": "Dummy equal to one if the person satisfies condition P or condition Q.",
  "transformations": "",
  "data_source": "[source if named]",
  "measure_label": "D dummy ([source])",
  "inclusion_quote": "",
  "inclusion_quote_location": "",
  "measurement_quote": "",
  "measurement_quote_location": "",
  "appears_in_tables": "Table 3",
  "notes": "Disjunctive category: the definition fires on P or Q, which fall in different trait families. <CONTENT_P> is listed first as the primary construct because the authors introduce D as a measure of P."
}
```

## Output

Your final message is exactly one JSON object with this shape, and no other text:

```json
{
  "paper": { "...paper block fields..." },
  "variables": [ { "...one record per variable..." } ],
  "coding_notes": { "...coding-notes block fields..." }
}
```

How LLM output is validated against the hand coding

  1. Parse the hand-coded sheet (article-categorization-manual-coding-080626.xlsx, 113 article rows × 33 category columns). Parsing rules: Federo et al. (2020) is excluded (its row was left uncoded, though the paper is still LLM-coded like the rest); Baik (2024)’s stray “2” is read as 1; the sheet’s own count column is ignored (the marked Yes cells are truth).
  2. Derive a comparison row per paper from the LLM's variable records, using the fixed (content, form) → column mapping below. Only records with level = director or board and scope = main enter; where a paper frames a variable nowhere, a documented defaults table supplies the content, and each record keeps a note of where its content came from.
  3. Score cell-level agreement per column (presence precision/recall, Cohen's kappa), under both scope variants (all main-specification variables vs focal-only) — the hand coding is internally mixed on controls, and the two scores reveal which rule it implicitly followed.
  4. Adjudicate every disagreement: a fresh agent re-reads just the disputed cell and issues a quote-backed verdict — model error / human error / genuinely ambiguous. Neither source is presumed right; only the ambiguous pile needs human review.

Example: LLM output to human-coded row of data

Three of the 60 records from Adams, Akyol & Verwijmeren (2018), “Director Skill Sets,” and the cells they produce. (Records abridged to the fields the mapping consumes.)

{ "name_in_paper": "Director tenure", "role": "control", "scope": "main", "content": "experience_tenure", "measurement_form": "individual", "level": "director", "measure_label": "Tenure" }
{ "name_in_paper": "Blau skill concentration", "role": "iv", "scope": "main", "content": "expertise", "measurement_form": "board_dispersion", "level": "board", "measure_label": "Skill Blau index" }
{ "name_in_paper": "Classified board", "role": "control", "scope": "main", "content": "board_structure", "measurement_form": "structural", "level": "board", "measure_label": "Classified board (0,1)" }
Sheet columnYesMeasure cell
Directorship Experience/Tenure1Tenure
Expertise Diversity/Levels1Skill Blau index
Staggered/Classified Board1

Each (content, form) pair looks up its column; the Measure cell concatenates the distinct measure labels of all records mapping to that column. The full row for this paper marks eleven columns and the hand-coded row marks eight. The differences are set out on the Disagreements screen.

The (content, form) → sheet-column mapping

Used for ground-truth comparison only; manuscript tables come from the records directly. Three columns are renamed relative to the spreadsheet so names reveal their level: Independence → Board Independence; Network Position → Board Network Position; Board Experience/Tenure → Directorship Experience/Tenure (a director-level category whose old name misread as board-level). Combinations not listed map to no column — new information the sheet could not distinguish. Counts (n) are how many of the 113 hand-coded articles mark the column.

Sheet columnnContent(s)Form(s)
Age31ageindividual
Gender31genderindividual
Nationality7nationalityindividual
Race/Ethnicity14race_ethnicityindividual
Political Ideology7political_ideologyindividual
Risk Aversion1risk_preferencesindividual
Directorship Experience/Tenure40experience_tenureindividual
Education19educationindividual
Expertise49expertiseindividual
Political Experience6political_experienceindividual
Busyness11busynessindividual
Distracted1distractionindividual
Meeting Attendance3attendanceindividual
Networks/Social Capital32networksindividual
Status2statusany
Stigma/Reputation3stigma_reputationindividual
Blockholder3blockholderindividual
Chairman/CEO Duality17ceo_dualityindividual, structural
Committee Chair8committee_chairindividual
Committee Member9committee_memberindividual
Current CEO Appointment7ceo_appointmentindividual
Election Outcome1election_outcomeindividual
Founder/Family3founder_familyindividual
Independent38independenceindividual
Investor Appointed1investor_appointedindividual
Lead Independent Director (LID)5lead_independent_directorindividual
Pay/Ownership11pay_ownershipindividual
Ascriptive Diversity24age, gender, race_ethnicity, nationalityany board form
Experience Diversity/Differences10experience_tenureany board form
Expertise Diversity/Levels17expertise, education, political_experienceany board form; also (expertise, individual + relative_to = other_directors)
Board Independence23independenceboard_share_count, board_mean_level
Board Network Position3networksany board form
Staggered/Classified Board5board_structure (staggered/classified)structural

Notes: status homogeneity (a board form of status) maps to Status per the hand coding's precedent. The Expertise Diversity/Levels extra clause covers director-level uniqueness measures relative to the rest of the board, which the hand coding files there. Multi-trait faultlines mark every diversity column their contents touch.

Tables for the paper

Table A — Existing major constructs

One row per construct: how much of the literature uses it, how it is operationalized, and why authors say they use it.

ConstructPapers using itTop operationalizationsCommon justifications
Counts from the records (papers with ≥1 qualifying record per content); operationalizations ranked by frequency of normalized measure labels, cut where usage share drops sharply; justifications synthesized by clustering the harvested inclusion/measurement quotes into recurring rationales, each backed by citable verbatim examples.

Table B — Papers measuring close to our approach

Detailed comparison for the constructs nearest our framework: expertise, experience, and networks. One row per variable variety: the varieties of measures, operationalization detail, and justification.

Table C — Measurement practice (enabled by the two-axis coding)

How the literature aggregates member traits: contents × measurement forms (individual vs mean vs share vs dispersion vs faultline vs overlap). Feeds the theory discussion on levels-vs-diversity measurement choices.

Open decisions for the tables

Decision 1. Composites and their components. Composites and their definitional components are each captured as their own variable record, with components linked to the composite via derived_from. Whether components count toward their own contents is therefore a per-table switch decided when each table is built, not a coding decision. The ground-truth comparison view currently includes components. The hand coding is not consistent on this: it marks the components of Acharya & Pollock's status composite but not those of You et al.'s power composite, which is why this needs settling. See Q2 on the Disagreements screen.
Decision 2. Do controls count, or only focal variables? Every variable record carries its role (DV, IV, moderator, mediator, control, instrument) and a robustness-only flag, so “do controls count toward a construct?” is a scope switch applied when outputs are derived; any table can be computed under either scope. The hand coding is internally mixed on this (Adams et al. 2018's committee-membership control is marked, but its outside-board-seats control is not), so validation is scored under both scopes. On the three papers coded so far, the scope that counts controls matches the hand coding better. The one question for the authors is which scope the manuscript's headline construct counts should use, and that can be decided as late as table construction.

Disagreements to check

The following presents cases where the LLM coding disagrees with the human coding in the 3 sample papers we coded using the LLM prompt. In most of these the human and LLM coding record the same construct and differ over the unit it should be recorded at (director or board). We need to review each of these disagreements to determine if the LLM prompt should be revised.

1 · Level of aggregation — You et al. (2023)

Cells: AgeGender Race/EthnicityExpertise Current CEO Appointment — all human coding only

The paper's control list names eleven board and director characteristics, and every one of them is an aggregate: a proportion, a heterogeneity index, or a similarity composite. Director age, gender and race never appear as per-director variables. They enter in one other place only, as inputs to the algorithm that sorts directors into identity subgroups.

One control shows the whole disagreement in a single row. For “the proportion of directors with CEO experience”, the human coding marked Expertise, a director-level column, with the Measure cell “Executive (CEO)”. The LLM coding recorded content expertise at form board_share_count, which the mapping sends to Expertise Diversity/Levels, a board-level column. Same variable, same construct, two different columns.

Across the sentence as a whole the two sides agree on all four board-level columns: Ascriptive Diversity, Experience Diversity/Differences, Expertise Diversity/Levels and Board Independence. The human coding's Measure cell for each of the three diversity columns reads “Faultine”, so both sides were looking at the same faultline measure. The disagreement is that the human coding also marked five director-level columns for those same traits, recording them twice; the LLM coding recorded them once, at the level the paper estimates.

One of the five is not a disagreement about level at all. “The proportion of directors appointed after the CEO” is a board-level share, but the sheet's only category for CEO appointment is director-level, so the mapping has nowhere to put it. The LLM coding did record the construct; the comparison simply cannot express it. The same is true of board-level job-title heterogeneity.

Human coding
Marks all five as director-level categories. Measure cells are blank except Expertise (“Executive (CEO)”).
LLM coding
Each trait is recorded, but only in the forms the paper uses:
  • age → Director age heterogeneity (board_dispersion), CEO age (level ceo)
  • gender → Proportion of female directors (board_share_count), Female CEO (ceo)
  • race → Director racial heterogeneity (board_dispersion), Minority CEO (ceo)
  • expertise → Proportion of directors with CEO experience (board_share_count)
  • ceo_appointment → Proportion of directors appointed after the CEO (board_share_count; no sheet column exists for this at board level)
All five also appear inside Faultline strength and Number of subgroups (board_faultline) and, for the demographics, CEO-board similarity (board_overlap).
“We controlled for 11 variables on board and director characteristics related to CEO dismissal or board subgroups (Westphal & Zajac, 1995): board size (the total number of directors), the proportion of independent directors, CEO-board similarity (a composite measure of age, gender, and race dissimilarities), board diversity (the proportion of female directors, director racial heterogeneity, director age heterogeneity using the coefficient of variation, the proportion of directors with CEO experience, the proportion of directors appointed after the CEO, director tenure heterogeneity, and director job title heterogeneity), faultline strength…” Data and Methods, p. 2823 — every board/director control in the paper, and every one of them is an aggregate
“we first calculated faultline strength using the average silhouette width (ASW) algorithm (Meyer & Glenz, 2013) and identified directors' membership of subgroups in each board based on five [director characteristics]” Data and Methods, p. 2822 — where director age, gender and race actually enter

2 · Ingredients of a composite index — Acharya & Pollock (2021) and You et al. (2023)

Cells: Directorship Experience/Tenure Education Expertise — human coding only, Acharya & Pollock  |  StatusCommittee Chair Pay/Ownership — LLM coding only, You et al.

Both papers build a headline variable out of named ingredients, and the human coding treats the two cases oppositely. That inconsistency is why this needs a decision rather than a ruling.

Acharya & Pollock. A director counts as “globally high status” if he or she holds any one of three credentials: a degree from an elite school, a vice-president-or-above post at an S&P 500 firm, or an outside directorship at an S&P 500 firm. That yes/no flag is then turned into a single board-level Blau index, and it is the index that enters the regressions. No model contains a director's education or occupation as a variable in its own right. The human coding marked all three credentials, and its Measure cells name them exactly: “Prestige”, “Occupation Employment prestige”, “Directorial prestige”. The LLM coding recorded the variable the paper estimates — one board-level record carrying all three constructs at once.

You et al. CEO subgroup power is built from three named indicators of an individual director's power: the director's title (chair of the board or of a committee), their number of external directorships, and their share ownership. Here the human coding marked none of the three, and did not mark the composite under Status either. The LLM coding recorded each named indicator under its own construct.

Two wrinkles sit inside the You et al. side. The count of external directorships is introduced only as an indicator of power, and the coding vocabulary has no entry for power; a directorship count counts as status only where the authors explicitly frame it as prestige, which they do not, which leaves that cell genuinely ambiguous. And “director title” runs together chairing the board with chairing a committee, two different roles, of which only the committee half has a column in the sheet.

Note that the Acharya & Pollock cells turn on the aggregation question in card 1 as well: the credentials are director-level, the estimated index is board-level.

Human coding — Acharya & Pollock
Marks the three credentials as director-level categories:
  • Directorship Experience/Tenure → “Directorial prestige (Proxies)”
  • Education → “Prestige (Proxies)”
  • Expertise → “Occupation Employment prestige (Proxies)”
It also marks Status, which the LLM coding marks too.
LLM coding — Acharya & Pollock
Records the variable the paper estimates:
  • Status Homogeneity (SH) — content [status, education, expertise], form board_dispersion, level board
  • Tenure Overlapexperience_tenure, board_overlap
  • Expertise Overlapexpertise, board_overlap
Human coding — You et al.
Marks none of the three indicators. The composite itself is not marked under Status either.
LLM coding — You et al.
One record per named indicator, each under its own construct:
  • director title[board_chair, committee_chair]
  • the number of external directorshipsstatus
  • director ownershippay_ownership
“we treated a director as globally high status if he or she possessed at least one of the following credentials: a degree from an elite educational institution listed in Finkelstein's (1992) list of prestigious institutions (educational prestige), experience as an executive at the level of vice president or above at an S&P 500 company (employment prestige), or experience as an outside director for an S&P 500 firm (directorial prestige).” Acharya & Pollock, Data and Methods, p. 1481
“we computed individual directors' power using three indicators (Zajac & Westphal, 1996): director title (coded as one for the chair of the board or a board committee and zero otherwise), the number of external directorships, and director ownership (the number of outstanding shares).”You et al., Data and Methods, p. 2822

3 · One dummy, two trait families — Adams, Akyol & Verwijmeren (2018)

Cells: Education — LLM coding only Political Experience — LLM coding only

The paper codes twenty director skills from proxy statements, each a yes/no dummy driven by a keyword list. The human coding treated the taxonomy as one measure and marked Expertise, with the Measure cell “Skills (Proxies)”. The LLM coding made one record per skill category, which agrees on Expertise but also picks up two categories whose definitions straddle a boundary.

“Academic” fires if a director is from academia or holds a higher degree. The first is an occupation, the second a credential, and the keyword list mixes them: dean, faculty and professor sit alongside doctorate, masters and Ph.D. Nothing in the data says which condition fired for a given director. “Government and policy” fires on any of four keywords — government, policy, politics, regulatory — so a director flagged on “regulatory” alone may have held no public office at all, yet the coding vocabulary deliberately separates political office from general expertise.

The LLM coding therefore records both constructs for these two dummies, primary one first. The question is whether a measure that genuinely conflates two things should count toward both, or only toward the one the authors lead with.

Human coding
Marks Expertise only, with the Measure cell “Skills (Proxies)” — all twenty proxy-statement skill categories treated as one measure under one construct.
LLM coding
One record per skill category. Two of the twenty have definitions that span two trait families, so they carry two contents, primary first:
  • Academic[expertise, education]
  • Government and policy[political_experience, expertise]
Both also mark Expertise, where the two agree.
Table 2, Skill categories (p. 645)
|Academic             |The director is from academia or has a higher degree (such as a Ph.D.).|
|Government and policy|The director has governmental, policy, or regulatory experience.       |

Appendix B.1, Skill dictionary (p. 660)
|Academic             |Academia, academic, dean, doctorate, education, faculty, graduate,
|                     | masters, Ph.D., professor, school environment                        |
|Government and policy|Government, policy, politics, regulatory                              |

4 · Apparent oversights — no standard needed, just confirmation

Cells: Board Independence — LLM coding only, Adams et al. Networks/Social Capital — human coding only, Acharya & Pollock

Board independence, Adams et al. The paper defines it in the variable appendix as the ratio of independent directors to board size, reports it in the descriptive statistics table, and carries it as a control in the main performance regressions (Table 5, columns 2, 4, 6 and 7, and Table 6 Panel C). The LLM coding recorded it; the human coding left the column blank. The same hand-coded row does mark the director-level Independent column for this paper, and that variable is also only a control, so controls were not being excluded as a class. The blank looks like an oversight rather than a convention.

Human coding
Board Independence left blank, though the director-level Independent column is marked.
LLM coding
Board independence — content independence, form board_share_count, level board, role control.

Networks, Acharya & Pollock. The human coding marked Networks/Social Capital with a blank Measure cell. The paper has no tie, interlock or centrality variable anywhere: its variables are tenure overlap, expertise overlap, status homogeneity, and firm-level controls. The one candidate is prior outside directorships at S&P 500 firms, but that is not a standalone variable — it is one of the three status credentials in card 2, which the paper labels directorial prestige and which the coding vocabulary sends to Status, a column both sides already mark. Nothing supports a separate networks mark, and the blank Measure cell is consistent with a stray tick.

Human coding
Networks/Social Capital marked yes, with a blank Measure cell.
LLM coding
No record for this paper carries networks. The paper has no tie, interlock or centrality variable.
Adams et al., Appendix A, Variable definitions (p. 659)
|Board independence|The ratio of independent directors on the board to the board size.|

Adams et al., Table 1, Panel A: Firm characteristics (p. 644)
|Variables         | Mean | Median| Std. dev.| Min. | Max. |
|Board independence| 0.795| 0.818 | 0.105    | 0.333| 0.941|

What we need to agree

Q1 — the aggregation question (card 1). When a paper uses a trait only in board-level form — a proportion, a heterogeneity index, a faultline input — should the director-level category be marked as well? This governs the five cells in card 1, and the three Acharya & Pollock cells in card 2 turn on it too.

Mark it: matches the sheet, and reads the categories as “does this paper engage with this trait?” Do not mark it: what the LLM coding does now, and keeps the individual and board-level columns meaning different things, which is what lets Table C distinguish levels from diversity. Note the sheet itself has separate board-level columns (Ascriptive Diversity, Board Independence), which suggests the distinction was intended.
Q2 — the composite-ingredient question (card 2). Should a named ingredient of a composite index count toward its own construct? The human coding does this in Acharya & Pollock and not in You et al., so no convention can be inferred from it. The LLM coding marks ingredients in both.
Q3 — the disjunctive-category question (card 3). When one dummy fires on conditions from two trait families and the data cannot show which fired, do both constructs count? The LLM coding records both, primary first, so either convention can be produced later without recoding.
Q4 — two straightforward errors (card 4). These need confirming, not deciding. If you agree, the remaining disagreement on these three papers is entirely Q1–Q3.

Bearing on Q1: on the three papers coded here, the set of constructs recorded per paper has been identical every time they were coded. Q1 does not change which constructs a paper counts toward — both sides record age, gender and race for You et al. It changes only which column, individual or board-level, records them. That makes it a presentation decision for the tables rather than a coding-accuracy one.