Exact instructions and definitions provided to the LLM are stated in the “LLM Prompt” section.
| Field | What it records |
|---|---|
| name_in_paper | The variable's label as the authors write it — the label from the variable-definitions table, or the row label in a regression table — with nothing appended by the coder. Table locations go in appears_in_tables, descriptions in operationalization. |
| role | The part the variable plays in the paper's main analysis: dv, iv, moderator, mediator, control, instrument. Other parts in other analyses go in notes. |
| scope | main (main specifications), robustness (robustness checks only), dropped (considered and excluded). Every variable in the empirical analyses gets a record — none skipped for seeming minor. |
| content | Which construct the authors intend the variable to measure, from the Content Vocabulary below. Identified from the paper's own framing — the variable's name, the hypothesis or argument it serves, the literature cited when introducing it — never from the measurement formula alone. Null when the paper provides no framing anywhere (no guessing). not_applicable for variables measuring no person or board trait (firm outcomes, industry, macro). Multi-valued (an ordered list) when one measure's definition spans more than one construct — a multi-trait faultline, or a disjunctive category that fires on conditions from different trait families. The primary construct is listed first, and a paper counts toward every construct in the list, not only the first. Multi-valued content applies only to variables with an empty derived_from: a composite takes the content its own framing supports, never the union of its constituents', which stays recoverable from the constituents' own records. |
| content_evidence | The framing evidence for the content choice: a short verbatim phrase plus location. Empty when content is null or not_applicable. |
| measurement_form | How the variable is measured over people, from the Form Vocabulary below. Decided by the unit the value varies over: differs across directors → individual; one value per board → a board_* form; a board feature that is not an aggregate of member traits → structural. Derived variables take a form only if one genuinely fits. |
| relative_to | The benchmark a relational measure is defined against: none, ceo, chair, other_directors, firm. The benchmark never changes the form, and benchmarking against the CEO or firm does not make the variable a CEO or firm variable. |
| level | Whose attribute the value is: director, board, ceo, firm, other. Person-trait contents apply at any person level — a CEO's age is (content: age, level: ceo). Board tables later filter to director/board. |
| derived_from | For variables built from other named measures (composite index, interaction term, ratio): the constituent variable names. Each constituent gets its own record; the derived record is coded in its own right (unframed derived variables take content = null like anything else). |
| operationalization | The precise definition of how the variable is constructed, as the authors define it. |
| transformations | Log, +1 offset, winsorization, reverse-coding, threshold dummy, standardization. |
| data_source | Where the variable's data comes from, as named by the authors. |
| measure_label | A short label: the operationalization phrase, plus the data source in parentheses where the paper names one. Full definitions stay in operationalization. |
| inclusion_quote (+ location) | Verbatim quote where the authors justify why this variable is in the analysis. Empty if none exists — never a paraphrase. |
| measurement_quote (+ location) | Verbatim quote justifying why the variable is measured this way. Empty if none. Quote policy for both: verbatim wording and word order; OCR spacing may be repaired; no word changed, added, or deleted. |
| appears_in_tables | The tables/panels where the variable appears. |
| notes | Ambiguities, secondary roles, suspected conversion artifacts, judgment calls. |
| ID | Definition | Boundary notes |
|---|---|---|
| age | A person's age: years, bins, retirement-age dummies. | |
| gender | Gender; female indicators; gender in matching or tie definitions. | Gender-based ties also code networks. |
| race_ethnicity | Race or ethnicity; minority indicators. | |
| nationality | The person's own nationality or foreign status. | International experience is expertise. |
| political_ideology | Political leanings, usually donation-based; ideological distance. | Office-holding is political_experience. |
| risk_preferences | Revealed risk preferences (e.g., option-exercise behavior). | |
| personality_psychology | Personality and cognition: narcissism, overconfidence, cognitive style. | Risk attitudes are risk_preferences. |
| experience_tenure | Experience as a director: tenure on the focal board, count of current/past board seats, first-time-director indicators, years served with the current CEO. | Seat counts framed as attention constraints are busyness. Executive work experience is expertise. |
| education | Degree level, institution prestige/elite lists, field of study, exam scores. | Education-based ties (same school) are networks. |
| expertise | Skills, functional background (finance, legal, accounting, media, medicine), industry experience, executive experience (including being a CEO elsewhere), country-specific/foreign experience. | Political office is political_experience; board seats as experience are experience_tenure. |
| political_experience | Held political office or worked in government; years in government. | Donation-based ideology is political_ideology. |
| busyness | Overcommitment: directorship-count thresholds, seat counts framed as attention constraints, outside employment. | Temporary shocks are distraction. |
| distraction | Temporary attention shocks (events at the person's other firms). | Chronic overload is busyness. |
| attendance | Meeting attendance records. | Count of meetings held is board_structure. |
| networks | Ties and network positions: interlocks, employment/education/social ties (including ties to the CEO or chair), centrality at any level. | |
| status | Standing in a prestige hierarchy, usually composite; directorship counts explicitly framed as status. | Negative standing is stigma_reputation. |
| stigma_reputation | Negative or at-risk public standing: fraud taint, lawsuit association, media reputation. | |
| independence | Independence classification (exchange/ISS definitions; inside/outside/affiliated). | |
| pay_ownership | Compensation; equity ownership; incentive structure. | Blockholder-scale stakes are blockholder. |
| blockholder | The person is, or represents, a large shareholder. | |
| founder_family | Founder or founding/controlling-family member. | |
| ceo_duality | The CEO also holds the board chair role, or the two roles are separated. | The chair role on its own, held by anyone, is board_chair. |
| board_chair | Holds the chair (or presiding chair) role of the board itself, as an attribute of a person. | Whether the CEO is the one holding it is ceo_duality; chairing a board committee is committee_chair. |
| committee_chair | Chairs a board committee. | |
| committee_member | Sits on a board committee; assignment used as a variable. | A paper using chair and membership variables codes both. |
| ceo_appointment | Director joined the board during the sitting CEO's tenure (co-opted directors, CEO allies). | “Currently a CEO elsewhere” is expertise. |
| election_outcome | Shareholder votes on the director: election/withhold results, vote shares. | |
| investor_appointed | Appointed by or representing an investor, typically an activist. | |
| lead_independent_director | Lead/presiding independent director role. | |
| board_structure | Structural board features that are not aggregates of member traits: staggered/classified board, board size, meeting frequency. | |
| not_applicable | The variable measures no person or board trait (firm characteristics and outcomes, industry and macro conditions). |
| ID | Definition |
|---|---|
| individual | Measured per person, including role dummies and dyadic measures that vary across directors. |
| board_mean_level | Board-level mean or level of a member trait. |
| board_share_count | Share or count of members with a trait; totals of member traits. |
| board_dispersion | Heterogeneity: SD, Blau/HHI, similarity indices, heterogeneity measures. |
| board_faultline | Faultline measures, including multi-trait ones (content is then multi-valued). |
| board_overlap | Pairwise overlap/similarity among members; uniqueness relative to the rest of the board (with relative_to = other_directors). |
| structural | Board features that are not aggregates of member traits. |
A variable can also have a benchmark, recorded separately in relative_to. The benchmark does not change the measurement form: average board tenure divided by CEO tenure is still a board mean — it is simply a board mean with relative_to = CEO.
| Block | Fields |
|---|---|
| paper | title, authors, year, journal; research_question (one sentence); theory_framing; samples (one entry per estimation sample: description, n, period); data_sources; empirical_methods; identification_strategy (with the authors' own caveats); measurement_innovation (reusable dictionaries, indices, coding protocols). |
| coding_notes | conversion_quality; ambiguities (judgment calls); possible_missing (content suspected lost in PDF conversion). |
Verbatim copy of the coding prompt (prompt body only; the coder sees raw markdown). {PAPER_MD_PATH} is filled in with the paper being coded. Every returned file is checked against the schema before it is accepted.
You are coding one academic paper for a structured literature database on board composition research.
Your final message must be a single JSON object and nothing else; it is parsed by a program. Do not
summarize the paper, evaluate its quality, or address a human reader.
The paper (a markdown conversion of the PDF) is at: `{PAPER_MD_PATH}`
## How to read the paper
- Read the ENTIRE file. Page through it with the Read tool (offset/limit) until you reach the end.
- If the paper has a variable-definitions appendix or table, read it FIRST and use it as the spine of
your inventory; then walk every regression table to catch variables the appendix omits; then read the
text for framing and justification quotes.
- Expect PDF-conversion artifacts: minus signs rendered as "2" (a table cell "20.30" may be −0.30), "×"
rendered as "3", "<" rendered as ",", dropped ligatures ("frm" for "firm"), collapsed whitespace, and
unreadable figures (captions survive). When you suspect an artifact, record the suspicion in
`coding_notes` instead of silently guessing.
- Never invent a variable, a number, or a quote. When you are unsure, say so in the `notes` field of the
affected record; uncertainty belongs in the data, not resolved by guessing.
## Which variables get a record
Create one record for every variable that appears in the paper's empirical analyses: dependent
variables, focal predictors, moderators, mediators, instruments, and controls. Set `role` to the part
the variable plays in the paper's main analysis; if it plays different parts in different analyses, use
the main analysis and note the others in `notes`. Set `scope` to `main` for variables in the main
specifications, `robustness` for variables appearing only in robustness checks, and `dropped` for
variables the authors considered but excluded. Do not skip a variable because it seems minor.
### What counts as one record
These two rules fix the boundaries of the inventory. Apply them literally.
- **Do not record estimation apparatus.** Fixed effects of any kind, intercepts, time dummies, offsets,
exposure terms, and weights are not variables. They measure no person or board trait and get no
record, however prominently the tables list them.
- **A variable and its transformation are one record.** Where variable X enters the analyses only in
transformed form — logged, winsorized, standardized, reverse-coded, or cut into a threshold dummy —
create one record, named as the authors label it, and put the transformation in `transformations`.
Create a second record only where the untransformed and transformed forms both enter as distinct
regressors, and say so in `notes`.
### Reconcile before you emit
Check your inventory against the paper, and fix any gap you find:
1. Every row of the variable-definitions appendix or table has a record.
2. Every row of every regression table has a record, or is estimation apparatus excluded above.
This check leaves no trace in your output; just complete it.
## Field guide
Every entry below is the rule for that field. Vocabulary values are defined once, in the tables after
this guide; nothing here redefines them.
**`name_in_paper`** — the variable's label as the authors write it: the label from the
variable-definitions table, or the row label in a regression table. Add nothing of your own devising —
not a description, not a table reference, not a clarifying parenthetical. Table locations belong in
`appears_in_tables` and descriptions in `operationalization`.
**`role`** — one of: `dv`, `iv`, `moderator`, `mediator`, `control`, `instrument`.
**`scope`** — one of: `main`, `robustness`, `dropped` (see "Which variables get a record").
**`content`** — which construct the authors intend the variable to measure, as a Content Vocabulary ID.
Identify the construct from the paper's own framing — the variable's name, the hypothesis or argument it
serves, and the literature cited when introducing it — never from the measurement formula alone. If the
paper provides no framing anywhere (typically a control that is listed but never discussed), set
`content` to null rather than guessing. Variables that measure no person or board trait — firm
characteristics and outcomes, industry and macro conditions — take `content` = `not_applicable`.
Some variables are defined DISJUNCTIVELY: one measure fires on any of several conditions that belong
to different constructs, and the data cannot show which condition fired for a given person.
Disclosure-based categories are the common case — a single dummy defined so that it fires when a
person satisfies either of two conditions drawn from different trait families, sometimes operationalized
with a keyword list spanning both. For these, `content` is a LIST naming every construct the definition
can fire on, with the PRIMARY construct FIRST. The primary construct is the one the authors' own framing
leads with; where the framing does not settle it, the one covering most of the definition. Do not
silently pick one condition and drop the others, and do not reorder the list alphabetically or by
vocabulary order; the order is meaningful, not cosmetic. A single-construct variable stays a plain
string, not a one-element list.
This applies ONLY to variables whose `derived_from` is empty. A variable built from other named
measures takes only the content its OWN framing supports, or null — never the union of its
constituents' contents; the constituents carry those in their own records. If you find yourself
listing three or more contents for one variable, check whether it is in fact built from constituents;
if it is, name them in `derived_from` and leave the extra contents off this record.
**`content_evidence`** — the framing evidence for your `content` choice: a short verbatim phrase or
sentence plus its location. Empty string when `content` is null or `not_applicable`.
**`measurement_form`** — how the variable is measured over people, as a Measurement Form ID. Decide by
the unit the value varies over: it differs across individual directors → `individual`; one value per
board → the matching `board_*` form; a board feature that is not an aggregate of member traits →
`structural`. For a derived variable (see `derived_from`), assign a form only if the derived variable
itself matches one; otherwise null, with a note.
**`relative_to`** — one of: `none`, `ceo`, `chair`, `other_directors`, `firm`. The benchmark a
relational measure is defined against. The benchmark never changes `measurement_form`, and using the CEO
or the firm as a benchmark does not make the variable a CEO or firm variable.
**`level`** — one of: `director`, `board`, `ceo`, `firm`, `other`. Content values describe traits of
people and boards; code every person-trait variable with its content regardless of whose trait it is —
`level` records whether the person is a director or the CEO.
**`derived_from`** — when the variable is built from other named measures (a composite index, an
interaction term, a ratio of two variables), list the constituent variable names here, and create a
record for each constituent as well. Code every record — including the derived one — in its own right
under this field guide; a derived variable with no framing of its own takes `content` = null like any
other unframed variable. Empty list otherwise.
**`operationalization`** — the precise definition of how the variable is constructed, as the authors
define it (include sample statistics only if the authors state them).
**`transformations`** — measurement transformations applied: log, +1 offset, winsorization,
reverse-coding, threshold dummy, standardization. Empty string if none.
**`data_source`** — where the variable's data comes from, as named by the authors.
**`measure_label`** — a short label of the form "[operationalization phrase] ([data source])", or just
the phrase where the paper names no source. Full definitions belong in `operationalization`, never in
the label.
**`inclusion_quote`, `inclusion_quote_location`** — verbatim quote where the authors justify WHY this
variable is in the analysis, and where it appears. Empty strings if no such justification exists — never
substitute a paraphrase.
**`measurement_quote`, `measurement_quote_location`** — verbatim quote where the authors justify why
the variable is measured THIS WAY, and where it appears. Empty strings if none.
Quote policy for both quote fields: verbatim in wording and word order; you may repair spacing lost in
PDF conversion; you may not change, add, or delete a word.
**`appears_in_tables`** — the tables/panels where the variable appears (e.g., "Table 2; Table 5 Panel A").
**`notes`** — ambiguities, secondary roles, suspected conversion artifacts, judgment calls.
### Paper block fields
`title`, `authors`, `year`, `journal`; `research_question` (one sentence); `theory_framing` (the
literature/theory motivating the paper, brief); `samples` (a list — one entry per estimation sample:
description, n, period); `data_sources`; `empirical_methods`; `identification_strategy` (the causal
strategy, or "descriptive/associational", with the authors' own caveats if any); `measurement_innovation`
(any reusable measurement contribution — a new dictionary, index, or coding protocol — or empty string).
Types in this block: every field above is a STRING except `samples`, which is a list of OBJECTS, one
per estimation sample, each with exactly the keys `description`, `n` and `period`. Do not write
`samples` as a list of sentences.
### Coding-notes block fields
`conversion_quality` (were tables and definitions readable?); `ambiguities` (judgment calls you made);
`possible_missing` (content you suspect exists in the PDF but was lost in conversion). All three are
strings.
## Content Vocabulary
Assign the ID whose definition matches the construct the paper's framing invokes. Boundary notes resolve
neighboring pairs. Multi-valued where one measure genuinely spans trait families — a faultline combining
demographic and skill traits, or a disjunctive category that fires on conditions from different families.
List the primary construct first.
| ID | Definition | Boundary notes |
|---|---|---|
| `age` | A person's age: years, bins, retirement-age dummies. | |
| `gender` | Gender; female indicators; gender in matching or tie definitions. | Gender-based ties also code `networks`. |
| `race_ethnicity` | Race or ethnicity; minority indicators. | |
| `nationality` | The person's own nationality or foreign status. | International *experience* is `expertise`. |
| `political_ideology` | Political leanings, usually donation-based; ideological distance. | Office-holding is `political_experience`. |
| `risk_preferences` | Revealed risk preferences (e.g., option-exercise behavior). | |
| `personality_psychology` | Personality and cognition: narcissism, overconfidence, cognitive style. | Risk attitudes are `risk_preferences`. |
| `experience_tenure` | Experience as a director: tenure on the focal board, count of current/past board seats, first-time-director indicators, years served with the current CEO. | Seat counts framed as attention constraints are `busyness`. Executive work experience is `expertise`. |
| `education` | Degree level, institution prestige/elite lists, field of study, exam scores. | Education-based *ties* (same school) are `networks`. |
| `expertise` | Skills, functional background (finance, legal, accounting, media, medicine), industry experience, executive experience (including being a CEO elsewhere), country-specific/foreign experience. | Political office is `political_experience`; board seats as experience are `experience_tenure`. |
| `political_experience` | Held political office or worked in government; years in government. | Donation-based ideology is `political_ideology`. |
| `busyness` | Overcommitment: directorship-count thresholds, seat counts framed as attention constraints, outside employment. | Temporary shocks are `distraction`. |
| `distraction` | Temporary attention shocks (events at the person's other firms). | Chronic overload is `busyness`. |
| `attendance` | Meeting attendance records. | Count of meetings held is `board_structure`. |
| `networks` | Ties and network positions: interlocks, employment/education/social ties (including ties to the CEO or chair), centrality at any level. | |
| `status` | Standing in a prestige hierarchy, usually composite; directorship counts explicitly framed as status. | Negative standing is `stigma_reputation`. |
| `stigma_reputation` | Negative or at-risk public standing: fraud taint, lawsuit association, media reputation. | |
| `independence` | Independence classification (exchange/ISS definitions; inside/outside/affiliated). | |
| `pay_ownership` | Compensation; equity ownership; incentive structure. | Blockholder-scale stakes are `blockholder`. |
| `blockholder` | The person is, or represents, a large shareholder. | |
| `founder_family` | Founder or founding/controlling-family member. | |
| `ceo_duality` | The CEO also holds the board chair role, or the two roles are separated. | The chair role on its own, held by anyone, is `board_chair`. |
| `board_chair` | Holds the chair (or presiding chair) role of the board itself, as an attribute of a person. | Whether the CEO is the one holding it is `ceo_duality`; chairing a board committee is `committee_chair`. |
| `committee_chair` | Chairs a board committee. | |
| `committee_member` | Sits on a board committee; assignment used as a variable. | A paper using chair and membership variables codes both. |
| `ceo_appointment` | Director joined the board during the sitting CEO's tenure (co-opted directors, CEO allies). | "Currently a CEO elsewhere" is `expertise`. |
| `election_outcome` | Shareholder votes on the director: election/withhold results, vote shares. | |
| `investor_appointed` | Appointed by or representing an investor, typically an activist. | |
| `lead_independent_director` | Lead/presiding independent director role. | |
| `board_structure` | Structural board features that are not aggregates of member traits: staggered/classified board, board size, meeting frequency. | |
| `not_applicable` | The variable measures no person or board trait (firm characteristics and outcomes, industry and macro conditions). | |
## Measurement Form Vocabulary
| ID | Definition |
|---|---|
| `individual` | Measured per person, including role dummies and dyadic measures that vary across directors. |
| `board_mean_level` | Board-level mean or level of a member trait. |
| `board_share_count` | Share or count of members with a trait; totals of member traits. |
| `board_dispersion` | Heterogeneity: SD, Blau/HHI, similarity indices, heterogeneity measures. |
| `board_faultline` | Faultline measures, including multi-trait ones (content is then multi-valued). |
| `board_overlap` | Pairwise overlap/similarity among members; uniqueness relative to the rest of the board (with `relative_to` = `other_directors`). |
| `structural` | Board features that are not aggregates of member traits. |
## Worked examples (schematic)
These examples use placeholder names (X, A, Z, D) and placeholder content tokens (`<CONTENT_C>` stands
for whichever Content Vocabulary ID the paper's framing supports). They demonstrate mechanics and format
only; they carry no information about which contents are common or likely.
**Example 1 — a framed focal variable.** The paper says: *"Our main independent variable is X, measured
as [definition] following Author (Year). We argue that X captures [construct C], and expect it to
increase Y because [mechanism]."*
```json
{
"name_in_paper": "X",
"role": "iv",
"scope": "main",
"content": "<CONTENT_C>",
"content_evidence": "\"We argue that X captures [construct C], and expect it to increase Y because [mechanism].\" (Hypotheses section, p. 7)",
"measurement_form": "individual",
"relative_to": "none",
"level": "director",
"derived_from": [],
"operationalization": "[definition as the authors state it]",
"transformations": "",
"data_source": "[source named by the authors]",
"measure_label": "[short phrase] ([source])",
"inclusion_quote": "\"We argue that X captures [construct C]...\"",
"inclusion_quote_location": "Hypotheses section, p. 7",
"measurement_quote": "\"measured as [definition] following Author (Year)\"",
"measurement_quote_location": "Data section, p. 12",
"appears_in_tables": "Table 3, Columns 1-4",
"notes": ""
}
```
**Example 2 — an unframed control.** The paper says only: *"Controls include A, B, and C."* and A is
never discussed again.
```json
{
"name_in_paper": "A",
"role": "control",
"scope": "main",
"content": null,
"content_evidence": "",
"measurement_form": "individual",
"relative_to": "none",
"level": "director",
"derived_from": [],
"operationalization": "[definition if given anywhere, e.g., a definitions appendix]",
"transformations": "",
"data_source": "[source if named]",
"measure_label": "[short phrase]",
"inclusion_quote": "",
"inclusion_quote_location": "",
"measurement_quote": "",
"measurement_quote_location": "",
"appears_in_tables": "Table 3, Columns 1-4",
"notes": "Listed among controls only; no framing anywhere in the paper."
}
```
**Example 3 — a derived variable and its constituents.** The paper interacts X with Z in its moderation
test. X and Z each get their own full record (as in Examples 1-2); the interaction gets:
```json
{
"name_in_paper": "X × Z",
"role": "moderator",
"scope": "main",
"content": null,
"content_evidence": "",
"measurement_form": "individual",
"relative_to": "none",
"level": "director",
"derived_from": ["X", "Z"],
"operationalization": "Product of X and Z.",
"transformations": "",
"data_source": "",
"measure_label": "X × Z",
"inclusion_quote": "\"Hypothesis 2. The relationship between X and Y is weaker when Z is high.\"",
"inclusion_quote_location": "Hypotheses section, p. 8",
"measurement_quote": "",
"measurement_quote_location": "",
"appears_in_tables": "Table 4",
"notes": "Interaction term; constituents X and Z recorded separately."
}
```
**Example 4 — a disjunctive category.** The paper defines a single dummy D as: *"D equals one if the
person satisfies condition P or condition Q"*, where P belongs to one trait family and Q to another,
and introduces D as a measure of P.
```json
{
"name_in_paper": "D",
"role": "iv",
"scope": "main",
"content": ["<CONTENT_P>", "<CONTENT_Q>"],
"content_evidence": "\"D equals one if the person satisfies condition P or condition Q.\" (Data section, p. 9)",
"measurement_form": "individual",
"relative_to": "none",
"level": "director",
"derived_from": [],
"operationalization": "Dummy equal to one if the person satisfies condition P or condition Q.",
"transformations": "",
"data_source": "[source if named]",
"measure_label": "D dummy ([source])",
"inclusion_quote": "",
"inclusion_quote_location": "",
"measurement_quote": "",
"measurement_quote_location": "",
"appears_in_tables": "Table 3",
"notes": "Disjunctive category: the definition fires on P or Q, which fall in different trait families. <CONTENT_P> is listed first as the primary construct because the authors introduce D as a measure of P."
}
```
## Output
Your final message is exactly one JSON object with this shape, and no other text:
```json
{
"paper": { "...paper block fields..." },
"variables": [ { "...one record per variable..." } ],
"coding_notes": { "...coding-notes block fields..." }
}
```
Three of the 60 records from Adams, Akyol & Verwijmeren (2018), “Director Skill Sets,” and the cells they produce. (Records abridged to the fields the mapping consumes.)
| Sheet column | Yes | Measure cell |
|---|---|---|
| Directorship Experience/Tenure | 1 | Tenure |
| Expertise Diversity/Levels | 1 | Skill Blau index |
| Staggered/Classified Board | 1 |
Each (content, form) pair looks up its column; the Measure cell concatenates the distinct measure labels of all records mapping to that column. The full row for this paper marks eleven columns and the hand-coded row marks eight. The differences are set out on the Disagreements screen.
Used for ground-truth comparison only; manuscript tables come from the records directly. Three columns are renamed relative to the spreadsheet so names reveal their level: Independence → Board Independence; Network Position → Board Network Position; Board Experience/Tenure → Directorship Experience/Tenure (a director-level category whose old name misread as board-level). Combinations not listed map to no column — new information the sheet could not distinguish. Counts (n) are how many of the 113 hand-coded articles mark the column.
| Sheet column | n | Content(s) | Form(s) |
|---|---|---|---|
| Age | 31 | age | individual |
| Gender | 31 | gender | individual |
| Nationality | 7 | nationality | individual |
| Race/Ethnicity | 14 | race_ethnicity | individual |
| Political Ideology | 7 | political_ideology | individual |
| Risk Aversion | 1 | risk_preferences | individual |
| Directorship Experience/Tenure | 40 | experience_tenure | individual |
| Education | 19 | education | individual |
| Expertise | 49 | expertise | individual |
| Political Experience | 6 | political_experience | individual |
| Busyness | 11 | busyness | individual |
| Distracted | 1 | distraction | individual |
| Meeting Attendance | 3 | attendance | individual |
| Networks/Social Capital | 32 | networks | individual |
| Status | 2 | status | any |
| Stigma/Reputation | 3 | stigma_reputation | individual |
| Blockholder | 3 | blockholder | individual |
| Chairman/CEO Duality | 17 | ceo_duality | individual, structural |
| Committee Chair | 8 | committee_chair | individual |
| Committee Member | 9 | committee_member | individual |
| Current CEO Appointment | 7 | ceo_appointment | individual |
| Election Outcome | 1 | election_outcome | individual |
| Founder/Family | 3 | founder_family | individual |
| Independent | 38 | independence | individual |
| Investor Appointed | 1 | investor_appointed | individual |
| Lead Independent Director (LID) | 5 | lead_independent_director | individual |
| Pay/Ownership | 11 | pay_ownership | individual |
| Ascriptive Diversity | 24 | age, gender, race_ethnicity, nationality | any board form |
| Experience Diversity/Differences | 10 | experience_tenure | any board form |
| Expertise Diversity/Levels | 17 | expertise, education, political_experience | any board form; also (expertise, individual + relative_to = other_directors) |
| Board Independence | 23 | independence | board_share_count, board_mean_level |
| Board Network Position | 3 | networks | any board form |
| Staggered/Classified Board | 5 | board_structure (staggered/classified) | structural |
Notes: status homogeneity (a board form of status) maps to Status per the hand coding's precedent. The Expertise Diversity/Levels extra clause covers director-level uniqueness measures relative to the rest of the board, which the hand coding files there. Multi-trait faultlines mark every diversity column their contents touch.
One row per construct: how much of the literature uses it, how it is operationalized, and why authors say they use it.
| Construct | Papers using it | Top operationalizations | Common justifications |
|---|---|---|---|
| Counts from the records (papers with ≥1 qualifying record per content); operationalizations ranked by frequency of normalized measure labels, cut where usage share drops sharply; justifications synthesized by clustering the harvested inclusion/measurement quotes into recurring rationales, each backed by citable verbatim examples. | |||
Detailed comparison for the constructs nearest our framework: expertise, experience, and networks. One row per variable variety: the varieties of measures, operationalization detail, and justification.
How the literature aggregates member traits: contents × measurement forms (individual vs mean vs share vs dispersion vs faultline vs overlap). Feeds the theory discussion on levels-vs-diversity measurement choices.
The following presents cases where the LLM coding disagrees with the human coding in the 3 sample papers we coded using the LLM prompt. In most of these the human and LLM coding record the same construct and differ over the unit it should be recorded at (director or board). We need to review each of these disagreements to determine if the LLM prompt should be revised.
The paper's control list names eleven board and director characteristics, and every one of them is an aggregate: a proportion, a heterogeneity index, or a similarity composite. Director age, gender and race never appear as per-director variables. They enter in one other place only, as inputs to the algorithm that sorts directors into identity subgroups.
One control shows the whole disagreement in a single row. For “the proportion of directors with CEO experience”, the human coding marked Expertise, a director-level column, with the Measure cell “Executive (CEO)”. The LLM coding recorded content expertise at form board_share_count, which the mapping sends to Expertise Diversity/Levels, a board-level column. Same variable, same construct, two different columns.
Across the sentence as a whole the two sides agree on all four board-level columns: Ascriptive Diversity, Experience Diversity/Differences, Expertise Diversity/Levels and Board Independence. The human coding's Measure cell for each of the three diversity columns reads “Faultine”, so both sides were looking at the same faultline measure. The disagreement is that the human coding also marked five director-level columns for those same traits, recording them twice; the LLM coding recorded them once, at the level the paper estimates.
One of the five is not a disagreement about level at all. “The proportion of directors appointed after the CEO” is a board-level share, but the sheet's only category for CEO appointment is director-level, so the mapping has nowhere to put it. The LLM coding did record the construct; the comparison simply cannot express it. The same is true of board-level job-title heterogeneity.
“We controlled for 11 variables on board and director characteristics related to CEO dismissal or board subgroups (Westphal & Zajac, 1995): board size (the total number of directors), the proportion of independent directors, CEO-board similarity (a composite measure of age, gender, and race dissimilarities), board diversity (the proportion of female directors, director racial heterogeneity, director age heterogeneity using the coefficient of variation, the proportion of directors with CEO experience, the proportion of directors appointed after the CEO, director tenure heterogeneity, and director job title heterogeneity), faultline strength…” Data and Methods, p. 2823 — every board/director control in the paper, and every one of them is an aggregate
“we first calculated faultline strength using the average silhouette width (ASW) algorithm (Meyer & Glenz, 2013) and identified directors' membership of subgroups in each board based on five [director characteristics]” Data and Methods, p. 2822 — where director age, gender and race actually enter
Both papers build a headline variable out of named ingredients, and the human coding treats the two cases oppositely. That inconsistency is why this needs a decision rather than a ruling.
Acharya & Pollock. A director counts as “globally high status” if he or she holds any one of three credentials: a degree from an elite school, a vice-president-or-above post at an S&P 500 firm, or an outside directorship at an S&P 500 firm. That yes/no flag is then turned into a single board-level Blau index, and it is the index that enters the regressions. No model contains a director's education or occupation as a variable in its own right. The human coding marked all three credentials, and its Measure cells name them exactly: “Prestige”, “Occupation Employment prestige”, “Directorial prestige”. The LLM coding recorded the variable the paper estimates — one board-level record carrying all three constructs at once.
You et al. CEO subgroup power is built from three named indicators of an individual director's power: the director's title (chair of the board or of a committee), their number of external directorships, and their share ownership. Here the human coding marked none of the three, and did not mark the composite under Status either. The LLM coding recorded each named indicator under its own construct.
Two wrinkles sit inside the You et al. side. The count of external directorships is introduced only as an indicator of power, and the coding vocabulary has no entry for power; a directorship count counts as status only where the authors explicitly frame it as prestige, which they do not, which leaves that cell genuinely ambiguous. And “director title” runs together chairing the board with chairing a committee, two different roles, of which only the committee half has a column in the sheet.
Note that the Acharya & Pollock cells turn on the aggregation question in card 1 as well: the credentials are director-level, the estimated index is board-level.
“we treated a director as globally high status if he or she possessed at least one of the following credentials: a degree from an elite educational institution listed in Finkelstein's (1992) list of prestigious institutions (educational prestige), experience as an executive at the level of vice president or above at an S&P 500 company (employment prestige), or experience as an outside director for an S&P 500 firm (directorial prestige).” Acharya & Pollock, Data and Methods, p. 1481
“we computed individual directors' power using three indicators (Zajac & Westphal, 1996): director title (coded as one for the chair of the board or a board committee and zero otherwise), the number of external directorships, and director ownership (the number of outstanding shares).”You et al., Data and Methods, p. 2822
The paper codes twenty director skills from proxy statements, each a yes/no dummy driven by a keyword list. The human coding treated the taxonomy as one measure and marked Expertise, with the Measure cell “Skills (Proxies)”. The LLM coding made one record per skill category, which agrees on Expertise but also picks up two categories whose definitions straddle a boundary.
“Academic” fires if a director is from academia or holds a higher degree. The first is an occupation, the second a credential, and the keyword list mixes them: dean, faculty and professor sit alongside doctorate, masters and Ph.D. Nothing in the data says which condition fired for a given director. “Government and policy” fires on any of four keywords — government, policy, politics, regulatory — so a director flagged on “regulatory” alone may have held no public office at all, yet the coding vocabulary deliberately separates political office from general expertise.
The LLM coding therefore records both constructs for these two dummies, primary one first. The question is whether a measure that genuinely conflates two things should count toward both, or only toward the one the authors lead with.
Table 2, Skill categories (p. 645) |Academic |The director is from academia or has a higher degree (such as a Ph.D.).| |Government and policy|The director has governmental, policy, or regulatory experience. | Appendix B.1, Skill dictionary (p. 660) |Academic |Academia, academic, dean, doctorate, education, faculty, graduate, | | masters, Ph.D., professor, school environment | |Government and policy|Government, policy, politics, regulatory |
Board independence, Adams et al. The paper defines it in the variable appendix as the ratio of independent directors to board size, reports it in the descriptive statistics table, and carries it as a control in the main performance regressions (Table 5, columns 2, 4, 6 and 7, and Table 6 Panel C). The LLM coding recorded it; the human coding left the column blank. The same hand-coded row does mark the director-level Independent column for this paper, and that variable is also only a control, so controls were not being excluded as a class. The blank looks like an oversight rather than a convention.
Networks, Acharya & Pollock. The human coding marked Networks/Social Capital with a blank Measure cell. The paper has no tie, interlock or centrality variable anywhere: its variables are tenure overlap, expertise overlap, status homogeneity, and firm-level controls. The one candidate is prior outside directorships at S&P 500 firms, but that is not a standalone variable — it is one of the three status credentials in card 2, which the paper labels directorial prestige and which the coding vocabulary sends to Status, a column both sides already mark. Nothing supports a separate networks mark, and the blank Measure cell is consistent with a stray tick.
Adams et al., Appendix A, Variable definitions (p. 659) |Board independence|The ratio of independent directors on the board to the board size.| Adams et al., Table 1, Panel A: Firm characteristics (p. 644) |Variables | Mean | Median| Std. dev.| Min. | Max. | |Board independence| 0.795| 0.818 | 0.105 | 0.333| 0.941|
Bearing on Q1: on the three papers coded here, the set of constructs recorded per paper has been identical every time they were coded. Q1 does not change which constructs a paper counts toward — both sides record age, gender and race for You et al. It changes only which column, individual or board-level, records them. That makes it a presentation decision for the tables rather than a coding-accuracy one.