Task gender-typing in humans: survey draft

An interactive draft of the survey for review. Read the overview, then test each section.

How gender-typed are everyday performance tasks?

A draft instrument for review. This page is a working mock-up: every task runs exactly as a participant would see it. Use the buttons on the right to read this overview, then try each section yourself.

What this study does

This is a standalone study with two goals. First, it measures how gender-typed a range of cognitive and physical tasks are perceived to be: for each task, do people think men or women tend to do better, and in what proportion. Second, for a few of those tasks it measures actual performance and the gender gap in performance. Putting the two together lets us ask where perceptions of gender-typing match real performance differences and where they diverge.

Why we are running it

The main ambition project uses incentivized search tasks (a word search and a number search) to study goal-setting and re-attempt behavior by gender. A natural question is whether those tasks, and related ones, are perceived as gender-typed, and whether any perceived typing lines up with actual performance gaps. Mental rotation is included as a task with a well-documented male advantage, and a number-grid counting task is included as a candidate that leaned slightly male in the synthetic data.

The survey, in order

A participant moves through these sections:

  1. Consent and introduction. The UT consent form with a short description of the study, how long it takes, and how the bonus works.
  2. Agreement to work without help. A brief pledge to complete the tasks using only one's own effort.
  3. Performance tasks. For each task: instructions with a worked example, then three short timed attempts with a new board each time. A quick comprehension check follows the first set of instructions. The best of the three attempts determines the bonus.
  4. Perception questions. Rate 17 tasks on two questions each: the share of the top 10% who are men (0 to 100), and who performs better on average (a 1 to 7 scale). Each task is shown with a small image of itself. Task order and question order are randomized.
  5. Demographic questions. Including the participant's own gender. These come last, so the performance tasks come before any questions about gender. A short attention check sits here.
  6. Debrief and payment. A closing note and how the bonus is paid.

We can randomize the order of these sections if we want.

The performance tasks

  • ZIP-code search: find five-digit codes hidden in a grid of digits.
  • Word search: find five-letter words in a grid (the anchor from the main study).
  • Number-grid counting: count how many numbers in a grid exceed a stated value.
  • Mental rotation: decide whether two shapes are the same shape rotated, or a mirror image (male-advantaged).
  • Digit-symbol coding: match symbols to numbers using a key, as fast as possible (female-advantaged).
  • Verbal fluency: type as many words as possible that start with a given letter (female-advantaged).
  • Finger tapping: tap a key as many times as possible in a few seconds (male-advantaged).

This is the full set assembled for review. The ability-typed tasks (mental rotation, digit-symbol coding, verbal fluency, finger tapping) give real performance gaps on both the male and female side to compare against perception; a real run would trim to about four or five tasks to fit the time.

Participants and incentives

  • Sample: about 250 participants on Prolific, balanced by sex (about 125 men and 125 women), on desktop or laptop only, so screen size and controls are the same for everyone.
  • Length and pay: about 20 minutes, paid at roughly $12 per hour, with a budget near $1,500.
  • Bonus: on each task, the top 10% of participants (by their best of three attempts) earn a flat $2 bonus, set after data collection. Only the performance tasks carry a bonus.

Other design choices

  • Mental rotation is scored as the number correct minus the number wrong, so accuracy matters more than speed.
  • Each perception question shows a small image of the task itself, drawn with objects and shapes.

What we expect to learn

  • A ranking of tasks by perceived gender-typing, with confidence intervals.
  • The gender gap in actual performance on the seven timed tasks participants complete.
  • For each of those tasks, the gap between perceived typing and the actual performance difference (the perceived-versus-actual map).

How to test this page

Use the buttons on the right to run the performance tasks and the perception module exactly as a participant will see them. Each task runs live with its real timer and scoring, and the panel on the right shows the data the survey records. Reload the page to start a section over.

The buttons cover the two hands-on parts of the survey. The remaining sections in the list above (consent, the AI-free pledge, the checks, demographics, and the debrief) are standard survey pages.

Background: how tasks become gender-typed

Whether a task is seen as more suited to men or to women has two roots in the research literature. One is about traits: tasks that call for assertive, competitive, achievement-oriented behavior read as masculine, and tasks that call for warmth, nurturing, and social sensitivity read as feminine (Heilman 2012; Eagly and Karau 2002). The other is about measured ability: on average men score higher on spatial and mechanical tasks, women score higher on verbal fluency, perceptual speed, and reading emotions, and most of these differences are small (Hyde 2005; Miller and Halpern 2014). People's beliefs about who performs better track these real differences closely and tend to exaggerate them (Bordalo et al. 2019).

Table 1 places the broad domains on that spectrum, drawing on major works in the literature, only some of which are experimental tasks. Table 2 turns to the specific tasks that experiments have actually used, with how each has been typed and a source for every claim, and notes which paradigms are used most.

Table 1. Gender-typed domains (broad literature).
DomainTyped towardBasisSources
Leadership, competition, assertiveness (agentic)MaleTrait stereotypeHeilman 2012; Eagly and Karau 2002
Physical strength, heavy liftingMalePhysical, agenticextreme anchor (Heilman 2012)
Spatial ability, mental rotationMaleMeasured ability (largest gap)Voyer et al. 1995; Linn and Petersen 1985
Reaction time, finger tappingMaleMeasured motor speedRoivainen 2011
MathematicsPerceived male; actual gap near zeroPerception vs. abilityBordalo et al. 2019 (perceived); Hyde et al. 1990, Lindberg et al. 2010 (actual)
Verbal ability, fluency, memoryFemaleMeasured abilityHyde and Linn 1988; Hirnstein et al. 2022
Perceptual speed, coding and clericalFemaleMeasured abilityRoivainen 2011; Feingold 1988
Emotion recognition, social sensitivityFemaleMeasured ability and traitMcClure 2000; Hall 1978
Nurturing, caregiving, supportiveness (communal)FemaleTrait stereotypeHeilman 2012; Eagly and Karau 2002
Fine motor dexterityFemaleMeasured abilityNicholson and Kimura 1996

Most gender differences are small overall (Hyde 2005), and perceived task typing has weakened over time (Cejka and Eagly 1999; Beggs and Doolittle 1993). Tasks can also be rated directly on a masculine-to-feminine scale (Shinar 1975).

Table 2. Tasks common in experimental economics and psychology, plus the tasks in our study, ordered from male-advantaged to female-advantaged by measured performance. "Perceived as" is the stereotype about the task; "Who actually performs better" is the measured gender difference. A dash means the source we found reports one measure but not the other, which is common: filling the perceived-typing gap across a shared task set is what our study adds. Quotes are verbatim from the cited source.
TaskPerceived asWho actually performs betterEvidence (quoted from the source)
Mental rotationMen, by a wide marginThe largest spatial sex difference, favoring men (meta-analysis of 286 effect sizes, Voyer, Voyer and Bryden 1995)
Raven's pattern matricesMen, slightly (adults)"males obtained a higher mean than females by between .22d and .33d" (Irwing and Lynn 2005)
Reaction time and finger tappingMen (faster)"males are faster on reaction time tests and finger tapping" (Roivainen 2011)
Mazes (spatial navigation)MaleAbout equal at baseline; men pull ahead under competition"the average performance of men was 11.2 mazes, whereas it was 9.7 for women" (Gneezy, Niederle and Rustichini 2003, in Niederle and Vesterlund 2011)
Adding two-digit numbersMale (math stereotype)No gap; women slightly better at computation"there are no gender differences in ability on easy math tests" (Niederle and Vesterlund 2007); "women are significantly better at the math task" (Kamas and Preston, in Niederle and Vesterlund 2011)
Ball-into-bucket (physical, field)Depends on culture"the gender gap in tournament entry reverses in the matrilineal society" (Gneezy, Leonard and List 2009, in Niederle and Vesterlund 2011)
Counting numbers in a gridNot reportedUsed as a tedious, skill-neutral effort task; no gender difference reported (Abeler et al. 2011)
Word/number encryption (decoding)Not reportedSkill-neutral real-effort task; no gender analysis reported (Erkal et al. 2011)
Slider taskNeutral by design"does not require or test pre-existing knowledge ... identical across repetitions" (Gill and Prowse 2012)
Number-string / ZIP-code searchMixed (paradigm-dependent)Women lead on perceptual speed (Roivainen 2011); men lead on visual search (Stoet 2011)
Stroop (color-word) taskWomen, slightly"women outperform men in the classic Stroop task" (d = 0.12) (Sjoberg et al. 2022)
Digit-symbol codingWomen (perceptual speed)"females outperform males on ... the Coding and Symbol Search subtests" of the WAIS (Roivainen, Suokas and Saari 2021)
Word search and anagramsFemaleWomen equal or slightly better"everyone believes that men are better at the math task and women at the word task. In actuality, women are significantly better at the math task and are slightly but not significantly better at the word task" (Kamas and Preston, in Niederle and Vesterlund 2011)
Verbal generation (fluency)FemaleNo gap; women predicted better"In the verbal task, however, there are no gender differences in performance under either incentive scheme" (Shurchkov 2012, in Niederle and Vesterlund 2011); phonemic fluency favors women (Hirnstein et al. 2022)

The pattern to notice across the two middle columns: tasks perceived as male often show no real performance gap. Adding numbers is seen as male yet has no ability gap (women are, if anything, better at computation). The gap that does appear on that task is in willingness to compete, which is largest for the adding-numbers task and negligible for verbal and field tasks (Markowsky and Beblo 2022).

Sections

Recorded data

runs when you test a section