Jeff Girard Measurement & evaluation me@jmgirard.com

Data

6 corpora and evaluation sets, 2014–2026. Each one says what it costs to get and what my hand in it was, because co-authored and led are not the same claim.

OPEN
Open. Downloadable now, no request needed. 1
GATED
Click-through. Public, behind terms you accept at download. 1
ON REQUEST
On request. Released to researchers after an access agreement. 4

  1. 2026 Benchmark

    FriendBench

    GATED

    Thin-slice social perception from dyadic interaction

    Can a model tell whether two people have met before from a 20-second clip of them talking? Every dyad answers the same ice-breaker prompt, so the content cannot give it away — the signal is in the manner. Ships the answer key, ~90 crowd raters per modality, and zero-shot predictions from 26 models.

    Baselines Chance is 50%. The best model reaches 66.7% on audio and on video; the human crowd reaches 71.9% on video. Text is near chance for everyone.

    Scale
    96 dyads
    Modalities
    Text · Audio · Video
    My role
    Led
    Released with
    Fluid Concepts Research
    License
    CC BY-NC 4.0

    Hugging Face MINT 2026 (in press)

  2. 2025 Dataset

    Dyadic Audiovisual Motion Modeling and Large-Scale Dataset

    Face-to-face interaction at a scale nothing else in this area approaches: over 4,000 hours of dyadic conversation across improvised and scripted contexts, with transcripts and extracted body and face motion. Built to train models that generate and understand how two people behave toward each other, not how one person emotes at a camera.

    Scale
    4,000+ hours · 4,000+ participants
    Modalities
    Video · Audio · Transcripts · Body & face motion
    My role
    Co-authored
    Released with
    Meta FAIR
    License
    CC BY-NC 4.0

    Download Agrawal et al. (2025)

  3. 2023 Dataset

    DynAMoS

    ON REQUEST

    Dynamic Affective Movie Clip Database for Subjectivity Analysis

    Affective movie clips selected to provoke disagreement rather than consensus, with metadata and emotion ratings from 83 participants. Built for work on subjectivity: when raters differ, that difference is the measurement, not noise in it.

    Scale
    22 clips · 83 raters
    Modalities
    Video · Ratings
    My role
    Led

    Request access Girard, Tie & Liebenthal (2023)

  4. 2017 Dataset

    GFT

    ON REQUEST

    Sayette Group Formation Task Spontaneous Facial Expression Database

    Spontaneous behavior from unscripted three-person social interactions, with frame-level Facial Action Coding System annotation. Group conversation, not posed expression in front of a camera.

    Scale
    96 participants
    Modalities
    Video · FACS
    My role
    Led

    Request access Girard et al. (2017)

  5. 2016 Dataset

    MMSE / BP4D+

    ON REQUEST

    Multimodal Spontaneous Emotion Corpus for Human Behavior Analysis

    An extension of BP4D adding thermal imaging and physiological recording to the 3D video, with frame-level FACS annotation — so a facial signal can be checked against what the body was doing underneath it.

    Scale
    140 participants
    Modalities
    3D · Thermal · Physiology · FACS
    My role
    Co-authored

    Request access Zhang et al. (2016)

  6. 2014 Dataset

    BP4D

    ON REQUEST

    Binghamton–Pittsburgh 4D Spontaneous Emotion Database

    Spontaneous facial behavior during emotion elicitation, captured as high-resolution 3D dynamic video and annotated frame by frame with FACS. Still one of the standard training corpora for automated action unit detection.

    Scale
    41 participants
    Modalities
    3D · Video · FACS
    My role
    Co-authored

    Request access Zhang et al. (2014)

Building one of these is most of the work in evaluating a model honestly, and it is what I am usually hired for. If you need a benchmark that measures what you actually claim to measure, say so.