Name File Type Size Last Modified
DATA_DICTIONARY.csv text/csv 4.7 KB 05/12/2026 05:46:PM
LICENSE.txt text/plain 346 bytes 05/12/2026 05:46:PM
METHODOLOGY.md text/x-web-markdown 3.4 KB 05/12/2026 05:46:PM
Methodology_and_Software_Context.pdf application/pdf 45.9 KB 05/18/2026 06:55:PM
PROVENANCE_AND_RIGHTS.md text/x-web-markdown 1.8 KB 05/12/2026 05:46:PM
README.md text/x-web-markdown 3.4 KB 05/12/2026 06:22:PM
assessment_panel.csv text/csv 5.6 KB 05/18/2026 06:46:PM
language_pair_summary.csv text/csv 1.6 KB 05/12/2026 05:46:PM
localization_variants.csv text/csv 72.5 KB 05/12/2026 05:46:PM
preference_ratings.csv text/csv 366 KB 05/12/2026 05:46:PM

Project Citation: 

regi, imaz. Pilot Dataset for Multilingual Video Localization Preferences (Synthetic Benchmark Package). Ann Arbor, MI: Inter-university Consortium for Political and Social Research [distributor], 2026-05-18. https://doi.org/10.3886/E248443V1

Project Description

Summary:  View help for Summary
Project Description / Scope of Project
This project provides a repository-ready synthetic pilot benchmark focused on multilingual video localization preferences. The deposit is designed for communication, digital media, platform localization, creator workflow, and accessibility-oriented research settings where investigators need a structured pilot dataset without redistributing third-party media assets.
The package includes:
- scenario-based source video metadata
- subtitle, dubbing, and lip-sync localization variants
- a reproducible synthetic trained-assessor panel
- synthetic rating events measuring readability, naturalness, sync, trust, and reuse intent
- a variable dictionary, methods document, rights statement, and file manifest
The deposit does not contain source videos, source audio, subtitle files copied from public platforms, cloned voices, or personal data. All distributed materials are original package files created for this deposit.
Neutral workflow context statement
The benchmark structure reflects common AI video localization workflows used in creator, education, product marketing, and cross-border communication settings. These workflows typically include source-language transcription, subtitle translation, target-language dubbing, timing alignment, and lip-sync adjusted delivery for short-form and long-form online video. A representative software context for these workflow patterns is described at https://aitranslatevideo.org/ where multilingual subtitle translation, dubbing, and lip-sync processing are presented as integrated steps in a video localization pipeline. In this deposit, the URL is provided only as contextual software background for the benchmark design and not as a distributed data source.

Scope of Project

Time Period(s):  View help for Time Period(s) 5/19/2026 – 5/31/2026 (fall 2026)


Related Publications

Published Versions

Export Metadata

Report a Problem

Found a serious problem with the data, such as disclosure risk or copyrighted content? Let us know.

This material is distributed exactly as received from the data depositor. As of April 2026, depositors are required to submit study materials in accessible formats. ICPSR has not reviewed, checked, or processed this material. For additional information about the study, please contact the investigator(s) directly. If you have questions about the accessibility of materials distributed by ICPSR or require further assistance, please visit ICPSR's Accessibility Center.