From digitized GI Bill records that reveal who was (and who wasn’t) helped to buy a home after World War II, to social media archives that help researchers study harassment, misinformation and political discourse, U-M is a national powerhouse for social science data. With support from MIDAS, our researchers are building the infrastructure, standards and training that turn messy data into tools for advancing democracy and evidence based policy.

The index card is yellowed with age, its typed entries faded but still legible. It records a mortgage guarantee issued by the Veterans Administration in 1947, including the loan number, property location, veteran’s name and amount borrowed. One card among thousands filling boxes at the National Archives, it is a remnant of the postwar housing boom when millions of returning World War II veterans bought homes with federally backed loans.
For decades, those cards sat underutilized. Historians could not analyze them systematically without investing enormous labor. The records were too voluminous to transcribe by hand, too inconsistent in format for simple scanning and too inaccessible for most researchers.
That changed when a team of University of Michigan researchers received MIDAS funding for a project called “Images to Integrated Data.” They developed machine learning tools to digitize and parse those index cards, converting tens of thousands mortgage guarantees from 1946 to 1954 into clean, analyzable data.
What emerged was the first public dataset documenting individual GI Bill mortgage guarantees, and it told a story that supported long-standing anecdotal accounts from veterans. The data showed exactly where loans went, who received them and who did not. When coded onto maps, the pattern became stark. Black veterans were systematically excluded from federally backed mortgages in many regions, perpetuating housing segregation that was widespread at the time.
“These records sat in boxes for decades,” said J. Trent Alexander, a researcher at the Inter-university Consortium for Political and Social Research (ICPSR) who led the digitization project. “Now that they’re digitized and linkable to other data, policymakers and communities have the potential to explore the role federal programs can play in creating or stifling economic opportunity and equality.”
That transformation from isolated index cards to usable evidence illustrates the value of social science data infrastructure. It is not glamorous work, but without it, many critical social science questions cannot be answered.
Why data infrastructure matters
Good policy requires good data. That is true whether the question is how to fund education, reform veterans’ benefits, protect voting rights, address housing discrimination or counter online extremism.
But useful data does not simply exist. Someone has to collect them, clean them, document what they mean, secure them against misuse and make them accessible to researchers who can analyze them rigorously. Such work often happens sporadically, driven by individual research projects or agency mandates. Datasets lived in scattered archives with inconsistent documentation and high barriers to access. Researchers often spent years just getting permission to use data, then more years cleaning and organizing them before any analysis could begin.
The result was enormous inefficiency and duplicate effort. Studies were difficult to reproduce because documentation was incomplete. Important research questions went unasked because the costs of assembling the necessary data were too high.
“Social science has always been data-intensive, but the infrastructure for managing that data lagged far behind what researchers actually needed,” said Margaret Levenstein, director of ICPSR. “We had surveys and administrative records, but we didn’t have the systems to make them findable, accessible, interoperable and reusable, what we now call FAIR principles.”
The challenge has grown more complex as new data sources have emerged. Researchers now also work with social media posts, mobile sensor data, digitized historical documents and other forms of digital trace information that reveal how people behave, not just what they report in surveys. These sources create powerful opportunities to understand social phenomena at unprecedented scale and resolution. They also raise difficult questions about privacy, consent, potential misuse and who owns such data.
Building infrastructure that addresses those challenges responsibly, protecting privacy while enabling research, documenting provenance and limitations and establishing ethical guidelines, requires sustained institutional investment.
The Michigan advantage
The University of Michigan has been a leader in social science data for decades, with ICPSR being an international leader that has archived and distributed social science data since 1962. ICPSR data catalog contains studies that are associated with more than 80,000 datasets covering topics from election studies and health surveys to international development indicators, and it provides data to researchers at 800 member institutions worldwide.
Even with its long history and deep expertise, ICPSR needed to modernize for an era of big data, digital trace information and AI powered analysis methods.
That is where MIDAS support came in. “We recognized that traditional surveys were being strained by societal changes, including declining response rates, higher costs and limited ability to capture fast-moving phenomena like social media dynamics,” said Dr. H. V. Jagadish, MIDAS director from 2019 to 2025. “At the same time, new data sources were emerging that could complement or even replace surveys for some purposes. Using those sources responsibly required new infrastructure, new methods and new norms.”
With MIDAS as a key collaborator, ICPSR secured $38M of funding from the National Science Foundation for its Research Data Ecosystem, to modernize how social science data are curated, discovered, accessed and analyzed.
The Research Data Ecosystem includes tools such as Explore Data, which allows researchers to browse thousands of datasets to identify relevant variables; Researcher Passport, which streamlines secure access to restricted-use data; and TurboCurator, which uses AI to suggest standardized metadata and keywords that make datasets easier to find and use.

“These aren’t just technical improvements,” Levenstein said. “They fundamentally change what’s possible. When researchers can quickly find and access the data they need, when datasets are well documented and follow common standards and when the infrastructure handles security and privacy protections automatically, researchers can focus on the questions that matter rather than logistics.”
Archiving Social Media
Another major infrastructure effort emerged as social media data became essential for understanding public discourse, misinformation, political mobilization and online harms. Most researchers, however, lacked the technical capacity or resources to collect and manage that data themselves.
Libby Hemphill, a professor at the School of Information, received a MIDAS pilot grant in 2021 to develop standards and infrastructure for ethical social media research. The project, “Ensuring FAIRness in Social Media Archives,” examined how to build archives that provide reliable, equitable access to social media data while protecting privacy and preventing misuse.
That work led to a partnership with Meta and ICPSR to launch SOMAR, the Social Media Archive. With $1.3 million in initial support, SOMAR is building a centralized repository for curated social media research datasets, complete with documentation, ethical guidelines and access controls.
The archive includes datasets such as #MeToo tweet IDs, congressional messaging collections and a white supremacist speech corpus combining Stormfront posts and Reddit comments.
These resources support research on content moderation, platform policies, political communication and the spread of extremist ideology.
“Social media platforms generate enormous amounts of data about public discourse, but that data is hard to access, unstable over time and often disappears when platforms change their APIs or policies,” Hemphill said. “SOMAR is working to create a stable archive of these data, democratizing access to powerful but sensitive digital trace data.”

One of 25,744 index card records from the administration of G.I. Bill Mortgages from 1946 to 1954, housed at the National Archives in College Park, Maryland.