With support from the eScience Institute and the UW Data Science Minor, the Humanities Data Science Summer Institute (HDSSI) entered its fourth year with a new twist. This summer, the program’s humanities research teams joined forces with social science researchers to expand the institute’s reach across campus. Now, HDSSI exists alongside the Social Science Data Science Summer Institute (SSDSSI). This interdisciplinary 5-week summer program brought together undergraduate students, graduate students, faculty, and staff to team up on research projects using data science methods.
This research and the collaborations that arise from it do not end when the summer is over. Previous institute projects regularly result in publications such as this recent article in the Journal of Cultural Analytics. Many of this year’s research teams are either submitting their work to conferences or have already had it accepted. Check out this year’s project descriptions below to learn more about how data-intensive approaches were used at the institute to understand social and cultural phenomena.

HDSSI Projects
Understanding Genre Among UW Student Writing Assignments
Faculty PIs: Stephanie Kerschbaum (English), Nancy Bou-Ayash (English)
Grad RA: Ashfaq Ahmed
Undergrad Fellows: Aditya Gajula, Suji Jang, Eliza Pinckney, Landon Saino
Fellows in this group worked with faculty from the English Department to analyze the corpus of writing portfolios submitted by students in the UW’s Program in Writing and Rhetoric. The content of these portfolios varies greatly from student to student, which brings with it a rich opportunity to look for trends in how undergraduates are learning to communicate. Over the summer, the group did a deep dive into these portfolios to understand how the concept of “genre” can be meaningfully represented in this context. Working together they devised a SQL database of student submissions with a web frontend that allowed faculty to query and collect statistics on the works. Moving forward, the program hopes to grow this database and use it to understand how students are choosing to express themselves.
Unifying Multimodal Collections Held by the American Institute of Physics
Faculty PI: Ben Lee (Information School)
Grad RA: Hannah Sun
Undergrad Fellows: Christian Alviz, Priya Devanesan, Phoebe Norton, Rebekah Song
Fellows in this group delved into the archives of the American Institute of Physics, a corpus of over 5000 images from the Emilio Segrè Visual Archives collection and 1754 interview transcripts from the Niels Bohr Library & Archives. In working with these records they hoped to better understand what types of images are represented in the AIP’s visual archives, and more generally, how computational humanities methods can support the analysis of large-scale archival collections. Over the summer term, the team experimented with using the DBSCAN clustering algorithm and named entity recognition to map who was represented in the archive and what the relationships between them. What they found offered a glimpse at, among other things, the way that wives of interviewed physicists were vital to their scientific work, but were often left unnamed in the archive.
The New Age: Extracting Poetry from Early 20th Century Periodicals
Faculty PI: Anna Preus (English)
Grad RA: Siddharth Bhogra
Undergrad Fellows: Ameli Graff, Chloe Osborn, Adithi Ramaswamy
Fellows in this group immersed themselves in the early 20th-century archives of the literary magazine The New Age, with a focus on the years 1907-1922. During this time, the periodical was helmed by A.R. Orage and published major authors like H.G. Wells, Ezra Pound, and G. Bernard Shaw. The team was particularly focussed on identifying and extracting poetry from these issues, which had not otherwise been processed by OCR (optical character recognition) algorithms to make them readable. The team did a technology review and ended up selecting the tool “llamaparse,” which allowed them to parse the issues into machine readable text and conduct analyses of things like topic frequency over time and the phenomenon of “shrinkflation” over a century before we had such a word for it.
What’s Seattle Reading?
Faculty PI: Melanie Walsh (Information School)
Grad RA: Neel Gupta
Undergrad Fellows: Alexis Martin, Bob Bao, Karalee Harris, Tiago Ramos
Fellows in this group spent the summer exploring the wealth of records in Seattle Public Library’s open checkout data, one of the most notable and open public resources available to track book circulation and popularity in the United States. Each of the team members pursued stories hidden beneath the surface of this data, looking for patterns that could reveal something about Seattle’s popular culture. In doing so, the team continued the work of previous summers to develop methods for identifying canonical names for authors and their works, as well as an interactive dataset exploration tool. Students investigated topics like how graphic novel popularity has changed over time, seasonal variations in readership behavior and how the library feeds (or fights against!) the “Seattle Freeze.” You can play with the interactive tool yourself at https://melaniewalsh.github.io/whats-seattle-reading-viewer.
SSDSSI Projects
The Environment, Civil Society, and Human Rights Law
Faculty PI: Rachel Cichowski (Political Science)
Graduate RA: Candela Arias Perez
Undergrad Fellows: Isabel Jones, Mia Mao, Ben Nguyen, Yanka Poznanski
Fellows in this group took a close focus to environmental law cases from the European Court of Human Rights (ECHR). In particular, they wanted to understand how ECHR environmental cases impacted national policy and human rights in practice. To do so, they developed web scraping tools and llamaparse to collect text from the corpus of ECHR environmental cases from 1990 to 2026 and national policy effects from 2010 to 2026. Using topic modeling and data visualizations, they were able to see how and where environmental law has had an increasing presence in Europe’s legal landscape over time. The group hopes this work can be extended to other areas of ECHR law in future projects.
Expanding the Corpus of Educational Research Available for Metastudy Analysis
Faculty PI: Jason Kerwin (Economics)
Graduate RA: Tynan Challenor
Undergrad Fellows: Brendan Sheehan, Navneeth Dhamotharan, Phoebe Thio, Yiqian Huang
Fellows in this group helped their PI build a technical pipeline for collecting and cleaning the results of educational research publications to be later used in meta-analyses of educational research effectiveness. The team used Apache Airflow to orchestrate the collection of papers with web scrapers, contacting relevant researchers through professional academic networks, and using LLMs to identify key metadata needed for proper data cleaning. They plan for this system to be used to collect data for a meta-analysis of treatment effect heterogeneity (that is, identifying how educational treatments affect individual students rather than their class as a whole).
Measuring Election Violence Through Machine Learning and AI Approaches to Measurement, Description, and Causal Inference
Faculty PI: James D. Long (Political Science) (with collaborators Luke Condra from the University of Pittsburgh and Danielle Jung from Emory University)
Graduate RA: Zachary Philip-Taylor Brown
Undergrad Fellows: Avery Jensen, Nathan Tu, Yuanxi Li
Fellows in this group worked with their PI and his collaborators at the University of Pittsburgh and Emory University to hone in on a definition and framework for the concept of “contested governance” to understand election violence around the world. To do this, they dove into a variety of country-specific news sources and bodies of literature and combined it with data from ACLED (Armed Conflict Location & Event Data), an NGO tracking global sectarian violence. From this, they trained machine learning models to identify contextual details of violent events and determine whether they represented efforts from rebel groups contesting mainstream governance.

