
After almost a year of work, we have finalized the first release of our comprehensive, open-access digital dataset comprising over 38,000 digitized historical records of Chinese lianhuanhua publications spanning the years 1949–1994!
Provided in a comma-separated values (CSV) format, the dataset digitizes over 1,200 printed pages of densely packed tables to document detailed publication information. As such, the dataset not only systematizes publication data as preserved in two decades-old catalogues, but also reveals novel, multilayered connections across tens of thousands of titles, deepening our understanding of what was one of the world’s most widely consumed reading materials throughout the twentieth century.
The ChinaComx Lianhuanhua Publication Dataset (v1.0.0)
Read all about the dataset on the GitHub repository page or download it via Zenodo
This dataset is an academic output of the ERC-ChinaComx project, led by Principal Investigator Lena Henningsen. Damian Mandzunowski conceptualized and supervised the development of the dataset. Tilen Zupan executed the main OCR extraction and data parsing, building upon initial digitization groundwork by Bettina Jin. We thank Matthias Arnold and Aijia Zhang for their valuable feedback and brainstorming throughout the work on the dataset.
Primary source materials such as the lianhuanhua publication catalogues were sourced from the Centre for Asian and Transcultural Studies (CATS) Library, Heidelberg University as well as from the ERC-ChinaComx collection at the CATS Library.
