Semantic deduplication.
A scraping, cleaning, and semantic deduplication workflow for large research datasets, developed in collaboration with Mridul Joshi, Researcher at Stanford.
7M+ duplicates removed in the deduplication project.
Illustrative workflow · explore the engineering above
The problem
Exact matching could not detect differently worded records that repeated the same meaning. These semantic duplicates reduced the quality and reliability of data used for downstream research and ML.
What I built
Engineered a scraping and cleaning workflow backed by MongoDB. Used HDBSCAN clustering to group semantically similar records and language models to verify duplicate meaning before removing redundant records.
The engineering thinking
Similarity alone does not establish duplication. Clustering groups candidate records; LLM verification helps distinguish repeated meaning from records that are merely related.
The outcome
Processed 33M+ news records and removed 7M+ duplicates, improving data quality for downstream ML training and research. This project is separate from the smartphone-use field experiment.