Field-experiment data infrastructure.
Engineering and auditing research-data infrastructure for a large-scale field experiment, in collaboration with Mridul Joshi, Researcher at Stanford.
Smartphone-use research data, independently validated.
Reliable data. Stronger research.
Developing and auditing the infrastructure behind a large-scale field experiment, from source records to research outputs.
Source-data discrepancies uncovered; reliability of the final research outputs strengthened.
Follow the data across environments.
I developed and audited research-data infrastructure across AWS, Stanford’s Sherlock computing cluster, and local analysis environments. The work involved millions of smartphone-use records and complex export and identity-linkage problems.
- AWS and Sherlock
- Local analysis environments
- Millions of smartphone-use records
Check what the matching pipeline actually produced.
I independently validated an ML-assisted friendship-network matching pipeline. Modern AI tools supported the work, alongside rigorous human review, provenance checks, and reproducibility safeguards.
- Independent ML-assisted matching validation
- Human review and provenance checks
- Reproducibility safeguards
Investigate discrepancies at the source.
The work uncovered important source-data discrepancies and materially strengthened the reliability of the final research outputs. Mridul Joshi’s recommendation specifically highlights the willingness to investigate rather than accept superficially plausible results.
- Source-data investigation
- Documented reasoning
- Evidence from the research collaborator’s recommendation