Diya Saraf

Diya Saraf

Ph.D. Student, Computer Science, Johns Hopkins University

dsaraf2 [AT] jhu.edu

Bio

Ph.D. student in CS at Johns Hopkins, studying the gap between what people say they want online and what recommenders learn they engage with. Previously M.S. at USC/ISI (coordinated inauthentic behavior, algorithmic amplification); B.S. at Santa Clara University.

Algorithmic auditing Recommender systems Computational social science Trust & safety

My research studies the gap between what people say they want to see online and what recommendation algorithms learn they engage with — and what that gap means for the design of social platforms. I work on algorithmic auditing, recommender systems and personalization, and computational social science, combining large-scale observational data, controlled audits, and field studies with instrumented data collection. Recent work has looked at coordinated influence operations [npj Complexity'25] and engagement dynamics around harmful content [DPSH'24].

I'm a Ph.D. student in Computer Science at Johns Hopkins University, where I'm part of the Social Computing group advised by Tiziano Piccardi. Before Hopkins, I completed an M.S. in Computer Science at the University of Southern California, where I worked in the HUMANS Lab at the Information Sciences Institute with Emilio Ferrara and Luca Luceri on coordinated inauthentic behavior, algorithmic amplification, and audits of YouTube and TikTok recommendations. I received my B.S. in Computer Science & Engineering with a minor in Mathematics from Santa Clara University, where I was a Clare Boothe Luce Scholar advised by Yuhong Liu. I've also spent time in industry as a software engineering intern at Dell Technologies.

Looking for Summer 2027 research internships

  • Multimodal LLMs detecting regulated-product content on TikTok at scale — publications ↓
  • LLM-verbalized user preference profiles vs. embedding-based personalization — research ↓
  • Interested in: responsible AI · trust & safety · recommender systems · CSS
Get in touch →

Research

LLM-verbalized preference profiles beat embedding retrieval at predicting what a user actually engages with — and stay human-editable.

Building and evaluating short, editable LLM-verbalized preference profiles against a user's held-out likes, and comparing them to embedding-based retrieval baselines for predicting future engagement. Across several rounds of prompt iteration and controls (permutation baselines, profile-swap tests, temporal train/test splits), verbalized profiles consistently outperform embedding centroids at matched recall — while staying human-readable and directly editable by the user, unlike a learned embedding.

Multimodal LLMs detecting nicotine-pouch and similar regulated-product content on TikTok at scale.

Using multimodal LLMs to detect and characterize regulated-product content — e.g. nicotine pouches — on TikTok at scale, extending computer-vision-based tobacco surveillance methods to handle video, caption, and on-screen text jointly (see Publications).

Comparing coverage and framing across knowledge platforms, e.g. Grokipedia vs. Wikipedia.

Systematically comparing coverage and framing across knowledge platforms — for example Grokipedia against Wikipedia — to surface structural differences in what gets documented and how.

Browser-based tooling for collecting real, consented feed data for field experiments.

Building browser-based tooling that collects real, consented feed data from study participants, enabling field experiments on live recommendation systems rather than simulated ones.

Computational social science Algorithmic auditing Recommender systems LLMs / NLP for user behavior Human–AI interaction Network science Trust & safety

Publications

Most recent publications on Google Scholar.
indicates equal contribution.

Scaling Surveillance of Nicotine Pouch Content on TikTok Using Multimodal Large Language Models

Artur Galimov, Luca Luceri, Reid C. Whaley, Diya Saraf, David Chu, Fred Morstatter, Yolanda Gil, Adam M. Leventhal, Jennifer B. Unger

Manuscript submitted for publication. 2026.

Large-scale Detection of Multilingual Coordinated Activity on Telegram

Leonardo Blas, Diya Saraf, Tanishq Salkar, Nora Adadurova, Luca Luceri, Emilio Ferrara

npj Complexity, vol. 2, 33. 2025.

The Impact of Emojis on User Engagement with Trolling Content in Online Platforms

Diya Saraf, Yuhong Liu, Hooria Jazaieri

IEEE DPSH'24: Digital Platforms and Societal Harms. 2024.

Scaling Surveillance of Nicotine Pouch Content on TikTok Using Multimodal Large Language Models

Artur Galimov, Luca Luceri, Reid C. Whaley, Diya Saraf, David Chu, Fred Morstatter, Yolanda Gil, Adam M. Leventhal, Jennifer B. Unger

Manuscript submitted for publication. 2026.

Large-scale Detection of Multilingual Coordinated Activity on Telegram

Leonardo Blas, Diya Saraf, Tanishq Salkar, Nora Adadurova, Luca Luceri, Emilio Ferrara

npj Complexity, vol. 2, 33. 2025.

The Impact of Emojis on User Engagement with Trolling Content in Online Platforms

Diya Saraf, Yuhong Liu, Hooria Jazaieri

IEEE DPSH'24: Digital Platforms and Societal Harms. 2024.

Vitæ

Blog

Notes on PhD applications, and running routes around the Bay Area and LA. All posts →

Sep 14, 2026 running

Running Spots: Bay Area, LA, and Beyond

A running list (pun intended) of routes worth the detour — starting with the Bay Area and greater LA, growing as I travel.

Sep 14, 2026 phd-advice

Notes on Applying to CS PhD Programs

What actually moved the needle on my applications, roughly in order of leverage — advisor fit, research statement, and the mechanics ever...

Outside Research

Crime documentaries Mountains Word puzzles Running — routes on the blog