Shaden Shaar

Shaden Shaar

I am a final-year PhD student in Computer Science at Cornell University, advised by Prof. Claire Cardie. My thesis, Long-Form Input Understanding across Vision-and-Language Tasks, asks what a model needs from a long input: what it has to locate, and when locating the right piece is not enough. I have studied this question in text (TACL 2025) and in short video, for which I built the MovieRecapsQA benchmark (CVPR 2026). I am now investigating long narrative videos. Alongside my thesis, I work with NewYork-Presbyterian Hospital on heart-failure projects, ranging from analyzing UNOS exception requests (JHLT 2026) to predictive medication models for heart-failure patients.

Prior to Cornell, I spent two years as a research assistant at the Qatar Computing Research Institute, working with Prof. Preslav Nakov and Prof. Giovanni Da San Martino on news fact-checking and propaganda detection, where I introduced the task of detecting previously fact-checked claims (ACL 2020). I received my BS in Computer Science from Carnegie Mellon University in 2019, with College and University Honors.

Shaden Shaar

News

  • Sep 2026

    Visiting the Berkeley NLP group this semester, working with Dr. Sewon Min.

  • Jun 2026

    MovieRecapsQA, our open-ended benchmark for question-answering over full-length movies, appeared at CVPR 2026.

  • May 2026

    Our JHLT paper with NewYork-Presbyterian uses an LLM for thematic analysis of accepted heart-transplant exception requests.

  • Aug 2025

    Wrapped up a summer as an Applied Scientist Intern at Zillow, building conversational real-estate agents with reinforcement learning.

  • May 2025

    "Are Triggers Needed for Document-Level Event Extraction?" was published in TACL. Also finished an ML research engineering internship at Scale AI.

Experience

  • May — Aug 2025

    Applied Scientist Intern · Zillow Group

    Remote, USA

    • Built a home-purchase co-pilot for Zillow: a multi-turn assistant that answers buyer questions, tracks purchase intent as it accumulates over a conversation, and narrows a large listing inventory toward properties that fit.
    • Post-trained the dialogue policy on long multi-turn conversation data, running DPO over preference pairs scored on whole dialogues, then GRPO with a trajectory-level reward so credit lands on the conversation rather than the individual turn.
    • Built the reward-modeling and evaluation pipeline behind that training, including rollout generation and conversation-level scoring for grounded property-search dialogue.
  • Jan — May 2025

    Machine Learning Research Engineer Intern · ScaleAI, Inc.

    New York City, NY, USA

    • Built the judge co-pilot for the Qatari judicial department: a set of case-analysis services, each exposed as an API, deployed together as a judge-facing dashboard.
    • Designed the Arabic information-extraction pipeline, running OCR over scanned filings, then event extraction that links each extracted event to the evidence spans supporting it in the case document, feeding a case summarizer built on that structure.
    • Built precedent lookup as a RAG system over constitutional bylaws and prior Qatari Supreme Court cassation rulings, combining dense embedding retrieval with sparse lexical matching and reranking the merged candidate set.
  • May — Aug 2022

    AI/ML Research Intern · Apple

    Seattle, WA, USA

    • Built a conditional synthetic dialogue generator for Siri training with Dr. Alex Churchill, producing multi-turn conversation data for scenarios with little or no real usage data to draw on.
    • Simulated user behaviour to generate the human side of each conversation, so the pipeline produced complete multi-turn dialogues rather than isolated utterances.
    • Conditioned generation on intent and user behaviour, letting a target dataset be specified to arbitrary requirements rather than freely sampled, and ran zero-shot without seed dialogues for the target scenarios.
  • Jul 2019 — Aug 2021

    Research Assistant · Qatar Computing Research Institute

    Doha, Qatar

    • Automated fact-checking: introduced the task of detecting previously fact-checked claims, matching an incoming claim against a database of prior verdicts so a checker could reuse existing work instead of starting over. Follow-up work extended it to document-level matching and to the role of surrounding context, with papers at ACL, EMNLP, RANLP, and IJCAI.
    • Propaganda and persuasion detection: built Prta, a public system that flags propaganda techniques in news articles (ACL 2020 Best Demo, Honorable Mention), then extended the work to multimodal memes where the technique is split across text and image, published at ACL and SemEval.
    • COVID-19 misinformation: built pipelines for factuality, harmfulness, and framing of pandemic claims across multiple languages and social platforms, with papers at ICWSM, EMNLP, NAACL, and COLING.
    • Emotion detection: cross-lingual classification for low-resource languages with little annotated data (LREC).
    • Co-organized shared tasks at SemEval, CLEF CheckThat!, and ACL workshops, releasing the datasets and evaluation setups these problems are still benchmarked on. Several of the news-facing systems were taken up by news agencies and fact-checking organizations.

Selected publications

All 32 publications →

A few representative papers. The full list is on the publications page and Google Scholar.

Education

  • 2021 — present

    PhD in Computer Science

    Cornell University · Ithaca, NY

    Advised by Prof. Claire Cardie · Minor in Applied Mathematics · Graduating Jan 2027

    University Fellowship (2021)

  • 2021 — 2024

    MS in Computer Science

    Cornell University · Ithaca, NY

  • 2015 — 2019

    BS in Computer Science

    Carnegie Mellon University · Doha, Qatar

    Minor in Mathematics · College and University Honors

    50% Academic Merit Scholarship (2015)

Contact

The best way to reach me is by email at ss2753 [at] cornell [dot] edu — whether about roles, collaborations, or anything related to the work. A printable CV is available as a PDF.