Qiyu Wu

Researcher at Sony / Research Lead

Ph.D. from UTokyo

Logo Photo

Home

Publications

Experiences

Email: wuqiyu576 [AT] gmail [DOT] com

qiyu.wu [AT] sony [DOT] com

qiyuw [AT] g.ecc.u-tokyo [DOT] ac.jp wuqiyu [AT] pku [DOT] edu.cn

Visitor Map

ClustrMaps visitor map

Bio

Hi there! This is Qiyu Wu, a Research Scientist on the Multimodal NLP team at Creative AI Lab, Sony. We conduct language-centric multimodal research to enhance content creation in music, film, and games. Feel free to contact for discussion or collaboration!

Before joining Sony, I received my Ph.D. from The University of Tokyo, advised by Yoshimasa Tsuruoka and supported by JSPS DC Fellowship. I have served as an Area Chair for ACL Rolling Review and NeurIPS, and as a program committee member (reviewer) for several top-tier conferences including ACL, EMNLP, NAACL, ICLR, NeurIPS, and ICML. Additionally, I co-organized the GenProCC Workshop at NeurIPS 2025. Here is my CV. Reach out to me by wuqiyu576 [AT] gmail [DOT] com, or LinkedIn.

Research Interests

My research focus lies in Multimodal NLP, mainly encompassing multimodal LLMs as well as better representing textual semantics in both monolingual and multilingual contexts. I have published papers More at conferences such as ACL, EMNLP, NAACL, NeurIPS, ICLR, ICML, AAAI, EACL, VLDB, etc.

News

Recent Research

Mixture of Probes thumbnail

Mixture of Probes

NeurIPS 2026

Learning from training-only privileged modalities through structured probing, improving multimodal LLMs when only one modality is available at inference.

MCA thumbnail

MCA

EMNLP 2026 Main, Oral

Modality composition awareness reduces modality shortcuts and improves composed multimodal retrieval under distribution shifts.

MLLMCLIP thumbnail

MLLMCLIP

EMNLP 2026 Main

Feature-level distillation transfers knowledge from multimodal LLMs to CLIP, improving compositional understanding and vision-language representations.

MusTBench thumbnail

MusTBench

Preprint 2026

A music-expert-validated benchmark with five temporal grounding tasks, paired with MusT training to improve temporal understanding in music LLMs.