Hello

Aalborg, Denmark, Spring 2023
Aalborg, Denmark, Spring 2023

About Me

I am a ZJU 100-talent Program Professor [百人计划研究员,博导] with the College of Computer Science and Technology at Zhejiang University. I am also an NSFC Excellent Young Scientist (Overseas) and an EU Marie Curie Individual Fellow [欧盟"玛丽·居里学者"].

Before rejoining Zhejiang University, I was an Assistant Professor at Aalborg University in Denmark (2020–2023). My experience spans both academia and industry, having previously worked as a Senior Engineer at Ant Group (2018–2019). I received my Ph.D. from Zhejiang University in 2018 and my B.Sc. from Sichuan University in 2012. I am honored to be a recipient of the ACM SIGMOD China Rising Star Award and the ACM China Rising Star Award (Honorable Mention).

Research

I am currently focused on research in data-centric, resource-efficient, and scalable AI, which I believe are essential for advancing the adoption of AI in real-world applications. My earlier work primarily focused on data management and AI techniques in the realm of time-series (TS) and spatiotemporal (ST) data. For now, these interests have broadened to include relational data, unstructured data, and multimodal data.

I'm lucky to work and grow with an amazing group of passionate and talented young people at Zhejiang University. Our research group is called Sustainable Data Intelligence and Data Systems [SuDIS]. Lately, we've been focusing on pushing the boundaries in two key areas.

Data + AI

  • Data preparation and governance for AI: data storage [HyperMR, SIGMOD'25; DeXOR, VLDB'25], knowledge amalgamation [BoKA, KDD'24], data protection [PIECK, ICDE'24], data selection [CHASe, TKDE'25; TAD, EMNLP'25; CoIDO, NeurIPS'25], and data augmentation and synthesis [CogSQL, AAAI'25; NotAllDataAreGoodLabels, NeurIPS'25; TabAug, CSUR'26].
  • AI for data quality and data systems: imputation and anomaly detection [MPIN, PVLDB'24; AdaCTSi, TKDE'26; CoAD, KDD'26], schema and column understanding [Filter-then-match, CIKM'26; LLM Knows, Speaks Not, EMNLP Findings'26], data discovery and BI [nlcTables, SIGIR'25; TableCopilot, VLDB'25; RedParrot, ICDE Industry'26; ChronosBI, SIGMOD Demo'26], data ingestion [Hippo, SIGMOD'25], and workload optimization [SafeLoad, VLDB'25; ScaleSense, VLDB Industry'26].

Efficient AI

  • Lightweight time-series models for edge computing [LightCTS, SIGMOD'23; E2USD, WWW'24; ReCTSi, KDD'24; LightCTS*, TKDE'24; PimShare, TCAD'25] and federated learning [LightTR, ICDE'24; FedBFPT, IJCAI'23; LightTR+, TKDE'25].
  • Efficient large and multimodal model training, inference, and serving: multi-tenant reuse [HMI, VLDBJ'25], speculative and parallel decoding [Draft & Verify, ACL'24; SpecVLM, EMNLP'25; Double, ACL'26; LVSpec, ACL'26; ParallelVLM, CVPR'26], KV-cache compression [HybridKV, ACL'26; HARD-KV, ICML'26], and memory-efficient fine-tuning [LoRAM, ICLR'25].
ZJU AAU DAISY SuDIS MALOT Indoor-LBS