Publication

Publications from our lab, organized by year.

2026

  1. From State to Action: OODA-Tool for Reliable Multi-Turn Tool Use
    Rongfeng Guo, Yinxuan Huang, Yusen Wu, Maoqing Zhong, Yunlu Chen, Meng Tang, Teng Long, Vincent Tao Hu
    arXiv 2026
  2. Diversifying Similar Subjects for Text-to-image Synthesis with Self-Cross Diffusion Guidance and Reward
    Weimin Qiu, Jieke Wang, Zhining Gu, Vincent Tao Hu, Meng Tang
    International Journal of Computer Vision 2026
  3. Guiding Token-Sparse Diffusion Models
    Felix Krause, Stefan Andreas Baumann, Johannes Schusterbauer, Olga Grebenkova, Ming Gui, Vincent Tao Hu, Björn Ommer
    CVPR 2026
  4. Purrception: Variational Flow Matching for Vector-Quantized Image Generation
    Răzvan-Andrei Matișan, Vincent Tao Hu, Grigory Bartosh, Björn Ommer, Cees G.M. Snoek, Max Welling, Jan-Willem Meent, Mohammad Mahdi Derakhshani, Floor Eijkelboom
    ICLR 2026
  5. Diffusion Models and Representation Learning: A Survey
    Michael Fuest, Pingchuan Ma, Ming Gui, Johannes Fischer, Vincent Tao Hu, Bjorn Ommer
    T-PAMI 2026

2025

  1. TREAD: Token Routing for Efficient Architecture-agnostic Diffusion Training
    Felix Krause, Timy Phan, Ming Gui, Stefan Andreas Baumann, Vincent Tao Hu, Björn Ommer
    ICCV 2025 Oral
  2. Stochastic Interpolants for Revealing Stylistic Flows across the History of Art
    Pingchuan Ma, Ming Gui, Johannes Schusterbauer, Xiaopei Yang, Olga Grebenkova, Vincent Tao Hu, Björn Ommer
    ICCV 2025
  3. Continuous, Subject-Specific Attribute Control in T2I Models by Identifying Semantic Directions
    Stefan Andreas Baumann, Felix Krause, Michael Neumayr, Nick Stracke, Vincent Tao Hu, Björn Ommer
    CVPR 2025
  4. MaskFlow: Discrete Flows for Flexible and Efficient Long Video Generation
    Michael Fuest, Vincent Tao Hu, Björn Ommer
    arXiv 2025
  5. ToddlerDiffusion: Flash Interpretable Controllable Diffusion Model
    Eslam Mohamed BAKR, Liangbing Zhao, Vincent Tao Hu, Matthieu Cord, Patrick Perez, Mohamed Elhoseiny
    ICLR 2025
  6. DepthFM: Fast Monocular Depth Estimation with Flow Matching
    Ming Gui, Johannes S. Fischer, Ulrich Prestel, Pingchuan Ma, Dmytro Kotovenko, Olga Grebenkova, Stefan A. Baumann, Vincent Tao Hu, Björn Ommer
    AAAI 2025 Oral
  7. Does VLM Classification Benefit from LLM Description Semantics?
    Pingchuan Ma, Lennart Rietdorf, Dmytro Kotovenko, Vincent Tao Hu, Björn Ommer
    AAAI 2025 Oral at Workshop
  8. Distillation of Diffusion Features for Semantic Correspondence
    Frank Fundel, Johannes Schusterbauer, Vincent Tao Hu, Björn Ommer
    WACV 2025

2024

  1. [MASK] is All You Need
    Vincent Tao Hu, Björn Ommer
    arXiv 2024
  2. Scaling Image Tokenizers with Grouped Spherical Quantization
    Jiangtao Wang, Zhen Qin, Yifan Zhang, Vincent Tao Hu, Björn Ommer, Rania Briq, Stefan Kesselheim
    arXiv 2024
  3. ZigMa: A DiT-style Zigzag Mamba Diffusion Model
    Vincent Tao Hu, Stefan Andreas Baumann, Ming Gui, Olga Grebenkova, Pingchuan Ma, Johannes Fischer, Bjorn Ommer
    ECCV 2024
  4. Boosting Latent Diffusion with Flow Matching
    Johannes S. Fischer, Ming Gui, Pingchuan Ma, Nick Stracke, Stefan A. Baumann, Vincent Tao Hu, Bjorn Ommer
    ECCV 2024 Oral
  5. Guided Flow Vision Transformer from Self-Supervised Diffusion Features
    Vincent Tao Hu, Yunlu Chen, Mathilde Caron, Yuki M. Asano, Cees G.M. Snoek, Björn Ommer
    arXiv 2024
  6. Motion Flow Matching for Human Motion Synthesis and Editing
    Vincent Tao Hu, Wenzhe Yin, Pingchuan Ma, Yunlu Chen, Basura Fernando, Yuki M. Asano, Efstratios Gavves, Pascal Mettes, Björn Ommer, Cees G.M. Snoek
    arXiv 2024
  7. Training Class-Imbalanced Diffusion Model Via Overlap Optimization
    Divin Yan, Lu Qi, Vincent Tao Hu, Ming-Hsuan Yang, Meng Tang
    arXiv 2024
  8. Latent Space Editing in Transformer-based Flow Matching
    Vincent Tao Hu, David W Zhang, Pascal Mettes, Meng Tang, Deli Zhao, Cees G.M. Snoek
    AAAI 2024. Also appear in ICML 2023 Workshop, New Frontiers in Learning, Control, and Dynamical Systems 2024
  9. Flow Matching for Conditional Text Generation in a Few Sampling Steps
    Vincent Tao Hu, Di Wu, Yuki M. Asano, Pascal Mettes, Basura Fernando, Björn Ommer, Cees G.M. Snoek
    EACL 2024

2023

  1. Self-Guided Diffusion Models
    Tao Hu*, David W Zhang*, Yuki M. Asano, Gertjan J. Burghouts, Cees G.M. Snoek
    CVPR 2023

2021

  1. Self-supervised Video Representation Learning with Cross-Stream Prototypical Contrasting
    Martine Toering, Ioannis Gatopoulos, Maarten Stol, Tao Hu
    WACV 2021

2020

  1. Localizing the Common Action Among a Few Videos
    Pengwan Yang*, Tao Hu*, Pascal Mettes, Cees G.M. Snoek
    European Conference on Computer Vision(ECCV) 2020
  2. Pointmixup: Augmentation for point clouds
    Yunlu Chen*, Tao Hu*, Efstratios Gavves, Thomas Mensink, Pascal Mettes, Pengwan Yang, Cees G.M. Snoek
    European Conference on Computer Vision(ECCV) 2020 Spotlight

2019

  1. Attention-based Multi-Context Guiding for Few-Shot Semantic Segmentation
    Tao Hu, Pengwan Yang, Chiliang Zhang, Gang Yu, Yadong Mu, Cees G.M. Snoek
    AAAI 2019 Spotlight
  2. SILCO: Show a Few Images, Localize the Common Object
    Tao Hu, Pascal Mettes, Jia-Hong Huang, Cees G.M. Snoek
    International Conference on Computer Vision(ICCV) 2019