BIT: Bidirectional Image to Text Diffusion Bridges for Multimodality Translation
Abstract
Multimodality translation (e.g., text-to-image) is a core generative AI task. However, existing approaches (1) have semantically meaningless generative paths, greatly limiting the flexibility of sampling algorithms; (2) are unidirectional, preventing inversion (e.g., image-to-text). We propose BIT: Bidirectional Image-Text Diffusion Bridges. In contrast to previous approaches, BIT starts directly from text and interpolates into images, providing (1) a semantically rich generative path that enables diverse and flexible sampling algorithms; (2) inherent invertibility of the model from image to text, providing a unified, bidirectional generative framework. BIT is derived through rigorous stochastic calculus and measure theory, yielding SDE forms that are friendly to simulation and tractable loss functions that scale well in high dimensions. Our theoretical analysis and empirical experiments further confirm the advantage over both noise-to-data diffusion baselines and deterministic flow model baselines; not only in vision-language tasks, but also in scientific domains like cell modeling.
Image-to-Text-to-Image Cycle Generation Results
Each video shows an image-to-text-to-image generation cycle, starting from the real image, converting it to text, and then converting the text back to a new image.
Text-to-Image-to-Text Cycle Generation Results
Each video shows a text-to-image-to-text generation cycle, starting from the real text, converting it to image, and then converting the image back to a new text.
BibTeX
@article{abc2026,
title={There and Back Again: Bidirectional Image-Text Diffusion Bridges for Multimodality Translation},
author={Anonymous Authors},
journal={Preprint},
year={2026},
url={https://bit-diffusion.github.io/}
}