The present application relates to a kind of
multimodal image fusion methods of inverse
diffusion model, comprising the following steps: step one, using inverse
diffusion technology, visible light image is reversed to
noise latent space, then using the feature of reversed visible light image, guide
infrared image is reversed;Step two, guide by the inverse process in
diffusion model, inject the appearance attribute of visible light into
infrared feature, and the feature can generate
infrared image with visible light style;Step three, design specific fusion rule, for the attention layer fusion of denoising process, reversed visible light and infrared feature are fused, the text interaction capability of model is retained, and language-driven fusion control is supported.The present application can directly generate high-quality
fusion image without additional training or fine-tuning.The obtained
fusion image is highly compatible with the base model, effectively solves the difference problem between data domains, and significantly improves the performance of downstream
machine perception tasks.The present application significantly reduces the training cost, and provides an efficient and innovative solution for cross-domain tasks.