基于扩散模型的水下场景多模态生成方法

By adopting a diffusion model-based method for underwater scene multimodal generation, the problems of difficult underwater data acquisition and insufficient multimodal annotation are solved. This method achieves unified modeling of multimodal information, improves the structural and semantic consistency of the generated results, and supports high-quality data generation for underwater vision tasks.

CN122176113BActive Publication Date: 2026-07-17OCEAN UNIV OF CHINA

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
OCEAN UNIV OF CHINA
Filing Date
2026-05-13
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing technologies face challenges such as difficulty in acquiring underwater data, insufficient multimodal annotation, and poor cross-modal consistency of generated results. In particular, severe light attenuation and significant scattering effects in the underwater environment lead to high costs and difficulties in acquiring high-quality data, making it difficult for existing methods to achieve unified modeling of multimodal information.

Method used

The underwater scene multimodal generation method based on diffusion model (UMDM-USG) is adopted. It generates depth map, semantic segmentation map, surface normal map and text information by collecting multiple raw underwater images, constructs five-tuple multimodal data, and performs feature encoding through text encoder and visual encoder. The multimodal alignment module realizes bidirectional interaction and semantic consistency between generated modalities and conditional modalities. The model is trained by combining reconstruction loss and representation alignment regularization.

Benefits of technology

It significantly improves the structural and semantic consistency of multimodal generation results, enabling the generation of high-quality multimodal data and supporting downstream tasks such as underwater semantic segmentation, depth estimation, and normal estimation, thereby improving the training stability and generation performance of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122176113B_ABST
    Figure CN122176113B_ABST
Patent Text Reader

Abstract

基于扩散模型的水下场景多模态生成方法(UMDM‑USG),属于计算机视觉与图像生成技术领域。针对水下场景数据采集成本高等不足,本发明首先对输入文本及多模态数据进行特征编码;其次通过角色分配模块随机将模态划分为生成模态与条件模态;然后在多模态对齐模块中构建生成条件对齐注意力,实现生成模态与条件模态之间的双向交互与对齐;最后,通过解码器输出各模态结果。此外引入了表征对齐正则化策略,利用预训练自监督模型的视觉先验作为监督信号,引导生成分布趋向真实场景,显著提升了几何保真度。实验结果表明,本发明在生成质量、语义一致性及多模态协同方面均优于现有方法,并且可以为下游视觉任务提供可靠的数据支持。
Need to check novelty before this filing date? Find Prior Art