基于扩散模型的水下场景多模态生成方法
By adopting a diffusion model-based method for underwater scene multimodal generation, the problems of difficult underwater data acquisition and insufficient multimodal annotation are solved. This method achieves unified modeling of multimodal information, improves the structural and semantic consistency of the generated results, and supports high-quality data generation for underwater vision tasks.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- OCEAN UNIV OF CHINA
- Filing Date
- 2026-05-13
- Publication Date
- 2026-07-17
AI Technical Summary
Existing technologies face challenges such as difficulty in acquiring underwater data, insufficient multimodal annotation, and poor cross-modal consistency of generated results. In particular, severe light attenuation and significant scattering effects in the underwater environment lead to high costs and difficulties in acquiring high-quality data, making it difficult for existing methods to achieve unified modeling of multimodal information.
The underwater scene multimodal generation method based on diffusion model (UMDM-USG) is adopted. It generates depth map, semantic segmentation map, surface normal map and text information by collecting multiple raw underwater images, constructs five-tuple multimodal data, and performs feature encoding through text encoder and visual encoder. The multimodal alignment module realizes bidirectional interaction and semantic consistency between generated modalities and conditional modalities. The model is trained by combining reconstruction loss and representation alignment regularization.
It significantly improves the structural and semantic consistency of multimodal generation results, enabling the generation of high-quality multimodal data and supporting downstream tasks such as underwater semantic segmentation, depth estimation, and normal estimation, thereby improving the training stability and generation performance of the model.
Smart Images

Figure CN122176113B_ABST