The invention relates to a multi-
modal image fusion method based on hierarchical semantic richness, and belongs to the technical field of multi-
modal image fusion. According to the method,
global information exchange is balanced by adopting multi-scale
feature aggregation and redistribution, meanwhile, fusion and segmentation tasks are dynamically bridged, a progressive semantic dense injection strategy is introduced, and global
semantics are injected into highly consistent
infrared features through dense connection, so that the global semantic fusion is realized. And then the semantic-
infrared mixed features are transmitted to visible light features. Secondly, two types of
feature fusion modules are introduced, one is to use a cross-
modal attention mechanism to carry out more comprehensive
feature fusion, the other is to use semantic features as third input to enhance
semantic representation of
image fusion, and
feature fusion is realized in a complex scene by dynamically balancing global
semantic consistency and fine-grained local detail representation. According to the method, layering and enriching of
semantic information are achieved through strategies such as semantic collection, distribution and injection, and the fusion visual effect and downstream
perception performance are enhanced.