The invention relates to an
underwater acoustic target
recognition system and method based on multi-
modal depth
feature fusion, and belongs to the field of
underwater acoustic target recognition. The method comprises the following steps: firstly, performing preprocessing on an obtained
underwater acoustic target original audio and associated
metadata, and constructing a multi-
modal data set; and then the constructed multi-
modal sample is input into an identification model, the model extracts deep representation of each modal through a multi-
branch feature coding network, depth alignment and complementary aggregation of different modal features are realized by using a cross-modal cross attention mechanism guided by potential query, and a target identification result is output based on a decision network of mixed experts. According to the method, experimental
verification is carried out on two disclosed underwater acoustic data sets, the experimental result verifies the effectiveness of the multi-modal deep fusion and
hybrid expert adaptive
decision strategy adopted by the method, and through mining the complementary advantages of acoustic features and semantic priori, the multi-modal deep fusion and
hybrid expert adaptive
decision strategy is obtained. And the robustness and generalization ability of the underwater acoustic target
recognition system in the strong-
noise and multi-working-condition environment are remarkably improved.