A multi-modal underwater acoustic target recognition method and system, and a storage medium

CN121142518BActive Publication Date: 2026-09-15SICHUAN JIUZHOU ELECTRIC GROUP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511243556.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-02
Publication Date
2026-09-15
Estimated Expiration
2045-09-02

AI Technical Summary

Technical Problem

[0004]本申请实施例提供了一种多模态水声目标识别方法、系统及存储介质,用于解决不同传感器的模态数据如何进行融合的问题以及如何提高多传感器水声目标识别的效果

Benefits of technology

(1)本发明提出的多模态水声目标识别模型,能够处理多传感器水声数据,避免了因识别结果决策融合导致的误差积累和信息损失。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121142518B_ABST
    Figure CN121142518B_ABST
Patent Text Reader

Abstract

The application provides a multimodal underwater acoustic target recognition method, comprising: data acquisition: collecting underwater acoustic data of different sensors and dividing into untagged underwater acoustic data and tagged underwater acoustic data; multimodal underwater acoustic target recognition model construction: the constructed multimodal underwater acoustic target recognition model comprises a feature extraction network and a recognition network; pre-training stage: the multimodal underwater acoustic target recognition model is pre-trained by using the untagged underwater acoustic data to determine the feature extraction network weight; supervised training stage: the feature extraction network weight of the pre-training stage is used as the initialization weight of the feature extraction network of the current stage, and the tagged underwater acoustic data is input into the feature extraction network to update the multimodal underwater acoustic target recognition model; target recognition: the real-time acquired underwater acoustic data of different sensors is input into the trained model to obtain a target recognition result. The application can process multi-sensor underwater acoustic data and improve the accuracy of underwater acoustic target recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of underwater acoustic target recognition, and in particular to a multimodal underwater acoustic target recognition method, system, and storage medium. Background Technology

[0002] In the field of underwater acoustic target identification, sonar plays a crucial role in marine monitoring, military reconnaissance, and environmental protection. In complex underwater environments, sonar sensors can be deployed to capture and analyze underwater acoustic signals in real time to identify submarines, torpedoes, and other surface and underwater targets. However, given the complex marine environment and diverse target types, the detection capabilities of existing single-platform, single-sensor systems are limited. Single sonar systems are constrained by limited spatiotemporal detection capabilities and accuracy, often resulting in inaccurate and uncertain target information, making it difficult to meet the demands for accurate detection and identification of underwater acoustic targets. Therefore, the comprehensive utilization of diverse information to construct and research multi-sensor, multi-platform, and multi-array underwater detection and identification technologies is a major future development trend.

[0003] Traditional multi-sensor underwater acoustic target recognition methods are mostly based on decision-level fusion. Decision-level fusion combines the decisions or results obtained from independent analysis of various data sources or feature sets to form the final decision, i.e., the underwater acoustic target recognition result. However, these methods often rely on the accuracy of each independent analysis. If the recognition result from a certain data source is inaccurate, it may directly affect the reliability of the final decision. Secondly, decision-level fusion usually ignores the correlation and complementarity between different data sources, which may lead to information loss. Therefore, methods that rely solely on decision-level fusion may not achieve optimal underwater acoustic target recognition performance in complex environments. Summary of the Invention

[0004] This application provides a multimodal underwater acoustic target recognition method, system, and storage medium to address the problem of how to fuse modal data from different sensors and how to improve the effectiveness of multi-sensor underwater acoustic target recognition.

[0005] Other features and advantages of this application will become apparent from the following detailed description, or may be learned in part from practice of this application.

[0006] According to a first aspect of the embodiments of this application, a multimodal underwater acoustic target recognition method is provided, comprising: Data acquisition: Collect underwater acoustic data from different sensors and divide it into unlabeled underwater acoustic data and labeled underwater acoustic data; Construction of a multimodal underwater acoustic target recognition model: The multimodal underwater acoustic target recognition model includes a feature extraction network and a recognition network; Pre-training phase: The multimodal underwater acoustic target recognition model is pre-trained using unlabeled underwater acoustic data to determine the weights of the feature extraction network; Supervised training phase: The weights of the feature extraction network in the pre-training phase are used as the initial weights of the feature extraction network in the current phase. Labeled underwater acoustic data is input into the feature extraction network to complete the update of the multimodal underwater acoustic target recognition model. Target recognition: The underwater acoustic data acquired in real time from different sensors is input into the multimodal underwater acoustic target recognition model that has completed the supervised training phase to obtain the target recognition result.

[0007] In one embodiment of this application, the pre-training phase specifically includes: The underwater acoustic data of two sensors were randomly selected from the unlabeled underwater acoustic data, namely the first underwater acoustic data and the second underwater acoustic data. The first underwater acoustic data is input into an online network to obtain a predicted value; the online network includes a feature extraction network, a projection network, and a prediction network. The second underwater acoustic data is input into the target network to obtain the target value; the target network includes a feature extraction network and a projection network; the feature extraction network in the online network and the target network has the same structure as the feature extraction network in the multimodal underwater acoustic target recognition model; both the projection network and the prediction network are composed of multiple fully connected layers; Training is completed by constraining the distance between predicted and target values ​​in the online and target networks using a loss function. During training, the gradient values ​​of the target network are continuously updated, and the weights of the feature extraction network in the online network are updated to the feature extraction network of the target network using the moving average method.

[0008] In one embodiment of this application, before inputting the first underwater acoustic data into the online network and the second underwater acoustic data into the target network, the first and second underwater acoustic data are further subjected to random enhancement.

[0009] In one embodiment of this application, the random enhancement includes pitch shifting, speed shifting, power gain, random clipping, and padding.

[0010] In one embodiment of this application, the supervised training phase specifically includes: The weights of the feature extraction network obtained in the pre-training stage are used as the initial weights of the feature extraction network in the current multimodal underwater acoustic recognition model. Labeled underwater acoustic data from multiple sensors are input into the feature extraction network to extract underwater acoustic data features. After integrating the underwater acoustic data features from multiple sensors into the recognition network of the multimodal underwater acoustic recognition model, the loss value between the recognition result and the underwater acoustic target type label is calculated using a loss function, thereby constraining the model update.

[0011] In one embodiment of this application, the target identification specifically includes: Acquire underwater acoustic data from multiple sensors; Feature maps of underwater acoustic data from each sensor are extracted using the feature extraction network of the multimodal underwater acoustic recognition model. The feature maps of water depth data from all sensors are fused to obtain multi-sensor fused underwater acoustic features. The underwater acoustic features fused from multiple sensors are input into the recognition network of a multimodal underwater acoustic recognition model to determine the type of underwater acoustic target.

[0012] In one embodiment of this application, fusing the feature maps of all sensor water depth data to obtain multi-sensor fused underwater acoustic features includes: calculating the mean of the feature maps of all sensor water depth data as the multi-sensor fused underwater acoustic features.

[0013] In one embodiment of this application, both the pre-training phase and the supervised training phase employ the Adam optimizer.

[0014] According to a second aspect of the embodiments of this application, a system is provided, including a memory and a processor, wherein the memory stores a computer program that can be loaded by the processor and executed corresponding to the multimodal underwater acoustic target recognition method described in the first aspect.

[0015] According to a third aspect of the embodiments of this application, a computer-readable storage medium is provided, on which computer program instructions are stored, which, when executed by a processor, are used to implement the process corresponding to the multimodal underwater acoustic target recognition method as described in the first aspect.

[0016] Compared with existing technologies, the beneficial effects of adopting the above technical solution are as follows: (1) The multimodal underwater acoustic target recognition model proposed in this invention can process underwater acoustic data from multiple sensors, avoiding error accumulation and information loss caused by decision fusion of recognition results.

[0017] (2) The multi-stage multimodal underwater acoustic target recognition model training method proposed in this invention can further improve the accuracy of the model in recognizing multimodal underwater acoustic targets. Attached Figure Description

[0018] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.

[0019] Figure 1 This is a flowchart of a multimodal underwater acoustic target recognition method according to an embodiment of this application.

[0020] Figure 2 This is a flowchart of the pre-training stage in an embodiment of this application.

[0021] Figure 3 This is a flowchart of the supervised training phase in an embodiment of this application.

[0022] Figure 4 This is a flowchart illustrating the target recognition process in an embodiment of this application.

[0023] Figure 5 This is a schematic diagram of an electronic device according to an embodiment of this application.

[0024] Figure 6 This is a schematic diagram of the structure of a computer system suitable for implementing the embodiments of this application. Detailed Implementation

[0025] The embodiments of this application are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar modules or modules having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this application, and should not be construed as limiting this application. Rather, the embodiments of this application include all variations, modifications, and equivalents falling within the spirit and scope of the appended claims.

[0026] Traditional methods relying solely on decision-level fusion may fail to achieve optimal underwater acoustic target recognition performance in complex environments. In recent years, artificial intelligence technologies, represented by deep learning and big data, have made rapid progress in multiple fields such as speech recognition and image processing. These achievements provide valuable insights for multi-sensor underwater acoustic target recognition.

[0027] In existing technologies, one approach proposes a Deformable Bayesian Network (DFBN) framework, utilizing a Dynamic Tree (DT) model to achieve multi-sensor data fusion. By using the DT model, measurement data from multiple sensing platforms are fused into a non-redundant representation. Corresponding features are extracted from seabed target images and corresponding seabed textures. The flexible structure of the DFBN is used to fuse common information from different sensors, enabling real-time fusion and recognition of sonar images. However, the Bayesian inference used in this approach is constrained by prior information, and the fusion result has a high initial probability dependence on multi-source information. Furthermore, this approach, based on sonar images, exhibits some bias. Another approach addresses the low recognition rate and high false alarm rate in underwater acoustic target recognition caused by a single target signal source by proposing a feature extraction and fusion method based on Long Short-Term Memory (LSTM) networks. By constructing a multi-layer LSTM model, multiple features such as the temporal envelope, DEMON line spectrum, and Mel-frequency cepstral coefficients of target noise are extracted. This leads to the establishment of a feature-level fusion recognition and classification model and a decision-level fusion model based on DS evidence theory. However, this scheme's identification result fusion is based on DS evidence theory. The evidence acquired by underwater sensors has significant uncertainty and a high probability of conflict, making it difficult to guarantee the reliability of the results. Another scheme extracts the short-time Fourier transform amplitude spectrum, phase spectrum, and bispectral features of the underwater acoustic signal to form the network input, and designs an integrated neural network to optimize and fuse the weight coefficients of the multi-feature input network. However, this scheme only acquires data from a single sensor, extracts different features, and then performs feature fusion, failing to handle data fusion from multiple sensors.

[0028] To address the shortcomings of existing technologies, this application proposes a multimodal underwater acoustic target recognition method. Based on underwater acoustic data from different sensor modes, a multimodal underwater acoustic target recognition model is constructed, and the features of the multimodal data are fused to effectively improve the underwater acoustic target recognition performance. This method can solve the following two technical problems: (1) How to fuse modal data from different sensors. This embodiment constructs a multimodal underwater acoustic target recognition model to achieve feature fusion of underwater acoustic data from different sensors, thus avoiding the situation where the reliability of each recognition result is insufficient when making a fusion strategy based on the recognition result.

[0029] (2) Improve the effect of multi-sensor underwater acoustic target recognition. Based on the constructed multi-modal underwater acoustic recognition model, the embodiments of this application propose a multi-stage training strategy, which can further improve the effect of multi-sensor underwater acoustic recognition.

[0030] For details, please refer to Figure 1 The multimodal underwater acoustic target recognition method specifically includes the following processes: S101. Data Acquisition: Collect underwater acoustic data from different sensors and divide it into unlabeled underwater acoustic data and labeled underwater acoustic data.

[0031] In this embodiment, training the multimodal underwater acoustic target recognition model requires a large amount of underwater acoustic data from different sensors, which is divided into unlabeled and labeled underwater acoustic data. Taking a sensor data volume of 3 as an example, each sensor collects 10 hours of audio data at a sampling rate of 16,000. One hour of manually labeled data is included, categorized into four types: large vessels, medium-sized vessels, small vessels, and marine environmental background noise. Finally, all audio data is segmented into 3-second samples, with each sample containing 48,000 sampling points. The final result includes 10,800 unlabeled underwater acoustic audio samples and 1,200 labeled underwater acoustic audio samples.

[0032] S102. Construction of a multimodal underwater acoustic target recognition model.

[0033] In this embodiment, the multimodal underwater acoustic target recognition model is mainly divided into a feature extraction network and a recognition network. The feature extraction network mainly consists of a series of convolutional layers, activation layers, and normalization layers, and its main objective is to extract feature maps from underwater acoustic data from multiple sensors through a neural network. The recognition network consists of multiple fully connected layers and is used to determine the type of underwater acoustic target based on the feature maps extracted by the feature extraction network.

[0034] In this embodiment, to further improve the model's accuracy in identifying underwater acoustic targets, the model training is divided into two stages: a pre-training stage and a supervised training stage, as detailed below: S103. Pre-training stage: The multimodal underwater acoustic target recognition model is pre-trained using unlabeled underwater acoustic data to determine the weights of the feature extraction network.

[0035] Please refer to Figure 2 In this embodiment, the pre-training stage uses unlabeled underwater acoustic data for training. This stage includes two parts: an online network and a target network. The online network comprises a feature extraction network, a projection network, and a prediction network, while the target network comprises a feature extraction network and a projection network. The feature extraction network and projection network in both the online and target networks have the same network structure. Furthermore, the feature extraction network structure is consistent with that in the multimodal underwater acoustic target recognition model. Both the projection network and the prediction network consist of multiple fully connected layers.

[0036] During specific training, sensors are randomly selected from unlabeled underwater acoustic data. The first underwater acoustic data of j Second underwater acoustic data In one embodiment, before inputting the first underwater acoustic data into the online network and the second underwater acoustic data into the target network, the method further includes randomly enhancing the first and second underwater acoustic data, and then feeding the enhanced data into a feature extraction network to extract features. Preferably, the random enhancement method includes pitch shifting, speed shifting, energy gain, random pruning, and padding.

[0037] Next, the first underwater acoustic data The data is fed into an online network, where a feature extraction network extracts features, which are then fed into a projection network and a prediction network to obtain predicted values. Its shape and size are [1, Second underwater acoustic data The data is fed into the target network, where features are extracted by the feature extraction network and then fed into the projection network to obtain the target value. Its shape and size are also [1, In this embodiment, the feature extraction network consists of four dilated convolutional layers with dilation coefficients of 1, 3, 5, and 7. Both the projection network and the prediction network consist of two fully connected layers. In this embodiment, the value is 512.

[0038] Finally, the predicted values ​​are constrained by the loss function. and target value The distance between them ensures that the target features extracted by the feature extraction network are highly similar when facing different sensors. In this embodiment, the loss function used is the Jensen-Shannon divergence, and its formula is:

[0039] in This represents the KL divergence (Kullback-Leibler Divergence).

[0040] It is important to note that during training, the gradient values ​​of the target network are not updated, and the weights of the feature extraction network in the online network are updated to the feature extraction network in the target network using a moving average method. These weights are then used as the initial weights for the supervised training phase. In this embodiment, the Adam optimizer is used in the pre-training phase, with a learning rate of... .

[0041] S104, Supervised Training Phase.

[0042] In the supervised training phase of this embodiment, the main focus is on fine-tuning the model. Please refer to... Figure 3In the supervised training phase, the weights of the feature extraction network obtained in the pre-training phase are used as the initial weights of the feature extraction network in the current multimodal underwater acoustic recognition model. Labeled underwater acoustic data from multiple sensors are input into the feature extraction network at this time to extract underwater acoustic data features from multiple sensors.

[0043] Then, after integrating the underwater acoustic data features from multiple sensors into the recognition network of the multimodal underwater acoustic recognition model, a loss function is used to calculate the loss value between the recognition result and the underwater acoustic target type label, thereby constraining the model update. In this embodiment, the loss function is cross-entropy loss. It is particularly important to note that the Adam optimizer will be used during the training phase, with a learning rate of... .

[0044] At this point, the trained multimodal underwater acoustic target recognition network is obtained.

[0045] S105, Target Recognition.

[0046] After obtaining the trained multimodal underwater acoustic target recognition model, real-time underwater acoustic data from different sensors is input into the model after the supervised training phase to obtain the target recognition result. Please refer to [reference needed]. Figure 4 The specific process is as follows: First, underwater acoustic data from multiple sensors is acquired. In this embodiment, underwater acoustic data from three sensors is used. For example, .

[0047] Then, feature maps of the underwater acoustic data from each sensor are extracted using the feature extraction network of the multimodal underwater acoustic recognition model. After passing through the feature extraction network, feature maps of the underwater acoustic data from the three sensors are extracted. Its shape is ,in The feature dimension is 512 in this embodiment.

[0048] Next, the feature maps of all sensor depth data are fused to obtain multi-sensor fused underwater acoustic features. In this embodiment, the fusion is performed by calculating the mean of the feature maps, that is:

[0049] Where D represents the multi-sensor fusion underwater acoustic features, and N represents the number of sensors.

[0050] Finally, the multi-sensor fused underwater acoustic features D are input into the recognition network of the multimodal underwater acoustic recognition model to determine the type of underwater acoustic target. In this embodiment, the recognition network consists of two fully connected layers, and the output of the recognition network is the probability of each category, i.e., the shape is... Where C represents the number of all target categories, which is 4 in this embodiment. The underwater acoustic target types can be identified based on the steps described above.

[0051] The above method was simulated and tested, and the following results were obtained: Table 1 Simulation Results

[0052] As can be seen, the multimodal underwater acoustic target recognition method in this application embodiment can effectively improve the recognition accuracy.

[0053] Please refer to Figure 5 According to one embodiment of this application, a system 200 includes a memory 201 and a processor 202. The memory 201 stores a computer program that can be loaded by the processor 202 and executed to correspond to the aforementioned multimodal underwater target recognition method. It should be noted that the electronic device also has a display screen for displaying a user interface (UI). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen is a touch screen, it also has the ability to collect touch signals on or above the surface of the display screen. The touch signals can be input to the processor as control signals for processing. In this case, the display screen can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, the display screen can be a single screen, the front panel of the electronic device; in other embodiments, there can be at least two screens, respectively disposed on different surfaces of the electronic device or in a folded design; in still other embodiments, the display screen can be a flexible screen, disposed on a curved surface or a folded surface of the electronic device. Furthermore, the display screen can also be configured as a non-rectangular irregular shape, i.e., a non-rectangular screen. The display screen can be made of materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0054] Figure 6 A schematic diagram of the structure of a computer system suitable for implementing the electronic device of the present application is shown.

[0055] It should be noted that, Figure 6 The computer system 300 of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0056] like Figure 6As shown, the computer system 300 includes a Central Processing Unit (CPU) 301, which can perform various appropriate actions and processes based on programs stored in Read-Only Memory (ROM) 302 or programs loaded from storage portion 308 into Random Access Memory (RAM) 303, such as performing the methods described in the above embodiments. The RAM 303 also stores various programs and data required for system operation. The CPU 301, ROM 302, and RAM 303 are interconnected via a bus 304. An Input / Output (I / O) interface 305 is also connected to the bus 304.

[0057] The following components are connected to I / O interface 305: an input section 306 including a keyboard, mouse, etc.; an output section 307 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 308 including a hard disk, etc.; and a communication section 309 including a network interface card such as a LAN (Local Area Network) card, modem, etc. The communication section 309 performs communication processing via a network such as the Internet. A drive 310 is also connected to I / O interface 305 as needed. Removable media 311, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 310 as needed so that computer programs read from them can be installed into storage section 308 as needed.

[0058] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 309, and / or installed from removable medium 311. When the computer program is executed by central processing unit (CPU) 301, it performs various functions defined in the system of this application.

[0059] It should be noted that the computer-readable medium shown in the embodiments of this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such transmitted data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.

[0060] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0061] The units described in the embodiments of this application can be implemented in software or hardware, and the described units can also be located in a processor. The names of these units do not necessarily limit the specific unit itself.

[0062] In another aspect, this application also provides a computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the multimodal underwater acoustic target recognition method described in the above embodiments.

[0063] In another aspect, this application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into the electronic device. The computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to implement the multimodal underwater acoustic target recognition method described in the above embodiments.

[0064] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0065] Through the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, touch terminal, or network device, etc.) to execute the method according to the embodiments of this application.

[0066] For those skilled in the art, the specific meanings of the above terms in this invention can be understood according to the specific circumstances; the accompanying drawings in the embodiments are used to clearly and completely describe the technical solutions in the embodiments of this invention. Obviously, the described embodiments are some embodiments of this invention, but not all embodiments. Generally, the components of the embodiments of this invention described and shown in the accompanying drawings can be arranged and designed in various different configurations.

[0067] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application.

Claims

1. A multimodal underwater acoustic target recognition method, characterized in that, include: Data acquisition: Collect underwater acoustic data from different sensors and divide it into unlabeled underwater acoustic data and labeled underwater acoustic data; Construction of a multimodal underwater acoustic target recognition model: The multimodal underwater acoustic target recognition model includes a feature extraction network and a recognition network; Pre-training phase: The multimodal underwater acoustic target recognition model is pre-trained using unlabeled underwater acoustic data to determine the weights of the feature extraction network; The pre-training phase specifically includes: randomly selecting underwater acoustic data from two sensors, namely, first underwater acoustic data and second underwater acoustic data, from unlabeled underwater acoustic data; inputting the first underwater acoustic data into an online network to obtain a predicted value; the online network includes a feature extraction network, a projection network, and a prediction network; inputting the second underwater acoustic data into a target network to obtain a target value; the target network includes a feature extraction network and a projection network; the feature extraction networks in the online network and the target network have the same structure as the feature extraction networks in the multimodal underwater acoustic target recognition model; both the projection network and the prediction network are composed of multiple fully connected layers; training is completed by constraining the distance between the predicted value and the target value in the online network and the target network using a loss function; during the training process, the gradient value of the target network is continuously updated, and the weights of the feature extraction network in the online network are updated to the feature extraction network of the target network using a moving average method; Supervised training phase: The weights of the feature extraction network in the pre-training phase are used as the initial weights of the feature extraction network in the current phase. Labeled underwater acoustic data is input into the feature extraction network to complete the update of the multimodal underwater acoustic target recognition model. Target recognition: The underwater acoustic data acquired in real time from different sensors is input into the multimodal underwater acoustic target recognition model that has completed the supervised training phase to obtain the target recognition result.

2. The multimodal underwater acoustic target recognition method according to claim 1, characterized in that, Before inputting the first underwater acoustic data into the online network and the second underwater acoustic data into the target network, the method further includes random augmentation of the first and second underwater acoustic data.

3. The multimodal underwater acoustic target recognition method according to claim 2, characterized in that, The random enhancements include pitch shifting, speed shifting, power gain, random clipping, and padding.

4. The multimodal underwater acoustic target recognition method according to claim 1, characterized in that, The supervised training phase specifically includes: The weights of the feature extraction network obtained in the pre-training stage are used as the initial weights of the feature extraction network in the current multimodal underwater acoustic recognition model. Labeled underwater acoustic data from multiple sensors are input into the feature extraction network to extract underwater acoustic data features. After integrating the underwater acoustic data features from multiple sensors into the recognition network of the multimodal underwater acoustic recognition model, the loss value between the recognition result and the underwater acoustic target type label is calculated using a loss function, thereby constraining the model update.

5. The multimodal underwater acoustic target recognition method according to claim 1, characterized in that, The target identification specifically includes: Acquire underwater acoustic data from multiple sensors; Feature maps of underwater acoustic data from each sensor are extracted using the feature extraction network of the multimodal underwater acoustic recognition model. The feature maps of water depth data from all sensors are fused to obtain multi-sensor fused underwater acoustic features. The underwater acoustic features fused from multiple sensors are input into the recognition network of a multimodal underwater acoustic recognition model to determine the type of underwater acoustic target.

6. The multimodal underwater acoustic target recognition method according to claim 5, characterized in that, The process of fusing the feature maps of all sensor water depth data to obtain multi-sensor fused underwater acoustic features includes: calculating the mean of the feature maps of all sensor water depth data as the multi-sensor fused underwater acoustic features.

7. The multimodal underwater acoustic target recognition method according to claim 1, characterized in that, Both the pre-training and supervised training phases employ the Adam optimizer.

8. A multimodal underwater acoustic target recognition system, characterized in that, It includes a memory and a processor, wherein the memory stores a computer program that can be loaded by the processor and executed corresponding to the multimodal underwater target recognition method as described in any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that, It stores computer program instructions, which, when executed by a processor, are used to implement the process corresponding to the multimodal underwater acoustic target recognition method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Underwater sound target identification method based on multi-modal fusion

    CN114420155A

  • Underwater sound target positioning method based on self-supervised learning

    CN115238783A