Underwater sound target identification method and device, related equipment and computer program product
The time-frequency characteristics of water acoustic signals are processed through the series residual network, and the shallow and deep characteristics are extracted and spliced, which solves the problem of difficulty in extracting water acoustic signals in the prior art, and improves the accuracy of water acoustic target recognition.
Patent Information
- Application Number
- CN202510333762.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art is difficult to effectively extract high-dimensional features in water acoustic signal processing, resulting in low accuracy of water acoustic target recognition results.
The time-frequency characteristics of the water acoustic signal are processed through more than two residual networks connected in series, and the shallow features extracted from the shallow residual network are spliced with the deep features extracted from the deep residual network to obtain the fusion features, and the water acoustic target is identified based on this.
The accuracy of the water acoustic target recognition results is improved, and local and global effective information of the water acoustic signal can be captured.
Smart Images

Figure CN120214768A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of sound signal processing, and more specifically, to an underwater acoustic target recognition method, device, related equipment and computer program product. Background Art
[0002] Underwater acoustic target recognition (detection) is an important part of the field of underwater acoustic signal processing, and the main research object is the radiation noise of underwater targets, including mechanical noise, propeller noise, hydrodynamic noise, etc.
[0003] Underwater acoustic target recognition, as an important means of modern ocean exploration, mainly includes two methods: based on traditional methods and based on deep learning. The traditional recognition method is currently commonly used, and the key lies in feature extraction and classifier design. Feature extraction mainly includes time-frequency features - time-frequency diagrams, auditory features - Mel frequency cepstrum coefficients (MFCC), etc.
[0004] The features extracted from underwater acoustic signals are usually modeled using neural networks. However, due to the different impacts of the depth of neural networks on features, it brings great difficulties to feature extraction and pattern recognition. Traditional neural network models generally extract high-dimensional feature representations through multi-level neural networks. As the number of network layers increases, the problem of gradient disappearance may occur. And for underwater acoustic signals in this specific scenario, the high-dimensional feature representations extracted by simply increasing the network levels cannot comprehensively capture the effective information in underwater acoustic signals, resulting in low accuracy of the final underwater acoustic target recognition results. Summary of the Invention
[0005] In view of the above problems, the present application is proposed to provide an underwater acoustic target recognition method, device, related equipment and computer program product to improve the accuracy of underwater acoustic target recognition results. The specific solutions are as follows:
[0006] In a first aspect, an underwater acoustic target recognition method is provided, including:
[0007] Obtain the time-frequency features of the underwater acoustic signal, where the underwater acoustic signal is the underwater acoustic signal collected from the target water area;
[0008] Process the time-frequency features through two or more cascaded residual networks, splice the shallow features extracted by the shallow residual network and the deep features extracted by the last residual network to obtain a fused feature;
[0009] Predict the recognition result of the underwater acoustic target based on the fused feature.
[0010] In a possible design, in another implementation manner of the first aspect of the embodiments of the present application, the process of splicing the shallow features extracted by the shallow residual network and the deep features extracted by the last residual network to obtain the fused features includes:
[0011] Send the shallow features extracted by the shallow residual network into the multi-scale feature selection layer, and the multi-scale feature selection layer extracts features of different scales from the shallow features respectively as multi-scale shallow features;
[0012] Splice the multi-scale shallow features and the deep features extracted by the last residual network to obtain the fused features.
[0013] In a possible design, in another implementation manner of the first aspect of the embodiments of the present application, the multi-scale feature selection layer includes two or more convolutional layers with different structural parameters, and different convolutional layers are used to extract features of different scales.
[0014] In a possible design, in another implementation manner of the first aspect of the embodiments of the present application, the process of the multi-scale feature selection layer extracting features of different scales from the shallow features respectively includes:
[0015] Each convolutional layer in the multi-scale feature selection layer processes the input shallow features respectively to obtain features of different scales output by each convolutional layer.
[0016] In a possible design, in another implementation manner of the first aspect of the embodiments of the present application, the process of the multi-scale feature selection layer extracting features of different scales from the shallow features respectively includes:
[0017] Each target convolutional layer in the multi-scale feature selection layer processes the input shallow features respectively to obtain features of different scales output by each target convolutional layer;
[0018] Wherein, the target convolutional layer is a specified convolutional layer adapted to the acquisition scenario of the underwater acoustic signal.
[0019] In a possible design, in another implementation manner of the first aspect of the embodiments of the present application, the process of processing the time-frequency features and predicting the recognition result of the underwater acoustic target based on the fused features is implemented by an underwater acoustic target recognition model, and the underwater acoustic target recognition model includes two or more residual networks connected in series, a feature fusion layer and a classification layer;
[0020] The two or more residual networks connected in series are used to process the time-frequency features of the input underwater acoustic signal to obtain the shallow features extracted by the shallow residual network and the deep features extracted by the last residual network;
[0021] The feature fusion layer is used to splice the shallow features and the deep features to obtain fused features;
[0022] The classification layer is used to predict the recognition result of the underwater acoustic target based on the fused features.
[0023] In a possible design, in another implementation manner of the first aspect of the embodiments of the present application, the process of obtaining the time-frequency features of the underwater acoustic signal includes:
[0024] Obtain the underwater acoustic signal collected from the target water area, and extract the LOFAR spectrogram of the underwater acoustic signal.
[0025] In a second aspect, an underwater acoustic target recognition device is provided, including:
[0026] A time-frequency feature acquisition unit, configured to acquire the time-frequency features of the underwater acoustic signal, where the underwater acoustic signal is the underwater acoustic signal collected from the target water area;
[0027] A feature extraction and fusion unit, configured to process the time-frequency features through two or more cascaded residual networks, splice the shallow features extracted by the shallow residual network and the deep features extracted by the last residual network to obtain fused features;
[0028] A target prediction unit, configured to predict the recognition result of the underwater acoustic target based on the fused features.
[0029] In a third aspect, an electronic device is provided, including: a memory and a processor;
[0030] The memory is used to store a program;
[0031] The processor is configured to execute the program to implement each step of the underwater acoustic target recognition method described in any one of the foregoing first aspects of the present application.
[0032] In a fourth aspect, a readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, each step of the underwater acoustic target recognition method described in any one of the foregoing first aspects of the present application is implemented.
[0033] In a fifth aspect, a computer program product is provided, including a computer program. When the computer program is executed by a processor, each step of the underwater acoustic target recognition method described in any one of the foregoing first aspects of the present application is implemented.
[0034] With the above technical solution, after obtaining the time-frequency characteristics of the underwater acoustic signal, the present application processes the time-frequency characteristics through two or more cascaded residual networks. The shallow residual network can extract shallow local and comprehensive feature information (shallow features) of the time-frequency characteristics, and the receptive field of the deep residual network is large enough to extract deep global information (deep features). Then, the shallow features and deep features are spliced and fused. The obtained fused features contain both fine-grained local information and overall global information, that is, they can capture the local and global effective information of the underwater acoustic signal. Based on this, the recognition result of the underwater acoustic target can be predicted more accurately based on the fused features, improving the recognition accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present application. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0036] Figure 1 It is a schematic diagram of an implementation system architecture of the underwater acoustic target recognition method provided by an embodiment of the present application;
[0037] Figure 2 It is a schematic diagram of a terminal structure provided by an embodiment of the present application;
[0038] Figure 3 It is a schematic diagram of a server structure provided by an embodiment of the present application;
[0039] Figure 4 It is a schematic diagram of the process of the underwater acoustic target recognition method provided by an embodiment of the present application;
[0040] Figure 5 It exemplifies a schematic diagram of the structure of an underwater acoustic target recognition model;
[0041] Figure 6 It exemplifies another schematic diagram of the structure of an underwater acoustic target recognition model;
[0042] Figure 7 It is a schematic diagram of the structure of an underwater acoustic target recognition device provided by an embodiment of the present application;
[0043] Figure 8 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0045] The present application provides an underwater acoustic target recognition method, which can be applied to a system architecture as shown in Figure 1 The system may include a terminal 100 and a server 200. The server 200 may include one or more servers ( Figure 1 In this example, it is described by taking one server as an example).
[0046] Either the terminal 100 or the server 200 can be used alone to execute the underwater acoustic target recognition method provided in the embodiments of the present application. In addition, the terminal 100 and the server 200 can also be used in cooperation to execute the underwater acoustic target recognition method provided in the embodiments of the present application.
[0047] Next, the product form of the terminal 100 will be described. Figure 1 in the product form of the terminal 100;
[0048] The terminal 100 in the embodiments of the present application may be a mobile phone, a tablet computer, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), etc. The embodiments of the present application do not make any restrictions on this.
[0049] Figure 2 shows an optional schematic diagram of the hardware structure of the terminal 100.
[0050] Referring to Figure 2 As shown, the terminal 100 may include a radio frequency unit 110, a memory 120, an input unit 130, a display unit 140, a camera 150 (optional), an audio circuit 160, a speaker 161, a microphone 162, a headphone jack 163 (optional), a processor 170, an external interface 180, a power supply 190, and other components. Those skilled in the art can understand that Figure 2 This is only an example of a terminal or a multifunctional device, and does not constitute a limitation on the terminal or the multifunctional device. It may include more or fewer components than those shown in the figure, or combine some components, or different components.
[0051] The input unit 130 can be used to receive input numerical or character information and generate key signal inputs related to the user settings and function control of the portable multifunctional device. Specifically, the input unit 130 can include a touch screen 131 and / or other input devices 132. The touch screen can detect a user's touch action on the touch screen, convert the touch action into a touch signal and send it to the processor 170, and can receive and execute commands sent by the processor 170; the touch signal at least includes contact coordinate information. The touch screen 131 can provide an input interface and an output interface between the terminal 100 and the user. In addition to the touch screen 131, the input unit 130 can also include other input devices. Specifically, the other input devices 132 can include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, power on / off keys, etc.), a trackball, a mouse, a joystick, etc.
[0052] The display unit 140 can be used to display information input by the user or information provided to the user, various menus of the terminal 100, an interactive interface, file display, and / or the playback of any multimedia file. In the embodiments of the present application, the display unit 140 can be used to display each interactive interface, processing result, etc. in the underwater acoustic target recognition method.
[0053] The memory 120 can be used to store instructions and data.
[0054] The processor 170 is the control center of the terminal 100, connects various parts of the entire terminal 100 using various interfaces and lines, and by running or executing instructions stored in the memory 120 and calling data stored in the memory 120, executes various functions of the terminal 100 and processes data, thereby performing overall control of the terminal device.
[0055] Among them, the memory 120 can be used to store software codes related to the underwater acoustic target recognition method, and the processor 170 can execute the steps of the underwater acoustic target recognition method or can also schedule other units (such as the above-mentioned input unit 130 and display unit 140) to implement corresponding functions.
[0056] The radio frequency unit 110 (optional) can be used to receive and send information or signals during a call. For example, after receiving the downlink information of the base station, it is given to the processor 170 for processing; in addition, it sends uplink data to the base station. In the embodiments of the present application, the radio frequency unit 110 can send data to the server 200 and receive the processing result sent by the server 200. Exemplarily, the radio frequency unit 110 receives the underwater acoustic signals collected by the sonar device and forwards them to the server 200. After the server 200 obtains the underwater acoustic signals, it performs recognition processing to obtain the recognition result of the underwater acoustic target and returns the recognition result to the terminal 100.
[0057] It should be understood that the radio frequency unit 110 is optional and can be replaced by other communication interfaces, such as a network interface.
[0058] Although not shown, the terminal 100 may further include a flash, a wireless fidelity (WiFi) module, a Bluetooth module, sensors with different functions, etc., which will not be elaborated here. Some or all of the methods described below can be applied to the Figure 2 terminal 100 as shown.
[0059] Next, the product form of the Figure 1 server 200 will be described;
[0060] Figure 3 A schematic structural diagram of a server 200 is provided. As Figure 3 shown, the server 200 includes a bus 201, a processor 202, a communication interface 203, and a memory 204. The processor 202, the memory 204, and the communication interface 203 communicate with each other through the bus 201.
[0061] The bus 201 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, Figure 3 only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0062] The processor 202 can be any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP), etc.
[0063] The memory 204 may include volatile memory, such as random access memory (RAM). The memory 204 may also include non-volatile memory, such as read-only memory (ROM), flash memory, a hard disk drive (HDD), or a solid state drive (SSD).
[0064] Among them, the memory 204 can be used to store software codes related to the underwater acoustic target recognition method, and the processor 202 can execute the steps of the underwater acoustic target recognition method or schedule other units to implement corresponding functions.
[0065] It should be understood that the above-mentioned terminal 100 and server 200 can be centralized or distributed devices, and the processors in the above-mentioned terminal 100 and server 200 (such as processor 170 and processor 202) can be hardware circuits (such as application specific integrated circuit (ASIC), field-programmable gate array (FPGA), general processor, digital signal processing (DSP), microprocessor or microcontroller, etc.), or a combination of these hardware circuits. For example, the processor can be a hardware system with the function of executing instructions, such as CPU, DSP, etc., or a hardware system without the function of executing instructions, such as ASIC, FPGA, etc., or a combination of the above-mentioned hardware system without the function of executing instructions and the hardware system with the function of executing instructions.
[0066] An embodiment of the present application provides an underwater acoustic target recognition method. Taking the application of this method to a computer device as an example, the computer device can specifically be Figure 1 the terminal 100 in or a system composed of the terminal 100 and the server 200. Referring to Figure 4 , the underwater acoustic target recognition method specifically includes the following steps:
[0067] Step S100: Obtain the time-frequency characteristics of the underwater acoustic signal, where the underwater acoustic signal is an underwater acoustic signal collected from the target water area.
[0068] The target water area is the water area to be monitored. The underwater acoustic signal of the target water area can be collected through sensors such as sonar devices. Among them, the sensors can be deployed in the target water area and send the collected underwater acoustic signals to the computer device in a wired or wireless form for subsequent recognition and processing. Or, the sensors can also be installed on water platforms such as ships to realize the collection of underwater acoustic signals of the target water area.
[0069] For subsequent recognition and processing of the underwater acoustic signal, the time-frequency characteristics of the underwater acoustic signal can be obtained in this step. The time-frequency characteristics refer to the signal characteristics obtained through time-frequency analysis, including but not limited to: LOFAR (Low Frequency Analysis Recording) spectrogram, characteristics of time-frequency transformation (STFT, wavelet, HHT, etc.).
[0070] Step S110: Process the time-frequency features through two or more residual networks connected in series, and concatenate the shallow features extracted by the shallow residual network with the deep features extracted by the last residual network to obtain fused features.
[0071] In this embodiment, in view of the fact that the shallow information of the time-frequency features extracted from the underwater acoustic signal in the neural network is comprehensive enough and the deep receptive field is large enough, a network model is designed that uses residual splicing to connect the shallow and deep features respectively.
[0072] Specifically, the input time-frequency features are processed by N residual networks connected in series. N is an integer greater than or equal to 2. In some possible implementations, experiments have found that when N is 4, a better recognition effect can be obtained.
[0073] The first layer of the N residual networks in series is defined as a shallow residual network, and the last layer is defined as a deep residual network. Generally, the number of layers of a shallow residual network does not exceed N / 2. Usually, the first or second residual network can be regarded as a shallow residual network, and the last residual network can be regarded as a deep residual network.
[0074] In each residual network connected in series, each residual network processes the features output by the previous layer in turn. The features output by the shallow residual network are defined as shallow features, and the features output by the deep residual network (such as the last residual network) are defined as deep features.
[0075] Shallow residual networks can extract shallow local and comprehensive feature information, which can also be called fine-grained local information. The receptive field of deep residual networks is large enough to extract deep global information.
[0076] In order to capture more comprehensive and effective information in the underwater acoustic signal in this embodiment, the extracted shallow features and deep features can be spliced and fused to obtain fused features, thereby ensuring that both local and global effective information can be monitored.
[0077] Step S120: predicting the recognition result of the hydroacoustic target based on the fusion feature.
[0078] The recognition result of the hydroacoustic target is to identify the type of the hydroacoustic target, such as mechanical noise, propeller noise, hydrodynamic noise or other types. The type label set can be preset by the user.
[0079] In this step, since the fused features cover both local and global effective information, the accuracy of the recognition results of the underwater acoustic target can be improved based on the fused features.
[0080] The underwater acoustic target recognition method provided by the embodiments of the present application processes the time-frequency features of the underwater acoustic signal through two or more cascaded residual networks after obtaining the time-frequency features. The shallow residual network can extract shallow local and comprehensive feature information (shallow features) of the time-frequency features, and the receptive field of the deep residual network is large enough to extract deep global information (deep features). Then, the shallow features and deep features are spliced and fused, and the obtained fusion features contain both fine-grained local information and overall global information, that is, they can capture the local and global effective information of the underwater acoustic signal. On this basis, the recognition result of the underwater acoustic target can be predicted more accurately based on the fusion features, improving the recognition accuracy.
[0081] Refer to Figure 5 As shown, an optional structure of an underwater acoustic target recognition model is provided.
[0082] The underwater acoustic target recognition model may include N cascaded residual networks Resnet ( Figure 5 Taking a total of 4 residual networks Resnet_1 - Resnet_4 as an example for illustration), a feature fusion layer, and a classification layer.
[0083] The time-frequency features of the obtained underwater acoustic signal ( Figure 5 Taking the LOFAR spectrogram as an example) are sent into the first residual network Resnet_1 to extract shallow features.
[0084] On the one hand, the first residual network Resnet_1 continues to pass the extracted shallow features backward, and after being processed by each layer of the residual network, the deep features output by the last residual network Resnet_4 are obtained, denoted as Embedding_2, and the deep features are sent into the feature fusion layer.
[0085] On the other hand, the shallow features are also sent into the feature fusion layer. Figure 5 In the network structure shown, a convolutional layer CNN can be further added, and the shallow features can be processed by the convolutional layer first to make the number of channels of the processed shallow features consistent with that of the deep features, which is convenient for feature splicing processing in the feature fusion layer.
[0086] The shallow features sent into the feature fusion layer are denoted as Embedding_1.
[0087] The feature fusion layer splices (concat) the shallow features Embedding_1 and the deep features Embedding_2 to obtain fusion features.
[0088] The classification layer predicts the recognition result of the underwater acoustic target based on the fusion features.
[0089] Figure 5The classification layer shown can adopt a structure in which a fully connected layer FC is connected in series with a Softmax layer, and can predict the classification labels of underwater acoustic targets.
[0090] During the training process of the underwater acoustic target recognition model provided in the above embodiment, the time-frequency features of the sample underwater acoustic signals can be used as training samples, and the known types of the underwater acoustic targets corresponding to the sample underwater acoustic signals can be used as sample labels for training. Given this structure of the model, the shallow residual network of the model can focus on fine-grained local information, and the deep layer can focus on the overall global information. The splicing and fusion of them ensure that both local and global features can be monitored, thus effectively improving the accuracy of the model in underwater acoustic target recognition.
[0091] In some embodiments of the present application, another implementation process of the underwater acoustic target recognition method is provided. For the foregoing step S110, the process of splicing the shallow features extracted by the shallow residual network and the deep features extracted by the last residual network to obtain the fused features can be implemented in the following manner:
[0092] Send the shallow features extracted by the shallow residual network into the multi-scale feature selection layer, and the multi-scale feature selection layer extracts features of different scales from the shallow features respectively as multi-scale shallow features.
[0093] Splice the multi-scale shallow features and the deep features extracted by the last residual network to obtain the fused features.
[0094] In this embodiment, the process of extracting multi-scale features from the shallow features by using the multi-scale feature selection layer is further added to obtain multi-scale shallow features. Then, the multi-scale shallow features and the deep features are spliced to obtain the fused features.
[0095] The multi-scale feature selection layer has the ability to extract several features of different scales from the input shallow features respectively.
[0096] Multi-scale features refer to extracting features from the input data through different scales (such as different spatial resolutions, different receptive field ranges) to capture semantic or structural information at different levels.
[0097] Considering that in different acquisition scenarios (different sonar device models and different water area environments), the distribution frequency bands of the effective information in the time-frequency characteristics of the acquired underwater acoustic signals may also be different. In order to extract more comprehensive effective information in the shallow features, in this embodiment, a multi-scale feature selection layer can be used to extract multi-scale features from the shallow features, so as to obtain multi-scale shallow features, which can contain more comprehensive effective information compared with the shallow features directly extracted by the shallow residual network. On this basis, the multi-scale shallow features are concatenated with the deep features to ensure the effectiveness of the fused features and further improve the accuracy of the predicted recognition results.
[0098] Among them, the multi-scale shallow features can include shallow features of multiple different scales. The shallow features of multiple different scales can be first concatenated together and then concatenated with the deep features to obtain the fused features. Or, the shallow features of multiple different scales can also be concatenated with the deep features together to obtain the fused features.
[0099] Combined with Figure 6 , an alternative structure of another underwater acoustic target recognition model is provided.
[0100] Compared with Figure 5 the model structure shown, Figure 6 a multi-scale feature selection layer is added to the model shown.
[0101] In a possible implementation, the multi-scale feature selection layer can include more than two convolutional networks with different structural parameters, and different convolutional layers are used to extract features of different scales.
[0102] Among them, the structural parameters of the convolutional network can include the convolutional kernel and the stride. Different combinations of convolutional kernels and strides can generate multi-scale features.
[0103] Based on the multi-scale feature selection layer including more than two convolutional layers with different structural parameters, in the above steps, the process of extracting features of different scales from the shallow features by the multi-scale feature selection layer can be implemented in any of the following ways:
[0104] The first:
[0105] Each convolutional layer in the multi-scale feature selection layer processes the input shallow features respectively to obtain features of different scales output by each convolutional layer.
[0106] Exemplarily, if the number of convolutional layers with different structural parameters included in the multi-scale feature selection layer is m, then by extracting one scale of shallow features through each convolutional layer respectively, m scales of shallow features can be obtained. Concatenate these m scales of shallow features together as Figure 6 the Embedding_1 shown.
[0107] By setting multiple convolutional layers with different structural parameters in the multi-scale feature selection layer, shallow features of multiple different scales can be extracted, ensuring that the shallow features can cover more comprehensive and effective information.
[0108] The second method:
[0109] Each target convolutional layer in the multi-scale feature selection layer processes the input shallow features respectively to obtain features of different scales output by each target convolutional layer. After that, the shallow features of different scales output by each target convolutional layer can be concatenated as Figure 6 Embedding_1 shown in
[0110] Among them, the target convolutional layer is a specified convolutional layer adapted to the acquisition scenario of underwater acoustic signals.
[0111] Specifically, several convolutional layers with different structural parameters can be set in the multi-scale feature selection layer to extract shallow features of different scales. During each recognition process, for the current acquisition scenario (such as environmental parameters such as the model of sonar equipment, water temperature, water flow velocity, salinity, etc.), a convolutional layer with structural parameters adapted to the current acquisition scenario can be selected as the target convolutional layer, and then the shallow features are processed by the target convolutional layer to obtain multi-scale shallow features extracted by each target convolutional layer.
[0112] Exemplarily, the user can obtain through experiments or based on expert experience in advance which convolutional layers with which structural parameters can extract more comprehensive and effective feature information under different acquisition scenarios. On this basis, the corresponding relationship between different acquisition scenario parameters and the selected target convolutional layer can be recorded. Each time an underwater acoustic signal is recognized, according to the acquisition scenario parameters of the current underwater acoustic signal, the corresponding relationship can be queried to determine the corresponding target convolutional layer, and then each target convolutional layer in the multi-scale feature selection layer processes the input shallow features respectively to obtain multi-scale shallow features output by each target convolutional layer and send them to the feature fusion layer.
[0113] In the solution provided in this embodiment, by pre-summarizing the target convolutional layers adapted to different acquisition scenarios, the target convolutional layer adapted to the current acquisition scenario can be located, and then each target convolutional layer in the multi-scale feature selection layer is used to extract multi-scale features. It can dynamically adjust the target convolutional layer following the change of the acquisition scenario, ensuring that more comprehensive and effective shallow feature information can be collected under various different acquisition scenarios, thereby improving the accuracy of subsequent recognition results.
[0114] Next, the underwater acoustic target recognition device provided in the embodiments of the present application will be described. The underwater acoustic target recognition device described below can be mutually referred to corresponding to the underwater acoustic target recognition method described above.
[0115] See Figure 7 , Figure 7 which is a schematic structural diagram of an underwater acoustic target recognition device disclosed in an embodiment of the present application.
[0116] As Figure 7 shown, the device may include:
[0117] A time-frequency feature acquisition unit 11, configured to acquire the time-frequency features of an underwater acoustic signal, where the underwater acoustic signal is an underwater acoustic signal collected from a target water area;
[0118] A feature extraction and fusion unit 12, configured to process the time-frequency features through two or more cascaded residual networks, and splice the shallow features extracted by the shallow residual network and the deep features extracted by the last residual network to obtain fused features;
[0119] A target prediction unit 13, configured to predict the recognition result of the underwater acoustic target based on the fused features.
[0120] In a possible implementation, the process of the feature extraction and fusion unit splicing the shallow features extracted by the shallow residual network and the deep features extracted by the last residual network to obtain fused features includes:
[0121] Feeding the shallow features extracted by the shallow residual network into a multi-scale feature selection layer, and respectively extracting features of different scales from the shallow features through the multi-scale feature selection layer as multi-scale shallow features;
[0122] Splicing the multi-scale shallow features and the deep features extracted by the last residual network to obtain fused features.
[0123] In a possible implementation, the multi-scale feature selection layer includes two or more convolutional layers with different structural parameters, and different convolutional layers are used to extract features of different scales.
[0124] In a possible implementation, the process of the feature extraction and fusion unit respectively extracting features of different scales from the shallow features through the multi-scale feature selection layer includes:
[0125] Processing the input shallow features respectively through each convolutional layer in the multi-scale feature selection layer to obtain features of different scales output by each convolutional layer.
[0126] In another possible implementation, the process of the feature extraction and fusion unit respectively extracting features of different scales from the shallow features through the multi-scale feature selection layer includes:
[0127] The input shallow features are respectively processed by each target convolutional layer in the multi-scale feature selection layer to obtain features of different scales output by each target convolutional layer;
[0128] Among them, the target convolutional layer is a specified convolutional layer adapted to the acquisition scenario of the underwater acoustic signal.
[0129] In a possible implementation, the process of the time-frequency feature acquisition unit for acquiring the time-frequency features of the underwater acoustic signal includes:
[0130] Acquire the underwater acoustic signal collected from the target water area, and extract the LOFAR spectrogram of the underwater acoustic signal.
[0131] In a possible implementation, the processing processes of the above-mentioned feature extraction and fusion unit and the target prediction unit are implemented by invoking an underwater acoustic target recognition model, and the underwater acoustic target recognition model includes two or more residual networks connected in series, a feature fusion layer, and a classification layer;
[0132] The two or more residual networks connected in series are used to process the time-frequency features of the input underwater acoustic signal to obtain shallow features extracted by the shallow residual network and deep features extracted by the last residual network;
[0133] The feature fusion layer is used to splice the shallow features and the deep features to obtain fused features;
[0134] The classification layer is used to predict the recognition result of the underwater acoustic target based on the fused features.
[0135] In the embodiments of the present application, an electronic device is further provided. Refer to Figure 8 As shown, it shows a schematic structural diagram of an electronic device suitable for implementing the electronic device in the embodiments of the present application. The electronic device in the embodiments of the present application may include, but is not limited to, fixed terminals such as mobile phones, tablet computers, and the like. Figure 8 The electronic device shown is only an example, and should not bring any limitation to the functions and usage ranges of the embodiments of the present application.
[0136] As Figure 8As shown in the figure, the electronic device may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 601, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603, so as to implement the underwater acoustic target recognition method in the foregoing embodiments of the present application. When the electronic device is powered on, various programs and data required for the operation of the electronic device are also stored in the RAM 603. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0137] Generally, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a memory card, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 8 an electronic device with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.
[0138] In an embodiment of the present application, there is also provided a computer program product including computer-readable instructions, which, when running on an electronic device, enable the electronic device to implement any of the underwater acoustic target recognition methods provided in the embodiments of the present application.
[0139] In an embodiment of the present application, there is also provided a computer-readable storage medium carrying one or more computer programs, which, when executed by an electronic device, can enable the electronic device to implement any of the underwater acoustic target recognition methods provided in the embodiments of the present application.
[0140] In addition, it should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the drawings of the device embodiments provided in the present application, the connection relationship between the modules indicates that they have a communication connection, which can be specifically implemented as one or more communication buses or signal lines.
[0141] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general hardware. Of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions accomplished by computer programs can easily be implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be various, such as analog circuits, digital circuits, or dedicated circuits. However, for this application, software program implementation is a better embodiment in more cases. Based on such an understanding, the technical solution of this application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disc of a computer, and includes several instructions for causing a computer device (which can be a personal computer, training device, or network device, etc.) to execute the methods described in various embodiments of this application.
[0142] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.
[0143] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of this application are generated in whole or in part. The computer can be a general-purpose computer, a dedicated computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, training device, or data center to another website, computer, training device, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can store, or a data storage device such as a training device or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).
[0144] The various embodiments in this specification are described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can be referred to each other.
Claims
1. A method for underwater acoustic target recognition, characterized in that: include: Acquiring time-frequency characteristics of a hydroacoustic signal, wherein the hydroacoustic signal is a hydroacoustic signal collected from a target water area; The time-frequency features are processed by two or more residual networks connected in series, and the shallow features extracted by the shallow residual network are spliced with the deep features extracted by the last residual network to obtain fused features; The recognition result of the hydroacoustic target is predicted based on the fusion feature.
2. The method according to claim 1, characterized in that The process of splicing the shallow features extracted by the shallow residual network with the deep features extracted by the last residual network to obtain fused features includes: The shallow features extracted by the shallow residual network are sent to the multi-scale feature selection layer, and the multi-scale feature selection layer is used to extract features of different scales from the shallow features as multi-scale shallow features; The multi-scale shallow features are concatenated with the deep features extracted by the last residual network to obtain fused features.
3. The method according to claim 2, characterized in that The multi-scale feature selection layer includes more than two convolutional layers with different structural parameters, and different convolutional layers are used to extract features of different scales.
4. The method according to claim 3, characterized in that The process of extracting features of different scales from the shallow features through the multi-scale feature selection layer includes: Each convolutional layer in the multi-scale feature selection layer processes the shallow features of the input respectively to obtain features of different scales output by each convolutional layer.
5. The method according to claim 3, characterized in that: The process of extracting features of different scales from the shallow features through the multi-scale feature selection layer includes: Each target convolution layer in the multi-scale feature selection layer processes the shallow features of the input respectively, and obtains features of different scales output by each target convolution layer; The target convolution layer is a specified convolution layer adapted to the acquisition scenario of the underwater acoustic signal.
6. The method according to claim 1, characterized in that The process of processing the time-frequency features and predicting the recognition result of the hydroacoustic target based on the fusion features is implemented by a hydroacoustic target recognition model, wherein the hydroacoustic target recognition model includes more than two residual networks, a feature fusion layer and a classification layer connected in series; The two or more residual networks connected in series are used to process the time-frequency characteristics of the input underwater acoustic signal to obtain shallow features extracted by the shallow residual network and deep features extracted by the last residual network; The feature fusion layer is used to splice the shallow features and the deep features to obtain fused features; The classification layer is used to predict the recognition result of the hydroacoustic target based on the fusion features.
7. The method according to any one of claims 1 to 6, characterized in that: The process of obtaining the time-frequency characteristics of underwater acoustic signals includes: Acquire the hydroacoustic signal collected from the target water area, and extract the LOFAR spectrogram of the hydroacoustic signal.
8. An underwater acoustic target recognition device, characterized in that: include: A time-frequency feature acquisition unit, used to acquire the time-frequency features of a hydroacoustic signal, wherein the hydroacoustic signal is a hydroacoustic signal collected from a target water area; A feature extraction and fusion unit, used to process the time-frequency features through two or more residual networks connected in series, and to concatenate the shallow features extracted by the shallow residual network with the deep features extracted by the last residual network to obtain fused features; A target prediction unit is used to predict the recognition result of the hydroacoustic target based on the fusion feature.
9. An electronic device, characterized in that: include: Memory and processor; The memory is used to store programs; The processor is used to execute the program to implement each step of the underwater acoustic target recognition method as described in any one of claims 1 to 7.
10. A readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, each step of the underwater acoustic target recognition method as described in any one of claims 1 to 7 is implemented.
11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, each step of the underwater acoustic target recognition method as described in any one of claims 1 to 7 is implemented.