Underwater acoustic signal noise reduction method and device, computer equipment and readable storage medium
The local and global attention mechanism of the target model is used to process the water acoustic signals, which solves the problem of noise interference of water acoustic signals in the underwater environment, and achieves efficient and accurate reconstruction of the signals.
Patent Information
- Application Number
- CN202510458340.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-04-11
AI Technical Summary
Water acoustic signals are disturbed by noise when propagating in underwater environments, resulting in reduced signal clarity and reliability, affecting the accurate reception and interpretation of the signal.
The target model is used for feature extraction, and the local and global noise characteristics are calculated through the attention mechanism of the department-controlled circulation unit and the attention mechanism of the global gated circulation unit, and the noise mask matrix is generated for noise filtering, and the target waveform signal is reconstructed.
It improves the accuracy of water acoustic signal reconstruction, effectively removes noise, retains signal characteristics, and avoids the problem of phase and amplitude decoupling in traditional methods.
Smart Images

Figure CN120496487A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of underwater acoustic signal processing, and in particular to an underwater acoustic signal noise reduction method, apparatus, computer equipment, and readable storage medium. Background Art
[0002] Hydroacoustic signals are sound wave signals that propagate underwater. They can be used in a variety of applications, including underwater communication, detection, navigation, and monitoring, such as underwater robots, submarines, and sonar systems. They are crucial for improving the intelligence of underwater systems, promoting marine scientific research, and protecting the marine environment. However, due to the complex underwater environment, hydroacoustic signals are often interfered with by various noises during propagation, such as water currents, marine life, and the noise of the equipment itself. This noise can reduce the clarity and reliability of the signal, hindering its accurate reception and interpretation. Therefore, noise reduction processing of hydroacoustic signals is necessary to ensure the efficiency and accuracy of underwater communication and detection systems.
[0003] In related technologies, a continuous time-domain waveform is generally segmented into multiple segments with short time intervals through a short-time Fourier transform (STFT), and each segment is Fourier transformed to convert the time-domain signal into a time-frequency diagram for signal reconstruction. However, when using the STFT to process underwater acoustic signals, an algorithm is required to estimate and reconstruct the phase. Factors such as multipath propagation, signal scattering, reverberation, and noise interference in the underwater acoustic environment may make the phase information extracted from the noise background inaccurate, resulting in phase estimation errors and, in turn, inaccurate reconstructed waveform signals. Summary of the Invention
[0004] The main purpose of the embodiments of the present application is to propose a method, device, computer equipment and readable storage medium for underwater acoustic signal noise reduction, which can improve the accuracy of the reconstructed waveform signal.
[0005] To achieve the above objectives, a first aspect of an embodiment of the present application provides a method for reducing noise of an underwater acoustic signal, the method comprising:
[0006] Acquiring a time domain signal of the underwater sound to be de-noised, and performing feature extraction on the time domain signal to be de-noised using a target model to obtain a feature tensor to be de-noised, wherein the feature tensor to be de-noised includes a plurality of local feature blocks;
[0007] By using the local gating recurrent unit attention mechanism of the target model, local position-dependent features and local time-dependent features between each local feature block and adjacent feature blocks in the feature tensor to be denoised are calculated in parallel to obtain local noise features;
[0008] Calculating the global position-dependent features and the global time-dependent features of the feature tensor to be denoised through the global gated recurrent unit attention mechanism of the target model to obtain a global noise feature;
[0009] Reconstructing a noise time domain signal of a target noise from the time domain signal to be denoised based on a noise mask matrix generated by fusing the local noise features and the global noise features through a decoder of the target model;
[0010] Based on the noise time domain signal, the time domain signal to be denoised is subjected to noise filtering to obtain a target waveform signal.
[0011] Accordingly, a second aspect of an embodiment of the present application provides an underwater acoustic signal noise reduction device, the device comprising:
[0012] an acquisition module, configured to acquire a time domain signal of the underwater sound to be de-noised, and perform feature extraction on the time domain signal to be de-noised using a target model to obtain a feature tensor to be de-noised, wherein the feature tensor to be de-noised includes a plurality of local feature blocks;
[0013] A first computing module is configured to parallelly compute local position-dependent features and local time-dependent features between each local feature block and adjacent feature blocks in the feature tensor to be denoised through an attention mechanism of a local gated recurrent unit of the target model to obtain local noise features;
[0014] A second computing module is configured to calculate the global position-dependent features and the global time-dependent features of the feature tensor to be denoised through the global gated recurrent unit attention mechanism of the target model to obtain a global noise feature;
[0015] A reconstruction module, configured to reconstruct a noise time domain signal of a target noise from the time domain signal to be denoised based on a noise mask matrix generated by fusing the local noise features and the global noise features through a decoder of the target model;
[0016] The filtering module is used to perform noise filtering on the time domain signal to be denoised based on the noise time domain signal to obtain a target waveform signal.
[0017] In some embodiments, the first computing module is further configured to:
[0018] By using each local attention head of the local gating recurrent unit attention mechanism of the target model, local time dependency features between each local feature block and adjacent feature blocks in the feature tensor to be denoised are calculated in parallel, thereby obtaining local temporal features corresponding to the multiple local feature blocks;
[0019] Concatenate multiple local time series features corresponding to multiple local attention heads to obtain a local time series feature tensor;
[0020] The local position-dependent features of the local temporal feature tensor are calculated by the local gating layer of the local gating recurrent unit attention mechanism to obtain local noise features.
[0021] In some embodiments, the first computing module is further configured to:
[0022] Fusing the local time series feature tensor and the feature tensor to be denoised to obtain a first feature;
[0023] calculating a local position-dependent feature of the first feature through a local gating layer of the local gating recurrent unit attention mechanism to obtain a second feature;
[0024] Performing a nonlinear transformation on the second feature through a leaky linear rectification function of the local gated recurrent unit attention mechanism to obtain a third feature;
[0025] The first feature and the third feature are fused to obtain a local noise feature.
[0026] In some embodiments, the second computing module is further configured to:
[0027] Calculating the global time-dependent features of the feature tensor to be denoised through each global attention head of the global gated recurrent unit attention mechanism of the target model, and obtaining the global temporal features corresponding to each global attention head;
[0028] Concatenate multiple global temporal features corresponding to multiple global attention heads to obtain a global temporal feature tensor;
[0029] The global position-dependent features of the global temporal feature tensor are calculated through the global gating layer of the global gated recurrent unit attention mechanism to obtain the global noise features.
[0030] In some embodiments, the underwater acoustic signal noise reduction device further includes a training module for:
[0031] Acquire a sample time domain signal of the underwater sound to be de-noised, and perform feature extraction on the sample time domain signal using a preset model to obtain a sample time domain signal tensor, wherein the sample time domain signal tensor includes a plurality of sample local feature blocks;
[0032] By using the local gating recurrent unit attention mechanism of the preset model, the sample local position dependence feature and the sample local time dependence feature between each sample local feature block and the adjacent sample feature blocks in the sample time domain signal tensor are calculated in parallel to obtain the sample local noise feature;
[0033] Calculating the sample global position-dependent feature and the sample global time-dependent feature of the sample time-domain signal tensor through the global gated recurrent unit attention mechanism of the preset model to obtain the sample global noise feature;
[0034] Reconstructing a predicted noise time domain signal of the sample noise from the sample time domain signal through a decoder of the preset model based on a sample noise mask matrix generated by fusing the sample local noise feature and the sample global noise feature;
[0035] Based on the predicted noise time domain signal, filtering the sample time domain signal to obtain a predicted waveform signal;
[0036] determining a target loss based on the predicted noise time domain signal and the predicted waveform signal;
[0037] The preset model is trained according to the target loss to obtain a target model.
[0038] In some embodiments, the training module is further configured to:
[0039] Obtaining a sample noise time domain signal and a sample waveform signal corresponding to the sample water sound to be de-noised;
[0040] Calculating a corresponding first similarity based on the predicted noise time domain signal and the sample noise time domain signal;
[0041] Calculating a corresponding second similarity based on the predicted waveform signal and the sample waveform signal;
[0042] A target loss is determined based on the first similarity and the second similarity.
[0043] In some embodiments, the acquisition module is further configured to:
[0044] Extracting features of the time domain signal to be denoised using a convolutional neural network of a target model to obtain corresponding features to be denoised;
[0045] Dividing the feature to be denoised into a plurality of local feature blocks according to a preset division length and division hop number;
[0046] The multiple local feature blocks are spliced to obtain a feature tensor to be denoised.
[0047] Correspondingly, the third aspect of the embodiments of the present application proposes a computer device, which includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it implements the underwater acoustic signal noise reduction method of any one of the embodiments of the first aspect of the present application.
[0048] Correspondingly, the fourth aspect of the embodiments of the present application proposes a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the underwater acoustic signal noise reduction method of any one of the embodiments of the first aspect of the present application.
[0049] The present application obtains a time domain signal of the underwater sound to be de-noised, and extracts features of the time domain signal to be de-noised through a target model to obtain a feature tensor to be de-noised, wherein the feature tensor to be de-noised includes multiple local feature blocks; through the local gated recurrent unit attention mechanism of the target model, the local position-dependent features and local time-dependent features between each local feature block in the feature tensor to be de-noised and the adjacent feature blocks are calculated in parallel to obtain local noise features; through the global gated recurrent unit attention mechanism of the target model, the global position-dependent features and global time-dependent features of the feature tensor to be de-noised are calculated to obtain global noise features; through the decoder of the target model, the noise mask matrix generated by the fusion of the local noise features and the global noise features is used to reconstruct the noise time domain signal of the target noise from the time domain signal to be de-noised; based on the noise time domain signal, the time domain signal to be de-noised is noise filtered to obtain a target waveform signal. In this way, by combining the local gated recurrent unit attention mechanism and the global gated recurrent unit attention mechanism, the model can simultaneously capture local and global dependencies in the sequence data, and can accurately capture transient noise and short-term signal details through the local gated recurrent unit attention mechanism, and can also model long-term dependencies through the global gated recurrent unit attention mechanism to accurately identify and utilize complex patterns in the signal, thereby more effectively retaining the characteristics of the target signal during the noise reduction process. At the same time, the present application directly reconstructs a clear signal waveform in the time domain and only learns and extracts the noise pattern to facilitate noise filtering of the signal, rather than extracting by enhancing the target waveform signal. It also avoids the decoupling problem of phase and amplitude in traditional methods, effectively improving the accuracy of the reconstructed waveform signal. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 Schematic diagram of the architecture of the underwater acoustic signal noise reduction system provided in an embodiment of the present application;
[0051] Figure 2 This is a flow chart of the underwater acoustic signal noise reduction method provided by an embodiment of the present application;
[0052] Figure 3 This is a specific structural diagram of a self-attention module provided in an embodiment of the present application;
[0053] Figure 4 It is the overall structural diagram of the model provided in the embodiment of the present application;
[0054] Figure 5This is a diagram showing the noise reduction effect of the present application provided in an embodiment of the present application;
[0055] Figure 6 This is a schematic diagram of the functional modules of the underwater acoustic signal noise reduction device provided in an embodiment of the present application;
[0056] Figure 7 This is a schematic diagram of the hardware structure of the computer device provided in the embodiment of the present application. DETAILED DESCRIPTION
[0057] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0058] It should be noted that although the device schematics illustrate functional module divisions and the flowcharts illustrate logical sequences, in certain circumstances, the steps shown or described may be performed in a sequence that differs from the module divisions in the device or the sequence in the flowcharts. The terms "first," "second," and so on, in the specification, claims, and drawings, are used to distinguish similar items and are not necessarily used to describe a specific sequence or precedence.
[0059] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0060] Hydroacoustic signals are sound wave signals that propagate underwater. They can be used in a variety of applications, including underwater communication, detection, navigation, and monitoring, such as underwater robots, submarines, and sonar systems. They are crucial for improving the intelligence of underwater systems, promoting marine scientific research, and protecting the marine environment. However, due to the complex underwater environment, hydroacoustic signals are often interfered with by various noises during propagation, such as water currents, marine life, and the noise of the equipment itself. This noise can reduce the clarity and reliability of the signal, hindering its accurate reception and interpretation. Therefore, noise reduction processing of hydroacoustic signals is necessary to ensure the efficiency and accuracy of underwater communication and detection systems.
[0061] In related technologies, a continuous time-domain waveform is generally segmented into multiple segments with short time intervals through a short-time Fourier transform (STFT), and each segment is Fourier transformed to convert the time-domain signal into a time-frequency diagram for signal reconstruction. However, when using the STFT to process underwater acoustic signals, an algorithm is required to estimate and reconstruct the phase. Factors such as multipath propagation, signal scattering, reverberation, and noise interference in the underwater acoustic environment may make the phase information extracted from the noise background inaccurate, resulting in phase estimation errors and, in turn, inaccurate reconstructed waveform signals.
[0062] Based on this, the embodiments of the present application provide a method, apparatus, computer device and readable storage medium for underwater acoustic signal noise reduction, which can improve the accuracy of the reconstructed waveform signal.
[0063] The underwater acoustic signal noise reduction method, device, computer equipment and readable storage medium provided in the embodiments of the present application are specifically illustrated through the following embodiments. First, the underwater acoustic signal noise reduction system in the embodiments of the present application is described.
[0064] Please refer to Figure 1 In some implementations, an embodiment of the present application provides an underwater acoustic signal noise reduction system, including a terminal 11 and a server 12 .
[0065] For example, terminal 11 can be an underwater acoustic signal acquisition and preprocessing device, which may include a hydrophone, embedded computer equipment, communications, etc. Terminal 11 can collect underwater acoustic signals in real time. Terminal 11 can also preprocess the collected signals locally, such as format conversion and simple filtering, to reduce the amount of data required for subsequent transmission, and then send the preprocessed signals to server 12 via wired or wireless means for further processing.
[0066] Furthermore, the server end 12 can be a high-performance server or cloud computing resource, which can train the preset model based on the sample, and after obtaining the target model through training, calculate the local and global noise features of the signal sent by the terminal 11 in parallel, and use the decoder to generate the noise time domain signal, and then obtain the target waveform signal through time domain subtraction, and finally return the target waveform signal to the terminal 11, thereby improving the efficiency and accuracy of the reconstructed waveform signal.
[0067] The method for reducing noise of underwater acoustic signals in the embodiments of the present application can be illustrated by the following embodiments.
[0068] It should be noted that in each specific embodiment of the present application, when it comes to the need to perform relevant processing based on data related to user identity or characteristics such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiment of the present application needs to obtain the user's sensitive personal information, the user's separate permission or consent will be obtained through a pop-up window or by jumping to a confirmation page. After clearly obtaining the user's separate permission or consent, the necessary user-related data for the normal operation of the embodiment of the present application will be obtained.
[0069] In the embodiment of the present application, the underwater acoustic signal noise reduction device will be described from the perspective of the underwater acoustic signal noise reduction device, which can be integrated into a computer device. Figure 2 , Figure 2 This is a flowchart of the steps of the underwater acoustic signal noise reduction method provided in an embodiment of the present application. In this embodiment of the present application, the underwater acoustic signal noise reduction device is specifically integrated into a terminal or server as an example. When the processor on the terminal or server executes the program instructions corresponding to the underwater acoustic signal noise reduction method, the specific process is as follows:
[0070] Step 101: obtain a time domain signal of the underwater sound to be de-noised, and perform feature extraction on the time domain signal to be de-noised using a target model to obtain a feature tensor to be de-noised, wherein the feature tensor to be de-noised includes multiple local feature blocks.
[0071] In some embodiments, in order to more effectively remove noise and restore the original signal, the time domain signal to be denoised can be extracted from the acquired water sound to be denoised, and deep feature extraction can be performed using the target model to obtain a feature tensor to be denoised that can represent the essential characteristics of the signal, so as to facilitate subsequent more detailed local feature analysis and global feature integration.
[0072] The water sound to be de-noised may be the original water sound collected from an underwater environment (e.g., an ocean environment) and containing noise (e.g., wind, rain, current, marine life, etc.) and detection targets (e.g., sounds emitted by ships, submarines, etc.).
[0073] The time domain signal to be de-noised may be time series data corresponding to the underwater sound to be de-noised, and is in a form that can be directly used for digital signal processing.
[0074] Among them, the target model can be an end-to-end underwater noise perception model, which is used to extract the characteristics of the target noise from the time domain signal to be de-noised and generate the target noise so as to facilitate the subsequent filtering of the target noise from the time domain signal to be de-noised.
[0075] Among them, the feature tensor to be denoised can be a high-dimensional data structure obtained by performing a convolution operation on the input time domain signal to be denoised through the target model, dividing it according to a preset division method and then splicing it, which contains the key feature information of the target noise.
[0076] Among them, the local feature block can be several sub-parts into which the feature tensor to be denoised is divided. Each sub-part (each local feature block) represents the local characteristics of the time domain signal to be denoised within a specific time period, which facilitates the subsequent learning of local dependencies and global dependencies.
[0077] For example, the time domain signal of the water sound to be reduced can be collected by a hydrophone or other equipment. The water sound to be reduced includes the sound from the ship (or other targets to be detected) and the background noise (such as the sound of wind, rain, water flow and marine life). The sampling rate and sampling time of the time domain signal to be reduced can be set according to the actual situation. For example, the sampling rate can be set to 16kHz and the sampling time can be set to 3 seconds.
[0078] For example, the pre-trained target model can be used to process the time domain signal x∈R to be denoised T Assume that the length of the input time domain signal to be denoised is T = 48,000 samples. After the one-dimensional convolution operation, the output feature dimension is F = 256. The time step remains unchanged, and the convolution kernel size k is set to 16. In this way, the local structural characteristics of the signal can be captured. The calculation formula is as follows: W = H(conv1d(x));
[0079] Where W∈R F×T is the feature to be denoised after the convolution operation, and H(·) represents the ReLU activation function.
[0080] Furthermore, after the above convolution operation, a feature to be denoised with a shape of [F, T] can be obtained. To facilitate subsequent local and global dependent feature processing, the feature to be denoised can be divided into multiple local feature blocks according to a preset partitioning method. These multiple local feature blocks are then concatenated to obtain a feature tensor to be denoised. The preset partitioning method can be to partition the feature to be denoised according to a preset partition length and number of partition hops.
[0081] By acquiring the time domain signal to be denoised and dividing it to obtain the feature tensor to be denoised, it is convenient to extract the local dependency features and the global dependency features subsequently.
[0082] In some embodiments, to enhance the ability to identify local patterns and dynamic changes in complex underwater acoustic environments, the features to be denoised can be divided into multiple local feature blocks according to a specific partitioning method. This allows the formation of a denoised feature tensor that integrates local details and global contextual information, providing rich contextual information for subsequent noise removal and signal recovery. For example, step 101 of "extracting features from the time-domain signal to be denoised using the target model to obtain a denoised feature tensor" may include:
[0083] (101.1) extracting features of the time domain signal to be denoised using a convolutional neural network of the target model to obtain corresponding features to be denoised;
[0084] (101.2) Dividing the feature to be denoised into a plurality of local feature blocks according to a preset division length and division hop number;
[0085] (101.3) Multiple local feature blocks are spliced to obtain a feature tensor to be denoised.
[0086] Among them, the convolutional neural network can be a one-dimensional convolutional neural network (Conv1D) in the target model for feature extraction of the denoised time domain signal, which can learn the high-dimensional representation of the input signal and capture the spatial structural characteristics of the input signal through a series of convolution kernels.
[0087] The feature to be denoised may be a feature representation obtained by processing the time domain signal to be denoised using a convolutional neural network, which effectively captures important information in the time domain signal to be denoised.
[0088] The preset division length may be a fixed length K of each block set when the feature to be denoised is divided into a plurality of local feature blocks, which determines the time step length or the number of data points contained in each local feature block.
[0089] Among them, the division hop number can be the time step or data point interval P between adjacent blocks when dividing the feature to be denoised into multiple local feature blocks, which affects the degree of overlap between local feature blocks and the target model's ability to capture sequence information.
[0090] For example, the pre-trained target model can be used to process the time domain signal x∈R to be denoised T Assume that the length of the input time-domain signal to be denoised is T = 48,000 samples. After performing a one-dimensional convolution operation on the denoised time-domain signal using a one-dimensional convolutional neural network (for example, with a convolution kernel size of 16, a stride of 1, and 256 output channels), the output feature dimension is F = 256. The time step remains unchanged, and the convolution kernel size k is set to 16. This allows the capture of local structural characteristics in the signal. The calculation formula is as follows: W = H(conv1d(x));
[0091] Where W∈R F×T is the feature to be denoised after the convolution operation, and H(·) represents the ReLU activation function.
[0092] Furthermore, after the above convolution operation, a feature to be denoised with a shape of [F, T] can be obtained, for example, the shape can be [256, 47984] (due to the reduction of the time step caused by the convolution operation). If the length of the local feature block (that is, the preset division length) K = 256 and the number of division hops P = 128 are set, the feature to be denoised is divided into multiple local feature blocks. Then, the coverage range of the first local feature block is [0, 256), the coverage range of the second local feature block is [128, 384), and so on, multiple local feature blocks can be obtained.
[0093] In some embodiments, overlapping between adjacent local feature blocks can be set to ensure data continuity and avoid information fragmentation. For example, a 50% overlap can be set, that is, the preset partition length is set to 256, the partition hop count is set to 128, etc.
[0094] Furthermore, all local feature blocks can be concatenated along the time dimension to form a feature tensor to be denoised. For example, if the shape of each local feature block is [256, 256], the shape of the concatenated feature tensor to be denoised is [256, 374×256], i.e., [256, 95744]. It is understood that the preset partition length and number of partition hops can be set according to actual conditions and are not specifically limited in this application.
[0095] In some embodiments, convolution kernels of different sizes can be used to capture features in different frequency ranges. For example, convolution kernels of sizes 16, 32, and 64 can be used simultaneously to capture short-term dynamic changes, medium-term changes, and long-term changes in the time domain signal to be denoised, respectively. The output results can be combined by splicing or weighted summation to form richer features to be denoised.
[0096] Through the above methods, the target model can carefully analyze different parts of the signal and enhance the understanding of local patterns, so as to lay a solid foundation for subsequent high-quality underwater acoustic signal processing.
[0097] In step 102, the local position-dependent features and the local time-dependent features between each local feature block and the adjacent feature blocks in the feature tensor to be denoised are calculated in parallel through the attention mechanism of the local gated recurrent unit of the target model to obtain the local noise features.
[0098] In some embodiments, in order to extract information that can describe local noise characteristics, the local position-dependent features and local time-dependent features between each local feature block and its adjacent feature blocks in the feature tensor to be denoised can be calculated in parallel through the attention mechanism of the local gated recurrent unit in the target model, thereby identifying local patterns and temporal dynamic changes within the signal, so as to more accurately remove noise and restore the original signal.
[0099] Among them, the local gated recurrent unit attention mechanism can be an attention mechanism designed in this application for processing sequence data, which combines the advantages of the gated recurrent unit and the attention mechanism and focuses on capturing local dependencies in sequence data.
[0100] The adjacent feature block may be a local feature block that is physically or temporally adjacent to the local feature block that currently needs to be weighted in the feature tensor to be denoised.
[0101] Among them, the local position-dependent feature can be the feature reflecting the relative position relationship between each local feature block and its adjacent feature blocks when analyzing them under the action of the attention mechanism of the local gated recurrent unit (such as transient pulse noise or high-frequency interference), which is used to characterize the spatial correlation between different local feature blocks.
[0102] Among them, the local time-dependent features can be the features reflecting the time series dependency between each local feature block and its adjacent feature blocks (such as the persistence or periodic fluctuation of low-frequency noise) extracted under the action of the attention mechanism of the local gated recurrent unit, and are used to characterize the dynamic change rules of different local feature blocks in the time dimension.
[0103] Among them, the local noise feature can be a comprehensive representation of the local position-dependent features and local time-dependent features within each local feature block and between each local feature block and its adjacent feature blocks, extracted from the feature tensor to be denoised after being processed by the attention mechanism of the local gated recurrent unit.
[0104] In some embodiments, each local attention head of the local gated recurrent unit attention mechanism can be used to parallelly calculate the local time dependency features between each local feature block and adjacent feature blocks in the feature tensor to be denoised, so that each local attention head can obtain a local temporal feature based on the feature tensor to be denoised.
[0105] Furthermore, the local temporal features corresponding to all local attention heads can be concatenated along the feature dimension to obtain a local temporal feature tensor. Assuming the output feature dimensions of each head are the same, they can be concatenated directly; if the output feature dimensions of different heads are different, they must first be aligned (e.g., through linear transformation) before concatenation. This allows information from multiple perspectives to be integrated, making the final feature representation richer and more comprehensive.
[0106] In some embodiments, since the position encoding part in the Transformer encoder is not applicable to acoustic sequences, in order to accurately obtain the position dependency information of the feature tensor to be denoised, the present application deletes the position encoding part in the Transformer encoder and replaces the first fully connected layer of the feedforward network with a local gating layer, that is, a gated recurrent unit (GRU), and calculates the local position dependency features of the local time series feature tensor through the local gating layer to obtain the local noise features.
[0107] Through the above method, not only the local patterns and temporal dynamic changes within the signal to be denoised are effectively captured, but also the ability to recognize noise characteristics in complex underwater acoustic environments is enhanced. By integrating the relationships between local feature blocks, this method significantly improves the accuracy of noise removal and the effect of target signal recovery, laying a solid foundation for subsequent high-quality underwater acoustic signal processing.
[0108] In some embodiments, in order to accurately capture the local patterns and temporal dynamic changes within the signal, the denoised feature tensor can be processed by the local attention head and the local gating layer in the attention mechanism of the local gating recurrent unit, ultimately obtaining a local noise feature that can focus on the local noise characteristics, thereby improving the accuracy of subsequent noise removal and signal recovery. For example, step 102 may include:
[0109] (102.1) By using each local attention head of the local gated recurrent unit attention mechanism of the target model, local temporal dependency features between each local feature block and adjacent feature blocks in the feature tensor to be denoised are calculated in parallel to obtain local temporal features corresponding to the multiple local feature blocks;
[0110] (102.2) Concatenate multiple local temporal features corresponding to multiple local attention heads to obtain a local temporal feature tensor;
[0111] (102.3) The local position-dependent features of the local temporal feature tensor are calculated through the local gating layer of the local gating recurrent unit attention mechanism to obtain the local noise features.
[0112] The local attention head can be the basic unit in the local gated recurrent unit attention mechanism responsible for parallel processing of the relationship between local feature blocks. Each local attention head focuses on learning the local time-dependent features in a specific subspace, that is, the characteristics of the interaction between its local feature block and adjacent feature blocks over time.
[0113] Among them, the local temporal feature can be obtained after processing by the local attention head, which reflects the comprehensive feature representation of the local time dependency between each local feature block and its adjacent feature blocks. After each local attention head processes the denoised feature tensor, a local temporal feature can be obtained.
[0114] The local temporal feature tensor is a high-dimensional data structure formed by concatenating the local temporal features output by multiple local attention heads. It integrates the learning results of all local attention heads on the temporal dependencies between local feature blocks, providing richer context for subsequent analysis.
[0115] Among them, the local gating layer can be located in the attention mechanism of the local gating recurrent unit, which is used to further process the local temporal feature tensor and extract the components of the local position-dependent features to capture and integrate the relative position relationship between local feature blocks, enhance the model's ability to understand local details, and provide support for accurate identification of local noise.
[0116] Please refer to Figure 3 , Figure 3 The structure diagram of the self-attention module designed for this application. The self-attention module can be represented by MV-MHSA. Each self-attention module contains a local gated recurrent unit attention mechanism and a global gated recurrent unit attention mechanism. Figure 3 , the process from (102.1) to (102.3) is given as an example.
[0117] In some embodiments, if the number of local attention heads is 4 (for example only, there may actually be more or fewer local attention heads), different local attention heads may capture the temporal dependencies of the feature tensors to be denoised in different representation subspaces. For example, local attention head 1 can focus on high-frequency noise components, local attention head 2 can focus on low-frequency target signals, local attention head 3 can focus on dynamic changes in a short period of time, and local attention head 4 can focus on long-term trends of signals; or local attention head 1 can focus on N adjacent feature blocks adjacent to the left and right of the local feature block whose weights are currently required to be calculated, local attention head 2 can focus on M adjacent feature blocks adjacent to the left of the local feature block whose weights are currently required to be calculated, local attention head 3 can focus on X adjacent feature blocks adjacent to the right of the local feature block whose weights are currently required to be calculated, local attention head 4 can focus on G adjacent feature blocks adjacent to the left and right of the local feature block whose weights are currently required to be calculated, and so on. The specific spatial scale of attention can be adjusted according to actual conditions, and the embodiments of the present application are not limited to this.
[0118] For example, assume that the feature tensor X to be denoised is a tensor of shape (N, L, D), where N is the number of samples, L is the sequence length (i.e., the number of local feature blocks), and D is the feature dimension. If there are queries, keys, and values corresponding to the query matrix, key matrix, and value matrix respectively, then for each local attention head i, the calculation process is as follows:
[0119]
[0120] Here, head i represents the local temporal feature, that is, the output of the i-th local attention head, which captures the local temporal dependency between each local feature block and the adjacent feature blocks in the feature tensor to be denoised; W i Q represents the query matrix, W i K represents the bond matrix, W i V represents the value matrix; d k Represents the dimension of the key vector; softmax represents the function that converts the input feature tensor to be denoised into a probability distribution.
[0121] Furthermore, the local temporal features corresponding to the multiple local attention heads can be concatenated to obtain the local temporal feature tensor. Specifically, assuming that there are h local attention heads, the local temporal features output by each local attention head are: head1, head2, ..., head h, these local time series features are spliced together to form the local time series feature tensor MultiHead(Q,K,V):
[0122] MultiHead(Q,K,V)=Concat(head1,head2,...,head h )W o ;
[0123] Among them, W o is another learnable weight matrix used to convert the concatenated data back to the original dimension and finally obtain the local time series feature tensor.
[0124] Furthermore, the local gating layer of the local gating recurrent unit attention mechanism can be used to calculate the local position dependency of the local temporal feature tensor. Specifically, the features of the local temporal feature tensor at each time step can be filtered out through the reset gate of the local gating layer to filter out the noise features that are redundant with the features of the current time step (e.g., suppressing historical noise residues unrelated to the current time step). The update gate of the local gating layer then dynamically fuses the input of the current time step with the historical features of the corresponding historical time steps (e.g., strengthening the transient correlation of burst noise) to obtain the local noise feature. In this way, the local position dependency feature can be accurately extracted, thereby obtaining the final local noise feature.
[0125] Through the above methods, the key information of the original signal and the deep dynamic characteristics of the time series can be effectively combined, which significantly improves the model's ability to capture complex patterns and the noise reduction effect, making the model more efficient and accurate in processing complex underwater acoustic signals.
[0126] In some embodiments, in order to more accurately identify and separate local noise components, the local position-dependent features of the local temporal features can be further calculated by the local gating layer of the local gating recurrent unit attention mechanism to obtain more accurate local noise features. For example, (102.3) may include:
[0127] (102.3.1) Fusing the local time series feature tensor and the feature tensor to be denoised to obtain the first feature;
[0128] (102.3.2) Calculating the local position-dependent features of the first feature through the local gating layer of the local gating recurrent unit attention mechanism to obtain the second feature;
[0129] (102.3.3) The third feature is obtained by performing a nonlinear transformation on the second feature through the leaky linear rectification function of the local gated recurrent unit attention mechanism;
[0130] (102.3.4) The first feature and the third feature are fused to obtain the local noise feature.
[0131] Among them, the first feature can be a feature representation obtained by fusing the feature tensor to be denoised and the local time series feature tensor through the residual connection result. It combines the key features of the original signal and the time series relationship between local feature blocks, providing rich contextual information for subsequent analysis.
[0132] Among them, the second feature can be the result obtained by processing the first feature through the local gating layer of the local gating recurrent unit attention mechanism, which is used to characterize the relative position relationship between different features in the first feature and its impact on local noise.
[0133] The leaky linear rectifier function is an activation function that can be used to solve the problem of ReLU units not being activated at all when receiving negative inputs. Specifically, when the input is positive, the Leaky ReLU is the same as the ReLU, and the output is equal to the input; however, for negative inputs, the Leaky ReLU does not completely suppress them to zero, but multiplies them by a small positive slope to increase the learning ability and expression power of the model.
[0134] The third feature may be an output result of a nonlinear transformation of the second feature using a leaky linear rectification function.
[0135] Please refer to Figure 3 In some embodiments, since the position encoding part of the original Transformer encoder is not suitable for the acoustic sequence, this application deletes the position encoding part when designing the model, and replaces the first fully connected layer of the feedforward network with a local gating layer to replace the position encoding module to learn the position information.
[0136] Furthermore, the local gating layer of the local gating recurrent unit attention mechanism can be used to model local dependencies between spatial positions. Specifically, the reset gate of the local gating layer can suppress local noise patterns unrelated to the current frame, avoiding redundant interactions in the global attention calculation. The update gate dynamically adjusts the weights of local features (such as changes in the intensity of sudden noise), retaining only key features that depend on the local position and filtering out irrelevant information. Furthermore, the weights of the reset and update gates can be dynamically adjusted to adapt to the rapid changes in non-stationary noise. For example, when the noise power suddenly changes, the update gate can quickly enhance the current input weight to avoid interference from historical noise features.
[0137] In some embodiments, the leaky linear rectification function can be designed as:
[0138]
[0139] Here, α is a small positive number that can be set based on actual conditions, such as 0.01, which determines the magnitude of the negative gradient. The leaky ReLU (Leaky ReLU) assigns a non-zero gradient to negative inputs, allowing neurons that were originally suppressed in the ReLU to continue participating in the training of the preset model. This improves the efficiency of the preset model training and the expressiveness and generalization performance of the trained target model.
[0140] In some embodiments, the denoised feature tensor X obtained after the target model processes the denoised time domain signal can be obtained. The specific method for obtaining the denoised feature tensor has been discussed above and will not be repeated here. Afterwards, the denoised feature tensor X is processed by the four local attention heads of the local gated recurrent unit attention mechanism to obtain the local temporal features output by each local attention head, and the multiple local temporal features corresponding to the four local attention heads are spliced to obtain a local temporal feature tensor, also known as MultiHead. For example, if the output dimension of a single head is T×64, the dimension after splicing is T×256.
[0141] Please refer to Figure 3 Furthermore, the denoised feature tensor X can be normalized by a normalization layer (LayerNorm), that is, X1=LayerNorm(X), to suppress abnormal amplitude fluctuations of the noise. The denoised feature tensor X1 obtained by the normalization process is then fused (Add) with the local temporal features to obtain the first feature, that is, X2=LayerNorm(X1+MultiHead), to retain the details of the original signal while introducing the context information extracted by the attention mechanism to alleviate the signal distortion caused by phase decoupling. Afterwards, the local position-dependent feature of the first feature X2 is calculated by the local gating layer of the local gating recurrent unit attention mechanism to obtain the second feature, that is, GRU(X2), to dynamically correct the spatiotemporal correlation of the noise features. Furthermore, the second feature can be nonlinearly transformed by the leaky linear rectification function (LeakyReLU) of the local gated recurrent unit to obtain the third feature FFN, that is, FFN = LeakyReLU (GRU (X2)); finally, the first and third features are residually fused through the residual connection result of the target model to obtain the local noise feature, that is, Output = LayerNorm (X2 + FFN), which can effectively prevent information loss.
[0142] Through the above method, not only important feature details are retained, but also the learning ability and stability of the model are enhanced by introducing nonlinear transformation and residual connection technology. In this way, the key information of the original signal and the deep sequence dynamic characteristics can be effectively combined, and the model's ability to capture complex patterns can be enhanced, thereby improving the noise reduction effect and the accuracy of signal restoration. It can be widely used in high-precision noise suppression scenarios such as underwater communications and sonar detection.
[0143] In step 103, the global position-dependent features and the global time-dependent features of the feature tensor to be denoised are calculated through the global gated recurrent unit attention mechanism of the target model to obtain the global noise features.
[0144] In some embodiments, in order to identify the overall pattern and long-range temporal dynamic changes of the signal, the global gated recurrent unit attention mechanism in the target model can be used to extract information that can describe the global noise characteristics in the entire signal, so as to improve the accuracy and completeness of noise identification and better restore the original signal.
[0145] Among them, the global gated recurrent unit attention mechanism can be used to analyze and learn the overall structure of the feature tensor to be denoised and its internal temporal dynamic changes.
[0146] Among them, the global position-dependent feature can be used to characterize the spatial correlation between each local feature block in the entire range of the time domain signal to be denoised, which helps to understand the overall structure of the time domain signal to be denoised. For example, in the time domain signal to be denoised, the noise in a certain frequency band may affect the energy distribution of other frequency bands.
[0147] Among them, the global time-dependent features can be used to characterize the dynamic change law of the entire time domain signal to be denoised on a long time scale, helping to identify and separate global noise components, such as the periodic change of noise or the temporal continuity of the time domain signal to be denoised.
[0148] The global noise feature is a comprehensive representation of the noise component by integrating global position-dependent and time-dependent information into the target model. It can include global characteristics such as the noise's spectral distribution and time-domain energy variation, and is used to identify noise from mixed signals.
[0149] In some embodiments, multiple global attention heads of a global gated recurrent unit attention mechanism can be used to parallelly calculate the global temporal dependency features between each global feature block in the feature tensor to be denoised and other feature blocks in the entire feature sequence. Each global attention head can obtain a global temporal feature based on the feature tensor to be denoised. It is understood that the number of global attention heads can be set according to actual conditions, for example, 8 global attention heads can be set, or 9, 12, and so on.
[0150] Furthermore, the global temporal features corresponding to all global attention heads can be concatenated along the feature dimension to obtain a global temporal feature tensor. Assuming that the output feature dimensions of each global attention head are the same, they can be concatenated directly; if the output feature dimensions of different global attention heads are different, they must first be aligned (e.g., through linear transformation) before concatenation. This allows information from multiple perspectives to be integrated, making the final feature representation richer and more comprehensive.
[0151] In some embodiments, similar to the local gated recurrent unit attention mechanism, since the position encoding part in the Transformer encoder is not applicable to acoustic sequences, in order to accurately obtain the global position dependency information of the feature tensor to be denoised, the present application deletes the position encoding part in the Transformer encoder and replaces the first fully connected layer of the feedforward network with a global gating layer, that is, a gated recurrent unit (GRU), and calculates the global position dependency features of the global time series feature tensor through the global gating layer to obtain the global noise features.
[0152] Through the above method, not only the global patterns and temporal dynamic changes within the signal to be denoised are effectively captured, but also the ability to recognize noise characteristics in complex underwater acoustic environments is enhanced, which helps to significantly improve the accuracy of noise removal and the effect of target signal recovery, laying a solid foundation for subsequent high-quality underwater acoustic signal processing.
[0153] In some embodiments, the purpose is to parallelly calculate the global time-dependent features of the feature tensor to be denoised through the global attention head in the global gated recurrent unit attention mechanism in the target model, and then splice these global temporal features into a global temporal feature tensor. Then, the global position-dependent features are further analyzed using the global gating layer, and finally a global noise feature that can describe the noise characteristics in the entire signal is extracted. This process helps to capture the overall pattern and long-range temporal dynamic changes of the signal, thereby improving the accuracy of noise removal and signal recovery. Step 103 may include:
[0154] (103.1) Calculate the global time-dependent features of the feature tensor to be denoised through each global attention head of the global gated recurrent unit attention mechanism of the target model, and obtain the global temporal features corresponding to each global attention head;
[0155] (103.2) Concatenate multiple global temporal features corresponding to multiple global attention heads to obtain a global temporal feature tensor;
[0156] (103.3) The global position-dependent features of the global temporal feature tensor are calculated through the global gating layer of the global gated recurrent unit attention mechanism to obtain the global noise features.
[0157] Among them, the global attention head can be the basic unit in the global gated recurrent unit attention mechanism responsible for parallel processing of the overall structure of the feature tensor to be denoised and its internal time dynamic relationship.
[0158] Among them, the global temporal features can be obtained after processing by the global attention head, which is used to reflect the long-term time series correlation of different local feature blocks in the feature tensor to be denoised within the range of the entire signal to be denoised, which helps to understand the overall temporal dynamic change law of the signal to be denoised.
[0159] Among them, the global temporal feature tensor can be a high-dimensional data structure formed by splicing the global temporal features output by multiple global attention heads.
[0160] The global gating layer can be a component within the global gated recurrent unit attention mechanism, further processing the global temporal feature tensor to extract global position-dependent features. The global gating layer effectively captures and integrates the relative positional relationships between different local feature blocks through the gating mechanism, enhancing the target model's ability to understand the global structure.
[0161] Please refer to Figure 3 , Figure 3 The structure diagram of the self-attention module designed for this application. The self-attention module can be represented by MV-MHSA. Each self-attention module contains a local gated recurrent unit attention mechanism and a global gated recurrent unit attention mechanism. Figure 3 , the process from (103.1) to (103.3) is given as an example.
[0162] In some embodiments, if the number of global attention heads is 8 (for example only, there may be more or less global attention heads in practice), it is assumed that the feature tensor to be denoised X is a tensor of shape (N, L, D), where N is the number of samples, L is the sequence length (i.e., the number of global feature blocks), and D is the feature dimension. If there is a query (Query), a key (Key), and a value (Value) corresponding to the query matrix W respectively i Q , bond matrix W i K Sum matrix W i V , where i represents the i-th global attention head.
[0163] Furthermore, we can calculate the dot product of the query matrix and the key matrix, apply the Softmax function after scaling to obtain the attention weights, and multiply the attention weights by the value matrix to obtain the global temporal features output by each global attention head for the feature tensor to be denoised. The specific calculation process is as follows:
[0164]
[0165] Here, head i represents the global temporal feature, that is, the output of the i-th global attention head, which captures the global temporal dependency between each global feature block and the adjacent feature blocks in the feature tensor to be denoised; W i Q represents the query matrix, W i K represents the bond matrix, W i V represents the value matrix; d k Represents the dimension of the key vector; softmax represents the function that converts the input feature tensor to be denoised into a probability distribution.
[0166] Furthermore, the global temporal features corresponding to the multiple global attention heads can be concatenated to obtain the global temporal feature tensor. Specifically, assuming that there are h global attention heads, the global temporal features output by each global attention head are: head1, head2, ..., head h , these global time series features are spliced together to form the global time series feature tensor MultiHead(Q,K,V):
[0167] MultiHead(Q,K,V)=Concat(head1,head2,...,head h )W o ;
[0168] Among them, W o is another learnable weight matrix used to convert the concatenated data back to the original dimension and finally obtain the global time series feature tensor.
[0169] Furthermore, the global position-dependent features of the global temporal feature tensor can be calculated through the global gating layer of the global gated recurrent unit attention mechanism to obtain the global noise features. Specifically, the features of the global temporal feature tensor at each time step can be filtered out through the reset gate of the global gating layer to filter out the redundant noise features of the features of the current time step (such as suppressing the historical noise residue irrelevant to the current time step), and then the input of the current time step is dynamically fused with the historical features of the corresponding historical time step (such as strengthening the transient correlation of burst noise) through the update gate of the global gating layer to obtain the global noise features. In this way, the global position-dependent features can be accurately extracted, thereby obtaining the final global noise features.
[0170] In some embodiments, in order to more accurately identify and separate the global noise component, the global position-dependent feature of the global temporal feature can be further calculated by the global gating layer of the global gated recurrent unit attention mechanism to obtain a more accurate global noise feature. Exemplarily, (103.3) may include: fusing the global temporal feature tensor and the feature tensor to be denoised to obtain a fourth feature; calculating the global position-dependent feature of the fourth feature by the global gating layer of the global gated recurrent unit attention mechanism to obtain a fifth feature; performing a nonlinear transformation on the fifth feature by the leaky linear rectification function of the global gated recurrent unit attention mechanism to obtain a sixth feature; and fusing the fourth and sixth features to obtain a global noise feature.
[0171] Please refer to Figure 3 In some embodiments, the global dependency between spatial positions can be modeled through the global gating layer of the global gated recurrent unit attention mechanism. Specifically, the reset gate of the global gating layer can be used to suppress global noise patterns that are irrelevant to the current frame, avoid redundant interactions in the global attention calculation, and dynamically adjust the weights of global features (such as changes in the intensity of sudden noise) through the update gate, retaining only key features that depend on the global position and filtering out irrelevant information. Furthermore, the weights of the reset gate and the update gate can be dynamically adjusted to adapt to the rapid changes of non-stationary noise. For example, when the noise power suddenly changes, the update gate can quickly enhance the current input weight to avoid interference from historical noise features.
[0172] It is understandable that the relevant content of the leaky linear rectifier function has been expanded above and will not be repeated here. Furthermore, the denoised feature tensor X can be processed by the eight global attention heads of the global gated recurrent unit attention mechanism to obtain the global temporal features output by each global attention head. The multiple global temporal features corresponding to the eight global attention heads are then concatenated to obtain the global temporal feature tensor, also known as MultiHead. For example, if the output dimension of a single head is T×64, the dimension after concatenation is T×512.
[0173] Please refer to Figure 3 Furthermore, the denoised feature tensor X can be normalized by the normalization layer (LayerNorm), that is, X1=LayerNorm(X), to suppress the abnormal amplitude fluctuation of the noise. Then, the denoised feature tensor X1 obtained by the normalization process is fused (Add) with the global temporal feature to obtain the fourth feature, that is, X2=LayerNorm(X1+MultiHead), to retain the details of the original signal while introducing the context information extracted by the attention mechanism to alleviate the signal distortion caused by phase decoupling; then, the global position dependency feature of the fourth feature X2 is calculated by the global gated layer of the global gated recurrent unit attention mechanism to obtain the fifth feature, that is, GRU(X2), to dynamically correct the spatiotemporal correlation of the noise feature. Furthermore, the fifth feature can be nonlinearly transformed by the leaky linear rectification function (LeakyReLU) of the global gated recurrent unit to obtain the sixth feature FFN, that is, FFN = LeakyReLU (GRU (X2)); finally, the fourth and sixth features are residually fused through the residual connection result of the target model to obtain the global noise feature, that is, Output = LayerNorm (X2 + FFN), which can effectively prevent information loss.
[0174] By using multiple global attention heads to capture long-term dependencies and global patterns across the entire sequence in parallel, the model can more accurately identify and separate noise. This makes the model more efficient and accurate when handling low signal-to-noise ratio scenarios, significantly improving the accuracy of noise reduction recognition.
[0175] Step 104 : reconstructing the noise time domain signal of the target noise from the time domain signal to be denoised based on the noise mask matrix generated by fusing the local noise features and the global noise features through the decoder of the target model.
[0176] In some embodiments, in order to accurately estimate the noise component and separate it from the original signal, the decoder in the target model can be used to reconstruct the noise time domain signal of the target noise from the time domain signal to be denoised, thereby providing a basis for subsequent noise removal and restoration of a clean target waveform signal.
[0177] The decoder can be the part of the target model that converts high-dimensional feature representations (such as noise mask matrices) back to time-domain signals. The decoder can be composed of a series of deconvolution or transposed convolution layers.
[0178] The noise mask matrix may be a matrix formed by adding local noise features and global noise features, performing feature normalization, and further compressing the matrix to [0, 1].
[0179] The target noise may be a signal representation obtained after being processed by a decoder and containing only noise components.
[0180] The noise time-domain signal can be a specific noise signal reconstructed by the decoder based on the noise mask matrix, expressed as a time series. The noise time-domain signal is used to accurately characterize the noise components in the time-domain signal to be denoised and can be used in the subsequent noise filtering process to remove noise from the original signal and restore a clear target waveform signal.
[0181] Please refer to Figure 4 In some embodiments, the target model includes multiple self-attention modules, each of which includes a local gated recurrent unit attention mechanism and a global gated recurrent unit attention mechanism. The results of multiple self-attention modules are fused module by module to finally obtain local noise features and global noise features. The local noise features and the global noise features are spliced along the channel dimension to obtain a mask matrix M, which can be in the shape of [B, T, 2C].
[0182] In some embodiments, the noise mask matrix Wm can be obtained by multiplying the feature to be denoised W and the noise mask matrix M. The noise mask matrix can be used to characterize the feature area dominated by noise. The specific formula is: Wm=M*W.
[0183] It can be understood that the method of obtaining the features to be denoised has been introduced above, that is, the time domain signal to be denoised (for example, the dimension is [B, T]) is input into the encoder of the target model, and processed by the convolution kernel and ReLU activation function to obtain the features to be denoised. This part has been expanded above and will not be repeated here.
[0184] Furthermore, based on the noise mask matrix, Wm can be mapped back to the time domain through transpose convolution (TransposeConv1D) to generate the noise time domain signal of the target noise. The specific formula is as follows: S t =TransposeConv1D(W m ).
[0185] By reconstructing the noise time domain signal from the time domain signal to be denoised through the learned noise mask matrix, we can bypass the phase distortion problem of traditional time-frequency domain methods and directly learn and process the noise features, significantly improving the noise separation accuracy and facilitating subsequent direct noise filtering to obtain the target waveform signal.
[0186] Step 105 : Based on the noisy time domain signal, perform noise filtering on the time domain signal to be de-noised to obtain a target waveform signal.
[0187] In some embodiments, in order to accurately estimate and remove the noise components in the original signal, the reconstructed noisy time domain signal can be used to perform noise filtering on the time domain signal to be de-noised, thereby obtaining a clean target waveform signal to ensure the validity and accuracy of the acquired underwater acoustic signal.
[0188] Among them, the target waveform signal can be a clean signal extracted from the time domain signal to be denoised after noise filtering processing. It represents the ideal output after removing noise interference from the original time domain signal to be denoised, retains the key features and information of the target signal, and has a high signal-to-noise ratio and clarity.
[0189] In some implementations, by performing a time-domain subtraction between the target noise and the target signal, a waveform signal representing the target in the actual underwater acoustic environment, such as the sounds of ships or marine life, can be directly obtained without the interference of background noise. This facilitates further analysis and identification of the target waveform signal, as well as other applications such as underwater target detection, positioning, and tracking, demonstrating its practicality.
[0190] The present application obtains a time domain signal of the underwater sound to be de-noised, and extracts features of the time domain signal to be de-noised through a target model to obtain a feature tensor to be de-noised, wherein the feature tensor to be de-noised includes multiple local feature blocks; through the local gated recurrent unit attention mechanism of the target model, the local position-dependent features and local time-dependent features between each local feature block in the feature tensor to be de-noised and the adjacent feature blocks are calculated in parallel to obtain local noise features; through the global gated recurrent unit attention mechanism of the target model, the global position-dependent features and global time-dependent features of the feature tensor to be de-noised are calculated to obtain global noise features; through the decoder of the target model, the noise mask matrix generated by the fusion of the local noise features and the global noise features is used to reconstruct the noise time domain signal of the target noise from the time domain signal to be de-noised; based on the noise time domain signal, the time domain signal to be de-noised is noise filtered to obtain a target waveform signal. In this way, by combining the local gated recurrent unit attention mechanism and the global gated recurrent unit attention mechanism, the model can simultaneously capture local and global dependencies in the sequence data, and can accurately capture transient noise and short-term signal details through the local gated recurrent unit attention mechanism, and can also model long-term dependencies through the global gated recurrent unit attention mechanism to accurately identify and utilize complex patterns in the signal, thereby more effectively retaining the characteristics of the target signal during the noise reduction process. At the same time, the present application directly reconstructs a clear signal waveform in the time domain and only learns and extracts the noise pattern to facilitate noise filtering of the signal, rather than extracting by enhancing the target waveform signal. It also avoids the decoupling problem of phase and amplitude in traditional methods, effectively improving the accuracy of the reconstructed waveform signal.
[0191] In some embodiments, to obtain a target model with better performance, a preset model can be trained using sample time domain signals corresponding to the sample underwater sound to be de-noised, so that it can accurately learn, identify and separate noise features and restore a clean predicted waveform signal, ensuring that the model can still effectively identify noise features when faced with complex signals. For example, the target model can be trained in the following ways:
[0192] (A.1) Obtaining a sample time domain signal of the underwater sound to be de-noised, and performing feature extraction on the sample time domain signal using a preset model to obtain a sample time domain signal tensor, wherein the sample time domain signal tensor includes a plurality of sample local feature blocks;
[0193] (A.2) Using the local gated recurrent unit attention mechanism of the preset model, the local position-dependent features and the local time-dependent features of each sample local feature block in the sample time domain signal tensor are calculated in parallel with the adjacent sample feature blocks to obtain the local noise features of the sample;
[0194] (A.3) Calculate the sample global position-dependent features and sample global time-dependent features of the sample time-domain signal tensor through the global gated recurrent unit attention mechanism of the preset model to obtain the sample global noise feature;
[0195] (A.4) reconstructing a predicted noise time domain signal of the sample noise from the sample time domain signal through a decoder of a preset model based on a sample noise mask matrix generated by fusing the sample local noise feature and the sample global noise feature;
[0196] (A.5) Based on the predicted noise time domain signal, filtering the sample time domain signal to obtain a predicted waveform signal;
[0197] (A.6) determining a target loss based on the predicted noise time domain signal and the predicted waveform signal;
[0198] (A.7) According to the target loss, the preset model is trained to obtain the target model.
[0199] The sample water sound to be de-noised can be the original water sound collected from an underwater environment (such as an ocean environment) containing noise (such as wind, rain, water flow, marine life, etc.) and detection targets (such as sounds emitted by ships, submarines, etc.). When training the preset model, each sample water sound to be de-noised corresponds to a sample noise time domain signal and a sample waveform signal. Alternatively, the sample water sound to be de-noised can also be directly obtained from a preset training data set.
[0200] The sample time domain signal may be time series data corresponding to the sample underwater sound to be de-noised, and is in a form that can be directly used for digital signal processing.
[0201] The preset model may be an initial model structure set at the beginning of training, and the target model may be obtained by training the preset model.
[0202] Among them, the sample time domain signal tensor can be a high-dimensional data structure obtained by performing a convolution operation on the input sample time domain signal through a preset model, dividing it according to a preset division method, and then splicing it together, which contains the key feature information of the sample noise.
[0203] Among them, the sample local feature block can be several sub-parts into which the sample time domain signal tensor is divided. Each sub-part (each sample local feature block) represents the local characteristics of the sample time domain signal in a specific time period, which facilitates the subsequent learning of local dependencies and global dependencies.
[0204] The sample adjacent feature block may be a sample local feature block that is physically or temporally adjacent to the sample local feature block that currently needs to be weighted in the sample time domain signal tensor.
[0205] Among them, the sample local position dependent features can be the features reflecting the relative position relationship between each sample local feature block and its adjacent sample feature blocks when analyzing them under the action of the local gated recurrent unit attention mechanism (such as transient pulse noise or high-frequency interference), which are used to characterize the spatial correlation between different sample local feature blocks.
[0206] Among them, the sample local time dependency feature can be the feature reflecting the time series dependency between each sample local feature block and its adjacent sample feature block, extracted under the action of the local gated recurrent unit attention mechanism (such as the persistence or periodic fluctuation of low-frequency noise), which is used to characterize the dynamic change law of different sample local feature blocks in the time dimension.
[0207] Among them, the local noise features of the samples can be extracted from the sample time domain signal tensor after being processed by the local gated recurrent unit attention mechanism, and can describe the comprehensive representation of the local position-dependent features and local time-dependent features within each sample local feature block and between it and the adjacent feature blocks of the sample.
[0208] Among them, the sample global position dependence feature can be used to characterize the spatial correlation between the local feature blocks of each sample in the entire sample time domain signal range, which helps to understand the overall structure of the sample time domain signal. For example, in the sample time domain signal, the noise in a certain frequency band may affect the energy distribution of other frequency bands.
[0209] Among them, the global time dependence characteristics of the sample can be used to characterize the dynamic change law of the entire sample time domain signal on a long time scale, helping to identify and separate global noise components, such as the periodic change of noise or the temporal continuity of the sample time domain signal.
[0210] The sample global noise signature can be a comprehensive representation of noise components by integrating global position-dependent and time-dependent information using a pre-set model. The sample global noise signature can include global characteristics such as the noise's spectral distribution and time-domain energy variation, and is used to identify noise from mixed signals.
[0211] The sample noise mask matrix may be a matrix formed by adding the sample local noise features and the sample global noise features, performing feature normalization, and further compressing the matrix to [0, 1].
[0212] The predicted noise time-domain signal can be the specific noise signal reconstructed by the decoder based on the sample noise mask matrix, expressed as a time series. The predicted noise time-domain signal accurately represents the noise component in the sample time-domain signal and can be used in subsequent noise filtering processes to remove noise from the original signal and restore a clear predicted waveform signal.
[0213] The predicted waveform signal may be a clean signal estimate obtained by performing noise filtering based on the predicted noise time domain signal, and represents an ideal output after removing noise interference.
[0214] Among them, the target loss can be a loss value calculated based on the difference between the predicted noise time domain signal and the actual sample noise time domain signal, and the predicted waveform signal and the actual sample waveform signal, which is used to guide the optimization process of the preset model.
[0215] For example, a comprehensive underwater noise reduction dataset covering a variety of scenarios and conditions can be constructed for training and evaluating models. The underwater noise reduction data of this application is gathered based on the public underwater acoustic signal dataset ShipsEar, and is generated using two types of ship signals: passenger ships and RORO (roll-on / roll-off ships), and two types of underwater environmental noise signals (wind noise, rain noise, flow noise). The underwater noise reduction dataset mixes ship radiation noise signals with randomly selected environmental noise signals to synthesize a total of 12,274 segments of sample time domain signals with extremely low signal-to-noise ratio ranges [-15dB, -10dB], [-10dB, -5dB], [-5dB, 0dB], of which 60% are used for training, 20% are used for training verification, and 20% are used for evaluation (the proportion can be adjusted according to actual conditions). All sample time domain signals are resampled to 16kHz and limited to 3s. By analyzing the spectrogram and power spectrum of the sample time-domain signals, we can see that the passenger ship has relatively clear spectral lines in the 0-500 Hz and 500-1 kHz bands. However, at low signal-to-noise ratios, the passenger ship signal is almost obscured, while noise exists at both low and high frequencies. Therefore, it is necessary to train the preset model so that it can accurately identify noise.
[0216] In some embodiments, (A.1) to (A.5), the process of processing the sample time domain signal through a preset model to finally obtain a sample noise mask matrix, and subsequently performing noise filtering on the sample time domain signal based on the sample noise mask matrix to obtain a predicted waveform signal is the same as the process introduced above of processing the denoised time domain signal through the target model to finally obtain a noise mask matrix, and subsequently performing noise filtering on the denoised time domain signal based on the noise mask matrix to obtain a target waveform signal. The two have the same algorithm architecture and operation steps, except that the stages of the signal processing models are different (the preset model and the target model, respectively), and the signal properties are different (the offline sample data to be labeled and the real-time input signal data, respectively). Therefore, the above processing flow can be referred to, and this application will not elaborate on it here.
[0217] In some embodiments, the noise prediction term and the signal prediction term can be combined to determine the target loss. Specifically, the present application designs a weighted noise loss function (r-nSI-SNR) to simultaneously consider the accuracy of noise prediction and the accuracy of signal prediction obtained by noise estimation based on the similarity between the predicted noise time domain signal and the sample noise time domain signal, as well as the predicted waveform signal and the sample waveform signal.
[0218] Furthermore, the backpropagation algorithm can be used to propagate the gradient of the loss function back to the various parameters of the preset model, update the weights and biases of the preset model, and iteratively optimize: repeat the forward propagation, target loss calculation, and backpropagation steps until the target loss of the preset model converges to a satisfactory level. For example, if the target loss is lower than the preset value (such as 0.01, etc.) for three consecutive times, or the number of training times reaches the preset number (such as 500 times, 1500 times, etc.), the training of the preset model can be stopped, and the optimized target model can be finally obtained.
[0219] Through target loss, the preset model can continuously adjust its own parameters to minimize the difference between predicted noise and real noise, while maximizing the similarity between predicted waveform and real waveform, so that the model can run more stably and efficiently in practical applications.
[0220] In some embodiments, to enable the model to more accurately learn noise characteristics and reconstruct the noise signal, the target loss can be determined by comparing the similarity between the predicted noise time domain signal and the sample noise time domain signal, as well as between the predicted waveform signal and the sample waveform signal, so as to improve the performance of the model, guide the optimization direction of the model, and subsequently accurately restore the clean target signal. For example, (A.6) may also include:
[0221] (A.6.1) Obtaining a sample noise time domain signal and a sample waveform signal corresponding to the sample water sound to be de-noised;
[0222] (A.6.2) calculating a corresponding first similarity based on the predicted noise time domain signal and the sample noise time domain signal;
[0223] (A.6.3) calculating a corresponding second similarity based on the predicted waveform signal and the sample waveform signal;
[0224] (A.6.4) Determine a target loss based on the first similarity and the second similarity.
[0225] Among them, the sample noise time domain signal can be a time series signal extracted from the sample water sound to be de-noised and containing only noise components. It is a known real noise signal and is used to compare with the predicted noise time domain signal generated by the preset model to evaluate the accuracy of the preset model's estimation of the noise component.
[0226] Among them, the sample waveform signal can be an ideal clean signal after noise is removed from the sample water sound to be de-noised, that is, a real target signal, which is used to compare with the predicted waveform signal generated by the preset model to evaluate the accuracy of the preset model in recovering the waveform signal.
[0227] The first similarity may be a measurement value obtained by comparing the similarity between the predicted noise time domain signal and the sample noise time domain signal. The first similarity may be calculated by using, for example, mean square error, cosine similarity, or other appropriate similarity measurement methods.
[0228] The second similarity may be a metric value obtained by comparing the similarity between the predicted waveform signal and the sample waveform signal, and may also be calculated using methods such as mean square error and cosine similarity.
[0229] For example, the sample noise time domain signal and sample waveform signal corresponding to the sample water sound to be de-noised can be obtained through signal separation technology (such as blind source separation technology, etc.), signal processing methods (such as filtering, spectral subtraction or statistical methods, etc.), expert knowledge, and pre-settings for the sample water sound to be de-noised. It is only necessary to ensure that the sample noise time domain signal and the sample waveform signal are clean and accurate signals separated from the sample water sound to be de-noised. This application does not limit the specific acquisition method.
[0230] In some embodiments, if the predicted noise time domain signal obtained by the preset model prediction is n′, the sample noise time domain signal is n, the predicted waveform signal is mix-n′, and the sample waveform signal is s, then the target loss r-nSISNR can be calculated using the following formula:
[0231] r-nSISNR=αSISNR(n',n)+βSISNR(mix-n',s);
[0232] Among them, α is the weight coefficient used to balance the importance of the noise prediction term in the target loss; β is the weight coefficient used to balance the importance of the signal prediction term in the target loss; SISNR(n',n) represents the first similarity, and SISNR(mix-n',s) represents the second similarity.
[0233] By calculating the target loss in the above way, the preset model can better learn the characteristics of noise from the mixed signal and accurately separate the clean target signal, thereby improving the noise reduction performance of the model.
[0234] Please refer to Figure 3 and Figure 4 In some embodiments, combined Figure 3 and Figure 4 , the overall process of this application is introduced. First, refer to Figure 4After obtaining the denoised signal of the underwater sound to be denoised, the time domain signal to be denoised can be encoded through a one-dimensional convolutional layer (Conv1D) and a rectified linear unit (ReLU) to generate the denoised features. The denoised features are then divided into multiple local feature blocks and spliced together to obtain the denoised feature tensor. In this way, the phase decoupling problem of STFT can be avoided through time domain convolution.
[0235] Furthermore, the feature tensor to be denoised can be input into a multi-view noise learning module composed of multiple self-attention modules (MV-MHSA). Each self-attention module contains a local gated recurrent unit attention mechanism and a global gated recurrent unit attention mechanism, which respectively extract local position / time dependent features (such as transient noise, etc.) and global position / time dependent features (such as background noise, etc.), and obtain the corresponding local noise features and global noise features.
[0236] Furthermore, the noise mask matrix (also known as the mask matrix) can be obtained by fusing local and global noise features through components such as the target model's parameterized rectified linear unit (PReLU), two-dimensional convolutional layer (Conv2D), and hyperbolic tangent function (Tanh). The decoder (one-dimensional transposed convolution) ultimately reconstructs the noisy time-domain signal, and the target waveform signal is obtained through time-domain subtraction. This allows an end-to-end time-domain processing framework to circumvent the phase estimation challenges of time-frequency domain methods, and combines a multi-perspective self-attention mechanism to separate the target waveform signal from the noise, significantly improving the signal-to-noise ratio and signal fidelity in low signal-to-noise ratio scenarios.
[0237] Further, Figure 3 This is a specific architecture diagram of the self-attention module (MV-MHSA). After the feature tensors to be denoised are input into the local gated recurrent unit attention mechanism and the global gated recurrent unit attention mechanism respectively, the target model can extract noise features in multiple subspaces in parallel through the multi-head mechanism, thereby enhancing the adaptability to complex noise patterns.
[0238] Furthermore, the outputs of the local gated recurrent unit attention mechanism and the global gated recurrent unit attention mechanism can be fused through superposition (Add) and layer normalization (LayerNorm), and combined with the leaky linear rectifier function (Leaky ReLU) to suppress gradient vanishing. Finally, the noise mask matrix is generated by the Sigmoid function to identify the noise-dominated area (weight value 0-1). Finally, the noise time domain signal is extracted by performing point multiplication processing on the feature to be denoised and the noise mask matrix.
[0239] Please refer to Figure 5 In some embodiments, a RORO signal in the test set of the underwater noise reduction dataset is taken as an example. Figure 5The power spectrum of the mixed signal shows that the time domain signal to be denoised ( Figure 5 The power spectrum of the target noise ( Figure 5 The power spectrum of the real waveform signal ( Figure 5 The target waveform signal (marked as B in the figure) is almost submerged in the noise and cannot be distinguished. The present application obtains the noise mask matrix through the target model, and reconstructs the noise time domain signal of the target noise based on the noise mask matrix, and obtains the target waveform signal ( Figure 5 The target model (marked as D in Figure 1) shows a striking resemblance to the true target waveform signal. Calculations show that, after noise reduction processing by the target model, the signal-to-interference plus noise ratio (SISNR) of the denoised time-domain signal increases by 21.7 dB (decibels), and the signal-to-distortion ratio (SDR) increases by 22.3 dB. This demonstrates that the target model can restore clean signal and noise components even at low signal-to-noise ratios.
[0240] See also Figure 6 The embodiment of the present application further provides an underwater acoustic signal noise reduction device, which can implement the above-mentioned underwater acoustic signal noise reduction method. The underwater acoustic signal noise reduction device includes:
[0241] An acquisition module 61 is configured to acquire a time domain signal of the underwater sound to be de-noised, and perform feature extraction on the time domain signal to be de-noised using a target model to obtain a feature tensor to be de-noised, wherein the feature tensor to be de-noised includes a plurality of local feature blocks;
[0242] A first computing module 62 is configured to concurrently compute local position-dependent features and local time-dependent features between each local feature block in the feature tensor to be denoised and its adjacent feature blocks through an attention mechanism of a local gated recurrent unit of the target model, thereby obtaining a local noise feature.
[0243] A second calculation module 63 is configured to calculate the global position-dependent features and the global time-dependent features of the feature tensor to be denoised through the global gated recurrent unit attention mechanism of the target model to obtain a global noise feature;
[0244] A reconstruction module 64 is configured to reconstruct a noise time domain signal of the target noise from the time domain signal to be denoised based on a noise mask matrix generated by fusing local noise features and global noise features through a decoder of the target model;
[0245] The filtering module 65 is configured to perform noise filtering on the time domain signal to be de-noised based on the noisy time domain signal to obtain a target waveform signal.
[0246] The specific implementation of the underwater acoustic signal noise reduction device is basically the same as the specific embodiment of the underwater acoustic signal noise reduction method described above, and will not be repeated here. Under the premise of meeting the requirements of the embodiment of the present application, the underwater acoustic signal noise reduction device can also be provided with other functional modules to implement the underwater acoustic signal noise reduction method in the above embodiment.
[0247] The present application also provides a computer device comprising a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described underwater acoustic signal noise reduction method. The computer device can be any intelligent terminal, including a tablet computer and an in-vehicle computer.
[0248] See also Figure 7 , Figure 7 The hardware structure of a computer device according to another embodiment is shown. The computer device includes:
[0249] The processor 71 may be implemented as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application.
[0250] The memory 72 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 72 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program codes are stored in the memory 72 and are called by the processor 71 to execute the underwater acoustic signal noise reduction method of the embodiments of this application.
[0251] Input / output interface 73, used for information input and output;
[0252] Communication interface 74, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);
[0253] bus 75 , which transmits information between the various components of the device (e.g., processor 71 , memory 72 , input / output interface 73 , and communication interface 74 );
[0254] The processor 71 , the memory 72 , the input / output interface 73 and the communication interface 74 are connected to each other in communication within the device via a bus 75 .
[0255] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned underwater acoustic signal noise reduction method is implemented.
[0256] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0257] The embodiments described in the embodiments of this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0258] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.
[0259] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.
[0260] Those skilled in the art will appreciate that all or some of the steps in the methods, systems, and functional modules / units in the devices disclosed above may be implemented as software, firmware, hardware, or appropriate combinations thereof.
[0261] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0262] It should be understood that in this application, "at least one (item)" and "several" refer to one or more, and "plurality" refers to two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0263] In the several embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the above units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0264] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0265] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0266] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes multiple instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: various media that can store programs, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0267] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but are not intended to limit the scope of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the present invention should be within the scope of the present invention.
Claims
1. A method for reducing noise of underwater acoustic signals, characterized in that: The method comprises: Acquiring a time domain signal of the underwater sound to be de-noised, and performing feature extraction on the time domain signal to be de-noised using a target model to obtain a feature tensor to be de-noised, wherein the feature tensor to be de-noised includes a plurality of local feature blocks; By using the local gating recurrent unit attention mechanism of the target model, local position-dependent features and local time-dependent features between each local feature block and adjacent feature blocks in the feature tensor to be denoised are calculated in parallel to obtain local noise features; Calculating the global position-dependent features and the global time-dependent features of the feature tensor to be denoised through the global gated recurrent unit attention mechanism of the target model to obtain a global noise feature; Reconstructing a noise time domain signal of a target noise from the time domain signal to be denoised based on a noise mask matrix generated by fusing the local noise features and the global noise features through a decoder of the target model; Based on the noise time domain signal, the time domain signal to be denoised is subjected to noise filtering to obtain a target waveform signal.
2. The underwater acoustic signal noise reduction method according to claim 1, characterized in that: The local noise features are obtained by parallelly calculating the local position-dependent features and the local time-dependent features between each local feature block and adjacent feature blocks in the feature tensor to be denoised through the local gated recurrent unit attention mechanism of the target model, including: By using each local attention head of the local gating recurrent unit attention mechanism of the target model, local time dependency features between each local feature block and adjacent feature blocks in the feature tensor to be denoised are calculated in parallel, thereby obtaining local temporal features corresponding to the multiple local feature blocks; Concatenate multiple local time series features corresponding to multiple local attention heads to obtain a local time series feature tensor; The local position-dependent features of the local temporal feature tensor are calculated by the local gating layer of the local gating recurrent unit attention mechanism to obtain local noise features.
3. The underwater acoustic signal noise reduction method according to claim 2, characterized in that: The calculating of the local position-dependent features of the local temporal feature tensor by the local gating layer of the local gating recurrent unit attention mechanism to obtain the local noise features includes: Fusing the local time series feature tensor and the feature tensor to be denoised to obtain a first feature; calculating a local position-dependent feature of the first feature through a local gating layer of the local gating recurrent unit attention mechanism to obtain a second feature; Performing a nonlinear transformation on the second feature through a leaky linear rectification function of the local gated recurrent unit attention mechanism to obtain a third feature; The first feature and the third feature are fused to obtain a local noise feature.
4. The underwater acoustic signal noise reduction method according to claim 1, characterized in that: The global gated recurrent unit attention mechanism of the target model is used to calculate the global position-dependent features and the global time-dependent features of the feature tensor to be denoised to obtain the global noise features, including: Calculating the global time-dependent features of the feature tensor to be denoised through each global attention head of the global gated recurrent unit attention mechanism of the target model, and obtaining the global temporal features corresponding to each global attention head; Concatenate multiple global temporal features corresponding to multiple global attention heads to obtain a global temporal feature tensor; The global position-dependent features of the global temporal feature tensor are calculated through the global gating layer of the global gated recurrent unit attention mechanism to obtain the global noise features.
5. The underwater acoustic signal noise reduction method according to claim 1, characterized in that: The target model is trained in the following way: Acquire a sample time domain signal of the underwater sound to be de-noised, and perform feature extraction on the sample time domain signal using a preset model to obtain a sample time domain signal tensor, wherein the sample time domain signal tensor includes a plurality of sample local feature blocks; By using the local gating recurrent unit attention mechanism of the preset model, the sample local position dependence feature and the sample local time dependence feature between each sample local feature block and the adjacent sample feature blocks in the sample time domain signal tensor are calculated in parallel to obtain the sample local noise feature; Calculating the sample global position-dependent feature and the sample global time-dependent feature of the sample time-domain signal tensor through the global gated recurrent unit attention mechanism of the preset model to obtain the sample global noise feature; Reconstructing a predicted noise time domain signal of the sample noise from the sample time domain signal through a decoder of the preset model based on a sample noise mask matrix generated by fusing the sample local noise feature and the sample global noise feature; Based on the predicted noise time domain signal, filtering the sample time domain signal to obtain a predicted waveform signal; determining a target loss based on the predicted noise time domain signal and the predicted waveform signal; The preset model is trained according to the target loss to obtain a target model.
6. The underwater acoustic signal noise reduction method according to claim 5, characterized in that: The determining of the target loss based on the predicted noise time domain signal and the predicted waveform signal includes: Obtaining a sample noise time domain signal and a sample waveform signal corresponding to the sample water sound to be de-noised; Calculating a corresponding first similarity based on the predicted noise time domain signal and the sample noise time domain signal; Calculating a corresponding second similarity based on the predicted waveform signal and the sample waveform signal; A target loss is determined based on the first similarity and the second similarity.
7. The underwater acoustic signal noise reduction method according to claim 1, characterized in that: The step of extracting features of the time domain signal to be denoised by using a target model to obtain a feature tensor to be denoised includes: Extracting features of the time domain signal to be denoised using a convolutional neural network of a target model to obtain corresponding features to be denoised; Dividing the feature to be denoised into a plurality of local feature blocks according to a preset division length and division hop number; The multiple local feature blocks are spliced to obtain a feature tensor to be denoised.
8. An underwater acoustic signal noise reduction device, characterized in that: The device comprises: an acquisition module, configured to acquire a time domain signal of the underwater sound to be de-noised, and perform feature extraction on the time domain signal to be de-noised using a target model to obtain a feature tensor to be de-noised, wherein the feature tensor to be de-noised includes a plurality of local feature blocks; A first computing module is configured to parallelly compute local position-dependent features and local time-dependent features between each local feature block and adjacent feature blocks in the feature tensor to be denoised through an attention mechanism of a local gated recurrent unit of the target model to obtain local noise features; A second computing module is configured to calculate the global position-dependent features and the global time-dependent features of the feature tensor to be denoised through the global gated recurrent unit attention mechanism of the target model to obtain a global noise feature; A reconstruction module, configured to reconstruct a noise time domain signal of a target noise from the time domain signal to be denoised based on a noise mask matrix generated by fusing the local noise features and the global noise features through a decoder of the target model; The filtering module is used to perform noise filtering on the time domain signal to be denoised based on the noise time domain signal to obtain a target waveform signal.
9. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the underwater acoustic signal noise reduction method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the underwater acoustic signal denoising method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Voice noise reduction method and device, equipment, storage medium and program product
CN114171038A
Speech enhancement model, electronic device, storage medium and related method
CN114333895A
Speech enhancement system based on time modeling generative adversarial network
CN114495958A
Data noise reduction and signal detection method, device and system and storage medium
CN117056680A
Self-supervised speech enhancement method and system based on efficient local attention
CN119028368A
Cited By
Direct wave signal suppression method of continuous wave active sonar detection system
CN120722331A