An end-to-end lightweight neural network time domain noise and interference suppression method for a remote underwater acoustic system
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- XIAMEN UNIV
- Filing Date
- 2026-05-15
- Publication Date
- 2026-08-07
AI Technical Summary
[0007]本发明目的在于克服现有技术中传统接收机级联架构误差传递严重、非高斯复合干扰处理能力弱,以及现有神经网络计算量大、处理延迟高的缺陷,提供一种面向远程水声系统的端到端轻量级神经网络时域噪声与干扰抑制方法
[0023]1. This invention adopts an end-to-end time-domain mapping architecture, which eliminates the need for complex calculations such as frequency domain Fourier transform and preprocessing steps such as additional manual filtering. It directly processes noisy underwater acoustic one-dimensional time-series signals, eliminates the delay and phase reconstruction error caused by frequency domain forward and inverse transform, and significantly reduces processing latency. It can fully adapt to the needs of underwater platforms with limited computing power and high real-time processing requirements.
Smart Images

Figure CN122293216B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of underwater acoustic communication technology, specifically relating to an end-to-end lightweight neural network time-domain noise and interference suppression method for long-range underwater acoustic systems, applicable to complex marine multipath transmission environments and strong non-Gaussian background impulse noise interference. Background Technology
[0002] With the rapid development of marine resource development, marine scientific research, and marine target detection, underwater communication technology has increasingly become a core support for marine information acquisition, interaction, and long-distance transmission. In complex underwater environments, electromagnetic wave transmission attenuation is extremely severe, making sound waves the only effective physical carrier for long-distance underwater information transmission. Various long-range underwater acoustic communication systems are core equipment for completing long-term underwater monitoring, detection, and information exchange tasks.
[0003] Due to the complex physical characteristics of the ocean, long-range underwater acoustic communication systems face multiple severe engineering challenges in practical applications. On the one hand, when underwater acoustic signals propagate in ocean waveguides with surface and seabed boundaries and non-uniform sound velocity profiles, strong refraction, reflection, and scattering effects occur, leading to significant time-varying multipath effects. This results in severe time delay spread and waveform distortion at the receiving end, compromising the integrity and demodulation of the communication signal. On the other hand, the marine environment in which the underwater acoustic communication receiving node is located generates non-Gaussian background impulse noise with obvious impulse characteristics due to factors such as ship navigation, sea surface disturbances, and marine biological activity. Extensive experimental data shows that this type of ocean noise has typical thick-tailed distribution characteristics. Signal processing methods based on the traditional Gaussian distribution assumption cannot achieve effective noise suppression in real marine environments, and may even fail. The combined effect of multipath transmission distortion and strong non-Gaussian background impulse noise directly leads to an extremely low signal-to-noise ratio at the long-range underwater acoustic communication receiving end, significantly reducing the transmission quality and reliability of the communication link.
[0004] To address the aforementioned harsh transmission environments, traditional underwater acoustic communication receivers generally employ a separate cascaded processing architecture. This involves first using a channel equalizer (such as a decision feedback equalizer) to eliminate waveform distortion caused by multipath transmission, and then cascading Wiener filters, wavelet denoising, and other traditional filtering modules to suppress background noise. However, this cascaded architecture has a fatal flaw in environments with extremely low signal-to-noise ratios and strong impulse noise: under strong noise interference, the channel equalizer is prone to misjudgment, the signal distortion generated by the front-end filtering process can disrupt the original phase structure of the multipath channel, leading to the progressive amplification and cumulative propagation of errors, ultimately resulting in ineffective signal recovery.
[0005] In recent years, deep learning technology has provided a novel data-driven solution for underwater acoustic communication signal processing due to its excellent nonlinear fitting capabilities. However, existing deep learning-based underwater acoustic signal processing solutions still have significant shortcomings. Conventional deep convolutional neural network models have a large number of parameters and extremely high floating-point computational complexity, making them difficult to adapt to the computing power and power consumption constraints of underwater embedded communication platforms. Furthermore, most existing network architectures heavily rely on frequency domain processing mechanisms, requiring frequent execution of short-time Fourier transforms and inverse transforms to complete time-frequency domain mapping. Such frequency domain conversion operations introduce significant computational latency, conflicting with the engineering requirements of real-time processing in underwater communication equipment.
[0006] Therefore, there is an urgent need for an end-to-end signal noise and interference suppression scheme that is suitable for long-range underwater acoustic communication, can simultaneously suppress multipath effects and strong non-Gaussian background impulse noise, and is lightweight and has low latency, in order to solve the core contradiction between the limited computing power of underwater communication platforms and signal processing in harsh environments. Summary of the Invention
[0007] The purpose of this invention is to overcome the shortcomings of existing technologies, such as severe error propagation in traditional receiver cascade architectures, weak non-Gaussian composite interference processing capabilities, and the high computational load and processing latency of existing neural networks. This invention provides an end-to-end lightweight neural network time-domain noise and interference suppression method for remote underwater acoustic systems. This invention achieves pure time-domain, low-latency, and high-precision anti-interference processing of underwater acoustic signals in harsh marine environments, simultaneously performing multipath channel distortion compensation and non-Gaussian background impulse noise suppression. This improves the transmission reliability and real-time performance of remote underwater acoustic communication in harsh sea conditions, meeting the engineering application requirements of underwater computing-constrained platforms.
[0008] To achieve the above-mentioned objectives, the present invention provides the following technical solution:
[0009] An end-to-end lightweight neural network-based time-domain noise and interference suppression method for remote underwater acoustic systems, targeting time-varying multipath effects in remote underwater acoustic transmission. For scenarios with stable distributed non-Gaussian background impulse noise interference and limited computing resources on underwater platforms, an end-to-end pure time-domain processing architecture is adopted. This architecture eliminates the need for preprocessing or post-processing steps such as frequency-domain Fourier transform and manual filtering, directly mapping noisy underwater acoustic one-dimensional time-series signals to clean underwater acoustic signals. Specifically, the following steps are included:
[0010] S1. Acquire the noisy one-dimensional observation signal from the receiver of the remote underwater acoustic system. The noisy one-dimensional observation signal Pure water acoustic signal underwater acoustic channel impulse response and obedience Stable distribution of non-Gaussian background impulse noise Composition, satisfaction ,in This represents a linear convolution operation;
[0011] S2. Construct an end-to-end lightweight noise and interference suppression neural network (referred to as 1D DGD-Net). This network sequentially includes a preprocessing module, a core noise reduction module, and a post-processing module. The core noise reduction module is composed of multiple robust, hole-separable residual blocks cascaded together. Each robust, hole-separable residual block integrates a differentiable dynamic amplitude gating (DDAG) module, a one-dimensional hole depth separable convolution operator, and local residual connections to achieve the synergistic effect of non-Gaussian impulse noise suppression, multi-scale temporal feature extraction, and signal detail preservation.
[0012] S3. The noisy one-dimensional observation signal is input into the end-to-end lightweight noise and interference suppression neural network. After the channel dimensionality increase, normalization and activation processing is completed by the preprocessing module, it is sent to the core noise reduction module. Non-Gaussian extreme impulse noise is suppressed by the differentiable dynamic amplitude gating module. Multi-scale time-domain features are extracted by the one-dimensional hole depth separable convolution operator. Combined with local residual connection and global noise residual mapping, the pure underwater acoustic signal estimate is output.
[0013] S4. Using one-dimensional time-series mean square error (MSE) as the loss function for network optimization, the loss is calculated based on the estimated value of the pure underwater acoustic signal and the corresponding pure reference signal. The network parameters are iteratively trained and optimized until the network converges, resulting in a trained end-to-end lightweight noise and interference suppression neural network.
[0014] S5. Input the noisy underwater acoustic signal to be processed at the receiving end into the trained end-to-end lightweight noise and interference suppression neural network, and output the final pure underwater acoustic signal to achieve multipath distortion compensation and non-Gaussian background impulse noise suppression, thus completing low-latency real-time signal purification.
[0015] In step S1, the pure underwater acoustic signal is a hyperbolic frequency modulation (HFM) signal, and the stable distribution is a zero-mean symmetric α stable distribution, which is used to accurately characterize the thick-tailed characteristics of marine non-Gaussian background impulse noise. Its characteristic index ranges from 1.2 to 2.0.
[0016] In step S2, the preprocessing module sequentially includes a one-dimensional up-dimensional convolution, a layer normalization layer, and a ReLU activation layer; the postprocessing module is a one-dimensional down-dimensional convolution layer, used to match the input and output dimensions of the one-dimensional underwater acoustic time-series signal.
[0017] In step S2, the differentiable dynamic amplitude gating module adopts a continuously differentiable nonlinear mapping mechanism, which can independently and adaptively learn threshold parameters for different feature channels to achieve lossless transmission of weak effective underwater acoustic signals and smooth amplitude limiting suppression of strong impulse noise. The one-dimensional hole depth separable convolution operator decouples the standard one-dimensional time-domain convolution into a one-dimensional hole depth convolution in the time dimension and a one-dimensional pointwise convolution in the channel dimension, which greatly reduces the number of model parameters and computational complexity while expanding the temporal receptive field. The hole rate of each level of robust hole separable residual block increases progressively in a geometric progression of 2, which is suitable for the extraction of multipath delay features in underwater acoustic channels.
[0018] In step S2, the number of robustly hollow separable residual blocks is 3 to 5, the number of basic feature channels of the network is set to 16, the convolution kernel size is fixed at 5, and the hole rate of each residual block is set to 1, 2, 4, and 8 respectively to expand the temporal receptive field for underwater acoustic multipath delay extraction; the threshold parameters and specific values of the hole rate of the micro-dynamic amplitude gating module can be adapted and adjusted according to the actual underwater acoustic scenario requirements.
[0019] In step S4, the network training uses the Adam optimizer and a one-dimensional temporal mean square error loss function, combined with gradient pruning and early stopping strategies to iteratively optimize until convergence.
[0020] This invention also provides a remote underwater acoustic system, comprising a signal receiving module, a noise and interference suppression module, and a signal output module connected in sequence. The signal receiving module acquires a noisy one-dimensional observation signal and transmits it to the noise and interference suppression module. The noise and interference suppression module incorporates the aforementioned end-to-end lightweight noise and interference suppression neural network, executes the aforementioned noise and interference suppression method, and purifies the noisy underwater acoustic signal. The signal output module receives the purified underwater acoustic signal output by the noise and interference suppression module and outputs it externally. This system is suitable for underwater computing resource-constrained scenarios, enabling real-time, high-precision transmission of remote underwater acoustic signals.
[0021] Furthermore, the signal receiving module can be a deep-sea hydrophone, and the signal output module can be a signal demodulator or a data transmission interface.
[0022] This invention can solve the signal denoising and interference suppression problems of underwater computing resource-constrained platforms in complex marine multipath transmission environments and under strong non-Gaussian background impulse noise interference. Compared with the prior art, this invention has the following outstanding advantages and technical effects:
[0023] 1. This invention adopts an end-to-end time-domain mapping architecture, which eliminates the need for complex calculations such as frequency domain Fourier transform and preprocessing steps such as additional manual filtering. It directly processes noisy underwater acoustic one-dimensional time-series signals, eliminates the delay and phase reconstruction error caused by frequency domain forward and inverse transform, and significantly reduces processing latency. It can fully adapt to the needs of underwater platforms with limited computing power and high real-time processing requirements.
[0024] 2. By introducing a differentiable dynamic amplitude gating (DDAG) mechanism, the threshold parameter can be adaptively learned to suppress non-Gaussian background impulse noise that follows an α-stable distribution. While preserving the details of the effective underwater acoustic signal to the maximum extent, it effectively isolates the damage of extreme impulses to network feature extraction, avoids feature extraction distortion, and significantly improves the robustness of the network in complex marine environments dominated by strong impulse noise. This addresses the engineering challenge of suppressing non-Gaussian impulse noise in marine underwater acoustics.
[0025] 3. A one-dimensional hole depth separable convolution operator is adopted as the core operator of the network. The standard convolution is innovatively decoupled into a depth convolution in the time dimension and a pointwise convolution in the channel dimension. While maintaining the signal representation capability, the number of model parameters and computational complexity are reduced, making it suitable for underwater computing resource-constrained scenarios.
[0026] 4. A residual learning paradigm combining global and local methods is adopted. Local residual connections are used to preserve the temporal details of the underwater acoustic signal, while global residual mapping is used to learn the composite components of noise and interference. Combined with a pure time-domain mean square error loss function, dual suppression of multipath interference and non-Gaussian background impulse noise is achieved, significantly improving waveform reconstruction accuracy. Numerical simulation and 100km sea trials have confirmed that a signal-to-noise ratio gain of over 10dB can still be achieved under extremely low input signal-to-noise ratio (-10dB), and the output signal-to-noise ratio is above 18dB in strong impulse noise environments, meeting the practical engineering application requirements of long-range underwater acoustic systems. Attached Figure Description
[0027] Figure 1 This is a schematic diagram of the overall architecture of an end-to-end one-dimensional time-domain network.
[0028] Figure 2 The graph shows the time-domain waveform and instantaneous frequency variation trend of the HFM signal; where (a) is the time-domain waveform of the HFM signal and (b) is the instantaneous frequency of the HFM signal.
[0029] Figure 3 The waveforms show the time-domain characteristics of α-stable distribution noise and Gaussian noise; where (a) is Gaussian distribution (α=2.0) and (b) is α-stable distribution (α=1.4).
[0030] Figure 4The flowchart and response curve of the Differentiable Dynamic Amplitude Gated (DDAG) processing mechanism are shown below; the upper figure is the DDAG nonlinear response curve, and the lower figure is the DDAG processing flow.
[0031] Figure 5 The diagram shows the structure of a lightweight residual network based on DDAG and multi-scale convolution.
[0032] Figure 6 This is a comparison chart of signal-to-noise ratio and gain in simulation experiments;
[0033] Figure 7 The following are time-domain and time-frequency domain comparison diagrams before and after signal processing in the simulation experiment; where (a) is the time-domain waveform of the pure HFM reference signal, (b) is the time-frequency domain feature diagram of the pure HFM reference signal, (c) is the time-domain waveform of the original received signal containing α-stable distribution non-Gaussian background impulse noise and multipath distortion, (d) is the time-frequency domain feature diagram of the original received signal containing α-stable distribution non-Gaussian background impulse noise and multipath distortion, (e) is the time-domain waveform of the reconstructed signal after processing by the 1D DGD-Net of this invention, and (f) is the time-frequency domain feature diagram of the reconstructed signal after processing by the 1D DGD-Net of this invention.
[0034] Figure 8 The graph shows the robustness evaluation curves for non-Gaussian background impulse noise.
[0035] Figure 9 This is a cross-comparison chart of algorithm complexity and inference latency;
[0036] Figure 10 Diagram showing the location of the transmitter and receiver for the sea trial;
[0037] Figure 11 This is a comparison chart of signal-to-noise ratio and gain in sea trials.
[0038] Figure 12 The images show a comparison of the full waveforms before and after signal processing in the sea trial experiment. (a) is the time-domain waveform of the transmitted reference signal, (b) is the time-frequency domain waveform of the transmitted reference signal, (c) is the time-domain waveform of the original received signal after 100km long-distance transmission, (d) is the time-frequency domain feature diagram of the original received signal after 100km long-distance transmission, (e) is the time-domain waveform of the reconstructed signal after processing according to this invention, and (f) is the time-frequency domain feature diagram of the reconstructed signal after processing according to this invention. Detailed Implementation
[0039] The following embodiments, in conjunction with the accompanying drawings, will further illustrate the present invention. These embodiments are for illustrative purposes only and are not intended to limit the scope of protection. Details not described in detail can be achieved using conventional techniques in the art.
[0040] 1. End-to-end one-dimensional time-domain network architecture
[0041] Figure 1 A specific end-to-end one-dimensional time-domain network architecture is presented. This network can directly perform signal denoising / enhancement tasks in the time domain without additional manual preprocessing steps such as frequency domain transformation, achieving direct mapping from the original noisy signal to a clean signal. Figure 1 In the diagram, the top dimension labels illustrate the shape changes of the feature data, where M represents the length of the temporal feature sequence, which remains constant throughout to ensure consistent input and output lengths; d model This represents the feature channel dimension of the model (the number of channels in the feature map after channel dimensionality upscaling); input [1×M] single-channel temporal features, which are preprocessed and upscaled to [d] model [×M], the core module processes the signal while maintaining this dimension, and the post-processing compresses it back to [1×M], finally outputting a clean [1×M] signal; the network consists of five modules from left to right: input, preprocessing, core noise reduction, post-processing, and output: the input stage inputs the noisy one-dimensional time-domain raw single-channel signal; the preprocessing stage includes one-dimensional convolution, layer normalization, and ReLU activation serial operations, and the one-dimensional convolution maps the input to d model The channel maintains the temporal feature sequence length M unchanged, layer normalization stabilizes training, and ReLU activation introduces nonlinearity; the core denoising module consists of three serial and structurally identical residual blocks, maintaining the feature dimension [d]. model [×M], the input features within the residual block are divided into two parallel paths: each residual block's main path has a differentiable dynamic amplitude gating module at the front end, which suppresses non-Gaussian background impulse noise through adaptive smoothing clamping, avoiding strong impulse outliers from damaging subsequent feature extraction; then, a cascaded one-dimensional dilated depth convolution and a one-dimensional pointwise convolution are performed for multi-scale temporal feature extraction; the skip branch path performs identity mapping; the outputs of the main path and the skip branch path are fused through ⊕ (element-wise addition operation) to achieve local residual mapping; the post-processing stage is a one-dimensional convolution with a large convolution kernel, compressing the channels to 1 and completing signal reconstruction; the output stage outputs a clean single-channel temporal signal, realizing end-to-end signal processing.
[0042] 2. Remote underwater acoustic signal and channel model
[0043] 2.1 Hyperbolic Frequency Modulation (HFM) Transmit Signal Model
[0044] In long-range underwater acoustic systems, the distortion resistance of the transmitted waveform is the primary physical condition for ensuring reliable transmission. Hyperbolic frequency modulation (HFM) signals, with their excellent Doppler invariance and large time-bandwidth product, can effectively cope with the severe Doppler frequency shift caused by the relative motion of underwater platforms. The instantaneous frequency of the HFM signal follows a hyperbolic variation law, and its frequency varies over a given duration of the HFM signal. From the initial frequency Smooth transition to termination frequency .
[0045] By performing rigorous integration on the instantaneous frequency, the phase function of the signal can be derived, thus obtaining its precise time-domain expression:
[0046]
[0047] in, Represents pure water acoustic signals. The radiation amplitude of the HFM signal. It is a cosine function. Let be the initial phase of the HFM signal, and 2π be a mathematical constant used to construct the phase term and realize the mapping between angle and phase period. The initial frequency of the HFM signal. This is the termination frequency of the HFM signal. The duration of the HFM signal. It is the natural logarithm function. The instantaneous time variable of the HFM signal; when affected by the Doppler effect, the waveform characteristics of the signal are mathematically represented only by the equivalent delay on the time axis, while the core frequency components remain stable; this characteristic enables the HFM signal to maintain an ideal time structure even after extremely long-distance transmission, providing a stable and easy-to-extract observation object for subsequent one-dimensional time-domain neural networks; Figure 2 The time-domain waveform and instantaneous frequency variation trend of the HFM signal are given. Figure 2 (a) in the figure is the time-domain waveform of the HFM signal (first 50ms). It can be seen that the time-domain waveform is a constant-amplitude high-frequency oscillation with stable amplitude and no obvious distortion, indicating that the signal transmission power is constant and the time-domain structure is regular. Figure 2 (b) in the figure represents the instantaneous frequency of the HFM signal. It can be seen that the instantaneous frequency increases monotonically with time, smoothly transitioning from the initial frequency of 1000Hz to the terminal frequency of about 1500Hz within 0~1s. This intuitively presents the hyperbolic frequency modulation characteristics of the HFM signal (which exhibits an approximately linear upward trend under this parameter). This modulation law endows the signal with excellent Doppler invariance.
[0048] 2.2 Underwater Acoustic Multipath Channel Model
[0049] The propagation of water waves in ocean waveguides is severely affected by sea surface reflection, seabed scattering, and the non-uniform distribution of sound velocity within the medium. In long-distance signal transmission scenarios, multipath effects can cause the received signal to spread drastically along the time axis. In order to accurately characterize the far-field propagation characteristics of sound waves in horizontally non-uniform media, this invention uses a parabolic equation (PE) model to describe the underwater acoustic channel.
[0050] The PE model, by removing backscattered energy, approximates the elliptic boundary value problem in the frequency domain wave equation as an initial value problem, thereby significantly improving the numerical solution efficiency while maintaining computational accuracy. In cylindrical coordinates, the equation of the unidirectional propagating parabola can be expressed as:
[0051]
[0052] in, The sign of the partial derivative. For complex sound pressure, For transmission distance, To meet The imaginary unit of ² = -1 For reference wavenumber, The angular frequency of the sound wave. For reference speed of sound, This is a depth operator related to depth and medium density; by solving this equation through step integration, the propagation distance of the sound wave at any distance can be obtained. The sound field distribution at the location; after converting the frequency domain sound field to the time domain, the impulse response of the long-range underwater acoustic channel can be equivalent to a non-stationary process composed of a large number of multipath components with different time delays, amplitudes and phases superimposed.
[0053] 2.3 Non-Gaussian Ocean Noise Model
[0054] In the real, vast ocean environment, background noise caused by wave disturbances, biological activity, and ship navigation often exhibits strong non-Gaussian and thick-tailed characteristics. The probability of extreme values appearing in the statistical distribution of this noise is much higher than that of the traditional Gaussian distribution, resulting in a large number of spike pulses being mixed in the received signal.
[0055] To overcome the theoretical blind spots of traditional Gaussian models in dealing with such physical phenomena, this invention introduces... Establish a noise model based on a stable distribution; due to The probability density function of a stable distribution typically lacks a closed-form analytical expression, and its statistical properties must be rigorously defined through characteristic functions. In engineering practice, ocean background noise is usually modeled using zero-mean and symmetric methods. Stable distribution ( This is achieved by ), and its characteristic function is simplified to:
[0056]
[0057] In the formula, The background noise characteristic index directly determines the thickness of the tail of the probability density curve (i.e., the intensity of the pulse). The scale parameter is used to quantify the degree of dispersion of noise samples relative to the center; when When this distribution no longer possesses finite second-order statistical moments (infinite variance), this is the fundamental reason why classical signal processing algorithms based on second-order moment theory fail in underwater acoustic environments. Figure 3 Provide typical underwater acoustic environments A comparison of the time-domain characteristics of stable distributed noise and Gaussian noise, among which... Figure 3 In the diagram, (a) is the time-domain waveform of Gaussian noise when α = 2.0. Figure 3 (b) in the middle is Time-domain waveform of α-stable distribution non-Gaussian background impulse noise; from Figure 3 It can be seen from this that The Gaussian distributed noise exhibits smooth amplitude fluctuations in the time domain waveform, without significant large impulse peaks, and generally displays stable, white noise-like characteristics; while The time-domain waveform of the α-stable distribution noise has a large number of spikes with amplitudes much higher than the background level, exhibiting significant impact characteristics. This is highly consistent with the characteristics of non-Gaussian impact noise caused by marine organisms, platform movement, and other factors in the underwater acoustic environment, demonstrating the modeling ability of the α-stable distribution for non-Gaussian heavy-tailed noise. The α-stable distribution modeling of underwater acoustic background noise in this invention is more consistent with the noise characteristics of the actual underwater acoustic channel than the Gaussian distribution.
[0058] 2.4 Remote underwater acoustic system observation model
[0059] Considering the multipath evolution of the underwater acoustic signal waveguide transmission stage and the non-Gaussian environmental noise in the receiving stage, the end-to-end one-dimensional physical observation model of the remote underwater acoustic system is expressed as follows:
[0060] (4)
[0061] in, The noisy multipath received signal observed at the receiving end. Represents pure water acoustic signals. This represents a linear convolution operation. It is the impulse response of the underwater acoustic channel. To show obedience Stable distribution of non-Gaussian background impulse noise; the noise and interference suppression problem of this invention is transformed into a highly nonlinear inverse mapping problem in the pure time domain; unlike existing technologies, it requires... This invention proposes to perform an extremely time-consuming short-time Fourier transform to extract the amplitude and phase spectra, and instead seeks a method in the pure time domain that utilizes network parameters. Dominated optimal residual mapping operator Its optimization objective is defined as:
[0062] (5)
[0063] in, For learnable network parameters governed by mapping operators, The goal is to optimize the network parameters (i.e., the parameter solution corresponding to the minimum value of the loss function). To find an optimization operator in the parameter space that minimizes the objective function, This is the loss function used to measure the error between the model output and the clean signal. For network parameters Dominant pure time-domain end-to-end mapping operator, For the mapping operator to the received signal The output (i.e., the signal estimate obtained by network recovery).
[0064] Under this model, the network does not need to explicitly estimate the parameters of channel distortion and non-Gaussian background impulse noise distribution; instead, it relies on massive amounts of noisy observation data. Direct driving, searching within the implicit domain for a value derived from network parameters. Dominated optimal residual mapping operator It is used to directly estimate the composite interference component of channel distortion and non-Gaussian background impulse noise.
[0065] This mechanism eliminates the delay and phase reconstruction error caused by frequency domain forward and inverse transformation, and achieves dual anti-interference recovery against transmission multipath and received noise while ensuring low processing latency.
[0066] 3. End-to-end lightweight noise and interference suppression neural network design
[0067] This invention proposes an end-to-end lightweight noise and interference suppression neural network for the transmission and reception links of remote underwater acoustic systems; this network architecture directly ingests noisy one-dimensional observation signals that have experienced channel fading and environmental disturbances in the time dimension. The pure underwater acoustic signal estimate is directly output through time-domain nonlinear mapping. In order to cope with the impulse response of the underwater acoustic channel defined in the aforementioned system model. Obey the receiving end Stable distribution of non-Gaussian background impulse noise To address the complex physical interference, this invention performs targeted acoustic physical mechanism mapping on the underlying operators and global topology.
[0068] 3.1 Micro-Dynamic Amplitude Gating (DDAG) Module
[0069] In response to complex marine environmental background noise, especially conforming to To address stable, distributed strong pulse interference, this invention designs a Differentiable Dynamic Amplitude Gating (DDAG) mechanism as a sub-module at the underlying operator level. Standard convolution operators perform linear accumulation within the receptive field; when encountering abnormal pulses with extremely high amplitudes, this can lead to feature overload and instantaneous amplification of errors. To fundamentally block this pulse propagation, the Differentiable Dynamic Amplitude Gating module (DDAG module) introduces an adaptive soft threshold before deep feature extraction. The mathematical model of the Differentiable Dynamic Amplitude Gating module is expressed as follows:
[0070]
[0071] in, It is a one-dimensional time series feature matrix. This is the robust time series feature matrix. It is the hyperbolic tangent function. For adaptive dynamic learning, the positive smoothing threshold parameter is used. The differentiable dynamic amplitude gating module, which can be independently learned for each feature channel during network training, employs a continuously differentiable nonlinear mapping mechanism to independently and adaptively learn threshold parameters for different feature channels. This enables lossless transmission of weak, effective underwater acoustic signals and smooth, amplitude-limited suppression of strong impulse noise. Specifically, when the input signal is a normal, weak, effective acoustic component (i.e.... According to Taylor expansion theory, this gating module is equivalent to near-lossless linear transmission, preserving the original waveform details of the underwater acoustic signal to a great extent; however, when encountering extreme non-Gaussian background impulse noise ( The gated output is smoothly and continuously differentially clamped in Near the boundary; this mechanism effectively isolates non-Gaussian outliers from disrupting subsequent network channel feature fusion at the physical level; Figure 4 The upper figure shows the nonlinear response curve of DDAG. The horizontal axis represents the input amplitude, and the vertical axis represents the output amplitude. The figure compares the response curves under five different threshold parameters. It can be seen that the adaptive dynamic learning positive smoothing threshold parameter learned by the method proposed in this invention... ( Figure 4 The curve in the image: Adaptive learning threshold It exhibits linear transmission characteristics in the central region (weak signal), while displaying smooth suppression and limiting characteristics in the large value regions on both sides (impulse noise), and compared to fixed soft and hard thresholds ( Figure 4 The curve in the middle: 1 soft threshold ( 1) 1 Hard threshold ( 1) 2 Hard threshold ( 2) and 3 Hard threshold ( 3) Its continuous differentiability is more conducive to the backpropagation of network gradients, which improves the stability and convergence speed of network training; Figure 4 The following diagram illustrates the processing flow of a DDAG: the input size is... Features First, the positive smoothing threshold parameter for adaptive dynamic learning is obtained through the Softplus activation function layer. Subsequently, the features of the original input Multiply Perform amplitude normalization, then input to The function layer performs amplitude limiting; finally, the limiting result is multiplied by the adaptively dynamically learned positive smoothing threshold parameter. Perform amplitude inverse normalization to obtain the final output robust time series feature matrix. .
[0072] 3.2 Multi-scale robust void separable residual block
[0073] After effectively filtering out extreme non-Gaussian background impulse noise using a differentiable dynamic amplitude gating (DDAG) module, the network needs to perform deep feature extraction on the long-term time-series evolution of the hyperbolic frequency modulation (HFM) signal to capture the temporal details and multipath propagation characteristics of the underwater acoustic signal. This invention employs a multi-scale feature extractor composed of multiple robust, dilated, separable residual blocks cascaded together to complete temporal feature mining. To obtain a sufficiently large temporal receptive field without increasing computational parameters and to adapt to the requirements of underwater acoustic multipath delay feature extraction, this invention uses a one-dimensional dilated, depth-separable convolution as the core feature extraction operator. The standard one-dimensional temporal convolution is mathematically decoupled into a depthwise convolution in the spatial (temporal) dimension and a pointwise convolution in the channel dimension. This decoupling design significantly reduces the model's computational complexity and inference latency, achieving lightweight and high-precision feature extraction. To clarify the operator's operational logic, let the sequence length of the input features be... The number of input channels is The number of output channels is The kernel size is The void ratio is One-dimensional dilated depthwise convolution operates independently on each input feature channel, capturing only the local temporal correlation within a single channel. Its discrete operation expression is:
[0074]
[0075] in, For the first Each channel in the time-series index The depthwise convolution output at that point; The kernel size; The weight parameters are those of the one-dimensional hole depth convolution kernel. For input feature number The time-series index on each channel is The characteristic values; ∑ represents the void ratio; ∑ is the summation symbol.
[0076] After completing the single-channel temporal feature extraction, one-dimensional pointwise convolution is used to utilize... The convolutional kernel performs cross-channel feature aggregation along the channel dimension, enabling information exchange between channels and feature dimension adjustment. Its operational expression is:
[0077]
[0078] in, For the first Each channel in the time-series index The pointwise convolution output at each point, and the channel index The range of values is to , Number of output channels; Input the number of channels; The weight parameters are for one-dimensional pointwise convolution. For the first Each channel in the time-series index The depthwise convolution output at the specified point; ∑ is the summation symbol.
[0079] To achieve multi-scale temporal feature extraction, this invention uses the hole rate... In consecutive residual blocks, the increments exponentially (e.g., as set in this embodiment). The network constructs an exponentially expanded temporal receptive field, effectively capturing multipath signal components with different time delays; compared to standard convolution, the number of parameters in this one-dimensional hole-depth separable convolution structure can be approximately reduced to [a smaller percentage of the original value]. This significantly reduces the memory usage and floating-point operations of underwater equipment, making it suitable for engineering scenarios where underwater computing resources are limited.
[0080] Figure 5 This paper presents a detailed lightweight residual network structure based on differentiable dynamic amplitude gating (DDAG) and multi-scale convolution; such as... Figure 5 As shown, this structure mainly includes a feature processing path for the feature input and a core residual connection path. These two paths work together to achieve robust feature extraction and noise suppression. The feature processing path consists of a one-dimensional temporal feature matrix. (dimension is) Combined with learning parameters Input the Softplus activation function to generate positive smoothing threshold parameters that are adaptively and dynamically learned for each channel. One-dimensional time series feature matrix Multiply by in sequence ,pass The function layer performs amplitude limiting and multiplies by an adaptively dynamically learned positive smoothing threshold parameter. Perform amplitude inverse normalization to output robust time series feature matrix. This process achieves secondary suppression of non-Gaussian background impulse noise. The core residual connection path consists of a differentiable dynamic amplitude gating module, a one-dimensional hole depth separable convolution module, and a one-dimensional pointwise convolution module cascaded together. The differentiable dynamic amplitude gating module is used to suppress strong impulse noise and avoid extreme values interfering with subsequent feature extraction. The one-dimensional hole depth separable convolution module contains a convolution kernel with a size of... Step size is The coefficient of thermal expansion is And the number of channels is The network employs deep convolutions followed by cascaded batch normalization (BN) layers and parameterized rectified linear unit (PReLU) activation functions to introduce non-linear features and stabilize network training; the one-dimensional pointwise convolution module uses... The convolution kernel reduces the number of feature channels from Mapped to The network then reconnects to the BN layer to achieve feature dimension adjustment and channel information fusion. To achieve multi-scale feature extraction, the network employs multi-scale dilation (dilation coefficient). = Cascaded blocks), lightweight channel count set to The kernel size is Finally, the robust temporal features extracted from the main path are added element-wise to the original input features through residual connections. ), output dimension is The feature matrix enables efficient preservation and fusion of temporal features.
[0081] In summary, a collaborative noise reduction structure is formed by a differentiable dynamic amplitude gating module, a one-dimensional hole-depth separable convolution module, a one-dimensional pointwise convolution module, and residual connections. The differentiable dynamic amplitude gating module first adaptively smooths and clamps non-Gaussian background impulse noise, eliminating the interference of extreme impulse outliers on feature extraction. Then, the one-dimensional hole-depth separable convolution operator expands the temporal receptive field while reducing computational cost, extracting multi-scale temporal features. Finally, residual connections fuse the gating and convolutional features with the original input features, preserving the temporal details of the underwater acoustic signal. The differentiable dynamic amplitude gating module, the one-dimensional hole-depth separable convolution operator, and the local residual connections work synergistically and closely together to simultaneously achieve non-Gaussian background impulse noise suppression and multipath distortion compensation in a pure temporal end-to-end architecture, ensuring both the depth and accuracy of feature extraction while also considering the model's lightweight and real-time performance.
[0082] 3.3 Residual Learning Paradigm and Loss Function
[0083] In the global mapping design of the network topology, this invention adopts a residual learning paradigm that combines global and local approaches. Addressing the weakness of target signals in remote underwater acoustic environments with extremely low signal-to-noise ratios, this invention abandons the traditional method of directly fitting the absolute waveform of a clean broadband underwater acoustic signal. Instead of forcibly fitting the complex absolute waveform of a broadband signal, this invention constructs a learning mechanism based on time-domain residual mapping, enabling the network to focus on learning the time-domain composite interference residual mapping between the input noisy observation signal and the clean underwater acoustic signal. The final end-to-end time-domain output relationship is defined as follows:
[0084]
[0085] in, This is an estimated value for the pure underwater acoustic signal; It is a noisy one-dimensional observation signal; For network parameters Dominant pure time-domain end-to-end mapping operator; The estimated composite components of noise and interference; These are the optimized parameters for the network; these are the parameters that drive the network. Through iterative optimization, this method abandons the hybrid loss function involving frequency domain transformation and instead constructs a purely time-domain supervised objective function. Considering gradient stability under residual impulse noise, the network adopts a one-dimensional time-series mean square error (MSE) based on a finite-length discrete sampling sequence as the final optimization criterion.
[0086]
[0087] in, This represents the one-dimensional time series mean square error. This represents the total number of signal sampling points. For sampling point index; The summation symbol; For pure water acoustic signals The actual value of each sampling point; The estimated value of the pure underwater acoustic signal is the first The numerical values of each sampling point; through the above-mentioned end-to-end mapping in the pure time domain and the in-depth innovation of the underlying operators, this network achieves efficient direct time-domain suppression of strong non-Gaussian ocean noise while maintaining low computational complexity.
[0088] 4. Simulation Experiment and Result Analysis
[0089] The simulation experiment was carried out using a hyperbolic frequency modulation (HFM) signal as the reference signal. The duration of a single frequency sweep signal was set to 1 second. Table 1 shows the simulation experiment parameters.
[0090] Table 1 Simulation Parameters
[0091]
[0092] 4.1 Simulation Environment and Dataset Construction
[0093] The far-field propagation characteristics of the underwater acoustic channel were solved numerically with high precision using the RAM (Range-dependent Acoustic Model) parabolic equation package. To simulate the non-homogeneity and diversity of the real ocean medium and seabed sediment, multiple random variables were introduced into the simulation: the propagation distance was randomly generated within a 100km range, and the water depth was randomly perturbed between 80m and 190m. The sound velocity profile was constructed as a typical negative gradient structure, while the longitudinal wave velocity (1500–1700 m / s), density (1.4–1.9 g / cm³), and attenuation coefficient (0.1–0.5 dB / s) of the seabed medium were also simulated. All samples are randomly combined within a reasonable range; by solving the frequency domain sound pressure under the above complex environment and performing inverse Fourier transform, the time domain impulse response with long delay extension is extracted.
[0094] The generation of noisy data employs a joint noise-channel simulation strategy; the clean HFM signal is discretized and convolved with the aforementioned time-varying PE channel impulse response, and then superimposed following a certain formula. Stable distribution of non-Gaussian background impulse noise; signal bandwidth B is defined as the difference between the termination frequency and the initial frequency of the HFM signal, in Hz, used to characterize the frequency coverage range of the HFM signal. In this simulation, the signal bandwidth B = 500Hz is set to match the center frequency of 1250Hz. The constructed HFM signal time-bandwidth product meets the requirements of long-distance transmission, ensuring both the signal's anti-interference capability and avoiding the increased signal attenuation caused by excessive bandwidth, adapting to the propagation characteristics of a 100km long-distance underwater acoustic channel; considering the statistical regularities of the real marine environment, the background noise characteristic index... Set to randomly select values within the interval [1.2, 2.0]; for Given the physical phenomenon of time-varying noise variance, the traditional global variance signal-to-noise ratio definition is no longer used. This experiment adopts a definition based on the energy ratio of signal intervals to calibrate the target signal-to-noise ratio:
[0095]
[0096] in, and These represent the time-domain power of the pure signal component and the separated noise component, respectively. The simulation generated 1000 channel and noise combinations with independent environmental parameters, of which 800 were used as the training set and 200 as an independent test set for performance evaluation. All simulation experiments were performed on a standard experimental server to ensure the stability and reproducibility of the simulation process. All experimental data are the average of three independent repeated simulations to reduce the impact of random errors on the experimental results.
[0097] 4.2 Network Training Setup
[0098] The network model is trained using the PyTorch deep learning framework. In this embodiment, the basic number of network channels is set to 16, matching the 500Hz bandwidth and 9600Hz sampling frequency of the HFM signal. This ensures accurate feature extraction while controlling the number of model parameters. The convolution kernel size is set to 5, adapting to the local feature extraction of one-dimensional time-series signals, avoiding the increased computation caused by excessively large kernels and the insufficient feature extraction caused by excessively small kernels. Multi-scale dilated residual blocks are used for one-dimensional depth-separable convolution operations. The Adam optimizer is used for parameter updates during training, with an initial learning rate set to 0.001. To drive the network to accurately approximate the target signal waveform in the pure time domain, this embodiment uses one-dimensional temporal mean square error (MSE) as the global loss function. The batch size is set during training. The size is set to 8, and the maximum number of training epochs is set to 100. At the same time, a gradient clipping mechanism (limiting the gradient norm to a maximum value of 10) and an early stopping strategy (terminating the model early if there is no significant improvement after 10 consecutive epochs) are introduced to effectively prevent the model from experiencing gradient explosion or overfitting under non-Gaussian strong impulse noise interference.
[0099] 4.3 Simulation Results and Performance Comparison
[0100] To fully verify the performance of the proposed end-to-end one-dimensional time-domain lightweight noise and interference suppression network (1DDGD-Net) in processing underwater acoustic signals, a detailed performance comparison was conducted with four benchmark algorithms, including: one-dimensional convolutional neural network (1D-CNN), Wiener filter, adaptive filter, and wavelet denoising.
[0101] Figure 6 This paper presents signal-to-noise ratio (SNR) gain evaluations under different input SNRs. This embodiment evaluates the noise and interference suppression performance of various methods within the input SNR range of -10dB to 10dB. Figure 6 It can be seen that traditional filtering algorithms based on statistics or fixed basis functions (Wiener filtering, adaptive filtering, wavelet denoising) have extremely limited output SNR improvement when facing complex non-Gaussian background impulse noise in underwater acoustic channels, and are difficult to effectively suppress strong non-Gaussian impulse noise. In contrast, deep learning methods (one-dimensional convolutional neural network 1D-CNN and the 1D DGD-Net of this invention) show significant advantages in nonlinear mapping. Among them, the lightweight 1D DGD-Net proposed in this invention achieves better performance than the 1D-CNN model in all input SNR frequency bands. This proves that while making the network lightweight, this invention not only retains its core deep feature fitting ability, but also further optimizes the suppression effect on non-Gaussian background impulse noise, effectively breaking through the performance bottleneck of traditional filtering algorithms. It has practicality and superiority in noise and interference suppression scenarios of remote underwater acoustic systems.
[0102] Figure 7 The time-domain and time-frequency domain comparisons of the signal before and after passing through the network are presented under extremely low signal-to-noise ratio and multipath interference. This time-frequency domain comparison diagram is only used for the visualization analysis of algorithm performance. Figure 7 (a) and (b) in the diagram are clean reference signals. It can be seen that the time-domain waveform of the transmitted reference signal is regular, the amplitude is stable, and its time-frequency domain energy distribution is concentrated, exhibiting clear hyperbolic frequency modulation (HFM) signal characteristics, providing an ideal reference for subsequent processing. Figure 7 (c) and (d) in the text are superimposed. The noisy signal with stable noise distribution clearly shows strong pulse interference and severe multipath tail distortion, and the original waveform of the signal is severely submerged by a large amount of noise. Figure 7 (e) and (f) in the figure are the output signals reconstructed by the method of the present invention; it can be seen that the time-domain envelope of the output signal of the present invention is similar to that of the original transmitted reference signal ( Figure 7 (a) in the text is highly consistent with the ideal state of the time-frequency domain energy distribution, which is clear and concentrated, and is consistent with the time-frequency characteristics of the transmitted reference signal. Figure 7The results are basically consistent with (b) in the original, with no obvious distortion. Experiments show that the method of the present invention not only suppresses high-amplitude pulse background noise, but also removes the trailing caused by multipath, and achieves high-precision recovery of the original pure signal in an environment with extremely low signal-to-noise ratio. This verifies the excellent robustness and feature learning ability of the differentiable dynamic amplitude gating (DDAG) module and multi-scale one-dimensional hole depth separable convolution proposed in the present invention in complex interference scenarios.
[0103] Figure 8 Give the background noise feature index of each algorithm The output SNR decay trend as the value decreases from 2.0 (degenerates to a Gaussian distribution) to 1.2 (strong impulses are non-Gaussian distributed); the figure contains five comparison curves, corresponding to the method of this invention, one-dimensional convolutional neural network (1D-CNN), Wiener filtering, adaptive filtering, and wavelet denoising, respectively. Figure 8 As can be seen, when the background noise characteristic index α = 2.0, the α-stable distribution noise degenerates into a Gaussian distribution. At this point, all algorithms can achieve a certain noise suppression effect, among which the method of this invention has the highest output signal-to-noise ratio, approximately 20 dB. As the background noise characteristic index α decreases from 2.0 to 1.2, the impulse characteristics of the noise increase sharply. This is because traditional filtering algorithms highly depend on the second-order statistical moments (variance within the background noise characteristic index). The noise suppression performance drops sharply due to the fact that the background noise feature index α approaches infinity, with wavelet denoising showing the most significant decline. When the background noise feature index α = 1.2, the output signal-to-noise ratio is only about 8.8 dB, which cannot effectively suppress strong impulse interference. At the same time, the output signal-to-noise ratio performance of the baseline one-dimensional convolutional neural network also shows a significant accelerated decline. This is because the standard linear convolution operator is prone to feature overload and gradient bias when encountering extremely large numerical impulses in the receptive field. However, the output signal-to-noise ratio curve of the method of this invention always maintains the highest level, and the decline trend is the most gradual and stable. When the background noise feature index α = 1.2, the output signal-to-noise ratio is still maintained at about 18.8 dB, which is significantly better than all other comparison algorithms. This fully verifies the strong isolation capability of the underlying differentiable dynamic amplitude gating (DDAG) mechanism of this invention against non-Gaussian extreme impulse noise, as well as the robustness of feature extraction of multi-scale one-dimensional hole depth separable convolution in strong interference environment. This proves that the method of this invention has excellent anti-interference performance and stability in strong impulse non-Gaussian underwater acoustic noise environment.
[0104] To demonstrate the feasibility of engineering deployment of this invention in underwater equipment with limited computing resources, Figure 9 A comprehensive cross-comparison of the computational complexity (parameters and floating-point operations per second) and single-frame inference latency of five types of heterogeneous algorithms was conducted to quantitatively evaluate the engineering adaptability of each algorithm; from Figure 9As can be seen, the baseline model, the one-dimensional convolutional neural network, has a parameter count of approximately 0.8M and a computational load of 1.5GFLOPs, causing its single inference latency to climb to around 15ms, which is insufficient to meet the microsecond-level real-time requirements of low-power underwater nodes. Traditional filtering algorithms (Wiener filtering, adaptive filtering, wavelet denoising) have extremely low parameter counts and floating-point computational loads, with inference latency in the low latency range of 2-4ms. However, due to their inherent principles, these algorithms have significant bottlenecks in noise suppression performance under non-Gaussian impulse environments, failing to effectively suppress extreme impulse interference with stable α distributions, making them unsuitable for the harsh interference scenarios of remote underwater acoustic systems. In contrast, the present invention... This method, through its rigorous mathematical decoupling design of one-dimensional depth-separable convolution with spatial (temporal) and channel operations, precisely controls the number of model parameters to approximately 0.05M and reduces floating-point operations to approximately 0.1GFLOPs, a reduction of more than an order of magnitude in both parameters and computational cost compared to the baseline 1D-CNN model. In terms of latency, the inference time of this invention is drastically reduced from 15ms for the baseline one-dimensional convolutional neural network to approximately 2.5ms, achieving a latency reduction of nearly 84%. While maintaining an extremely fast inference speed comparable to traditional filtering algorithms, it achieves significantly better noise and interference suppression gains than traditional algorithms, solving the core pain point of traditional algorithms being low latency but low performance. Single-frame inference latency testing was conducted on an underwater embedded low-computing-power platform using a lightweight model adapted to underwater conditions. The above comparison results fully verify that the present invention achieves strict mathematical decoupling of time and channel operations through one-dimensional depth-separable convolution with one-dimensional holes. While ensuring the ability to extract deep features, it significantly reduces the computational complexity and inference latency of the model. The present invention achieves a balance between low computing power, low system latency, and high anti-interference performance. It improves waveform reconstruction accuracy by more than 20% at extremely low signal-to-noise ratios (-10dB), which can fully meet the real-time processing requirements of underwater low-power platforms and has excellent engineering deployment feasibility and practical application value.
[0105] 5. Remote underwater acoustic sea trial verification
[0106] To verify the practical feasibility of this invention in a remote underwater acoustic communication system, a remote underwater acoustic signal transmission and reception experiment was conducted in the Mediterranean Sea based on numerical simulation (based on an underwater acoustic channel model). The physical configuration and acoustic parameters of the experimental system are shown in Table 2.
[0107] Table 2 Sea Trial Experiment Parameter Settings
[0108]
[0109] The signal bandwidth B = 500Hz, consistent with the simulation experiment. Combined with a center frequency of 1250Hz, this effectively balances the signal's anti-attenuation and anti-interference capabilities, avoiding the problems of insufficient information transmission due to excessively narrow bandwidth and excessively rapid signal attenuation due to excessively wide bandwidth. This ensures effective HFM signal transmission and the accuracy of subsequent network feature extraction. The center frequency of 1250Hz is the optimal frequency band for long-distance underwater acoustic communication, balancing signal propagation distance and anti-attenuation capabilities, and avoiding the drawbacks of low transmission rate at low frequencies and excessively rapid attenuation at high frequencies. The sampling frequency of 9600Hz is much higher than the highest frequency of the signal. The Nyquist sampling theorem is satisfied to ensure that the time-domain waveform and frequency characteristics of the HFM signal are not lost or distorted; the sweep time is 1s and the pulse interval is 0.5s, which is adapted to the frequency modulation characteristics of the HFM signal, ensuring that the signal carries sufficient information while preventing interference caused by pulse superposition; the sound source depth is 20m and the receiver depth is 100m, which is set according to the water depth environment of the Mediterranean Sea, which can reduce the impact of sea surface reflection and seabed scattering on signal transmission, optimize multipath channel conditions, and ensure the signal reception quality of 100km long-distance transmission; the positions of the transmitter and receiver and the transmission path in the experiment are as follows. Figure 10 As shown; Figure 10 The transmitting vessel on the left has a transducer deployed underwater as the transmitting node, while the receiving vessel on the right has a deep-water hydrophone deployed underwater as the receiving node. The straight-line transmission span between the transmitting and receiving nodes is 100 km. The solid arrows in the figure indicate the direct transmission path of the sound wave, while the dashed arrows indicate the multipath transmission path formed by the sound wave's reflection from the sea surface and scattering from the seabed. This schematic diagram intuitively reflects the complex multipath effect that causes severe signal delay spread in real ocean waveguides. The above numerical simulation uses an underwater acoustic channel model to model the simulated channel. However, when facing long-range underwater acoustic signal systems, the results of sea trials are of great significance for verifying the feasibility of the proposed method.
[0110] The performance of the pre-trained network model was verified based on 100km of real sea trial data; Figure 11Under sea trial conditions with an input signal-to-noise ratio (SNR) ranging from -10 dB to 10 dB, deep learning methods (one-dimensional convolutional neural networks and the method of this invention) still demonstrate advantages over traditional filtering algorithms (Wiener filtering, adaptive filtering, and wavelet denoising). In real non-stationary ocean backgrounds and complex multipath interference, the performance of traditional filtering algorithms deteriorates sharply in the low SNR range, especially under extremely low input conditions of -10 dB, where the output SNR of wavelet denoising remains negative, making it difficult to effectively suppress non-Gaussian background impulse noise and multipath interference in real ocean environments. In contrast, the method of this invention consistently maintains the highest level of SNR gain, slightly outperforming the baseline one-dimensional convolutional neural network. Particularly under extremely low input conditions of -10 dB, this invention can still improve the output SNR to above 0 dB, fully verifying the strong suppression capability of the dynamic gating (DDAG) mechanism at the bottom layer of the network for real ocean non-Gaussian background impulse noise. This proves that the method of this invention has excellent noise and interference suppression performance and engineering practicality in real long-range underwater acoustic systems.
[0111] Figure 12 The full waveform comparison of a 2.0-second physical frame including the guard interval before and after actual transmission over 100km is shown. Figure 12 (a) in the figure is the time-domain waveform of the transmitted reference signal, showing the pure core HFM pulse signal centered within the 0.5-second silence protection interval before and after; Figure 12 (b) in the figure is the time-frequency domain diagram of the transmitted reference signal, which shows the clear and smooth hyperbolic frequency modulation characteristics of the ideal HFM signal; Figure 12 (c) in the figure is the time-domain waveform of the received noisy signal. Actual sea trial observation data shows that the target HFM signal has been completely submerged by severe geometric spread attenuation and extremely strong sharp impact pulses. The red dashed box in the figure clearly shows the severe "multipath tail" distortion caused by the complex underwater acoustic channel. Figure 12 (d) in the figure is the time-frequency domain diagram of the received noisy signal, which can intuitively reflect the strong background noise energy and the time-frequency energy dispersion and blurring caused by multipath effect; Figure 12 (e) is the time-domain waveform of the output signal after processing by the network of this invention. After pure time-domain end-to-end mapping by the network, it not only accurately reconstructs the macroscopic envelope and timing phase of the core HFM pulse, but also completely eliminates the extremely bad spike interference and multipath tail. Furthermore, it maintains a flat baseline within the 0.5-second silent protection interval before and after, without introducing any parasitic artifacts. Figure 12 In the diagram, (f) is the time-frequency domain plot of the network output signal. (Compare) Figure 12 As shown in Figure (d), background frequency band noise is significantly suppressed, and the pure time-frequency characteristics of the HFM signal are recovered with high fidelity; this verifies the high fidelity and physical reliability of the architecture proposed in this invention in complex underwater acoustic engineering applications.
[0112] 6. Summary of Technological Innovations
[0113] This invention addresses the challenges faced by long-range underwater acoustic systems, including ocean background noise, multipath interference, and limited computational resources. It proposes an end-to-end lightweight noise and interference suppression neural network. This method avoids complex frequency domain transformation processes, instead directly performing end-to-end mapping of the clean waveform of the one-dimensional noisy time-series signal in the time dimension. Its core incorporates a differentiable dynamic amplitude gating (DDAG) mechanism in conjunction with multi-scale one-dimensional dilated depth-separable convolution. To significantly reduce the computational burden while maintaining feature extraction depth, this invention constructs a network architecture with one-dimensional dilated depth-separable convolution as the core operator. This operator mathematically decouples the standard convolution operation into a depthwise convolution that handles local temporal correlations and a pointwise convolution that handles channel-dimensional information fusion. Both theory and practice demonstrate that this structure… While maintaining strong signal representation capabilities, the number of model parameters and floating-point operations can be reduced exponentially. Simulation and 100km sea trial experiments confirm that the network performs excellently in harsh marine environments. Even with extremely low input signal-to-noise ratio (-10dB), the sea trial data can still achieve a signal-to-noise ratio gain of over 10dB, and the output signal-to-noise ratio is above 18dB in strong impulse noise environments. It has high waveform reconstruction accuracy and anti-interference robustness. At the same time, its parameter count is only 0.05M, and the single-frame inference latency is as low as 2.5ms. With low computing power overhead, it meets the actual needs of real-time processing of underwater platforms. Ultimately, it provides a low-latency, high-precision, and robust noise and interference suppression technology for remote underwater acoustic systems under limited computing power conditions, so as to complete robust and reliable underwater signal transmission.
[0114] The above embodiments are merely preferred embodiments of the present invention and should not be considered as limiting the scope of the present invention. All equivalent variations and improvements made within the scope of the present invention should still fall within the patent coverage of the present invention.
Claims
1. A lightweight end-to-end neural network-based time-domain noise and interference suppression method for remote underwater acoustic systems, characterized in that, Includes the following steps: S1. Acquire noisy one-dimensional underwater acoustic time-series observation signals from the receiver of the remote underwater acoustic system. The noisy one-dimensional underwater acoustic time-series observation signals are obtained by linear convolution superposition of pure underwater acoustic signals after multipath transmission distortion through the underwater acoustic channel, and follow zero-mean symmetry. It consists of a stable distribution of non-Gaussian background impulse noise; S2. Construct an end-to-end lightweight noise and interference suppression neural network. The network adopts a pure time-domain architecture to complete the end-to-end mapping of one-dimensional underwater acoustic time-series signals in the time domain. The network includes a preprocessing module, a core noise reduction module, and a post-processing module connected in sequence. The core noise reduction module is composed of multiple robustly hollow separable residual blocks cascaded together. Each robustly hollow separable residual block integrates a differentiable dynamic amplitude gating module and a one-dimensional hole depth separable convolution operator, and is configured with local residual connections. The hole rate of each level of robustly hollow separable residual block is set to increase exponentially. The differentiable dynamic amplitude gating module employs a continuously differentiable nonlinear mapping mechanism to independently and adaptively learn threshold parameters for different feature channels. The working mechanism of the differentiable dynamic amplitude gating module is implemented through the following mathematical model: in, It is a one-dimensional time series feature matrix. This is the robust time series feature matrix. It is the hyperbolic tangent function. For adaptive dynamic learning, the positive smoothing threshold parameter is used. S3. The noisy one-dimensional underwater acoustic time-series observation signal is input into the end-to-end lightweight noise and interference suppression neural network. After the preprocessing module completes channel dimensionality upscaling, normalization, and nonlinear activation, it is sent to the core noise reduction module. The non-Gaussian background impulse noise is adaptively smoothed and clamped and suppressed by the differentiable dynamic amplitude gating module to learn the smoothing threshold. Then, multi-scale time-domain features are extracted by the one-dimensional hole depth separable convolution operator, and the time-domain detail features of the underwater acoustic signal are fused by local residual connection. The network learns the noise and distortion composite interference components through global residual mapping and outputs the pure underwater acoustic signal estimate. S4. Construct a one-dimensional time-series mean square error loss function based on the pure underwater acoustic signal estimate and the corresponding pure reference signal. Iteratively train and optimize the network parameters until the network converges, and obtain the trained end-to-end lightweight noise and interference suppression neural network. S5. Input the noisy one-dimensional underwater acoustic time-series signal to be processed at the receiver of the remote underwater acoustic system into the trained network, and directly output the reconstructed pure underwater acoustic signal in the time domain, simultaneously completing underwater acoustic signal multipath distortion compensation and non-Gaussian background impulse noise suppression.
2. The end-to-end lightweight neural network time-domain noise and interference suppression method for remote underwater acoustic systems according to claim 1, characterized in that, In step S1, the pure underwater acoustic signal is a hyperbolic frequency modulated (HFM) signal, whose time-domain characteristics are characterized by radiation amplitude, initial phase, initial frequency, termination frequency, and duration. The specific time-domain operation relationships are as follows: in, Represents pure water acoustic signals. The radiation amplitude of the HFM signal. It is a cosine function. The initial phase of the HFM signal is given, and 2π is a mathematical constant. The initial frequency of the HFM signal. This is the termination frequency of the HFM signal. The duration of the HFM signal. It is the natural logarithm function. For the instantaneous time variable of the HFM signal; The non-Gaussian background impulse noise follows a zero-mean symmetric pattern. A stable distribution, characterized by background noise characteristic exponents and a scale parameter, exhibits statistical distribution properties; its characteristic function is: ,in This is a background noise characteristic index. For scale parameters, Represented by natural constant An exponential function with base 0. This is the instantaneous time variable of the HFM signal.
3. The end-to-end lightweight neural network time-domain noise and interference suppression method for remote underwater acoustic systems according to claim 1, characterized in that, In step S2, the preprocessing module includes one-dimensional up-dimensional convolution, layer normalization, and activation layers; the postprocessing module uses one-dimensional down-dimensional convolution to compress the network feature channels into a single channel, matching the input and output dimensions of the one-dimensional underwater acoustic time-series signal.
4. The end-to-end lightweight neural network time-domain noise and interference suppression method for remote underwater acoustic systems according to claim 1, characterized in that, In step S2, the void ratio of the robust void separable residual blocks at each level increases exponentially by a fixed multiple of 2, gradually expanding the network's temporal receptive field and adapting to the extraction of multipath delay features in the underwater acoustic channel.
5. The end-to-end lightweight neural network time-domain noise and interference suppression method for remote underwater acoustic systems according to claim 1, characterized in that, In step S2, the differentiable dynamic amplitude gating module, the one-dimensional hole depth separable convolution operator, and the local residual connection form a collaborative noise reduction structure. The differentiable dynamic amplitude gating module first performs adaptive smoothing clamping on non-Gaussian background impulse noise to eliminate the interference of extreme impulses on feature extraction. Then, the one-dimensional hole depth separable convolution operator expands the temporal receptive field with low parameter quantity to extract multi-scale temporal features. Finally, the features after gating and convolution are fused with the original input features through the local residual connection to preserve the temporal details of the underwater acoustic signal, thus achieving the simultaneous completion of non-Gaussian background impulse noise suppression and high-precision reconstruction of temporal waveform.
6. The end-to-end lightweight neural network time-domain noise and interference suppression method for remote underwater acoustic systems according to claim 1, characterized in that, In step S2, the one-dimensional hole depth separable convolution operator decouples the standard one-dimensional temporal convolution into a one-dimensional hole depth convolution in the time dimension and a one-dimensional pointwise convolution in the channel dimension. The two independently realize single-channel temporal feature extraction and cross-channel feature aggregation, respectively. The discrete operation expression for the one-dimensional dilated depth convolution is: in, For sampling point index; For the first The depthwise convolution output of each channel; ∑ is the summation symbol; The kernel size; The weight parameters are those of the one-dimensional hole depth convolution kernel. For input feature number The time-series index on each channel is The characteristic values; Void ratio; The operational expression for the one-dimensional pointwise convolution is: in, For the first Each channel in the time-series index The pointwise convolution output is at the specified position; ∑ is the summation symbol. Input the number of channels; These are the weight parameters for a one-dimensional pointwise convolution; For the first Each channel in the time-series index The depthwise convolution output at that point.
7. The end-to-end lightweight neural network time-domain noise and interference suppression method for remote underwater acoustic systems according to claim 1, characterized in that, In step S3, the estimated value of the pure underwater acoustic signal is obtained by mapping the residual between the noisy one-dimensional observation signal and the noise and interference composite components extracted by the network. The noise and interference composite components are adaptively extracted by the end-to-end lightweight noise and interference suppression neural network through feature learning. The computational expression for the estimated value of the pure underwater acoustic signal is: in, This is an estimated value for the pure underwater acoustic signal; It is a noisy one-dimensional observation signal; For network parameters Dominant pure time-domain end-to-end mapping operator; The estimated composite components of noise and interference; These are the parameters after network optimization.
8. The end-to-end lightweight neural network time-domain noise and interference suppression method for remote underwater acoustic systems according to claim 1, characterized in that, In step S4, the one-dimensional time-series mean square error is used as the loss function for network training. Loss constraints are constructed based on the error between the true and estimated values of the signal at discrete time-series sampling points, thus completing supervised optimization of the network parameters. The expression for the one-dimensional time-series mean square error is: in, The mean square error of the time series is 1; ∑ represents the summation symbol. This represents the total number of signal sampling points. For sampling point index; For pure water acoustic signals The actual value of each sampling point; The estimated value of the pure underwater acoustic signal is the first The values of each sampling point are used; combined with gradient clipping mechanism and early stopping strategy, the network parameters are iteratively trained and optimized until the network converges.
Citation Information
Patent Citations
Methods and systems for spectral analysis of sonar data
CA2968209A1
Sound source separation algorithm of end-to-end time domain multi-scale convolutional neural network
CN113314140A