A high signal-to-noise ratio distributed optical fiber underwater target acoustic print reconstruction method and system

CN122821987APending Publication Date: 2026-09-25STATE GRID ZHEJIANG ELECTRIC POWER CO LTD ZHOUSHAN POWER SUPPLY CO
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611303289.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-26
Publication Date
2026-09-25

AI Technical Summary

Benefits of technology

第一,显著提高信噪比。本技术方案通过引入相位跳变掩码引导的注意力机制,使神经网络能够精准聚焦于信号发生2π跳变的候选位置进行重点修复,同时抑制平稳区域噪声。结合物理约束损失函数,确保解卷绕过程不仅符合数据规律,更严格遵循声波传播的物理极限,从源头避免了误差累积,最终输出高保真、高信噪比的解卷绕相位序列。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821987A_ABST
    Figure CN122821987A_ABST
Patent Text Reader

Abstract

The application discloses a high signal-to-noise ratio distributed optical fiber underwater target acoustic print reconstruction method and relates to the technical fields of optical fiber sensing and underwater acoustic signal processing. The application aims at solving the problems of large error and poor physical consistency of the existing phase unwrapping under low signal-to-noise ratio. The application comprises the following steps: obtaining an initial phase sequence by preprocessing the original wrapped phase; generating a phase jump mask through time-frequency analysis; inputting a deep learning network to guide the attention to focus on the jump position, extract the weighted deep features and time sequence hidden features; in the training process, a loss function containing a physical constraint term is used to limit the adjacent phase difference within the maximum allowed value determined by the sound velocity and the sampling interval; and the unwrapped phase is generated after the training to complete the acoustic print reconstruction. The application deeply integrates the jump physical priori and the sound velocity physical constraint into the deep learning framework, and significantly improves the noise resistance robustness and physical authenticity of the acoustic print reconstruction under low signal-to-noise ratio.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of fiber optic sensing and underwater acoustic signal processing technology, and in particular to a high signal-to-noise ratio distributed fiber optic method for underwater target acoustic pattern reconstruction. Background Technology

[0002] Distributed fiber optic acoustic sensing (DAS), with its long-range continuous sensing, high sensitivity, and excellent electromagnetic interference resistance, has become an important technology in the field of marine monitoring. This system can virtualize existing submarine fiber optic cable infrastructure into a high-density acoustic sensor array, thereby providing high-dimensional data support for submarine earthquake observation, pipeline safety early warning, and ocean dynamics research.

[0003] Applying Direct Aspect-Oriented Satellite (DAS) to ship identification, by detecting the radiated noise generated by the vibration of the ship's propeller and main engine, to track and identify navigational targets, is a promising direction for the development of this technology. However, the complex acoustic environment of the seabed poses a significant challenge to DAS signal processing. Specifically, the physical quantity directly acquired by the DAS system is the phase difference of backscattered light. Due to the periodicity of trigonometric functions, this phase value is limited to... The unwinding interval is referred to as the "wound phase." In practical applications, the complex underwater acoustic environment presents three core challenges to unwinding phase: First, the low-speed guided waves (such as Scholte waves, with a speed of approximately 200-1200 m / s) on the seabed overlap with the frequency band of ship-radiated noise, and their energy is often stronger than the target signal, causing the phase signal to be submerged and the signal-to-noise ratio to deteriorate. Second, traditional phase unwinding algorithms (such as the least squares method and the Goldstein branching method) rely on the assumption of phase continuity. When noise causes outliers or discontinuities in the local phase, the unwinding error will propagate along the time axis, resulting in unwinding errors. Third, although deep learning-based phase reconstruction methods have alleviated noise interference to some extent in recent years, most of them treat the data as an abstract tensor, ignoring the prior physical continuity constraint of phase changes between adjacent sampling points on the speed of sound in the medium. This leads to the model easily generating "pseudo-phase" signals that conform to the data distribution but violate physical laws, failing to meet the requirements of high-precision underwater acoustic signature recognition.

[0004] While DAS technology has been established as an important tool in the field of seabed monitoring, its practical application is hampered by multiple bottlenecks, including low signal-to-noise ratio, difficulty in phase decoupling, lack of physical mechanisms in traditional methods, and the neglect of physical priors by deep learning. How to integrate physical constraints with deep learning to overcome noise interference and decoupling errors has become a core problem that urgently needs to be solved. Summary of the Invention

[0005] The purpose of this invention is to provide a high signal-to-noise ratio distributed optical fiber underwater target acoustic pattern reconstruction method to solve the problems of large noise interference and poor physical consistency in the phase dewinding of existing technologies.

[0006] In a first aspect, a high signal-to-noise ratio distributed optical fiber acoustic signature reconstruction method for underwater targets is provided, which includes the following steps: The raw wound phase data collected by the distributed fiber optic acoustic wave sensing system is acquired and preprocessed to generate an initial phase sequence. Perform time-frequency analysis on the initial phase sequence to detect phase transition positions and generate phase transition masks; The initial phase sequence is input into a pre-constructed deep learning network, which includes a cascaded feature extraction module and a temporal modeling module. The feature extraction module embeds an attention mechanism, using the phase transition mask as a guiding factor, to assign higher feature attention weights to transition positions than to stable positions, and outputs weighted deep features. The temporal modeling module extracts the temporal dependencies of the weighted deep features to obtain temporal hidden state features. Based on the temporal hidden state features, a predicted phase sequence is generated, and the deep learning network is trained using a total loss function that includes physical constraints to optimize the network parameters; wherein, the physical constraints constrain the amplitude of the phase difference between adjacent time steps so that it does not exceed the maximum physically permissible phase difference determined by the medium sound speed propagation limit and the sampling interval. The trained deep learning network generates an unwound phase sequence based on the temporal hidden state features, thereby achieving high signal-to-noise ratio voiceprint reconstruction.

[0007] This technical solution generates a phase transition mask through time-frequency analysis and uses it as a guiding factor for the attention mechanism. This allows the network to accurately locate and focus on the 2π transition position for key repair, effectively avoiding error propagation caused by noise and transition confusion in traditional methods at low signal-to-noise ratios, and significantly improving unwinding accuracy. A physical constraint term based on the speed of sound propagation limit is introduced into the training loss function, forcing the amplitude of the phase difference between adjacent time steps to not exceed the maximum allowable value determined by the medium's sound speed and sampling interval. This fundamentally eliminates pseudo-phases that are "mathematically correct but physically impossible," ensuring the physical consistency of the reconstructed voiceprint. A cascaded feature extraction module and a temporal modeling module work together to extract both deep local phase features and capture long-range temporal dependencies, taking into account both local signal details and global context, providing a high-quality feature foundation for high-fidelity voiceprint reconstruction. Deeply integrating prior knowledge of phase transitions from signal processing with the deep learning model eliminates the need for the network to blindly search for transition patterns from noise, achieving a complementary advantage between physical model guidance and data-driven learning, improving the overall performance and interpretability of the model.

[0008] As a preferred technical means: the preprocessing includes drift removal processing and bandpass filtering; wherein, the drift removal processing is used to remove low-frequency baseline drift caused by laser frequency drift or slow changes in ambient temperature, and the bandpass filtering has a frequency range of 20 Hz to 200 Hz.

[0009] By employing de-drift processing, interference from non-acoustic factors causing gradual phase changes on subsequent signal processing is avoided, ensuring that the initial phase sequence accurately reflects the changes in the underwater sound field. Bandpass filtering from 20 Hz to 200 Hz preserves the main energy frequency bands of ship radiated noise while suppressing low-frequency wave noise and high-frequency electronic noise, achieving focused extraction of the target signal frequency band.

[0010] As a preferred technical means, the time-frequency analysis adopts continuous wavelet transform, selects Morlet wavelet as mother wavelet, and identifies the phase jump position based on the distribution of modulus maxima of transform coefficients at different scales.

[0011] Continuous wavelet transform possesses variable time-frequency resolution, enabling the simultaneous capture of local details and global trends of signals at different scales. Phase transitions manifest as instantaneous high-frequency energy bursts in the time-frequency domain. By detecting the modulus maxima of wavelet coefficients, the accurate location of the transition can be determined comprehensively across multiple scales, avoiding misjudgments or missed detections caused by noise at a single scale. The waveform of the Morlet wavelet is highly similar to the local oscillation characteristics caused by phase transitions; selecting it as the mother wavelet yields a stronger response to transition signals, improving the detection rate of weak transitions. Discrimination based on the multi-scale modulus maxima distribution can effectively distinguish random noise from genuine transition events, enhancing the reliability of transition detection in low signal-to-noise ratio environments.

[0012] As a preferred technical means: the generation of the phase transition mask specifically involves: processing the modulus maxima distribution using a threshold method to identify candidate positions where a 2π phase transition occurs, and generating a binary phase transition mask of the same length as the initial phase sequence; wherein, the position with a value of the first value corresponds to a phase transition candidate point, and the position with a value of the second value corresponds to a stationary point.

[0013] The multi-scale modulus maxima distribution is directly converted into a binary mask using a thresholding method. This approach offers low computational complexity and fast processing speed, making it suitable for real-time processing of large-scale time-series data. The generated mask is of the same length as the initial phase sequence, and the first and second values ​​clearly distinguish between transition candidate points and stationary points, providing a clear spatial guidance signal for the subsequent attention module and reducing the learning burden on the network to autonomously identify transition locations. The binary mask quantifies prior physical knowledge into structured guidance information, enabling the attention mechanism to directly focus on transition candidate regions for feature enhancement, achieving a seamless integration of signal processing and deep learning.

[0014] As a preferred technical means: the cascaded feature extraction module includes at least two cascaded residual blocks, each residual block including a convolutional layer, a batch normalization layer, an activation function layer, and residual connections that directly superimpose the input to the output; the attention mechanism is a channel attention mechanism or a spatial attention mechanism, which uses the phase transition mask as a weight guiding factor to assign a higher weight to the feature map channel or spatial region corresponding to the transition candidate point than to the stationary point, so as to output the weighted deep features.

[0015] By constructing deep networks using at least two cascaded residual blocks, complex phase features, ranging from local details to global structure, can be extracted layer by layer. Residual connections allow gradients to propagate directly across layers, effectively mitigating the vanishing gradient problem during deep network training and enabling the model to maintain stable convergence while deepening the network. The attention mechanism directly uses phase transition masks as weight guides, allowing the network to selectively enhance the expression of transition-related features while suppressing irrelevant features in stable regions. This concentrates processing efforts on critical transition regions that determine the success or failure of unwinding, improving the specificity of feature extraction and overall computational efficiency.

[0016] As a preferred technical means, the temporal modeling module is a bidirectional long short-term memory network, used to extract temporal hidden state features containing forward and backward contextual information.

[0017] Bidirectional Long Short-Term Memory (LSTM) networks capture temporal dependencies between the past and future simultaneously through hidden layers in both forward and backward directions. The identification of phase transition points relies not only on the preceding signal trend but also on subsequent signal changes, thus more accurately confirming the authenticity and type of the transition and reducing the lag or misjudgment that may occur with unidirectional models. The gating mechanism of LSM networks effectively captures dependencies over long time spans, associating long-range stable phase information with local transition regions, ensuring the continuity and consistency of the recovered unwound phase sequence globally. The bidirectional structure allows the network to comprehensively consider the global context of the entire sequence when processing each time step; this helps the network accurately repair transitions while maintaining phase smoothness in stable regions, avoiding the introduction of new discontinuities due to over-focusing on local transitions, thereby outputting more natural and coherent temporal hidden state features.

[0018] As a preferred technical means: the total loss function is composed of the reconstruction loss and the physical constraint term weighted sum; the physical constraint term is used to penalize the amplitude of adjacent phase differences in the predicted phase sequence that exceed the maximum physical allowable phase difference.

[0019] The overall loss function enables the network to simultaneously pursue two optimization objectives during training: data fitting accuracy and physical compliance. The weighting mechanism allows for flexible adjustment between these two objectives. The physical constraint term specifically penalizes the amplitude of adjacent phase differences exceeding the maximum physically permissible phase difference, rather than imposing indiscriminate constraints on all phase changes. This allows the network to preserve high-frequency phase changes within the physically permissible range while suppressing spurious transitions exceeding the speed of sound limit. The physical constraint term provides explicit physical boundaries for the model output, significantly enhancing the reliability and interpretability of the unwinding results.

[0020] As a preferred technical means: the generation of the unwound phase sequence is performed by the decoder network in the deep learning network. The decoder network adaptively adjusts the reconstruction step size according to the value of the phase transition mask: when the mask value indicates a transition candidate point, a large step size correction strategy is adopted to compensate for the 2π transition; when the mask value indicates a stable point, a small step size smoothing strategy is adopted to maintain phase continuity.

[0021] The decoder employs different strategies for transition candidate points and stationary points based on the mask value, enabling targeted processing of these two types of regions. At transition points, it effectively compensates for transitions that are multiples of 2π, while at stationary points, it avoids introducing unnecessary correction disturbances, thereby improving the accuracy of dewinding and the natural continuity of the phase. A large-step correction strategy is used for transition candidate points, which can quickly overcome 2π phase discontinuities, significantly improving correction efficiency. A small-step smoothing strategy is used for stationary points, which can preserve the detailed information in the original signal to the greatest extent possible, avoiding over-smoothing or loss of detail caused by global large-step correction.

[0022] The second aspect: Provides a high signal-to-noise ratio distributed fiber optic underwater target acoustic signature reconstruction system, comprising: The preprocessing module is used to acquire the raw wound phase data collected by the distributed fiber optic acoustic wave sensing system, and to preprocess it to generate the initial phase sequence of the wound phase. The mask generation module is used to perform time-frequency analysis on the initial phase sequence, detect the phase jump positions caused by phase entanglement, and generate a phase jump mask based on the detection results. The deep learning network module includes a cascaded feature extraction unit and a temporal modeling unit. The feature extraction unit embeds an attention mechanism and is configured to introduce the phase transition mask as a guiding factor so that the network pays more attention to the features corresponding to the phase transition positions than to the stable positions, and outputs weighted deep features. The temporal modeling unit is used to extract the temporal dependencies in the weighted deep features to obtain temporal hidden state features. The training module is used to obtain a predicted phase sequence based on the temporal hidden state features during the training process of the deep learning network module, and to optimize the network parameters using a total loss function; wherein, the total loss function includes a physical constraint term, which is used to constrain the amplitude of the phase difference between adjacent time steps in the predicted phase sequence, so that it does not exceed the maximum physically permissible phase difference determined by the medium sound speed propagation limit and the system sampling interval; The voiceprint reconstruction module is used to generate an unwound phase sequence based on the temporal hidden state features through the trained deep learning network module, thereby completing high signal-to-noise ratio voiceprint reconstruction.

[0023] The mask generation module transforms the phase transition positions extracted from time-frequency analysis into structured phase transition masks. The deep learning network module introduces the phase transition masks as guiding factors through an attention mechanism. Simultaneously, the training module embeds the physical constraint of the sound speed propagation limit into the loss function, achieving an organic integration of signal processing prior knowledge and data-driven learning at the system level. The system covers the complete processing chain from raw wound phase data input to high signal-to-noise ratio dewound phase sequence output, automatically completing high-quality reconstruction of underwater target acoustic signatures without manual intervention.

[0024] As a preferred technical means: the preprocessing module is used to perform drift removal processing and bandpass filtering, wherein the frequency range of the bandpass filter is 20 Hz to 200 Hz; The mask generation module uses Morlet wavelet as the mother wavelet in continuous wavelet transform, and generates a binarized phase transition mask of the same length as the initial phase sequence by means of a threshold method based on the distribution of the modulus maxima of the transform coefficients. The feature extraction unit contains at least two cascaded residual blocks, each residual block including a convolutional layer, a batch normalization layer, an activation function layer and a residual connection, and the attention mechanism is a channel attention mechanism or a spatial attention mechanism. The time-series modeling unit is a bidirectional long short-term memory network; The total loss function is composed of a weighted sum of the reconstruction loss and the physical constraint term; The voiceprint reconstruction module includes a decoder network. The decoder network adaptively adjusts the reconstruction step size according to the value of the phase transition mask. When the mask value indicates a transition candidate point, a large step size correction strategy is adopted, and when it indicates a stable point, a small step size smoothing strategy is adopted.

[0025] The modules work together to form a complete optimization chain from data purification, transition localization, feature enhancement, temporal modeling to physical constraint reconstruction. From the mask generation module using Morlet wavelet continuous wavelet transform and generating a binary mask based on the modulus maxima thresholding method, to the feature extraction unit using the binary mask to guide the attention mechanism for differentiated weighting, and then to the voiceprint reconstruction module adaptively selecting a large step size correction or a small step size smoothing strategy based on the mask value, the three stages of transition detection, focusing, and compensation are organically linked, achieving precise processing of phase transitions across the entire chain. The total loss function balances prediction accuracy and physical compliance during the training phase, while the decoder network adaptively adjusts the reconstruction step size based on the transition mask during the inference phase. This ensures that the physical constraint capabilities learned during training are continued in the application phase. Training and inference work together to guarantee the quality of the output voiceprint.

[0026] Beneficial effects: First, it significantly improves the signal-to-noise ratio. This technical solution introduces a phase-jump mask-guided attention mechanism, enabling the neural network to precisely focus on the signal generation point. π Candidate positions with abrupt transitions are carefully repaired while noise in stable regions is suppressed. By combining a physical constraint loss function, the unwinding process is ensured to not only conform to data patterns but also strictly adhere to the physical limits of sound wave propagation, thus avoiding error accumulation at the source and ultimately outputting a high-fidelity, high signal-to-noise ratio unwinding phase sequence.

[0027] Second, the model's adaptability and robustness in complex scenarios are enhanced. This technical solution adopts a collaborative framework of "mask generation + physical constraints + adaptive decoding". The phase-jump mask provides prior knowledge for processing; the combination of cascaded residual networks and bidirectional long short-term memory networks fully exploits the local features and long-range temporal dependencies of the signal; the decoder adaptively adjusts the reconstruction step size according to the mask, realizing differentiated processing of transition points and stationary points, which can flexibly cope with complex underwater signals with different signal-to-noise ratios and different transition modes, and significantly enhances generalization ability and environmental robustness.

[0028] Third, it has strong engineering applicability. The high-quality voiceprint output by this method can be directly input into downstream ship identification systems, which can significantly improve the probability of target detection and reduce the false alarm rate. Attached Figure Description

[0029] Figure 1 This is a schematic diagram of the process of the present invention.

[0030] Figure 2 This is the time-frequency image of the acoustic signature reconstructed from three types of underwater targets using the method of this invention. Detailed Implementation

[0031] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings.

[0032] Example 1 This embodiment provides a high signal-to-noise ratio distributed fiber optic underwater target acoustic signature reconstruction method, which can be implemented using a distributed fiber optic acoustic sensing system. The distributed fiber optic acoustic sensing system utilizes existing submarine optical cables as the sensing medium, sensing underwater acoustic signals by detecting phase changes in backscattered light. This embodiment uses a DAS system with a spatial resolution of 0.2 meters and a sampling rate of 200 Hz as an example for illustration. Figure 1 As shown, this embodiment includes the following steps: S1: Data Preprocessing Acquire the raw wound phase data collected by the DAS system. The raw wound phase data is a one-dimensional phase sequence that varies with time. Due to the periodicity of the trigonometric function, all phase values ​​are wound within the interval [-π, π].

[0033] Preprocessing includes two sub-steps: drift removal and bandpass filtering.

[0034] S1.1: De-drift processing is used to remove low-frequency baseline drift caused by laser frequency drift and slow changes in ambient temperature. Specifically, a polynomial fitting method can be used, that is, a low-order polynomial is used to fit the original phase sequence and the fitting trend term is subtracted to retain the dynamic components in the signal.

[0035] S1.2: The passband frequency of the bandpass filter is set to 20 Hz to 200 Hz. This passband frequency covers the main energy frequency band of radiated noise generated by the vibration of the ship's propeller and main engine, while effectively suppressing low-frequency wave noise below 20 Hz and high-frequency electronic noise above 200 Hz. The bandpass filter can be implemented using a Butterworth filter or a Chebyshev filter.

[0036] After the above preprocessing, an initial phase sequence for winding is generated. .

[0037] S2: Jump Detection and Mask Generation Based on Continuous Wavelet Transform Perform a continuous wavelet transform on the initial phase sequence generated in step S1. Select the Morlet wavelet with a center frequency of 1 Hz as the mother wavelet.

[0038] After performing a multi-scale continuous wavelet transform on the initial phase sequence, time-frequency coefficient matrices at different scales are obtained. Since phase jumps in the time-frequency domain manifest as instantaneous high-frequency energy bursts spanning multiple scales, the jump locations can be determined by detecting the modulus maxima of the wavelet coefficients at each scale. Specifically, local maxima of the wavelet coefficient amplitude are searched along the time axis at each scale, and the time locations where modulus maxima occur simultaneously at multiple scales are marked as candidate jump locations.

[0039] Based on the aforementioned modulus maxima distribution, a binarized phase transition mask of the same length as the initial phase sequence is generated using a thresholding method. The threshold is set to the global maximum value of the wavelet coefficient amplitude across all scales. 0.8 times, that is:

[0040] When the maximum modulus at a certain time point exceeds the threshold δ, the corresponding position of the mask is set to 1, indicating that the point is a candidate point for phase transition; when it does not exceed the threshold, the value is 0, indicating that the point is a stationary point.

[0041] This generates a binary sequence consisting of 1s and 0s. Its length is exactly the same as the initial phase sequence, and each time sampling point corresponds to a clear transition or stationary marker.

[0042] S3: Construction and Forward Propagation of Deep Learning Networks The deep learning network constructed in this embodiment includes a cascaded feature extraction module and a temporal modeling module. The specific structure and function of each module are as follows.

[0043] S3.1 Feature Extraction Module The feature extraction module consists of two cascaded residual blocks, with an attention module embedded between the two residual blocks.

[0044] The first residual block receives the initial phase sequence as input and includes a one-dimensional convolutional layer with a kernel size of 7, a batch normalization layer, and a ReLU activation function layer. It also establishes residual connections that directly connect the input to the output. Through these residual connections, the input signal can bypass the convolutional layers and be directly transmitted to the output, effectively mitigating the gradient vanishing problem in deep networks and ensuring training stability.

[0045] The output of the first residual block enters the attention module. This embodiment employs the SE attention mechanism. The attention module uses the phase transition mask generated in step S2. As a guiding factor for channel weights, specifically: the mask... As a bias term or multiplication factor in the attention weights, the feature map channels corresponding to jump candidate points are assigned higher weights than those of stationary points. For example, for a jump candidate position marked as 1 in the mask, the channel weight of its corresponding feature is multiplied by an enhancement factor greater than 1; for a stationary position marked as 0 in the mask, the original weights are maintained or multiplied by a suppression factor less than 1. In this way, the network is explicitly guided during the learning process to focus on regions with 2π jumps, without having to discover the jump patterns on its own amidst a large amount of noise.

[0046] The attention-weighted features are input into the second residual block. The structure of the second residual block is similar to the first, also including convolutional layers, batch normalization layers, activation function layers, and residual connections, used to further extract deeper semantic features. After cascading processing of two residual blocks and an attention module, the final output is a weighted deep feature.

[0047] S3.2 Timing Modeling Module The weighted deep features output by the feature extraction module are input into the temporal modeling module. In this embodiment, the temporal modeling module uses a bidirectional long short-term memory network.

[0048] The bidirectional Long Short-Term Memory (LSTM) network consists of two parts: a forward LSTM and a backward LSTM. The forward LSTM processes the weighted deep feature sequence in ascending chronological order, while the backward LSTM processes the same sequence in reverse chronological order. The outputs of the two LSTMs are concatenated along the feature dimension. The hidden layer dimension is set to 128, and the number of layers is set to 2.

[0049] By utilizing a bidirectional long short-term memory network, the model can simultaneously capture contextual information from both past and future moments in the phase sequence. This is particularly important for phase decoupling tasks: a correct 2π transition correction decision depends not only on the signal trend before the transition but also on the signal trajectory after the transition to confirm the type and direction of the transition. Ultimately, the temporal modeling module outputs temporal hidden state features that integrate forward and backward contextual information.

[0050] S4: Network Training and Physical Constraints Based on the above network structure, the deep learning network is trained.

[0051] First, based on the temporal hidden state features, a predicted phase sequence is generated through a fully connected layer or decoder network. Then, based on the difference between the predicted phase sequence and the true unwound phase sequence, as well as the physical compliance of the predicted phase sequence itself, the total loss function is calculated and the network parameters are updated via backpropagation.

[0052] The total loss function L is composed of the reconstruction loss. With physical constraints The weighted summation consists of:

[0053] in, The balancing coefficient is used to adjust the weight between the reconstruction loss and the physical constraint term; reconstruction loss Mean square error can be used to measure the point-by-point deviation between the predicted phase sequence and the actual unwound phase.

[0054] The physical constraint term L_phy is used to penalize the amplitude of adjacent phase differences in the predicted phase sequence that exceed the maximum physically permissible phase difference. Its expression is:

[0055] in, , representing the phase difference between adjacent time steps; These are the weighting coefficients (hyperparameters) of the physical constraint term, used to control the penalty intensity of that term; This is the maximum physically permissible phase difference determined based on the medium's sound velocity propagation limit and the system sampling interval. Specifically, if the system sampling interval is... The speed of sound in water is The wavelength of the sensing light wave is ,but It can be calculated as Alternatively, it can be obtained from the calibration of the actual system.

[0056] It should be noted that the physical constraint term only applies when the absolute value of the adjacent phase difference exceeds [a certain threshold]. The penalty is applied only when necessary, while no constraints are imposed on real high-frequency phase changes within the physical limits. This allows the model to retain real fast phase changes in the signal while suppressing spurious transitions that exceed the speed of sound and are physically impossible.

[0057] In each training round, the preprocessed initial phase sequence and the corresponding transition mask are input into the network as a pair. After forward propagation, the predicted phase sequence is obtained. The total loss function is calculated, and then the network parameters are updated via backpropagation. After training, the network has the ability to physically and legally unwind the input phase sequence.

[0058] S5: Adaptive Decoding and Voice Reconstruction Voiceprint reconstruction is performed using a trained deep learning network. The temporal hidden state features output in step S3 are fed into the decoder network, which generates the final unwound phase sequence.

[0059] In this embodiment, the decoder network has the ability to adaptively adjust the reconstruction step size, and the adaptive mechanism relies on the phase transition mask generated in step S2. The values ​​are used to determine and switch: When the mask value is 1 (indicating a transition candidate point), the decoder employs a large step size correction strategy. This strategy applies a correction amount close to an integer multiple of 2π to the predicted phase value to quickly compensate for phase discontinuities caused by winding. Large step size correction is achieved by increasing the response amplitude of the decoder's output layer to the input features, enabling the model to correct the phase value from the transition side to the correct continuous interval with a larger step size.

[0060] When the mask value is 0 (indicating a stable point), the decoder employs a small-step smoothing strategy. This strategy only fine-tunes the predicted phase value, maintaining the natural continuity of the phase by limiting the amplitude of output changes. It avoids introducing unnecessary correction disturbances into the stable region, preserving the detailed features of the original underwater acoustic signal to the greatest extent possible.

[0061] Through the aforementioned adaptive decoding mechanism, the decoder network can perform differentiated processing on different regions, ultimately generating a high signal-to-noise ratio, phase-continuous, and physically realistic unwound phase sequence to complete the acoustic reconstruction of underwater targets.

[0062] Verification effect The method described in this embodiment is used to process the DAS signals of three typical underwater targets (underwater robots, released fish, and divers). Figure 2 It contains three independent time-frequency maps, with the horizontal axis representing time and the vertical axis representing frequency. The color bars on the right represent the signal energy amplitude, with warm-colored areas representing target radiated noise characteristics and cool-colored areas representing seabed environmental noise. The top ROV acoustic signature map: 20s in duration, with a large area of ​​continuous high-energy red bands appearing after 10s, corresponding to the stable radiated noise generated by the continuous vibration of the ROV propeller. The characteristic spectral lines in the low-to-mid frequency range are clear, and the wave base noise is significantly suppressed. The middle fish release acoustic signature map: 20s in duration, with discrete pulsed high-energy characteristics from 0 to 6s, corresponding to the disturbance signal of fish swimming. The overall background is uniform and smooth, without random noise interference. The bottom diver acoustic signature map: 5s in duration, with multiple continuous fine spectral lines distributed across the entire frequency band, depicting the weak vibration signal caused by the diver's limb movements. The weak target characteristics are still completely preserved even at low signal-to-noise ratios.

[0063] The acoustic pattern reconstructed by the method of this invention has significantly reduced background noise, and the unique spectral lines of the three types of targets are clearly distinguishable. Compared with the original noisy coiled phase signal, the overall signal-to-noise ratio is significantly improved, which can provide a high-quality data foundation for subsequent underwater target identification and classification.

[0064] Example 2 This embodiment provides a high signal-to-noise ratio distributed fiber optic underwater target acoustic signature reconstruction system to implement the method described in Embodiment 1. This system can be deployed and run on a server, embedded processing platform, or cloud computing node, receiving raw data collected by the DAS system and outputting a high signal-to-noise ratio unwound phase sequence.

[0065] This system includes a preprocessing module, a mask generation module, a deep learning network module, a training module, and a voiceprint reconstruction module. The following is a detailed description of each module.

[0066] I. Preprocessing Module The preprocessing module connects to the data output interface of the DAS system and receives the raw wound phase data acquired by the DAS system. The preprocessing module internally includes a drift-removal processing unit and a bandpass filtering unit.

[0067] The de-drift processing unit removes low-frequency baseline drift caused by laser frequency drift and slow changes in ambient temperature. During long-term continuous operation of the DAS system, the laser output wavelength drifts slowly with temperature and time, and changes in seabed temperature also cause thermal expansion and contraction of the sensing fiber. These factors result in a low-frequency trend term in the phase data. The de-drift processing unit extracts and subtracts this trend term using a moving average method or polynomial fitting, ensuring that the dynamic components in the output data accurately reflect the underwater acoustic signal rather than the gradual changes in the environment.

[0068] The bandpass filter unit is connected after the drift removal processing unit. Its passband frequency is set to 20 Hz to 200 Hz, covering the frequency range of ship propeller cavitation noise and main engine mechanical vibration noise. At the same time, it suppresses low-frequency environmental noise such as ocean waves and currents below 20 Hz, as well as electronic noise and system background noise above 200 Hz.

[0069] The preprocessing module outputs the initial phase sequence of the winding.

[0070] II. Mask Generation Module The mask generation module receives the initial phase sequence output by the preprocessing module. Internally, the mask generation module includes a continuous wavelet transform unit, a modulus maxima detection unit, and a threshold decision unit.

[0071] The continuous wavelet transform unit selects the Morlet wavelet as the mother wavelet to perform multi-scale continuous wavelet transform on the initial phase sequence. The waveform characteristics of the Morlet wavelet highly match the local high-frequency oscillations caused by phase jumps, and can generate strong coefficient responses at the jump positions. By changing the scale parameter, the continuous wavelet transform unit performs wavelet transform on a set of continuous scales, outputting time-frequency coefficient matrices at different scales.

[0072] The modulus maxima detection unit performs time-point-by-time amplitude detection on the time-frequency coefficients at each scale, searching for local maxima points of the wavelet coefficient amplitude. Since a 2π phase transition manifests as an instantaneous high-frequency energy burst spanning multiple scales in the time-frequency domain, the true transition location will simultaneously exhibit modulus maxima characteristics at multiple scales, while random noise usually only appears at isolated scales.

[0073] The threshold decision unit processes the modulus maxima detection results using a threshold method. A threshold is set. ,in This represents the global maximum value of the wavelet coefficient amplitude across all scales. Time points where the modulus maxima exceed the threshold are identified as candidate locations for a 2π phase transition and marked as 1 in the mask; time points where the modulus maxima do not exceed the threshold are marked as 0. This generates a binary phase transition mask of the same length as the initial phase sequence.

[0074] The phase-jump mask output by the mask generation module is simultaneously sent to the attention unit in the deep learning network module and the decoder network in the speaker reconstruction module as a physical guiding signal throughout the entire processing flow.

[0075] III. Deep Learning Network Module The deep learning network module is the core processing unit of the system, which includes cascaded feature extraction units and temporal modeling units.

[0076] The feature extraction unit contains at least two cascaded residual blocks, with attention units embedded between the residual blocks.

[0077] The first residual block consists of convolutional layers, batch normalization layers, and activation function layers, and includes residual connections that directly stack the input to the output. The convolutional layers use one-dimensional convolutions with a kernel size of 7, responsible for extracting local waveform features from the initial phase sequence of the input. The batch normalization layers normalize the features, stabilizing the training process. The activation function layers use the ReLU function. The residual connections allow gradients to propagate directly across layers during backpropagation, effectively mitigating the vanishing gradient problem in deep networks.

[0078] The attention unit is located between two residual blocks and employs either channel attention or spatial attention mechanisms. A key input to the attention unit is the phase transition mask output by the mask generation module. The attention unit is configured to use the phase transition mask as a weight guide factor. When calculating attention weights, higher attention weights are assigned to the feature map channels or spatial regions corresponding to positions marked as transition candidate points (value 1) in the mask, while lower weights are assigned to regions corresponding to stationary points (value 0). Through this physically-guided approach, the attention unit focuses the network's feature extraction capabilities on the critical transition regions that determine the success or failure of unwinding.

[0079] The second residual block has the same structure as the first residual block. It receives the weighted features output by the attention unit, further extracts deeper semantic information, and finally outputs weighted deep features.

[0080] The temporal modeling unit receives weighted deep features from the feature extraction unit. The temporal modeling unit employs a bidirectional long short-term memory (LSTM) network with a hidden layer dimension of 128 and a layer count of 2. A forward LSTM processes the feature sequence in forward chronological order, while a backward LSTM processes the same sequence in reverse chronological order. The outputs of both at each time step are concatenated along the feature dimension. The bidirectional LSTM network effectively controls long-range memory and forgetting of information through gating mechanisms such as forget gates, input gates, and output gates. This allows it to capture long-range dependencies across time steps in the weighted deep features and extract temporal hidden state features containing forward and backward contextual information.

[0081] IV. Training Module The training module is used to train the deep learning network module offline before it is deployed in practical applications. The training module receives the temporal hidden state features output by the deep learning network module and generates a predicted phase sequence through the decoder network.

[0082] The core of the training module lies in its total loss function. The total loss function is a weighted sum of the reconstruction loss and the physical constraint term. The reconstruction loss measures the deviation between the predicted phase sequence and the true unwound phase label. The physical constraint term is constructed based on the physical constraints of underwater acoustic propagation: the phase difference between adjacent time steps in the predicted phase sequence is calculated, and the maximum physically permissible phase difference determined by the medium's sound speed propagation limit and the system sampling interval is used as the constraint boundary. Penalties are imposed on adjacent phase difference amplitudes exceeding the constraint boundary. In this embodiment, the upper limit of the sound speed in water is taken as 1500 m / s, and combined with a sampling rate of 200 Hz, the theoretical maximum phase change between adjacent sampling points can be determined.

[0083] The training module minimizes the total loss function using the backpropagation algorithm and the Adam optimizer, updating all trainable parameters in the deep learning network module. During iterative training, the network gradually learns how to output an unwound phase that conforms to the physical laws of underwater acoustic propagation while ensuring data fitting accuracy.

[0084] V. Voiceprint Reconstruction Module The voiceprint reconstruction module is the final output module of the system. It generates a high signal-to-noise ratio unwound phase sequence based on the temporal hidden state features through a trained deep learning network module, thus completing the voiceprint reconstruction.

[0085] The voiceprint reconstruction module includes a decoder network with adaptive step size adjustment capability. When generating the unwound phase sequence, the decoder network synchronously receives the phase transition mask output by the mask generation module as a control signal and performs differentiated processing based on the mask value: when the mask value indicates that the current time point is a transition candidate point (value 1), the decoder network adopts a large step size correction strategy, applying a correction amount close to an integer multiple of 2π to the predicted phase to compensate for phase discontinuities caused by phase winding; when the mask value indicates that the current time point is a stationary point (value 0), the decoder network adopts a small step size smoothing strategy, making only minor adjustments to the predicted phase to maintain the natural continuity of the phase and preserve the original details of the signal.

[0086] The unwound phase sequence output by the voiceprint reconstruction module has a high signal-to-noise ratio and physical consistency, and can be directly used for voiceprint mapping or as input features for target recognition algorithms.

[0087] The above modules work together to achieve fully automated processing from raw DAS winding phase data to high-quality unwound phase sequences. The system's output acoustic signature can significantly improve the probability of underwater target detection and reduce the false alarm rate.

[0088] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Those skilled in the art should understand that various modifications and equivalent substitutions can be made to the present invention within the spirit and scope defined by the claims, and these modifications and substitutions also fall within the protection scope of the present invention.

Claims

1. A high signal-to-noise ratio distributed optical fiber underwater target acoustic signature reconstruction method, characterized in that, Includes the following steps: The raw wound phase data collected by the distributed fiber optic acoustic wave sensing system is acquired and preprocessed to generate an initial phase sequence. Perform time-frequency analysis on the initial phase sequence to detect phase transition positions and generate phase transition masks; The initial phase sequence is input into a pre-constructed deep learning network, which includes a cascaded feature extraction module and a temporal modeling module. The feature extraction module embeds an attention mechanism, using the phase transition mask as a guiding factor, to assign higher feature attention weights to transition positions than to stable positions, and outputs weighted deep features. The temporal modeling module extracts the temporal dependencies of the weighted deep features to obtain temporal hidden state features. Based on the temporal hidden state features, a predicted phase sequence is generated, and the deep learning network is trained using a total loss function that includes physical constraints to optimize the network parameters; wherein, the physical constraints constrain the amplitude of the phase difference between adjacent time steps so that it does not exceed the maximum physically permissible phase difference determined by the medium sound speed propagation limit and the sampling interval. The trained deep learning network generates an unwound phase sequence based on the temporal hidden state features, thereby achieving high signal-to-noise ratio voiceprint reconstruction.

2. The high signal-to-noise ratio distributed optical fiber underwater target acoustic signature reconstruction method according to claim 1, characterized in that: The preprocessing includes drift removal and bandpass filtering; wherein the drift removal is used to remove low-frequency baseline drift caused by laser frequency drift or slow changes in ambient temperature, and the bandpass filtering has a frequency band range of 20 Hz to 200 Hz.

3. The high signal-to-noise ratio distributed optical fiber underwater target acoustic signature reconstruction method according to claim 1, characterized in that: The time-frequency analysis employs continuous wavelet transform, selecting the Morlet wavelet as the mother wavelet, and identifying the phase transition position based on the distribution of modulus maxima of the transform coefficients at different scales.

4. The high signal-to-noise ratio distributed optical fiber underwater target acoustic signature reconstruction method according to claim 3, characterized in that: The generation of the phase transition mask specifically involves: processing the modulus maxima distribution using a threshold method to identify candidate positions where a 2π phase transition occurs, and generating a binary phase transition mask of the same length as the initial phase sequence; wherein, positions with a first value correspond to candidate phase transition points, and positions with a second value correspond to stationary points.

5. The high signal-to-noise ratio distributed optical fiber underwater target acoustic signature reconstruction method according to claim 1, characterized in that: The cascaded feature extraction module includes at least two cascaded residual blocks. Each residual block includes a convolutional layer, a batch normalization layer, an activation function layer, and residual connections that directly superimpose the input to the output. The attention mechanism is a channel attention mechanism or a spatial attention mechanism, which uses the phase transition mask as a weight guiding factor to assign a higher weight to the feature map channel or spatial region corresponding to the transition candidate point than to the stationary point, so as to output the weighted deep features.

6. The high signal-to-noise ratio distributed optical fiber underwater target acoustic signature reconstruction method according to claim 1, characterized in that: The temporal modeling module is a bidirectional long short-term memory network used to extract temporal hidden state features containing forward and backward contextual information.

7. The high signal-to-noise ratio distributed optical fiber underwater target acoustic signature reconstruction method according to claim 1, characterized in that: The total loss function is composed of a weighted sum of the reconstruction loss and the physical constraint term; the reconstruction loss is used to measure the deviation between the predicted phase sequence and the true unwound phase, and the physical constraint term is used to penalize the magnitude of adjacent phase differences in the predicted phase sequence that exceed the maximum physical allowable phase difference.

8. The high signal-to-noise ratio distributed optical fiber underwater target acoustic signature reconstruction method according to claim 1, characterized in that: The generation of the unwound phase sequence is performed by the decoder network in the deep learning network. The decoder network adaptively adjusts the reconstruction step size according to the value of the phase transition mask: when the mask value indicates a transition candidate point, a large step size correction strategy is adopted to compensate for the 2π transition; when the mask value indicates a stable point, a small step size smoothing strategy is adopted to maintain phase continuity.

9. A high signal-to-noise ratio distributed optical fiber underwater target acoustic signature reconstruction system, characterized in that, include: The preprocessing module is used to acquire the raw wound phase data collected by the distributed fiber optic acoustic wave sensing system, and to preprocess it to generate the initial phase sequence of the wound phase. The mask generation module is used to perform time-frequency analysis on the initial phase sequence, detect the phase jump positions caused by phase entanglement, and generate a phase jump mask based on the detection results. The deep learning network module includes a cascaded feature extraction unit and a temporal modeling unit. The feature extraction unit embeds an attention mechanism and is configured to introduce the phase transition mask as a guiding factor so that the network pays more attention to the features corresponding to the phase transition positions than to the stable positions, and outputs weighted deep features. The temporal modeling unit is used to extract the temporal dependencies in the weighted deep features to obtain temporal hidden state features. The training module is used to obtain a predicted phase sequence based on the temporal hidden state features during the training process of the deep learning network module, and to optimize the network parameters using a total loss function; wherein, the total loss function includes a physical constraint term, which is used to constrain the amplitude of the phase difference between adjacent time steps in the predicted phase sequence, so that it does not exceed the maximum physically permissible phase difference determined by the medium sound speed propagation limit and the system sampling interval; The voiceprint reconstruction module is used to generate an unwound phase sequence based on the temporal hidden state features through the trained deep learning network module, thereby completing high signal-to-noise ratio voiceprint reconstruction.

10. The high signal-to-noise ratio distributed optical fiber underwater target acoustic signature reconstruction system according to claim 9, characterized in that: The preprocessing module is used to perform drift removal and bandpass filtering, wherein the frequency range of the bandpass filter is 20 Hz to 200 Hz; The mask generation module uses Morlet wavelet as the mother wavelet in continuous wavelet transform, and generates a binarized phase transition mask of the same length as the initial phase sequence by means of a threshold method based on the distribution of the modulus maxima of the transform coefficients. The feature extraction unit contains at least two cascaded residual blocks, each residual block including a convolutional layer, a batch normalization layer, an activation function layer and a residual connection, and the attention mechanism is a channel attention mechanism or a spatial attention mechanism. The time-series modeling unit is a bidirectional long short-term memory network; The total loss function is composed of a weighted sum of the reconstruction loss and the physical constraint term; The voiceprint reconstruction module includes a decoder network. The decoder network adaptively adjusts the reconstruction step size according to the value of the phase transition mask. When the mask value indicates a transition candidate point, a large step size correction strategy is adopted, and when it indicates a stable point, a small step size smoothing strategy is adopted.