A distributed optical fiber acoustic wave sensing system and pattern recognition method

CN122384963BActive Publication Date: 2026-09-29SHENZHEN INST OF GUANGDONG OCEAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610845478.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-12
Publication Date
2026-09-29
Estimated Expiration
2046-06-12

AI Technical Summary

Technical Problem

[0008]本发明的目的是为了解决现有分布式光纤声波传感系统(DAS)在复杂环境及低样本量条件下的模式识别应用瓶颈,提供一种分布式光纤声波传感系统和模式识别方法

Benefits of technology

[0108]本发明的技术方案,在硬件方面采用偏振分集接收架构平抑偏振衰落;在软件方面引入内存级在线增强策略迫使模型学习物理不变量,并利用改进的SE(Squeeze-and-Excitation Network,挤压与激励网络)通道降噪与TimeFormer全局编码器协同作业,在剔除环境噪声的同时保留声波时间轴。整套方案实现了从信号感知到复杂模式分类的全流程优化,显著降低了系统的误报率与部署门槛,为周界安防、管道监测等领域提供了一种兼具高可靠性与高泛化能力的分布式声波监测新方法,具有极高的工程推广价值。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122384963B_ABST
    Figure CN122384963B_ABST
Patent Text Reader

Abstract

The application provides a distributed optical fiber acoustic wave sensing system and a pattern recognition method. The application relates to the fields of optical fiber sensing and pattern recognition, and the system comprises the following steps: continuous coherent light emitted by a narrow-line-width laser is split into two paths by a beam splitting fiber coupler, one path is used as local reference light, and the other path is used as probe light; the probe light is modulated into optical pulses in an acousto-optic modulator driven by a waveform generator, and the optical pulses are injected into a measured sensing optical fiber through a fiber loop; Rayleigh backscattering light returned through the measured sensing optical fiber is mixed with the local reference light through the fiber loop, and the mixed light is decomposed into two orthogonal polarized light beams in a polarization beam splitter and is transmitted to a first balanced photodetector and a second balanced photodetector respectively to be converted into electrical signals; and the two orthogonal interference electrical signals are synchronously collected by a data acquisition card and are sent to a host computer for demodulation. The application provides a new distributed acoustic wave monitoring method with high reliability and high generalization ability, and has extremely high engineering popularization value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of fiber optic sensing and pattern recognition, specifically to a distributed fiber optic acoustic sensing system and a pattern recognition method. Background Technology

[0002] Distributed Acoustic Sensing (DAS) systems primarily rely on detecting the intensity and phase evolution of Rayleigh backscattered light within the optical fiber to achieve continuous dynamic monitoring of acoustic and vibration signals over long distances and large linear areas. Compared to traditional electrical point sensors, DAS technology uses ordinary optical fiber itself as the sensing medium, naturally possessing significant advantages such as passive explosion-proof capability, resistance to electromagnetic interference, long monitoring distance, and strong concealment.

[0003] When faced with massive and complex DAS monitoring data, the core challenge for the technology's engineering application lies in how to quickly and accurately identify specific acoustic event types. Currently, mainstream solutions in the industry tend to incorporate deep learning frameworks, especially two-dimensional image classification algorithms represented by convolutional neural networks (CNNs), which convert one-dimensional time-domain signals or phase data into two-dimensional time-frequency maps for feature extraction and classification. However, in complex and ever-changing real-world application environments, existing conventional deep learning methods reveal the following significant technical limitations:

[0004] First, DAS systems in real-world operating conditions are highly susceptible to severe contamination from complex environmental background noise. Existing conventional CNN models often treat all feature channels equally during feature extraction, lacking an effective channel-level signal filtering mechanism. This results in the network's inability to autonomously suppress invalid frequency bands submerged by noise, making it difficult to meet the reliability requirements of engineering applications in terms of event recognition accuracy under low signal-to-noise ratio conditions.

[0005] Secondly, typical man-made or natural physical acoustic events often exhibit significant long-term durations and dynamic evolution patterns. Traditional convolutional networks, in the process of performing multi-layer spatial pooling to compress feature dimensions, are highly susceptible to irreversible loss of global contextual information along the time axis. This limitation of the local receptive field makes it difficult for the model to accurately grasp the evolutionary "rhythm" and long-term temporal dependencies of complex sound waves, leading to serious cross-class misjudgments.

[0006] Finally, constructing a large-scale and accurately labeled field DAS acoustic database faces extremely high labor costs and implementation difficulties. Under conditions of scarce effective samples, conventional deep neural networks are prone to overfitting, meaning the model tends to memorize specific distortion features of the training set samples rather than learning the acoustic essence. Although some existing solutions use methods such as image translation and scaling to increase the training set size, this approach easily leads to data traversal during cross-validation, resulting in artificially inflated model evaluation metrics. Furthermore, in actual deployments facing complex physical delays and unknown waveform distortions, their generalization ability and robustness remain inadequate.

[0007] In order to overcome the technical bottlenecks of existing deep learning demodulation methods in terms of noise resistance, temporal information preservation, and small sample generalization, there is an urgent need for a distributed fiber optic acoustic wave monitoring method that combines high reliability and high generalization ability. Summary of the Invention

[0008] The purpose of this invention is to address the bottleneck of existing distributed optical fiber acoustic sensing systems (DAS) in pattern recognition applications under complex environments and low sample size conditions, and to provide a distributed optical fiber acoustic sensing system and pattern recognition method.

[0009] A distributed fiber optic acoustic wave sensing system includes: a narrow linewidth laser 1, a beam splitter fiber coupler 2, a waveform generator 3, an acousto-optic modulator 4, a fiber amplifier 5, a fiber optic circulator 6, a sensing fiber under test 7, a beam combiner fiber coupler 8, a polarization beam splitter 9, a first balanced photodetector 10, a second balanced photodetector 11, a data acquisition card 12, and a host computer 13.

[0010] The continuous coherent light emitted by the narrow linewidth laser 1 is split into two paths by the beam splitter fiber coupler 2. One path is used as a local reference light and connected to the beam combiner fiber coupler 8. The other path is used as a probe light and enters the acousto-optic modulator 4 driven by the waveform generator 3 to be modulated into an optical pulse. The power is then amplified by the fiber amplifier 5. The amplified probe light pulse is injected into the sensing fiber 7 under test through the fiber circulator 6.

[0011] Rayleigh backscattered light returning from the sensing fiber 7 passes through the fiber optic circulator 6 and mixes with the local reference light at the fiber optic coupler 8. The mixed optical signal enters the polarization beam splitter 9 and is decomposed into two orthogonal polarized beams, which are then transmitted to the first balanced photodetector 10 and the second balanced photodetector 11 respectively and converted into electrical signals. The two orthogonal interference electrical signals are synchronously acquired by the data acquisition card 12 and sent to the host computer 13 for demodulation.

[0012] Secondly, the present invention also provides a distributed fiber optic acoustic pattern recognition method, comprising:

[0013] S101. Perform digital down-conversion on the acquired interferometric electrical signal to obtain the baseband in-phase component. Orthogonal components of baseband The baseband complex envelope signal is constructed from the baseband in-phase component and the baseband quadrature component. The baseband complex envelope signal is matched and filtered using a digital transmit pulse reference template corresponding to the acousto-optic modulator drive waveform to obtain the pulse-compressed baseband complex Rayleigh scattering signal.

[0014] S102. Extract the spatial differential phase of the pulse-compressed baseband complex Rayleigh scattering signal along the distance axis by complex conjugate multiplication, and determine the target spatial position of the acoustic event in the tested sensing fiber by combining the frequency band energy positioning algorithm. Perform bandpass filtering and downsampling on the slow time phase sequence of the target spatial position to obtain a one-dimensional time-domain acoustic signal.

[0015] S103. The one-dimensional time-domain acoustic signal is mapped to a two-dimensional Mel spectrum through short-time Fourier transform, and the local frequency energy distribution and time evolution trajectory corresponding to the two-dimensional Mel spectrum are extracted.

[0016] S104. Perform dynamic data augmentation on the two-dimensional Mel spectrogram and perform random spectrum masking and soft label data aliasing to obtain a dynamically augmented feature matrix.

[0017] S105. The background noise is suppressed by adaptively calculating the weights, and the dynamic enhancement feature matrix is ​​downsampled by using an asymmetric max pooling operator to obtain a frequency dimension compressed feature matrix.

[0018] S106. Using a self-attention mechanism, capture the long-range dynamic evolution beat of the acoustic events corresponding to the frequency dimension compressed feature matrix across the time axis, and determine the category of the acoustic events of the tested sensing fiber through a fully connected network.

[0019] In one implementation, step S101 includes:

[0020] The acquired interferometric electrical signals are added and fused using the X-polarization and Y-polarization channels to obtain the polarization-fused interferometric beat frequency electrical signal. , represented as:

[0021] ;

[0022] in, ;

[0023] Represents electron charge, Indicates the quantum efficiency of the detector. This represents Planck's constant. Indicates the photon frequency; and These represent the AC beat frequency current signals on the two positive traffic channels X and Y, respectively; and These represent the electric field intensity components of the probe light in the X and Y polarization directions, respectively; and This represents the electric field intensity components of the local reference light in the X and Y polarization directions;

[0024] The signal of the returned probe light is represented as follows:

[0025] ;

[0026] in, Indicates the optical frequency of the incident light. Indicates the frequency shift of the acousto-optic modulator. This represents the intensity of the backscattered Rayleigh light. This indicates the modulation phase caused by external sound waves or vibrations;

[0027] The reference light is represented as:

[0028] ;

[0029] in, Indicates the intensity of the reference light;

[0030] Let the equivalent interference amplitude after fusion be Interference beat frequency electrical signal after polarization fusion Simplified representation:

[0031] ;

[0032] The generated two mutually orthogonal digital reference signals are in sync with the frequency shift of the acousto-optic modulator. and The polarization-fused interferometric beat frequency electrical signal respectively with the digital reference signal and Multiply;

[0033] After multiplication, the in-phase branch signal is obtained. and orthogonal branch signals Represented as:

[0034]

[0035] A digital low-pass filter is used for the in-phase branch signal. and the orthogonal branch signal The process involves filtering out high-frequency harmonic components to obtain the equivalent interference amplitude. With slowly changing phase Baseband in-phase component of information Orthogonal components to baseband :

[0036]

[0037] in, Indicates a direct proportional relationship;

[0038] Baseband complex envelope signal From baseband in-phase components Orthogonal components of baseband The complex signal is calculated using the following formula: ;

[0039] Using a digital transmit pulse reference template corresponding to the acousto-optic modulator drive waveform For the baseband complex envelope signal Matched filtering is performed to obtain the pulse-compressed baseband complex Rayleigh scattering signal. , represented as: ;

[0040] in, This represents the baseband complex Rayleigh scattering signal after pulse compression and spatial resolution enhancement. This represents a digital transmit pulse reference template corresponding to the transmit probe light pulse waveform. This represents the convolution operation. This indicates the complex conjugate operation.

[0041] In one implementation, step S102 includes:

[0042] The baseband Rayleigh scattering signal after pulse compression The detection light pulses are framed according to their emission period to obtain the pulse-compressed Rayleigh scattering complex traces corresponding to different emission periods. Based on the round-trip propagation delay corresponding to each fast time sampling point in the complex traces, the round-trip propagation delay is mapped to the spatial sampling position of the sensing fiber under test along the distance axis, thus obtaining a two-dimensional complex signal characterized by both slow time and spatial position. ;

[0043] Where t represents the slow-time sampling sequence formed by continuously emitting probe light pulses, z represents the spatial sampling position along the length of the sensing fiber under test, and the spatial sampling position z is determined by the round-trip propagation delay corresponding to the fast-time sampling point;

[0044] The two-dimensional complex signal Represented as a complex number consisting of a real part and an imaginary part:

[0045]

[0046] in, This represents the in-phase component of the pulse-compressed baseband complex Rayleigh scattering signal at spatial position z. This represents the orthogonal component of the pulse-compressed baseband complex Rayleigh scattering signal at spatial position z;

[0047] Complex conjugate multiplication of the pulse-compressed baseband complex Rayleigh scattering signals at two spatial sampling positions z and z+L, separated by a fixed gauge length L, yields the differential complex signal C(t,z), expressed as:

[0048]

[0049] in, Represents the complex conjugate operation;

[0050] The argument of the differential complex signal C(t,z) is calculated to obtain the differential phase in the winding space. , represented as:

[0051]

[0052] Where Re[C(t,z)] and Im[C(t,z)] represent the real part and imaginary part of the differential complex signal, respectively, and atan2 is used to determine the corresponding phase angle based on the real part and imaginary part;

[0053] Combined with the phase unwinding algorithm, the differential phase of the winding space is analyzed. Perform 2kπ jump compensation to obtain spatial differential phase. ;

[0054] According to the spatial difference phase The spatial location of the acoustic event is determined by the response distribution along the distance axis of the sensing fiber under test, thus identifying the target spatial location z0 corresponding to the acoustic event.

[0055] Extract the slow-time spatial difference phase sequence at the target spatial location z0. and the slow time-space differential phase sequence Digital bandpass filtering and audio downsampling are performed to obtain a one-dimensional time-domain physical sound wave signal. , represented as:

[0056]

[0057] in,

[0058]

[0059] z0 represents the target spatial location where the acoustic event occurs, and z0+L represents another spatial location along the length of the sensing fiber being measured, which is a fixed gauge length L away from z0, where L is the phase difference gauge length.

[0060] In one implementation, step S103 includes:

[0061] The one-dimensional time-domain physical acoustic signal at the time domain level Perform random waveform enhancement;

[0062] Enhanced one-dimensional time domain signal It can be represented as:

[0063]

[0064] in, A random gain factor that follows a uniform distribution is used to simulate the random fluctuations of sound source energy; The dynamically generated data follows a normal distribution based on a set signal-to-noise ratio. Gaussian white noise;

[0065] For the enhanced one-dimensional time-domain signal Perform feature extraction and map the extracted features into a two-dimensional log-Mel spectrum feature matrix that characterizes the time-frequency distribution of sound energy.

[0066] In one implementation, the enhanced one-dimensional time-domain signal Perform feature extraction, mapping the extracted features to a two-dimensional log-Mel spectrum feature matrix characterizing the time-frequency distribution of sound energy, including:

[0067] The time-domain signal is converted into a linear time-frequency complex matrix by the short-time Fourier transform, and the discrete calculation formula is as follows:

[0068]

[0069] in, For the number of points in the Fast Fourier Transform, The frame shift step size, For smoothing window functions, and These represent the time frame index and the frequency point index, respectively; n is the index of the discrete-time sampling point within each frame's short-time analysis window, with a value range of 0 ≤ n ≤ N. FFT -1;

[0070] The linear time-frequency complex matrix is ​​obtained by... Mel filter bank consisting of triangular filters Calculate the energy of each Mel frequency band to form the Mel energy spectrum. The calculation formula is as follows:

[0071]

[0072] Where k is the index of the linear frequency point after the short-time Fourier transform, and its value range is related to the number of points in the fast Fourier transform; for the complete spectrum, k = 0, 1, 2, ..., N FFT -1; For the one-sided spectrum of a real signal, k=0,1,2,…,N FFT / 2;

[0073] Taking the logarithm of the Mel energy spectrum and using the maximum value in the matrix as the reference point... It is converted to a decibel scale to obtain a two-dimensional log-Mel spectrum feature matrix. The calculation formula is:

[0074]

[0075] in, To prevent tiny extreme values ​​at logarithmic zero poles.

[0076] In one implementation, step S104 includes:

[0077] The two-dimensional logarithmic Mel-frequency characteristic matrix As a feature sample, another feature sample X is randomly selected from the training dataset or the current training batch. j and its corresponding one-hot encoded label Y j , where X i and X j Y represents the two-dimensional logarithmic Mel-frequency characteristic matrix obtained after time-frequency mapping of two one-dimensional time-domain acoustic signals. i and Y j These represent the corresponding acoustic event category labels; the features and labels are linearly fused according to the following formula to obtain a hybrid two-dimensional feature matrix and hybrid soft labels:

[0078]

[0079] Among them, the random weights of soft label aliasing are generated according to the Beta distribution. , To control the hyperparameters of the Beta distribution shape, and These represent the fused hybrid feature matrix and the hybrid soft label, respectively.

[0080] A dynamic spectral masking operator is used to randomly generate time masking blocks and frequency masking blocks of a preset width on the two-dimensional feature matrix, and the energy values ​​within the time masking blocks and frequency masking blocks are forcibly set to zero to obtain a dynamically enhanced feature matrix.

[0081] In one implementation, step S105 includes:

[0082] An improved squeeze-excitation channel denoising process is applied to the dynamically enhanced feature matrix;

[0083] The improved squeeze-excitation channel denoising process includes three cascaded feature extraction layers. In each special zone extraction layer, the input feature matrix is ​​sequentially processed by two-dimensional convolution, batch normalization, and ReLU activation function.

[0084] The denoised feature matrix is ​​subjected to SE channel attention processing, and each feature channel is calculated using global average pooling. Global spatial statistics The calculation formula is:

[0085]

[0086] in, Representation of the characteristic matrix In the Each channel, spatial coordinates are Pixel feature value at that location and These represent the frequency dimension and time dimension of the feature matrix, respectively;

[0087] For the global spatial statistics An activation network containing fully connected layers for dimensionality reduction and expansion adaptively learns and outputs the weight distribution of each channel. The calculation formula is:

[0088]

[0089] in, For ReLU function, For the Sigmoid function, The weight matrix of the channel-reduced fully connected layer. To restore the original number of channels in the weight matrix of the upgraded fully connected layer;

[0090] The weight distribution Multiply the original feature matrix channel by channel to obtain the effective acoustic channel purified features;

[0091] In each layer of feature extraction, an asymmetric max-pooling operator with a preset kernel size is used to downsample the purified features of the effective acoustic channels to obtain a feature matrix with compressed frequency dimension.

[0092] In one implementation, step S106 includes:

[0093] The mean value of the frequency-dimensional compressed feature matrix is ​​calculated along the remaining frequency dimension, flattened, and transposed into a one-dimensional time series feature vector of a preset size.

[0094] The one-dimensional time-series feature vector is used to preserve the temporal location features of the acoustic events using global temporal coding, and a learnable location coding matrix containing the temporal location features is then used. Injected into the feature sequence after linear projection In the formula, the calculation method is:

[0095]

[0096] in, The input is a one-dimensional time series feature after the previous flattening operation. To map the one-dimensional time series feature vector to a linear projection weight matrix of the global temporal coding hidden layer dimension, The sequence after injecting location information;

[0097] Sequence after injection of position information Encoding based on a multi-head self-attention mechanism to obtain a fused feature sequence includes: encoding the sequence after injecting positional information... Linear mapping to query matrix Key matrix Sum matrix Furthermore, global similarity calculation and context aggregation are performed along the time dimension. The calculation formula is as follows:

[0098]

[0099] in, All are sequences after injection of position information Multiply by the corresponding learnable weight matrix to obtain, The feature dimension of the key vector is used to scale the dot product result to prevent gradient vanishing.

[0100] The fused feature sequence is layer-normalized and global average pooling is performed along the time dimension to obtain a global acoustic event representation vector. ;

[0101] The global acoustic event representation vector A multilayer perceptron classification with Dropout regularization is performed, and the predicted probability distribution of the category of the external acoustic event at the current fiber optic sensor point is output using the Softmax function. The calculation formula is as follows:

[0102]

[0103] in, This is the global acoustic event representation vector. and These are the weight matrix and bias vector of the final fully connected classification layer, respectively. This represents the predicted probability distribution output by the model.

[0104] In one implementation, the sequence after the injection of location information Linear mapping to query matrix Key matrix Sum matrix include:

[0105] The query matrix is ​​determined using the following formula. Key matrix Sum matrix :

[0106]

[0107] in, and Three independent, trainable weight matrices are used to weight the sequence after the position information is injected, respectively. Projected onto Q, K, V space.

[0108] The technical solution of this invention employs a polarization diversity receiver architecture to mitigate polarization fading in hardware. In software, it introduces a memory-level online enhancement strategy to force the model to learn physical invariants and utilizes an improved Squeeze-and-Excitation Network (SE) channel for noise reduction in collaboration with a TimeFormer global encoder to remove environmental noise while preserving the acoustic time axis. This entire solution achieves end-to-end optimization from signal perception to complex pattern classification, significantly reducing the system's false alarm rate and deployment threshold. It provides a novel distributed acoustic monitoring method with high reliability and high generalization capability for perimeter security, pipeline monitoring, and other fields, possessing significant engineering application value. Attached Figure Description

[0109] Figure 1 This is a schematic diagram of a distributed fiber optic acoustic wave sensing system according to an embodiment of this application;

[0110] Figure 2 This is a flowchart of a distributed fiber optic acoustic pattern recognition method according to an embodiment of this application;

[0111] Figure 3 This is a schematic diagram of the dual attention network structure and algorithm flow according to an embodiment of this application;

[0112] Figure 4This is a one-dimensional physical waveform and a two-dimensional log-Mel spectrum feature mapping diagram of a typical acoustic event according to an embodiment of this application;

[0113] Figure 5 This is a diagram illustrating the online dynamic spectrum masking data enhancement effect according to an embodiment of this application.

[0114] Figure 6 It is a graph showing the accuracy and loss reduction of the identification based on the baseline model of existing technology;

[0115] Figure 7 It is a confusion matrix diagram for acoustic pattern recognition based on the baseline model of existing technology;

[0116] Figure 8 This is a graph showing the decrease in recognition accuracy versus loss of the dual attention model according to an embodiment of this application;

[0117] Figure 9 This is a confusion matrix diagram for acoustic pattern recognition based on the dual attention model according to an embodiment of this application. Detailed Implementation

[0118] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, unless otherwise specified, the following embodiments and features described therein can be combined with each other.

[0119] Figure 1 This is a schematic diagram of a distributed fiber optic acoustic wave sensing system according to an embodiment of this application. The system includes:

[0120] Narrow linewidth laser 1, beam splitter fiber coupler 2, waveform generator 3, acousto-optic modulator 4, fiber amplifier 5, fiber optic circulator 6, sensor fiber under test 7, bundled fiber coupler 8, polarization beam splitter 9, first balanced photodetector 10, second balanced photodetector 11, data acquisition card 12, and host computer 13.

[0121] The continuous coherent light emitted by the narrow linewidth laser 1 is split into two paths by the beam splitter fiber coupler 2. One path is used as a local reference light and connected to the beam combiner fiber coupler 8. The other path is used as a probe light and enters the acousto-optic modulator 4 driven by the waveform generator 3 to be modulated into an optical pulse. The power is then amplified by the fiber amplifier 5. The amplified probe light pulse is injected into the sensing fiber 7 under test through the fiber circulator 6.

[0122] Rayleigh backscattered light returning from the sensing fiber 7 passes through the fiber optic circulator 6 and mixes with the local reference light at the fiber optic coupler 8. The mixed optical signal enters the polarization beam splitter 9 and is decomposed into two orthogonal polarized beams, which are then transmitted to the first balanced photodetector 10 and the second balanced photodetector 11 respectively and converted into electrical signals. The two orthogonal interference electrical signals are synchronously acquired by the data acquisition card 12 and sent to the host computer 13 for demodulation.

[0123] In this embodiment of the invention, the distributed fiber optic acoustic sensing system comprises two main parts: the first part is the front-end optical path sensing and data acquisition, mainly including a narrow-linewidth laser 1, an acousto-optic modulator 4, an optical fiber amplifier 5, an optical fiber circulator 6, a sensing fiber under test 7, a first balanced photodetector 10, a second balanced photodetector 11, and a high-speed data acquisition card 12. The continuous coherent light emitted by the narrow-linewidth laser 1 is modulated into pulsed light by the acousto-optic modulator 4, amplified by the erbium-doped fiber amplifier 5, and then injected into the sensing fiber under test 7 through the optical fiber circulator 6. The sensing fiber under test 7 experiences micro-strain along its path due to external sound waves or vibrations, causing the returned Rayleigh backscattered light to carry corresponding acoustic characteristic information. After the backscattered light is coherently mixed with the local reference light, it is converted into a one-dimensional analog electrical signal by the first balanced photodetector 10 and the second balanced photodetector 11. The high-speed data acquisition card 12 performs analog-to-digital conversion and transmits the signal to the host computer 13 in real time. The signal is then restored by the demodulation algorithm for subsequent intelligent analysis. The second part is the back-end artificial intelligence signal processing. On the host computer 13, the acquired data is demodulated and pattern recognized through feature mapping and preprocessing, online dynamic data enhancement, channel denoising feature extraction based on asymmetric pooling, and global temporal coding based on multi-head self-attention.

[0124] This invention addresses the bottleneck of existing distributed optical fiber acoustic sensing (DAS) systems in pattern recognition applications under complex environments and low sample size conditions. By employing a polarization diversity receiver architecture at the hardware level to mitigate polarization fading, it significantly reduces the false alarm rate and deployment threshold, providing a distributed acoustic monitoring method with both high reliability and high generalization capability for perimeter security, pipeline monitoring, and other fields, thus possessing significant engineering application value.

[0125] Distributed fiber optic acoustic sensing system such as Figure 1As shown, the system comprises a narrow-linewidth laser 1, a beam splitter fiber coupler 2, a waveform generator 3, an acousto-optic modulator 4, a fiber amplifier 5, a fiber optic circulator 6, a sensing fiber under test 7, a beam combiner fiber coupler 8, a polarization beam splitter 9, a first balanced photodetector 10, a second balanced photodetector 11, a data acquisition card 12, and a host computer 13. The continuous coherent light emitted from the narrow-linewidth laser 1 is split into two paths by the 95:5 (probe light (95%) and local reference light (5%)) beam splitter fiber coupler 2. One path serves as the local reference light and is directly connected to the 50:50 beam combiner fiber coupler 8; the other path serves as the probe light, entering the acousto-optic modulator 4 driven by the arbitrary waveform generator 3 to be modulated into an optical pulse, which is then amplified by the erbium-doped fiber amplifier 5. The amplified probe light pulse is injected into the sensing fiber under test 7 via the fiber optic circulator 6. The sensing fiber 7 under test experiences micro-strain along its length due to external sound waves or vibrations, causing changes in its refractive index and length. This results in the returned Rayleigh backscattered light carrying corresponding acoustic phase information. This backscattered light returns via the fiber circulator 6 and is remixed with the local reference light at the fiber coupler 8. The mixed optical signal enters the polarization beamsplitter 9, where it is decomposed into two orthogonal polarization states. These states are then fed into the first balanced photodetector 10 and the second balanced photodetector 11, respectively, and converted into electrical signals. This achieves polarization diversity reception, effectively overcoming the polarization fading effect caused by fiber birefringence.

[0126] like Figure 2 As shown, this embodiment of the invention also provides a distributed fiber optic acoustic pattern recognition method, comprising:

[0127] S101. Perform digital down-conversion on the acquired interferometric electrical signal to obtain the baseband in-phase component. Orthogonal components of baseband The baseband complex envelope signal is constructed from the baseband in-phase component and the baseband quadrature component. The baseband complex envelope signal is matched and filtered using a digital transmit pulse reference template corresponding to the acousto-optic modulator drive waveform to obtain the pulse-compressed baseband complex Rayleigh scattering signal.

[0128] S102. Extract the spatial differential phase of the pulse-compressed baseband complex Rayleigh scattering signal along the distance axis by complex conjugate multiplication, and determine the target spatial position of the acoustic event in the tested sensing fiber by combining the frequency band energy positioning algorithm. Perform bandpass filtering and downsampling on the slow time phase sequence of the target spatial position to obtain a one-dimensional time-domain acoustic signal.

[0129] S103. The one-dimensional time-domain acoustic signal is mapped to a two-dimensional Mel spectrum through short-time Fourier transform, and the local frequency energy distribution and time evolution trajectory corresponding to the two-dimensional Mel spectrum are extracted.

[0130] S104. Perform dynamic data augmentation on the two-dimensional Mel spectrogram and perform random spectrum masking and soft label data aliasing to obtain a dynamically augmented feature matrix.

[0131] S105. The noise floor is suppressed by adaptively calculating the channel weights, and the dynamic enhancement feature matrix is ​​downsampled by using an asymmetric max pooling operator to obtain a frequency dimension compressed feature matrix.

[0132] S106. Using a self-attention mechanism, capture the long-range dynamic evolution beat of the acoustic events corresponding to the frequency dimension compressed feature matrix across the time axis, and determine the category of the acoustic events of the tested sensing fiber through a fully connected network.

[0133] In this embodiment of the invention, two orthogonal interference electrical signals are synchronously acquired by the data acquisition card 12 and then sent to the host computer 13 for demodulation. The host computer 13 first performs acoustic wave physical reconstruction on the acquired interference electrical signals. The specific process is as follows: digital down-conversion is performed on the acquired interference electrical signals to extract the baseband complex envelope signal. Then, matched filtering is performed using the transmitted pulse reference signal to achieve pulse compression and spatial resolution enhancement, obtaining the baseband complex Rayleigh scattering signal. Next, the spatial differential phase of the baseband complex Rayleigh scattering signal is extracted along the distance axis by complex conjugate multiplication, and the target spatial position of the acoustic event is locked by combining the frequency band energy positioning algorithm. Then, the slow time phase sequence of the target spatial position is bandpass filtered and downsampled to reconstruct the real one-dimensional time-domain acoustic wave signal along the optical fiber with high fidelity. After the acoustic wave reconstruction is completed, this embodiment of the invention uses a demodulation algorithm based on channel spatiotemporal dual attention and dynamic enhancement to classify the one-dimensional time-domain acoustic wave signal into modes. First, a preprocessing algorithm is used to map the reconstructed one-dimensional time-domain acoustic signal into a two-dimensional Mel spectrogram via a short-time Fourier transform, in order to extract the local frequency energy distribution and time evolution trajectory. Second, an online dynamic data augmentation mechanism is introduced. During each forward propagation of the network, random spectral masking and soft-label data aliasing are directly performed on the two-dimensional Mel spectrogram in memory to obtain a dynamically enhanced feature matrix. This simulates frequency band interference and multi-source overlap environments, eliminating the data traversal risks caused by offline augmentation. Next, the enhanced high-dimensional features are fed into an improved SE channel denoising module. Adaptive calculation of channel weights suppresses background noise, and an asymmetric max-pooling mechanism with a kernel of (2,1) is used to downsample the dynamically enhanced feature matrix, obtaining a frequency-dimensional compressed feature matrix that fully preserves the time axis length of the acoustic signal while compressing the frequency dimension. Finally, the denoised and time-complete feature matrix is ​​fed into the TimeFormer global temporal coding module based on multi-head self-attention. Utilizing the self-attention mechanism, it accurately captures the long-range dynamic evolution of acoustic events corresponding to the frequency-dimensional compressed feature matrix across the time axis, and outputs precise acoustic event category decisions through a fully connected network. This complete processing flow effectively smooths polarization fading and accurately reconstructs the true sound field at the physical layer, while overcoming demodulation failures caused by strong noise and time delay distortion at the algorithmic layer, achieving highly robust and high-precision intelligent sensing.

[0134] The method of this invention abandons traditional offline data augmentation and conventional two-dimensional convolutional networks, and innovatively introduces a memory-based online dynamic enhancement and asymmetric pooling temporal tracking algorithm to achieve accurate discrimination of complex acoustic events. Addressing the low robustness problem caused by complex background noise, traditional networks lack channel-level filtering mechanisms and are easily misjudged by background noise interference; addressing the loss of long-term temporal features, traditional convolution irreversibly disrupts the dynamic evolution rhythm of sound waves during global spatial downsampling; and addressing the risk of overfitting under small sample conditions, conventional offline static augmentation easily leads to data crossing from the same source, causing a sharp drop in generalization ability during actual deployment. To solve these problems, this invention introduces a memory-level online enhancement strategy to force the model to learn physical invariants, and utilizes improved SE channel denoising and the TimeFormer global encoder in collaboration to remove environmental noise while preserving the sound wave time axis. The entire solution achieves end-to-end optimization from signal perception to complex pattern classification.

[0135] In this embodiment of the invention, step S101 includes:

[0136] The acquired interferometric electrical signals are added and fused using the X-polarization and Y-polarization channels to obtain the polarization-fused interferometric beat frequency electrical signal. , represented as:

[0137] ;

[0138] in, ;

[0139] Represents electron charge, Indicates the quantum efficiency of the detector. This represents Planck's constant. Indicates the photon frequency; and These represent the AC beat frequency current signals on the two positive traffic channels X and Y, respectively; and These represent the electric field intensity components of the probe light in the X and Y polarization directions, respectively; and This represents the electric field intensity components of the local reference light in the X and Y polarization directions;

[0140] The signal of the returned probe light is represented as follows:

[0141] ;

[0142] in, Indicates the optical frequency of the incident light. Indicates the frequency shift of the acousto-optic modulator. This represents the intensity of the backscattered Rayleigh light. This indicates the modulation phase caused by external sound waves or vibrations;

[0143] The reference light is represented as:

[0144] ;

[0145] in, Indicates the intensity of the reference light;

[0146] Let the equivalent interference amplitude after fusion be Interference beat frequency electrical signal after polarization fusion Simplified representation:

[0147] ;

[0148] The generated two mutually orthogonal digital reference signals are in sync with the frequency shift of the acousto-optic modulator. and The polarization-fused interferometric beat frequency electrical signal respectively with the digital reference signal and Multiply;

[0149] After multiplication, the in-phase branch signal is obtained. and orthogonal branch signals Represented as:

[0150]

[0151] A digital low-pass filter is used for the in-phase branch signal. and the orthogonal branch signal The process involves filtering out high-frequency harmonic components to obtain the equivalent interference amplitude. With slowly changing phase Baseband in-phase component of information Orthogonal components to baseband :

[0152]

[0153] in, Indicates a direct proportional relationship;

[0154] Baseband complex scattering signal From baseband in-phase components Orthogonal components of baseband The complex signal is calculated using the following formula: ;

[0155] Using a digital transmit pulse reference template corresponding to the acousto-optic modulator drive waveform For the baseband complex envelope signal Matched filtering is performed to obtain the pulse-compressed baseband complex Rayleigh scattering signal. , represented as: ;

[0156] in, This represents the baseband complex Rayleigh scattering signal after pulse compression and spatial resolution enhancement. This represents a digital transmit pulse reference template corresponding to the transmit probe light pulse waveform. This represents the convolution operation. This indicates the complex conjugate operation.

[0157] exist Figure 1 In this process, continuous light generated by a narrow-linewidth laser is split into probe light (95%) and local reference light (5%) after passing through a 95:5 coupler (beam splitter fiber coupler). The probe light enters an acousto-optic modulator (AOM), is modulated into pulsed light by an arbitrary waveform generator (AWG), amplified by an erbium-doped fiber amplifier, and then enters a fiber circulator before entering the sensor fiber under test. Backscattered Rayleigh light generated in the sensor fiber under test, after passing through the fiber circulator, interferes with the local reference light at a 50:50 coupler (beam combiner fiber coupler). The interference signal is split into two orthogonally polarized beams by a polarization beam splitter (PBS), which are received by a first balanced photodetector and a second balanced photodetector (BPD), respectively, and then collected by a data acquisition card and demodulated by a host computer.

[0158] The signal of the returned probe light can be represented as:

[0159]

[0160] in, Indicates the optical frequency of the incident light. Indicates the frequency shift of the acousto-optic modulator. This represents the intensity of the backscattered Rayleigh light. This indicates the modulation phase caused by external sound waves or vibrations.

[0161] The local reference light can be represented as:

[0162]

[0163] in, This indicates the intensity of the reference light.

[0164] In long-distance optical fibers under test, the polarization state of backscattered Rayleigh light evolves randomly due to the random birefringence effect. Direct beat frequency can cause drastic fluctuations in the amplitude of the interference signal, i.e., polarization fading. To eliminate this effect, polarization diversity reception technology is employed. The mixed optical signal after interference is decomposed into two mutually orthogonal polarization states by a polarization beam splitter (PBS). Typically, the polarization states are adjusted in the local oscillator path to ensure that the projected power of the local oscillator light in the two orthogonal directions of the PBS is equal.

[0165] Unlike ordinary single-ended detectors, this embodiment of the invention employs a dual-balanced photodetector (BPD). The dual-balanced detector effectively filters out the DC components generated by the local reference light and probe light, as well as the common-mode intensity noise of the light source, through differential subtraction, retaining only the high-frequency AC interference term. Therefore, after photoelectric conversion, the AC beat frequency current signals on the X and Y positive traffic channels are... and They can be represented as follows:

[0166]

[0167] in, Represents electron charge, Indicates the quantum efficiency of the detector. This represents Planck's constant. This represents the photon frequency. and These represent the electric field intensity components of the probe light in the X and Y polarization directions, respectively; and This represents the electric field intensity components of the reference light in the X and Y polarization directions. The interference signals from the X and Y polarization channels are added and fused. The resulting interference beat frequency electrical signal is obtained after polarization fusion. It can be represented as:

[0168]

[0169] To simplify the subsequent formula expression, let the equivalent interference amplitude after fusion be... Then we have:

[0170]

[0171] Subsequently, the fused high-frequency carrier signal is digitally quadrature demodulated. First, it enters the mixing stage, converting the baseband complex signal... Two mutually orthogonal digital reference signals generated by the computer and with the same frequency shift as the acousto-optic modulator. and Multiplication. After multiplication, the output in-phase branches are... and orthogonal branches for:

[0172]

[0173] The mixed signal then enters the filtering stage. A digital low-pass filter is used for filtering. and During processing, the high-frequency harmonic components are effectively filtered out by the filter. At this point, the high-frequency carrier in the signal has been eliminated, leaving only the baseband in-phase component carrying the original signal amplitude and slowly varying phase information. Orthogonal components to baseband :

[0174]

[0175] After obtaining a clean baseband signal, the phase demodulation stage begins.

[0176] The baseband complex envelope signal is composed of the baseband in-phase component. Orthogonal components of baseband The complex signal formed by the combination:

[0177]

[0178] Where t represents the slow-time sampling dimension, which reflects the sound wave vibrations at the same spatial location as they change over time.

[0179] Subsequently, the baseband complex envelope signal is obtained after digital down-conversion; then it is compared with the transmit pulse reference template. After matched filtering, the output is no longer the original high-frequency fused interference signal, but rather the baseband complex Rayleigh scattering signal after pulse compression. Among them, the baseband complex Rayleigh scattering signal It can reflect the amplitude and phase information of the Rayleigh backscattered signal, making it suitable as an object for subsequent spatial differential phase extraction.

[0180] In this embodiment of the invention, the transmit pulse reference template This refers to the digital reference template corresponding to the transmit probe pulse generated by the waveform generator driving the acousto-optic modulator, i.e., a digital copy of the transmit pulse waveform. This digital reference template is used in the host computer to perform matched filtering on the baseband complex envelope signal, thereby achieving pulse compression and spatial resolution enhancement.

[0181] The enhanced baseband complex Rayleigh scattering signal is obtained after matched filtering. .

[0182] In this embodiment of the invention, step S102 includes:

[0183] The baseband Rayleigh scattering signal after pulse compression The detection light pulses are framed according to their emission period to obtain the pulse-compressed Rayleigh scattering complex traces corresponding to different emission periods. Based on the round-trip propagation delay corresponding to each fast time sampling point in the complex traces, the round-trip propagation delay is mapped to the spatial sampling position of the sensing fiber under test along the distance axis, thus obtaining a two-dimensional complex signal characterized by both slow time and spatial position. ;

[0184] Where t represents the slow-time sampling sequence formed by continuously emitting probe light pulses, z represents the spatial sampling position along the length of the sensing fiber under test, and the spatial sampling position z is determined by the round-trip propagation delay corresponding to the fast-time sampling point;

[0185] The two-dimensional complex signal Represented as a complex number consisting of a real part and an imaginary part:

[0186]

[0187] in, This represents the in-phase component of the pulse-compressed baseband complex Rayleigh scattering signal at spatial position z. This represents the orthogonal component of the pulse-compressed baseband complex Rayleigh scattering signal at spatial position z;

[0188] Complex conjugate multiplication of the pulse-compressed baseband complex Rayleigh scattering signals at two spatial sampling positions z and z+L, separated by a fixed gauge length L, yields the differential complex signal C(t,z), expressed as:

[0189]

[0190] in, Represents the complex conjugate operation;

[0191] The argument of the differential complex signal C(t,z) is calculated to obtain the differential phase in the winding space. , represented as:

[0192]

[0193] Where Re[C(t,z)] and Im[C(t,z)] represent the real part and imaginary part of the differential complex signal, respectively, and atan2 is used to determine the corresponding phase angle based on the real part and imaginary part;

[0194] Combined with the phase unwinding algorithm, the differential phase of the winding space is analyzed. Perform 2kπ jump compensation to obtain spatial differential phase. ;

[0195] According to the spatial difference phase The spatial location of the acoustic event is determined by the response distribution along the distance axis of the sensing fiber under test, thus identifying the target spatial location z0 corresponding to the acoustic event.

[0196] Extract the slow-time spatial difference phase sequence at the target spatial location z0. and the slow time-space differential phase sequence Digital bandpass filtering and audio downsampling are performed to obtain a one-dimensional time-domain physical sound wave signal. , represented as:

[0197]

[0198] in,

[0199]

[0200] z0 represents the target spatial location where the acoustic event occurs, and z0+L represents another spatial location along the length of the sensing fiber being measured, which is a fixed gauge length L away from z0, where L is the phase difference gauge length.

[0201] In this embodiment of the invention, the equivalent interference amplitude can be eliminated. To mitigate the effects of random fluctuations, the entangled phase information corresponding to each sampling point on the tested sensing fiber is demodulated. This is then combined with a phase unwinding algorithm. The jump compensation yields the optical phase change (i.e., spatial differential phase). :

[0202] The extracted entanglement phase information still contains low-frequency thermal noise and common-mode phase drift from the laser. To reconstruct the true acoustic signal, a spatial difference algorithm needs to be introduced along the fiber distance axis. This is achieved by calculating the fixed gauge length on the tested sensing fiber. Spatial differential phase between two sampling points This not only cancels out global common-mode noise, but also the spatial differential phase The axial dynamic strain of the optical fiber induced by external physical sound waves exhibits a strictly linear proportional relationship. After passing this spatial differential phase sequence through slow-time-axis digital bandpass filtering and audio downsampling, it is reconstructed into a one-dimensional time-domain physical sound wave signal. :

[0203] ;

[0204] in,

[0205] ;

[0206] Thus, the extremely fragile beat frequency signal, which is prone to fading, has been successfully and robustly converted into a high-quality one-dimensional time-domain physical acoustic signal dataset.

[0207] In this embodiment of the invention, step S103 includes:

[0208] The one-dimensional time-domain physical acoustic signal at the time domain level Perform random waveform enhancement;

[0209] Enhanced one-dimensional time domain signal It can be represented as:

[0210]

[0211] in, A random gain factor that follows a uniform distribution is used to simulate the random fluctuations of sound source energy; The dynamically generated data follows a normal distribution based on a set signal-to-noise ratio. Gaussian white noise;

[0212] For the enhanced one-dimensional time-domain signal Perform feature extraction and map the extracted features into a two-dimensional log-Mel spectrum feature matrix that characterizes the time-frequency distribution of sound energy.

[0213] In this embodiment of the invention, in order to construct a high-dimensional feature space suitable for subsequent deep spatiotemporal attention networks and effectively overcome the overfitting problem caused by the scarcity of real samples in distributed fiber optic sensing, the one-dimensional time-domain physical acoustic signal needs to undergo multi-dimensional physical data enhancement and two-dimensional time-frequency feature mapping before the data is fed into the deep neural network. First, the one-dimensional time-domain physical acoustic signal is... Random waveform enhancement is performed. To simulate environmental distortions under complex real-world conditions, enhancement operations such as amplitude scaling, Gaussian white noise injection, time shifting, time-domain stretching, and fundamental frequency shifting are introduced through random combinations. Taking noise addition and amplitude scaling as an example, the enhanced one-dimensional time-domain signal... It can be represented as:

[0214]

[0215] in, A random gain factor that follows a uniform distribution is used to simulate the random fluctuations of sound source energy; The dynamically generated data follows a normal distribution based on a set signal-to-noise ratio. Gaussian white noise. This one-dimensional enhancement effectively expands the sample abundance of the original data.

[0216] In this embodiment of the invention, the enhanced one-dimensional time-domain signal Perform feature extraction, mapping the extracted features to a two-dimensional log-Mel spectrum feature matrix characterizing the time-frequency distribution of sound energy, including:

[0217] The time-domain signal is converted into a linear time-frequency complex matrix by the short-time Fourier transform, and the discrete calculation formula is as follows:

[0218]

[0219] in, For the number of points in the Fast Fourier Transform, The frame shift step size, For smoothing window functions, and These represent the time frame index and the frequency point index, respectively; n is the index of the discrete-time sampling point within each frame's short-time analysis window, with a value range of 0 ≤ n ≤ N. FFT -1;

[0220] The linear time-frequency complex matrix is ​​obtained by... Mel filter bank consisting of triangular filters Calculate the energy of each Mel frequency band to form the Mel energy spectrum. The calculation formula is as follows:

[0221]

[0222] Where k is the index of the linear frequency point after the short-time Fourier transform, and its value range is related to the number of points in the fast Fourier transform; for the complete spectrum, k = 0, 1, 2, ..., N FFT -1; For the one-sided spectrum of a real signal, k=0,1,2,…,N FFT / 2;

[0223] Taking the logarithm of the Mel energy spectrum and using the maximum value in the matrix as the reference point... It is converted to a decibel scale to obtain a two-dimensional log-Mel spectrum feature matrix. The calculation formula is:

[0224]

[0225] in, To prevent tiny extreme values ​​at logarithmic zero poles.

[0226] In this embodiment of the invention, the enhanced one-dimensional time-domain signal Feature extraction is performed, mapping the signal to a two-dimensional log-Mel spectrum representing the time-frequency distribution of sound energy. First, the time-domain signal is converted into a linear time-frequency complex matrix using a short-time Fourier transform; its discrete calculation formula is as follows:

[0227]

[0228] in, For the number of points in the Fast Fourier Transform, The frame shift step size, For smoothing window functions, and These represent the time frame index and the frequency point index, respectively. To make the feature distribution more consistent with acoustic perception characteristics, the linear time-frequency complex matrix is ​​transformed by... Mel filter bank consisting of triangular filters Calculate the energy of each Mel frequency band:

[0229]

[0230] To further compress the dynamic range of the data and highlight acoustic texture features, the logarithm of the Mel energy spectrum was taken, and the maximum value in the matrix was used as a reference point to convert it to a decibel (dB) scale:

[0231]

[0232] in, To prevent tiny extrema at logarithmic zeros and poles, after truncation and zero-padding, the matrix size of the output two-dimensional log-Mel spectrum characteristic matrix is ​​fixed at [value missing]. This matrix not only compresses redundant data but also preserves the characteristic beats of acoustic events as they evolve over time.

[0233] In this embodiment of the invention, step S104 includes:

[0234] The two-dimensional logarithmic Mel-frequency characteristic matrix As a feature sample, another feature sample X is randomly selected from the training dataset or the current training batch. j and its corresponding one-hot encoded label Y j , where X i and X j Y represents the two-dimensional logarithmic Mel-frequency characteristic matrix obtained after time-frequency mapping of two one-dimensional time-domain acoustic signals. i and Y j These represent the corresponding acoustic event category labels; the features and labels are linearly fused according to the following formula to obtain a hybrid two-dimensional feature matrix and hybrid soft labels:

[0235] ;

[0236] Among them, the random weights of soft label aliasing are generated according to the Beta distribution. , To control the hyperparameters of the Beta distribution shape, and These represent the fused hybrid feature matrix and the hybrid soft label, respectively.

[0237] A dynamic spectral masking operator is used to randomly generate time masking blocks and frequency masking blocks of a preset width on the two-dimensional feature matrix, and the energy values ​​within the time masking blocks and frequency masking blocks are forcibly set to zero to obtain a dynamically enhanced feature matrix.

[0238] In this embodiment of the invention, to achieve efficient supervised training and performance evaluation of the deep dual-attention network, a distributed fiber optic acoustic wave sensing public dataset containing real complex working condition samples is introduced for model construction. This dataset comprehensively covers 10 typical categories of natural and man-made acoustic physical events in the field, specifically including: vehicle horns, electric drills, footsteps, hand hammering, hand sawing, pneumatic drilling, natural rainfall, shoveling, thunderstorms, and electrical welding. The independent original samples for each category in the dataset are stored in the form of digital matrix files, which record the intensity and phase sequences captured by the tested sensing fiber under micro-strain. First, all sample files in the aforementioned 10 category folders are traversed and read, and then, using the aforementioned demodulation algorithm, they are converted one by one into a two-dimensional log-Mel spectrum matrix. Based on the physical category of the real acoustic event to which each sample belongs, a corresponding unique thermal coding label is assigned. Therefore, a dataset for model training was constructed. N is a positive integer representing the number of datasets.

[0239] During network training, to further smooth decision boundaries and prevent overfitting of the model to a limited number of samples, the system performs memory-level online dynamic data augmentation directly on the batch input two-dimensional feature matrix during each forward propagation. This mechanism includes soft label aliasing and dynamic spectral masking. For two randomly selected feature samples and their one-hot encoded labels... and Soft label aliasing uses random weights generated according to a Beta distribution. Perform linear fusion of features and labels:

[0240] ;

[0241] in, To control the hyperparameters of the Beta distribution shape, and These represent the fused feature matrix and the fused soft label, respectively. The two feature samples here are two complete training samples randomly selected from the training dataset or the current training batch. Each feature sample X... i or X j Both are two-dimensional logarithmic Mel-frequency characteristic matrices obtained after short-time Fourier transform and Mel-filtering of a one-dimensional time-domain acoustic signal, with corresponding labels Y and Y respectively. i and Y jThe soft label aliasing operation is a linearly weighted fusion of the feature matrices and their labels from two complete samples.

[0242] Meanwhile, the dynamic spectral masking operator randomly generates time and frequency masking blocks of specific widths on the two-dimensional feature matrix and forces the energy values ​​within these blocks to zero. Since this masking operation is dynamically generated during each network iteration, it forces the model to learn robust acoustic features from spectra with varying degrees of damage each time, eliminating the risk of permanent loss of key features due to static masking.

[0243] In this embodiment of the invention, step S105 includes:

[0244] An improved squeeze-excitation channel denoising process is applied to the dynamically enhanced feature matrix;

[0245] The improved squeeze-excitation channel denoising process includes three cascaded feature extraction layers. In each special zone extraction layer, the input feature matrix is ​​sequentially processed by two-dimensional convolution, batch normalization, and ReLU activation function.

[0246] The denoised feature matrix is ​​subjected to SE channel attention processing, and each feature channel is calculated using global average pooling. Global spatial statistics The calculation formula is:

[0247]

[0248] in, Representation of the characteristic matrix In the Each channel, spatial coordinates are Pixel feature value at that location and These represent the frequency dimension and time dimension of the feature matrix, respectively;

[0249] For the global spatial statistics An activation network containing fully connected layers for dimensionality reduction and expansion adaptively learns and outputs the weight distribution of each channel. The calculation formula is:

[0250]

[0251] in, For ReLU function, For the Sigmoid function, The weight matrix of the channel-reduced fully connected layer. To restore the original number of channels in the weight matrix of the upgraded fully connected layer;

[0252] The weight distribution Multiply the original feature matrix channel by channel to obtain the effective acoustic channel purified features;

[0253] In each layer of feature extraction, an asymmetric max-pooling operator with a preset kernel size is used to downsample the purified features of the effective acoustic channels to obtain a feature matrix with compressed frequency dimension.

[0254] In this embodiment of the invention, a dynamically enhanced high-dimensional feature matrix is ​​input into a dual attention network. The first layer of attention in the dual attention network consists of an improved squeeze-excitation channel denoising module (SE-CNN). This improved squeeze-excitation channel denoising module comprises three cascaded feature extraction blocks. In each feature extraction block, the input features first undergo two-dimensional convolution, batch normalization, and ReLU activation, and then enter the SE channel attention mechanism. The SE channel attention mechanism first calculates the value of each feature channel through global average pooling. Global spatial statistics :

[0255]

[0256] in, Representation of the characteristic matrix In the Each channel, spatial coordinates are Pixel feature value at that location and These represent the frequency and time dimensions of the feature matrix, respectively. Then, through an activation network containing fully connected layers for dimensionality reduction and expansion, the weight distribution of each channel is adaptively learned and output. :

[0257]

[0258] in, For ReLU function, For the Sigmoid function, The weight matrix of the channel-reduced fully connected layer. To restore the weight matrix of the upgraded fully connected layer with the original number of channels, the weights are distributed as follows. Channel-by-channel multiplication with the original feature matrix achieves suppression of the noise floor channels and purification of the effective acoustic channels. Then, an asymmetric max-pooling operator (with a kernel size of up to [missing information]) is used in each cascaded block. The purified features are downsampled. This operator only halves the frequency dimension of the feature matrix, while strictly and losslessly preserving the time axis length of the acoustic signal. This provides a complete feature foundation for subsequent long-range time series modeling.

[0259] In this embodiment of the invention, cascaded feature extraction blocks are used in the backbone of the neural network. Specifically, in the improved SE-CNN channel denoising module, multiple layers of continuously connected feature extraction blocks are set up. Each feature extraction block includes two-dimensional convolution, batch normalization, ReLU activation, SE channel attention weighting, and asymmetric max pooling operation in sequence. The output of the previous feature extraction block is used as the input of the next feature extraction block, hence the term "cascaded". The asymmetric max pooling operator refers to a pooling window with different sizes in the frequency and time dimensions. For example, when using a max pooling operator with a kernel size of (2,1), only a 2x downsampling is performed in the frequency dimension, while the length remains unchanged in the time dimension. Its function is to compress the frequency dimension and reduce the computational load while preserving as much of the dynamic evolution information of the acoustic signal along the time axis as possible.

[0260] In this embodiment of the invention, step S106 includes:

[0261] The mean value of the frequency-dimensional compressed feature matrix is ​​calculated along the remaining frequency dimension, flattened, and transposed into a one-dimensional time series feature vector of a preset size.

[0262] The one-dimensional time-series feature vector is used to preserve the temporal location features of the acoustic events using global temporal coding, and a learnable location coding matrix containing the temporal location features is then used. Injected into the feature sequence after linear projection In the formula, the calculation method is:

[0263]

[0264] in, The input is a one-dimensional time series feature after the previous flattening operation. To map the one-dimensional time series feature vector to a linear projection weight matrix of the global temporal coding hidden layer dimension, The sequence after injecting location information;

[0265] Sequence after injection of position information Encoding based on a multi-head self-attention mechanism to obtain a fused feature sequence includes: encoding the sequence after injecting positional information... Linear mapping to query matrix Key matrix Sum matrix Furthermore, global similarity calculation and context aggregation are performed along the time dimension. The calculation formula is as follows:

[0266]

[0267] in, All are sequences after injection of position information Multiply by the corresponding learnable weight matrix to obtain, The feature dimension of the key vector is used to scale the dot product result to prevent gradient vanishing.

[0268] The fused feature sequence is layer-normalized and global average pooling is performed along the time dimension to obtain a global acoustic event representation vector. ;

[0269] The global acoustic event representation vector A multilayer perceptron classification with Dropout regularization is performed, and the predicted probability distribution of the category of the external acoustic event at the current fiber optic sensor point is output using the Softmax function. The calculation formula is as follows:

[0270]

[0271] in, This is the global acoustic event representation vector. and These are the weight matrix and bias vector of the final fully connected classification layer, respectively. This represents the predicted probability distribution output by the model.

[0272] After channel denoising and frequency dimension compression by multiple SE-CNN modules, this embodiment of the invention calculates the mean along the residual frequency dimension of the output high-dimensional feature matrix, flattens it, and transposes it into a shape of... The sequence is then processed into a one-dimensional time-series feature vector. This sequence subsequently enters the network's second attention module, the global temporal encoding module (TimeFormer). To preserve the chronological order of acoustic events, a learnable positional encoding matrix containing temporal positional features is generated. Injected into the feature sequence after linear projection middle:

[0273]

[0274] in, The input is a one-dimensional time series feature after the previous flattening operation. To map it to a linear projection weight matrix of the Transformer hidden layer dimension, The final sequence contains location information. The sequence, after being injected with temporal location features, is fed into a Transformer encoder layer based on a multi-head self-attention mechanism. The feature sequence is linearly mapped to a query matrix. Key matrix Sum matrix And perform global similarity calculation and context aggregation in the time dimension:

[0275]

[0276] in, All are composed of characteristic sequences Multiply by the corresponding learnable weight matrix to obtain, The feature dimension of the key vector is used to scale the dot product result to prevent gradient vanishing. This mechanism breaks the receptive field limitation of local convolution, enabling feature tracing directly on the global time axis across extremely long time steps, accurately capturing the dynamic evolutionary rhythm of complex acoustic events.

[0277] The fused feature sequence, after being encoded by multiple TimeFormer layers, is finally fed into the pattern classification decision head. First, it undergoes layer normalization, and then global average pooling is performed along the time dimension to losslessly compress the variable-length temporal features into a fixed-length global acoustic event representation vector. Finally, this vector is passed through a multilayer perceptron classifier with Dropout regularization, and the Softmax function is used to output the predicted probability distribution of the category of the external acoustic event at the current fiber optic sensing point.

[0278]

[0279] in, This is the compressed, fixed-length global acoustic event representation vector. and These are the weight matrix and bias vector of the final fully connected classification layer, respectively. This represents the predicted probability distribution output by the model. During the model optimization phase, the predicted probabilities are calculated using the cross-entropy loss function with a label smoothing factor. The error between the actual label and the target label is addressed by combining an adaptive moment estimation weight decay optimizer with a cosine annealing learning rate scheduling strategy to complete end-to-end deep network training. This achieves a complete closed loop from the physical extraction of weak phases in the underlying optical fiber to intelligent tracking of high-dimensional spatiotemporal joint features, significantly improving the robustness and accuracy of distributed optical fiber acoustic pattern recognition.

[0280] In this embodiment of the invention, the sequence after the injection of position information is... Linear mapping to query matrix Key matrix Sum matrix include:

[0281] The query matrix is ​​determined using the following formula. Key matrix Sum matrix :

[0282]

[0283] in, and Three independent, trainable weight matrices are used to weight the sequence after the position information is injected, respectively. Projected onto Q, K, V space.

[0284] In this embodiment of the invention, The temporal feature matrix after injection position encoding; and Three independent, trainable weight matrices are used to project the input sequence into the Q, K, and V spaces, respectively. In multi-head attention mechanisms, typically... , where h is the number of attention heads. The resulting Q, K, and V are then used in the global similarity calculation and context aggregation formula to perform global temporal self-attention calculation.

[0285] like Figure 3 The diagram shows a flowchart of a distributed fiber optic acoustic pattern recognition method with a channel-spatiotemporal dual attention mechanism and dynamic enhancement provided by an embodiment of the present invention. As can be seen from the diagram, the intelligent recognition model has a modular cascaded architecture, mainly comprising three core modules from top to bottom: First, a data input and online enhancement module, which takes the preprocessed two-dimensional Mel-frequency spectrum matrix as input and randomly introduces soft label aliasing (Mixup) and dynamic spectral masking (SpecAugment) during the forward propagation of network training to achieve online distortion and augmentation of the data; second, a channel-spatiotemporal dual attention backbone network, where the enhanced features first pass through a three-layer cascaded SE-CNN channel denoising module (including two-dimensional convolution, SE channel attention, and asymmetric max pooling with unidirectional frequency compression). The process involves several steps: first, high-frequency noise is removed while preserving time-axis features; then, dimensionality transformation, linear projection, and positional encoding injection are performed before the data is fed into the TimeFormer global temporal encoding module (which includes a multi-head self-attention mechanism and a feedforward network) to capture long-range beats of acoustic features in the global time dimension; finally, a pattern classification decision head is used, where high-dimensional features encoded with dual attention are passed through layer normalization, global average pooling (GAP), and a multilayer perceptron (MLP) in sequence, and finally, a fully connected classification layer combined with a Softmax function outputs the predicted probability distribution of the fiber optic sensing data belonging to each acoustic event category.

[0286] like Figure 4The figure shows a comparison of data preprocessing and feature mapping for four typical physical acoustic events provided in this embodiment of the invention. The figure illustrates four different types of acoustic signals: (a) car horn, (b) drilling, (c) footsteps, and (d) thunderstorm. In each sub-figure, the upper part is the high-fidelity one-dimensional time-domain waveform signal demodulated and restored by the system front-end, reflecting the dynamic strain evolution of different physical events in the time domain; the lower part is the two-dimensional logarithmic Mel spectrum generated after short-time Fourier transform and Mel filter bank processing, with time on the horizontal axis and Mel frequency on the vertical axis. As can be seen from the figure, after mapping, redundant environmental noise is effectively filtered out, and the unique frequency distribution and transient beats of different acoustic events are clearly displayed in the two-dimensional spectrum matrix, providing characteristic input data for subsequent deep attention networks.

[0287] like Figure 5 The image shown is a diagram illustrating the spectral masking effect in the online dynamic data augmentation operation provided by an embodiment of the present invention. Figure 5 (a) in the figure is the original two-dimensional logarithmic Mel spectrum, which contains the complete time-frequency energy distribution of the sound wave; Figure 5 (b) shows the damaged spectrum after dynamically introducing the masking operator during the forward propagation phase of network training. The system randomly generates horizontal frequency masking blocks and vertical temporal masking blocks on the spectrum, and the energy values ​​of the masked regions are forced to zero. Since this masking operation is dynamically generated in each iteration of the network, it effectively forces the dual attention model to avoid over-reliance on a single local feature, and instead requires robust decision-making based on global temporal and frequency context, thereby fundamentally eliminating the overfitting risk caused by static small samples.

[0288] like Figure 6 The figure shows the performance curve of a traditional pure convolutional neural network baseline model without the dual attention mechanism of this invention during the training process. Figure 6 (a) shows how the accuracy on the training and validation sets changes with the number of iterations. Figure 6 Figure (b) illustrates the decrease in the training set loss value. As can be seen from the figure, the traditional baseline model exhibits significant bottlenecks in extracting complex fiber acoustic features: the model converges slowly, the loss decrease is accompanied by severe oscillations, and the validation set accuracy stops increasing after reaching approximately 80%, making it difficult to achieve higher-precision pattern classification. This fully exposes the technical limitations of traditional locally particulate convolutional networks in handling long-range temporal distortions and strong background noise.

[0289] like Figure 8The figure shows the training performance curves of the SE-TimeFormer core algorithm architecture (channel-spatiotemporal dual attention mechanism combined with dynamic enhancement) of this invention. Figure 8 (a) in the diagram demonstrates an exceptionally high accuracy contrast: due to the significant learning difficulty introduced by the online dynamic augmentation mechanism on the training set, the training set accuracy exhibits a certain degree of low-level oscillation; however, this forces the model to learn extremely fundamental physical acoustic features, enabling it to demonstrate astonishing generalization ability on a clean validation set, with the validation set accuracy rapidly climbing and stabilizing close to 100% within a very short number of iterations. Combined with... Figure 8 The smooth and rapidly converging loss descent curve in (b) strongly demonstrates that the dual attention network proposed in this invention has extremely high feature purification efficiency and anti-overfitting ability.

[0290] like Figure 9 The image shows the pattern recognition normalized confusion matrix generated by the SE-TimeFormer model of this invention on the final test set. (And...) Figure 7 In stark contrast to traditional models, the confusion matrix output by the present invention exhibits an extremely dark color distribution along its main diagonal. The correct recognition rate for almost all categories reaches 1.00 (100%) or 0.99 (99%), while the misclassification probability in the off-diagonal region is almost entirely eliminated. This result intuitively and conclusively demonstrates that, thanks to the adaptive noise reduction of the SE module and the global beat-locking capability of the TimeFormer module across extremely long distances, the method of this invention overcomes the recognition challenges caused by strong background noise interference and long-term feature distortion in traditional distributed fiber optic acoustic sensing. It achieves high-precision, highly robust intelligent discrimination for complex working condition data, representing a significant technological advancement.

[0291] The aforementioned processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0292] The aforementioned memory can be an internal storage unit of the electronic device, such as a hard drive or RAM. Alternatively, it can be an external storage device, such as a plug-in hard drive, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card. Furthermore, the memory can include both internal and external storage units. The memory is used to store the computer program and other programs and data required by the electronic device. It can also be used to temporarily store data that has been output or will be output.

[0293] The various techniques described herein can be implemented in combination with hardware or software, or a combination thereof. The methods and apparatus of the present invention, or certain aspects or portions thereof, may take the form of program code (i.e., instructions) embedded in a tangible medium, such as a removable hard disk, USB flash drive, floppy disk, CD-ROM, or any other machine-readable storage medium, wherein when the program is loaded into and executed by a machine such as a computer, the machine becomes an apparatus for practicing the present invention.

[0294] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of this disclosure. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and embodiments are to be considered exemplary only.

[0295] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

[0296] The various techniques described herein can be implemented in combination with hardware or software, or a combination thereof. Thus, the methods and apparatus of the present invention, or certain aspects or portions thereof, can take the form of program code (i.e., instructions) embedded in a tangible medium, such as a removable hard disk, USB flash drive, floppy disk, CD-ROM, or any other machine-readable storage medium, wherein when the program is loaded into and executed by a machine such as a computer, the machine becomes an apparatus for practicing the present invention.

Claims

1. A distributed fiber optic acoustic pattern recognition method, characterized in that, include: S101. Perform digital down-conversion on the acquired interferometric electrical signal to obtain the baseband in-phase component. Orthogonal components of baseband The baseband complex envelope signal is constructed from the baseband in-phase component and the baseband quadrature component. ; The baseband complex envelope signal is matched and filtered using a digital transmit pulse reference template corresponding to the acousto-optic modulator drive waveform to obtain the pulse-compressed baseband complex Rayleigh scattering signal. S102. Extract the spatial differential phase of the pulse-compressed baseband complex Rayleigh scattering signal along the distance axis by complex conjugate multiplication, and determine the target spatial position of the acoustic event in the tested sensing fiber by combining the frequency band energy positioning algorithm. Perform bandpass filtering and downsampling on the slow time phase sequence of the target spatial position to obtain a one-dimensional time-domain acoustic signal. S103. The one-dimensional time-domain acoustic signal is mapped to a two-dimensional Mel spectrum through short-time Fourier transform, and the local frequency energy distribution and time evolution trajectory corresponding to the two-dimensional Mel spectrum are extracted. S104. Perform dynamic data augmentation on the two-dimensional Mel spectrogram and perform random spectrum masking and soft label data aliasing to obtain a dynamically augmented feature matrix. S105. The background noise is suppressed by adaptively calculating the weights, and the dynamic enhancement feature matrix is ​​downsampled by using an asymmetric max pooling operator to obtain a frequency dimension compressed feature matrix. S106. Using a self-attention mechanism, capture the long-range dynamic evolution rhythm of the acoustic events corresponding to the frequency dimension compressed feature matrix across the time axis, and determine the category of the acoustic events of the tested sensing fiber through a fully connected network. Step S101 includes: The acquired interferometric electrical signals are added and fused using the X-polarization and Y-polarization channels to obtain the polarization-fused interferometric beat frequency electrical signal. , is represented as: ; in, ; e Represents electron charge, μ Indicates the quantum efficiency of the detector. h This represents Planck's constant. w g Indicates the photon frequency; and These represent the AC beat frequency current signals on the two positive traffic channels X and Y, respectively; and These represent the electric field intensity components of the probe light in the X and Y polarization directions, respectively; and This represents the electric field intensity components of the local reference light in the X and Y polarization directions; The signal of the returned probe light is represented as follows: ; in, ν Δ represents the optical frequency of the incident light. ν Indicates the frequency shift of the acousto-optic modulator. This represents the intensity of the backscattered Rayleigh light. This indicates the modulation phase caused by external sound waves or vibrations; The reference light is represented as: ; in, Indicates the intensity of the reference light; Let the equivalent interference amplitude after fusion be Interference beat frequency electrical signal after polarization fusion Simplified representation: ; The generated two mutually orthogonal digital reference signals are in sync with the frequency shift of the acousto-optic modulator. and The polarization-fused interferometric beat frequency electrical signal respectively with the digital reference signal and Multiply; After multiplication, the in-phase branch signal is obtained. and orthogonal branch signals Represented as: ; A digital low-pass filter is used for the in-phase branch signal. and the orthogonal branch signal The process involves filtering out high-frequency harmonic components to obtain the equivalent interference amplitude. With slowly changing phase Baseband in-phase component of information Orthogonal components to baseband : ; in, Indicates a direct proportional relationship; Baseband complex envelope signal From baseband in-phase components Orthogonal components of baseband The complex signal is calculated using the following formula: ; Using a digital transmit pulse reference template corresponding to the acousto-optic modulator drive waveform For the baseband complex envelope signal Matched filtering is performed to obtain the pulse-compressed baseband complex Rayleigh scattering signal. , is represented as: ; in, This represents the baseband complex Rayleigh scattering signal after pulse compression and spatial resolution enhancement. This represents a digital transmit pulse reference template corresponding to the transmit probe light pulse waveform. This represents the convolution operation. Represents the complex conjugate operation; Step S102 includes: The baseband Rayleigh scattering signal after pulse compression The detection light pulses are framed according to their emission period to obtain the pulse-compressed Rayleigh scattering complex traces corresponding to different emission periods. Based on the round-trip propagation delay corresponding to each fast time sampling point in the complex traces, the round-trip propagation delay is mapped to the spatial sampling position of the sensing fiber under test along the distance axis, thus obtaining a two-dimensional complex signal characterized by both slow time and spatial position. ; Where t represents the slow-time sampling sequence formed by continuously emitting probe light pulses, z represents the spatial sampling position along the length of the sensing fiber under test, and the spatial sampling position z is determined by the round-trip propagation delay corresponding to the fast-time sampling point; The two-dimensional complex signal Represented as a complex number consisting of a real part and an imaginary part: ; in, This represents the in-phase component of the pulse-compressed baseband complex Rayleigh scattering signal at spatial position z. This represents the orthogonal component of the pulse-compressed baseband complex Rayleigh scattering signal at spatial position z; Complex conjugate multiplication is performed on the pulse-compressed baseband complex Rayleigh scattering signals at two spatial sampling positions z and z+L, which are separated by a fixed gauge length L, to obtain the differential complex signal. C ( t , z ), represented as: ; For the differential complex signal C ( t , z To obtain the argument, the differential phase in the winding space is obtained. , is represented as: ; Among them, Re[ C ( t , z )] and Im[ C ( t , z )] represent the real part and imaginary part of the differential complex signal, respectively, and atan2 is used to determine the corresponding phase angle based on the real part and imaginary part; Combined with the phase unwinding algorithm, the differential phase of the winding space is analyzed. Perform 2kπ jump compensation to obtain spatial differential phase. ; According to the spatial difference phase The spatial location of the acoustic event is determined by the response distribution along the distance axis of the sensing fiber under test, thus identifying the target spatial location z0 corresponding to the acoustic event. Extract the slow-time spatial difference phase sequence at the target spatial location z0. and the slow time-space differential phase sequence Digital bandpass filtering and audio downsampling are performed to obtain a one-dimensional time-domain physical sound wave signal. , is represented as: ; in, ; z0 represents the target spatial location where the acoustic event occurs, and z0+L represents another spatial location along the length of the measured sensing fiber that is a fixed gauge length L away from z0, where L is the phase difference gauge length. Step S103 includes: The one-dimensional time-domain physical acoustic signal at the time domain level Perform random waveform enhancement; Enhanced one-dimensional time domain signal It can be represented as: ; in, A random gain factor that follows a uniform distribution is used to simulate the random fluctuations of sound source energy; The dynamically generated data follows a normal distribution based on a set signal-to-noise ratio. Gaussian white noise; For the enhanced one-dimensional time-domain signal Perform feature extraction and map the extracted features into a two-dimensional log-Mel spectrum feature matrix representing the time-frequency distribution of sound energy; For the enhanced one-dimensional time-domain signal Perform feature extraction, mapping the extracted features to a two-dimensional log-Mel spectrum feature matrix characterizing the time-frequency distribution of sound energy, including: The time-domain signal is converted into a linear time-frequency complex matrix by the short-time Fourier transform, and the discrete calculation formula is as follows: ; in, For the number of points in the Fast Fourier Transform, HOP The frame shift step size, For smoothing window functions, m and k These represent the time frame index and the frequency point index, respectively. n This is the index of the discrete-time sampling point within each frame's short-time analysis window, with a value range of 0 ≤ n ≤ N FFT -1; The linear time-frequency complex matrix is ​​obtained by... Mel filter bank consisting of triangular filters Calculate the energy of each Mel frequency band to form the Mel energy spectrum. The calculation formula is as follows: ; Where k is the index of the linear frequency point after the short-time Fourier transform, and its value range is related to the number of points in the fast Fourier transform; for the complete spectrum, k = 0, 1, 2, ... N FFT -1; For the one-sided spectrum of a real signal, k=0,1,2,…, N FFT / 2; Taking the logarithm of the Mel energy spectrum and using the maximum value in the matrix as the reference point... It is converted to a decibel scale to obtain a two-dimensional log-Mel spectrum feature matrix. The calculation formula is: ; in, To prevent tiny extrema at the logarithmic zero poles; Step S104 includes: The two-dimensional logarithmic Mel spectrum feature matrix As a feature sample, another feature sample is randomly selected from the training dataset or the current training batch. X j and its corresponding uniquely coded label Y j ,in, X i and X j These represent the two-dimensional logarithmic Mel-frequency characteristic matrices obtained after time-frequency mapping of two one-dimensional time-domain acoustic signals. Y i and Y j These represent the corresponding acoustic event category labels; the features and labels are linearly fused according to the following formula to obtain a hybrid two-dimensional feature matrix and hybrid soft labels: ; Among them, the random weights of soft label aliasing are generated according to the Beta distribution. λ , α To control the hyperparameters of the Beta distribution shape, and These represent the fused hybrid feature matrix and the hybrid soft label, respectively. A dynamic spectral masking operator is used to randomly generate time masking blocks and frequency masking blocks of a preset width on the two-dimensional feature matrix, and the energy values ​​within the time masking blocks and frequency masking blocks are forcibly set to zero to obtain a dynamically enhanced feature matrix; Step S105 includes: An improved squeeze-excitation channel denoising process is applied to the dynamically enhanced feature matrix; The improved squeeze-excitation channel denoising process includes three cascaded feature extraction layers. In each special zone extraction layer, the input feature matrix is ​​sequentially processed by two-dimensional convolution, batch normalization, and ReLU activation function. The denoised feature matrix is ​​subjected to SE channel attention processing, and each feature channel is calculated using global average pooling. c Global spatial statistics z c The calculation formula is: ; in, Representation of the characteristic matrix U In the c Each channel, spatial coordinates are ( i , j The pixel feature value at ) F and T These represent the frequency dimension and time dimension of the feature matrix, respectively; For the global spatial statistics z c An activation network containing fully connected layers for dimensionality reduction and expansion adaptively learns and outputs the weight distribution of each channel. s c The calculation formula is: ; in, For ReLU function, For the Sigmoid function, The weight matrix of the channel-reduced fully connected layer. To restore the original number of channels in the weight matrix of the upgraded fully connected layer; The weight distribution Multiply the original feature matrix channel by channel to obtain the effective acoustic channel purified features; In each layer of feature extraction, an asymmetric max pooling operator with a preset kernel size is used to downsample the purified features of the effective acoustic channel to obtain a feature matrix with frequency dimension compression. Step S106 includes: The mean value of the frequency-dimensional compressed feature matrix is ​​calculated along the remaining frequency dimension, flattened, and transposed into a one-dimensional time series feature vector of a preset size. The one-dimensional time-series feature vector is used to preserve the temporal location features of the acoustic events using global temporal coding, and a learnable location coding matrix containing the temporal location features is then used. Injected into the feature sequence after linear projection In the formula, the calculation method is: ; in, The input is a one-dimensional time series feature after the previous flattening operation. To map the one-dimensional time series feature vector to a linear projection weight matrix of the global temporal coding hidden layer dimension, The sequence after injecting location information; Sequence after injection of position information Encoding based on a multi-head self-attention mechanism to obtain a fused feature sequence includes: encoding the sequence after injecting positional information... Linear mapping to query matrix Q Key matrix K Sum matrix V Furthermore, global similarity calculation and context aggregation are performed along the time dimension. The calculation formula is as follows: ; in, Q , K , V All are sequences after injection of position information Multiply by the corresponding learnable weight matrix to obtain, The feature dimension of the key vector is used to scale the dot product result to prevent gradient vanishing. The fused feature sequence is layer-normalized and global average pooling is performed along the time dimension to obtain a global acoustic event representation vector. ; The global acoustic event representation vector A multilayer perceptron classification with Dropout regularization is performed, and the predicted probability distribution of the category of the external acoustic event at the current fiber optic sensor point is output using the Softmax function. The calculation formula is as follows: ; in, This is the global acoustic event representation vector. and These are the weight matrix and bias vector of the final fully connected classification layer, respectively. This represents the predicted probability distribution output by the model.

2. The pattern recognition method according to claim 1, characterized in that, The sequence after injecting the location information Linear mapping to query matrix Q Key matrix K Sum matrix V include: The query matrix is ​​determined using the following formula. Q Key matrix K Sum matrix V : ; in, and Three independent, trainable weight matrices are used to weight the sequence after the position information is injected, respectively. Project to Q , K , V space.

3. A distributed fiber optic acoustic wave sensing system, characterized in that, Implementing the distributed fiber optic acoustic pattern recognition method according to any one of claims 1-2, comprising: Narrow linewidth laser (1), fiber optic splitter (2), waveform generator (3), acousto-optic modulator (4), fiber amplifier (5), fiber optic circulator (6), sensing fiber under test (7), fiber optic coupler (8), polarization beam splitter (9), first balanced photodetector (10), second balanced photodetector (11), data acquisition card (12), and host computer (13). The continuous coherent light emitted by the narrow linewidth laser (1) is split into two paths by the beam splitter fiber coupler (2). One path is connected to the beam combiner fiber coupler (8) as a local reference light, and the other path is used as a probe light. It is modulated into an optical pulse by the acousto-optic modulator (4) driven by the waveform generator (3) and then amplified by the fiber amplifier (5). The amplified probe light pulse is injected into the sensing fiber (7) under test through the fiber circulator (6). Rayleigh backscattered light returning from the sensing fiber (7) passes through the fiber circulator (6) and mixes with the local reference light at the fiber coupler (8). The mixed light signal enters the polarization beam splitter (9) and is decomposed into two orthogonal polarized beams, which are then transmitted to the first balanced photodetector (10) and the second balanced photodetector (11) respectively to be converted into electrical signals. The two orthogonal interference electrical signals are synchronously acquired by the data acquisition card (12) and sent to the host computer (13) for demodulation.

Citation Information

Patent Citations

  • High-resolution distributed acoustic vibration demodulation method based on sweep frequency light depth fading resistance

    CN122171010A

  • DiffOptics: A conditional diffusion model for optics sensing data imputation

    US20260110572A1