Double-branch sound-vibration fusion event identification and positioning method based on DAS and AI

Through the dual-branch sound and vibration fusion method combined with DAS and AI, high-precision event recognition and positioning of the infrastructure is achieved, and the problems of insufficient recognition accuracy, positioning error and robustness in the existing technology are solved, and an efficient and accurate monitoring solution is provided.

CN120508911AActive Publication Date: 2025-08-19ZHILIAN XINNENG POWER TECH CO LTD

Patent Information

Application Number
CN202510977594.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-08-19
Estimated Expiration
2045-07-16

AI Technical Summary

Technical Problem

The existing DAS technology has problems in infrastructure monitoring, such as insufficient event recognition accuracy, poor positioning accuracy and stability, insufficient environmental adaptability and robustness, limited intelligence, and low real-time and computational efficiency, especially in complex noise environments and long-distance monitoring.

Method used

The dual-branch sound and vibration fusion event recognition and positioning method based on DAS and AI is adopted to realize intelligent event classification through sound recognition branches, the vibration positioning branches achieve high-precision spatio-temporal positioning, and intelligent collaboration of multimodal data is achieved through the fusion and verification module. The specific steps include collecting sound and vibration data from the DAS system, performing multi-scale feature extraction and weighted fusion, using CNN and Bi-LSTM for timing modeling, and combining four-point spatiotemporal weighting optimization method and dynamic likelihood ratio for event positioning and verification.

Benefits of technology

It breaks through the bottlenecks of insufficient identification accuracy and large positioning errors in complex environments by traditional technologies, realizes efficient, accurate and stable infrastructure monitoring, and improves the robustness and real-time processing capabilities of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120508911A_ABST
    Figure CN120508911A_ABST
Patent Text Reader

Abstract

The invention discloses a double-branch sound-vibration fusion event recognition and positioning method based on DAS and AI, and the method comprises the steps: synchronously collecting sound and vibration data, and constructing a sound recognition branch and a vibration positioning branch; the sound branches extract multi-scale joint features through time-frequency domain feature fusion, weights of CNN and BiLSTM are dynamically adjusted to realize adaptive fusion, and event types and confidence coefficients are output; and the vibration branch calculates an initial coordinate by using a four-point space-time weighted optimization method, performs position fine tuning in combination with an ASTCN network, and outputs a final event position. Performing probability distribution verification by fusing double branch results and adopting a dynamic likelihood ratio evaluation mechanism, and outputting a final judgment result; intelligent event classification is achieved through the voice recognition branch, high-precision event space-time positioning is achieved through the vibration positioning branch, intelligent cooperation and system optimization of multi-modal data are achieved through the fusion and verification module, and the bottlenecks of the traditional technology in the aspects of recognition precision, positioning errors and system robustness are effectively broken through.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of infrastructure monitoring, and in particular to a dual-branch acoustic-vibration fusion event recognition and positioning method based on DAS and AI. Background Art

[0002] Distributed acoustic sensing (DAS) technology has been widely used in infrastructure monitoring in recent years, demonstrating significant advantages in scenarios such as cable tunnels, oil pipelines, railways, and border security. This technology uses optical fibers as distributed sensors, detecting phase changes in optical signals along the line to capture sound and vibration signals in real time, enabling the detection of abnormal events such as cable faults, mechanical impacts, or human intrusion. Compared to traditional point sensors, DAS offers wide coverage (up to tens of kilometers), high sensitivity, and the absence of a power supply, making it a key technology in intelligent monitoring.

[0003] However, despite the great potential of DAS technology in event monitoring, it still has many defects in practical applications, which limit its further performance improvement. Specifically, the current status and defects of similar technology applications are mainly reflected in the following aspects: insufficient event recognition accuracy, especially in high-noise environments, it is difficult to effectively separate event signals from background noise, resulting in high false alarm and missed alarm rates; poor positioning accuracy and stability, affected by signal attenuation along the optical fiber, multipath effects and environmental interference, especially in long-distance monitoring, the error is more obvious; insufficient environmental adaptability and robustness, poor adaptability to dynamic environmental changes such as temperature fluctuations, optical fiber aging and external electromagnetic interference, and system performance degrades over time; limited intelligence, relying mostly on manually set thresholds or rules, and lacking adaptive learning capabilities; low real-time performance and computational efficiency. The existing deep learning models have high computational complexity, making it difficult to achieve real-time processing on edge devices, and the response speed cannot meet the needs of rapid monitoring. The existence of these problems has greatly restricted the application effect of DAS technology in complex scenarios.

[0004] Specifically, in the context of distributed optical sensing (DAS) technology being applied to infrastructure monitoring, existing technologies face a series of prominent problems in event identification and location, as follows: 1. Inadequate event recognition accuracy: Existing technologies struggle to accurately distinguish event signals from background noise in complex noisy environments, such as traffic or wind noise in urban cable tunnels. This is particularly true when multiple events overlap or transient events occur, significantly increasing the false alarm and missed alarm rates. Traditional signal processing methods have limited processing capabilities for high-dimensional, heterogeneous DAS data and struggle to capture the temporal dynamic characteristics of sound and vibration. Traditional machine learning methods rely on manual feature extraction and have limited adaptability, resulting in poor recognition performance.

[0005] 2. Poor positioning accuracy and stability: DAS systems typically calculate event locations based on signal arrival time difference or phase difference. However, due to signal attenuation along the fiber, multipath effects, and environmental interference, positioning errors can be significant, especially over long distances. Existing positioning algorithms often rely on static assumptions and lack the ability to dynamically track the spatiotemporal evolution of events. For example, positioning results for mobile intrusions or persistent faults are prone to drift or ambiguity.

[0006] 3. Inadequate environmental adaptability and robustness: Existing technologies are poorly adaptable to dynamic environmental changes. Factors such as temperature fluctuations, fiber aging, and external electromagnetic interference can cause system performance to degrade over time. Furthermore, reliance on single-modal data processing (e.g., using only sound or vibration signals) is susceptible to interference in complex scenarios. For example, strong winds can mask sound signals, while ground vibrations can be confused by other mechanical activity, making it difficult to meet diverse monitoring needs.

[0007] 4. Limited Intelligence: Current DAS systems often rely on manually set thresholds or rules, lacking adaptive learning capabilities and unable to dynamically optimize model parameters based on real-time data. While deep learning methods have been tried in DAS applications, their application in diverse scenarios is significantly hampered by insufficient training data and a single model structure.

[0008] Low real-time performance and computational efficiency: Existing deep learning models have high computational complexity, making real-time processing difficult on edge devices. This limits the deployment of DAS systems in low-power scenarios. Furthermore, existing methods are slow when processing large-scale DAS data, making it difficult to meet the practical needs of real-time monitoring and rapid response. Summary of the Invention

[0009] The technical problem to be solved by the present invention is to provide a dual-branch acoustic vibration fusion event recognition and positioning method based on DAS and AI to overcome the deficiencies in the above-mentioned prior art.

[0010] The present invention solves the above technical problems with the following technical solution: a dual-branch vibroacoustic fusion event recognition and positioning method based on DAS and AI, comprising the following steps: Step S01: Collect sound data from the DAS system; capture multi-scale features including time domain features and frequency domain features from the sound data, perform weighted and adaptive fusion on the time domain features and frequency domain features to obtain joint features ; Step S02: From the joint features Extracting high-level spatial features , perform time series modeling on the feature sequence output by CNN to generate time series features , dynamically adjust the fusion weights of CNN and Bi-LSTM to obtain fusion features , the fusion features Input the fully connected layer, through The function outputs the event type and its confidence, which is recorded as the output result of the sound recognition branch ; Step S03: Use the DAS system to collect vibration intensity data; pre-process and standardize the vibration intensity data to obtain standardized vibration intensity data ; Based on the four-point spatiotemporal weighted optimization positioning method, calculate the initial positioning coordinates of the target point ; Initial positioning coordinates Make fine adjustments and output the final event position , and get the vibration positioning branch output result ; Step S04: Output the result of the voice recognition branch And the vibration positioning branch output results Perform fusion verification to obtain fusion probability distribution , consistency assessment is performed through dynamic likelihood ratio, and the likelihood ratio is calculated , combining the probability distributions of the two branches to generate the final fusion result ; Setting judgment thresholds , and the likelihood ratio and judgment threshold Compare and determine the final .

[0011] The beneficial effects of the present invention are as follows: The present invention aims to overcome the bottlenecks of traditional technologies in complex environments, such as insufficient event recognition accuracy, large spatiotemporal positioning errors, and poor system robustness. Intelligent event classification is achieved through the "sound recognition branch," high-precision event spatiotemporal positioning is achieved through the "vibration positioning branch," and intelligent collaboration and system optimization of multimodal data are achieved through the "fusion and verification" module. This effectively overcomes the bottlenecks of traditional technologies in recognition accuracy, positioning errors, and system robustness, providing an efficient, accurate, and stable monitoring solution for cable tunnels and other critical infrastructure.

[0012] On the basis of the above technical solution, the present invention can also be improved as follows.

[0013] Furthermore, step S01 specifically includes the following steps: Step S11: extracting multi-scale features simultaneously in the time domain and frequency domain; Step S111: Use sliding windows of different lengths or CNN to extract features of different time spans; the output is expressed as ,in is the scale index.

[0014] Step S112: Extract features of different frequency resolutions through wavelet transform or STFT; the output is expressed as ; Step S12: Calculate the attention weights in the time domain and frequency domain respectively through the dual-branch attention module, and weight the features: Step S121: Time domain features of each scale Calculate attention weights: ; in, is a learnable weight matrix; weighted features: ; in, Represents element-wise multiplication; Step S122: Frequency domain features of each scale Calculate attention weights: ; Weighted features: ; in, Represents element-wise multiplication; Step S13: Adaptively adjust the scale weight according to the input signal characteristics, fuse the multi-scale weighted features, and achieve time-frequency joint optimization: Step S131: weighted sum of the weighted features in the time domain and frequency domain respectively: ; ; in, and is a learnable scale weight that represents the contribution of each scale; Step S132: Integrate the time domain and frequency domain features through the joint attention module to obtain the joint feature : .

[0015] Furthermore, step S02 specifically includes the following steps: Step S21: CNN from joint features Extracting high-level spatial features , achieved through multi-layer convolution operations, the calculation formula is as follows: ; ; The parameters in the formula are described as follows: No. The input features of the layer (for the first layer, ); No. The convolution kernel of the layer is usually a four-dimensional tensor ,in are the height and width of the convolution kernel, is the number of input channels, is the number of output channels; Convolution operation, which calculates the sliding window convolution of the input features and the convolution kernel; No. The bias term of the layer has the shape ; Activation function; No. The output feature map of the layer has the shape ,in Determined by the input size, convolution kernel size and stride; Step S22: Bi-LSTM performs time series modeling on the feature sequence output by CNN to generate time series features ; The calculation process is as follows: ; ; ; ; ; ; ; in, and are the forward and backward hidden states respectively; Represents element-wise multiplication; Step S23: Dynamically adjust the fusion weights of CNN and Bi-LSTM according to the spatiotemporal characteristics of the input signal to improve the adaptability of feature fusion; the calculation formula is as follows: ; Among them, the adaptive weight Calculated by the following formula: ; in, and is a learnable parameter; Step S24: Extract features from the CNN middle layer and interact with the Bi-LSTM hidden state to obtain , the calculation formula is as follows: ; in, CNN Layer features, For Bi-LSTM in time The hidden state of Step S25: or Input the fully connected layer of CNN, through The function outputs the event type and its confidence, which is recorded as the output result of the sound recognition branch .

[0016] Furthermore, step S131 further includes: Use a lightweight neural network to predict the best scale combination based on the global characteristics of the input signal: ; in, are the original input features, is a global pooling operation, is the scale selection probability; use Update scale weights and , to achieve dynamic optimization.

[0017] Furthermore, step S03 specifically includes the following steps: Step S31: Preprocess and standardize the vibration intensity data to obtain standardized vibration intensity data ; Step S32: Using the four-point spatiotemporal weighted optimization positioning method, the spatiotemporal information of four adjacent measurement points is used to calculate the position coordinates of the target point through weighted optimization. ; Step S321: Obtain data from four adjacent test points, marked as , the position coordinates of each test point are represent; Step S322: define signal parameters, is the initial propagation speed of the signal; The signal reaches The time of each test point; Step S323: Arrival time of signals from four test points Find the earliest arrival time , the formula is: ; Step S324: For each test point Calculate weighting factors , the calculation formula is as follows: ; Where, For the Signal strength or reliability factor at each test point; is the attenuation coefficient; The signal reaches The time of each test point; is the earliest arrival time of the signals at the four test points; Step S325: Using weighting factors and test point locations , calculate the initial positioning coordinates of the target point , the calculation formula is as follows: ; Where, For the The position coordinates of the test points; is the weighting factor; is the sum of the weighted positions of all test points; is the sum of all weights, used for normalization; is the initial positioning coordinate of the target point; Step S33: normalize the signal Perform wavelet transform and decompose it into multi-scale time-frequency features. The calculation formula is as follows: ; in, is the wavelet basis function, and are scale and displacement parameters, respectively; Step S34: Initial positioning coordinates by ASTCN Fine-tune the calculation formula: ; Where, is the initial position calculated by the four-point spatiotemporal weighted optimization method; is the final event position; To calibrate the offset, the dynamic calibration function The calculation formula is as follows: ; Where, Feature vectors generated for multi-scale spatiotemporal correction; For real-time monitoring of ambient noise levels.

[0018] Furthermore, the specific steps of preprocessing and standardization in step S03 include: The original vibration signal is de-noised using wavelet threshold denoising technology. to process; The vibration intensity is standardized and the calculation formula is as follows: ; Where, is the mean value of the vibration intensity, is the standard deviation, is the normalized vibration intensity data.

[0019] Furthermore, step S04 specifically includes the following steps: Step S41: Map the event type probability of the sound recognition branch and the spatiotemporal features of the vibration localization branch to a unified event hypothesis space to form a joint probability distribution, as shown in the following formula: ; Where: Output the results for the sound recognition branch; Output results for the vibration positioning branch; is the prior probability of the event; is the normalization factor to ensure the validity of the probability distribution; Step S42: Introduce the dynamic likelihood ratio to quantify the consistency of the sound and vibration branch outputs and identify potential conflicts. The formula is: ; Where, Output consistent hypotheses for both branches; The output of the two branches is inconsistent with the hypothesis; is the likelihood ratio, and a larger value indicates a higher consistency; Step S43: Combine the probability distributions of the two branches to generate the final fusion result , the formula is: ; Where: is the branch weight, which is dynamically adjusted according to the data quality; is the normalization function, ensuring ; Step S44: Dynamically adjust the judgment threshold according to the likelihood ratio , the formula is as follows: ; Where, is the basic threshold; is the adjustment coefficient; is the standard deviation of ambient noise; The likelihood ratio and judgment threshold For comparison: when , the system accepts the current As the judgment result is final, no update is required; when , the system triggers the secondary analysis mechanism, re-evaluates the data, and further updates .

[0020] Furthermore, step S41 further includes: an adaptive Bayesian fusion network, which dynamically optimizes the fusion process through an environmental adaptive factor, and the formula is as follows: ; Where, is the environmental adaptive factor, defined as: ; Where, is the environmental feature vector; are parameters optimized through machine learning. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 This is a diagram of the time-frequency joint modeling architecture of the present invention; Figure 2 is an execution diagram of the DSTCO of the present invention; Figure 3 This is a flow chart of the four-point spatiotemporal weighted optimization positioning method of the present invention; Figure 4 This is a technical module diagram of the multi-scale spatiotemporal correction of the present invention; Figure 5 Schematic diagram of the Adaptive Spatiotemporal Calibration Network (ASTCN) of the present invention; Figure 6 This is a flowchart of the intelligent collaboration of the fusion and verification modules of the present invention. DETAILED DESCRIPTION

[0022] The principles and features of the present invention are described below with reference to the accompanying drawings. The examples given are only used to explain the present invention and are not used to limit the scope of the present invention.

[0023] like Figures 1 to 6 As shown in Example 1, a dual-branch vibroacoustic fusion event recognition and positioning method based on DAS and AI includes the following steps: Step S01: Collect sound data from the DAS system (such as cable discharge sound, mechanical impact sound, and human activity sound, stored in WAV format); sound signals have complex multi-scale characteristics in the time and frequency dimensions, such as short-term high-frequency transients (such as discharge sound) and long-term low-frequency continuous signals (such as mechanical vibration). Traditional attention mechanisms usually operate on a single scale or a single dimension (time domain or frequency domain), making it difficult to fully capture these characteristics, resulting in limited recognition performance in complex scenarios; capture multi-scale features including time domain features and frequency domain features from the sound data, weight and adaptively fuse the time domain features and frequency domain features to obtain joint features To address the multi-scale nature of sound signals, MSTFAA introduces a dynamic scale selection mechanism. Using a lightweight neural network, it predicts the optimal scale combination and adaptively adjusts the weights of time and frequency domain features. Compared to traditional fixed-scale methods, this mechanism dynamically optimizes feature fusion based on the input signal characteristics, improving the model's generalization and adaptability in diverse scenarios. This intelligent design demonstrates remarkable flexibility in handling complex sound events and demonstrates significant originality.

[0024] Step S02: CNN (Convolutional Neural Network) module is used to extract the joint features Extracting high-level spatial features , through the Bi-LSTM (bidirectional long short-term memory network) module, the feature sequence output by CNN is modeled to generate time series features According to the spatiotemporal characteristics of the input signal, the fusion weights of CNN and Bi-LSTM are dynamically adjusted to obtain the fusion features. , the fusion features Input the fully connected layer, through The function outputs the event type and its confidence, which is recorded as the output result of the sound recognition branch Dynamic Spatiotemporal Co-Optimization (DSTCO) is an adaptive feature fusion technology that dynamically adjusts the fusion weights of the convolutional neural network (CNN) and the bidirectional long short-term memory network (Bi-LSTM) outputs based on the spatiotemporal characteristics of the sound signal. Traditional methods usually use a fixed fusion strategy, while DSTCO introduces adaptive weights. , dynamically optimizes the feature fusion process according to the characteristics of the input signal, thereby significantly improving the model's adaptability and recognition accuracy to different sound events.

[0025] Step S03: Use the DAS system to collect vibration intensity data (stored in VIB format) and achieve high-precision spatiotemporal positioning of the event location through advanced signal processing and optimization algorithms. This branch effectively overcomes the challenges of optical fiber signal attenuation and environmental interference; pre-process and standardize the vibration intensity data to obtain standardized vibration intensity data. ; Using the four-point time-space weighted optimization positioning method, the time-space information (position and signal arrival time) of four adjacent measurement points is used to calculate the initial positioning coordinates of the target point through weighted optimization ; Through the "distance-time double weighting" mechanism , taking into account the time delay and signal characteristics , which is more comprehensive and has improved robustness compared to traditional distance-based methods. The weight design cleverly quantifies the impact of time delay as nonlinear attenuation, giving priority to early signals and improving noise resistance. The introduction of makes the method dynamically adjustable according to the signal quality, and the weighted average formula maintains low computational complexity, making it suitable for real-time applications. Using four test points instead of the traditional three increases redundant information and effectively resists single measurement errors; the initial positioning coordinates are adjusted by the adaptive space-time calibration network (ASTCN). Make fine adjustments and output the final event position , and get the vibration positioning branch output result . Output results of vibration positioning branch This includes vibration intensity, location, and time. It breaks through the limitations of traditional fixed calibration methods and achieves dynamic adaptive calibration through the combined input of multi-scale features and ambient noise levels. The introduction of intelligent optimization mechanisms enables the system to maintain high-precision positioning even in scenarios with poor signal quality or strong interference.

[0026] In order to ensure the high consistency of the output results of the two branches of sound recognition and vibration positioning, the present invention designs a fusion and verification module to achieve deep collaboration of multimodal data through probability fusion, conditional probability calculation and weight adaptive adjustment.

[0027] Step S04: Output the result of the voice recognition branch And the vibration positioning branch output results Input the Bayesian network fusion mechanism for fusion verification to obtain the fusion probability distribution , consistency evaluation is performed through dynamic likelihood ratio (Likelihood Ratio, LR), and the likelihood ratio is calculated , combining the probability distributions of the two branches to generate the final fusion result ; The result of event recognition is a probability value that mainly reflects the confidence of the event category and does not directly include the positioning result. is the positioning result, which indirectly affects Calculation of Setting judgment thresholds , and the likelihood ratio and judgment threshold For comparison: when , the system accepts the current As the judgment result is final, no update is required; when , the system triggers the secondary analysis mechanism, re-evaluates the data, and further updates .

[0028] This invention aims to overcome the bottlenecks of traditional technologies in complex environments, including insufficient event recognition accuracy, large spatiotemporal positioning errors, and poor system robustness. By implementing intelligent event classification through a "sound recognition branch" and high-precision spatiotemporal event positioning through a "vibration positioning branch," and intelligent multimodal data collaboration and system optimization through a "fusion and verification" module, this technology effectively overcomes the bottlenecks of traditional technologies in recognition accuracy, positioning errors, and system robustness, providing an efficient, accurate, and stable monitoring solution for cable tunnels and other critical infrastructure.

[0029] Example 2: This example is a further improvement on Example 1, and its details are as follows: Step S01 specifically includes the following steps: Step S11: Extract multi-scale features simultaneously in the time domain and frequency domain to cover the diverse characteristics of the sound signal. Extract multi-scale features simultaneously in the time domain and frequency domain to cover the diverse characteristics of the sound signal. In order to capture the multi-scale characteristics of the sound signal, first extract features in the time domain and frequency domain separately: Step S111: Multi-scale feature extraction: Multi-scale features in the time domain are used to extract features of different time spans using sliding windows of different lengths or CNN (multi-layer convolutional neural network); for example, short windows capture transient events and long windows analyze continuous signals; the output is represented as ,in is the scale index; Step S112: Frequency domain multi-scale features: Through wavelet transform or STFT (multi-resolution short-time Fourier transform), features with different frequency resolutions are extracted; for example, low resolution captures low-frequency trends, and high resolution focuses on high-frequency details; the output is represented as ; These features provide rich input for subsequent attention allocation; Step S12: To achieve collaborative optimization in the time domain and frequency domain, a dual-branch attention module is used to calculate attention weights in the time domain and frequency domain, and weight the features: Step S121: Temporal attention: Temporal features at each scale Calculate attention weights: ; in, is a learnable weight matrix; weighted features: ; in, Represents element-wise multiplication; Step S122: Frequency domain attention: frequency domain features of each scale Calculate attention weights: ; Weighted features: ; in, represents element-wise multiplication; in this way, the model can focus on key areas in the time and frequency domains.

[0030] MSTFAA's unique dual-branch time-frequency attention module calculates attention weights in both the time and frequency domains, achieving coordinated optimization of spatiotemporal information. Time-domain attention focuses on multi-scale dynamic changes, while frequency-domain attention focuses on multi-resolution frequency patterns. This dual-branch design enables the model to simultaneously capture transient events and high-frequency details, improving recognition accuracy and robustness in complex scenarios, demonstrating the technology's forward-looking and original nature.

[0031] Step S13: Adaptively adjust the scale weight according to the input signal characteristics to improve the generalization ability of the model; fuse the multi-scale weighted features and achieve time-frequency joint optimization: Step S131: Scale fusion: weighted sum of the weighted features in the time domain and frequency domain respectively: ; ; in, and is a learnable scale weight that represents the contribution of each scale; Step S132: Time-frequency joint features: The time domain and frequency domain features are integrated through the joint attention module to obtain joint features : ; Here It can be a simple weighted sum or a more complex multi-head self-attention mechanism.

[0032] Using this mechanism, we get It can be used as the first convolutional layer in the "CNN-BiLSTM sound event recognition fusion architecture" The input features of .

[0033] To address the multi-scale nature of sound signals, MSTFAA introduces a dynamic scale selection mechanism. Using a lightweight neural network, it predicts the optimal scale combination and adaptively adjusts the weights of time and frequency domain features. Compared to traditional fixed-scale methods, this mechanism dynamically optimizes feature fusion based on the input signal characteristics, improving the model's generalization and adaptability in diverse scenarios. This intelligent design demonstrates remarkable flexibility in handling complex sound events and demonstrates significant originality.

[0034] Example 3: This example is a further improvement on Example 2, and its details are as follows: Step S02 specifically includes the following steps: Step S21: CNN from joint features Extracting high-level spatial features , achieved through multi-layer convolution operations, the calculation formula is as follows: ; ; The parameters in the formula are described as follows: : No. The input features of the layer (for the first layer, ); : No. The convolution kernel (weight matrix) of the layer, usually a four-dimensional tensor ,in are the height and width of the convolution kernel, is the number of input channels, is the number of output channels; Convolution operation, which calculates the sliding window convolution of the input features and the convolution kernel; No. The bias term of the layer has the shape ; The activation function, here the ReLU function, is defined as , used to introduce nonlinearity; No. The output feature map of the layer has the shape ,in Determined by the input size, convolution kernel size and stride; Step S22: Bi-LSTM performs time series modeling on the feature sequence output by CNN to generate time series features ; The calculation process is as follows:

[0035]

[0036] in, and They are the forward and backward hidden states respectively, achieving bidirectional capture of temporal dependencies; Represents element-wise multiplication; Step S23: Dynamic spatiotemporal collaborative optimization mechanism (DSTCO): Dynamically adjust the fusion weights of CNN and Bi-LSTM according to the spatiotemporal characteristics of the input signal to improve the adaptability of feature fusion; the calculation formula is as follows:

[0037] in, and To learn parameters, ensure the intelligence and efficiency of the fusion process; Figure 2 The implementation process of DSTCO is dynamic spatiotemporal collaborative optimization (DSTCO), which is an adaptive feature fusion technology that dynamically adjusts the fusion weights of the convolutional neural network (CNN) and the bidirectional long short-term memory network (Bi-LSTM) output according to the spatiotemporal characteristics of the sound signal. Traditional methods usually adopt a fixed fusion strategy, while DSTCO introduces adaptive weights. , dynamically optimizes the feature fusion process according to the characteristics of the input signal, thereby significantly improving the model's adaptability and recognition accuracy to different sound events.

[0038] Step S24: Extract features from the CNN middle layer and interact with the Bi-LSTM hidden state to enhance the integration effect of spatiotemporal information. , the calculation formula is as follows:

[0039] in, CNN Layer features, For Bi-LSTM in time The hidden state of Cross-layer feature interaction (CLFI) is a novel feature integration mechanism that directly interacts the spatial features of CNN intermediate layers with the temporal hidden states of Bi-LSTM to achieve multi-level spatiotemporal information fusion. Traditional models typically perform feature fusion at a single level, but CLFI's cross-layer design enhances the model's ability to capture complex spatiotemporal patterns in sound signals. Step S25: or Input the fully connected layer of CNN, through The function outputs the event type and its confidence, which is recorded as the output result of the sound recognition branch .

[0040] Example 4: This example is a further improvement based on either Example 2 or 3, and its details are as follows: Step S131 also includes: Use the scale selection module to make the model adaptive to different input signals: Scale selection prediction: Use a lightweight neural network to predict the best scale combination based on the global characteristics of the input signal:

[0041] in, are the original input features, is a global pooling operation, is the scale selection probability; Adaptive adjustment: Use Update scale weights and , achieving dynamic optimization. This mechanism ensures that the model can select the most appropriate feature scale in different scenarios.

[0042] Example 5: This example is a further improvement on Example 1, and its details are as follows: The vibration location branch uses vibration intensity data collected by the DAS system (stored in VIB format) and, through advanced signal processing and optimization algorithms, achieves high-precision spatiotemporal location of the event location. This branch effectively overcomes the challenges of optical fiber signal attenuation and environmental interference. Step S03 specifically includes the following steps: Step S31: Preprocess and standardize the vibration intensity data to obtain standardized vibration intensity data ; This technology proposes an innovative "four-point spatiotemporal weighted optimization positioning method" that utilizes the spatiotemporal information (position and signal arrival time) of four adjacent measurement points to calculate the target point's position coordinates through weighted optimization. Its core approach is to introduce composite weights to comprehensively consider the impact of signal arrival time delay, thereby reducing the inaccuracies caused by measurement errors in traditional positioning methods. By weighting the vibration intensity and time delay of four adjacent measurement points, high-precision positioning of the event is achieved; see [1]. Figure 3 Flowchart of the four-point spatiotemporal weighted optimization positioning method; Step S32: using a four-point spatiotemporal weighted optimization positioning method to calculate the position coordinates of the target point through weighted optimization; Step S321: Obtain data from four adjacent test points, marked as , the position coordinates of each test point are represent; Step S322: define signal parameters, is the initial propagation speed of the signal; The signal reaches This step collects the raw data required for positioning, the test point location For subsequent calculation of target coordinates, and Used to analyze the spatiotemporal characteristics of signal propagation; Step S323: Arrival time of signals from four test points Find the earliest arrival time , the formula is:

[0043] Yes all The minimum value among is used as the benchmark for time delay calculation. It reflects the test point where the signal arrives fastest, which usually means that the measurement at this point is least affected by interference and has higher reliability.

[0044] Step S324: For each test point Calculate weighting factors , the calculation formula is as follows:

[0045] Where, For the Signal strength or reliability factor at each test point; is the attenuation coefficient, a constant used to control the decay speed of the weight due to time delay; The signal reaches The time of each test point; is the earliest arrival time of the signals at the four test points; Indicates the The delay of each test point relative to the earliest arrival time, the greater the delay, The smaller, The lower the weight. It is an exponential decay function, ensuring that the test points with earlier signal arrival times receive higher weights; Introducing additional flexibility, allowing weights to be adjusted based on signal quality (e.g., strength). Dynamically reflects the measurement reliability of each test point, with points with smaller time delays contributing more; Step S325: Using weighting factors and test point locations , calculate the initial positioning coordinates of the target point , the calculation formula is as follows:

[0046] Where, For the The position coordinates of the test points; is the weighting factor; is the sum of the weighted positions of all test points; is the sum of all weights, used for normalization; is the initial positioning coordinate of the target point; This formula is a mathematical expression of weighted average, ensuring that test points with high weights contribute more to the final position; the normalization factor Make the results independent of the absolute value of the weight and only reflect the relative contribution; By weighted averaging and integrating the information of the four test points, the optimal estimated position of the target point is obtained.

[0047] Through the "distance-time double weighting" mechanism , taking into account the time delay and signal characteristics , which is more comprehensive and has improved robustness compared to traditional distance-based methods. The weight design cleverly quantifies the impact of time delay as nonlinear attenuation, giving priority to early signals and improving noise resistance. The introduction of [ ] enables the method to dynamically adjust based on signal quality, while the weighted average formula maintains low computational complexity, making it suitable for real-time applications. Using four test points instead of the traditional three increases redundant information, effectively preventing single measurement errors.

[0048] To further optimize positioning results, this technology introduces a "multi-scale spatiotemporal correction" mechanism to improve signal quality through time-frequency analysis and interference suppression.

[0049] Step S33: Wavelet transform time-frequency analysis: normalized signal Perform wavelet transform and decompose it into multi-scale time-frequency features. The calculation formula is as follows:

[0050] in, is the wavelet basis function, and are scale and displacement parameters, respectively. Through this step, the system can capture the comprehensive characteristics of the vibration signal from low-frequency trends to high-frequency transients.

[0051] This two-dimensional function reflects the energy distribution of a signal in the scale-time domain. It effectively reflects the dynamic characteristics of a signal at multiple scales. Based on the main energy position, frequency band energy distribution, and statistical features extracted by this function, a multi-scale feature vector can be constructed. This serves as the key input to the ASTCN, effectively correcting spatiotemporal features and improving positioning accuracy.

[0052] Beamforming technology is based on a directional enhancement algorithm to suppress multipath effects and environmental noise interference, further optimize feature representation, and ensure the stability of positioning results.

[0053] See also Figure 4 , showing the technical module diagram of multi-scale spatiotemporal correction; This solution overcomes the limitations of traditional Fourier transforms in terms of time-frequency locality, employing wavelet transforms to achieve multi-resolution analysis, more accurately capturing the dynamic changes in vibration signals. Combined with beamforming technology, it effectively suppresses interference and improves feature robustness, significantly outperforming traditional single-scale analysis methods.

[0054] Step S34: Initial positioning coordinates are adjusted by ASTCN (Adaptive Spatio-Temporal Calibration Network) Fine-tune the calculation formula:

[0055] Where, is the initial position calculated by the four-point spatiotemporal weighted optimization method; is the final event position; To calibrate the offset, the dynamic calibration function The calculation formula is as follows:

[0056] Where, The feature vector generated for multi-scale spatiotemporal correction is Structured representation after compression, aggregation, and dimensionality reduction in spatial (sensor distribution) and frequency (scale) dimensions; For real-time monitoring of ambient noise levels.

[0057] It is implemented through a lightweight neural network, and the positioning results are dynamically adjusted according to environmental parameters to ensure high robustness under signal attenuation or noise interference.

[0058] After ASTCN calibration, the system outputs the final event position , achieving sub-meter precision positioning.

[0059] Figure 5 This is the architecture diagram of the Adaptive Spatiotemporal Calibration Network (ASTCN); This solution overcomes the limitations of traditional fixed calibration methods by integrating multi-scale features and ambient noise levels into dynamic adaptive calibration. It also incorporates intelligent optimization mechanisms, enabling the system to maintain high-precision positioning even in scenarios with poor signal quality or strong interference.

[0060] Example 6: This example is a further improvement on Example 1, and its details are as follows: The specific steps of preprocessing and standardization in step S03 include: Data preprocessing is a fundamental step in high-precision positioning, aiming to extract effective information from raw vibration data and eliminate interference. The preprocessing process is as follows: Wavelet denoising: Using wavelet threshold denoising technology, the original vibration signal Processing is performed to effectively separate the target signal from the environmental noise; Data standardization: Standardize the vibration intensity to ensure data consistency and facilitate subsequent calculations. The calculation formula is as follows: ; Where, is the mean value of the vibration intensity, is the standard deviation, is the normalized vibration intensity data.

[0061] Example 7: This example is a further improvement on Example 1, and its details are as follows: Step S04 specifically includes the following steps: To ensure high consistency between the outputs of sound recognition and vibration localization, the present invention designs a fusion and verification module. This module achieves deep collaboration of multimodal data through probabilistic fusion, conditional probability calculation, and adaptive weight adjustment. This module, with an improved Bayesian network as its core, achieves collaborative optimization of sound and vibration data through probabilistic fusion. The specific steps are as follows: Step S41: Map the event type probability of the sound recognition branch (e.g., "cable fault: 95%) and the spatiotemporal characteristics of the vibration location branch (e.g., "abnormal vibration intensity at the 500th meter of the optical fiber") into a unified event hypothesis space to form a joint probability distribution. The formula is as follows: ; The results of the above two branches are used as input for fusion verification. Where: Output the results (event type and its confidence) for the sound recognition branch; Output results for the vibration positioning branch (vibration intensity, location, time, etc.); is the prior probability of the event; is the normalization factor to ensure the validity of the probability distribution; Step S42: Introduce the dynamic likelihood ratio (LR) to quantify the consistency of the sound and vibration branch outputs and identify potential conflicts. The formula is: ; Where, Output consistent hypotheses for both branches; The output of the two branches is inconsistent with the hypothesis; is the likelihood ratio, and a larger value indicates a higher consistency; Step S43: Combine the probability distributions of the two branches to generate the final fusion result , the formula is: ; Where: is the branch weight, which is dynamically adjusted according to the data quality; is the normalization function, ensuring ; The improved Bayesian network fusion mechanism achieves effective collaboration of multi-source data through event hypothesis space mapping and weighted probability optimization. The dynamic likelihood ratio assessment method quantifies the consistency between multi-source data by dynamically calculating likelihood ratios, breaking through the limitations of traditional static thresholds.

[0062] Step S44: Dynamically adjust the judgment threshold according to the likelihood ratio , the formula is as follows: ; Where, is the basic threshold; is the adjustment coefficient; is the standard deviation of ambient noise; The likelihood ratio and judgment threshold For comparison: when , indicating that the consistency of the fusion results is high enough, and the system accepts the current As the judgment result is final, no update is required; when , indicating insufficient consistency, the system triggers a secondary analysis mechanism to re-evaluate the data and further update .

[0063] This threshold is used to decide whether to accept the fusion result, which directly affects the system's decision sensitivity and accuracy. For example, in a noisy environment, the threshold may be appropriately increased to reduce false positives.

[0064] Principle of the secondary analysis mechanism: When the consistency is lower than the threshold, in-depth analysis is triggered to re-evaluate the data.

[0065] Implementation: Combine historical data and context features to update and output a final event probability .

[0066] Function: This output ensures that the system can still provide reliable judgment results under complex or uncertain conditions, thereby enhancing the robustness of the system.

[0067] Introducing environmental adaptive factors , achieving dynamic optimization of fusion strategies. The system can adjust parameters according to real-time environmental changes, significantly improving robustness and high-precision performance under changing conditions.

[0068] Example 8: This example is a further improvement on Example 7, and its details are as follows: Step S41 also includes: Adaptive Bayesian Fusion Network (ABFN), which dynamically optimizes the fusion process through the environmental adaptive factor, and the formula is as follows: ; Where, is the environmental adaptive factor, defined as: ; Where, is the environmental characteristic vector (including noise level, signal attenuation rate, etc.); are parameters optimized through machine learning.

[0069] Enables the system to adjust fusion strategies according to real-time environmental changes to ensure high robustness; original quantitative consistency technology improves system reliability.

[0070] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.

Claims

1. A dual-branch acoustic vibration fusion event recognition and positioning method based on DAS and AI, characterized by: The steps include: Step S01: Collect sound data from the DAS system; capture multi-scale features including time domain features and frequency domain features from the sound data, perform weighted and adaptive fusion on the time domain features and the frequency domain features to obtain joint features ; Step S02: From the joint features Extracting high-level spatial features Perform time series modeling on the feature sequence output by CNN to generate time series features Dynamically adjust the fusion weights of CNN and the Bi-LSTM to obtain fusion features , the fusion features Input the fully connected layer, through The function outputs the event type and its confidence, which is recorded as the output result of the sound recognition branch ; Step S03: Use the DAS system to collect vibration intensity data; pre-process and standardize the vibration intensity data to obtain vibration intensity data ; Based on the four-point spatiotemporal weighted optimization positioning method, calculate the initial positioning coordinates of the target point ; Initial positioning coordinates Make fine adjustments and output the final event position , and get the vibration positioning branch output result ; Step S04: Output the result of the voice recognition branch And the vibration positioning branch output results Perform fusion verification to obtain fusion probability distribution, perform consistency assessment through dynamic likelihood ratio, and calculate likelihood ratio , combining the probability distributions of the two branches to generate the final fusion result ; Setting judgment thresholds , and the likelihood ratio With the judgment threshold Compare and determine the final .

2. The dual-branch vibroacoustic fusion event recognition and positioning method based on DAS and AI according to claim 1 is characterized in that: The step S01 specifically includes the following steps: Step S11: extracting multi-scale features simultaneously in the time domain and the frequency domain; Step S111: Use sliding windows of different lengths or CNN to extract features of different time spans; the output is expressed as ,in is the scale index; Step S112: extracting features of different frequency resolutions through wavelet transform or STFT; The output is represented as ; Step S12: Calculate the attention weights in the time domain and frequency domain respectively through the dual-branch attention module, and weight the features: Step S121: Time domain features of each scale Calculate attention weights: ; in, Learnable weight matrix; weighted features: ; in, Represents element-wise multiplication; Step S122: Frequency domain features of each scale Calculate attention weights: ; Weighted features: ; in, Represents element-wise multiplication; Step S13: Adaptively adjust the scale weight according to the input signal characteristics, fuse the multi-scale weighted features, and achieve time-frequency joint optimization: Step S131: weighted sum of the weighted features in the time domain and frequency domain respectively: ; ; in, and is a learnable scale weight that represents the contribution of each scale; Step S132: Integrate the time domain and frequency domain features through the joint attention module to obtain the joint feature : 。 3. The method for identifying and locating dual-branch vibroacoustic fusion events based on DAS and AI according to claim 2, characterized in that: The step S02 specifically includes the following steps: Step S21: CNN from joint features Extracting high-level spatial features , achieved through multi-layer convolution operations, the calculation formula is as follows: ; ; The parameters in the formula are described as follows: : No. The input features of the layer (for the first layer, ); : No. The convolution kernel of the layer is usually a four-dimensional tensor ,in are the height and width of the convolution kernel, is the number of input channels, is the number of output channels; : Convolution operation, calculating the sliding window convolution of the input features and the convolution kernel; : No. The bias term of the layer has the shape ; σ: activation function; : No. The output feature map of the layer has the shape ,in Determined by the input size, convolution kernel size and stride; Step S22: Bi-LSTM performs time series modeling on the feature sequence output by CNN to generate time series features ; The calculation process is as follows: ; ; ; ; ; ; ; in, and are the forward and backward hidden states respectively; Represents element-wise multiplication; Step S23: Dynamically adjust the fusion weights of CNN and Bi-LSTM according to the spatiotemporal characteristics of the input signal to improve the adaptability of feature fusion; the calculation formula is as follows: ; Among them, the adaptive weight Calculated by the following formula: ; in, and is a learnable parameter; Step S24: Extract features from the CNN middle layer and interact with the Bi-LSTM hidden state to obtain , the calculation formula is as follows: ; in, CNN Layer features, For Bi-LSTM in time The hidden state of Step S25: or Input the fully connected layer of CNN, through The function outputs the event type and its confidence, which is recorded as the output result of the sound recognition branch .

4. A dual-branch vibroacoustic fusion event recognition and positioning method based on DAS and AI according to any one of claims 2 or 3, characterized in that: The step S131 further includes: Use a lightweight neural network to predict the best scale combination based on the global characteristics of the input signal: ; in, are the original input features, is a global pooling operation, is the scale selection probability; use Update scale weights and , to achieve dynamic optimization.

5. The method for identifying and locating dual-branch vibroacoustic fusion events based on DAS and AI according to claim 1, characterized in that: The step S03 specifically includes the following steps: Step S31: Preprocess and standardize the vibration intensity data to obtain standardized vibration intensity data. ; Step S32: Using the four-point spatiotemporal weighted optimization positioning method, the spatiotemporal information of four adjacent measurement points is used to calculate the position coordinates of the target point through weighted optimization. ; Step S321: Obtain data from four adjacent test points, marked as , The position coordinates of each test point are represent; Step S322: define signal parameters, is the initial propagation speed of the signal; The signal reaches The time of each test point; Step S323: Arrival time of signals from four test points Find the earliest arrival time , the formula is: ; Step S324: For each test point Calculate weighting factors , the calculation formula is as follows: ; Where, For the Signal strength or reliability factor at each test point; is the attenuation coefficient; The signal reaches The time of each test point; is the earliest arrival time of the signals at the four test points; Step S325: Using weighting factors and test point locations , calculate the initial positioning coordinates of the target point , the calculation formula is as follows: ; Where, For the The position coordinates of the test points; is the weighting factor; is the sum of the weighted positions of all test points; is the sum of all weights, used for normalization; is the initial positioning coordinate of the target point; Step S33: normalize the signal Perform wavelet transform and decompose it into multi-scale time-frequency features. The calculation formula is as follows: ; in, is the wavelet basis function, and are scale and displacement parameters, respectively; Step S34: Initial positioning coordinates by ASTCN Fine-tune the calculation formula: ; Where, is the initial position calculated by the four-point spatiotemporal weighted optimization method; is the final event position; To calibrate the offset, the dynamic calibration function The calculation formula is as follows: ; Where, Feature vectors generated for multi-scale spatiotemporal correction; For real-time monitoring of ambient noise levels.

6. The method for identifying and locating dual-branch vibroacoustic fusion events based on DAS and AI according to claim 1, characterized in that: The specific steps of preprocessing and standardization in step S03 include: The original vibration signal is de-noised using wavelet threshold denoising technology. to process; The vibration intensity is standardized and the calculation formula is as follows: ; Where, is the mean value of the vibration intensity, is the standard deviation, is the normalized vibration intensity data.

7. The method for identifying and locating dual-branch vibroacoustic fusion events based on DAS and AI according to claim 1, characterized in that: The step S04 specifically includes the following steps: Step S41: Map the event type probability of the sound recognition branch and the spatiotemporal features of the vibration localization branch to a unified event hypothesis space to form a joint probability distribution, as shown in the following formula: ; Where: Output the results for the sound recognition branch; Output results for the vibration positioning branch; is the prior probability of the event; is the normalization factor to ensure the validity of the probability distribution; Step S42: Introduce the dynamic likelihood ratio to quantify the consistency of the sound and vibration branch outputs and identify potential conflicts. The formula is: ; Where, Output consistent hypotheses for both branches; The output of the two branches is inconsistent with the hypothesis; is the likelihood ratio, and a larger value indicates a higher consistency; Step S43: Combine the probability distributions of the two branches to generate the final fusion result , the formula is: ; Where: is the branch weight, which is dynamically adjusted according to the data quality; is the normalization function, ensuring ; Step S44: Dynamically adjust the judgment threshold according to the likelihood ratio , the formula is as follows: ; Where, is the basic threshold; is the adjustment coefficient; is the standard deviation of ambient noise; The likelihood ratio With the judgment threshold For comparison: when , the system accepts the current As the judgment result is final, no update is required; when , the system triggers the secondary analysis mechanism, re-evaluates the data, and further updates .

8. The method for identifying and locating dual-branch vibroacoustic fusion events based on DAS and AI according to claim 7, characterized in that: The step S41 further includes: an adaptive Bayesian fusion network, which dynamically optimizes the fusion process through an environmental adaptive factor, and the formula is as follows: ; Where, is the environmental adaptive factor, defined as: ; Where, is the environmental feature vector; are parameters optimized through machine learning.

Citation Information

Patent Citations

  • Underwater target recognition method based on multi-feature fusion

    CN112183582A

  • Bridge expansion joint damage detecting and positioning method based on multi-modal signal fusion

    CN119885069A

  • Cable vibration event identification method and system based on deep neural network

    CN120030494A

  • Vibration space-time diagram event identification method based on distributed optical fiber sensing

    CN120123869A

  • Sound identification systems

    US20150106095A1

Cited By

  • Power cable fault sound recognition method and system based on multi-network fusion

    CN120847555A

  • Power cable fault sound recognition method and system based on multi-network fusion

    CN120847555B

  • Electric power well environment detection method, system and equipment based on image processing and medium

    CN121482461A

  • Vibration frequency extraction method based on event stream adaptive binning and multi-modal fusion

    CN122571062B