A vibration event classification method and system based on distributed fiber sensing
By using distributed fiber optic sensing technology, combined with wavelet transform, Fourier transform and deep convolutional networks, efficient and accurate monitoring of bridge vibration events has been achieved, solving the problems of low efficiency and poor stability in existing technologies and meeting the monitoring needs of bridges in high-risk environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- UNIV OF SCI & TECH BEIJING
- Filing Date
- 2025-09-26
- Publication Date
- 2026-07-31
AI Technical Summary
Existing manual inspections are inefficient and costly, making it difficult to detect sudden damage in a timely manner. Traditional sensors are susceptible to electromagnetic interference and environmental factors, resulting in insufficient stability and reliability of bridge monitoring data, which cannot meet the high-precision monitoring needs in high-risk environments.
Distributed fiber optic sensing technology is used to collect vibration signals, perform noise reduction processing, extract time-domain, frequency-domain, and audio-domain features, and combine wavelet transform, Fourier transform, and Gram angle difference field coding. Deep convolutional networks are used for multi-level feature learning to achieve multi-modal feature fusion and classification.
It improves the real-time and comprehensiveness of bridge monitoring, enhances the ability to identify damage and abnormal events, and improves the stability and robustness of monitoring results. It can accurately identify bridge vibration events and meet the needs of efficient and accurate monitoring in high-risk environments.
Smart Images

Figure CN121302000B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of structural health monitoring technology, and in particular to a vibration event classification method and system based on distributed optical fiber sensing. Background Technology
[0002] With the rapid development of global modern transportation networks, bridges, as a crucial component of transportation infrastructure, are directly related to the stable operation of the transportation system and the sustainable development of the social economy. In recent years, ship collisions with bridges have become a prominent safety threat in bridge operation, potentially causing partial or complete damage to the bridge structure, and even leading to collapse accidents, resulting in significant casualties and economic losses. Especially in the past two decades, hundreds of major ship-bridge collision accidents have occurred in my country, showing a trend of frequent accidents and severe consequences, urgently requiring more efficient and accurate monitoring and prevention technologies.
[0003] Currently, bridge safety monitoring mainly relies on two types of technologies: manual inspection and fixed sensor monitoring. Manual inspection relies on professionals to periodically check and evaluate the bridge structure; its advantage lies in its ability to make qualitative judgments based on experience. Fixed sensor monitoring, on the other hand, involves deploying sensors at key locations on the bridge to achieve real-time data acquisition and status monitoring of local areas. These methods, to a certain extent, improve the safety of bridges during operation and form the main technical support for the current bridge monitoring system.
[0004] However, existing manual inspections are inefficient, costly, and struggle to detect sudden damage in a timely manner, easily leaving safety hazards due to human oversight. Furthermore, traditional sensors are highly susceptible to external environmental factors such as electromagnetic interference and temperature and humidity changes, resulting in insufficient stability and reliability of monitoring data, making it difficult to meet the urgent need for high-precision, safe monitoring of bridges in high-risk environments. Summary of the Invention
[0005] To address the shortcomings of existing manual inspection methods, such as low efficiency, high cost, difficulty in timely detection of sudden damage, and susceptibility to safety hazards due to human error, as well as the vulnerability of traditional sensors to electromagnetic interference, temperature and humidity changes, leading to insufficient stability and reliability of monitoring data and failing to meet the urgent need for high-precision and safe monitoring of bridges in high-risk environments, this invention provides a vibration event classification method and system based on distributed optical fiber sensing.
[0006] The technical solutions provided by the embodiments of the present invention are as follows: First aspect: This invention provides a vibration event classification method based on distributed optical fiber sensing, comprising: S1: Acquire vibration signals; S2: Denoise the vibration signal to obtain a denoised vibration signal; S3: Extract time-domain features, frequency-domain features, and audio-domain features from the denoised vibration signal to form statistical features; S4: The denoised vibration signal is subjected to feature extraction by continuous wavelet transform and short-time Fourier transform respectively, and the extraction results are fused based on the entropy weighting mechanism to obtain a preliminary fused feature map; S5: Based on the Gram difference field coding principle, feature extraction is performed on the denoised vibration signal to obtain Gram difference field image features; S6: Perform a convolution operation on the preliminary fused feature map to extract local features, and then perform a weighted fusion of the local features and the Gram angle difference field image features to obtain fused image features; S7: Perform a depthwise separable convolution operation on the fused image features to extract image depth features; S8: Based on the image depth features, multi-scale dilated convolution is used to capture the long temporal dependence of vibration events to obtain the final sequence features; S9: Perform multimodal fusion on the statistical features, the deep image features, and the final sequence features to obtain the final fused features; S10: Based on the final fusion features, output the category classification results of vibration events through the fully connected layer.
[0007] The second aspect: This invention provides a vibration event classification system based on distributed optical fiber sensing, comprising: processor; A memory storing computer-readable instructions, which, when executed by the processor, implement the vibration event classification method based on distributed optical fiber sensing as described in the first aspect.
[0008] Third aspect: The present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the vibration event classification method based on distributed optical fiber sensing as described in the first aspect.
[0009] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following: In this embodiment of the invention, the real-time performance and comprehensiveness of bridge monitoring are improved through vibration signal acquisition, denoising, and multi-domain feature extraction. Combined with multi-level feature learning using wavelet transform, Fourier transform, Gram difference field coding, and deep convolutional networks, the ability to identify damage and abnormal events is enhanced, effectively overcoming the shortcomings of low efficiency, high cost, and susceptibility to oversights in manual inspections. Simultaneously, denoising processing, multi-modal feature fusion, and deep convolutional structures improve the stability and robustness of monitoring results, enhance resistance to electromagnetic interference and environmental changes, and ultimately achieve accurate event identification through classification output, meeting the urgent need for efficient, accurate, and intelligent bridge monitoring in high-risk environments. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 A flowchart illustrating a vibration event classification method based on distributed optical fiber sensing provided in an embodiment of the present invention; Figure 2 This is a schematic diagram illustrating the effect of a signal preprocessing noise reduction method provided in an embodiment of the present invention; Figure 3 This is a schematic diagram comparing depthwise separable convolution and ordinary convolution, provided as an embodiment of the present invention. Figure 4 This is a schematic diagram of a vibration event classification system based on distributed optical fiber sensing, provided as an embodiment of the present invention. Detailed Implementation
[0012] The technical solution of the present invention will now be described with reference to the accompanying drawings.
[0013] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.
[0014] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning.
[0015] In this embodiment of the invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.
[0016] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.
[0017] Reference manual attached Figure 1 The diagram shows a flowchart of a vibration event classification method based on distributed optical fiber sensing provided by an embodiment of the present invention.
[0018] This invention provides a vibration event classification method based on distributed optical fiber sensing. This method can be implemented by a vibration event classification device based on distributed optical fiber sensing, which can be a terminal or a server. The processing flow of the vibration event classification method based on distributed optical fiber sensing may include the following steps: S1: Acquire vibration signals.
[0019] Specifically, distributed fiber optic vibration sensing cables are laid in critical, vulnerable parts of the bridge, such as piers, decks, and towers, forming a high-density vibration monitoring network using phase-sensitive optical time-domain reflectometry (DSI). The distributed fiber optic vibration sensing system boasts a spatial resolution down to the meter level and a sampling frequency exceeding 10 kHz, enabling real-time capture of vibration waves propagating along the optical fiber. The fiber optic detection signals are analyzed and processed using a phase-sensitive optical time-domain reflectometer (Φ-OTDR) and a DAS system to obtain amplitude data for corresponding spatial locations and event sequences, constructing a spatiotemporal matrix of the vibration signal. This matrix represents the vibration signal at different locations along the fiber in the spatial dimension and the change of the vibration signal over time in the temporal dimension, containing both spatiotemporal information to aid in subsequent event type identification. When the bridge is impacted, the high-density vibration monitoring network captures the vibration signal.
[0020] Reference manual attached Figure 2 The diagram illustrates the effect of a signal preprocessing noise reduction method provided in an embodiment of the present invention.
[0021] S2: Denoise the vibration signal to obtain a denoised vibration signal.
[0022] It should be noted that, for situations where the original signal contains high noise and complex information under complex working conditions, this solution uses a wavelet-morphological joint denoising method to reduce the noise in the data.
[0023] In one possible implementation, S2 specifically includes: S201: Decompose the vibration signal using wavelet basis functions to extract low-frequency and high-frequency components:
[0024] in, Indicates low-frequency components. Represents high-frequency components. Represents discrete wavelet transform. x ( t ) represents a vibration signal. This represents the Symlet-5 wavelet basis function.
[0025] Wavelet basis function denoising is a commonly used signal processing method. It uses wavelet transform to decompose the original signal into low-frequency and high-frequency components of different scales. The low-frequency component mainly preserves the overall trend of the signal, while the high-frequency component contains noise and detailed information. By selecting appropriate wavelet basis functions (such as Symlet, Daubechies, etc.) and the number of decomposition levels, thresholding (soft thresholding or hard thresholding) is applied to the high-frequency components to suppress noise interference. Then, an inverse wavelet transform is performed together with the low-frequency components to reconstruct a smooth and near-real signal.
[0026] S202: Calculate the kurtosis value of the high-frequency components and normalize the kurtosis value.
[0027] It should be noted that, in order to achieve adaptive optimization of noise filtering, this scheme dynamically adjusts the threshold based on signal characteristics. The signal characteristics are measured using kurtosis, which characterizes the steepness of the signal distribution; higher kurtosis indicates a greater amount of impulse components. Simultaneously, to avoid thresholds being too large or too small, the kurtosis is normalized to enhance robustness.
[0028] S203: Calculate the adaptive threshold based on the normalized kurtosis value:
[0029] in, Indicates an adaptive threshold. Indicates the standard deviation of noise. N Indicates signal length. Indicates eLogarithmic function with base 0. Represents the hyperbolic tangent function. This refers to the high-frequency components.
[0030] S204: Thresholding is performed on the high-frequency components to retain the high-frequency elements, resulting in the noise-reduced high-frequency components:
[0031] in, Indicates the high-frequency components of noise reduction. The function represents the sign function, and max() represents the maximum value function. This indicates taking the absolute value.
[0032] S205: Perform inverse discrete wavelet transform on the low-frequency component and the noise-reducing high-frequency component to obtain the reconstructed signal, and then perform morphological filtering on the reconstructed signal to obtain the denoised vibration signal.
[0033] in, This represents the denoised vibration signal. This represents the closing operation. This indicates the opening operation. This represents the inverse discrete wavelet transform.
[0034] Specifically, morphological filtering involves first performing an opening operation (erosion + dilation) to remove small-scale noise such as spikes. Then, a closing operation (dilation + erosion) is performed to fill holes and eliminate impulse noise.
[0035] In this embodiment of the invention, wavelet basis function decomposition is used to divide the signal into low-frequency and high-frequency components. A kurtosis-driven adaptive threshold is then used to threshold the high-frequency components, effectively suppressing noise interference and avoiding the over-filtering or noise leakage problems inherent in fixed-threshold methods. Furthermore, by reconstructing the signal through inverse wavelet transform and applying morphological opening and closing operations in the time domain, both spikes and isolated noise points can be removed simultaneously, maintaining the continuity and integrity of the signal.
[0036] Compared with existing single wavelet denoising or traditional filtering methods, this invention significantly improves the signal-to-noise ratio and robustness while ensuring that key information (such as transient impacts and low-frequency structural responses) is not distorted. This provides more reliable input data for subsequent feature extraction and classification, thereby enhancing the system's adaptability and overall monitoring accuracy in complex environments.
[0037] S3: Extract time-domain features, frequency-domain features, and audio-domain features from the denoised vibration signal to form statistical features.
[0038] In one possible implementation, S3 specifically includes: S301: Extract time-domain features from the denoised vibration signal, including mean, variance, kurtosis, mean absolute value, and zero-crossing rate.
[0039] S302: Extract the frequency domain features of the denoised vibration signal, including the energy proportion of 0-50Hz, peak frequency, and energy second derivative at characteristic frequencies.
[0040] S303: Extract the audio domain features of the denoised vibration signal, including Mel frequency cepstral coefficients, spectral center, and spectral change rate.
[0041] S304: Combine time-domain features, frequency-domain features, and audio-domain features to form statistical features:
[0042] in, Indicates statistical characteristics, This represents the mean. Represents variance. Indicates kurtosis, Represents the absolute value of the average. Indicates the zero-crossing rate. This indicates the energy percentage in the 0-50Hz range. Indicates peak frequency. This represents the second derivative of the energy at the characteristic frequency. This represents the cepstral coefficients of the first 5 Mel frequencies. Indicates the spectral center, This represents the rate of change of the spectrum.
[0043] Specifically, since the selection of features directly affects the classification effect of vibration signals, artificial features of vibration signals are extracted from the denoised vibration signals to effectively distinguish vibration signals corresponding to different disturbance events. These artificial features include various time-domain features, frequency-domain features, and audio-domain features of the vibration signals. Specifically: time-domain features include mean, variance, kurtosis, mean absolute value, and zero-crossing rate; frequency-domain features include energy proportion of 0-50Hz, peak frequency, and second derivative of energy at characteristic frequencies; and audio-domain features include cepstral coefficients of the first five Mel frequencies, spectral center, and spectral change rate, extracting 11 statistical features.
[0044] In this embodiment of the invention, by extracting multidimensional statistical features in the time domain, frequency domain, and audio domain from the denoised vibration signal, the characteristics of the signal in different dimensions can be comprehensively characterized. Specifically, time domain features can reflect the overall trend, fluctuation, and impulsivity of the signal; frequency domain features can reveal the energy distribution and main vibrational components; and audio domain features can capture the spectral morphology and dynamic change patterns. The combination of these three types of features forms an 11-dimensional statistical feature vector, which not only ensures the diversity and representativeness of the features but also has good interpretability and low computational complexity.
[0045] S4: Features of the denoised vibration signal are extracted by continuous wavelet transform and short-time Fourier transform, respectively, and the extraction results are fused based on the entropy weighting mechanism to obtain a preliminary fused feature map.
[0046] Continuous wavelet transform (CWT) is a time-frequency analysis method suitable for non-stationary signal analysis. Its core idea is to convolve the signal with wavelet functions of different scales and displacements, thus simultaneously unfolding the signal in both the time and scale domains. CWT can provide good frequency resolution at low frequencies and good time resolution at high frequencies, thereby taking into account both long-term trends and short-term abrupt changes. It is particularly suitable for detecting local features such as transient shocks and abrupt change points.
[0047] Short-Time Fourier Transform (STFT) is a time-frequency analysis method that performs Fourier transform on a signal within a short time window. By sliding a window across the signal, the overall non-stationary signal is decomposed into a series of locally approximately stationary segments, and the spectrum is calculated within each window to obtain the frequency distribution of the signal over time. STFT maintains a fixed balance between frequency resolution and time resolution, making it suitable for analyzing local periodicity and energy distribution patterns.
[0048] In one possible implementation, S4 specifically includes: S401: Perform continuous wavelet transform on the denoised vibration signal to obtain the continuous wavelet time-frequency diagram:
[0049] in, This represents a continuous wavelet time-frequency plot. Indicates the translation parameter. Indicates the scale parameter. This represents the denoised vibration signal. Describing the wavelet function, Represents the integral symbol.
[0050] S402: Using the Hanning window method, a short-time Fourier transform is performed on the denoised vibration signal to obtain the short-time Fourier frequency diagram:
[0051] in, Represents the short-time Fourier time-frequency plot. Represents the window function. It represents a complex sine wave.
[0052] The Hanning window is a commonly used windowing function, typically used in spectral analysis such as short-time Fourier transform. It is defined as a smooth curve with a cosine shape within the window length, which gradually decays to zero in the time domain.
[0053] S403: Calculate the adaptive entropy weights:
[0054] in, Represents the adaptive entropy value weight. This represents the information entropy function.
[0055] S404: Based on adaptive entropy weights, the continuous wavelet time-frequency map and the short-time Fourier time-frequency map are weighted and fused to obtain a preliminary fused feature map:
[0056] in, This indicates the initial fusion of feature maps. This represents a continuous wavelet time-frequency plot. This represents a short-time Fourier frequency plot.
[0057] In this embodiment of the invention, by performing continuous wavelet transform (CWT) and short-time Fourier transform (STFT) on the denoised vibration signal, the local variation characteristics of the signal in both the time and frequency domains can be obtained simultaneously, thus fully presenting the dynamic evolution law of the vibration signal in a two-dimensional time-frequency space. Furthermore, by using an entropy weighting mechanism to adaptively weight and fuse the results of CWT and STFT, the time-frequency graph with greater information content and more complete feature expression can occupy a higher proportion in the fused result, thereby improving the effectiveness and robustness of feature expression.
[0058] S5: Based on the Gram difference field coding principle, feature extraction is performed on the denoised vibration signal to obtain Gram difference field image features.
[0059] It should be noted that the Gramian Angular Field (GAF) encoding principle is to map one-dimensional time series data to a two-dimensional polar coordinate system. These mapped values are used to construct an angle matrix, where each value corresponds to a pixel, thus generating an image. Constructing the GAF matrix first requires normalization, normalizing the one-dimensional time series data to the interval [-1, 1] so that it can be mapped onto the unit circle.
[0060] In one possible implementation, S5 specifically includes: S501: Represent the denoised vibration signal as one-dimensional time series data:
[0061] in, Represents one-dimensional time series data. x u Indicates the first u A spatial point, u =1,2,…, T , T This indicates the total length of the time series.
[0062] S502: Normalize the one-dimensional time series data to obtain the normalized sequence:
[0063] in, Represents a normalized sequence. This indicates that the value of the smallest element in the array is obtained. This indicates that the value of the maximum element in the array is obtained.
[0064] S503: Map the normalized sequence to a polar coordinate system and calculate the polar angle and radius in the polar coordinate system:
[0065]
[0066] in, Represents the polar angle in the polar coordinate system. Represents the inverse cosine function. Represents the radius in the coordinate system.
[0067] S504: Based on the angle and radius in polar coordinates, calculate the angular difference relationship between all time points and construct the Gram angular difference field matrix:
[0068] in, The angular difference field matrix of each channel represents the first... i OK j Column elements, Represents the cosine function. Indicates a point in time j The radius in polar coordinates. Indicates a point in time j The polar angle in the polar coordinate system.
[0069] It should be noted that GADF(i,j) is the i-th element in the GADF matrix.i OK j The matrix elements of the column represent the nth column in the original time series. i The time point and the j A special relationship exists between n time points, where each element in the matrix represents the relationship between two time points. Therefore, the nth time point of this matrix... i OK j The element calculation for the column is based on the time point. i , j Solve for the polar radius and polar angle in polar coordinates.
[0070] It should be noted that the Gramian Angular Difference Field (GADF) in GAF is more sensitive to local time differences, such as abrupt changes and noise, and is suitable for vibration events such as ship collisions with bridges.
[0071] S505: Output the Gram angular difference field matrix as the Gram angular difference field image feature.
[0072] It should be noted that by employing Gram Angular Difference Field (GADF) encoding, the denoised vibration time series is mapped to a two-dimensional polar coordinate space, and an angle difference matrix is constructed, transforming a one-dimensional signal into two-dimensional image features. This process not only preserves the temporal order of the signal but also encodes global temporal dependence and local variation patterns into geometric relationships between pixels, thereby achieving a visual representation of the time series. Compared to conventional time-frequency mapping methods, GADF features are more sensitive to local abrupt changes, anomalous fluctuations, and short-term shocks in the time series, effectively capturing the unique temporal evolution patterns of events such as ship collisions with bridges.
[0073] S6: Perform a convolution operation on the preliminary fused feature map to extract local features, and then perform weighted fusion of the local features and Gram angle difference field image features to obtain the fused image features.
[0074] In one possible implementation, S6 specifically includes: S601: Perform a 3×3 convolution operation on the preliminary fused image to extract local features.
[0075] S602: Using the ReLU activation function, local features are weighted and fused with Gram angular difference field image features to obtain fused image features:
[0076]
[0077] in, Indicates fused image features, Indicates the final fusion weight. This indicates the activation of the Han function. This represents a 3×3 convolution operation. Indicates the features of the Gram angular difference field image. This represents the information entropy function.
[0078] In this embodiment of the invention, by employing convolution operations on the initial fused time-frequency map, local energy distribution and texture features can be effectively extracted, enhancing the expression of short-term impacts and local patterns. Based on this, the local features obtained from convolution are weighted and fused with Gram Angular Difference Field (GADF) features, and an entropy weighting mechanism is introduced. This allows for adaptive allocation of the fusion ratio according to the amount of information contained in different features, thereby avoiding the problem of one-sided representation by a single feature.
[0079] S7: Perform depthwise separable convolution on the fused image features to extract image depth features.
[0080] Reference manual attached Figure 3 The diagram illustrates a comparison between depth-separable convolution and ordinary convolution provided by an embodiment of the present invention.
[0081] It should be noted that traditional methods for deep feature extraction suffer from excessive parameters and low computational efficiency. To address this drawback, this scheme employs depthwise separable convolution to extract spatial details from the fused feature map. Operationally, depthwise separable convolution performs convolution on each channel individually, rather than fusing across channels. This method reduces the number of parameters, improves computational efficiency, and enhances the real-time performance of identifying and classifying ship-bridge collision vibration events.
[0082] In one possible implementation, S7 specifically includes: S701: Uses depthwise separable convolution to extract spatial detail features from the fused feature map.
[0083] in, To represent spatial details This represents depthwise separable convolution. This represents the features of the fused image.
[0084] Optionally, in order to efficiently and selectively extract key spatial-channel depth features from the fused feature map, this scheme introduces a channel and spatial attention mechanism.
[0085] Among them, the channel attention mechanism is an adaptive feature weight allocation method. Its core idea is to perform global statistics on the feature map in the spatial dimension to obtain the global representation of each channel, and then generate weight coefficients through nonlinear mapping to measure the importance of different channels in the overall task. By multiplying these weights with the original features channel by channel, it is possible to highlight the channel features that contribute more to classification or recognition, suppress invalid or redundant channels, and thus improve the discriminative power and robustness of the model.
[0086] Spatial attention is a method that focuses on the "positional importance" of feature maps. Its core idea is to compress the feature map along the channel dimension and then generate a spatial weight matrix through convolution operations to represent the degree of attention at different spatial locations. After multiplying this weight element-wise with the original feature map, the model can focus more on regions with concentrated energy and significant changes, while weakening background or noise interference, thereby enhancing the ability to capture local regional patterns.
[0087] S702: Utilizing a channel attention mechanism, global average pooling is performed on spatial detail features. The pooling result is input into a multilayer perceptron, and the sigmoid function is used to generate channel attention weights.
[0088] in, Indicates channel attention weights. This represents the Sigmoid function. This represents a multilayer perceptron. This indicates global average pooling.
[0089] S703: Multiply the channel attention weights by the spatial detail features channel by channel to obtain the channel-weighted features:
[0090] in, Indicates channel weighting characteristics, This indicates multiplication by channel. It represents spatial detail features.
[0091] S704: Global average pooling and global max pooling are performed on the channel-weighted features respectively. The pooling results are concatenated along the channel dimension and then subjected to a 7×7 convolution operation. The sigmoid activation function is used to generate spatial attention weights.
[0092] in, Represents spatial attention weights. This represents the Sigmoid activation function. This represents a 7×7 convolution operation. This indicates global max pooling.
[0093] S705: Perform pointwise convolution on the channel-weighted features, and multiply the spatial attention weights elementwise with the pointwise convolution result to obtain the image depth features:
[0094] in, Represents image depth features, This represents pointwise convolution.
[0095] In this embodiment of the invention, by employing depthwise separable convolution on the fused image features, spatial detail features can be effectively extracted while significantly reducing the number of network parameters and computational complexity, meeting the lightweight requirements of real-time monitoring scenarios. Based on this, a channel attention mechanism is introduced, using global average pooling and a multilayer perceptron to model the importance of each channel, enabling the model to adaptively highlight feature channels that contribute significantly to classification. Simultaneously, combined with a spatial attention mechanism, global pooling and convolution operations are performed on the feature map to generate a spatial weight distribution, which can highlight areas of concentrated vibration energy and suppress irrelevant background.
[0096] S8: Based on image depth features, multi-scale dilated convolution is used to capture the long temporal dependence of vibration events, resulting in the final sequence features.
[0097] It should be noted that when a ship collides with a bridge pier, high-frequency vibration waves propagate upwards along optical fibers to the bridge tower, while low-frequency structural responses diffuse to the bridge deck, forming unique spatiotemporal amplitude distribution characteristics. Therefore, this solution introduces sequence feature processing on the basis of deep feature extraction, focusing on the time dimension to capture the dynamic laws and temporal dependencies of vibration signals as they change over time, and paying attention to the correlation between features at different times, such as the impact changes of vibration signals in a short period of time and the continuous trend over a long period of time, so as to more accurately identify and classify ship-bridge collision vibration events.
[0098] Multi-scale dilated convolution is a convolutional modeling method used to capture short- and long-term dependencies. Its basic idea is to introduce different dilation rates into the convolution operation, expanding the receptive field by inserting holes between the sampling points of the convolution kernel. With the same computational cost, convolutional branches with smaller dilation rates can extract local short-term details, while those with larger dilation rates can cover global patterns over longer time spans, thus achieving joint modeling of information at different time scales. Through multi-scale parallel dilated convolution structures, both short-term impact features and long-term trend features of vibration signals can be captured simultaneously in a single computation, improving the model's ability to represent complex dynamic events.
[0099] In one possible implementation, S8 specifically includes: S801: Perform multi-scale dilation convolution on the image depth features, and then perform layer normalization on the dilation convolution results to obtain preliminary sequence features under different dilation rates:
[0100] in, Indicates the first l Layer expansion rate d sequence features, Representation layer normalization, This represents multi-scale dilated convolution. Indicates the expansion rate d The corresponding input image features.
[0101] S802: Concatenate the preliminary sequence features under different expansion rates to obtain the concatenated sequence features:
[0102] in, Indicates the features of the spliced sequence. Indicates the first l Features of spliced sequences with a layer dilation rate of 1 Indicates the first l Features of spliced sequences with a layer dilation rate of 2 Indicates the first l Features of spliced sequences with a layer dilation rate of 4 Indicates the first l Features of spliced sequences with a layer expansion rate of 8.
[0103] S803: Using the learnable first gating parameter matrix, a linear mapping is performed on the features of the concatenated sequence to obtain dynamic gating weights.
[0104] in, Indicates dynamic gating weights, This represents the Sigmoid activation function. This represents the first gating parameter matrix.
[0105] S804: Using a learnable second gating parameter matrix, convolution is performed on the concatenated sequence features, and after activation by an activation function, the features are multiplied element-wise with dynamic gating weights to obtain the final sequence features.
[0106] in, Indicates the final sequence features, This indicates element-wise multiplication. express Activation function.
[0107] In this embodiment of the invention, by introducing a multi-scale dilated convolutional structure based on image depth features, the receptive field can be significantly expanded with the same computational cost. Parallel branches with different dilation rates are used to simultaneously capture both short-term impact features and long-term trend features of vibration signals, thereby achieving joint modeling of the short- and long-term dependencies of vibration events. Furthermore, a gating mechanism is employed to dynamically filter the stitched multi-scale features, using learnable gating weights to highlight information useful for classification and suppress irrelevant or redundant features.
[0108] S9: Perform multimodal fusion of statistical features, deep image features, and final sequence features to obtain the final fused features.
[0109] In one possible implementation, S9 specifically includes: S901: Utilizing inter-information technology to measure the nonlinear correlation between statistical features and deep image features:
[0110] in, Indicates mutual information, Indicates the first i Statistical characteristics of each channel dimension Indicates the first j Deep image features in one channel dimension This represents a double integral. Represents the logarithmic function. This represents the integration operator. This represents the joint probability distribution between statistical features and deep image features.
[0111] Mutual information (MI) is an information-theoretic metric that measures the correlation and dependence between two random variables. Essentially, it quantifies the amount of information shared between variables by calculating the difference between their joint distribution and their respective marginal distributions. A higher mutual information value indicates a stronger correlation between the two variables; zero mutual information indicates that the variables are independent.
[0112] S902: Determine the adaptive threshold based on the mutual information between statistical features and deep image features.
[0113] in, Indicates an adaptive threshold. This indicates taking the maximum value.
[0114] S903: Filter statistical features that have a non-linear correlation with any dimension of deep image features greater than or equal to an adaptive threshold.
[0115] in, This represents the statistical characteristics after filtering. Indicates statistical characteristics, This indicates retrieving all channel dimensions of the image features. j The maximum value in.
[0116] S904: Principal component analysis is used to reduce the dimensionality of deep image features, resulting in dimensionality-reduced deep image features.
[0117] in, Represents deep features of a dimensionality-reduced image. This indicates that dimensionality reduction was achieved through principal component analysis. Represents deep features of an image. k This indicates the number of principal components retained after dimensionality reduction.
[0118] S905: Map the filtered statistical features, dimensionality-reduced deep image features, and final sequence features to the same dimension, and use the mapped statistical features, dimensionality-reduced deep image features, and final sequence features as the query matrix and key matrix, respectively, to calculate the corresponding self-attention scores:
[0119] in, Represents the self-attention score. Represents the query matrix. Represents the key matrix. T This indicates the transpose operation. D The dimension representing the feature.
[0120] S906: Perform a softmax operation on the self-attention scores to obtain the self-attention weights:
[0121] in, This represents the self-attention weight.
[0122] S907: Based on self-attention weights, the mapped statistical features, the deep features of the reduced-dimensional image, and the final sequence features are fused to obtain the final fused features:
[0123] in, Indicates the final fusion characteristics, This represents the weights of the statistical features after mapping. This represents the statistical characteristics after mapping. This represents the weights of the deep features in the dimensionality-reduced image after mapping. This represents the deep features of the mapped, dimensionality-reduced image. The weights represent the features of the final sequence after mapping. This represents the final sequence features after mapping.
[0124] In this embodiment of the invention, by introducing a multimodal fusion mechanism based on statistical features, fused image features, and sequence features, the complementary advantages of different feature dimensions can be fully utilized.
[0125] S10: Based on the final fusion features, the vibration event category classification results are output through a fully connected layer.
[0126] Specifically, the fused features Convert to a fixed-length vector This means compressing the spatial and temporal dimensions to one. A fixed-length vector can be represented as:
[0127] in, Represents the global feature vector. Indicates global average pooling. This indicates the final fusion feature.
[0128] right Batch normalization is obtained The result can be expressed as:
[0129] in, This represents the eigenvectors after batch normalization. This indicates batch normalization.
[0130] Again A dropout operation is performed with a dropout probability p = 0.3 to prevent overfitting. The result can be expressed as:
[0131] in, This represents the feature vector after the discard operation. This indicates a discard operation.
[0132] Finally, through a fully connected layer, combined with learnable weights... and bias The data is then processed using the softmax function to output the probability of each vibration event belonging to a specific category. The classification probability vector can be represented as:
[0133] in, Represents the classification probability vector. This represents the softmax function. Represents the learnable weights. Indicates bias. ∈ , N This indicates the number of vibration event categories.
[0134] In this embodiment of the invention, by batch normalizing and dropout processing the final fused features, not only can global discriminative information be preserved while reducing feature dimensionality and computational complexity, but feature distribution shifts can also be effectively suppressed and model overfitting can be prevented, thereby improving the stability and generalization ability of the model. Furthermore, by combining a fully connected layer with a softmax classifier to output multi-class probability distributions, different perturbation events can be accurately distinguished. Compared with traditional classification methods based on a single feature or fixed threshold, the classification output mechanism of this scheme significantly improves the classification accuracy and reliability of complex perturbation events such as ship collisions, traffic loads, and wind-induced vibrations while ensuring real-time performance.
[0135] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following: In this embodiment of the invention, the real-time performance and comprehensiveness of bridge monitoring are improved through vibration signal acquisition, denoising, and multi-domain feature extraction. Combined with multi-level feature learning using wavelet transform, Fourier transform, Gram difference field coding, and deep convolutional networks, the ability to identify damage and abnormal events is enhanced, effectively overcoming the shortcomings of low efficiency, high cost, and susceptibility to oversights in manual inspections. Simultaneously, denoising processing, multi-modal feature fusion, and deep convolutional structures improve the stability and robustness of monitoring results, enhance resistance to electromagnetic interference and environmental changes, and ultimately achieve accurate event identification through classification output, meeting the urgent need for efficient, accurate, and intelligent bridge monitoring in high-risk environments.
[0136] Reference manual attached Figure 4 The diagram shows a structural schematic of a vibration event classification system based on distributed optical fiber sensing provided by the present invention.
[0137] The present invention also provides a vibration event classification system 20 based on distributed optical fiber sensing, applied to the above-mentioned vibration event classification method based on distributed optical fiber sensing, comprising: Processor 201.
[0138] The memory 202 stores computer-readable instructions, which, when executed by the processor 201, implement the vibration event classification method based on distributed optical fiber sensing as described in the method embodiment.
[0139] The vibration event classification system 20 based on distributed optical fiber sensing provided by the present invention can perform the vibration event classification method based on distributed optical fiber sensing described above and achieve the same or similar technical effects. To avoid repetition, the present invention will not elaborate further.
[0140] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following: In this embodiment of the invention, the real-time performance and comprehensiveness of bridge monitoring are improved through vibration signal acquisition, denoising, and multi-domain feature extraction. Combined with multi-level feature learning using wavelet transform, Fourier transform, Gram difference field coding, and deep convolutional networks, the ability to identify damage and abnormal events is enhanced, effectively overcoming the shortcomings of low efficiency, high cost, and susceptibility to oversights in manual inspections. Simultaneously, denoising processing, multi-modal feature fusion, and deep convolutional structures improve the stability and robustness of monitoring results, enhance resistance to electromagnetic interference and environmental changes, and ultimately achieve accurate event identification through classification output, meeting the urgent need for efficient, accurate, and intelligent bridge monitoring in high-risk environments.
[0141] It should be understood that the processor in the embodiments of the present invention can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0142] It should also be understood that the memory in the embodiments of the present invention can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced synchronous DRAM (ESDRAM), synchronous linked DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0143] The above embodiments can be implemented, in whole or in part, by software, hardware (such as circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.
[0144] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. A and B can be singular or plural. Additionally, the character " / " in this article generally indicates an "or" relationship between the preceding and following related objects, but it can also represent an "and / or" relationship. Please refer to the context for a more accurate understanding.
[0145] In this invention, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of a single item or a plurality of items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be a single item or multiple items.
[0146] It should be understood that, in various embodiments of the present invention, the order of the above-mentioned process numbers does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0147] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0148] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices, apparatuses, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0149] In the several embodiments provided by this invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0150] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0151] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0152] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0153] This invention provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the vibration event classification method based on distributed optical fiber sensing as described in the method embodiment.
[0154] The present invention provides a computer-readable storage medium that can implement the steps and effects of the vibration event classification method based on distributed optical fiber sensing in the above-described method embodiments. To avoid repetition, the present invention will not repeat them.
[0155] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following: In this embodiment of the invention, the real-time performance and comprehensiveness of bridge monitoring are improved through vibration signal acquisition, denoising, and multi-domain feature extraction. Combined with multi-level feature learning using wavelet transform, Fourier transform, Gram difference field coding, and deep convolutional networks, the ability to identify damage and abnormal events is enhanced, effectively overcoming the shortcomings of low efficiency, high cost, and susceptibility to oversights in manual inspections. Simultaneously, denoising processing, multi-modal feature fusion, and deep convolutional structures improve the stability and robustness of monitoring results, enhance resistance to electromagnetic interference and environmental changes, and ultimately achieve accurate event identification through classification output, meeting the urgent need for efficient, accurate, and intelligent bridge monitoring in high-risk environments.
[0156] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
[0157] The following points need to be explained: (1) The accompanying drawings of the embodiments of the present invention only involve the structures involved in the embodiments of the present invention. Other structures can refer to the general design.
[0158] (2) For clarity, the thickness of layers or regions is enlarged or reduced in the drawings used to describe embodiments of the invention, i.e., these drawings are not drawn to scale. It is understood that when an element such as a layer, film, region or substrate is referred to as being “above” or “below” another element, the element may be “directly” located “above” or “below” the other element or there may be intermediate elements.
[0159] (3) Where there is no conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other to obtain new embodiments.
[0160] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. The scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A vibration event classification method based on distributed optical fiber sensing, characterized in that, include: S1: Acquire vibration signals; S2: Denoise the vibration signal to obtain a denoised vibration signal; S3: Extract time-domain features, frequency-domain features, and audio-domain features from the denoised vibration signal to form statistical features; S4: The denoised vibration signal is subjected to feature extraction by continuous wavelet transform and short-time Fourier transform respectively, and the extraction results are fused based on the entropy weighting mechanism to obtain a preliminary fused feature map; S5: Based on the Gram difference field coding principle, feature extraction is performed on the denoised vibration signal to obtain Gram difference field image features; S6: Perform a convolution operation on the preliminary fused feature map to extract local features, and then perform a weighted fusion of the local features and the Gram angle difference field image features to obtain fused image features; S7: Perform a depthwise separable convolution operation on the fused image features to extract image depth features; S8: Based on the image depth features, multi-scale dilated convolution is used to capture the long temporal dependence of vibration events to obtain the final sequence features; S9: Perform multimodal fusion on the statistical features, the image depth features, and the final sequence features to obtain the final fused features; S10: Based on the final fusion features, output the category classification results of vibration events through the fully connected layer.
2. The distributed fiber optic sensing based vibration event classification method of claim 1, wherein, S2 specifically includes: S201: Decompose the vibration signal using wavelet basis functions to extract low-frequency and high-frequency components: ; in, Indicates low-frequency components. Represents high-frequency components. Represents discrete wavelet transform. x ( t ) represents a vibration signal. Represents the Symlet-5 wavelet basis functions; S202: Calculate the kurtosis value of the high-frequency component and normalize the kurtosis value; S203: Calculate the adaptive threshold based on the normalized kurtosis value: ; in, Indicates an adaptive threshold. Indicates the standard deviation of noise. N Indicates signal length. Indicated by e Logarithmic function with base 0. Represents the hyperbolic tangent function. Kurtosis of high-frequency components; S204: Threshold the high-frequency components to retain the high-frequency elements and obtain the noise-reduced high-frequency components: ; wherein denotes a noise-reduced high-frequency component, denotes a sign function, max( ) denotes a maximum function, denotes taking an absolute value; S205: Perform inverse discrete wavelet transform on the low-frequency component and the denoised high-frequency component to obtain the reconstructed signal, and perform morphological filtering on the reconstructed signal to obtain the denoised vibration signal. ; in, This represents the denoised vibration signal. This represents the closing operation. This indicates the opening operation. This represents the inverse discrete wavelet transform.
3. The vibration event classification method based on distributed optical fiber sensing according to claim 1, characterized in that, S3 specifically includes: S301: Extract the time-domain features of the denoised vibration signal, including mean suos, variance, kurtosis, mean absolute value, and zero-crossing rate. S302: Extract the frequency domain features of the denoised vibration signal, including the energy proportion of 0-50Hz, peak frequency, and energy second derivative at characteristic frequencies; S303: Extract the audio domain features of the denoised vibration signal, including Mel frequency cepstral coefficients, spectral center, and spectral change rate; S304: Combine the time-domain features, the frequency-domain features, and the audio-domain features to form the statistical features: ; in, Indicates statistical characteristics, This represents the mean. Represents variance. Indicates kurtosis, Represents the absolute value of the average. Indicates the zero-crossing rate. This indicates the energy percentage in the 0-50Hz range. Indicates peak frequency. This represents the second derivative of the energy at the characteristic frequency. This represents the cepstral coefficients of the first 5 Mel frequencies. Indicates the spectral center, This represents the rate of change of the spectrum.
4. The distributed fiber optic sensing based vibration event classification method of claim 1, wherein, S4 specifically includes: S401: Perform continuous wavelet transform on the denoised vibration signal to obtain the continuous wavelet time-frequency diagram: ; in, This represents a continuous wavelet time-frequency plot. Indicates the translation parameter. Indicates the scale parameter. This represents the denoised vibration signal. Describing the wavelet function, Indicates the integral symbol; S402: Using the Hanning window method, perform a short-time Fourier transform on the denoised vibration signal to obtain the short-time Fourier time-frequency diagram: ; wherein denotes a short-time Fourier time-frequency diagram, denotes a window function, denotes a complex sinusoidal wave; S403: Calculate the adaptive entropy weights: ; wherein, denotes an adaptive entropy weight, denotes an information entropy function; S404: Based on the adaptive entropy weights, the continuous wavelet time-frequency map and the short-time Fourier time-frequency map are weighted and fused to obtain the preliminary fused feature map: ; wherein, represents the preliminary fused feature map.
5. The vibration event classification method based on distributed optical fiber sensing according to claim 4, characterized in that, S5 specifically includes: S501: Represent the denoised vibration signal as one-dimensional time series data: ; in, Represents one-dimensional time series data. x u Indicates the first u A spatial point, u =1,2,…, T , T Indicates the total length of the time series; S502: Normalize the one-dimensional time series data to obtain a normalized sequence: ; in, Represents a normalized sequence. This indicates that the value of the smallest element in the array is obtained. This indicates that the value of the maximum element in the array is obtained; S503: Map the normalized sequence to a polar coordinate system, and calculate the polar angle and radius in the polar coordinate system: ; ; in, Represents the polar angle in the polar coordinate system. Represents the inverse cosine function. Represents the radius in the coordinate system; S504: Based on the angle and radius of the polar coordinates, calculate the angle difference relationship between all time points and construct the Gram angle difference field matrix: ; in, Represents time points in a normalized time series i and time point j The relationship between them Represents the cosine function. Indicates a point in time j The radius in polar coordinates. Indicates a point in time j Polar angle in polar coordinates; S505: Output the Gram angular difference field matrix as the Gram angular difference field image feature.
6. The distributed fiber optic sensing based vibration event classification method of claim 4, wherein, S6 specifically includes: S601: Perform a 3×3 convolution operation on the preliminary fused feature map to extract local features; S602: Using the ReLU activation function, the local features and the Gram angular difference field image features are weighted and fused to obtain the fused image features: ; ; in, Indicates fused image features, Indicates the final fusion weight. This indicates the activation of the Han function. This represents a 3×3 convolution operation. Indicates the features of the Gram angular difference field image. This represents the information entropy function.
7. The distributed fiber optic sensing based vibration event classification method of claim 1, wherein, Specifically, S7 includes: S701: Using depthwise separable convolution, extract the spatial detail features of the fused image features: ; wherein, denotes spatial detail features, denotes depthwise separable convolution, denotes fusing image features; S702: Using a channel attention mechanism, global average pooling is performed on the spatial detail features. The pooling result is input into a multilayer perceptron, and the channel attention weights are generated using the sigmoid function. ; wherein, denotes a channel attention weight, denotes a Sigmoid function, denotes a multi-layer perceptron, denotes a global average pooling; S703: Multiply the channel attention weights by the spatial detail features channel by channel to obtain the channel-weighted features: ; wherein, represents a channel weighting feature, represents a channel-wise multiplication; S704: Perform global average pooling and global max pooling on the channel-weighted features respectively. After concatenating the pooling results along the channel dimension, perform a 7×7 convolution operation and use the Sigmoid activation function to generate spatial attention weights. ; wherein, denotes spatial attention weights, denotes a Sigmoid activation function, denotes a 7x7 convolution operation, denotes global max pooling; S705: Perform pointwise convolution on the channel-weighted features, and multiply the spatial attention weights elementwise with the pointwise convolution result to obtain the image depth features: ; wherein, denotes an image depth feature, denotes a point-wise convolution.
8. The distributed fiber optic sensing based vibration event classification method of claim 1, wherein, S8 specifically includes: S801: Perform multi-scale dilation convolution on the image depth features, and perform layer normalization on the dilation convolution results to obtain preliminary sequence features under different dilation rates: ; in, Indicates the first l Layer expansion rate d sequence features, Representation layer normalization, This represents multi-scale dilated convolution. Indicates the expansion rate d The corresponding input image features; S802: Concatenate the preliminary sequence features under different expansion rates to obtain concatenated sequence features: ; in, Indicates the features of the spliced sequence. Indicates the first l Features of spliced sequences with a layer dilation rate of 1 Indicates the first l Features of spliced sequences with a layer dilation rate of 2 Indicates the first l Features of spliced sequences with a layer dilation rate of 4 Indicates the first l Features of spliced sequences with a layer dilation rate of 8; S803: Using the learnable first gating parameter matrix, the features of the concatenated sequence are linearly mapped to obtain dynamic gating weights. ; wherein, denotes a dynamic gating weight, denotes a sigmoid activation function, denotes a first gating parameter matrix; S804: Using a learnable second gating parameter matrix, perform a convolution operation on the concatenated sequence features, and after activation by an activation function, multiply them element-wise with dynamic gating weights to obtain the final sequence features: ; wherein, represents the final sequence feature, represents element-wise multiplication, represents an activation function.
9. The distributed optical fiber sensing based vibration event classification method of claim 1, wherein, S9 specifically includes: S901: Using inter-information technology, measure the nonlinear correlation between the statistical features and the image depth features: ; in, Indicates mutual information, Indicates the first i Statistical characteristics of each channel dimension Indicates the first j Image depth features in one channel dimension This represents a double integral. Represents the logarithmic function. This represents the integration operator. This represents the joint probability distribution between statistical features and image depth features; S902: Based on the mutual information between the statistical features and the image depth features, determine the adaptive threshold: ; wherein denotes an adaptive threshold, denotes taking the maximum value; S903: Filter the statistical features whose nonlinear correlation with any dimension of the image depth features is greater than or equal to the adaptive threshold: ; wherein, denotes the filtered statistical feature, denotes the statistical feature, denotes taking the maximum value over all channel dimensions of the image feature j over all channel dimensions of the image feature S904: The image depth features are reduced in dimension using principal component analysis to obtain the reduced image depth features: ; wherein, represents a reduced dimension image depth feature, represents a dimension reduction process by principal component analysis, represents an image depth feature, k represents the number of principal components retained after dimension reduction; S905: Map the filtered statistical features, the reduced-dimensional image depth features, and the final sequence features to the same dimension, and use the mapped statistical features, reduced-dimensional image depth features, and final sequence features as the query matrix and key matrix to calculate the corresponding self-attention scores: ; wherein, denotes a self-attention score, denotes a query matrix, denotes a key matrix, T denotes a transpose operation, D denotes a dimension of a feature; S906: Perform a softmax operation on the self-attention scores to obtain the self-attention weights: ; wherein, denotes a self-attention weight; S907: Based on the self-attention weights, the mapped statistical features, image depth features, and final sequence features are fused to obtain the final fused features: ; wherein, denotes the final fusion feature, denotes the weight of the mapped statistical feature, denotes the mapped statistical feature, denotes the weight of the mapped reduced image depth feature, denotes the mapped image depth feature, denotes the weight of the mapped final sequence feature, denotes the mapped final sequence feature.
10. A distributed fiber optic sensing based vibration event classification system, characterized by, include: processor; A memory storing computer-readable instructions, which, when executed by the processor, implement the vibration event classification method based on distributed optical fiber sensing as described in any one of claims 1 to 9.