Vibration event classification method and system based on distributed optical fiber sensing
By using distributed fiber optic sensing technology, combined with multi-domain feature extraction and deep convolutional networks, real-time and accurate identification of bridge vibration events was achieved, solving the problems of low efficiency and insufficient stability in existing technologies, and meeting the high-precision monitoring needs of bridges in high-risk environments.
Patent Information
- Application Number
- CN202511385237.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-26
- Publication Date
- 2026-01-09
- Estimated Expiration
- 2045-09-26
AI Technical Summary
Existing manual inspections are inefficient and costly, making it difficult to detect sudden damage in a timely manner. Traditional sensors are susceptible to electromagnetic interference and environmental factors, resulting in insufficient stability and reliability of bridge monitoring data, which cannot meet the high-precision safety monitoring needs in high-risk environments.
By employing distributed fiber optic sensing technology, vibration signals are collected and denoised. Combined with multi-level feature learning using wavelet transform, Fourier transform, Gram difference field coding, and deep convolutional networks, multi-modal features are extracted and fused to achieve real-time and accurate identification of bridge vibration events.
It improves the real-time and comprehensiveness of bridge monitoring, enhances the ability to identify damage and abnormal events, and improves the stability and robustness of monitoring results, thus meeting the needs for efficient, accurate and intelligent bridge monitoring in high-risk environments.
Smart Images

Figure CN121302000A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of structural health monitoring, in particular to a vibration event classification method and system based on distributed optical fiber sensing. BACKGROUND
[0002] With the rapid development of global modern transportation network, bridges as an important part of transportation infrastructure, their safety is directly related to the stable operation of the transportation system and the sustainable development of social economy. In recent years, ship impact on bridge (referred to as "ship collision") has become a prominent safety threat in bridge operation, which not only may cause local or overall damage of bridge structure, but also may cause collapse accident, resulting in serious casualties and economic losses. Especially in the past two decades, hundreds of ship collision accidents have occurred in China, showing a trend of frequent accidents and serious consequences, which urgently needs more efficient and accurate monitoring and prevention technology.
[0003] At present, bridge safety monitoring mainly relies on two technical means: manual inspection and fixed sensor monitoring. The manual inspection method relies on professional personnel to check and evaluate the bridge structure regularly, which has the advantage of qualitative judgment according to experience; the fixed sensor monitoring method realizes real-time data acquisition and state monitoring of local area by laying sensors at key parts of the bridge. The above methods have improved the safety of the bridge during operation to a certain extent, and have formed the main technical support of the current bridge monitoring system.
[0004] However, the existing manual inspection is low in efficiency and high in cost, and it is difficult to find sudden damage in time, which is easy to leave safety hazards due to human negligence. At the same time, the traditional sensor is easily affected by external environmental factors such as electromagnetic interference, temperature and humidity changes, resulting in insufficient stability and reliability of monitoring data, which is difficult to meet the urgent needs of high-precision and safe monitoring of bridges in high-risk environments. SUMMARY
[0005] In order to solve the technical problems that the existing manual inspection is low in efficiency and high in cost, and it is difficult to find sudden damage in time, which is easy to leave safety hazards due to human negligence. At the same time, the traditional sensor is easily affected by external environmental factors such as electromagnetic interference, temperature and humidity changes, resulting in insufficient stability and reliability of monitoring data, which is difficult to meet the urgent needs of high-precision and safe monitoring of bridges in high-risk environments, the present application provides a vibration event classification method and system based on distributed optical fiber sensing.
[0006] The technical scheme provided by the embodiments of the present application is as follows:
[0007] First aspect:
[0008] The vibration event classification method based on distributed optical fiber sensing provided by the embodiments of the present application comprises:
[0009] S1: collect a vibration signal;
[0010] S2: denoising processing is performed on the vibration signal to obtain a denoised vibration signal;
[0011] S3: time domain features, frequency domain features and audio domain features are extracted from the denoised vibration signal to form statistical features;
[0012] S4: feature extraction is respectively performed on the denoised vibration signal through continuous wavelet transform and short-time Fourier transform, and the extraction results are fused based on an entropy weight mechanism to obtain a preliminary fusion feature map;
[0013] S5: feature extraction is performed on the denoised vibration signal based on a Gram angle difference field coding principle to obtain a Gram angle difference field image feature;
[0014] S6: a convolution operation is performed on the preliminary fusion feature map to extract local features, and the local features and the Gram angle difference field image feature are weighted and fused to obtain a fusion image feature;
[0015] S7: a deep separable convolution operation is performed on the fusion image feature to extract image deep features;
[0016] S8: based on the image deep features, a multi-scale dilated convolution is used to capture long-time sequence dependence of a vibration event to obtain final sequence features;
[0017] S9: multi-modal fusion is performed on the statistical features, the image deep features and the final sequence features to obtain final fusion features;
[0018] S10: based on the final fusion features, a class classification result of a vibration event is output through a full connection layer.
[0019] The second aspect:
[0020] The embodiment of the present application provides a vibration event classification system based on distributed optical fiber sensing, which comprises:
[0021] a processor;
[0022] a memory, wherein the memory stores computer readable instructions, and the computer readable instructions are executed by the processor to realize the vibration event classification method based on distributed optical fiber sensing.
[0023] The third aspect:
[0024] The embodiment of the present application provides a computer readable storage medium, which stores a computer program, and the program is executed by a processor to realize the vibration event classification method based on distributed optical fiber sensing.
[0025] The technical scheme provided by the embodiment of the present application has at least the following beneficial effects:
[0026] In the embodiment of the present application, the real-time performance and comprehensiveness of bridge monitoring are improved through the collection, denoising and multi-domain feature extraction of the vibration signal, the multi-level feature learning of the wavelet transform, the Fourier transform, the Gram angle difference field coding and the deep convolution network is combined, and the identification capability of the damage and abnormal events is enhanced, so that the defects of low efficiency, high cost and easy omission of artificial inspection are effectively overcome. Meanwhile, the stability and robustness of the monitoring result are improved by using the denoising processing, the multi-modal feature fusion and the deep convolution structure, the anti-electromagnetic interference and environmental change capability is enhanced, and finally the accurate identification of the event is realized through the classification output, and the urgent needs of efficient, accurate and intelligent monitoring of the bridge in a high-risk environment are met. BRIEF DESCRIPTION OF DRAWINGS
[0027] In order to more clearly illustrate the technical scheme in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0028] Figure 1 A flowchart of a vibration event classification method based on distributed optical fiber sensing provided by the embodiment of the present application is shown in the figure.
[0029] Figure 2 An effect diagram of a signal preprocessing and denoising method provided by the embodiment of the present application is shown in the figure.
[0030] Figure 3 A deep separable convolution and ordinary convolution comparison diagram provided by the embodiment of the present application is shown in the figure.
[0031] Figure 4 A structure diagram of a vibration event classification system based on distributed optical fiber sensing provided by the embodiment of the present application is shown in the figure. DETAILED DESCRIPTION
[0032] The technical scheme in the present application will be described below with reference to the drawings.
[0033] In the embodiments of the present application, the words such as "exemplary", "for example", etc. are used to represent an example, an illustration or an explanation. Any embodiment or design solution described as "exemplary" in the present application should not be interpreted as being more preferred or having more advantages than other embodiments or design solutions. In fact, the word "exemplary" is used in the sense of presenting a concept in a concrete manner. In addition, in the embodiments of the present application, the meaning expressed by "and / or" can be both, or can be either one of the two.
[0034] In the embodiments of the present application, "image" and "picture" can be used interchangeably at times. It should be pointed out that the meanings expressed are consistent when the distinction is not emphasized. "Of", "corresponding" and "corresponding" can be used interchangeably at times. It should be pointed out that the meanings expressed are consistent when the distinction is not emphasized.
[0035] In the embodiments of the present application, sometimes the subscript such as W1 can be written in the form of non-subscript such as W1. The meanings expressed are consistent when the distinction is not emphasized.
[0036] In order to make the technical problems, technical solutions and advantages of the present application clearer, specific embodiments will be described in detail below with reference to the accompanying drawings.
[0037] Reference is made to the accompanying drawings and specific embodiments described in the specification Figure 1 , which shows a flowchart of a vibration event classification method based on distributed optical fiber sensing provided by an embodiment of the present application.
[0038] The embodiment of the present application provides a vibration event classification method based on distributed optical fiber sensing. The method can be realized by a vibration event classification device based on distributed optical fiber sensing. The vibration event classification device based on distributed optical fiber sensing can be a terminal or a server. The processing flow of the vibration event classification method based on distributed optical fiber sensing can include the following steps:
[0039] S1: Collecting a vibration signal.
[0040] Specifically, the sensing cable of the distributed optical fiber vibration sensing system is laid on the key parts of the bridge, such as the pier, the bridge deck, the bridge tower and the like, a high-density vibration monitoring network is formed by using the phase optical time domain reflection technology. The spatial resolution of the distributed optical fiber vibration sensing system can reach the order of meters, the sampling frequency can reach more than 10 kHz, and the vibration wave propagation process along the way can be captured in real time. The optical fiber detection signal is analyzed and processed by using the phase-sensitive optical time domain reflectometer (Φ-OTDR) and the DAS system, the amplitude data corresponding to the spatial position and the event sequence are obtained, and the vibration signal space-time matrix is constructed. The time-space matrix represents the vibration signals at different positions on the optical fiber in the spatial dimension, and represents the change of the vibration signal with time in the time dimension, and contains the space-time two-dimensional information, which is helpful for subsequent identification of the event type. When the bridge is hit, the high-density vibration monitoring network captures the vibration signal.
[0041] Reference is made to the accompanying drawings Figure 2 , which shows an effect schematic diagram of the signal preprocessing denoising method provided by the embodiment of the application.
[0042] S2: performing denoising processing on the vibration signal to obtain a denoised vibration signal.
[0043] It should be noted that, in view of the case that the original signal has high noise and complex information under complex working conditions, the wavelet-morphology joint denoising method is used to denoise the data.
[0044] In a possible implementation, S2 specifically includes:
[0045] S201: decomposing the vibration signal using a wavelet basis function to extract a low-frequency component and a high-frequency component:
[0046] [cA, cD] = DWT(x(t); Symlet-5)
[0047] wherein cA represents the low-frequency component, cD represents the high-frequency component, DWT represents the discrete wavelet transform, x(t) represents the vibration signal, and Symlet-5 represents the Symlet-5 wavelet basis function.
[0048] Among them, the wavelet basis function denoising is a common signal processing method, which uses wavelet transform to decompose the original signal into low-frequency components and high-frequency components of different scales, and the low-frequency part mainly retains the overall trend of the signal, and the high-frequency part contains noise and detail information. By selecting appropriate wavelet basis function (such as Symlet, Daubechies, etc.) and decomposition level, thresholding processing (soft threshold or hard threshold) is applied to the high-frequency component to suppress the noise interference, and then inverse wavelet transform is performed with the low-frequency component to reconstruct the smooth and close-to-real signal.
[0049] S202: Calculate the kurtosis value of the high-frequency component, and normalize the kurtosis value.
[0050] It should be noted that in order to realize the adaptive optimization of noise filtering, the threshold is dynamically adjusted according to the signal characteristics in this scheme. The kurtosis of the detail coefficient is selected as the signal characteristic, which represents the steepness of the signal distribution. The higher the kurtosis, the more pulse components it contains. At the same time, in order to avoid too large or too small threshold, the kurtosis is normalized to enhance the robustness.
[0051] S203: Calculate the adaptive threshold based on the normalized kurtosis value:
[0052]
[0053] Where λ represents the adaptive threshold, σ represents the noise standard deviation, N represents the signal length, ln represents the logarithmic function with e as the base, tanh represents the hyperbolic tangent function, and kurtosis(cD) represents the high-frequency component of the high-frequency component.
[0054] S204: Thresholding processing is performed on the high-frequency component to retain the high-frequency component, and the denoising high-frequency component is obtained:
[0055] cD' = sign(cD) max(|cD|-λ,0)
[0056] Where cD' represents the denoising high-frequency component, sign() represents the sign function, max() represents the maximum value function, and || represents the absolute value.
[0057] S205: Perform inverse discrete wavelet transform on the low-frequency component and the denoising high-frequency component to obtain the reconstructed signal, and perform morphological filtering on the reconstructed signal to obtain the denoising vibration signal:
[0058] x denoised (t) = γ close (γ open (IDWT(cA,cD'))
[0059] Where x denoised (t) represents the denoising vibration signal, γclose denotes a closing operation, γ open denotes an opening operation, and IDWT denotes an inverse discrete wavelet transform.
[0060] Specifically, the morphological filtering process includes first performing an opening operation (erosion + dilation) to remove small-scale noise such as spikes. Then, a closing operation (dilation + erosion) is performed to fill holes and eliminate impulse noise.
[0061] In the embodiment of the present application, the signal is divided into low-frequency components and high-frequency components by wavelet basis function decomposition, and the high-frequency components are thresholded by combining the kurtosis-driven adaptive thresholding process. Not only can the noise interference be effectively suppressed, but also the problems of excessive filtering or noise leakage existing in the fixed threshold method can be avoided. Further, the signal is reconstructed by inverse wavelet transform, and the morphological opening and closing operations are used in the time domain to filter, which can simultaneously remove spikes and isolated noise points, and maintain the continuity and integrity of the signal.
[0062] Compared with the existing single wavelet denoising or traditional filtering method, the present application can significantly improve the signal-to-noise ratio and robustness while ensuring the integrity of key information (such as transient impact and low-frequency structural response), providing more reliable input data for subsequent feature extraction and classification recognition, thereby enhancing the adaptability and overall monitoring accuracy of the system in complex environments.
[0063] S3: Extracting time-domain features, frequency-domain features, and audio-domain features from the denoised vibration signal to form statistical features.
[0064] In one possible implementation, S3 specifically includes:
[0065] S301: Extracting time-domain features including mean, variance, kurtosis, average absolute value, and zero-crossing rate from the denoised vibration signal.
[0066] S302: Extracting frequency-domain features including energy proportion of 0-50Hz, peak frequency, and energy second derivative at characteristic frequency from the denoised vibration signal.
[0067] S303: Extracting audio-domain features including mel-frequency cepstral coefficient, spectral center, and spectral change rate from the denoised vibration signal.
[0068] S304: Combining the time-domain features, frequency-domain features, and audio-domain features to form statistical features:
[0069]
[0070] wherein F stat denotes statistical features, μ denotes mean, σ 2 denotes variance, k denotes kurtosis, denotes average absolute value, and ZCR denotes zero-crossing rate, represents the energy proportion of 0-50Hz, f max represents the peak frequency, represents the energy second derivative at the characteristic frequency, MFCC 1:5 represents the first five mel-frequency cepstral coefficients, SC represents the spectral center, and Δ spec represents the spectral change rate.
[0071] Specifically, since the selection of features directly affects the classification effect of the vibration signal, in order to effectively distinguish the vibration signals corresponding to different disturbance events, artificial features of the vibration signal are extracted from the denoised vibration signal. The artificial features of the vibration signal include time domain features, frequency domain features and audio domain features of the vibration signal, which are respectively: the time domain features of the vibration signal include: mean, variance, kurtosis, average absolute value, zero-crossing rate; the frequency domain features of the vibration signal include: energy proportion of 0-50Hz, peak frequency, energy second derivative at the characteristic frequency; the audio domain features of the vibration signal include: the first five mel-frequency cepstral coefficients, spectral center, spectral change rate, and 11-dimensional statistical features are extracted.
[0072] In the embodiment of the application, by extracting multi-dimensional statistical features of time domain, frequency domain and audio domain from the denoised vibration signal, the characteristics of the signal in different dimensions can be fully described. Specifically, the time domain features can reflect the overall trend, volatility and impulsivity of the signal; the frequency domain features can reveal the energy distribution and main vibration components; the audio domain features can capture the spectral form and dynamic change rule. The combination of the three types of features forms an 11-dimensional statistical feature vector, which not only ensures the diversity and representativeness of the features, but also has good interpretability and low computational complexity.
[0073] S4: by continuous wavelet transform and short-time Fourier transform, feature extraction is performed on the denoised vibration signal respectively, and the fusion result is fused based on an entropy weight mechanism to obtain a preliminary fusion feature map.
[0074] Among them, the continuous wavelet transform is a time-frequency analysis method suitable for non-stationary signal analysis, and its core idea is to expand the signal in the time domain and scale domain by convolving the signal with wavelet functions of different scales and displacements. CWT can provide good frequency resolution at low frequency and good time resolution at high frequency, so as to balance long-term trend and short-term mutation, and is particularly suitable for detecting local features such as transient impact and mutation points.
[0075] Wherein, the short-time Fourier transform is a time-frequency analysis method of Fourier transform of signals in a short time window, by sliding window operation on the signal, the whole non-stationary signal is decomposed into a series of local approximate stationary segments, and the frequency spectrum is calculated in each window, so as to obtain the frequency distribution of the signal with time. STFT maintains a fixed balance between frequency resolution and time resolution, which is suitable for analyzing local periodicity and energy distribution law.
[0076] In one possible implementation, S4 specifically includes:
[0077] S401: Continuous wavelet transform is performed on the denoised vibration signal to obtain a continuous wavelet time-frequency diagram:
[0078]
[0079] Wherein, S cwt represents the continuous wavelet time-frequency diagram, τ represents the translation parameter, s represents the scale parameter, x denoised (t) represents the denoised vibration signal, ψ represents the wavelet function, and ∫ represents the integral symbol.
[0080] S402: Short-time Fourier transform is performed on the denoised vibration signal in the manner of Hanning window to obtain a short-time Fourier time-frequency diagram:
[0081] S stft (τf)=∫x denoised (t)·w(t-τ)·e -2πjft dt
[0082] Wherein, S stft represents the short-time Fourier time-frequency diagram, w() represents the window function, and e -2πjft represents a complex sinusoidal wave.
[0083] Wherein, the Hanning window is a commonly used windowing function, which is usually used in spectral analysis such as short-time Fourier transform, and its definition is a smooth curve in the form of cosine within the window length, which can gradually decay to zero in the time domain.
[0084] S403: Adaptive entropy weight is calculated:
[0085]
[0086] Wherein, θ represents the adaptive entropy weight, and entropy() represents the information entropy function.
[0087] S404: Based on the adaptive entropy weight, the continuous wavelet time-frequency diagram and the short-time Fourier time-frequency diagram are weighted and fused to obtain a preliminary fusion feature map:
[0088] S tf= · θ · S cwt + (1- θ) · S stft
[0089] wherein S tf denotes the preliminary fusion feature map, S cwt denotes the continuous wavelet time-frequency map, S stft denotes the short-time Fourier time-frequency map.
[0090] In the embodiment of the present application, by performing continuous wavelet transform (CWT) and short-time Fourier transform (STFT) on the denoised vibration signal respectively, the local variation characteristics of the signal in the time domain and the frequency domain can be obtained simultaneously, so that the dynamic evolution law of the vibration signal is fully presented in the two-dimensional time-frequency space. Further, by using the entropy weight mechanism to adaptively weight the results of CWT and STFT, the time-frequency map with larger information quantity and more sufficient feature expression can occupy a higher proportion in the fusion result, thereby improving the effectiveness and robustness of feature expression.
[0091] S5: performing feature extraction on the denoised vibration signal based on the Gram angle difference field coding principle to obtain a Gram angle difference field image feature.
[0092] It should be noted that the Gram angle field (Gramian Angular Field, GAF) coding principle is to map one-dimensional time series data to a two-dimensional polar coordinate system, and to construct an angle matrix using these mapping values, each value in the matrix corresponding to a pixel point, thereby generating an image. To construct the GAF matrix, first, normalization processing is required, which normalizes the one-dimensional time series data to the interval [-1, 1] so as to be mapped to the unit circle.
[0093] In one possible implementation, S5 specifically includes:
[0094] S501: representing the denoised vibration signal as one-dimensional time series data:
[0095] X = [x1, x2, …, xT] T ]
[0096] wherein X denotes one-dimensional time series data, x u denotes the u-th spatial point, u = 1, 2, …, T, and T denotes the total length of the time series.
[0097] S502: performing normalization processing on the one-dimensional time series data to obtain a normalized sequence:
[0098]
[0099] wherein, denotes the normalized sequence, min() denotes the value of the minimum element in the array, and max() denotes the value of the maximum element in the array.
[0100] S503: Map the normalized sequence to the polar coordinate system, and calculate the polar angle and radius in the polar coordinate system:
[0101]
[0102] wherein, denotes the polar angle in the polar coordinate system, arccos denotes the inverse cosine function, and r i denotes the radius in the coordinate system.
[0103] S504: According to the angle and radius of the polar coordinate, calculate the angle difference relationship of all time point pairs, and construct the Gram angle difference field matrix:
[0104]
[0105] wherein, GADF(i,j) denotes the element of the i-th row and j-th column of the Gram angle difference field matrix of each channel, cos denotes the cosine function, and r j denotes the radius of the time point j in the polar coordinate system, denotes the polar angle of the time point j in the polar coordinate system.
[0106] It should be noted that GADF(i,j) is the matrix element of the i-th row and j-th column in the GADF matrix, which represents a special relationship between the i-th time point and the j-th time point in the original time sequence, i.e. each element in the matrix represents the relationship between two time points. Then, the element of the i-th row and j-th column of the matrix is solved according to the polar coordinate radius and the polar angle of the time points i and j.
[0107] It should be noted that the Gram angle difference field (Gramian Angular Difference Field, GADF) in GAF is more sensitive to local time differences such as mutations and noise, and is suitable for ship-bridge collision vibration events.
[0108] S505: Output the Gram angle difference field matrix as the Gram angle difference field image feature.
[0109] It should be noted that by adopting the Gram angle difference field (GADF) encoding mode, the vibration time sequence after denoising is mapped to a two-dimensional polar coordinate space, and an angle difference matrix is constructed, so that the one-dimensional signal can be converted into a two-dimensional image feature. This process not only preserves the sequential relationship of the signal in the time dimension, but also encodes the global time dependence and local change pattern into the geometric relationship between pixels, thereby realizing the image expression of the time sequence. Compared with the conventional time-frequency diagram method, the GADF feature is more sensitive to local mutations, abnormal fluctuations and short-time impacts in the time sequence, and can effectively capture the unique pattern of events such as ship collision with the bridge in the time sequence evolution.
[0110] S6: performing convolution operation on the preliminary fusion feature map, extracting local features, and performing weighted fusion of the local features and the Gram angle difference field image feature to obtain a fusion image feature.
[0111] In a possible implementation, S6 specifically includes:
[0112] S601: performing 3*3 convolution operation on the preliminary fusion image to extract local features.
[0113] S602: performing weighted fusion of the local features and the Gram angle difference field image feature by using a ReLu activation function to obtain a fusion image feature:
[0114] I fusion =·α·ReLu(CNN 3×3 (S tf ))+(1-α)·F GADF
[0115]
[0116] Wherein, i fusion represents the fusion image feature, a represents the final fusion weight, ReLu represents the activation function, CNN 3×3 represents 3*3 convolution operation, F GADF represents the Gram angle difference field image feature, and entropy() represents the information entropy function.
[0117] In the embodiment of the application, by using convolution operation on the preliminary fusion time-frequency diagram, local energy distribution and texture features can be effectively extracted, and the expression of short-time impact and local pattern is strengthened. On this basis, the local features obtained by convolution are weighted and fused with the Gram angle difference field (GADF) features, and an entropy weight mechanism is introduced, so that the fusion proportion can be adaptively allocated according to the amount of information contained in different features, thereby avoiding the problem of one-sidedness of single feature expression.
[0118] S7: performing depth separable convolution operation on the fusion image feature to extract image depth features.
[0119] With reference to the accompanying drawings, the embodiments of the present application are illustrated. Figure 3 , which shows a deep separable convolution and ordinary convolution comparison diagram provided by an embodiment of the present application.
[0120] It should be noted that the deep feature extraction of the traditional method has the problems of excessive parameter quantity and low calculation efficiency. In view of this defect, the spatial details of the fused feature map are extracted by using the deep separable convolution in the present scheme. The deep separable convolution is individually convolved on each channel in operation, rather than cross-channel fusion. This method reduces the parameter quantity, improves the calculation efficiency, and enhances the real-time performance of the ship-bridge collision vibration event recognition and classification.
[0121] In a possible implementation, S7 specifically includes:
[0122] S701: using deep separable convolution, extracting spatial detail features of the fused feature map:
[0123] F depth = DepthwiseConv(i fusion )
[0124] wherein F depth represents the spatial detail features, DepthwiseConv represents the deep separable convolution, and I fusion represents the fused image features.
[0125] Optionally, in order to efficiently and selectively extract key spatial-channel depth features from the fused feature map, the present scheme introduces a channel and spatial attention mechanism.
[0126] The channel attention mechanism is an adaptive feature weight distribution method, and its core idea is to obtain the global expression of each channel by performing global statistics on the spatial dimension of the feature map, and then generate weight coefficients through nonlinear mapping, so as to measure the importance of different channels in the overall task. By multiplying these weights with the original features channel by channel, the channel features that contribute more to classification or recognition can be highlighted, and invalid or redundant channels can be suppressed, thereby improving the discriminability and robustness of the model.
[0127] The spatial attention mechanism is a method of paying attention to the importance of positions in the feature map, and its core idea is to compress the feature map in the channel dimension, and then generate a spatial weight matrix through convolution operation, which is used to represent the attention degree of different spatial positions. After multiplying the weight with the original feature map element by element, the model can focus more on the areas with concentrated energy and significant changes, while weakening the background or noise interference, thereby enhancing the ability to capture local area patterns.
[0128] S702: Global average pooling is performed on the spatial detail feature by using the channel attention mechanism, the pooling result is input into a multi-layer perceptron, and a channel attention weight is generated by using a Sigmoid function:
[0129] A c =·σ·(MLP(GlobalAvgPool(F depth· ))
[0130] wherein A c represents the channel attention weight, σ represents the Sigmoid function, MLP represents the multi-layer perceptron, and GlobalAvgPool represents the global average pooling.
[0131] S703: The channel attention weight is multiplied with the spatial detail feature in a channel-by-channel manner to obtain a channel weighted feature:
[0132] F c = A c ⊙F depth
[0133] wherein F c represents the channel weighted feature, ⊙ represents the channel-by-channel multiplication, and F depth represents the spatial detail feature.
[0134] S704: The channel weighted feature is respectively subjected to global average pooling and global maximum pooling, the pooled results are concatenated in the channel dimension, and then subjected to a 7×7 convolution operation, and a spatial attention weight is generated by using a Sigmoid activation function:
[0135] A s =σ(Conv 7×7 ([GlobalAvgPool(F c );GlobalMaxPool(F c )]))
[0136] wherein A s represents the spatial attention weight, σ represents the Sigmoid activation function, Conv 7×7 represents the 7×7 convolution operation, and GlobalMaxPool represents the global maximum pooling.
[0137] S705: The channel weighted feature is subjected to pointwise convolution, and the spatial attention weight is multiplied with the pointwise convolution result in an element-by-element manner to obtain an image depth feature:
[0138] F img = A s ⊙PointwiseConv(F c )
[0139] wherein F img represents the image depth feature, and PointwiseConv represents a pointwise convolution.
[0140] In the embodiments of the present application, by using a depth separable convolution on the fused image features, spatial detail features can be effectively extracted while significantly reducing the network parameter quantity and computational complexity, meeting the lightweight demand in real-time monitoring scenarios. On this basis, a channel attention mechanism is introduced, the importance of each channel is modeled through global average pooling and a multilayer perceptron, so that the model can adaptively highlight the feature channels with greater contribution to classification; at the same time, a spatial attention mechanism is combined, a spatial weight distribution is generated through global pooling and convolution operations on the feature map, which can highlight the vibration energy concentration area and suppress irrelevant backgrounds.
[0141] S8: Based on the image depth feature, a multi-scale dilated convolution is used to capture the long-time sequence dependence of the vibration event, to obtain the final sequence feature.
[0142] It should be noted that, since when the ship collides with the pier, high-frequency vibration waves will be transmitted to the bridge tower along the optical fiber, while low-frequency structural responses will spread to the bridge deck, forming a unique spatio-temporal amplitude distribution feature. Therefore, on the basis of depth feature extraction, the present scheme introduces sequence feature processing, focuses on the time dimension, captures the dynamic law and time sequence dependence of the vibration signal changing with time, and pays attention to the correlation between features at different times, such as the impact change of the vibration signal in a short time and the continuous trend in a long time, so as to more accurately identify and classify the ship-bridge collision vibration event.
[0143] wherein the multi-scale dilated convolution is a convolution modeling method for capturing long and short term dependencies, and the basic idea is to introduce different dilation rates in the convolution operation, and to expand the receptive field by inserting holes between the convolution kernel sampling points. Under the same computational load, the convolution branch with a smaller dilation rate can extract local short-time detail features, while the convolution branch with a larger dilation rate can cover longer time span global patterns, thereby realizing joint modeling of information of different time scales. Through the multi-scale parallel dilated convolution structure, short-time impact features and long-term trend features of the vibration signal can be captured at one time, and the representation ability of the model for complex dynamic events can be improved.
[0144] In one possible implementation, S8 specifically includes:
[0145] S801: Perform a multi-scale dilated convolution operation on the image depth feature, and perform layer normalization processing on the dilated convolution result to obtain preliminary sequence features under different dilation rates:
[0146]
[0147] wherein, represents the sequence feature of the expansion rate d in the lth layer, LayerNorm represents layer normalization, DilatedConv() represents a multi-scale dilated convolution, F img represents the input image feature corresponding to the expansion rate d.
[0148] S802: splicing the preliminary sequence features under different expansion rates to obtain spliced sequence features:
[0149]
[0150] wherein, represents the spliced sequence feature, represents the spliced sequence feature under the expansion rate 1 in the lth layer, represents the spliced sequence feature under the expansion rate 2 in the lth layer, represents the spliced sequence feature under the expansion rate 4 in the lth layer, represents the spliced sequence feature under the expansion rate 8 in the lth layer.
[0151] S803: performing linear mapping on the spliced sequence feature by using a first learnable gating parameter matrix to obtain dynamic gating weights:
[0152]
[0153] wherein, G represents the dynamic gating weights, σ represents a Sigmoid activation function, W g represents the first gating parameter matrix.
[0154] S804: performing convolution operation on the spliced sequence feature by using a second learnable gating parameter matrix, and performing element-wise multiplication on the activated feature and the dynamic gating weights after activation by an activation function to obtain final sequence features:
[0155]
[0156] wherein, F seq represents the final sequence feature, represents element-wise multiplication, and tanh represents a tanh activation function.
[0157] In the embodiments of the present application, by introducing a multi-scale dilated convolution structure on the basis of image depth features, the receptive field can be significantly expanded under the same amount of calculation, and the short-term impact features and long-term trend features of the vibration signal can be captured simultaneously by using parallel branches with different expansion rates, so as to realize joint modeling of long-term and short-term dependencies of the vibration event. On this basis, a gating mechanism is further used to dynamically select the spliced multi-scale features, and the learnable gating weights are used to highlight the information useful for classification and suppress irrelevant or redundant features.
[0158] S9: Multi-modal fusion is performed on the statistical features, the image deep features and the final sequence features to obtain final fusion features.
[0159] In a possible implementation, S9 specifically includes:
[0160] S901: The non-linear correlation between the statistical features and the image deep features is measured by using a mutual information technology.
[0161]
[0162] wherein MI represents mutual information, represents the statistical features of the i-th channel dimension, represents the image deep features of the j-th channel dimension, represents a double integral, log represents a logarithmic function, d represents an integral operator, and p() represents a joint probability distribution between the statistical features and the image deep features.
[0163] The mutual information (Mutual Information, MI) is an information theory index for measuring the correlation and dependence between two random variables. The essence is to quantify the amount of information shared between variables by calculating the difference between the joint distribution and the respective marginal distribution. The greater the mutual information value, the stronger the correlation between the two variables. If the mutual information is zero, the variables are independent of each other.
[0164] S902: An adaptive threshold is determined based on the mutual information between the statistical features and the image deep features.
[0165] ρ = 0.5 max (MI (F stat ,F img ))
[0166] wherein ρ represents the adaptive threshold, and max represents the maximum value.
[0167] S903: The statistical features with non-linear correlation greater than or equal to the adaptive threshold are filtered from the statistical features and the image deep features.
[0168]
[0169] wherein F′ stat represents the filtered statistical features, F stat represents the statistical features, and max j represents the maximum value in all channel dimensions j of the traversal image features.
[0170] S904: The image deep features are processed by dimension reduction through principal component analysis to obtain reduced image deep features.
[0171] F′img = PCA (F img , k = 32)
[0172] wherein F' img represents a reduced dimension image deep feature, PCA represents a dimension reduction processing by a principal component analysis method, F img represents an image deep feature, and k represents a number of principal components reserved after dimension reduction.
[0173] S905: Map the filtered statistical features, the reduced dimension image deep features and the final sequence features to the same dimension, and take the mapped statistical features, the reduced dimension image deep features and the final sequence features as a query matrix and a key matrix, and calculate corresponding self-attention scores:
[0174]
[0175] wherein a (Q, K) represents a self-attention score, Q represents a query matrix, k represents a key matrix, T represents a transposition operation, and D represents a dimension of a feature.
[0176] S906: Perform a softmax operation on the self-attention scores to obtain self-attention weights:
[0177] β = softmax (a (Q, k))
[0178] wherein β represents a self-attention weight.
[0179] S907: Based on the self-attention weights, fuse the mapped statistical features, the reduced dimension image deep features and the final sequence features to obtain final fused features:
[0180] F fused = β1·F1+ β2·F2+ β3·F3
[0181] wherein F fused represents final fused features, β1represents a weight of the mapped statistical features, F1represents the mapped statistical features, β2represents a weight of the mapped reduced dimension image deep features, F2represents the mapped reduced dimension image deep features, β3represents a weight of the mapped final sequence features, and F3represents the mapped final sequence features.
[0182] In the embodiment of the present application, by introducing a multi-modal fusion mechanism on the basis of the statistical features, the fused image features and the sequence features, the complementary advantages of different feature dimensions can be fully exerted.
[0183] S10: Based on the final fused features, output a class classification result of the vibration event through a fully connected layer.
[0184] Specifically, the fused features F fusedTransforming into fixed-length vector F global , i.e. compressing the spatial and temporal dimensions into 1. The fixed-length vector can be expressed as:
[0185] F global = GlobalAveragePooling(F fused )
[0186] wherein F global represents a global feature vector, GlobalAveragePooling represents global average pooling, and F fused represents a final fusion feature.
[0187] Batch normalization is performed on F global to obtain F bn , and the result can be expressed as:
[0188] F bn = BatchNorm(F global )
[0189] wherein F bn represents a feature vector after batch normalization, and BatchNorm represents batch normalization.
[0190] Dropout is then performed on F bn , and the dropout probability p = 0.3 is used to prevent overfitting, and the result can be expressed as:
[0191] F drop = Dropout(F bn , p = 0.3)
[0192] wherein F drop represents a feature vector after dropout, and Dropout represents the dropout operation.
[0193] Finally, through a fully connected layer, the learned weight W cls and the bias b cls are combined, and the softmax function is used for processing, to output the probability of the vibration event belonging to each category. The classification probability vector can be expressed as:
[0194]
[0195] wherein P represents the classification probability vector, softmax represents the softmax function, W cls represents the learned weight, b cls represents the bias, and N represents the number of vibration event categories.
[0196] In the embodiment of the present application, by batch normalization and dropout processing on the final fusion features, not only the global discriminant information can be preserved while reducing the feature dimension and computational complexity, but also the feature distribution deviation can be effectively suppressed and the model overfitting can be prevented, thereby improving the stability and generalization ability of the model. Further, combining the full connection layer and the softmax classifier outputs the multi-class probability distribution, so that different disturbance events can be accurately distinguished. Compared with the traditional classification method based on single feature or fixed threshold, the classification output mechanism of the present scheme significantly improves the classification accuracy and reliability of complex disturbance events such as ship collision, traffic load and wind-induced vibration, while ensuring real-time performance.
[0197] The technical scheme provided by the embodiment of the present application has at least the following beneficial effects:
[0198] In the embodiment of the present application, by collecting, denoising and multi-domain feature extraction of vibration signals, the real-time performance and comprehensiveness of bridge monitoring are improved, and by combining wavelet transform, Fourier transform, Gram angle difference field coding and multi-level feature learning of deep convolutional network, the recognition ability of damage and abnormal events is enhanced, thereby effectively overcoming the defects of low efficiency, high cost and easy omission of artificial inspection. At the same time, by using denoising processing, multi-modal feature fusion and deep convolutional structure, the stability and robustness of the monitoring result are improved, the ability to resist electromagnetic interference and environmental changes is enhanced, and finally the accurate recognition of events is realized through classification output, which meets the urgent needs of efficient, accurate and intelligent monitoring of bridges in high-risk environments.
[0199] Referring to the accompanying drawings Figure 4 , a structure schematic diagram of a vibration event classification system based on distributed optical fiber sensing provided by the present application is shown.
[0200] The present application also provides a vibration event classification system 20 based on distributed optical fiber sensing, which is applied to the vibration event classification method based on distributed optical fiber sensing described above, and comprises:
[0201] A processor 201.
[0202] A memory 202, the memory 202 stores computer readable instructions, and when the computer readable instructions are executed by the processor 201, the vibration event classification method based on distributed optical fiber sensing as in the method embodiment is realized.
[0203] The vibration event classification system 20 based on distributed optical fiber sensing provided by the present application can execute the vibration event classification method based on distributed optical fiber sensing described above, and achieve the same or similar technical effects. To avoid repetition, the present application will not be described again.
[0204] The technical scheme provided by the embodiment of the present application has at least the following beneficial effects:
[0205] In the embodiments of the present application, the real-time and comprehensiveness of bridge monitoring are improved through the collection, denoising and multi-domain feature extraction of vibration signals, the multi-level feature learning of wavelet transform, Fourier transform, Gram angle difference field coding and deep convolution network is combined, the recognition ability of damage and abnormal events is enhanced, thereby effectively overcoming the defects of low efficiency, high cost and easy to miss of artificial inspection. At the same time, the stability and robustness of the monitoring result are improved by using denoising processing, multi-modal feature fusion and deep convolution structure, the ability of resisting electromagnetic interference and environmental change is enhanced, and finally the accurate recognition of events is realized through classification output, which meets the urgent needs of efficient, accurate and intelligent monitoring of bridges in high-risk environments.
[0206] It should be understood that the processor in the embodiments of the present application can be a central processing unit (CPU), and the processor can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.
[0207] It should also be understood that the memory in the embodiments of the present application can be volatile or nonvolatile memory, or can include both volatile and nonvolatile memory. The nonvolatile memory can be read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), electrically EPROM (EEPROM), or flash memory. The volatile memory can be random access memory (RAM) used as external cache. By way of example, and not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous dynamic RAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), Synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0208] The above-described embodiments can be implemented in whole or in part by software, hardware (e.g., circuitry), firmware, or any combination thereof. When implemented in software, the above-described embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are wholly or partially generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transferred from one website, computer, server, or data center to another website, computer, server, or data center through wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible by a computer or a data storage device such as a server, data center, etc. containing one or more available medium collections. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. The semiconductor medium can be a solid-state disk.
[0209] It should be understood that the term "and / or" herein merely describes an association relationship of associated objects, which means that there can be three relationships, for example, A and / or B can represent the following three cases: A exists alone, A and B exist together, and B exists alone, where A and B can be singular or plural. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after it, but it can also represent an "and / or" relationship, which can be understood according to the context before and after it.
[0210] In the present application, "at least one" means one or more, and "multiple" means two or more. "At least one of the following" or similar expressions means any combination of these items, including any combination of single or multiple items. For example, at least one of a, b, or c can represent a, b, c, a-b, a-c, b-c, or a-b-c, where a, b, and c can be single or multiple.
[0211] It should be understood that in various embodiments of the present application, the size of the sequence number of each process described above does not mean the order of execution, and the execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0212] Those skilled in the art can clearly understand that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0213] Those skilled in the art can clearly understand that, for the convenience and brevity of the description, the specific working processes of the devices, apparatuses and units described above can refer to the corresponding processes in the foregoing method embodiments, which will not be repeated here.
[0214] In several embodiments provided by the present application, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely schematic, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0215] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.
[0216] In addition, each functional unit in each embodiment of the present application can be integrated into a processing unit, or each unit can exist physically independently, or two or more units can be integrated into one unit.
[0217] If the functions are realized in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the parts of the prior art that make contributions or parts of the technical solutions can be embodied in the form of software products. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the various embodiments of the present application. The aforementioned storage medium includes a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0218] The embodiment of the present application provides a computer readable storage medium, which stores a computer program. The program is executed by a processor to realize the vibration event classification method based on distributed optical fiber sensing as described in the method embodiment.
[0219] The computer readable storage medium provided by the present application can realize the steps and effects of the vibration event classification method based on distributed optical fiber sensing of the above-mentioned method embodiment. To avoid repetition, the present application will not be described again.
[0220] The technical solutions provided by the embodiment of the present application have at least the following beneficial effects:
[0221] In the embodiment of the present application, through the collection, denoising and multi-domain feature extraction of the vibration signal, the real-time and comprehensiveness of the bridge monitoring are improved. Combined with the multi-level feature learning of wavelet transform, Fourier transform, Gram angle difference field coding and deep convolution network, the recognition ability of damage and abnormal events is enhanced, so as to effectively overcome the defects of low efficiency, high cost and easy omission of artificial inspection. At the same time, by using denoising processing, multi-modal feature fusion and deep convolution structure, the stability and robustness of the monitoring result are improved, the ability of resisting electromagnetic interference and environmental change is enhanced, and finally the accurate recognition of events is realized through classification output, which meets the urgent needs of efficient, accurate and intelligent monitoring of bridges in high-risk environments.
[0222] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical range disclosed by the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
[0223] The following points need to be explained:
[0224] (1) The drawings of the embodiments of the present application only relate to the structures involved in the embodiments of the present application, and other structures can be referred to the general design.
[0225] (2) In the drawings used to describe the embodiments of the present application, the thickness of a layer or region is exaggerated or reduced for clarity, i.e., the drawings are not drawn according to the actual scale. It can be understood that when an element such as a layer, film, region or substrate is referred to as being "on" or "under" another element, it can be "directly" on or under the other element or there can be an intermediate element.
[0226] (3) In the case of no conflict, the embodiments of the present application and the features in the embodiments can be combined with each other to obtain new embodiments.
[0227] The above is only a specific embodiment of the present application, but the protection scope of the present application is not limited thereto, and the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method of classifying vibration events based on distributed fiber optic sensing, the method comprising: Comprise: S1: collect vibration signals; S2: denoising processing is carried out to the vibration signal, and denoising vibration signal is obtained; S3: time domain feature, frequency domain feature and audio domain feature are extracted from the denoising vibration signal, and statistical feature is formed; S4: the denoising vibration signal is respectively extracted by continuous wavelet transform and short-time Fourier transform, and the extraction result is fused based on entropy weight mechanism, and preliminary fusion feature map is obtained; S5: the denoising vibration signal is extracted based on the principle of Gram angle difference field coding, and Gram angle difference field image feature is obtained; S6: the preliminary fusion feature map is subjected to convolution operation, local feature is extracted, and the local feature and the Gram angle difference field image feature are weighted and fused, and fusion image feature is obtained; S7: the fusion image feature is subjected to depth separable convolution operation, and image depth feature is extracted; S8: based on the image depth feature, a multi-scale dilated convolution is used to capture the long-time sequence dependence of the vibration event, and a final sequence feature is obtained; S9: the statistical feature, the image depth feature and the final sequence feature are subjected to multi-modal fusion, and a final fusion feature is obtained; S10: based on the final fusion feature, a class classification result of the vibration event is output through a fully connected layer.
2. The distributed fiber optic sensing based vibration event classification method of claim 1, wherein, The S2 specifically comprises: S201: the vibration signal is decomposed using a wavelet basis function, and low-frequency components and high-frequency components are extracted: [cA, cD] = DWT (x (t) ; Symlet-5) ; Wherein, cA represents the low-frequency component, cD represents the high-frequency component, DWT represents the discrete wavelet transform, x (t) represents the vibration signal, and Symlet-5 represents the Symlet-5 wavelet basis function; S202: the kurtosis value of the high-frequency component is calculated, and the kurtosis value is normalized; S203: based on the normalized kurtosis value, an adaptive threshold is calculated: Wherein, λ represents the adaptive threshold, σ represents the noise standard deviation, N represents the signal length, ln represents the logarithmic function with e as the base, tanh represents the hyperbolic tangent function, and kurtosis (cD) represents the high-frequency component of the high-frequency component; S204: the high-frequency component is thresholded, and the high-frequency component is retained, to obtain a denoising high-frequency component: cD'= sign (cD) · max (|cD|- λ, 0) ; Wherein, cD' represents the denoising high-frequency component, sign () represents the sign function, max () represents the maximum function, and || represents the absolute value; S205: the low-frequency component and the denoising high-frequency component are subjected to inverse discrete wavelet transform, to obtain a reconstructed signal, and the reconstructed signal is subjected to morphological filtering, to obtain the denoising vibration signal: x denoised (t) = γ close (γ open (IDWT(cA,cD'))); where x denoised (t) denotes the denoised vibration signal, γ close denotes a close operation, γ open denotes an open operation, and IDWT denotes an inverse discrete wavelet transform.
3. The distributed fiber optic sensing based vibration event classification method of claim 1, wherein, The S3 specifically comprises: S301: the time domain feature including mean suos, variance, kurtosis, average absolute value and zero-crossing rate in the denoising vibration signal is extracted; S302: the frequency domain feature including 0-50Hz energy ratio, peak frequency and energy second derivative at characteristic frequency in the denoising vibration signal is extracted; S303: extracting audio domain features of the denoised vibration signal, including mel frequency cepstral coefficient, spectral center and spectral change rate; S304: combining the time domain features, the frequency domain features and the audio domain features to form the statistical features: where F stat denotes statistical features, μ denotes mean, σ 2 denotes variance, k denotes kurtosis, denotes mean absolute value, ZCR denotes zero-crossing rate, denotes energy ratio of 0-50Hz, f max denotes peak frequency, denotes energy second derivative at feature frequency, MFCC 1:5 denotes first 5 mel-frequency cepstral coefficients, SC denotes spectral center, Δ spec denotes spectral change rate.
4. The distributed fiber optic sensing based vibration event classification method of claim 1, wherein, The S4 specifically comprises: S401: performing continuous wavelet transform on the denoised vibration signal to obtain a continuous wavelet time-frequency graph: where S cwt represents a continuous wavelet time-frequency plot, τ represents a translation parameter, s represents a scale parameter, x denoised (t) represents a denoised vibration signal, ψ represents a wavelet function, and ∫ represents an integral sign; S402: performing short-time Fourier transform on the denoised vibration signal in the manner of Hanning window to obtain a short-time Fourier time-frequency graph: S stft (t) = ∫x denoised (t) · w(t - τ) · e -2πjft dt; where S stft represents a short-time Fourier time-frequency diagram, w() represents a window function, e -2πjft represents a complex sinusoidal wave; S403: calculating an adaptive entropy weight: Wherein, θ represents the adaptive entropy weight, and entropy() represents an information entropy function; S404: weighting and fusing the continuous wavelet time-frequency graph and the short-time Fourier time-frequency graph based on the adaptive entropy weight to obtain a preliminary fusion feature graph: S tf = · θ · S cwt + (1 - θ) · S stft ; where S tf represents the preliminary fusion feature map, S cwt represents the continuous wavelet time-frequency map, S stft represents the short-time Fourier time-frequency map.
5. The distributed fiber optic sensing based vibration event classification method of claim 1, wherein, The S5 specifically comprises: S501: representing the denoised vibration signal as one-dimensional time series data: X = [x1, x2,..., x T ]; wherein X represents a one-dimensional time series data, x u represents the u-th spatial point, u = 1, 2, …, T, and T represents the total length of the time series. S502: performing normalization processing on the one-dimensional time series data to obtain a normalized sequence: wherein represents a normalized sequence, min( ) represents a value of obtaining the minimum element in an array, and max( ) represents a value of obtaining the maximum element in an array. S503: mapping the normalized sequence to a polar coordinate system to calculate a polar angle and a radius in the polar coordinate system: wherein denotes the polar angle in polar coordinates, arccos denotes the inverse cosine function, r i denotes the radius in a coordinate system; S504: calculating an angle difference relationship of all time point pairs according to the angle and the radius of the polar coordinate, and constructing a Gram angle difference field matrix: wherein GADF(i,j) represents the relationship between time point i and time point j in the normalized time series, cos represents the cosine function, r j represents the radius of time point j in the polar coordinate system, represents the polar angle of time point j in the polar coordinate system; S505: taking the Gram angle difference field matrix as the Gram angle difference field image feature output.
6. The distributed fiber optic sensing based vibration event classification method of claim 1, wherein, The S6 specifically comprises: S601: performing a 3*3 convolution operation on the preliminary fusion graph to extract local features; S602: weighting and fusing the local features and the Gram angle difference field image feature through a ReLu activation function to obtain the fusion image feature: I fusion = a ReLu(CNN 3×3 (S tf ) + (1 - a) F GADF ; where I fusion denotes the fused image feature, a denotes the final fusion weight, ReLu denotes the activation function, CNN 3×3 denotes the 3x3 convolution operation, F GADF denotes the Gram angle difference field image feature, and entropy() denotes the information entropy function.
7. The distributed fiber optic sensing based vibration event classification method of claim 1, wherein, The S7 specifically comprises: S701: using a depth separable convolution to extract spatial detail features of the fusion image feature: F depth = DepthwiseConv(I fusion ); wherein F depth represents a spatial detail feature, DepthwiseConv represents a depthwise separable convolution, I fusion represents a fused image feature; S702: using a channel attention mechanism to perform global average pooling on the spatial detail features, inputting a pooling result into a multi-layer perceptron, and using a Sigmoid function to generate a channel attention weight: A c = σ(MLP(GlobalAvgPool(F depth· ))) ; where A c denotes the channel attention weight, denotes the Sigmoid function, MLP denotes a multi-layer perceptron, and GlobalAvgPool denotes a global average pooling. S703: multiplying the channel attention weight and the spatial detail features channel by channel to obtain a channel weighted feature: F c = A c O F depth ; where F c denotes the channel weighting feature, denotes the element-wise multiplication, F depth denotes the spatial detail feature; S704: performing global average pooling and global maximum pooling on the channel weighted feature respectively, performing a 7*7 convolution operation on a channel dimension spliced pooling result, and using a Sigmoid activation function to generate a spatial attention weight: A s = σ(Conv 7×7 ([GlobalAvgPool(F c ) ; GlobalMaxPool(F c )])) ; wherein A s denotes the spatial attention weight, σ denotes the Sigmoid activation function, Conv 7×7 denotes a 7x7 convolution operation, and GlobalMaxPool denotes a global max pooling; S705: performing point-by-point convolution on the channel weighted feature, and multiplying the spatial attention weight and the point-by-point convolution result element by element to obtain the image depth feature: F img = A s ⊙PointwiseConv(F c ); where F img denotes the image depth feature, and PointwiseConv denotes a pointwise convolution.
8. The distributed fiber optic sensing based vibration event classification method of claim 1, wherein, The S8 specifically comprises: S801: performing a multi-scale dilated convolution operation on the image depth feature, and performing layer normalization processing on a dilated convolution result to obtain preliminary sequence features under different dilation rates: wherein, represents the sequence feature of the l-th layer with dilation rate d, LayerNorm represents layer normalization, DilatedConv() represents a multi-scale dilated convolution, F img , d represents the input image feature corresponding to the dilation rate d. S802: splicing the preliminary sequence features under different dilation rates to obtain a spliced sequence feature: wherein, represents the splicing sequence feature, represents the splicing sequence feature at the inflation rate of 1 at the lth layer, represents the splicing sequence feature at the inflation rate of 2 at the lth layer, represents the splicing sequence feature at the inflation rate of 4 at the lth layer, represents the splicing sequence feature at the inflation rate of 8 at the lth layer; S803: using a first learnable gating parameter matrix to linearly map the spliced sequence feature to obtain a dynamic gating weight: where G represents a dynamic gating weight, σ represents a sigmoid activation function, W g represents a first gating parameter matrix; S804: Convolution operation is performed on the splicing sequence feature by using a learnable second gating parameter matrix, and after being activated by an activation function, the splicing sequence feature is element-wise multiplied by a dynamic gating weight to obtain the final sequence feature: where F seq denotes the final sequence feature, denotes element-wise multiplication, and tanh denotes the tanh activation function.
9. The distributed optical fiber sensing based vibration event classification method of claim 1, wherein, The S9 specifically includes: S901: The non-linear correlation between the statistical features and the image deep features is measured by using mutual information technology: where MI denotes mutual information, denotes the statistical feature of the i-th channel dimension, denotes the image deep feature of the j-th channel dimension, denotes double integral, log denotes logarithm function, d denotes integral operator, and p() denotes the joint probability distribution between the statistical feature and the image deep feature. S902: Based on the mutual information result between the statistical features and the image deep features, an adaptive threshold is determined: p = 0.5 max(MI(F stat , F img )); Wherein, ρ represents the adaptive threshold, and max represents the maximum value; S903: The statistical features with non-linear correlation greater than or equal to the adaptive threshold with any dimension of the image deep features are filtered: where F' = F - F stat represents the filtered statistical feature, F stat represents the statistical feature, max j represents the maximum value in all channel dimensions j of the traversed image feature; S904: The image deep features are processed by dimension reduction by principal component analysis to obtain reduced image deep features: F′ img = PCA(F img , k = 32); wherein F' = F - F0 img represents a deep feature of a reduced dimension image, PCA represents a dimension reduction process by principal component analysis, F img represents a deep feature of an image, and k represents the number of principal components retained after dimension reduction. S905: The filtered statistical features, the reduced image deep features and the final sequence features are mapped to the same dimension, and the mapped statistical features, reduced image deep features and final sequence features are taken as query matrix and key matrix to calculate corresponding self-attention score: Wherein, a(Q, K) represents the self-attention score, Q represents the query matrix, K represents the key matrix, T represents the transposition operation, and D represents the dimension of the feature; S906: The self-attention score is subjected to softmax operation to obtain self-attention weight: β=softmax(a(Q,K)); Wherein, β represents the self-attention weight; S907: Based on the self-attention weight, the mapped statistical features, image deep features and final sequence features are fused to obtain final fusion features: F fused = β1·F1+ β2·F2+ β3·F3; wherein F fused represents the final fusion feature, β1 represents the weight of the mapped statistical feature, F1 represents the mapped statistical feature, β2 represents the weight of the mapped deep image feature, F2 represents the mapped deep image feature, β3 represents the weight of the mapped final sequence feature, and F3 represents the mapped final sequence feature.
10. A distributed fiber optic sensing based vibration event classification system, characterized in that, It includes: A processor; A memory, the memory has computer readable instructions stored thereon, and the computer readable instructions are executed by the processor to implement the vibration event classification method based on distributed optical fiber sensing according to any one of claims 1 to 9.
Citation Information
Patent Citations
Rolling bearing fault diagnosis method based on FFT coding and L-CNN
CN117473872A
Bearing fault diagnosis method based on dual-channel multi-scale attention feature fusion
CN118918424A
Unbalanced small sample event rapid classification method and system suitable for phi-OTDR (Optical Time Domain Reflectometer) system
CN119691588A
Abnormal driving behavior identification method based on multi-modal sensor data fusion
CN120279532A
Aerospace bearing fault diagnosis method based on GADF-driven KAN-Swin Transformer double-branch network
CN120404149A
Cited By
Road intersection vulnerable group type classification method and system facing road surface disturbance
CN121786596A
Road intersection vulnerable group type classification method and system facing road surface disturbance
CN121786596B