An edge recognition method and device for structural vibration events based on DWT-CNN multi-modal spectrum fusion
Patent Information
- Application Number
- CN202611311759.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-27
- Publication Date
- 2026-09-29
AI Technical Summary
[0006]为解决现有边缘监测装备依赖固定阈值触发而车辆振动误报率高、单方向特征表达不足以及一维信号分类难以充分利用三轴频域特征的问题,本发明提供了一种基于DWT-CNN多模态频谱融合的结构振动事件边缘识别方法及装置
第一,通过离散小波变换对三轴振动信号进行多尺度分解,并采用保留目标频段系数、其余系数置零后经离散小波逆变换重构频段信号的方式,能够将原始复杂振动信号拆解为具有明确频段含义的多尺度重构信号。由于地震信号通常在较低频段能量较为集中,车辆振动常在中高频范围出现局部增强,而环境噪声多表现为低幅值或随机频带分布,该方法有效增强了三类振动事件在时频域上的可区分性,为后续识别提供了特征差异化的频段数据基础。
Smart Images

Figure CN122839299A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the fields of structural health monitoring, intelligent sensing devices, edge computing and engineering digital twin technology, and particularly relates to a method and device for edge recognition of structural vibration events based on DWT-CNN multimodal spectrum fusion. Background Technology
[0002] In engineering structural health monitoring systems, acceleration and vibration sensors such as strong-motion seismometers, triaxial accelerometers, and MEMS accelerometers are typically used in conjunction with edge computers, communication modules, and monitoring platforms to collect vibration responses from structural foundations, dam galleries, bridge piers, tunnel linings, and other components. Edge computers typically need to determine near the sensors whether to initiate seismic event reporting, structural dynamic response analysis, damage accumulation calculation, and digital twin linkage processes.
[0003] In real-world engineering environments, vibration sensors are often deployed near roads, construction areas, around electromechanical equipment, or in complex terrain. The acquired signals are easily affected by heavy vehicles, construction disturbances, unit operation, and environmental noise. If edge devices rely solely on fixed peak threshold triggers, vehicle vibrations and environmental interference may be frequently misinterpreted as seismic events, leading to false start-ups of the monitoring system, communication link occupancy, redundant cloud storage, wasted dynamic computing resources, and inaccurate structural condition assessment results.
[0004] Traditional thresholding methods and single-frequency-domain feature methods rely on manually set parameters, making it difficult to simultaneously adapt to the differences in time scale, frequency components, energy distribution, and triaxial coupling relationships among seismic waves, vehicle vibrations, and environmental noise. While directly inputting one-dimensional time-history signals into neural networks can enable end-to-end classification, it is prone to problems such as insufficient feature representation, low efficiency in utilizing training samples, insufficient model generalization ability, and low efficiency in edge deployment in scenarios with limited sample size, non-stationary signals, significant differences in directional response, and limited computing power on the edge side.
[0005] Therefore, there is an urgent need for a method and device for edge recognition of structural vibration events that can utilize triaxial data collected by acceleration vibration sensors to complete DWT-IDWT band reconstruction, PSD data generation, multimodal spectrum fusion, and CNN recognition on an edge computing device. Summary of the Invention
[0006] To address the problems of existing edge monitoring equipment relying on fixed threshold triggering, resulting in high false alarm rates for vehicle vibration, insufficient unidirectional feature representation, and difficulty in fully utilizing three-axis frequency domain features in one-dimensional signal classification, this invention provides a method and device for structural vibration event edge recognition based on DWT-CNN multimodal spectrum fusion.
[0007] This invention provides a method for edge recognition of structural vibration events based on DWT-CNN multimodal spectral fusion, comprising: Receives multi-directional vibration time history signals of engineering structures collected by acceleration vibration sensors; Based on the multi-directional vibration time history signal, preprocessed sample segments in each direction are obtained; Perform discrete wavelet transform on sample segments in each direction to obtain multi-scale wavelet coefficients in each direction. Based on the preset frequency band, retain the target wavelet coefficients and set the remaining wavelet coefficients to zero. Reconstruct the zeroed wavelet coefficients into time domain signals of the corresponding frequency bands through inverse discrete wavelet transform to obtain the reconstructed frequency band signals in each direction. Power spectral density is estimated based on the reconstructed signals in each frequency band, power spectral density spectrum data in each direction is obtained, and power spectral density spectrum images in each direction are generated based on the power spectral density spectrum data. The power spectral density spectrum images in each direction are mapped to different color channels, and weighted fusion is performed according to the channel weights of each color channel to obtain a multimodal spectrum image; The multimodal spectral image is input into a convolutional neural network recognition model to obtain the category probabilities of earthquake signals, vehicle vibrations, and environmental noise, and the category of structural vibration events is determined based on the category probabilities.
[0008] Optionally, based on the multi-directional vibration time history signal, preprocessed sample segments in each direction are obtained, specifically including: Vibration time history signals in multiple directions are continuously acquired according to a preset sampling frequency, and sample segments are divided by a preset time window. Linear detrending and constant mean removal processes are sequentially performed on the sample segments in each direction to remove the linear trend term and constant offset of the sample segments in each direction. The sample segments in each direction after detrending and mean removal are zeroed out to obtain preprocessed sample segments in each direction.
[0009] Optionally, a discrete wavelet transform is performed on the sample segments in each direction to obtain multi-scale wavelet coefficients in each direction. The target wavelet coefficients are retained according to a preset frequency band, and the remaining wavelet coefficients are set to zero. The zeroed wavelet coefficients are then reconstructed into time-domain signals of the corresponding frequency bands using an inverse discrete wavelet transform to obtain the reconstructed frequency band signals in each direction. Specifically, this includes: A multi-level discrete wavelet decomposition of sample segments in each direction is performed using a preset wavelet basis function to obtain low-frequency approximation coefficients and multi-level high-frequency detail coefficients. The wavelet coefficient levels to be retained are determined based on the target frequency band. The low-frequency approximation coefficients or high-frequency detail coefficients of the wavelet coefficient levels to be retained are taken as the target wavelet coefficients, and all wavelet coefficients of the other levels except the target wavelet coefficients are set to zero. The target wavelet coefficients and the remaining level wavelet coefficients after being set to zero are input into the discrete wavelet inverse transform to reconstruct the time domain signal of the corresponding target frequency band, and the reconstructed signal of each direction frequency band is obtained.
[0010] Optionally, power spectral density estimation is performed based on the reconstructed signals in each frequency band to obtain power spectral density spectrum data in each direction, and power spectral density spectrum images in each direction are generated based on the power spectral density spectrum data, specifically including: Welch method is used to estimate the power spectral density of the reconstructed signal in each frequency band and obtain the power spectral density spectrum data in each direction. The power spectral density values in the power spectral density spectrum data of each direction are normalized to the pixel intensity range, and the normalized spectrum data is plotted as a filled two-dimensional power spectral density spectrum image.
[0011] Optionally, the power spectral density spectrum images in each direction are mapped to different color channels, and weighted fusion is performed according to the channel weights of each color channel to obtain a multimodal spectrum image, specifically including: Map the east-west power spectral density spectrum image to the red channel, the north-south power spectral density spectrum image to the green channel, and the vertical power spectral density spectrum image to the blue channel; The three color channels are weighted and summed according to the preset weights for the red, green, and blue channels to obtain a multimodal spectrum image.
[0012] Optionally, after determining the structural vibration event category based on the category probability, the method further includes: When the structural vibration event is classified as a seismic signal, the control edge computing device uploads the original vibration segment, power spectral density spectrum data, multimodal spectrum image and recognition result, and triggers the digital twin system to perform structural dynamic response calculation; When the structural vibration event is classified as vehicle vibration or environmental noise, the control edge computing device records the event log and suppresses false triggering of earthquake events.
[0013] Optionally, the convolutional neural network recognition model includes five sequentially connected convolutional blocks, one flattening layer, two fully connected layers, and one Softmax classification layer. Each convolutional block contains a convolutional layer and a max pooling layer. The convolutional layer is used to extract local convolutional features from the multimodal spectral image, and the max pooling layer is used to preserve local salient responses and reduce the feature map size.
[0014] The present invention also provides a structural vibration event edge recognition device, comprising: Accelerometer vibration sensor, data acquisition interface, edge computer, memory and communication module; The acceleration vibration sensor is used to collect time history signals of multi-directional vibration of engineering structures; The data acquisition interface is connected to the acceleration vibration sensor and the edge computer, and is used to transmit the multi-directional vibration time history signal to the edge computer. The memory stores computer programs and convolutional neural network recognition models; The edge computer is connected to the memory and the communication module to execute the computer program to implement the method, and uploads event data or triggers digital twin linkage calculations through the communication module according to the identification results.
[0015] Optionally, the edge computer includes: The data receiving module is used to receive the multi-directional vibration time history signal transmitted by the data acquisition interface; The preprocessing module is used to perform detrending, zeroing, and fixed-length segmentation on the multi-directional vibration time history signal to obtain sample segments in each direction; The Discrete Wavelet Transform and Reconstruction Module is used to perform Discrete Wavelet Transform and Inverse Discrete Wavelet Transform on sample segments in each direction to obtain reconstructed signals in each frequency band. The power spectral density generation module is used to estimate the power spectral density of the reconstructed signal in each frequency band and generate power spectral density spectrum data and power spectral density spectrum images in each direction. The weighted fusion module is used to map the power spectral density spectrum images of each direction to different color channels and fuse them according to channel weights to generate a multimodal spectrum image; A convolutional neural network recognition module is used to deploy a convolutional neural network recognition model to recognize the multimodal spectral image and output the category probabilities of seismic signals, vehicle vibrations, and environmental noise. The control module is used to control event uploading, digital twin triggering, or false triggering suppression based on the probability of the category.
[0016] Optionally, the acceleration vibration sensor is at least one of a strong vibration meter, a triaxial accelerometer, or a MEMS acceleration sensor; The communication module is at least one of an Ethernet module, a 4G communication module, a 5G communication module, an optical fiber communication module, or a wireless local area network module.
[0017] Compared with the prior art, the present invention has the following advantages and technical effects: First, by performing multi-scale decomposition of the triaxial vibration signal using discrete wavelet transform, and reconstructing the frequency band signal by retaining the target frequency band coefficients and setting the remaining coefficients to zero, the original complex vibration signal can be decomposed into multi-scale reconstructed signals with clear frequency band meanings. Since seismic signals typically have concentrated energy at lower frequencies, vehicle vibrations often exhibit localized enhancement in the mid-to-high frequency range, and environmental noise is mostly characterized by low amplitude or random frequency band distribution, this method effectively enhances the distinguishability of the three types of vibration events in the time-frequency domain, providing a frequency band data foundation with differentiated characteristics for subsequent identification.
[0018] Second, by estimating the power spectral density of the reconstructed frequency band signal to generate power spectral density spectrum data, and encoding the power spectral density spectrum data into a two-dimensional spectrum image, the one-dimensional time series is converted into a two-dimensional texture image, enabling the convolutional neural network to effectively learn the spectral energy distribution features through local convolutional kernels, thus avoiding the problems of insufficient feature expression and low efficiency of training sample utilization when directly performing end-to-end classification of one-dimensional time history signals.
[0019] Third, the power spectral density spectrum images in the east-west, north-south, and vertical directions are mapped to the red, green, and blue channels respectively, and then fused into a multimodal spectrum image by weighting according to the channel weights. While preserving the complementary relationship and directional coupling information between the three-axis signals, the frequency domain features of the three directions are compressed into a single image, reducing the data redundancy caused by inputting the three-directional images separately, reducing the inference computation load of the edge computing device, and improving the real-time performance of edge deployment.
[0020] Fourth, by deploying the convolutional neural network recognition model on edge computing devices, vibration event identification and category determination are completed near the sensors, enabling edge devices to autonomously decide whether to report earthquake events based on the identification results. When an earthquake signal is identified, digital twin linkage calculations are triggered; when vehicle vibrations or environmental noise are identified, false triggers are suppressed. This effectively reduces communication link occupation, cloud storage redundancy, and waste of digital twin computing resources caused by false triggers, significantly improving the operating efficiency of the structural health monitoring system. Attached Figure Description
[0021] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a schematic diagram of the method flow according to an embodiment of the present invention; Figure 2 This is a schematic diagram illustrating the generation of PSD data in the DWT and IDWT frequency bands according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the weighted fusion of the RGB channels of the three-axis PSD spectrum according to an embodiment of the present invention; Figure 4 This is a flowchart of the CNN model PSD fusion spectrum recognition process according to an embodiment of the present invention; Figure 5 This is a schematic diagram of the structural vibration event edge recognition device according to an embodiment of the present invention. Detailed Implementation
[0022] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.
[0023] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.
[0024] Example 1 like Figure 1 As shown, this embodiment provides a structural vibration event edge recognition method based on DWT-CNN multimodal spectral fusion, including: Receives multi-directional vibration time history signals of engineering structures collected by acceleration vibration sensors; Based on the multi-directional vibration time history signal, preprocessed sample segments in each direction are obtained; Perform discrete wavelet transform on sample segments in each direction to obtain multi-scale wavelet coefficients in each direction. Based on the preset frequency band, retain the target wavelet coefficients and set the remaining wavelet coefficients to zero. Reconstruct the zeroed wavelet coefficients into time domain signals of the corresponding frequency bands through inverse discrete wavelet transform to obtain the reconstructed frequency band signals in each direction. Power spectral density is estimated based on the reconstructed signals in each frequency band, power spectral density spectrum data in each direction is obtained, and power spectral density spectrum images in each direction are generated based on the power spectral density spectrum data. The power spectral density spectrum images in each direction are mapped to different color channels, and weighted fusion is performed according to the channel weights of each color channel to obtain a multimodal spectrum image; The multimodal spectral image is input into a convolutional neural network recognition model to obtain the category probabilities of earthquake signals, vehicle vibrations, and environmental noise, and the category of structural vibration events is determined based on the category probabilities.
[0025] As a feasible implementation method, the specific steps include: like Figure 1As shown, the structural vibration event edge recognition method includes steps such as triaxial vibration time history acquisition, signal preprocessing, DWT and IDWT frequency band reconstruction, PSD data generation, triaxial PSD spectrum RGB channel weighted fusion, CNN model edge recognition, and event uploading or false triggering suppression.
[0026] Specifically, the system receives multi-directional vibration time-history signals of the engineering structure collected by the acceleration vibration sensor; performs detrending, zeroing, and fixed-length segmentation on the multi-directional vibration time-history signals to obtain sample segments for each direction; performs discrete wavelet transform on the sample segments for each direction to obtain low-frequency approximation coefficients and multi-layer high-frequency detail coefficients; retains at least one set of wavelet coefficients according to a preset frequency band and sets the remaining wavelet coefficients to zero, and reconstructs the frequency band signal through inverse discrete wavelet transform; performs power spectral density estimation on the frequency band signal to generate PSD spectrum data and PSD spectrum images for the corresponding directions; maps the PSD spectrum images of multiple directions to different color channels and fuses them into a multimodal spectrum image according to the channel weights; inputs the multimodal spectrum image into a CNN recognition model deployed on the edge computing device, outputs the recognition probabilities of seismic signals, vehicle vibrations, and environmental noise, and determines the structural vibration event category based on the recognition probabilities.
[0027] The structural vibration event edge identification device may include an acceleration vibration sensor, a data acquisition interface, an edge computer, a memory, a processor, and a communication module. The acceleration vibration sensor may be a strong vibration meter, a triaxial accelerometer, or a MEMS accelerometer, and may be deployed at the foundation of the engineering structure, dam gallery, key parts of the bridge, tunnel lining, or other structural parts that require monitoring.
[0028] The edge computer receives triaxial vibration time history signals via a wired acquisition interface, a wireless transmission interface, or a data acquisition card, and saves them in a local cache according to the sensor number, sampling timestamp, and direction identifier. The three axes can be denoted as East-West (EW), North-South (NS), and Vertical (UD). The system continuously samples at a preset sampling frequency fs and segments the samples into time windows T; the time window for the fixed-length segments is 30 to 120 seconds, preferably 60 seconds; the sampling frequency is 50 Hz to 200 Hz, preferably 100 Hz; the detrending process includes linear detrending and constant value mean removal processing.
[0029] The acceleration vibration sensor includes at least one of a strong vibration meter, a triaxial accelerometer, or a MEMS acceleration sensor, and the multi-directional vibration time history signal includes east-west, north-south, and vertical triaxial vibration signals.
[0030] Sample segments in each direction can be represented as: The edge computer sequentially performs linear detrending, constant mean reduction, and zeroing processing on sample segments from each direction to reduce the impact of sensor drift, zero-point offset, and low-frequency baseline terms on frequency domain analysis. The preprocessed sample can be represented as: ; In the formula, X d Indicates direction d The set of sample fragments on; d Indicates the direction of vibration. d You can choose an east-west orientation (EW), a north-south orientation (NS), or a vertical orientation (UD); x d [ n [Indicates direction] d Upper n The acceleration vibration amplitude at each sampling point; n Indicates the sampling point number; N This represents the total number of sampling points within a time window.
[0031] ; In the formula, Indicates direction d Upper n Vibration amplitude values at each sampling point after preprocessing; This indicates the vibration amplitude before preprocessing; and They represent directions respectively. d The slope and intercept of the linear trend term on the graph; Indicates direction d The mean of the sample segment; n Indicates the sampling point number; N This represents the total number of sampling points within a time window.
[0032] like Figure 2 As shown, a discrete wavelet transform is performed on each preprocessed directional signal using wavelet basis functions to obtain low-frequency approximation coefficients and multiple high-frequency detail coefficients. In a preferred embodiment, a 5-level decomposition is performed using the db6 wavelet basis to obtain a set of low-frequency approximation coefficients and five sets of high-frequency detail coefficients; the frequency band signal includes at least one time-domain reconstructed component obtained by inverse discrete wavelet transform of the low-frequency approximation coefficients or high-frequency detail coefficients. The discrete wavelet coefficients can be expressed as: To extract independent features from different frequency bands, the edge computer retains the target wavelet coefficients and sets the remaining wavelet coefficients to zero, constructing a target frequency band coefficient set. Then, it uses inverse discrete wavelet transform to reconstruct the time-domain components of the corresponding frequency bands. In this way, the original complex vibration signal is decomposed into a multi-scale reconstructed signal with clear frequency band meaning. Seismic signals usually have more concentrated energy in the lower frequency band, vehicle vibrations often show local enhancement in the mid-to-high frequency range, and environmental noise often exhibits low amplitude or random frequency band distribution.
[0033] The Welch method is used to estimate the power spectral density of the reconstructed frequency band signal, generating frequency band PSD data. The power spectral density curve is then plotted as a filled two-dimensional PSD spectrum image. The power spectral density can be represented by the Fourier transform of the autocorrelation function: During image encoding, the power spectral density values are normalized to the pixel intensity range, and the PSD spectrum image is scaled or cropped to 240×360 pixels to meet the input size requirements of the subsequent CNN recognition model.
[0034] ; In the formula, W x ( j , k ) indicates a signal x At scale layer number j Translation position k Wavelet coefficients at; This represents the preprocessed discrete vibration signal; n Indicates the sampling point number; ψ j,k * [ n ] represents wavelet basis functions ψ j,k [ n The conjugate function of ]; ∑ represents the sampling point index. n Sum.
[0035] ; In the formula, ψ j,k [ n ] indicates that it is composed of the mother wavelet ψ (·) Discrete wavelet basis functions obtained by scaling and translation; ψ (·) denotes the mother wavelet function; j Indicates the scale layer number; k Indicates the translation position; n Indicates the sampling point number; 2 -j / 2 2 is the scale normalization factor; -j n - kThis represents the independent variable of the function after scaling and translation transformations.
[0036] ; In the formula, x m [ n ] indicates the first m The time-domain components of each target frequency band are reconstructed using IDWT; IDWT(·) represents the inverse discrete wavelet transform. c 0,m Indicates the first m The set of low-frequency approximation coefficients corresponding to each target frequency band; c 1,m to c L,m They represent the first to the second level, respectively. L Layer detail factor at the 1st m The set of coefficients after constructing each target frequency band; L Indicates the wavelet decomposition level; m Indicates the target frequency band number; n Indicates the sampling point number.
[0037] ; In the formula, c i,m Indicates the first m During the construction of the first target frequency band, the first i Layer wavelet coefficients; c i The first value obtained from DWT decomposition is... i Layer original wavelet coefficients; i Indicates the wavelet coefficient layer number; m Indicates the target frequency band layer number that is reserved; when i = m When retaining the original coefficients c i ,when i Not equal to m The corresponding coefficient will be set to zero.
[0038] ; In the formula, R x ( t ) indicates a signal x The autocorrelation function; t E[·] represents the time delay; E[·] represents the expected value. x ( t )express t The amplitude of the continuous vibration signal at any given moment; x ( t + t ) indicates delay tThe amplitude of the vibration signal afterward; t Represents a time variable.
[0039] ; In the formula, S x ( f ) indicates a signal x In frequency f Power spectral density at; R x ( t ) indicates a signal x The autocorrelation function; t Indicates a time delay; f The frequency is represented by 'e'; 'e' represents the base of the natural exponential function; in the formula... j Represents the imaginary unit, and is related to the wavelet scale layer number. j Different meanings; d t Indicates time delay t integral.
[0040] ; In the formula, P ( u , v ) represents the pixel coordinates in the PSD spectral image. u , v The grayscale or channel intensity value at () u Represents the horizontal pixel coordinates of the image; v Represents the vertical pixel coordinates of the image; S ( f v ) indicates frequency f v The power spectral density value at that location; f v Indicates the relationship with the first v The frequency corresponding to each vertical pixel position; S min and S max These represent the minimum and maximum values in the current PSD data, respectively. e This represents a small positive constant used to avoid a denominator of zero; 255 represents the upper limit of pixel intensity for an 8-bit image.
[0041] like Figure 3 As shown, the PSD spectral images corresponding to the east-west, north-south, and vertical directions within the same time window are used as three modal inputs. To preserve the complementary relationship between the three-axis signals, the three PSD spectral images are mapped to the R, G, and B channels in the RGB color space, respectively, and then weighted and fused according to the channel weights to obtain the multimodal spectral image. Wherein, PEW, PNS, and PUD represent the east-west, north-south, and vertical PSD spectrum images, respectively, and αE, αN, and αU are the channel weights for the three directions, with αE+αN+αU=1. In a preferred embodiment, the weights for each of the three channels are all set to 1 / 3, mapping the east-west PSD spectrum image to the red channel, the north-south PSD spectrum image to the green channel, and the vertical PSD spectrum image to the blue channel. For structures with obvious directional characteristics, the weights can also be adjusted based on the sensor installation direction, the main vibration direction of the structure, historical classification results, or the proportion of signal energy in each direction.
[0042] Through the above fusion, the frequency domain features of the three directions are compressed into a single RGB multimodal spectrum image, which not only preserves cross-directional physical information but also reduces data redundancy caused by separate input of the three-directional images, making it easier to perform real-time inference on edge computing devices.
[0043] ; In the formula, C fuse This represents the fused RGB multimodal spectrum image; P EW , P NS and P UD These represent the PSD spectral images in the east-west, north-south, and vertical directions, respectively. α E , α N and α U These represent the channel weights corresponding to the east-west, north-south, and vertical directions, respectively. R , G , B The three channels are sequentially from α E · P EW , α N · P NS and α U · P UD Composition; the weights of each channel satisfy α E + α N + α U =1.
[0044] like Figure 4As shown, multimodal spectral images are used as input to a CNN recognition model deployed on an edge computing device, with output categories including seismic signals, vehicle vibrations, and environmental noise. In one embodiment, the input size is 240×360×3, and the network includes five convolutional blocks, flattened layers, two fully connected layers, and a Softmax three-class classification output layer. Table 1 shows the CNN deep convolutional network structure: Table 1 Convolutional layers extract PSD spectral texture and three-axis fusion features through local convolutional kernels: Where H(l) is the feature map of the l-th layer, K(l) is the convolution kernel weight, b(l) is the bias, and σ is the ReLU activation function. The max-pooling layer is used to preserve local salient responses and reduce the feature map size. The Softmax classification layer outputs the class probabilities of seismic signals, vehicle vibrations, and environmental noise.
[0045] During training, multimodal spectral images are used as samples, and seismic signals, vehicle vibrations, and environmental noise are used as labels. A multi-class cross-entropy loss function is employed for training. In one embodiment, the optimizer uses a stochastic gradient descent optimizer with a learning rate of 0.0001, a momentum of 0.8, and 100 training epochs. The above network structure and training parameters are used to illustrate possible implementations of the present invention and do not constitute a limitation on the number of convolutional network layers, optimizer type, or number of training epochs.
[0046] During edge inference, the edge computing device uses the category corresponding to the highest category probability as the category of structural vibration event. When identified as an earthquake signal, the edge computing device uploads the original vibration segment, PSD spectrum data, multimodal spectrum image, and recognition result, and triggers the digital twin system to perform structural dynamic response calculation; when identified as vehicle vibration or environmental noise, the edge computing device logs and suppresses false earthquake event triggering.
[0047] ; In the formula, H l Indicates the first l Convolutional layers output feature maps; H l-1 Indicates the first l -1 layer input feature map; K l Indicates the first l Layer convolution kernel weights; b l Indicates the first l Layer bias term; * indicates convolution operation; ReLU(·) represents the rectified linear activation function; l Indicates the sequence number of the convolutional layer or convolutional block.
[0048] ; In the formula, p k This indicates that the output of the CNN recognition model is the first... k The probability of structural vibration events; z k Indicates the first k The output value of the fully connected layer corresponding to the class; z i Indicates the first i The output value of the fully connected layer corresponding to the class; K Indicates the total number of event categories; i Indicates a category summation index; k This indicates the current category number; exp(·) represents the exponential function.
[0049] ; In the formula, This indicates the predicted category output by the CNN recognition model; Indicates the category number k The probability of taking the upper pass The largest category; Indicates the first k Predicted probability of structural vibration events; k Indicates the category number.
[0050] ; In the formula, L This represents the cross-entropy loss function for multi-class classification. B This indicates the number of samples in a training batch. b Indicates the sample sequence number within the batch; K Indicates the total number of event categories; k Indicates the category number; y b,k Indicates the first b The sample at the th k The actual label value on the class; p b,k The model represents the first b The sample belongs to the first k The predicted probability of the class; log(·) represents the natural logarithm function.
[0051] like Figure 5As shown, this embodiment also provides a structural vibration event edge recognition device. This device, serving as the hardware carrier of the above method, may include an acceleration vibration sensor, a data acquisition interface, an edge computer, a memory, a processor, and a communication module. The edge computer in the device can integrate functions such as data reception, preprocessing, DWT-IDWT frequency band reconstruction, PSD data generation, RGB weighted fusion, CNN edge recognition, and edge triggering. Table 2 shows the composition of the edge recognition device in this embodiment: Table 2 The system includes an acceleration vibration sensor for acquiring multi-directional vibration time-history signals of the engineering structure; a data acquisition interface for inputting multi-directional vibration time-history signals into an edge computer; a memory for storing computer programs, CNN recognition models, sample fragments, PSD spectrum data, and event logs; a processor for executing the aforementioned structural vibration event edge recognition method; and a communication module for uploading event data, triggering digital twin calculations, or suppressing false triggers based on the recognition results.
[0052] The above-described device is only used to illustrate the implementable hardware carrier of this method and does not constitute a limitation on the model of the acceleration vibration sensor, the model of the edge computer, the type of processor, the type of storage medium, or the communication method.
[0053] In one experimental embodiment, the dataset includes actual vibration data recorded by acceleration vibration sensors on engineering structures. The training set, validation set, and test set are used for model training, parameter validation, and recognition performance evaluation, respectively. Test results can be evaluated using confusion matrix, accuracy, seismic signal recall rate, and vehicle vibration false trigger rate. These evaluation metrics can be expressed as follows: ; In the formula, Acc represents the overall recognition accuracy of the test set; Oh Represents the set of test samples; | Oh | indicates the total number of test samples; i Indicates the test sample number; I(·) represents the indicator function, which takes the value 1 when the condition in parentheses is true, and 0 otherwise; Indicates the first i The predicted category for each test sample; yes Indicates the first i The true category of each test sample.
[0054] ; In the formula, Recall EQ The recall rate represents the seismic signal category; EQ represents the seismic signal category. TP EQ This represents the number of samples that are actually seismic signals and have been correctly identified as seismic signals. FN EQ This represents the number of samples that are actually seismic signals but were not identified as such.
[0055] ; In the formula, FAR WC→EQ This indicates the false trigger rate where vehicle vibrations are misidentified as seismic signals; WC represents the vehicle vibration category; EQ represents the seismic signal category. FP WC→EQ This represents the number of samples that were actually vehicle vibrations but were mistakenly identified as seismic signals. TN WC This represents the number of samples that are actually vehicle vibrations and were not mistakenly identified as seismic signals.
[0056] Experimental results show that the method in this embodiment can effectively distinguish three types of vibration events, the earthquake signal recall rate is significantly higher than that of the traditional threshold method, and the false trigger rate of vehicle vibration is greatly reduced.
[0057] Compared with the prior art, this embodiment has at least the following beneficial effects: First, multi-scale frequency band features are extracted through DWT-IDWT band reconstruction, enhancing the distinguishability of seismic signals, vehicle vibrations, and environmental noise in the time-frequency domain. Second, one-dimensional time series are converted into two-dimensional texture images through PSD data generation and spectral image encoding, enabling CNN models to learn spectral energy distribution characteristics. Third, RGB channel weighted fusion preserves the complementary relationships and directional coupling information between the three-axis signals, while reducing data redundancy caused by separate input of images in the three directions. Fourth, identification and triggering decisions are completed near the sensor using edge computing devices, which improves the real-time performance of the monitoring system and reduces communication, storage, and digital twin computing overhead caused by false triggers. Fifth, the method can form a complete device scheme with acceleration vibration sensors, edge computers, and monitoring platforms, facilitating deployment and implementation in structural health monitoring systems.
[0058] The above are merely preferred embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for edge recognition of structural vibration events based on DWT-CNN multimodal spectral fusion, characterized in that, include: Receives multi-directional vibration time history signals of engineering structures collected by acceleration vibration sensors; Based on the multi-directional vibration time history signal, preprocessed sample segments in each direction are obtained; Perform discrete wavelet transform on sample segments in each direction to obtain multi-scale wavelet coefficients in each direction. Based on the preset frequency band, retain the target wavelet coefficients and set the remaining wavelet coefficients to zero. Reconstruct the zeroed wavelet coefficients into time domain signals of the corresponding frequency bands through inverse discrete wavelet transform to obtain the reconstructed frequency band signals in each direction. Power spectral density is estimated based on the reconstructed signals in each frequency band, power spectral density spectrum data in each direction is obtained, and power spectral density spectrum images in each direction are generated based on the power spectral density spectrum data. The power spectral density spectrum images in each direction are mapped to different color channels, and weighted fusion is performed according to the channel weights of each color channel to obtain a multimodal spectrum image; The multimodal spectral image is input into a convolutional neural network recognition model to obtain the category probabilities of earthquake signals, vehicle vibrations, and environmental noise, and the category of structural vibration events is determined based on the category probabilities.
2. The method according to claim 1, characterized in that, Based on the multi-directional vibration time history signal, preprocessed sample segments in each direction are obtained, specifically including: Vibration time history signals in multiple directions are continuously acquired according to a preset sampling frequency, and sample segments are divided by a preset time window. Linear detrending and constant mean removal processes are sequentially performed on the sample segments in each direction to remove the linear trend term and constant offset of the sample segments in each direction. The sample segments in each direction after detrending and mean removal are zeroed out to obtain preprocessed sample segments in each direction.
3. The method according to claim 1, characterized in that, Perform discrete wavelet transform on sample segments in each direction to obtain multi-scale wavelet coefficients in each direction. Based on a preset frequency band, retain the target wavelet coefficients and set all other wavelet coefficients to zero. Reconstruct the zeroed wavelet coefficients into time-domain signals for the corresponding frequency bands using inverse discrete wavelet transform, obtaining the reconstructed frequency band signals for each direction. Specifically, this includes: A multi-level discrete wavelet decomposition of sample segments in each direction is performed using a preset wavelet basis function to obtain low-frequency approximation coefficients and multi-level high-frequency detail coefficients. The wavelet coefficient levels to be retained are determined based on the target frequency band. The low-frequency approximation coefficients or high-frequency detail coefficients of the wavelet coefficient levels to be retained are taken as the target wavelet coefficients, and all wavelet coefficients of the other levels except the target wavelet coefficients are set to zero. The target wavelet coefficients and the remaining level wavelet coefficients after being set to zero are input into the discrete wavelet inverse transform to reconstruct the time domain signal of the corresponding target frequency band, and the reconstructed frequency band signals in each direction are obtained.
4. The method according to claim 1, characterized in that, Power spectral density estimation is performed based on the reconstructed signals in each frequency band, power spectral density spectrum data in each direction is obtained, and power spectral density spectrum images in each direction are generated based on the power spectral density spectrum data, specifically including: Welch method is used to estimate the power spectral density of the reconstructed signal in each frequency band and obtain the power spectral density spectrum data in each direction. The power spectral density values in the power spectral density spectrum data of each direction are normalized to the pixel intensity range, and the normalized spectrum data is plotted as a filled two-dimensional power spectral density spectrum image.
5. The method according to claim 1, characterized in that, The power spectral density spectrum images in each direction are mapped to different color channels, and then weighted and fused according to the channel weights of each color channel to obtain a multimodal spectrum image, specifically including: Map the east-west power spectral density spectrum image to the red channel, the north-south power spectral density spectrum image to the green channel, and the vertical power spectral density spectrum image to the blue channel; The three color channels are weighted and summed according to the preset weights for the red, green, and blue channels to obtain a multimodal spectrum image.
6. The method according to claim 1, characterized in that, After determining the category of structural vibration events based on the aforementioned category probabilities, the process further includes: When the structural vibration event is classified as a seismic signal, the control edge computing device uploads the original vibration segment, power spectral density spectrum data, multimodal spectrum image and recognition result, and triggers the digital twin system to perform structural dynamic response calculation; When the structural vibration event is classified as vehicle vibration or environmental noise, the control edge computing device records the event log and suppresses false triggering of earthquake events.
7. The method according to claim 1, characterized in that, The convolutional neural network recognition model includes five sequentially connected convolutional blocks, one flattening layer, two fully connected layers, and one Softmax classification layer. Each convolutional block contains a convolutional layer and a max pooling layer. The convolutional layer is used to extract local convolutional features from the multimodal spectral image, and the max pooling layer is used to preserve local salient responses and reduce the feature map size.
8. A structural vibration event edge recognition device, characterized in that, include: Accelerometer vibration sensor, data acquisition interface, edge computer, memory and communication module; The acceleration vibration sensor is used to collect time history signals of multi-directional vibration of engineering structures; The data acquisition interface is connected to the acceleration vibration sensor and the edge computer, and is used to transmit the multi-directional vibration time history signal to the edge computer. The memory stores computer programs and convolutional neural network recognition models; The edge computer is connected to the memory and the communication module to execute the computer program to implement the method as described in any one of claims 1-7, and to upload event data or trigger digital twin linkage calculations through the communication module based on the identification result.
9. The apparatus according to claim 8, characterized in that, The edge computer includes: The data receiving module is used to receive the multi-directional vibration time history signal transmitted by the data acquisition interface; The preprocessing module is used to perform detrending, zeroing, and fixed-length segmentation on the multi-directional vibration time history signal to obtain sample segments in each direction; The Discrete Wavelet Transform and Reconstruction Module is used to perform Discrete Wavelet Transform and Inverse Discrete Wavelet Transform on sample segments in each direction to obtain reconstructed signals in each frequency band. The power spectral density generation module is used to estimate the power spectral density of the reconstructed signal in each frequency band and generate power spectral density spectrum data and power spectral density spectrum images in each direction. The weighted fusion module is used to map the power spectral density spectrum images of each direction to different color channels and fuse them according to channel weights to generate a multimodal spectrum image; A convolutional neural network recognition module is used to deploy a convolutional neural network recognition model to recognize the multimodal spectral image and output the category probabilities of seismic signals, vehicle vibrations, and environmental noise. The control module is used to control event uploading, digital twin triggering, or false triggering suppression based on the probability of the category.
10. The apparatus according to claim 8, characterized in that, The acceleration vibration sensor is at least one of a strong vibration meter, a triaxial accelerometer, or a MEMS acceleration sensor; The communication module is at least one of an Ethernet module, a 4G communication module, a 5G communication module, an optical fiber communication module, or a wireless local area network module.