High-rigidity structure target vibration recognition method and system based on computer vision, terminal and medium
By constructing a bridge model training dataset and training an EDVR model, and combining a pyramid cascade deformable alignment and spatiotemporal attention fusion network, the problem of image detail loss in vibration recognition of high-stiffness structures is solved, achieving high-precision and efficient recognition results, and solving the problem of insufficient recognition accuracy in existing technologies.
Patent Information
- Application Number
- CN202511519356.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-23
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2045-10-23
AI Technical Summary
Existing technologies struggle to effectively identify vibration details in high-rigidity structures, especially when cameras capture images under environmental interference, as they are prone to losing vibration details at feature points. Furthermore, image super-resolution processing methods can easily introduce false textures and jagged edges, reducing recognition accuracy.
A training dataset based on a bridge model was constructed to train the EDVR model, which includes a pyramid cascade deformable alignment network and a spatiotemporal attention fusion network. The image sequence was processed by super-resolution, and the target displacement and vibration frequency were identified by fitting an ellipse.
It improves the accuracy and sensitivity of vibration identification in high-rigidity structures, overcomes environmental interference and hardware limitations, and expands the application scope of computer vision technology in civil engineering.
Smart Images

Figure CN121033631A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of building vibration identification, and in particular to a high-rigidity structure target point vibration identification method, system, terminal and medium based on computer vision. BACKGROUND
[0002] Structural health monitoring using computer vision technology is increasingly popular in the engineering industry. Structural vibration analysis based on computer vision generally works well for structures with low rigidity. Low-rigidity structures have large deformation displacement. When a camera is used to capture the vibration process, the vibration displacement of the feature points can be captured. However, for high-rigidity structures, the displacement response is very small due to the high natural frequency. When a camera is used to capture the vibration, the vibration details of the feature points are easily lost under environmental interference. Due to the limitation of resolution, it is difficult to distinguish the very small displacement changes, and they are easily overwhelmed by noise. In addition, some subtle vibration textures that can be identified are easily smoothed out as noise in image processing.
[0003] When a camera is used to identify structural vibration, in order to obtain clearer images and higher shooting frame rates, one of the methods is to use a camera with better hardware conditions. However, the cost of the camera will also be a concern. When the hardware conditions of the camera are limited, the shooting accuracy of the camera to the target will also be limited. In the case of limited camera hardware conditions, in order to improve the accuracy of target vibration displacement identification, there is a certain constraint relationship between the distance from the camera to the target. When the camera is far away from the target, the proportion of the target in the image will be small, and it will be more difficult to identify the subtle displacement changes of the target. The obtained image sequence will introduce too much noise, and the subtle displacement of the target contained in the noise will be removed in the noise smoothing of image processing.
[0004] In the field of computer vision, image super-resolution processing is an effective processing method when the original image quality is poor. However, the image super-resolution processing method is to improve the resolution of a single image. When facing the image sequence formed by the vibration video, the time effect of the image sequence will be ignored when only a single image is processed by the super-resolution processing method. Especially when the original image quality is poor and the target displacement is very small, the generated high-resolution image will have problems such as false texture, jagged edges, ghosting, etc. The original small vibration details are smoothed out in the generated high-resolution image. The effect will be worse when extracting the target vibration displacement from the image sequence processed by the image super-resolution processing method.
[0005] Therefore, the prior art still has defects. SUMMARY
[0006] The technical problems to be solved by the present application are to provide a high-rigidity structure target point vibration identification method, system, terminal and medium based on computer vision to solve the above defects of the prior art. In a first aspect, the present application provides a high-rigidity structure target point vibration identification method based on computer vision, wherein the method comprises: Based on the bridge model, a training data set for target point vibration identification is constructed, and an EDVR model is trained based on the training data set, wherein the EDVR model is a video super-resolution network, and the training of the EDVR model includes the training of a pyramid cascaded deformable alignment network and the training of a spatiotemporal attention fusion network. An image sequence to be identified is obtained, the image sequence to be identified is subjected to super-resolution processing based on the trained EDVR model, and a fitting ellipse is obtained according to the image sequence to be identified subjected to super-resolution processing. Based on the fitting ellipse, the identification of structure target point displacement and structure vibration frequency is performed to obtain the vertical physical displacement of the target point at each moment and the vertical vibration acceleration data of the target point at each moment, and the deflection of the bridge and the frequency of the bridge are obtained based on the vertical physical displacement and the vertical vibration acceleration data.
[0007] In an implementation manner, based on the bridge model, a training data set for target point vibration identification is constructed, comprising: A plurality of target points are arranged at different positions on the bridge model, and a camera is placed at the end of the bridge model to shoot along the bridge direction of the bridge model to obtain a structure vibration video, wherein the distance between the camera and each target point is different. The structure vibration video is converted into a.mxf format file and split into a training image sequence. Based on the training image sequence, the interested regions where each target point is located are framed to obtain an interested region image sequence of each target point. Based on the interested region image sequence, a training data set for target point vibration identification is obtained.
[0008] In an implementation manner, based on the interested region image sequence, a training data set for target point vibration identification is obtained, comprising: The interested image sequence corresponding to each target point is subjected to blur processing, brightness change processing, symmetry processing, noise processing, rotation processing and resolution reduction processing to obtain a low-resolution data set. The interested image sequence corresponding to each target point is subjected to symmetry processing, noise processing and rotation processing to obtain a high-resolution data set. According to the low-resolution data set and the high-resolution data set, a training data set for target point vibration recognition is obtained, wherein the low-resolution data set and the high-resolution data set are one-to-one corresponding.
[0009] In an implementation manner, the training process of the pyramid cascaded deformable alignment network comprises: Based on the training data set, 2 times down-sampling is performed on the current frame and each of the previous and subsequent frames to form a pyramid of feature maps with different resolutions; The offset of the target frame feature map and each of the previous and subsequent frame feature maps is calculated, and the offset matrix of the current layer is sequentially up-sampled and fused with the offset matrix of the previous layer; The previous and subsequent frame feature maps of the target frame feature map are deformable convoluted with the offset matrix of the current layer to form an alignment feature matrix of the current layer, and the alignment feature matrix of the previous layer is sequentially up-sampled and fused to form the offset of each layer level; The offset of each layer level is deformable convoluted with the alignment feature matrix of the first layer to obtain the alignment feature matrix of the current frame and the previous and subsequent frames.
[0010] In an implementation manner, the training process of the spatio-temporal attention fusion network comprises: When training the temporal attention part, the temporal feature map correlation of the target frame alignment feature and the alignment features of the previous and subsequent frames is calculated to obtain the temporal attention weight of each frame; The temporal attention weight is element-wise multiplied with the original alignment feature to obtain the temporal attention modulation feature of each frame, and a fusion convolution layer is used to fuse the temporal attention modulation features of each frame; When training the spatial attention part, the fusion result of the temporal attention modulation features of each frame is taken as input, a pyramid structure is used to obtain the spatial feature maps of each layer, and the up-sampling processing, addition processing and multiplication processing are combined to obtain the feature map of the fused spatial attention.
[0011] In an implementation manner, according to the to-be-recognized image sequence subjected to super-resolution processing, a fitting ellipse is obtained, comprising: According to the to-be-recognized image sequence subjected to super-resolution processing, a to-be-recognized region of interest image is determined; The to-be-recognized region of interest image is subjected to grayscale processing, and the to-be-recognized region of interest image subjected to grayscale processing is subjected to Otsu adaptive threshold binarization processing to obtain a binary image, the binary image being used to highlight the edge features of the target point; The Canny algorithm is used to perform edge detection on the target point in the binary image to obtain edge pixel information; According to the edge pixel information, an edge contour fitting is performed by using a least square ellipse fitting method to obtain a fitting ellipse.
[0012] In an implementation manner, based on the fitting ellipse, identification of structural target point displacement and structural vibration frequency is performed to obtain vertical physical displacement of the target point at each moment and vertical vibration acceleration data of the target point at each moment, including: Based on the ellipse center coordinates of the fitting ellipse, pixel coordinates of a static point in the image sequence to be identified are calculated, and based on the ellipse center coordinates and the pixel coordinates of the static point, pixel vertical coordinates of the target point are obtained; The pixel vertical coordinates of the target point are subjected to one forward difference to obtain a pixel displacement change amount of the target point at each moment, and the pixel displacement change amount is converted to obtain the vertical physical displacement of the target point at each moment; The pixel vertical coordinates of the target point are subjected to two forward differences to obtain the vertical vibration acceleration data of the target point at each moment.
[0013] In a second aspect, the embodiments of the present application further provide a high-rigidity structure target point vibration identification system based on computer vision, wherein the system is used to implement the steps of the high-rigidity structure target point vibration identification method based on computer vision in any one of the above schemes, and the system comprises: A model training module is configured to construct a training data set for target point vibration identification based on a bridge model, and train an EDVR model based on the training data set, wherein the EDVR model is a video super-resolution network, and the training of the EDVR model includes training of a pyramid cascading deformable alignment network and training of a spatiotemporal attention fusion network; A super-resolution processing module is configured to obtain an image sequence to be identified, perform super-resolution processing on the image sequence to be identified based on the trained EDVR model, and obtain a fitting ellipse according to the image sequence to be identified subjected to the super-resolution processing; A target point vibration identification module is configured to perform identification of structural target point displacement and structural vibration frequency based on the fitting ellipse to obtain vertical physical displacement of the target point at each moment and vertical vibration acceleration data of the target point at each moment, and obtain deflection of the bridge and frequency of the bridge based on the vertical physical displacement and the vertical vibration acceleration data.
[0014] In a third aspect, the embodiments of the present application further provide a terminal, wherein the terminal comprises a memory, a processor, and a high-rigidity structure target point vibration identification program based on computer vision stored in the memory and executable on the processor, and when the processor executes the high-rigidity structure target point vibration identification program based on computer vision, the steps of the high-rigidity structure target point vibration identification method based on computer vision in any one of the above schemes are implemented.
[0015] In a fourth aspect, the embodiments of the present application also provide a computer readable storage medium, wherein a computer vision-based high-rigidity structure target point vibration identification program is stored on the computer readable storage medium, and the computer vision-based high-rigidity structure target point vibration identification program implements the steps of the computer vision-based high-rigidity structure target point vibration identification method in any of the above solutions on the computer readable storage medium.
[0016] Beneficial effects: Compared with the prior art, the present application provides a computer vision-based high-rigidity structure target point vibration identification method, which firstly constructs a training data set for target point vibration identification based on a bridge model, trains an EDVR model based on the training data set, wherein the EDVR model is a video super-resolution network, and the training of the EDVR model includes the training of a pyramid cascading deformable alignment network and the training of a spatiotemporal attention fusion network. Then, an image sequence to be identified is obtained, the image sequence to be identified is subjected to super-resolution processing based on the trained EDVR model, and a fitting ellipse is obtained according to the image sequence to be identified subjected to super-resolution processing. Finally, based on the fitting ellipse, the identification of structure target point displacement and structure vibration frequency is carried out, the vertical physical displacement of the target point at each moment and the vertical vibration acceleration data of the target point at each moment are obtained, and the deflection of the bridge and the frequency of the bridge are obtained based on the vertical physical displacement and the vertical vibration acceleration data.
[0017] The present application applies the EDVR model in the field of computer video super-resolution to the vibration identification of high-rigidity structures in civil engineering, enhances the precision and sensitivity of computer vision technology for vibration identification of high-rigidity structures, overcomes the problem that the vibration texture in the original image sequence is not clear and is smoothed in the case that the camera hardware and environmental conditions are limited, and expands the use range of computer vision technology in civil engineering. BRIEF DESCRIPTION OF DRAWINGS
[0018] Figure 1 The flowchart of the preferred embodiment of the computer vision-based high-rigidity structure target point vibration identification method provided by the embodiments of the present application.
[0019] Figure 2 The schematic diagram of the pyramid cascading deformable alignment network in the computer vision-based high-rigidity structure target point vibration identification method provided by the embodiments of the present application.
[0020] Figure 3 The schematic diagram of the spatiotemporal attention fusion network in the computer vision-based high-rigidity structure target point vibration identification method provided by the embodiments of the present application.
[0021] Figure 4 The effect comparison diagram of the super-resolution processing of the trained EDVR model of the embodiments of the present application.
[0022] Figure 5 The target point vibration acceleration frequency domain graph identified based on the original image sequence without using the EDVR model processing after the structural target point vibration video is converted into the.mxf format video.
[0023] Figure 6 The target point vibration acceleration frequency domain graph identified based on the EDVR model trained in the embodiment after the structural target point vibration video is converted into the.mxf format video.
[0024] Figure 7 The target point vibration acceleration frequency domain graph identified by the acceleration sensor.
[0025] Figure 8 The target point vibration acceleration frequency domain graph identified based on the original image sequence without using the EDVR model processing after the structural target point vibration video directly shot by the camera.
[0026] Figure 9 The target point vibration acceleration frequency domain graph identified based on the traditional EDVR pre-trained model processing after the structural target point vibration video directly shot by the camera.
[0027] Figure 10 The target point vibration acceleration frequency domain graph identified based on the EDVR model trained in the embodiment after the structural target point vibration video directly shot by the camera.
[0028] Figure 11 The schematic diagram of the high-rigidity structural target point vibration identification system based on computer vision provided by the embodiment of the application.
[0029] Figure 12 The schematic diagram of the terminal provided by the embodiment of the application. DETAILED DESCRIPTION
[0030] To make the purpose, technical scheme and effect of the application more clear and explicit, the application is further described in detail below with reference to the drawings and examples. It should be understood that the specific examples described herein are only used to explain the application and are not used to limit the application.
[0031] The flowchart shown in the drawings is only an example and does not necessarily include all the contents and operations or steps, nor does it necessarily be executed in the order described. For example, some operations or steps can be decomposed, combined or partially merged, so that the actual execution order can be changed according to the actual situation.
[0032] It is to be understood that the terms used in the specification and the appended claims are intended to describe particular embodiments by way of example and not meant in a limiting sense. It should be understood that, in order to facilitate the clear description of the technical solutions of the embodiments of the present application, in the embodiments of the present application, the terms "first", "second", etc. are used to distinguish the same or similar items with basically the same function and effect. For example, the first control information and the second control information are only used to distinguish different control information, and do not limit the order. It should be understood that the terms "first", "second", etc. do not limit the quantity and execution order, and the terms "first", "second", etc. also do not necessarily mean different. It should also be understood that the term "and / or" used in the specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes these combinations.
[0033] Image resolution is a major concern in the field of computer vision. In recent years, many scholars have devoted to restoring low-quality images to high-quality images while preserving details and textures. A series of super-resolution models have been developed through deep learning. Super-resolution models can be divided into image super-resolution models and video super-resolution models. Image super-resolution model refers to considering only the pixel features of the current image when performing super-resolution processing on low-resolution images, that is, when training an image super-resolution model, for a certain frame of low-resolution image, only the mapping features of the low-resolution image to the corresponding high-resolution image in the training data set are learned. Video super-resolution model refers to considering the time effect of the current image and the upper and lower frame images when performing super-resolution processing on low-resolution images, that is, when training a video super-resolution model, for a certain frame of low-resolution image, not only the mapping features of the low-resolution image to the corresponding high-resolution image in the training data set are learned, but also the mapping features of the upper and lower frame low-resolution images to the corresponding upper and lower frame high-resolution images are learned.
[0034] Structural health monitoring using computer vision technology is increasingly popular in the engineering industry. Compared with vibration sensors, the structural monitoring method using cameras is more affected by the environment, and the resolution and texture accuracy of the image are also limited by the cost of the camera. In addition, the most critical problem is that vibration sensors are more sensitive to high-frequency vibrations, while cameras can only identify low-frequency vibrations of the structure due to hardware limitations. Therefore, under the condition that the camera hardware is limited, it is difficult to identify the vibration frequency of the high-stiffness structure.
[0035] To solve the problems in the prior art, the embodiment provides a high-rigidity structure target point vibration identification method based on computer vision, applies an EDVR (Video Restoration framework with Enhanced Deformable convolutions) model in the field of computer video super resolution to vibration identification of a high-rigidity structure in civil engineering, enhances the precision and sensitivity of computer vision technology in vibration identification of the high-rigidity structure, overcomes the problem that vibration texture in an original image sequence is not clearly identified and is smoothed in the case that camera hardware and environmental conditions are limited, and expands the use range of computer vision technology in civil engineering. In a specific application, the embodiment first constructs a training data set for target point vibration identification based on a bridge model, trains an EDVR model based on the training data set, wherein the training of the EDVR model includes training of a pyramid cascading deformable alignment network and training of a spatiotemporal attention fusion network. Then, an image sequence to be identified is acquired, the image sequence to be identified is subjected to super resolution processing based on the trained EDVR model, and a fitted ellipse is obtained according to the image sequence to be identified subjected to super resolution processing. Finally, structure target point displacement and structure vibration frequency are identified based on the fitted ellipse, vertical physical displacement of a target point at each moment and vertical vibration acceleration data of the target point at each moment are obtained, and deflection of a bridge and frequency of the bridge are obtained based on the vertical physical displacement and the vertical vibration acceleration data.
[0036] The high-rigidity structure target point vibration identification method based on computer vision of the embodiment can be applied to a terminal, and the terminal includes a computer and other intelligent product terminals. Figure 1 The method of the embodiment includes the following steps: In step S100, a training data set for target point vibration identification is constructed based on a bridge model, and an EDVR model is trained based on the training data set, wherein the EDVR model is a video super resolution network, and the training of the EDVR model includes training of a pyramid cascading deformable alignment network and training of a spatiotemporal attention fusion network.
[0037] Specifically, the embodiment arranges multiple target points at different positions on a bridge model for generating simulated high-rigidity structure vibration video data, and the target points are key parts in the bridge model, such as piers, bearings, etc. A camera is placed at the end of the bridge model to shoot along the bridge model in the bridge direction to obtain a structure vibration video, wherein the distances from the camera to the target points are different. The structure vibration video is converted into a.mxf format file and split to obtain a training image sequence. Then, based on the training image sequence, a region of interest (ROI) in which each target point is located is framed to obtain a region of interest image sequence of each target point. Then, based on the region of interest image sequence, a training data set for target point vibration recognition is obtained.
[0038] The training data set of the EDVR model needs to prepare a high-resolution data set and a low-resolution data set, and the high-resolution data set and the low-resolution data set correspond to each other, so the image numbers in the two data sets need to correspond to each other. Specifically, in the construction of the low-resolution data set, in the embodiment, the interest image sequence corresponding to each target point is subjected to blurring processing, brightness change processing, symmetry processing, noise processing, rotation processing, and resolution reduction processing to obtain the low-resolution data set. In the construction of the high-resolution data set, based on the one-to-one correspondence relationship between the high-resolution data set and the low-resolution data set, the blurring processing and the high brightness processing in the low-resolution data set do not need to be performed when constructing the high-resolution data set, so only the interest image sequence corresponding to each target point needs to be subjected to symmetry processing, noise processing, and rotation processing to obtain the high-resolution data set.
[0039] The EDVR model needs to be trained in two stages, the first stage being the training of the pyramid cascading deformable alignment network and the second stage being the training of the spatiotemporal attention fusion network. In the training process of the two stages, the same set of training data sets is used. Specifically, the pyramid cascading deformable alignment network in the first stage is as shown in Figure 2 The embodiment first performs 2 times downsampling on the current frame and the frames before and after the current frame based on the training data set to form a pyramid of feature maps of different resolutions. Then, the offset of the target frame feature map and the feature maps of the frames before and after the target frame is calculated, and the offset matrix of the current layer is sequentially upsampled and fused with the offset matrix of the previous layer. Then, the feature maps of the frames before and after the target frame and the offset matrix of the current layer are subjected to deformable convolution to form an alignment feature matrix of the current layer, and the alignment feature matrix of the previous layer is sequentially upsampled and fused until the alignment feature matrix of the first layer is fused to form the offset of each layer. Finally, the offset of each layer and the alignment feature matrix of the first layer are subjected to deformable convolution to obtain the alignment feature matrix of the current frame and the frames before and after the current frame. In the embodiment, the alignment features of each frame are represented as:
[0040] wherein, is a raw alignment feature at the position; is an alignment feature of each frame, denotes a target frame or a current frame, refers to a frame position before or after the target frame or the current frame, i.e. a neighboring frame, , is the number of image frames in the training data set; is a weight at the k position; is a pre-specified offset at the k position; is a learnable offset; is a modulation scalar.
[0041] The spatio-temporal attention fusion network in the second stage is as shown in Figure 3 The spatio-temporal attention fusion network is divided into a temporal attention part and a spatial attention part. When training the temporal attention part, the embodiment calculates the temporal feature map correlation of the target frame alignment feature and the alignment features of the frames before and after the target frame, to obtain the temporal attention weight of each frame. The expression of the temporal attention weight is:
[0042] wherein, is the similarity of the alignment feature matrices of two frames of images; and are the alignment feature matrices of two frames of images, respectively. Each frame of image can form an alignment feature matrix. is the alignment feature matrix of the first frame of image to the target frame or the current frame is the alignment feature matrix of oneself to oneself, is the alignment feature matrix of oneself to oneself, and are convolution calculations, is a transposed matrix.
[0043] Then, the temporal attention weight is multiplied element by element with the raw alignment feature, to obtain the temporal attention modulation feature of each frame, and a fusion convolution layer is used to fuse the temporal attention modulation features of each frame. The expression of the temporal attention modulation feature is: wherein, is the multiplication of the elements at the corresponding positions in the two matrices.
[0044] The expression of the fusion of the temporal attention modulation features of each frame is: .
[0045] In the training of the spatial attention part, the embodiment takes the fusion result of the time attention modulation features of each frame as input, adopts a pyramid structure to obtain spatial feature maps of each layer, and combines upsampling processing, addition processing and multiplication processing to obtain a feature map of the fused spatial attention. The embodiment cross-fuses the video super-resolution field in computer vision and the structural vibration identification field in civil engineering, applies the EDVR model to structural vibration identification in civil engineering, can fully utilize the time sequence information in the structural target point vibration image sequence, and through the pyramid cascaded deformable alignment and spatio-temporal attention fusion mechanism, is conducive to realizing super-resolution reconstruction and high-precision vibration feature extraction under the condition of low-resolution images.
[0046] Step S200, obtaining an image sequence to be identified, performing super-resolution processing on the image sequence to be identified based on the trained EDVR model, and obtaining a fitted ellipse according to the image sequence to be identified after super-resolution processing.
[0047] The embodiment first acquires the structural target point vibration video of the bridge to be detected through the camera, then builds an OpenCV-Python platform, and acquires the acquired structural target point vibration video. Since the format derived by the camera shooting video is generally.mp4 format, the EDVR model can be directly used in the multimedia field to restore the video of the.mp4 file, but for the identification of the micro-vibration of the high-rigidity structure in civil engineering, the EDVR model cannot be directly used to process the.mp4 format file, and the video shot by the camera needs to be additionally converted into.mxf format. The.mxf format has a smaller or even no compression amount of image than the.mp4 format, and has a better effect of saving micro-vibration texture. Therefore, the embodiment needs to convert the format of the acquired structural target point vibration video into.mxf format, then performs frame processing on the structural target point vibration video, splits it into an image sequence to be identified, and saves it to a specified folder, so that the video data can be avoided from being compressed and the details can be more completely retained. The OpenCV-Python platform is a Python binding library for solving computer vision problems. OpenCV is a cross-platform computer vision and machine learning software library, and Python is a widely used scripting language that can be applied to web crawling, artificial intelligence, data analysis and data visualization.
[0048] Further, the embodiment can perform super-resolution processing on the to-be-recognized image sequence based on the trained EDVR model, and the trained EDVR model has a function of simultaneously considering 5-frame image time relationship and 4-fold super-resolution. Specifically, the embodiment builds a deep learning running environment such as pytorch, which is an open-source Python machine learning library, constructs an inference script of the EDVR model, modifies a.yml file corresponding to each EDVR model, specifies a path of the to-be-recognized image sequence, specifies a used EDVR model, specifies an image saving path of an output, and then performs 4-fold super-resolution processing on the to-be-recognized image sequence. The super-resolution processing effect of the embodiment is shown in FIGS. 8A, 8B and 8C. Figure 4 Figure 4 In FIGS. 8A, 8B and 8C, (a) is an image processed by using a traditional EDVR pre-trained model, (b) is an image processed by using the trained EDVR model of the embodiment, and (c) is an original image without processing. Figure 4 As can be seen in FIGS. 8A, 8B and 8C, the super-resolution processing effect of the to-be-recognized image sequence by using the trained EDVR model of the embodiment is the best, and the processing of the traditional EDVR pre-trained model introduces ghosting, pseudo texture and the like.
[0049] Further, in order to reduce the amount of calculation and improve the processing efficiency, a single to-be-recognized image sequence is preprocessed, and a region of interest (ROI) is selected in combination with the position of the target point in the image, so as to obtain a to-be-recognized region of interest image. Then, the embodiment performs grayscale processing on the to-be-recognized region of interest image, converts three-dimensional image data into one-dimensional image data, and reduces the calculation complexity. Then, the to-be-recognized region of interest image after the grayscale processing is subjected to Otsu adaptive threshold binarization processing, so as to obtain a binary image, which is used to highlight the edge features of the target point. Then, the Canny algorithm (a multi-level edge detection algorithm) is used to perform edge detection on the target point in the binary image, so as to obtain edge pixel information. In the Canny algorithm, a Sobel operator is used to perform convolution calculation on the image, and the gradients in the horizontal and vertical directions of the image are calculated respectively. The Sobel operator is one of the most important operators in pixel image edge detection, and plays an important role in the fields of machine learning, digital media, computer vision and the like. Specifically, the Sobel operator in the horizontal direction is:
[0050] the Sobel operator in the vertical direction is:
[0051] Finally, according to the edge pixel information, an ellipse fitting method of the least square method is used to fit the edge contour, and an error function is constructed and a minimum processing is performed on the same to obtain a fitting ellipse, specifically as follows:
[0052] In the formula of the fitting ellipse, , respectively are the semi-major axis length and the semi-minor axis length of the ellipse, , respectively are the horizontal coordinate and the vertical coordinate of the center of the ellipse, is the rotation angle of the ellipse, , respectively are the horizontal coordinate and the vertical coordinate of the data points in the image.
[0053] In actual application, the embodiment can first sort the image sequence to be recognized in the folder. Then a for loop body is established to realize the processing of each frame of image. In a single loop of the for, the selection of the region of interest, the gray dimension reduction, and the binarization are sequentially performed. In the binarization part, the adaptive Otsu threshold method is used to highlight the edge features of the target points. Then the binarized image highlighting the edge features of the target points is subjected to Canny edge detection to realize the acquisition of the outermost contour information of the edge. Finally, the ellipse is fitted, and the coordinates of the center of the ellipse are output.
[0054] Step S300, based on the fitting ellipse, the recognition of the structural target point displacement and the structural vibration frequency is performed to obtain the vertical physical displacement of the target point at each moment and the vertical vibration acceleration data of the target point at each moment, and based on the vertical physical displacement and the vertical vibration acceleration data, the deflection of the bridge and the frequency of the bridge are obtained.
[0055] In the recognition of the structural target point displacement, after obtaining the above-mentioned coordinates of the center of the ellipse, in order to eliminate the camera jitter caused by the environment, based on the coordinates of the center of the ellipse of the fitting ellipse, the pixel coordinates of the stationary point in the image sequence to be recognized are calculated. The difference between the coordinates of each center of the ellipse and the coordinates of the stationary point is obtained to obtain the vertical pixel coordinates of the target point of the image sequence to be recognized, and the formula is as follows:
[0056] Among them, is the actual displacement after deducting the camera jitter. Then, the vertical pixel coordinates of the target point are subjected to a forward difference to obtain the pixel displacement change amount of the target point at each moment, and the pixel displacement change amount is converted to obtain the vertical physical displacement of the target point at each moment. Further, the deflection of the bridge can be obtained based on the analysis and calculation of the vertical physical displacement. The formula of the forward difference is as follows:
[0057] The formula of the first forward difference is as follows: k is the image sequence number, The coordinates of the point in the image with the image sequence number k are represented, and the essence of the formula of the forward difference is that the coordinates calculated from the next frame image are subtracted from the coordinates calculated from the previous frame image to obtain the displacement.
[0058] The formula for converting the pixel displacement change amount is as follows:
[0059] In the formula for converting the pixel displacement change amount, is the actual displacement of the target point, is the pixel displacement of the target point in the image, is the actual distance from the camera to the target point plane, is the focal length of the camera, is the angle between the camera optical axis and the horizontal line, is the size of a unit pixel.
[0060] In the identification of the structural vibration frequency, the vertical coordinates of the target point pixels in the image sequence to be identified are subjected to second forward difference to obtain the vertical vibration acceleration data of the target point at each time, and the formula is as follows:
[0061] In the formula of the second forward difference, is the difference operator, k is the image sequence number.
[0062] After obtaining the vertical vibration acceleration data of the target point at each time, the frequency of the bridge can be obtained by performing fast Fourier transform on the vertical vibration acceleration data, and the formula is as follows:
[0063] In the formula of the fast Fourier transform, is the length or total number of input data, i.e., the number of time domain data in the fast Fourier transform, is the acceleration data in the time domain, is the acceleration signal in the frequency domain, is the sequence number of the time domain data, each corresponds to the vertical vibration acceleration data at one time, is the sequence number of the frequency domain component, each corresponds to an acceleration signal in the frequency domain.
[0064] In actual application, the embodiment can traverse all image sequences to be identified, that is, the y-axis coordinate of the center of the target point in each frame of image is obtained, and the actual y-axis coordinate of the target point is obtained by subtracting the y-axis coordinate of the static point from the y-axis coordinate of the center of the target point , then the obtained time domain data is subjected to one forward difference by using the np.diff() function, and based on the size conversion factor of physical displacement and pixel displacement, the actual change amount of the bridge deflection is obtained. The time domain data is subjected to two forward differences, that is, the np.diff() function is used twice, and the vibration acceleration information of each target point is obtained. The bridge frequency is obtained by performing fast Fourier transform on the time domain acceleration data, and an image is drawn. The application effect of the embodiment is shown in Figure 5 , Figure 6 and Figure 7 . Since the image compression amount of the.mxf format is smaller than that of the.mp4 format, even no compression, the effect of saving the tiny vibration texture is better. Therefore, the structure target point vibration video captured by the camera is converted into a.mxf format file, and then processed based on the trained EDVR model and the target point vibration frequency domain image is identified. Based on this, Figure 5 , after the structure target point vibration video is converted into a.mxf format video, the vibration acceleration frequency domain image of the target point identified based on the original image sequence which is not processed by using the EDVR model. Figure 6 , after the structure target point vibration video is converted into a.mxf format video, the vibration acceleration frequency domain image of the target point identified based on the EDVR model trained in the embodiment. Figure 7 , the vibration acceleration frequency domain image of the target point identified by using the acceleration sensor.
[0065] Computer vision technology for identifying structural vibrations requires capturing minute displacements at target points on the structure. For structures with high flexibility, the deformation under external forces is significant, making them relatively easy to capture with a camera. However, for high-stiffness structures, the deformation is smaller, and the displacement of a single vibration is less than the size of a single pixel in the image, making it difficult to capture with a camera. This difficulty increases further, especially for high-order frequency vibrations. This invention applies an existing EDVR model to the vibration identification of civil engineering structures, requiring additional processing throughout the vibration identification process; it is not simply a resizing of the EDVR model for a different scenario. The results of using traditional EDVR pre-trained models for structural target point vibration identification are not ideal. Furthermore, the video exported from cameras is generally in .mp4 format. In the multimedia field, the EDVR model can be directly used to recover video from .mp4 files. However, when identifying minute vibrations of high-stiffness structures in civil engineering, for .mp4 format structural target vibration videos, directly processing the image sequence to generate acceleration data and then performing a Fast Fourier Transform (FFT) results in many singular values in the time domain data. After removing these singular values one by one, the first-order frequency can be identified. For .mp4 format structural target vibration videos, if the traditional EDVR model is used to generate acceleration data, many singular values exist in the time domain. If a FFT is performed directly, no frequency can be identified. If these singular values are removed one by one, the second-order frequency can be identified. However, if the EDVR model trained in this embodiment is used, the generated acceleration data has no singular values, and a FFT can be performed directly to obtain the bridge frequency. Specific comparison results are as follows: Figure 8 , Figure 9 and Figure 10 As shown. Figure 8 The target vibration frequency domain map is obtained by directly using video of structural target vibration (i.e., .mp4 format video) and based on the original image sequence without processing using the EDVR model. Figure 9 The frequency domain map of target vibration acceleration is obtained by directly using video of structural target vibration captured by a camera and processed based on a traditional EDVR pre-trained model. Figure 10 The target vibration frequency domain map is obtained by directly using video of structural target vibration captured by a camera and processing it based on the EDVR model trained in this embodiment.
[0066] The embodiment is the first time to cross and integrate the field of video super-resolution in computer vision and the field of structural vibration identification in civil engineering. The EDVR model is applied to structural vibration identification in civil engineering, which can fully utilize the time sequence information in the structural target point vibration image sequence, realize super-resolution reconstruction and high-precision vibration feature extraction under the condition of low-resolution image through the pyramid cascaded deformable alignment and spatio-temporal attention fusion mechanism. The method has good adaptability to complex environments such as blur and light change, and can reduce the dependence on high-performance visual sensors, significantly reducing the monitoring cost. The training data set constructed by the civil engineering structural target point vibration image sequence has good generalization for the vibration identification of civil engineering structures, and is more in line with the engineering application requirements, and can be widely applied to the vibration monitoring and health assessment of structures such as bridges, tunnels and high-rise buildings. In addition, at the present stage, computer vision is applied to vibration identification in civil engineering, which has higher requirements for shooting environment and camera parameters. Since the camera shooting position and site are used for shooting in many cases, the camera cannot be monitored in practice. In addition, the arrangement of fixed cameras on the structure will constrain the arrangement of the target point. The method proposed in the embodiment has great improvement in the field of computer vision vibration identification in civil engineering at the present stage. Under the condition of limited camera hardware, higher-order vibration textures can be obtained, and the identification accuracy of structural vibration is improved.
[0067] Based on the above embodiment, the application further provides a high-rigidity structural target point vibration identification system based on computer vision. The system is used to implement the steps in the method embodiments, as shown in Figure 11 The system includes a model training module 10, a super-resolution processing module 20 and a target point vibration identification module 30. Specifically, the model training module 10 is used to construct a training data set for target point vibration identification based on a bridge model, and train an EDVR model based on the training data set, wherein the EDVR model is a video super-resolution network, and the training of the EDVR model includes the training of a pyramid cascaded deformable alignment network and the training of a spatio-temporal attention fusion network. The super-resolution processing module 20 is used to obtain a to-be-identified image sequence, perform super-resolution processing on the to-be-identified image sequence based on the trained EDVR model, and obtain a fitted ellipse according to the to-be-identified image sequence subjected to super-resolution processing. The target point vibration identification module 30 is used to identify the structural target point displacement and the structural vibration frequency based on the fitted ellipse, obtain the vertical physical displacement of the target point at each moment and the vertical vibration acceleration data of the target point at each moment, and obtain the deflection of the bridge and the frequency of the bridge based on the vertical physical displacement and the vertical vibration acceleration data.
[0068] The working principles of the various functional modules in the system embodiment of the present embodiment are the same as those of the various steps in the computer vision-based high-rigidity structure target point vibration identification method described above, and thus will not be described again here.
[0069] The various modules in the computer vision-based high-rigidity structure target point vibration identification system described above can be implemented wholly or partially by software, hardware, or a combination thereof. The various modules described above can be embedded in or independent of a processor in the terminal in hardware form, or can be stored in a memory in the terminal in software form so as to be invoked and executed by the processor to perform the operations corresponding to the various modules.
[0070] Based on the embodiments described above, the present application further provides a terminal, a principle block diagram of which can be as shown in Figure 12 The terminal can include one or more processors 100 (only one is shown in Figure 12 ), a memory 101, and a computer program 102 stored in the memory 101 and executable on the one or more processors 100. For example, a computer vision-based high-rigidity structure target point vibration identification program. The one or more processors 100 can implement the various steps in the computer vision-based high-rigidity structure target point vibration identification method embodiment when executing the computer program 102. Alternatively, the one or more processors 100 can implement the functions of the various modules / units in the computer vision-based high-rigidity structure target point vibration identification system embodiment when executing the computer program 102, which is not limited here.
[0071] In one embodiment, the processor 100 can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.
[0072] In one embodiment, the storage 101 can be an internal storage unit of the electronic device, such as a hard disk or a memory of the electronic device. The storage 101 can also be an external storage device of the electronic device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, or the like equipped on the electronic device. Further, the storage 101 can include both an internal storage unit and an external storage device of the electronic device. The storage 101 is used to store a computer program and other programs and data required by the terminal. The storage 101 can also be used to temporarily store data that has been output or will be output.
[0073] Those skilled in the art can understand that, Figure 12 The person skilled in the art can understand that the principle block diagram shown in the above embodiment is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the terminal to which the scheme of the present application is applied. The specific terminal can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0074] The person skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer readable storage medium, and when executed, can include the processes of the above-mentioned embodiments. Any reference to memory, storage, operating database or other medium used in the embodiments provided by the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct RAM bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0075] It should be pointed out finally that the above embodiments are only used to illustrate the technical solutions of the present application, but not to limit the same; and although the present application has been described in detail with reference to the foregoing embodiments, it should be appreciated by those skilled in the art that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features thereof can be replaced equivalently; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for identifying vibrations of high-stiffness structural targets based on computer vision, characterized in that, The method includes: A training dataset for target vibration identification is constructed based on a bridge model. The EDVR model is trained based on the training dataset. The EDVR model is a video super-resolution network. The training of the EDVR model includes the training of a pyramid cascade deformable alignment network and the training of a spatiotemporal attention fusion network. Obtain the image sequence to be identified, perform super-resolution processing on the image sequence to be identified based on the trained EDVR model, and obtain the fitted ellipse based on the super-resolution processed image sequence to be identified. Based on the fitted ellipse, the displacement of the structural target point and the vibration frequency of the structure are identified, and the vertical physical displacement of the target point and the vertical vibration acceleration data of the target point at each moment are obtained. Based on the vertical physical displacement and the vertical vibration acceleration data, the deflection of the bridge and the frequency of the bridge are obtained.
2. The high-stiffness structural target vibration recognition method based on computer vision according to claim 1, characterized in that, A training dataset for target vibration identification was constructed based on a bridge model, including: Multiple target points were placed at different locations on the bridge model, and a camera was placed at the end of the bridge model to take pictures along the longitudinal direction of the bridge model to obtain structural vibration videos. The distance between the camera and each target point was different. The vibration video of the structure was converted into an .mxf format file and split into a training image sequence; Based on the training image sequence, the region of interest where each target point is located is selected to obtain the region of interest image sequence for each target point. Based on the image sequence of the region of interest, a training dataset for target vibration recognition is obtained.
3. The high-stiffness structural target vibration recognition method based on computer vision according to claim 2, characterized in that, Based on the image sequence of the region of interest, a training dataset for target vibration recognition is obtained, including: The image sequences of interest corresponding to each target point are subjected to blurring, brightness variation processing, symmetry processing, noise reduction, rotation processing, and resolution reduction processing to obtain a low-resolution dataset. Symmetry processing, noise reduction, and rotation processing are performed on the image sequences of interest corresponding to each target point to obtain a high-resolution dataset. A training dataset for target vibration identification is obtained based on the low-resolution dataset and the high-resolution dataset, wherein the low-resolution dataset and the high-resolution dataset correspond one-to-one.
4. The high-stiffness structural target vibration recognition method based on computer vision according to claim 1, characterized in that, The training process of the pyramid cascade deformable alignment network includes: Based on the training dataset, the current frame and the frames before and after it are downsampled by a factor of 2 to form a pyramid of feature maps with different resolutions; Calculate the offset between the target frame feature map and the feature maps of the preceding and following frames, and then upsample and fuse them with the offset matrix of the previous layer. The feature maps of the target frame before and after each frame are deformably convolved with the offset matrix of the current layer to form the alignment feature matrix of the current layer. Then, the matrix is upsampled and fused with the alignment feature matrix of the previous layer to form the offset of each layer. The offsets of each layer are deformably convolved with the alignment feature matrix of the first layer to obtain the alignment feature matrix of the current frame and the frames before and after.
5. The high-stiffness structural target vibration recognition method based on computer vision according to claim 4, characterized in that, The training process of the spatiotemporal attention fusion network includes: When training the temporal attention part, the correlation between the temporal feature maps of the alignment features of the target frame and the alignment features of the preceding and following frames is calculated to obtain the temporal attention weights of each frame. The temporal attention weights are multiplied element-wise with the original alignment features to obtain the temporal attention modulation features of each frame, and a fusion convolutional layer is used to fuse the temporal attention modulation features of each frame. When training the spatial attention part, the fusion result of the temporal attention modulation features of each frame is used as input. A pyramid structure is used to obtain the spatial feature map of each layer. Combined with upsampling, addition and multiplication processing, the feature map of fused spatial attention is obtained.
6. The high-stiffness structural target vibration recognition method based on computer vision according to claim 1, characterized in that, Based on the super-resolution processed image sequence to be identified, the fitted ellipse is obtained, including: Based on the super-resolution processed image sequence, determine the region of interest image to be identified; The image of the region of interest to be identified is processed into grayscale, and the grayscale processed image of the region of interest to be identified is subjected to Otsu adaptive threshold binarization to obtain a binary image, which is used to highlight the edge features of the target point. The Canny algorithm is used to perform edge detection on the target points in the binary image to obtain edge pixel information; Based on the edge pixel information, the edge contour is fitted using the least squares ellipse fitting method to obtain a fitted ellipse.
7. The high-stiffness structural target vibration recognition method based on computer vision according to claim 1, characterized in that, Based on the fitted ellipse, the displacement of the structural target point and the vibration frequency of the structure are identified, obtaining the vertical physical displacement of the target point at each moment and the vertical vibration acceleration data of the target point at each moment, including: Based on the coordinates of the center of the fitted ellipse, the pixel coordinates of the stationary points in the image sequence to be identified are calculated, and the vertical coordinates of the target pixel are obtained based on the coordinates of the center of the ellipse and the pixel coordinates of the stationary points. Perform a forward difference on the vertical coordinates of the target pixel to obtain the pixel displacement change of the target at each time step, and transform the pixel displacement change to obtain the vertical physical displacement of the target at each time step. The vertical coordinates of the target pixel are subjected to a second forward difference to obtain the vertical vibration acceleration data of the target at each moment.
8. A high-rigidity structural target vibration recognition system based on computer vision, characterized in that, The system is used to implement the steps of the computer vision-based high-rigidity structural target vibration recognition method according to any one of claims 1-7, and the system includes: The model training module is used to construct a training dataset for target vibration recognition based on the bridge model, and to train the EDVR model based on the training dataset. The EDVR model is a video super-resolution network, and the training of the EDVR model includes the training of a pyramid cascade deformable alignment network and the training of a spatiotemporal attention fusion network. The super-resolution processing module is used to acquire the image sequence to be identified, perform super-resolution processing on the image sequence to be identified based on the trained EDVR model, and obtain the fitted ellipse based on the super-resolution processed image sequence to be identified. The target vibration identification module is used to identify the structural target displacement and structural vibration frequency based on the fitted ellipse, obtain the vertical physical displacement of the target at each moment and the vertical vibration acceleration data of the target at each moment, and obtain the bridge deflection and bridge frequency based on the vertical physical displacement and the vertical vibration acceleration data.
9. A terminal, characterized in that, The terminal includes a memory, a processor, and a computer vision-based high-stiffness structural target vibration recognition program stored in the memory and executable on the processor. When the processor executes the computer vision-based high-stiffness structural target vibration recognition program, it implements the steps of the computer vision-based high-stiffness structural target vibration recognition method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a high-stiffness structural target vibration recognition program based on computer vision, and the high-stiffness structural target vibration recognition program based on computer vision implements the steps of the high-stiffness structural target vibration recognition method based on computer vision as described in any one of claims 1-7 on the computer-readable storage medium.
Citation Information
Patent Citations
Video super-resolution method and system based on data simulation, equipment and storage medium
CN113469884A
Cited By
Steel tower vibration mode and displacement identification method based on video tracking
CN122176644A