Structural Displacement Visual Recognition Method and System Incorporating a 3D Deformable Model
Through deep learning technology and three-dimensional deformable grid model, combined with semantic segmentation and dense optical flow estimation, consumer-grade monocular cameras are used to identify three-dimensional structural displacement, solving the problems of high-cost sensors and low-precision consumer-grade cameras in the existing technology, and achieving high-precision and low-cost structural health monitoring.
Patent Information
- Application Number
- CN202510141794.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-08
AI Technical Summary
The existing structural health monitoring technology relies on expensive high-precision sensors, has high installation and maintenance costs, and is difficult to promote and apply in large-scale engineering structures. Consumer-grade monocular cameras have poor imaging quality, large environmental interference, and insufficient recognition accuracy in micro displacement recognition.
Deep learning technology is used to combine three-dimensional deformable grid model, semantic segmentation, dense optical flow estimation and cross-attention mechanism to identify three-dimensional structural displacements through consumer-grade monocular cameras, overcome picture distortion and environmental interference, and improve the accuracy of displacement recognition.
It significantly improves the identification accuracy of small displacements of engineering structures, reduces equipment costs, does not require contact monitoring, is easy to install and maintain, and is suitable for health monitoring of large-scale engineering structures.
Smart Images

Figure CN119648922B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of structural health monitoring, and particularly relates to a method and system for visually identifying structural displacement by integrating a three-dimensional deformable model. By using computer vision technology and deep learning algorithms, and using a consumer-grade monocular camera, combined with three-dimensional deformable mesh modeling and dense optical flow estimation, the method realizes the accurate identification of minute displacements of buildings, and is applicable to the health monitoring and vibration monitoring of high-rise engineering structures. Background Art
[0002] Structural deformation measurement can timely detect abnormal deformations of structures, prevent structural failures and collapses, and thus ensure the safety and stability of engineering structures. For example, engineering structures such as buildings and bridges are subjected to various external forces and environmental effects during their service life, such as earthquakes, winds, vehicle loads, temperature and humidity changes, etc. These loads and environmental effects will cause structural deformations. If the deformation exceeds the allowable range of the structural design, it may lead to the failure or even collapse of the engineering structure, resulting in serious personnel and property losses. Through a deformation monitoring system, the structural deformation can be monitored in real time and early warnings can be given in a timely manner to ensure the safety and stability of the engineering structure and prevent the occurrence of disaster events.
[0003] Currently, structural health monitoring technologies mainly rely on physical sensors, such as strain gauges, accelerometers, laser displacement sensors, GPS, etc. These sensors can provide high-precision displacement and vibration data, but they have relatively high costs in terms of installation and maintenance. Especially in large-scale engineering structures or infrastructure such as bridges, the economic cost of deploying these sensors is relatively high. In addition, these sensor systems usually require complex wiring, calibration and maintenance, and are particularly vulnerable to environmental factors such as temperature and humidity during long-term monitoring, thus affecting the measurement accuracy and long-term data stability. Traditional vision sensing technologies, such as using high-end industrial cameras or lidar for displacement monitoring, although they can provide a non-contact measurement method, their equipment costs are high, data processing is complex, and there are certain difficulties in popularizing and applying them in the health monitoring of large-scale structures. This has prompted researchers to start exploring the use of relatively inexpensive and easily deployable consumer-grade vision sensing devices, such as monocular cameras or mobile phone cameras, for engineering structure displacement monitoring and analysis. The displacement monitoring method based on a monocular camera has attracted wide attention in recent years. Monocular cameras have the advantages of low price and easy installation, and can perform non-contact visual monitoring without disturbing the normal use of buildings. However, due to the imaging principle and hardware limitations of monocular cameras, their application in structural health monitoring still faces multiple technical challenges. First, the imaging quality of consumer-grade monocular cameras is lower than that of industrial cameras. Limited by the resolution and focal length, it is difficult to accurately capture minute displacements of engineering structures. Second, the monocular vision system lacks depth information and requires additional algorithms to compensate its perception ability of structural motion in three-dimensional space, which increases the complexity of data processing.
[0004] At present, there are still some deficiencies in the research on displacement recognition based on monocular consumer cameras in structural health monitoring. First, the existing consumer monocular cameras have low pixel counts, resulting in poor-quality captured videos, which limits the accuracy and reliability of displacement recognition. Second, displacement recognition based on computer vision is easily interfered by factors such as the background, introducing a large amount of noise in the image and affecting the recognition accuracy. In addition, consumer cameras have limited contrast. In environments with drastic weather or lighting changes, especially in high-contrast scenes, the image quality will significantly decline, resulting in the loss of some displacement information and affecting the recognition effect. Therefore, to solve the above problems, this paper intends to use a monocular camera and deep learning technology, make full use of the limited perspective of the monocular camera, establish a three-dimensional deformable mesh model, perform dense optical flow estimation, and generate a mask in cooperation with semantic segmentation to control the influence of the background, weather, and lighting on the image quality. Use a deep learning neural network for pose estimation, and innovatively transfer the cross-attention mechanism applied to medical images to the structural dense displacement estimation method to achieve precise displacement recognition of specific engineering structures, and use a consumer monocular camera to achieve high-precision monitoring effects.
[0005] In summary, although monocular cameras have broad application prospects in the field of structural health monitoring, they still face problems such as poor imaging quality, large environmental interference, and insufficient recognition accuracy in micro-displacement recognition. This invention is proposed based on the solution of these problems, aiming to significantly improve the micro-displacement recognition accuracy based on consumer monocular cameras through advanced technical means such as deep learning, three-dimensional modeling, and dense optical flow estimation, and provide a low-cost, efficient, and easy-to-deploy solution for the health monitoring of engineering structures. Summary of the Invention
[0006] The purpose of the present invention is to address the limitations in the existing structural health monitoring (SHM) systems, where it is difficult to recognize the displacement of engineering structures due to issues such as the inability of monocular devices to recognize three-dimensional structures, the high cost of high-precision cameras, the low monitoring accuracy of consumer cameras, and environmental (background, weather, lighting, etc.) interference, especially in micro-displacement recognition, where existing methods are difficult to meet the actual needs. To overcome these technical challenges, a method for recognizing the displacement of engineering structures based on a monocular consumer camera is proposed. This method combines deep learning technology, three-dimensional deformable mesh modeling, semantic segmentation, dense optical flow estimation, and a cross-attention mechanism to use a monocular consumer camera to recognize the three-dimensional structural displacement, overcome problems such as image distortion and environmental interference, and significantly improve the accuracy of displacement recognition, especially having high sensitivity and accuracy in recognizing the micro-displacement of engineering structures.
[0007] In view of the above deficiencies or improvement requirements of the prior art, as the first aspect of the present invention, the present invention provides a structural displacement visual recognition method integrating a three-dimensional deformable model, including:
[0008] S1. Collect engineering structure vibration video data through a consumer-grade monocular camera, establish a three-dimensional deformable mesh model, and generate pose rendering images and displacement value data;
[0009] S2. Remove the interference of environmental noise on displacement recognition;
[0010] S3. Estimate the dense optical flow data in the rendering image through an optical flow model, and construct an optical flow feature dataset containing pose information;
[0011] S4. Based on the structural dense displacement estimation method of the cross-attention mechanism, through the fusion of local and global features, output the model displacement prediction value. The specific steps are as follows:
[0012] Input two groups of data sequences, the overall structure image sequence and the local structure image sequence. After multi-layer convolution operations of the convolutional layer, generate the corresponding overall features and local features ;
[0013] The overall features and local features Generate their feature query vectors, key vectors, and value vectors through linear transformation, and establish dynamic weights between the two features based on these vectors to obtain two fusion features and ;
[0014] Determine the loss function, and perform global average pooling on the two obtained fusion features in sequence, and then output the pose estimation value through a multi-layer perceptron.
[0015] Further, in step S1, the displacement value is randomly generated and obeys a uniform distribution, and the extreme values of the displacement value can be adjusted according to the actual engineering project requirements and experimental requirements.
[0016] Further, in step S2, the specific method for removing the interference of environmental noise on displacement recognition is as follows:
[0017] Generate a semantic segmentation mask of the engineering structure through a semantic segmentation model to remove the interference of environmental noise on displacement recognition. Its calculation formula is:
[0018]
[0019] Apply the mask to the image. Its calculation formula is:
[0020]
[0021] In the formula, is the original image, is the structural area in the zero-displacement state at time zero, is the image after denoising.
[0022] Furthermore, the specific method of the step S3 includes:
[0023] Estimate the dense optical flow data in the rendered image through an optical flow model;
[0024] Extract the most significant local feature regions from the overall displacement optical flow data , then magnify and preprocess them to obtain local displacement optical flow data ;
[0025] Obtain the displacement optical flow data after fusing local features and overall features .
[0026] Furthermore, the specific method for estimating the dense optical flow data in the rendered image through the optical flow model is:
[0027] For two frames of images at time zero and time t, the calculation formula is:
[0028]
[0029] In the formula, represents the image at time zero, represents the image at time respectively represent the components of the optical flow at time zero in the and directions, respectively represent the components of the optical flow at time and directions;
[0030] According to the Taylor expansion, expand at to the first order, and the following can be obtained:
[0031]
[0032] Perform Taylor expansion on this equation and ignore the high-order terms to obtain the optical flow constraint equation:
[0033]
[0034] In the formula, , represent the components of the optical flow in and Component in the
[0035] Calculate the optical flow between the 0th frame and the frame, and the average optical flow between the 0th frame and the frame can be expressed as: , which can be expressed as:
[0036]
[0037] Substitute this definition into the optical flow constraint equation to obtain the optical flow expression for the multi-frame interval as:
[0038]
[0039] In the formula, and represent the gradients of the pixel brightness of the 0th frame in the and directions;
[0040] represents the temporal gradient of the image.
[0041] Furthermore, the calculation method of the fused displacement optical flow data in step S3 is: The calculation method is:
[0042] The semantic segmentation model is improved based on the model, and the optical flow model is improved based on the ; among which, the calculation process of the displacement optical flow data of the overall displacement of the th frame image can be expressed as:
[0043]
[0044]
[0045] The calculation process of the displacement optical flow data of the local displacement of the th frame image can be expressed as:
[0046]
[0047]
[0048] Expression of the fused dataset:
[0049]
[0050] In the formula, respectively represent the components of the optical flow in the and directions at time represents the output overall displacement mask, Indicates the output local displacement mask, Indicates the operation of extracting the mask using the U-Net network model, Indicates the optical flow image, Indicates the fused optical flow data after mask denoising, Indicates the overall displacement optical flow data after mask denoising, Indicates the overall local optical flow data after mask denoising, Indicates the optical flow model function, whose input is the pose rendering image and the output is the optical flow data of the corresponding structure.
[0051] Furthermore, in step S4, the method of establishing dynamic weights between these two features according to these vectors to obtain the fused feature is as follows:
[0052] The overall feature and the local feature After being input into the network, query, key, and value vectors are generated through linear transformation; taking the overall feature as an example, it generates corresponding query, key, and value vectors through a set of learnable weight matrices:
[0053]
[0054] Similarly, the local feature can also generate query, key, and value vectors:
[0055]
[0056] Among them, respectively represent the overall feature and the local feature; respectively represent the query vector, key vector, and value vector of the overall feature; respectively represent the query vector, key vector, and value vector of the local feature; are the learnable weight matrices for generating query vectors, key vectors, and value vectors; these matrices can automatically learn the most suitable parameters from the training data of the model to generate vector representations that contribute to displacement estimation on different features;
[0057] In the cross-attention mechanism, the similarity or correlation is calculated between the generated query vector and key vector, and then the attention weight is generated; the calculation of similarity uses the vector dot product and is scaled to maintain a stable gradient. The specific formula is:
[0058]
[0059] In the formula, represents the dimension of the key vector, which is used for scaling to prevent the vector dot product value from being too large; respectively represent the query vector, key vector, and value vector, represents matrix transpose, represents the activation function, represents the calculation of the cross-attention mechanism;
[0060] The core of this formula lies in calculating the similarity between the query vector and the key vector, dividing by the scaling factor and passing it through to normalize it into attention weights; then, through the weighted sum of value vectors, to extract features; in the cross-attention mechanism, this calculation process is carried out bidirectionally between the global features and local features to capture detailed information; using the global features of the query vector can be used to focus on the key position information of local features:
[0061]
[0062] Query of local features is used to extract the context information of global features:
[0063]
[0064] In the formula, respectively represent the query vector, key vector, and value vector of global features; respectively represent the query vector, key vector, and value vector of local features.
[0065] Furthermore, the loss function expression in step S4 is:
[0066]
[0067] In the formula, is the actual displacement value data, is the model-predicted displacement value data, is the total number of samples, is the regularization coefficient; represents the regularization weight, represents the i-th actual displacement value data, represents the i-th model-predicted displacement value data.
[0068] As the second aspect of the present invention, there is also provided a structural displacement visual recognition system integrating a three-dimensional deformable model, including:
[0069] A three-dimensional deformable grid model construction unit for collecting engineering structure vibration video data through a consumer-grade monocular camera, establishing a three-dimensional deformable grid model, and generating pose rendering images and displacement value data;
[0070] A data preprocessing unit for removing the interference of environmental noise on displacement recognition;
[0071] An optical flow feature dataset construction unit for estimating dense optical flow data in a rendered image through an optical flow model and constructing an optical flow feature dataset containing pose information;
[0072] A model learning unit for a structure - dense displacement estimation method based on a cross - attention mechanism. Through the fusion of local and global features, it outputs a model displacement prediction value. The specific steps are as follows:
[0073] Input two groups of data sequences, the overall structure image sequence and the local structure image sequence. After multi - layer convolution operations of the convolutional layer, corresponding overall features and local features are generated respectively;
[0074] The overall features and local features are linearly transformed to generate their feature query vectors, key vectors, and value vectors, and dynamic weights are established between the two features based on these vectors to obtain two fused features and ;
[0075] Determine the loss function, and perform global average pooling on the two obtained fused features successively, and then output the pose estimation value through a multi - layer perceptron.
[0076] As a third aspect of the present invention, it also relates to a computer - readable storage medium on which a computer program is stored. The computer program is executed by a processor to perform the above - mentioned method for visually recognizing the structural displacement of a three - dimensional deformable model.
[0077] Generally speaking, compared with the prior art through the above - mentioned technical solutions conceived by the present invention, the following beneficial effects can be achieved:
[0078] 1. The method for visually recognizing the structural displacement of a three - dimensional deformable model of the present invention introduces feature fusion methods such as a cross - attention mechanism to fuse features from different sources (such as overall feature maps, local feature maps, optical flow data, etc.). This mechanism can fully explore the correlations between different features, enabling the model to comprehensively judge structural displacement from multiple dimensions such as global and local, space and motion, overcoming the limitations of traditional methods that only analyze from a single angle. These data respectively reflect the state of the engineering structure from different angles. For example, the overall image provides the macroscopic structural pose, the local image captures the details of key parts, the optical flow data reflects pixel - level motion information, and the semantic segmentation mask is used to exclude environmental interference. The fusion of such multi - source data is an innovative idea in the field of structural displacement recognition and can more comprehensively describe the displacement of the engineering structure.
[0079] 2. The structural displacement visual recognition method integrating a three-dimensional deformable model of the present invention realizes the recognition of three-dimensional vibration displacement through a three-dimensional deformable grid model rendering, a cross-attention mechanism, and a dual-attention mechanism, completing the transformation of two-dimensional plane monitoring to three-dimensional space monitoring only through a consumer-grade monocular camera. Compared with traditional high-precision sensors and industrial cameras, the equipment cost is greatly reduced, and this method does not require contact monitoring of the engineering structure, reducing potential damage to the engineering structure, facilitating installation and maintenance, and being applicable to the health monitoring of large-scale engineering structures. BRIEF DESCRIPTION OF THE DRAWINGS
[0080] Figure 1 It is a flowchart of the method according to a preferred embodiment of the present invention;
[0081] Figure 2 It is a diagram of the method steps according to a preferred embodiment of the present invention;
[0082] Figure 3 It is a schematic diagram of the details and dimensions of the physical model according to a preferred embodiment of the present invention;
[0083] Figure 4 It is a schematic diagram of the three-dimensional deformable grid model according to a preferred embodiment of the present invention;
[0084] Figure 5 It is a vibration simulation diagram according to a preferred embodiment of the present invention;
[0085] Figure 6 It is an overall structure pose rendering diagram according to a preferred embodiment of the present invention;
[0086] Figure 7 It is a local structure pose rendering diagram according to a preferred embodiment of the present invention;
[0087] Figure 8 It is a schematic diagram of the overall structure semantic segmentation mask according to a preferred embodiment of the present invention;
[0088] Figure 9 It is a schematic diagram of the local structure semantic segmentation mask according to a preferred embodiment of the present invention;
[0089] Figure 10 It is a visualization diagram of the overall structure displacement image optical flow data according to a preferred embodiment of the present invention;
[0090] Figure 11 It is a visualization diagram of the local structure displacement image optical flow data according to a preferred embodiment of the present invention;
[0091] Figure 12 It is a schematic diagram of the loss function curve during the training process of the invalid data detection model according to a preferred embodiment of the present invention;
[0092] Figure 13Schematic diagram of the result comparison between the actual vibration data and the predicted values of the model output in the preferred embodiment of the present invention;
[0093] Figure 14 Schematic diagram of the system unit in the preferred embodiment of the present invention. Detailed implementation manners
[0094] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0095] Embodiment 1
[0096] Please refer to Figure 1 , Embodiment 1 of the present invention provides a structural displacement visual recognition method integrating a three-dimensional deformable model, including:
[0097] S1. Collect engineering structure vibration video data through a consumer-grade monocular camera, establish a three-dimensional deformable mesh model, and generate pose rendering images and displacement value data;
[0098] S2. Remove the interference of environmental noise on displacement recognition;
[0099] S3. Estimate the dense optical flow data in the rendering image through an optical flow model, and construct an optical flow feature dataset containing pose information;
[0100] S4. Based on the structural dense displacement estimation method of the cross-attention mechanism, through the fusion of local and global features, output the model displacement prediction value, and the specific steps are as follows:
[0101] Input two groups of data sequences, the overall structure image sequence and the local structure image sequence. After multi-layer convolution operations of the convolutional layer, generate the corresponding overall features and local features ;
[0102] The overall features and local features are linearly transformed to generate their feature query vectors, key vectors and value vectors, and dynamic weights are established between the two features based on these vectors to obtain two fusion features and ;
[0103] Determine the loss function, perform global average pooling on the two obtained fusion features in sequence, and then output the pose estimation value through a multi-layer perceptron.
[0104] For a further expansion of the above method steps of this embodiment, please refer to Figure 2 , the purpose of this embodiment is to propose an engineering structure displacement recognition method based on deep learning and a three-dimensional deformable mesh model, which is applied to the processing and analysis of predicting and recognizing the displacement of an engineering structure by a consumer-grade monocular camera.
[0105] The core of the present invention is to realize the recognition of three-dimensional vibration displacement based on three-dimensional deformable mesh model rendering, cross-attention mechanism and dual-attention mechanism. This method is mainly divided into two steps. The first step is the training of the model. In this process, a three-dimensional deformable mesh model is required for modeling and vibration simulation. The output pose rendering image is subjected to semantic segmentation and optical flow conversion, and then input into the model so that the model can learn the corresponding structural displacements of different poses. The second step is the displacement recognition of the observed video, that is, using the trained model to process the captured video data, calculate and predict the displacement data of the captured engineering structure, so as to realize the displacement recognition of the engineering structure.
[0106] Specifically, step S1 is specifically as follows:
[0107] (1a) Create a 1:1 model of the facade and feature structure of the engineering structure to be monitored. Insert and bind "bones" in the "columns" of the model, adjust the material, length and related attributes, and set the hierarchical relationship of "parent class" and "subclass". At the same time, to limit the horizontal and vertical displacement ranges of the floors, a follow effect and constraint conditions are added, and the common section S and the control section S are set c to achieve precise control of the deformation of the model;
[0108] (1b) Write a script in the built-in driver of Blender so that it can read the data in the displacement file and assign the displacement values at the corresponding positions to the key frame control sections, so as to realize the dynamic displacement simulation of the model. The displacement data file can be in Excel format, and the displacement values in it are randomly generated and follow a uniform distribution. The extreme values can be adjusted according to the actual engineering project requirements and experimental requirements.
[0109] (1c) Render all key frames and output the final simulated pose rendering image. The output rendering image has a width of 384 and a height of 704.
[0110] Specifically, step S2 uses a deep learning-based semantic mask and dense optical flow estimation method to generate a semantic segmentation mask of the engineering structure through a pre-trained semantic segmentation model to remove the interference of environmental noise on displacement recognition. Specifically:
[0111] The semantic segmentation model generates a semantic segmentation mask of the engineering structure to remove the interference of environmental noise on displacement recognition. Its calculation formula is:
[0112]
[0113] Apply a mask to an image, and its calculation formula is:
[0114]
[0115] In the formula, is the original image, is the structural area in the state of zero displacement at time 0, is the image after denoising.
[0116] Specifically, the specific steps of step S3 include:
[0117] Estimate the dense optical flow data in the rendered image through an optical flow model;
[0118] Extract the most significant local feature regions from the overall displacement optical flow data and then perform magnification and preprocessing to obtain local displacement optical flow data ;
[0119] Obtain the displacement optical flow data after fusing local features and overall features .
[0120] (3a) Estimate the dense optical flow data in the rendered image through an optical flow model, and construct an optical flow feature dataset containing pose information. For two frames of images at time 0 and time t, its calculation formula is:
[0121]
[0122] In the formula, represents the image at time 0, represents the image at time; respectively represent the components of the optical flow at time 0 in the and directions, respectively represent the components of the optical flow at time and directions;
[0123] According to the Taylor expansion, expand at to the first order, and the following can be obtained:
[0124]
[0125] Perform Taylor expansion on this equation and ignore the high-order terms to obtain the optical flow constraint equation:
[0126]
[0127] In the formula,
[0128] and represent the gradients of the pixel luminance of the 0th frame in the and directions;
[0129] represents the temporal gradient of the image;
[0130] , are respectively the components of the optical flow in the and directions.
[0131] Calculate the optical flow between the 0th frame and the th frame, and the average optical flow between the 0th frame and the th frame , which can be expressed as:
[0132]
[0133] Substitute this definition into the optical flow constraint equation to obtain the optical flow expression at the multi-frame interval as:
[0134]
[0135] (3b) Construct an optical flow feature dataset containing pose information, extract the most significant local feature regions from the overall displacement optical flow data , then perform magnification and preprocessing to obtain the local displacement optical flow data , which can be expressed by the following formula:
[0136]
[0137] Specifically, the semantic segmentation model is improved based on the U-Net model, and the optical flow model is improved based on Flownet2. Among them, the calculation process of the displacement optical flow data of the overall displacement th frame image can be expressed as:
[0138]
[0139]
[0140] The calculation process of the displacement optical flow data of the local displacement th frame image can be expressed as:
[0141]
[0142]
[0143] The expression of the dataset after fusion:
[0144]
[0145] In the formula, respectively represent the optical flow at time and the components in the represents the overall displacement mask of the output, represents the local displacement mask of the output, represents the operation of extracting the mask using the U-Net network model, represents the optical flow image, represents the fused optical flow data after mask denoising, represents the overall displacement optical flow data after mask denoising, represents the overall local optical flow data after mask denoising, represents the optical flow model function, whose input is the pose rendering map and the output is the optical flow data of the corresponding structure.
[0146] Further expand step S4, and its specific steps are:
[0147] (4a) In the overall architecture, the model first inputs two sets of data sequences, namely the overall structure image sequence and the local structure image sequence. After the multi-layer convolution operation of EfficientNet-B0, these two sets of images respectively generate the corresponding overall feature map and local feature map. At this time, the overall feature map contains the macroscopic pose information of the engineering structure, while the local feature map contains the subtle local pose changes. The expression of the optical flow sequence input to the model is:
[0148]
[0149] (4b) The core of the cross-attention mechanism is to establish dynamic weights between different features through the interaction of the query vector Q, the key vector K, and the value vector V, so as to more accurately capture the subtle local changes. First, the overall feature and the local feature are converted into three main vectors, namely the query vector Q, the key vector K, and the value vector V. These three vectors work together to achieve feature interaction and information capture. The overall feature and the local feature After input into the network, they generate query, key, and value vectors through linear transformation. Taking the overall feature as an example, it generates the corresponding query, key, and value vectors through a set of learnable weight matrices:
[0150]
[0151] Similarly, local features can also generate query, key, and value vectors:
[0152]
[0153] where, represent the global feature and local feature respectively; represent the query vector, key vector, and value vector of the global feature respectively; represent the query vector, key vector, and value vector of the local feature respectively; are learnable weight matrices for generating query vectors, key vectors, and value vectors; these matrices can automatically learn the most suitable parameters from the training data of the model to generate vector representations that contribute to displacement estimation on different features;
[0154] In the cross - attention mechanism, the generated query vector and key vector calculate the similarity or correlation between them, and then generate attention weights. The calculation of similarity uses the vector dot - product and scaling processing to maintain a stable gradient. The specific formula is:
[0155]
[0156] In the formula, represents the dimension of the key vector, which is used for scaling to prevent the vector dot - product value from being too large; represent the query vector, key vector, and value vector respectively, represents matrix transpose, represents the activation function, represents the calculation of the cross - attention mechanism;
[0157] The core of this formula is to calculate the similarity between the query vector and the key vector, divide by the scaling factor and normalize it to attention weights through Then, the weighted sum of the value vectors is used to extract features. In the cross - attention mechanism, this calculation process is carried out bidirectionally between the global feature and the local feature to capture detailed information. For example, using the query vector of the global feature can be used to focus on the key position information of the local feature:
[0158]
[0159] Similarly, the query of the local feature can also be used to extract the context information of the global feature:
[0160]
[0161] Through two-way cross-attention calculation, the global and local features not only retain their respective important information but also complement the details not captured in each other. This two-way information interaction design ensures that the model can accurately capture displacement changes at different scales when extracting features.
[0162] (4c) The fused features obtained by the cross-attention mechanism are then first subjected to global average pooling, and the expression is as follows:
[0163]
[0164]
[0165] The obtained sequence is passed to the multi-layer perceptron (MLP) layer. The MLP layer contains several fully connected layers and non-linear activation functions, which can be used to further process the features and abstract complex relationships. Through non-linear transformation, the high-dimensional fused features are mapped to a low-dimensional space, thereby effectively compressing information redundancy and improving the accuracy of feature expression. After being processed by the MLP layer, the model outputs a pose estimation value, representing the displacement state of the overall structure or local structure in a specific environment or perspective. The predicted displacement value The expression is as follows:
[0166]
[0167]
[0168] In addition, in this embodiment, the expression of the loss function in step S4 is:
[0169]
[0170] In the formula, is the actual displacement value data, is the displacement value data predicted by the model, is the total number of samples, is the regularization coefficient.
[0171] This embodiment also provides examples in actual application scenarios, as follows:
[0172] This embodiment is completed by a computer program written in the Python language, and the data used is the data provided by the Third International Structural Health Monitoring Competition. The following specific examples illustrate the effects of the present invention.
[0173] First, a three-dimensional deformable mesh model was created using Blender software. The physical model was set up on a shaking table at Tongji University for dynamic tests under unidirectional seismic action. The main structure of this model is 6.5 meters high, with a floor height of 1.5 meters for each floor and a base height of 0.35 meters. The column spacing along the vibration direction is 3.0 meters, and the column spacing perpendicular to the vibration direction is 1.5 meters. The size of the shaking table is 4.0 meters × 4.0 meters, with a load-bearing capacity of 25 tons. During the test, various monitors were installed on each floor of the model. Absolute acceleration and relative displacement were measured conventionally using accelerometers and linear displacement sensors. The test camera (SONY HDR - PJ220) was arranged 3.0 meters in front of the model frame, and a fixed camera (Black Magic Design URSA mini 4K) was set up 8.0 meters away from the frame model. Referring to the dimensions of the model and the camera layout perspective of the third International Structural Health Monitoring Competition, a 1:1 scale model was established in Blender. The model is as Figure 3 , Figure 4 shown.
[0174] "Bones" were inserted and bound in the "columns" of the model, and the material, length, and related properties were adjusted. The hierarchical relationship between the "parent class" and "subclass" was set. At the same time, to limit the horizontal and vertical displacement ranges of the floors, a follow effect and constraint conditions were added. A script was written in the built-in driver of Blender to enable it to read the data in the displacement file and assign the displacement values at the corresponding positions to the keyframe control sections, thereby realizing the dynamic displacement simulation of the model. The displacement data file can be in Excel format, and the displacement values in it are randomly generated and follow a uniform distribution. The extreme values can be adjusted according to the actual engineering project requirements and experimental requirements. All keyframes were rendered to output the final simulated renderings. The output renderings have a width of 384 and a height of 704. A total of 80,001 frames of pose renderings were generated, including one frame of the initial pose rendering with 0 displacement at time 0. The most significant local feature regions were extracted, then magnified and preprocessed to obtain local structure pose renderings.
[0175] The generated overall structure and local structure pose rendering images were input into a deep learning-based semantic mask and dense optical flow estimation model to generate semantic segmentation masks and participate in the calculation, generating an overall structure and local structure optical flow dataset for training. A total of 80,001×2 data were generated, which were used for model training respectively. The actual vibration data provided by the third International Structural Health Monitoring Competition were preprocessed in the same way and used as the test set. The specific data division method is shown in Table 1.
[0176] Table 1 Datasets used in the experiment
[0177]
[0178] When sampling sequence data using a sliding window, it is necessary to set an appropriate sliding window size and sliding step. At the same time, to make the model converge better, learning rate decay is used in this model. Specifically, the learning rate is adjusted to 0.1 times the original value after each iteration. Before training the model, the detection data should be preprocessed, that is, each detection index data is standardized.
[0179] Some parameter settings during the experiment are shown in Table 2.
[0180] Table 2 Model-related parameter settings
[0181]
[0182] The model is trained using the divided training set and validation set. The loss function curve during training is as Figure 12 shown. During the training process, the loss functions of the training set and the validation set rapidly drop below 0.03 after 4 iterations. As the iteration progresses, the loss functions of both tend to stabilize after 40 iterations, that is, the model reaches convergence. Finally, the loss function stabilizes at around 0.08. In addition, the overall trends of the loss function curves of the training set and the validation set are the same, which conforms to the learning curve law during the deep learning training process, and the model does not show overfitting.
[0183] After the model training is completed, it is necessary to deeply analyze the predicted values and actual true values of the model to ensure that the model has good generalization ability and prediction accuracy. Multiple evaluation metrics are used to measure the performance of the model, including mean squared error (MSE), mean absolute error (MAE), coefficient of determination (R²), and root mean squared error (RMSE). The evaluation results are shown in the following table:
[0184] Table 3 Evaluation results of multiple metrics
[0185]
[0186] Based on these metrics, it can be seen that the error between the predicted values and the actual values of the model is very small, the fitting effect is excellent, and it has good accuracy and robustness.
[0187] Please refer to Figure 13 , in the result comparison of the curve fitting between the actual vibration data and the predicted values output by the model, it can fit well for both large displacements and small displacements. It realizes the prediction of displacements at the order of 10^-2 mm at about 10 m, which is equivalent to realizing the displacement identification with centimeter-level accuracy for engineering structures at 1000 m, and has the possibility of practical application value.
[0188] As can be seen from the above analysis, the engineering structure displacement recognition method based on deep learning and three-dimensional deformable mesh model used in this paper adopts a consumer-grade monocular camera and realizes the recognition of tiny displacements of engineering structures through computer vision and deep learning technologies. By establishing a corresponding three-dimensional deformable mesh model, 80,000 pose rendering images of random vibrations are output as the training set and validation set for model testing. The model is trained by setting appropriate parameters. The loss function in the training process shows that the model converges after 4 iterations and the training process is stable. Using various error analysis methods including mean square error (MSE), mean absolute error (MAE), coefficient of determination (R²), and root mean square error (RMSE) and performing result curve fitting, the results show that the mean square error (MSE) of the model is 0.08; the mean absolute error (MAE) is 0.19; the coefficient of determination (R²) is 0.99; the root mean square error (RMSE) is 0.28. The detection effect of the model is good and stable, realizing the prediction of displacements at the order of 10⁻² millimeters at a distance of about 10 m, which is equivalent to realizing the displacement recognition with centimeter-level accuracy for engineering structures at 1000 m, having the possibility of practical application value.
[0189] Example 2
[0190] Please refer to Figure 14 , this Example 2 provides a structural displacement visual recognition system integrating a three-dimensional deformable model, including:
[0191] A three-dimensional deformable grid model construction unit for collecting engineering structure vibration video data through a consumer-grade monocular camera, establishing a three-dimensional deformable grid model and generating pose rendering images and displacement value data;
[0192] A data preprocessing unit for removing the interference of environmental noise on displacement recognition;
[0193] An optical flow feature dataset construction unit for estimating dense optical flow data in the rendering image through an optical flow model and constructing an optical flow feature dataset containing pose information;
[0194] A model learning unit for the structural dense displacement estimation method based on the cross-attention mechanism, and outputting the model displacement prediction value through the fusion of local and global features. The specific steps are as follows:
[0195] Input two groups of data sequences, the overall structure image sequence and the local structure image sequence. After multi-layer convolution operations of the convolutional layer, the corresponding overall features and local features ;
[0196] The overall features and local features Generate its feature query vector, key vector, and value vector through linear transformation, and establish dynamic weights between the two features based on these vectors to obtain two fused features and ;
[0197] Determine the loss function, perform global average pooling on the two obtained fused features successively, and then output the pose estimation value through a multi-layer perceptron.
[0198] Embodiment 3
[0199] Embodiment 3 of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of any one of the above-mentioned structure displacement visual recognition methods for fusing three-dimensional deformable models can be implemented.
[0200] The computer-readable storage medium may include: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.
[0201] For the introduction of the computer-readable storage medium provided in the present application, please refer to the above method embodiments, and the present application will not elaborate herein.
[0202] Those skilled in the art can easily understand that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for visual recognition of structural displacements by integrating a three-dimensional deformable model, characterized in that: include: S1. Collect engineering structure vibration video data through a consumer-grade monocular camera, build a 3D deformable mesh model and generate pose rendering images and displacement value data; S2. Remove the interference of environmental noise on displacement recognition; S3. Estimating dense optical flow data in the rendered image through the optical flow model, and constructing an optical flow feature dataset containing pose information; S4. A structured dense displacement estimation method based on a cross-attention mechanism outputs the model displacement prediction value by fusing local and global features. The specific steps are as follows: Input two sets of data sequences: the overall structure image sequence and the local structure image sequence. After multi-layer convolution operations of the convolution layer, the corresponding overall features are generated respectively. With local features ; Overall characteristics and local features After linear transformation, the feature query vector, key vector and value vector are generated, and dynamic weights are established between the two features based on these vectors to obtain two fusion features. as well as ; Determine the loss function, perform global average pooling on the two fusion features obtained, and then output the pose estimation value through a multi-layer perceptron.
2. According to the method of claim 1, the structure displacement visual recognition method integrating a three-dimensional deformable model is characterized in that: The displacement value in step S1 is randomly generated and obeys uniform distribution, and the extreme value of the displacement value can be adjusted according to the actual engineering project requirements and experimental requirements.
3. According to the method of claim 1, the structure displacement visual recognition method integrating a three-dimensional deformable model is characterized in that: The specific method for removing the interference of environmental noise on displacement recognition in step S2 is: The semantic segmentation mask of the engineering structure is generated through the semantic segmentation model to remove the interference of environmental noise on displacement recognition. The calculation formula is: , Apply Mask For images, the calculation formula is: , In the formula, is the original image, is the structural area at time 0 and 0 displacement state, is the denoised image.
4. According to claim 3, a method for visually recognizing structural displacements by integrating a three-dimensional deformable model, characterized in that: The specific method of step S3 includes: Estimate dense optical flow data in rendered images through optical flow models; From the overall displacement optical flow data The most significant local feature area is extracted, and then enlarged and preprocessed to obtain local displacement optical flow data ; Obtain the displacement optical flow data after fusion of local features and overall features .
5. According to claim 4, a method for visually recognizing structural displacements by integrating a three-dimensional deformable model, characterized in that: The dense optical flow data in the rendered image is estimated by the optical flow model, and the specific method is as follows: For two frames of images at time 0 and time t, the calculation formula is: , In the formula, represents the image at time 0, express images of moments; Respectively represent the optical flow at time 0 and The weight in direction, Respectively The light flows at all times and Directional weight; According to Taylor expansion, exist Performing a first-order expansion at , we can obtain: , Taylor expansion of the equation and ignoring higher-order terms gives the optical flow constraint equation: , In the formula, , Denotes the optical flow in and Directional weight; Calculate the 0th and 1st frames The optical flow between frames, frame 0 and frame Average optical flow between frames , which can be expressed as: , Substituting this definition into the optical flow constraint equation, the optical flow expression at multi-frame intervals is obtained: , In the formula, and Indicates the pixel brightness of frame 0 and Directional gradient; Represents the temporal gradient of an image.
6. According to claim 4, a method for visually recognizing structural displacements by integrating a three-dimensional deformable model, characterized in that: The fused displacement optical flow data in step S3 The calculation method is: The semantic segmentation model is based on The optical flow model is improved from the Improved; the overall displacement The calculation process of the frame image displacement optical flow data can be expressed as: , , Local displacement The calculation process of the frame image displacement optical flow data can be expressed as: , , The expression of the dataset after fusion: , In the formula, Respectively The light flows at all times and The weight in direction, represents the overall displacement mask of the output, Represents the output local displacement mask, Indicates the operation of extracting masks based on the U-Net network model. represents the optical flow image, represents the fused optical flow data after mask denoising, Represents the overall displacement optical flow data after mask denoising, represents the overall local optical flow data after mask denoising, Represents the optical flow model function, whose input is the pose rendering image and output is the optical flow data of the corresponding structure.
7. According to claim 1, a method for visually recognizing structural displacements by integrating a three-dimensional deformable model, characterized in that: In step S4, a dynamic weight is established between the two features according to these vectors to obtain a fusion feature. The specific method is: Overall characteristics and local features After input into the network, linear transformation is performed to generate query, key and value vectors; For example, it generates the corresponding query, key, and value vectors through a set of learnable weight matrices: , Similarly, local features It is also possible to generate query, key, and value vectors: , in, Represent overall features and local features respectively; The query vector, key vector and value vector represent the overall features respectively; The query vector, key vector and value vector represent the local features respectively; are learnable weight matrices used to generate query vectors, key vectors, and value vectors; these matrices can automatically learn the most suitable parameters through the model's training data to generate vector representations that contribute to displacement estimation on different features; In the cross-attention mechanism, the similarity or correlation between the generated query vector and the key vector is calculated to generate the attention weight; the similarity is calculated using the vector dot product and scaled to maintain a stable gradient. The specific formula is: , In the formula, Indicates the dimension of the key vector, used for scaling to prevent the vector dot product value from being too large; denote the query vector, key vector and value vector respectively, represents the matrix transpose, represents the activation function, Represents the calculation of the cross-attention mechanism; The core of this formula is to calculate the similarity between the query vector and the key vector, divided by the scaling factor and pass it through Normalized to the attention weight; then, through the weighted sum vector, To extract features; in the cross attention mechanism, this calculation process is carried out bidirectionally between the overall features and the local features to capture detailed information; using the overall features The query vector Can be used to focus on key location information of local features: , Local feature query Contextual information used to extract overall features: , In the formula, The query vector, key vector and value vector represent the overall features respectively; They represent the query vector, key vector and value vector of local features respectively.
8. According to claim 1, a method for visually recognizing structural displacements by integrating a three-dimensional deformable model, characterized in that: The loss function expression in step S4 is: , In the formula, is the actual displacement value data, Predict displacement value data for the model, is the total number of samples, is the regularization coefficient; represents the regularization weight, Represents the ith actual displacement value data, Represents the displacement value data predicted by the i-th model.
9. A structural displacement visual recognition system integrating a three-dimensional deformable model, characterized in that: include: A three-dimensional deformable mesh model building unit is used to collect engineering structure vibration video data through a consumer-grade monocular camera, build a three-dimensional deformable mesh model, and generate pose rendering images and displacement value data; A data preprocessing unit, used to remove the interference of environmental noise on displacement recognition; An optical flow feature data set construction unit, used to estimate dense optical flow data in a rendered image through an optical flow model, and to construct an optical flow feature data set containing pose information; The model learning unit is used for the structure-intensive displacement estimation method based on the cross-attention mechanism. By fusing local and global features, the model displacement prediction value is output. The specific steps are as follows: Input two sets of data sequences: the overall structure image sequence and the local structure image sequence. After multi-layer convolution operations at the volume base layer, the corresponding overall features are generated respectively. With local features ; Overall characteristics and local features After linear transformation, the feature query vector, key vector and value vector are generated, and dynamic weights are established between the two features based on these vectors to obtain two fusion features. as well as ; Determine the loss function, perform global average pooling on the two fusion features obtained, and then output the pose estimation value through a multi-layer perceptron.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: The computer program is executed by a processor to implement the structural displacement visual recognition method of fused three-dimensional deformable model as described in any one of claims 1-8.
Citation Information
Patent Citations
Monocular vision and deep learning-based structure full-field displacement dense measurement method and device, equipment and storage medium
CN113989699A
Structural dense displacement identification method and system based on deformable three-dimensional model and optical flow representation learning
CN116433755A