A Structural Vibration Displacement Identification Method Based on Dense Matching and Prior Knowledge Enhancement

Through a method based on dense matching and prior knowledge enhancement, the trained image feature network and optical flow propagation strategy are used, combined with structural linear vibration theory, the problem of fast dense recognition in visual displacement measurement is solved, and efficient structural vibration displacement recognition is achieved in harsh environments.

CN119006522BActive Publication Date: 2025-07-25HARBIN INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411091247.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-09
Publication Date
2025-07-25
Estimated Expiration
2044-08-09

AI Technical Summary

Technical Problem

The existing visual displacement measurement methods cannot achieve fast and dense structural vibration displacement recognition at the same time, and are not robust under harsh environmental conditions, so they cannot effectively utilize the motion priors and physical knowledge in monitoring videos.

Method used

Using a method based on dense matching and prior knowledge enhancement, we train the image feature enhancement network, use dense matching and optical flow propagation strategies based on attention mechanism, combined with the physical knowledge of structural linear vibration theory, non-iteration optical flow estimation and displacement conversion are carried out, and the measurement results are improved using motion priors and physical prior constraints.

Benefits of technology

It realizes rapid and dense structural vibration displacement recognition under harsh environmental conditions, improves recognition accuracy and robustness, and is suitable for structural health monitoring systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119006522B_ABST
    Figure CN119006522B_ABST
Patent Text Reader

Abstract

The present invention proposes a method for identifying structural vibration displacement based on dense matching and prior knowledge enhancement. The method includes the architecture design and training of an image feature enhancement network, the establishment of a non-iterative optical flow estimation model based on dense matching, the improvement of the optical flow estimation result based on the motion prior information in the monitoring video, the conversion of pixel motion to structural displacement, and the improvement of the displacement recognition result based on physical prior knowledge, etc. The method of the present invention uses a trained deep learning model to enhance image features, and uses a dense matching and optical flow propagation strategy based on the attention mechanism to obtain the full-field pixel motion of the monitoring video. This method realizes fast and dense motion estimation and solves the problems of existing methods. The method has unique advantages in terms of the accuracy, density, and speed of identifying structural vibration displacement under low-quality monitoring videos, and also has strong robustness to harsh environmental conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of signal processing and structural health monitoring, and particularly relates to a method for identifying structural vibration displacement based on dense matching and prior knowledge enhancement. Background Art

[0002] The measurement of the structural vibration displacement time history is very important in civil engineering. It not only reveals the static and dynamic characteristics of civil structures, but also directly reflects the structural state. For example, indicators such as excessive deflection in bridge structures and excessive inter-story drift angles in building structures directly indicate potential damage to in-service structures. Current measurement techniques include: contact sensors connected to monitoring points, traditional non-contact sensors, and vision-based non-contact sensors.

[0003] Contact sensors measure the relative motion between two physically connected points to obtain displacement. Although such sensors have high displacement measurement accuracy, they require fixed reference points, which is not feasible for actual large-scale structures. In addition, the wired connection makes the deployment and maintenance costs of such systems high. Wireless transmission technology can reduce the complexity of wiring, but contact systems usually only deploy sparse "point-type" sensors, which means that the displacement measurement results have only low spatial resolution. Therefore, structural information may be lost, resulting in the subsequent inability to accurately identify local damage. Traditional non-contact sensors obtain displacement by measuring satellite positioning, laser, radar, and other remote sensing signals. Among them, satellite positioning has low measurement accuracy and frequency and is suitable for static displacement monitoring. Laser measurement has high accuracy and frequency, but the laser power under safety restrictions cannot achieve long-distance displacement measurement. In addition, both of the above non-contact systems need to install markers on the structure as measurement points, so they also have sparse measurement defects similar to contact systems. Radar can achieve dense measurement, but its cost and power consumption are very high and it is not suitable for the actual application of civil engineering.

[0004] Thanks to the progress of vision sensors, it is now possible to perform low-cost, accurate, and relatively high-speed long-distance video recording. For example, a combination of consumer cameras or smartphones, zoom lenses, and tripods, or even drones can be used for video recording. The obtained image sequence records the motion of all pixels in the field of view, which contains dense structural displacement information. Vision-based portable measurement systems are not only easy to operate but can also quickly resume work after major disasters to obtain structural information and ensure timely emergency response. In the past few decades, a series of vision tracking and transformation algorithms have been applied to civil engineering scenarios to achieve non-contact displacement measurement based on monitored video data. The advantages of low cost, dense measurement, and easy operation and maintenance make vision-based systems an effective supplement to traditional displacement measurement systems. In a stable imaging and monitoring environment, vision measurement systems are even a highly competitive alternative.

[0005] Vision-based displacement measurement includes two main steps: pixel motion estimation and transformation from motion to displacement. Existing motion estimation algorithms are divided into two categories: non-iterative and iterative. Among them, non-iterative algorithms are suitable for quickly tracking the motion of specific pixel points and require additional calculations to achieve quasi-dense measurement and sub-pixel accuracy. Iterative algorithms are limited to estimating the motion field with a small displacement amplitude and require multi-layer image pyramid stacking and multiple iterations to improve measurement accuracy. Currently, a single motion estimation algorithm cannot achieve both fast estimation and dense estimation simultaneously, which hinders the real-time identification of structural displacement and the maximum utilization of monitored video data. In addition, existing methods also require manual adjustment of the hyperparameters of the measurement algorithm to cope with harsh environmental conditions. The above problems are very unfavorable to the application and automation of vision-based displacement measurement technology.

[0006] Currently, there is no research and application of vision displacement recognition methods that simultaneously consider the speed and density of motion estimation in the field of structural health monitoring. The potential of using the motion prior and physical knowledge implicit in the monitored data to achieve accurate and robust identification of structural vibration displacement under low-quality monitored video has not been explored. Summary of the Invention

[0007] The object of the present invention is to solve the problems in the prior art to meet the actual detection needs, and a structural vibration displacement recognition method based on dense matching and prior knowledge enhancement is proposed.

[0008] The present invention is realized through the following technical solutions. The present invention proposes a structural vibration displacement recognition method based on dense matching and prior knowledge enhancement, and the method includes the following steps:

[0009] Step 1: Establish a large dataset of videos containing diverse monitoring scenarios and inter-frame optical flow through simulation rendering technology for training the image feature enhancement network

[0010] Step 2. Establish a non-iterative optical flow estimation model based on dense matching: Select the initial network architecture hyperparameters and establish an image feature enhancement network. Input two related images (I1, I2) in the dataset, and calculate to output enhanced features (F1, F2) of the same resolution; Denote the positions of all pixel points in image I1 as P1, and use the dense matching strategy based on the attention mechanism for non-iterative feature matching to obtain the corresponding positions P2 of these pixel points in image I2, and then obtain the initial inter-frame optical flow estimation result O = P2 - P1; Use the optical flow propagation strategy based on the attention mechanism to propagate the initial optical flow to the occluded and out-of-boundary regions that cannot be matched in the image, and obtain the optical flow estimation value O. * Compare the optical flow estimation value with the ground truth in the dataset, and train the network through error backpropagation and gradient descent methods. After that, perform hyperparameter tuning to obtain an image feature enhancement network that has learned the matching description ability. Thereafter, fix its parameters θ and do not change them anymore.

[0011] Step 3. Extract the motion prior information implicit in the actual structural vibration monitoring video and improve the optical flow estimation result.

[0012] Step 4. Set a reference frame where the structure is stationary in the monitoring video, take the stationary frame and the current frame together as inputs, and calculate the full-field pixel motion between the two frames based on the optical flow estimation model in Step 2 and the two types of motion prior constraints in Step 3.

[0013] Step 5. Based on the physical knowledge of the structural linear vibration theory, improve the displacement measurement result; Compose the vibration displacement time history vectors of all measurement nodes into a displacement matrix; Use singular value decomposition to approximately perform modal decomposition of the structural vibration response, and approximate the low-rank property by only retaining the main singular values to perform noise reduction on the displacement measurement result and reduce the excessive error level caused by the poor recording quality of some structural node regions in the monitoring video; Recombine using the main singular value components to obtain a more accurate structural vibration displacement identification result.

[0014] Furthermore, the inter-frame optical flow is the full-field pixel motion that occurs between two video images.

[0015] Further, in step three, on the one hand, a non-iterative motion estimation algorithm is used to quickly obtain an estimate of the upper bound of the motion amplitude of the structural nodes to be monitored, which is used as a motion prior to limit the range of dense matching; on the other hand, the forward optical flow and the backward optical flow between video frames are compared to estimate the occluded and out-of-boundary regions in the monitored video, and the overlapping degree between these regions and the structural node regions is calculated, which is used as a motion prior to determine whether optical flow propagation is required and to limit the range of optical flow propagation.

[0016] Further, in step four, the structural node regions in the video that need to identify vibration displacements are boxed, the homography matrix for converting the captured video view to the front view of the structure is obtained based on feature point matching and perspective transformation, and the scale factor for unit conversion is calculated in the front view; the homography matrix and the scale factor are used to convert pixel motion to actual displacement, and the displacements of all pixel points in the node region are averaged to obtain the node vibration displacement; the above process is repeated to obtain the time history of the node vibration displacement.

[0017] Further, step two is specifically as follows:

[0018] Step 2.1: Establish an initial image feature enhancement network to process two frames of images (I1, I2):

[0019]

[0020] Among them, the enhanced image features have the same spatial resolution as the original images;

[0021] Step 2.2: Denote the positions of all pixel points in image I1 as P1, and denote the corresponding positions of the paired pixel points in image I2 as P2. Based on the pixel-level feature descriptors, a dense matching strategy based on the attention mechanism is used to achieve optical flow estimation:

[0022]

[0023] O = P2 - P1

[0024] The optical flow estimation result based on dense matching is O; among them, the two spatial coordinate axes of the image feature matrix are combined into one, and the softmax function is used to normalize the feature matching transfer weights; and the original positions of the pixel points are recombined and converted into new positions, realizing sub-pixel accuracy optical flow estimation;

[0025] Step 2.3: Utilize the self-similarity of the input image feature F1 to propagate the optical flow estimation result of the successfully matched region to the unmatched region; implement the optical flow propagation strategy based on the attention mechanism:

[0026]

[0027] Finally, the optical flow estimation result O considering the propagation strategy is obtained. * ;

[0028] Step 2.4: Compare the optical flow estimation result with the ground truth in the dataset, establish an objective function, train through error backpropagation and gradient descent methods, and perform hyperparameter tuning to obtain an image feature enhancement network that has learned the matching description ability. After that, its parameters θ are fixed.

[0029] Furthermore, in step three, a non-iterative motion estimation algorithm is used to quickly obtain an estimate of the upper bound of the motion amplitude of the structural nodes to be monitored as a motion prior, restricting the range of dense matching for each pixel point:

[0030]

[0031] Among them, the mask function changes with the pixel position, setting the feature transfer weight outside the matching area constrained by the motion prior to zero, achieving pixel-level shielding of information that is not conducive to accurate matching.

[0032] Furthermore, in step three, the input order of the two frames of images is swapped to obtain the reverse optical flow. The forward optical flow and the backward optical flow between video frames are compared to estimate the occluded and out-of-boundary regions existing in the monitored video, and the overlap degree between these regions and the structural node regions is calculated as a motion prior, restricting the range of optical flow propagation for each pixel point:

[0033]

[0034] Among them, the mask function changes with the pixel position, setting the feature transfer weight outside the propagation area constrained by the motion prior to zero, achieving pixel-level shielding of information that is not conducive to accurate propagation.

[0035] Furthermore, the specific content of step five is as follows:

[0036] Step 5.1: Denote the time history vectors of the vibration displacements of each node identified in the video as d i (i = 1,..., M), and form a structural displacement matrix D = [d1... d i ... d M ; According to the modal superposition principle in the theory of structural linear vibration, the modal decomposition of the structural vibration response can be approximated using singular value decomposition:

[0037] D = U·diag(s)·V T

[0038] Among them, U and V are singular vector matrices, and s approximately represents the amplitudes of each modal component in the structural displacement response;

[0039] Step 5.2: According to physical knowledge, that is, the structural displacement matrix D has the property of low rank; calculate the ratio of adjacent singular values in s, set a threshold, and set the too small singular values to zero to obtain a new singular value vector s * , and recombine to obtain:

[0040] D * = U·diag(s * )·V T

[0041] where D * is the identified result of the structural vibration displacement after denoising.

[0042] The present invention also provides an electronic device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the method for identifying structural vibration displacement based on dense matching and prior knowledge enhancement are implemented.

[0043] The present invention also provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the steps of the method for identifying structural vibration displacement based on dense matching and prior knowledge enhancement are implemented.

[0044] Advantages of the present invention:

[0045] 1. Compared with the existing methods, the non-iterative method based on dense matching of the present invention obtains the estimation result of the full-field pixel motion through a single-step calculation by a matching strategy based on the attention mechanism, solves the problem that the existing methods cannot simultaneously achieve fast displacement identification and dense displacement identification, and makes the method more suitable for application in the structural health monitoring system.

[0046] 2. The optical flow propagation strategy of the method of the present invention can utilize the surrounding information of the image to recover the pixel motion in the occluded and out-of-boundary regions in the monitoring video, so that the method can still calculate a relatively accurate displacement identification result even when the monitoring area is partially occluded; in addition, the method of the present invention enhances the image features using a deep learning model and uses the enhanced image features for pixel motion estimation, so that the method has better robustness under harsh environmental conditions and is more suitable for practical applications.

[0047] 3. The method of the present invention utilizes the prior knowledge of the motion of the monitoring area implicitly contained in the video data, and restricts the range of dense matching and optical flow propagation based on the prior knowledge, further improving the accuracy of displacement identification.

[0048] 4. The present invention utilizes the physical prior knowledge that the structural vibration response is usually dominated by a small number of modes. Based on the approximate relationship between the modal decomposition and singular value decomposition of the structural vibration response in linear vibration theory, the displacement identification result is denoised by only retaining the main singular value components, improving the accuracy of structural vibration displacement identification and enabling the method to resist the relatively high measurement noise in some structural node regions caused by poor monitoring video quality. Description of the Drawings

[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the provided drawings.

[0050] Figure 1 It is a flowchart of the method for identifying structural vibration displacement based on dense matching and prior knowledge enhancement of the present invention.

[0051] Figure 2 It is a schematic diagram of the optical flow estimation method.

[0052] Figure 3 It is a schematic diagram of the monitoring video used in the example of the present invention, and there are 16 structural nodes in this video.

[0053] Figure 4 It is a schematic diagram of the estimated results of the full-field pixel motion (i.e., optical flow) of the method of the present invention under the condition that monitoring node 5 is partially blocked and under random illumination conditions, respectively.

[0054] Figure 5 It is a schematic diagram of the comparison result of the error between the structural vibration displacement identification result of the method of the present invention and the true value before and after using the motion prior knowledge to constrain the optical flow estimation process. The error shown is the average error of 16 monitoring nodes.

[0055] Figure 6 It is a schematic diagram of the comparison result of the structural vibration displacement identification result of the method of the present invention and the true value before and after using the physical prior knowledge to denoise the displacement identification result. The corresponding result of monitoring node 9 is shown. Detailed Embodiments

[0056] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0057] To solve the difficult problem of fast, dense, and robust structural displacement recognition based on vision, the present invention implements a structural vibration displacement recognition method based on dense matching and prior knowledge enhancement. First, this method trains a deep learning model to enhance image features and uses a dense matching method based on the attention mechanism to obtain the full-field pixel motion of the monitored video. This method is non-iterative, but achieves fast and dense motion estimation, solving the problems of existing algorithms. In addition, the enhanced image features are more robust to harsh environmental conditions. Further, combined with a method that converts the pixel motion in all structural node regions into structural displacement at one time, the present invention quickly and robustly measures the dense structural vibration displacement. The present invention also proposes a strategy based on motion prior and physical knowledge enhancement to improve the measurement results. Finally, the method of the present invention realizes accurate structural vibration displacement recognition. Except for selecting the structural node regions for displacement conversion, the method of the present invention does not require additional manual intervention, which is conducive to promoting the realization of automated structural vibration displacement recognition.

[0058] The object of the present invention is to solve the difficult problem that existing vision-based structural vibration displacement recognition methods cannot simultaneously achieve fast displacement recognition and dense displacement recognition, and to propose a non-iterative method based on dense matching, which can estimate the full-field pixel motion through a single-step non-iterative matching, and then realizes fast and dense displacement recognition. In addition, when the monitoring environmental conditions are relatively harsh (for example, the light is unstable), or the target monitoring area in the structural vibration video data is partially occluded, the method of the present invention uses the image features enhanced by the deep learning model for recognition and uses the optical flow propagation strategy to estimate the pixel motion in the occluded area, which has a wider application range in practical engineering. The present invention also proposes to use prior knowledge to improve the displacement recognition results. For example Figure 2 The process principles of optical flow estimation, structural vibration displacement recognition, and prior knowledge constraint shown. For the video data in the actual structural health monitoring system, it is required to analyze the structural state information in real time and comprehensively. Therefore, a method for quickly recognizing dense structural vibration displacement is needed. At the same time, due to the high variability of the outdoor environment, the displacement recognition method needs to have high robustness, so as to be applicable to the method of the present invention.

[0059] Combined Figure 1 With the above, the present invention proposes a structural vibration displacement recognition method based on dense matching and prior knowledge enhancement, which specifically includes the following steps:

[0060] Step 1: Establish a large dataset containing videos with diverse monitoring scenarios and the ground truth of inter-frame optical flow (i.e., the full-field pixel motion occurring between two video images) through simulation rendering technology for training the image feature enhancement network

[0061] Step 2: Establish a non-iterative optical flow estimation model based on dense matching. Select the initial network architecture hyperparameters and establish an image feature enhancement network. Input two related images (I1, I2) in the dataset, and after calculation, output enhanced features (F1, F2) with the same resolution. Denote the positions of all pixel points in image I1 as P1, and use the dense matching strategy based on the attention mechanism for non-iterative feature matching to obtain the corresponding positions P2 of these pixel points in image I2, and then obtain the initial inter-frame optical flow estimation result O = P2 - P1. Use the optical flow propagation strategy based on the attention mechanism to propagate the initial optical flow to the occluded and out-of-boundary regions in the image where matching cannot be performed, and obtain the optical flow estimation value O. * Compare the optical flow estimation value with the ground truth in the dataset, and train the network through error backpropagation and gradient descent methods. Perform hyperparameter tuning to obtain an image feature enhancement network that has learned the matching description ability. After that, fix its parameters θ and do not change them anymore.

[0062] Step 3: Extract the motion prior information implicit in the actual structural vibration monitoring video and improve the optical flow estimation result. First, use the existing non-iterative motion estimation algorithm to quickly obtain an estimate of the upper bound of the motion amplitude of the structural nodes to be monitored as the motion prior to limit the dense matching range. Second, compare the forward optical flow and backward optical flow between adjacent video image frames, estimate the occluded and out-of-boundary regions in the monitoring video, calculate the overlap degree between these regions and the structural node regions as the motion prior, and judge whether optical flow propagation is required to limit the optical flow propagation range.

[0063] Step 4: Set the reference frame where the structure is stationary in the monitoring video, use the stationary frame and the current frame together as the input, and based on the optical flow estimation model in Step 2 and the two types of motion prior constraints in Step 3, calculate the full-field pixel motion between the two frames. Select the structural node region in the video where the vibration displacement needs to be identified, and based on existing technologies such as feature point matching and perspective transformation, obtain the homography matrix for converting the video capture view to the front view of the structure, and calculate the scale factor for unit conversion in the front view. Use the homography matrix and the scale factor to realize the conversion of pixel motion to actual displacement, and average the displacements of all pixel points in the node region to obtain the node vibration displacement. Repeat the above process to obtain the time history of the node vibration displacement.

[0064] Step 5. Improve the displacement measurement results based on the physical knowledge of structural linear vibration theory. The vibration displacement time history vectors of all measurement nodes are composed into a displacement matrix. According to physical knowledge, this matrix generally has a low-rank property. Singular value decomposition is used to approximately decompose the modal of the structural vibration response. By only retaining the main singular values to approximate the low-rank property, noise reduction of the displacement measurement results is performed, reducing the excessive error level caused by the poor recording quality of some structural node regions in the monitoring video. The main singular value components are recombined to obtain a more accurate structural vibration displacement identification result.

[0065] For the optical flow estimation process, Step 2 is specifically as follows:

[0066] Step 2.1. A pixel-level feature descriptor with high distinctiveness is the key to the success of the optical flow estimation method based on dense matching in the present invention. An initial image feature enhancement network is established to process two frames of images (I1, I2):

[0067]

[0068] Among them, the enhanced image features have the same spatial resolution as the original image;

[0069] Step 2.2. The optical flow estimation problem is essentially to find the positions of all paired pixel points between two frames of images. Denote the positions of all pixel points in image I1 as P1, and the corresponding positions of the paired pixel points in image I2 as P2. Based on a reliable pixel-level feature descriptor, an optical flow estimation is realized using a dense matching strategy based on an attention mechanism:

[0070]

[0071] O = p2 - p1

[0072] The optical flow estimation result based on dense matching is O. Among them, for the convenience of matrix multiplication, the two spatial coordinate axes of the image feature matrix are combined into one. The softmax function is used to normalize the feature matching transfer weights. Considering the correlation of image features, the original positions of the pixel points are recombined and converted into new positions, realizing optical flow estimation with sub-pixel accuracy. This dense matching process only needs to be performed once. Therefore, the method of the present invention is non-iterative and realizes fast optical flow estimation;

[0073] Step 2.3. Considering that there may be occluded and out-of-bound pixel regions in the actual monitoring video, which will lead to matching failures. The self-similarity of the input image feature F1 is utilized to propagate the optical flow estimation results of the successfully matched regions to the unmatched regions. Similarly, the above optical flow propagation strategy is realized based on an attention mechanism:

[0074]

[0075] Finally, the estimated result O considering the optical flow propagation strategy is obtained. * ;

[0076] Step 2.4: The above feature enhancement, optical flow estimation, and optical flow propagation processes are all differentiable. Compare the optical flow estimation result with the ground truth in the dataset, establish an objective function, train through error backpropagation and gradient descent methods, and perform hyperparameter tuning to obtain an image feature enhancement network that has learned the matching description ability. After that, fix its parameters θ.

[0077] The matching strategy in Step 2.2 and the propagation strategy in Step 2.3 both consider the global pixel features, and the global features may contain local information that is not conducive to the motion estimation of each pixel point. The motion prior knowledge can be used to limit the range of the attention mechanism and improve the optical flow estimation result. The specific steps of Step 3 are as follows:

[0078] Step 3.1: Use an existing non-iterative motion estimation algorithm to quickly obtain an estimate of the upper bound of the motion amplitude of the structural nodes to be monitored as a motion prior, and limit the range of dense matching for each pixel point:

[0079]

[0080] Among them, the mask function changes with the pixel position, sets the feature transfer weight outside the matching area to zero, and realizes pixel-level shielding of information that is not conducive to accurate matching.

[0081] Step 3.2: Exchange the input order of the two frames of images to obtain the reverse optical flow. Compare the forward optical flow and the backward optical flow between video frames, estimate the occluded and out-of-boundary areas in the monitored video, and calculate the overlap degree between these areas and the structural node area as a motion prior to limit the range of optical flow propagation for each pixel point:

[0082]

[0083] Among them, the mask function changes with the pixel position, sets the feature transfer weight outside the propagation area to zero, and realizes pixel-level shielding of information that is not conducive to accurate propagation.

[0084] Considering that the shooting quality of different regions in the same monitored video may not be uniform, resulting in a relatively high error level, use physical prior knowledge to improve the displacement measurement result obtained in Step 4. The specific steps of Step 5 are as follows:

[0085] Step 5.1: Denote the vibration displacement time history vectors of each node identified in the video as d i (i = 1, …, M), and form a structural displacement matrix D = [d1 … d i…d M . According to the modal superposition principle in the structural linear vibration theory, the modal decomposition of the structural vibration response can be approximated using singular value decomposition:

[0086] D = U·diag(s)·V T

[0087] where U and V are singular vector matrices, and s approximately represents the amplitudes of each modal component in the structural displacement response;

[0088] Step 5.2. According to physical knowledge, the structural vibration response is usually dominated by a small number of modes, that is, the structural displacement matrix D has the property of low rank. This constraint of physical prior knowledge can be imposed by only retaining the main singular value components to remove the noise in the vision-based displacement recognition results. Calculate the ratio of adjacent singular values in s, set a threshold, set the too-small singular values to zero, and obtain a new singular value vector s * , and recombine to obtain:

[0089] D * = U·diag(s * )·V T

[0090] where D * is the denoised structural vibration displacement recognition result.

[0091] The present invention proposes a method for identifying structural vibration displacement based on dense matching and prior knowledge enhancement. The method includes the architecture design and training of an image feature enhancement network, the establishment of a non-iterative optical flow estimation model based on dense matching, the improvement of the optical flow estimation result based on the motion prior information in the monitoring video, the conversion of pixel motion to structural displacement, and the improvement of the displacement recognition result based on physical prior knowledge. The method of the present invention uses a trained deep learning model to enhance image features and uses a dense matching and optical flow propagation strategy based on the attention mechanism to obtain the full-field pixel motion of the monitoring video. This method realizes fast and dense motion estimation and solves the problems of existing methods. In addition, the enhanced image features are more robust to harsh environmental conditions and have a wider application range. Combined with a method that converts the pixel motion in all structural node regions into structural displacement at one time, the present invention quickly and robustly identifies the dense structural vibration displacement in the monitoring video. The present invention also proposes a strategy based on motion prior and physical knowledge enhancement to improve the recognition result and realizes accurate identification of structural vibration displacement. The method has unique advantages in terms of the accuracy, density, and speed of identifying structural vibration displacement under low-quality monitoring videos and is also relatively robust to harsh environmental conditions.

[0092] Embodiment

[0093] Figure 3Schematic diagram of the monitoring video used in this example. There are 16 structural nodes in the video, and the structure mainly undergoes vertical vibration. The resolution of the video is 1920 pixels × 1080 pixels, the shooting duration and frequency are 10 seconds and 120 Hz respectively, and the video shooting angle is perpendicular to the structural movement plane. Combining Figure 4 , Figure 5 , Figure 6 , the method based on dense matching and prior knowledge enhancement of the present invention is used for structural vibration displacement identification.

[0094] Specifically, step one is: establishing a large dataset of videos containing diverse scenarios (different environmental light, resolution, etc.) and optical flow ground truth between frames through simulation rendering software such as Blender for training the image feature enhancement network

[0095] Specifically, step two is: in order to capture as much information as possible in the input image data that is helpful for dense matching, an image feature enhancement network is established based on the Transformer architecture It consists of three parts: a self-attention layer applied to the first input, a cross-attention layer applied to the first input and the second input, and a fully connected layer. Two related images (I1, I2) in the input dataset are input, and after calculation, enhanced features of the same resolution (F1, F2) are output. Based on the image features (F1, F2), dense matching and optical flow propagation strategies are used to estimate the optical flow between image frames. Using the first norm error between the optical flow estimated value and the ground truth as the objective function, the Adam optimization algorithm is used to initialize various training parameters (batch size, number of training epochs, learning rate, etc.) to train the network And perform hyperparameter tuning for training to obtain an image feature enhancement network that effectively learns the matching description ability Fix its network parameters and no longer change.

[0096] After obtaining the optical flow estimation model, in Figure 4 , partial occlusion is artificially applied to the video area where monitoring node 5 is located, and the brightness of each frame of the monitoring video is randomly adjusted, and the method of the present invention is used for full-field pixel motion estimation. The estimated full-field pixel motion is visualized as an optical flow image, where the color represents the motion direction and the saturation represents the motion amplitude. Figure 4 The good optical flow estimation results shown indicate that the method of the present invention has extremely strong robustness to harsh environmental conditions.

[0097] Step 3 specifically includes: on the one hand, the KLT feature tracking algorithm is used to quickly obtain that the upper bound of the motion amplitude of the target structure nodes is 6 pixels, which is used as a motion prior to limit the matching range radius of each pixel point to 6 pixels; on the other hand, by comparing the forward optical flow and the backward optical flow between adjacent video image frames, it is observed that there are no occlusion and out-of-boundary problems in the monitored video used, which is used as a motion prior to remove the optical flow propagation step and improve the optical flow estimation result.

[0098] Step 4 specifically includes: the structure in the monitored video is in a static state in the first frame. The first frame and the current frame are used as inputs together, and the full-field pixel motion between the two frames is calculated based on Step 2 and Step 3. Since the monitored video used is taken from a front view, the scale factor for converting the pixel length to displacement can be directly calculated based on the known structure size. The structural node area where the vibration displacement needs to be identified in the video is boxed, and the motion estimation results of all pixel points within the node area are multiplied by the scale factor and averaged to obtain the node vibration displacement. Repeat the above process to identify the vibration displacement time history of 16 nodes;

[0099] In Figure 5 after applying the two types of motion prior constraints in Step 3, the error of the method of the present invention for identifying the structural vibration displacement decreases significantly, verifying the effectiveness of the strategy proposed by the present invention.

[0100] Step 5 specifically includes: the video used consists of 1200 frames of images. The displacement time history vectors identified for 16 structural nodes respectively form a 16-dimensional × 1200-dimensional structural displacement matrix. Singular value decomposition is performed, and it is observed that the main singular value components of this matrix are the first three orders. The first three orders of singular values are retained for noise reduction to obtain the final structural vibration displacement identification result.

[0101] In Figure 6 after applying the physical prior constraint in Step 5, the identification result of the vibration displacement time history of monitoring node 9 is significantly closer to the true value both in the time domain and the frequency domain, verifying the effectiveness of the strategy proposed by the present invention.

[0102] Combined with the above results, it can be seen that the method of the present invention can quickly and accurately identify the dense structural vibration displacement from the monitored video, and the identification result has good robustness to problems such as harsh environmental conditions, occlusion, and video recording noise.

[0103] The present invention also proposes an electronic device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the method for identifying structural vibration displacement based on dense matching and prior knowledge enhancement are implemented.

[0104] The present invention also provides a computer-readable storage medium for storing computer instructions, which, when executed by a processor, implement the steps of the method for identifying structural vibration displacement based on dense matching and prior knowledge enhancement.

[0105] The memory in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read only memory (ROM), a programmable ROM (PROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DRRAM). It should be noted that the memory for the method described in the present invention is intended to include but not limited to these and any other suitable types of memory.

[0106] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a high-density digital video disc (DVD)), or a semiconductor medium (such as a solid state disc (SSD)), etc.

[0107] In the implementation process, the steps of the above method can be completed by the integrated logic circuit of the hardware in the processor or the instructions in the form of software. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed by the hardware processor or executed by the combination of the hardware and software modules in the processor. The software module can be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, registers, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.

[0108] It should be noted that the processor in the embodiments of the present application can be an integrated circuit chip with signal processing capabilities. In the implementation process, the steps of the above method embodiments can be completed by the integrated logic circuit in the hardware of the processor or instructions in the form of software. The above-mentioned processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.

[0109] The above has introduced in detail a method for identifying structural vibration displacement based on dense matching and prior knowledge enhancement proposed by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A method for identifying structural vibration displacement based on dense matching and prior knowledge enhancement, characterized in that The method includes the following steps: Step 1: Establish a large dataset of videos containing diverse monitoring scenarios and inter-frame optical flow through simulation rendering technology for training the image feature enhancement network Step 2. Establish a non-iterative optical flow estimation model based on dense matching: Select initial network architecture hyperparameters and establish an image feature enhancement network Input two related images I1 and I2 in the dataset, and calculate to output enhanced features F1 and F2 with the same resolution; Denote the positions of all pixel points in image I1 as P1, and use the dense matching strategy based on the attention mechanism for non-iterative feature matching to obtain the corresponding positions P2 of these pixel points in image I2, and then obtain the initial inter-frame optical flow estimation result 0 = P2 - P1; Use the optical flow propagation strategy based on the attention mechanism to propagate the initial optical flow to the occluded and out-of-boundary regions in the image where matching cannot be performed, and obtain the optical flow estimation value O * ; Compare the optical flow estimation value with the ground truth in the dataset, and train the network through error backpropagation and gradient descent methods After that, perform hyperparameter tuning to obtain an image feature enhancement network that has learned the matching description ability After that, fix its parameters θ and do not change them anymore; Step 3: Extract the motion prior information implicit in the actual structural vibration monitoring video to improve the optical flow estimation result; In Step 3, on the one hand, use a non-iterative motion estimation algorithm to quickly obtain an estimate of the upper bound of the motion amplitude of the structural nodes to be monitored as a motion prior to limit the range of dense matching; on the other hand, compare the forward optical flow and the backward optical flow between video frames, estimate the occluded and out-of-boundary regions in the monitoring video, calculate the overlap degree between these regions and the structural node regions as a motion prior, and determine whether optical flow propagation is required to limit the range of optical flow propagation; Step 4: Set a reference frame where the structure is stationary in the monitoring video, use the stationary frame and the current frame as inputs, and calculate the full-field pixel motion between the two frames based on the optical flow estimation model in Step 2 and the two types of motion prior constraints in Step 3; Step 5: Improve the displacement measurement result based on the physical knowledge of the structural linear vibration theory; form a displacement matrix with the vibration displacement time history vectors of all measurement nodes; use singular value decomposition to approximately perform modal decomposition of the structural vibration response, approximate the low-rank property by only retaining the main singular values, perform noise reduction on the displacement measurement result, and reduce the excessive error level caused by the poor recording quality of some structural node regions in the monitoring video; recombine with the main singular value components to obtain a more accurate structural vibration displacement identification result.

2. The method according to claim 1, characterized in that, The inter-frame optical flow is the full-field pixel motion that occurs between two video images.

3. The method according to claim 1, characterized in that, In Step 4, select the structural node regions in the video for which the vibration displacement needs to be identified, obtain the homography matrix for converting the captured video view to the front view of the structure based on feature point matching and perspective transformation, and calculate the scale factor for unit conversion in the front view; Use the homography matrix and the scale factor to achieve the conversion of pixel motion to actual displacement, and average the displacements of all pixel points in the node region to obtain the node vibration displacement; Repeat the process of Step 4 to obtain the node vibration displacement time history.

4. The method according to claim 1, wherein The specific content of Step 2 is as follows: Step 2.1: Establish an initial image feature enhancement network to process two frames of images I1 and I2: Among them, the enhanced image features have the same spatial resolution as the original image; Step 2.2: Denote the positions of all pixel points in image I1 as P1, and denote the corresponding positions of the paired pixel points in image I2 as P2. Based on the pixel-level feature descriptors, use a dense matching strategy based on the attention mechanism to achieve optical flow estimation: O = P2 - P1 The optical flow estimation result based on dense matching is O; among them, the two spatial coordinate axes of the image feature matrix are combined into one, and the softmax function is used to normalize the feature matching transfer weights; and the original positions of the pixel points are recombined and converted into new positions to achieve sub-pixel accuracy optical flow estimation; Step 2.3: Use the self-similarity of the input image feature F1 to propagate the optical flow estimation result of the successfully matched region to the unmatched region; implement the optical flow propagation strategy based on the attention mechanism: Finally, the optical flow estimation result O considering the propagation strategy is obtained * ; Step 2.4: Compare the optical flow estimation results with the ground truth in the dataset, establish an objective function, train it through error backpropagation and gradient descent methods, perform hyperparameter tuning, and obtain an image feature enhancement network that has learned the matching description ability. After that, fix its parameters θ.

5. The method according to claim 1, characterized in that, In Step 3, use a non-iterative motion estimation algorithm to quickly obtain an estimate of the upper bound of the motion amplitude of the structural nodes to be monitored as a motion prior to limit the range of dense matching for each pixel point: Among them, the mask function changes with the pixel position, sets the feature transfer weight outside the matching area constrained by the motion prior to zero, and realizes pixel-level masking of information that is not conducive to accurate matching.

6. The method according to claim 1, characterized in that In step three, the input order of the two frames of images is exchanged to obtain the reverse optical flow. The forward optical flow and the backward optical flow between video frames are compared to estimate the occluded and out-of-boundary regions existing in the monitoring video, and the overlap degree between these regions and the structural node region is calculated as the motion prior to limit the range of optical flow propagation for each pixel point: Among them, the mask function changes with the pixel position, sets the feature transfer weight outside the propagation area constrained by the motion prior to zero, and realizes pixel-level masking of information that is not conducive to accurate propagation.

7. The method according to claim 1, characterized in that The specific content of step five is as follows: Step 5.1: Denote the time history vectors of the vibration displacements of each node identified in the video as d i (i = 1, …, M), and form the structural displacement matrix D = [d1 … d i … d M ; According to the modal superposition principle in the theory of structural linear vibration, the modal decomposition of the structural vibration response is approximated using singular value decomposition: D = U·diag(s)·V T Among them, U and V are singular vector matrices, and s approximately represents the amplitude of each order modal component in the structural displacement response; Step 5.2: According to physical knowledge, that is, the structural displacement matrix D has the property of low rank; calculate the ratio of adjacent singular values in s, set a threshold, set the too-small singular values to zero, and obtain a new singular value vector s * , and recombine to obtain: D * = U·diag(s * )·V T Among them, D * is the identification result of the structural vibration displacement after denoising.

8. An electronic device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1-7.

9. A computer-readable storage medium for storing computer instructions, characterized in that, When the computer instructions are executed by the processor, it implements the steps of the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Structural modal parameter identification method and device, computer equipment and storage medium

    CN113901920A

  • Structural vibration displacement identification method and system based on deep recurrent neural network optical flow estimation model

    CN114485417A