Structural Displacement Measurement Method and System Based on Video Frame Interpolation and Computer Vision

By introducing attention mechanism and edge detection technology with local area effect in the video interpolation algorithm, the interframe relationship and weight allocation are improved, the accuracy and accuracy of structural displacement measurement are improved, the shortcomings in the existing technology are solved, and efficient contactless structural displacement measurement is achieved.

CN117308791BActive Publication Date: 2025-07-22FUJIAN AGRI & FORESTRY UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311244174.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-25
Publication Date
2025-07-22
Estimated Expiration
2043-09-25

AI Technical Summary

Technical Problem

The existing video interpolation algorithm has shortcomings when considering interframe relationships and weight allocation, which affects the quality of intermediate frames and structural displacement measurement accuracy. Traditional high-frame rate cameras are expensive and inconvenient to use, while ordinary cameras have limited frame rates, which limits the application of contactless structural displacement measurement.

Method used

The attention mechanism is introduced to improve the video interpolation algorithm, and the attention mechanism is designed on the optical flow network by enhancing the secondary video interpolation and TSA modules, assigning aggregate weights to each group of optical flows, and combining edge detection and edge tracking technologies based on local area effects to improve edge detection accuracy and displacement monitoring accuracy.

Benefits of technology

The quality and structural displacement measurement accuracy of the video interpolation algorithm generates intermediate frames, can effectively monitor the vibration response of the laboratory model and identify modal parameters, showing its potential in contactless structural vibration testing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117308791B_ABST
    Figure CN117308791B_ABST
Patent Text Reader

Abstract

The present application provides a structural displacement measurement method and system based on video interpolation and computer vision. The structural displacement measurement method acquires structural vibration video data through video interpolation technology based on computer vision, and introduces an attention mechanism into the video interpolation algorithm to obtain an improved video interpolation algorithm. By shooting the vibration video of a cantilever beam model as a data set, the performance of the improved video interpolation algorithm in structural vibration testing is evaluated. Finally, based on the interpolated structural vibration video, the time history data of the measured point displacement is obtained by detecting in the measurement area through edge detection technology and edge tracking technology based on the local area effect, the vibration response of the laboratory model is monitored, and modal parameter identification is carried out. The structural displacement measurement method processes the inter-frame relationship and weight allocation more comprehensively and accurately through the improved video interpolation algorithm, improving the quality of the generated intermediate frames and the monitoring accuracy of structural vibration displacement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of structural displacement measurement, and particularly to a structural displacement measurement method and system based on video frame interpolation and computer vision. Background Art

[0002] Structural vibration testing and modal parameter identification are basic methods for structural health monitoring based on the dynamic characteristics of structures. Traditional contact testing methods have many inconveniences, including cumbersome installation, adding additional mass to the structure under test, and high costs. In contrast, non-contact testing methods can effectively overcome these problems. Among them, the vibration testing method based on computer vision has attracted much attention. In traditional vibration testing, in order to maintain the accuracy of the signal, a sufficiently high sampling frequency is required to record the vibration signal. Similarly, non-contact methods based on computer vision also require high-frame-rate videos that meet the needs of signal analysis. However, high-frame-rate cameras are expensive and inconvenient to use, while the frame rate of ordinary consumer cameras is limited, which restricts the application of such technologies in engineering practice. Video frame interpolation technology can insert synthetic frames between the frames of the original video to increase the video frame rate, thereby solving the problem of the low frame rate of the original video.

[0003] Currently, the video frame interpolation algorithms used for structural displacement measurement have deficiencies in considering the inter-frame relationship and weight allocation, which affect the quality of the generated intermediate frames and the accuracy of structural displacement measurement, and there is room for improvement. Summary of the Invention

[0004] The purpose of the embodiments of this application is to provide a structural displacement measurement method and system based on video frame interpolation and computer vision to improve the problem that the video frame interpolation algorithms used for structural displacement measurement in related technologies have deficiencies in considering the inter-frame relationship and weight allocation, which affect the quality of the generated intermediate frames and the accuracy of structural displacement measurement.

[0005] In the first aspect of the present invention, a structural displacement measurement method based on video frame interpolation and computer vision is provided. The method includes:

[0006] Collecting structural vibration video data through non-contact measurement video frame interpolation technology based on computer vision;

[0007] Introducing an attention mechanism into the video frame interpolation algorithm to obtain an improved video frame interpolation algorithm;

[0008] Shooting the vibration video of the cantilever beam model as a data set to evaluate the performance of the improved video frame interpolation algorithm in structural vibration testing;

[0009] Based on the interpolated structural vibration video, edge detection technology and edge tracking technology based on the local area effect are used to detect in the measurement area to obtain the time history data of the measuring point displacement;

[0010] Monitor the vibration response of the laboratory model and identify the modal parameters.

[0011] Preferably, an attention mechanism is introduced into the video interpolation algorithm to obtain an improved video interpolation algorithm, including:

[0012] Design an attention mechanism on the optical flow network based on enhanced quadratic video interpolation and the TSA module, assign aggregation weights to each group of optical flows, and obtain an improved attention module;

[0013] By defining the parameters of the improved attention module and adjusting the input size to adapt to the dataset under the engineering vibration test task, that is, performing tensor stacking during input and tensor splitting during output, an improved video interpolation algorithm is obtained.

[0014] Preferably, obtaining the improved attention module includes: the improved attention module is used to calculate the similarity of different frames within a certain time range in the embedding space, and based on this, assign appropriate weights to different frames. For the input image sequence in the vision-based engineering vibration test task, adjacent frames that are more similar to the reference frame are taken. The similarity distance of each group of structural vibration image sequences can be calculated as:

[0015]

[0016] Where and are the image embedding vectors of the comparison frame image and the reference frame image respectively. This image embedding vector can be obtained through a simple convolutional network. The Sigmoid activation function is used to limit the input result within the range of [0,1]. For each specific spatial position, there is its specific spatial attention value. The spatial size of the similarity distance is the same as the size of .

[0017] After calculating the similarity distance of each group of structural vibration image sequences, multiply the distance feature map by the original input image in the way of pixel dot product, and use an additional convolutional layer to aggregate these attention-modulated features

[0018]

[0019] Preferably, based on the interpolated structural vibration video, using edge detection technology based on the local area effect to detect in the measurement area to obtain the time history data of the measuring point displacement, including:

[0020] The edge detection technique based on the local area effect generates a gradient image by calculating the partial derivatives of the image on the x-axis and y-axis, and manually selects a threshold to detect edge pixels. If edge pixels are detected, a straight line is fitted to determine the sub-pixel position of each edge, and the intensity ratio on both sides of the edge in each frame of the image is calculated to improve the accuracy of edge detection.

[0021] Preferably, for the structural vibration video after frame interpolation, the edge tracking technique for detecting and obtaining the time history data of the measuring point displacement in the measurement area by the edge tracking technique includes:

[0022] The edge tracking technique adopts a feature point tracking algorithm based on the edge detection method;

[0023] The feature point tracking algorithm selects an edge area that may appear throughout the vibration process in the current frame, fits sub-pixel points by the least squares method, reconstructs two approximate straight line edges, calculates the intersection point of these two straight lines as the feature point for subsequent displacement calculation. The feature point is located in the main vibration direction of the sensor to reduce errors. During the entire feature point tracking process, the first frame is used as the reference frame, and each frame is compared with the reference frame to calculate the displacement, obtaining the sub-pixel position sequence of the feature points at each moment.

[0024] Preferably, after obtaining the sub-pixel position sequence of the feature points at each moment, it includes: using matrix splicing technology in Matlab to merge all the sub-pixel information of the feature points into a matrix. The elements of each column of the matrix represent the displacement time-domain signal of a single pixel point, calculate the distance between all vector pairs in the matrix and return it to the distance matrix, and calculate the time history of the feature points according to the distance matrix.

[0025] Preferably, the monitoring of the vibration response of the laboratory model and the identification of modal parameters include: setting measuring points or visual targets on the structure, obtaining the vibration time history information of the structure through the sensing system and the signal acquisition system, and performing modal analysis on the vibration time history signal to identify the structural modal parameters. To reduce the leakage caused by the transformation from the time domain to the frequency domain, a weighting function is used to process the displacement time history.

[0026] Preferably, to reduce the leakage caused by the transformation from the time domain to the frequency domain, using a weighting function (windowing) to process the displacement time history includes:

[0027] Obtaining the Fourier spectrum Y i (f) through the fast Fourier transform (FFT), and performing mathematical decomposition to obtain the amplitude spectrum |Y i (f)|, the phase spectrum and the imaginary part spectrum Im[Y i (f)];

[0028] The peak of the amplitude spectrum corresponds to the modal order of the structure. Each peak represents the first-order mode. The abscissa of the peak represents the frequency of the mode, and the ordinate of the peak represents the amplitude of the modal shape. The amplitude sign of the measurement point is determined by the phase spectrum and the imaginary part spectrum. For the nth-order modal shape, the ith measurement point S i (f n )'s modal displacement:

[0029]

[0030] Then the nth-order modal coordinate of the ith measurement point is (x i , S i (f n ))), where x i is the position of the ith measurement point. In addition, at the support point, the ordinate of the modal shape is zero. By integrating the nth modal coordinates of all measurement points and support points, the nth modal shape is obtained.

[0031] In a second aspect of the present invention, there is provided a structural displacement measurement system based on video interpolation and computer vision, including a dynamic displacement measurement system. The dynamic displacement measurement system includes an ordinary smartphone, a photographic tripod, a YHD-100 type displacement meter, a computer, and a DHDAS test system. The dynamic displacement measurement system combines a slender aluminum alloy cantilever beam model to verify the video interpolation and computer vision-based structural displacement measurement method through experiments.

[0032] In summary, the beneficial effects of the present invention are as follows: The structural displacement measurement method and system based on video interpolation and computer vision effectively improve the sampling frequency of the video collected in non-contact measurement based on computer vision through video interpolation technology, so as to obtain vibration response data closer to the true value. In the video interpolation algorithm, an attention mechanism is introduced to improve the quality of the generated intermediate frames. When extracting the structural vibration displacement time history from the video, sub-pixel-based edge detection and improved edge tracking technology are used to improve the monitoring accuracy of the structural vibration displacement; In addition, in order to evaluate the improved video interpolation algorithm, a set of cantilever beam vibration videos are collected as a data set and compared with the linear interpolation algorithm and the quadratic interpolation algorithm. The results show that the improved algorithm has a significant improvement in the output picture quality. By using the same laboratory cantilever beam model as the measured object, displacement response monitoring and modal parameter identification tasks are carried out, and the displacement response results are compared with the data collected by the displacement meter, which proves the effectiveness and accuracy of the proposed dynamic displacement measurement method. In addition, the proposed algorithm can better identify the natural frequency and modal shape of the cantilever beam, showing potential in non-contact structural vibration testing tasks. Description of the Drawings

[0033] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly describe some of the accompanying drawings in the embodiments of the present application. It should be understood that the following drawings only show some embodiments of the present application and should not be considered as limiting the scope of the present application.

[0034] Figure 1 It is a schematic flowchart of a structural displacement measurement method based on video frame interpolation and computer vision provided by an embodiment of the present application;

[0035] Figure 2 It is a region / pixel map in which the edge is divided into two regions;

[0036] Figure 3 It is a main flowchart of the time-frequency domain transformation of the vibration signal;

[0037] Figure 4 It is a comparison of the time history curves of the vertical displacement at the free end of the cantilever beam model test: (a) Displacement time history within 10 seconds under four excitations; (b) Enlarged view of the displacement time history during the first excitation;

[0038] Figure 5 It is the correlation coefficient and the average mean absolute error under different algorithms in the cantilever beam vibration experiment. Specific embodiments

[0039] The following combines the embodiments and the attached Figure 1 To Figure 5 The present invention will be further described in detail, but the embodiments of the present invention are not limited thereto.

[0040] Please refer to Figure 1 , an embodiment of the present invention provides a structural displacement measurement method based on video frame interpolation and computer vision. This method includes the following steps:

[0041] Step 101: Collect structural vibration video data through non-contact measurement video frame interpolation technology based on computer vision;

[0042] Step 102: Introduce an attention mechanism into the video frame interpolation algorithm to obtain an improved video frame interpolation algorithm EQVI-T;

[0043] Step 103: Shoot the vibration video of the cantilever beam model as a data set to evaluate the performance of the improved video frame interpolation algorithm EQVI-T in structural vibration testing;

[0044] Step 104: Based on the interpolated structural vibration video, detect and obtain the time history data of the measured point displacement in the measurement area through edge detection technology and edge tracking technology based on the local area effect;

[0045] Step 105: Monitor the vibration response of the laboratory model and perform modal parameter identification.

[0046] Among them, in step 101, structure vibration video data is collected through video frame interpolation technology based on computer vision and non-contact measurement. Structure vibration testing is a method for evaluating and analyzing the vibration behavior of structures, which is widely used in fields such as structural dynamic performance evaluation, structural design improvement and optimization, and structural health monitoring. Compared with traditional contact measurement methods, the non-contact testing technology based on computer vision has the advantages of long-term stability, strong remote monitoring ability, controllable shooting speed, and the ability to perform non-destructive testing. However, when using the non-contact method based on computer vision technology to detect the vibration displacement of structures, it is necessary to ensure that the captured video has a sufficient frame rate. High-speed cameras can collect high-frame-rate vibration videos of structures, but their usage environment requirements are harsh and the cost is relatively high, which is not conducive to popularization in actual engineering. At present, when consumer cameras shoot at high resolutions, the upper limit of the frame rate is relatively low, making it difficult to meet the requirements of structure vibration testing. Video frame interpolation technology can insert synthetic frames between the frames of the original video to increase the video frame rate, thus solving the problem of the low frame rate of the original video.

[0047] Among them, in step 102, an attention mechanism is introduced into the video frame interpolation algorithm to obtain an improved video frame interpolation algorithm. By designing an attention mechanism on the optical flow network based on enhanced quadratic video frame interpolation and the TSA module, aggregation weights are assigned to each group of optical flows to obtain an improved attention module, which is also called the TA module. By defining the parameters of the TA module and adjusting the input size to adapt to the dataset under the engineering vibration testing task, that is, performing tensor stacking during input and tensor splitting during output, an improved video frame interpolation algorithm is obtained, which is also called EQVI-T;

[0048] The TA module is used to calculate the similarity of different frames within a certain time range in the embedding space and, based on this, allocate appropriate weights to different input frames. For the input image sequence in the vision-based engineering vibration testing task, more attention should be paid to the adjacent frames that are more similar to the reference frame. The similarity distance of each group of structure vibration image sequences can be calculated as:

[0049]

[0050] where and are the image embedding vectors of the comparison frame image and the reference frame image respectively, which can be obtained through a simple convolutional network. The Sigmoid activation function is used to limit the input result within the range of [0,1]. For each specific spatial position, there is its specific spatial attention value. Specifically, the spatial size of is the same as the size of

[0051] After the above calculations, the distance feature map is multiplied with the original input image in a pixel dot product manner and an additional convolutional layer is used to aggregate these attention-modulated features

[0052]

[0053] By redefining the parameters of the TA module and adjusting the input size to fit the dataset under the engineering vibration test task, i.e., performing tensor stacking during input and tensor splitting during output, the improved algorithm is called EQVI-T.

[0054] In some embodiments, the attention mechanism is a machine learning technique that mimics the human attention process. It allows the model to focus on specific parts of the input data according to the task requirements, thereby improving the model performance. In recent years, there have been many successful applications in computer vision tasks. In video frame interpolation, by introducing a suitable attention mechanism to improve the model, it is possible to better focus on key information to improve the image quality and accuracy of video frame interpolation, and to improve the deficiencies of existing quadratic video interpolation (QVI) and enhanced quadratic video interpolation (EQVI) in considering inter-frame relationships and weight allocation. In addition, appropriate improvements may help enhance the model's ability to process high-order motion information, making it perform better when processing structural vibration time-series videos.

[0055] Among them, in step 103, by shooting the vibration video of a typical cantilever beam model in the engineering vibration test task as the dataset and combining with a dynamic displacement measurement system, the performance of the improved video frame interpolation algorithm EQVI-T in structural vibration tests is evaluated. The improved video frame interpolation algorithm EQVI-T is evaluated through frame interpolation network training and testing, frame interpolation performance evaluation, displacement response detection results, and modal parameter identification. The specific evaluation content will be elaborated in detail later.

[0056] Among them, in step 104, based on the interpolated structural vibration video, the edge tracking technology and the edge detection technology based on the local area effect are used to detect in the measurement area to obtain the time history data of the measuring point displacement. Among them, due to the traditional pixel-level edge detection methods, such as Sobel, Prewitt and other operators, which use convolution operations to calculate gradient values and identify object edges by setting thresholds, this method can only provide approximate integer edge positions and is not ideal for applications that require higher precision. The edge detection method based on the local area effect adopted in this application assumes that the edge signal has a certain discontinuity, rather than simply describing the maximum gradient of the continuous signal. It generates a gradient image by calculating the partial derivatives of the image on the x-axis and y-axis, and manually selects a threshold to detect edge pixels. Once the edge pixels are detected, this method will fit a straight line to accurately determine the sub-pixel level position of each edge, and calculate the intensity ratio on both sides of the edge in each frame of the image to improve the accuracy of edge detection;

[0057] In some embodiments, referring to the attached Figure 2 , to achieve this goal, the method calculates the intensity ratio on both sides of the edge in each frame. Assuming that each pixel is divided into two regions by the edge y = ax + b, the regions on both sides of the edge have intensity values A and B respectively. Assuming that a, b, A, and B are unknown, the pixel value of the gray part may be a pixel with an intermediate value between A and B. This assumption can be expressed as the following equation:

[0058]

[0059] Among them, β(i, j) is the pixel value at position (i, j), which can have an intermediate value between A and B. h represents the length of the pixel side. α(i, j) is the area under the edge line within the pixel coordinates (i, j). Calculating the area under this edge line can determine the coefficients of the straight line representing the edge, so that the sub-pixel position coordinates of the edge and some other edge features can be extracted. For the curved edges that need to be processed in many cases, the second-order curve y = a + bx + cx 2 is used for approximation.

[0060] Among them, in step 104, the edge tracking technology adopts a feature point tracking algorithm based on the edge detection method. The feature point tracking algorithm includes: selecting an edge region in the frame that may appear during the entire vibration process, and fitting sub-pixel points by the least square method to reconstruct two approximate straight line edges, calculating the intersection point of these two straight lines as the feature point for subsequent displacement calculation; among them, the feature point is located in the main vibration direction of the sensor to reduce errors. During the entire feature point tracking process, the first frame is used as the reference frame, and each frame is compared with the reference frame to calculate the displacement;

[0061] In some embodiments, after obtaining the sub-pixel position sequences of feature points at each moment, the matrix splicing technique is used in Matlab to merge all the sub-pixel information of feature points into a matrix. The elements of each column of the matrix represent the displacement time-domain signal of a single pixel point. Calculate the distances between all vector pairs in the matrix and return them to the distance matrix. Finally, calculate the time history of the feature points according to the distance matrix. The above improvement can effectively track the displacement of feature points, reduce the errors of non-interested objects, and improve the robustness of target point recognition and tracking in the image sequence.

[0062] Among them, in step 105, monitor the vibration response of the laboratory model and identify the modal parameters. Set measuring points or visual targets on the structure, obtain the vibration time history information of the structure through the sensing system and the signal acquisition system, and perform modal analysis on the vibration time history signal to identify the structural modal parameters. When using computer vision technology to monitor the structural displacement, the camera records the video of the visual target during the entire vibration process, and then interpolates the video and extracts the displacement time history of the structural vibration.

[0063] In some embodiments, refer to the appendix Figure 3 , to reduce the leakage caused by the transformation from the time domain to the frequency domain, a weighting function (windowing) is used to process the displacement time history, and the Fourier spectrum Y i (f) is obtained through the fast Fourier transform (FFT), and further mathematical decomposition can be performed to obtain the amplitude spectrum Y i (f)|, the phase spectrum and the imaginary part spectrum Im[Y i (f)];

[0064] The peak of the amplitude spectrum corresponds to the modal order of the structure. Each peak represents the first-order mode. The abscissa of the peak represents the frequency of the mode, and the ordinate of the peak reflects the amplitude of the modal shape. The amplitude symbol of the measuring point is determined by the phase spectrum and the imaginary part spectrum. For the nth-order modal vibration mode, the modal displacement of the ith measuring point S i (f n ) can be determined according to the imaginary part spectrum:

[0065]

[0066] Then the nth-order modal coordinate of the ith measuring point is (x i , S i (f n ))), where x i is the position of the ith measuring point. In addition, at the support point, the ordinate of the modal vibration mode is zero. By integrating the nth modal coordinates of all measuring points and support points, the nth modal vibration mode is obtained.

[0067] In the second aspect of the present invention, the present invention provides a structural displacement measurement system based on video frame interpolation and computer vision. The system includes a dynamic displacement measurement system, which includes an ordinary smartphone, a photographic tripod, a YHD-100 type displacement meter, a computer, and a DHDAS test system. The dynamic displacement measurement system combines a slender aluminum alloy cantilever beam model to conduct a verification experiment on the structural displacement measurement method of video frame interpolation and computer vision;

[0068] In some embodiments, the longitudinal length L of the cantilever beam is 1000 mm, the cross-sectional width B is 10 mm, the cross-sectional height H is 5 mm, the Young's modulus E of the beam is 70 GPa, the Poisson's ratio is 0.33, and the density is 2750 kg / m 3 , before the test, the beam is equally divided into 10 segments, and a total of 9 measuring points C1 - C9 are sequentially set at each equal division point. The C9 measuring point is selected as the reference point, and an artificial cross mark is placed on the side at this position. The side of the thickness of the cantilever beam structure is used as the visual measurement plane. By setting artificial targets, it is ensured that the displacements of these marked points can be accurately detected and measured, thereby reducing the uncertainty caused by systematic errors or noise and improving the accuracy of displacement measurement.

[0069] The entire dynamic displacement measurement system consists of an ordinary smartphone, a photographic tripod, a YHD-100 type displacement meter, a computer, and a DHDAS test system. The DHDAS test system has data acquisition and signal processing functions. The left end of the cantilever beam model is used as a fixed end, a heavy object is placed on it and fixed with two G-clamps. Under a stable indoor light source, a smartphone is used to collect the motion video of the cross mark in the vertical direction. The distance from the camera to the mark is 30 cm. The excitation tool is a force hammer. The fixed-point excitation method is used in the experiment. The free vibration of the cantilever beam is excited by the force hammer. The sampling frequency of the DHDAS system is set to 500 Hz, and the shooting parameters of the mobile phone camera are set to FPS120 and resolution 1920*1080. After the force hammer strikes the surface of the cantilever beam vertically, the vibration video is collected for subsequent vision-based displacement measurement analysis.

[0070] In some embodiments, the experimental environment in this experiment was configured with 4 NVIDIA TITAN Xp GPUs. The operating system used was Ubuntu, and the neural network framework was the PyTorch platform. The detailed parameter settings for the improved method model during training and testing were set and read. Additionally, to evaluate the improved model, a comparative experiment was conducted, in which two different algorithm models in the field of video interpolation were used: the real-time monitoring intermediate optical flow algorithm model RIFE and the original EQVI model as references. RIFE uses traditional dual-frame input to generate intermediate frame images, while EQVI fuses four adjacent consecutive images as the input to the network and uses quadratic interpolation to interpolate the images. To ensure objectivity, both models were trained with the same amount of training data, used the same training method as the original authors of the reference models, and maintained the same hyperparameter settings. The following three evaluation methods were used to evaluate the generated image results on the cantilever beam model vibration dataset: mean squared error (MSE), peak signal-to-noise ratio (PSNR), and structural similarity index (SSIM). The evaluation results showed that EQVI-T demonstrated comparable generation capabilities to other algorithms on this dataset.

[0071] In some embodiments, referring to the appendix Figure 4 , for the research and verification of the working performance of the video interpolation method in structural displacement measurement, first, the original 120-frame low-frame-rate video collected by the smartphone was downsampled to 100 frames for subsequent interpolation comparison. Then, a self-written Matlab program was used to decompose the downsampled video into an image sequence and crop the ROI. Next, several video interpolation algorithms such as RIFE, EQVI, and the proposed EQVI-T in this paper were used to interpolate the image sequence to artificially increase the frame rate. As the basis for further analysis, the required part of the image sequence was intercepted and sent to the improved sub-pixel edge detection program. Through the above feature point tracking calculation process, the sub-pixel displacement time history of the feature points at any time was obtained. Finally, the obtained pixel coordinates were converted to physical coordinates through a scale factor.

[0072] A force hammer was used to strike at the free end of the beam to induce free vibration. The displacement response of the C9 measuring point of the cantilever beam was tracked and monitored using a measurement system based on video interpolation. Each vibration decayed within 2 s. The displacement time history of the beam in the vertical direction during four consecutive hammer strikes is shown in the appendix Figure 4 . The data of the free vibration response caused by the first hammer strike was selected and redrawn in the appendix Figure 4 to expand the time scale for clearer visualization.

[0073] Referring to the appendix Figure 4, four hammer blows caused four free vibrations. The data during each free vibration decay period (i.e., the data between [1.3, 2.4], [3.9, 4.5], [6.2, 6.9], and [8.3, 8.9] s) was used to evaluate the accuracy of video interpolation. The true displacement values for the remaining time periods were zero, and the non-zero measurement data could be regarded as measurement errors. Attached Figure 4 It can be seen that the algorithm based on linear interpolation only increases the sampling frequency and cannot restore some of the non-linear motion trajectories during the structural vibration process, just like the detection results based on low-frame-rate videos. The displacement time-history curves measured based on the EQVI and the improved EQVI-T models are relatively more consistent with the true values of the displacement gauges. Especially at the peak and trough parts, the motion trajectories of the measuring points are more restored. In addition, the detection results based on the EQVI-T model not only perform well like the EQVI during the obvious vibration decay process of the structure, but also produce less noise and have relatively smoother overall fluctuations when the structure is not vibrating.

[0074] To further quantify the application advantages of the non-linear interpolation algorithm for structural vibration videos, this paper uses two evaluation parameters as performance indicators. One is the Pearson-correlation coefficient, and the other is the mean absolute error (MAE). The Pearson-correlation coefficient is the quotient of the covariance and standard deviation of two sets of displacement time-history data, which is used to measure the correlation degree between the algorithm data and the displacement sensor data. The larger the correlation coefficient, the higher the correlation degree. The mean absolute error is used to measure the average deviation degree between each point of two displacement time-history curves. The smaller the MAE, the smaller the deviation degree. The calculation expressions for the Pearson-correlation coefficient and the mean absolute error are as follows:

[0075]

[0076]

[0077] Refer to the attached Figure 5 By analyzing the average results of multiple groups of experiments, it can be seen that the improved EQVI-T model has better correlation and better applicability in the task of increasing the frame rate of structural vibration videos.

[0078] The implementation principle of a structural displacement measurement method and system based on video frame interpolation and computer vision in an embodiment of this application is as follows: In the frame interpolation algorithm of this structural displacement measurement method based on video frame interpolation and computer vision, an attention mechanism is introduced to explore its applicability in structural vibration testing. First, by introducing the attention mechanism into the optical flow estimation network, different input frames or information are weighted to improve the model's attention to frames with higher information content, so that the generated intermediate frames are more accurate. Second, the vibration video of a typical cantilever beam model in an engineering vibration test task is taken as a data set to evaluate the performance of the improved video frame interpolation algorithm in structural vibration testing. Finally, based on the interpolated structural vibration video, the time history data of the measured point displacement is obtained by detecting in the measurement area through edge detection technology and edge tracking technology based on the local area effect, monitoring the vibration response of the laboratory model, and performing modal parameter identification.

[0079] The above are all preferred embodiments of this application, and the protection scope of this application is not limited accordingly. Therefore, all equivalent changes made according to the structure, shape, and principle of this application should be covered within the protection scope of this application.

Claims

1. A structural displacement measurement method based on video frame interpolation and computer vision, characterized in that The method includes: Collecting structural vibration video data through video frame interpolation technology based on computer vision with non-contact measurement; Introducing an attention mechanism into the video frame interpolation algorithm to obtain an improved video frame interpolation algorithm; Shooting the vibration video of the cantilever beam model as a data set to evaluate the performance of the improved video frame interpolation algorithm in structural vibration testing; Based on the interpolated structural vibration video, obtaining the time history data of the measuring point displacement by detecting in the measurement area through edge detection technology and edge tracking technology based on the local area effect; Monitoring the vibration response of the laboratory model and performing modal parameter identification; The introducing an attention mechanism into the video frame interpolation algorithm to obtain an improved video frame interpolation algorithm includes: Designing an attention mechanism on the optical flow network based on enhanced quadratic video frame interpolation and TSA module, assigning aggregation weights to each group of optical flow, and obtaining an improved attention module; By defining the parameters of the improved attention module and adjusting the input size to adapt to the data set under the engineering vibration test task, that is, performing tensor stacking during input and tensor splitting during output, an improved video frame interpolation algorithm is obtained; The obtaining the improved attention module includes: The improved attention module is used to calculate the similarity of different frames within a certain time range in the embedding space, and based on this, assign appropriate weights to different frames. For the input image sequence in the vision-based engineering vibration test task, adjacent frames that are more similar to the reference frame are selected. The similarity distance of each group of structural vibration image sequences can be calculated as: where and are the image embedding vectors of the comparison frame image and the reference frame image respectively. The image embedding vector can be obtained through a simple convolutional network. The Sigmoid activation function is used to limit the input result within the range of [0, 1]. For each specific spatial position, there is a specific spatial attention value. The spatial size of the similarity distance is the same as that of ; After calculating the similarity distance of each group of structural vibration image sequences the distance feature map is multiplied with the original input image in the form of pixel dot product and an additional convolutional layer is used to aggregate these attention-modulated features 2. The structural displacement measurement method based on video frame interpolation and computer vision according to claim 1, characterized in that The obtaining the time history data of the measuring point displacement by detecting in the measurement area through edge detection technology based on the local area effect based on the interpolated structural vibration video includes: Based on the edge detection technology based on the local area effect, generating a gradient image by calculating the partial derivatives of the image on the x-axis and y-axis, and manually selecting a threshold to detect edge pixels. If edge pixels are detected, fitting a straight line to determine the sub-pixel level position of each edge, and calculating the intensity ratio on both sides of the edge in each frame of the image to improve the accuracy of edge detection.

3. The structural displacement measurement method based on video frame interpolation and computer vision according to claim 1, characterized in that, The edge tracking technology in the obtaining the time history data of the measuring point displacement by detecting in the measurement area through edge tracking technology based on the interpolated structural vibration video includes: The edge tracking technology adopts a feature point tracking algorithm based on the edge detection method; The feature point tracking algorithm selects an edge area in the current frame that may appear throughout the vibration process, and fits sub-pixel points by the least squares method to reconstruct two approximate straight edges. The intersection point of these two straight lines is calculated as the feature point for subsequent displacement calculation. The feature point is located in the main vibration direction of the sensor to reduce errors. During the entire feature point tracking process, the first frame is used as the reference frame, and each frame is compared with the reference frame to calculate the displacement, obtaining the sequence of sub-pixel positions of the feature points at each moment.

4. The method for measuring structural displacement based on video frame interpolation and computer vision according to claim 3, wherein After obtaining the sequence of sub-pixel positions of the feature points at each moment, it includes: In Matlab, matrix splicing technology is used to merge the sub-pixel information of all feature points into a matrix. Each column element of the matrix represents the displacement time domain signal of a single pixel point. The distance between all vector pairs in the matrix is calculated and returned to the distance matrix. The time history of the feature points is calculated based on the distance matrix.

5. The structural displacement measurement method based on video frame interpolation and computer vision according to claim 1, wherein The monitoring of the laboratory model vibration response and the identification of modal parameters include: Measuring points or visual targets are set on the structure, and the vibration time history information of the structure is obtained through the sensing system and signal acquisition system. The modal parameters of the structure can be identified by performing modal analysis on the vibration time history signals. In order to reduce the leakage caused by the transformation from time domain to frequency domain, the displacement time history is processed by a weighted function.

6. The method for measuring structural displacement based on video frame interpolation and computer vision according to claim 5, characterized in that In order to reduce the leakage caused by the transformation from the time domain to the frequency domain, a weighted function (windowing) is used to process the displacement time history, including: The Fourier spectrum Y is obtained by fast Fourier transform (FFT) i (f), and through mathematical decomposition, the amplitude spectrum |Y i (f)|, the phase spectrum and the imaginary part spectrum Im[Y i (f)] are obtained; The peak of the amplitude spectrum corresponds to the modal order of the structure. Each peak represents the first-order mode. The abscissa of the peak represents the frequency of the mode, and the ordinate of the peak represents the amplitude of the modal shape. The amplitude sign of the measurement point is determined by the phase spectrum and the imaginary part spectrum. For the nth-order modal shape, the ith measurement point S can be determined according to the imaginary part spectrum. i (f n ) modal displacement: Then the n-th modal coordinate of the i-th measurement point is (x i , S i (f n ))), where x i is the position of the i-th measurement point. In addition, at the support points, the ordinate of the modal shape is zero. By integrating the n-th modal coordinates of all measurement points and support points, the n-th modal shape is obtained.

7. A structural displacement measurement system based on video frame interpolation and computer vision, characterized in that, The invention comprises a dynamic displacement measurement system, which comprises an ordinary smart phone, a photography tripod, a YHD-100 displacement meter, a computer and a DHDAS test system. The dynamic displacement measurement system is combined with a slender aluminum alloy cantilever beam model to conduct a verification experiment on a structural displacement measurement method of video interpolation and computer vision.