Semi-dense sub-pixel displacement measurement method based on LoFTR

Through the semi-dense sub-pixel displacement measurement method based on LoFTR, the problem of insufficient robustness of visual displacement measurement in complex scenes is solved, and high-precision and stable displacement measurement is achieved, which is suitable for complex and changeable scenes.

CN120689383APending Publication Date: 2025-09-23NANJING BRIDGE & TUNNEL INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510655940.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing visual displacement measurement methods lack robustness in complex scenes, especially in feature-less or dynamic environments, where their accuracy decreases. Their reliance on specific landmarks limits their versatility and flexibility.

Method used

A semi-dense sub-pixel displacement measurement method based on LoFTR is adopted. By improving the camera housing design, image preprocessing and LoFTR matching framework, combined with the RANSAC algorithm and linear interpolation method, sub-pixel displacement measurement is achieved without relying on specific landmarks.

Benefits of technology

High-precision measurements can be achieved on both textured and texture-poor surfaces, improving the robustness and continuity of the measurement and enhancing the stability of the method under different displacement change conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689383A_ABST
    Figure CN120689383A_ABST
Patent Text Reader

Abstract

The invention discloses a semi-dense sub-pixel displacement measurement method based on LoFTR, and the method comprises the steps: carrying out the front shooting of a target structure through an improved camera, obtaining a corresponding image, and carrying out the preprocessing of the image, and obtaining a matching pair; processing the matching pairs by using an LoFTR matching framework, and establishing a semi-dense key point matching relationship between frame pairs to obtain a feature point matching set; and based on the feature point matching set, determining the sub-pixel-level relative displacement of the target structure by using an RANSAC algorithm and a linear interpolation method, and completing the measurement of the sub-pixel displacement. According to the invention, high-precision semi-dense sub-pixel displacement measurement can be realized on surfaces with abundant textures and surfaces without textures.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of structural health monitoring, and in particular to a semi-dense sub-pixel displacement measurement method based on LoFTR. Background Art

[0002] SHM (Structural Health Monitoring) is an important means of ensuring the safety and reliability of civil engineering structures and is widely used in the performance evaluation and safety monitoring of critical infrastructure such as bridges and buildings. By providing real-time observation of the dynamic response and deformation characteristics of a structure during its service life, SHM provides critical data support for structural safety assessment, maintenance strategy optimization, and disaster warning. Displacement measurement is a core component of the SHM system, reflecting the deformation response of the structure under external loads and providing an important basis for identifying structural performance degradation and potential risks.

[0003] The template matching-based method is a regional benchmark image matching method that achieves target positioning and displacement measurement by searching for the area in the image sequence that is most similar to the predefined template. Generally, this method uses grayscale value similarity metrics (such as normalized cross-correlation and mean square error) to evaluate the degree of match between the template and the target area, and further improves the accuracy through sub-pixel interpolation technology. The template matching-based method has been widely used in bridge deformation monitoring and high-rise building vibration analysis. Its advantage is that it can achieve high-precision point displacement measurement. However, since template matching relies on predefined target templates, its performance is easily affected by illumination changes, target occlusion and background interference, especially in long-term field monitoring or dynamic environments. It shows insufficient robustness.

[0004] Methods based on sparse keypoint matching extract salient feature points from an image and estimate displacement information using the matching relationships between these feature points. This method typically exploits geometric transformation invariance and illumination invariance, making it highly robust to changes in target appearance or complex environments. However, sparse keypoint matching requires a high level of feature point abundance in the target region, and its matching accuracy may decline in the absence of distinct features or when features are sparse. Furthermore, keypoint matching methods suffer from poor real-time performance in dynamic target tracking and have limited applicability to irregular or low-feature target regions.

[0005] In summary, while traditional visual displacement measurement methods (such as template matching, optical flow estimation, and sparse keypoint matching) have met the displacement measurement needs of structural health monitoring to a certain extent, their limitations are particularly evident in complex scenes. Deep learning-based methods have recently demonstrated strong potential in the field of visual displacement measurement. Leveraging the feature extraction capabilities of multi-layer neural networks, deep learning can automatically learn complex nonlinear relationships, overcoming the traditional methods' reliance on specific templates or salient features, making displacement measurement more robust and accurate in complex scenes. However, current deep learning-based methods still have several limitations. First, most methods focus on detecting specific targets (such as nodes or artificial landmarks), requiring network training tailored to specific scenes and landmarks. Furthermore, while existing methods use deep learning as an auxiliary tool to improve image quality or enhance specific features, these techniques often still rely on traditional methods (such as template matching) to achieve sub-pixel displacement accuracy. This reliance on specific landmarks limits the methods' versatility and flexibility, particularly in scenes without landmarks or where target features are less prominent.

[0006] Meanwhile, deep learning-based dense feature point detection and matching techniques have made significant progress in visual localization and 3D reconstruction, demonstrating superior performance in feature extraction and matching tasks. However, these techniques have yet to be fully applied to displacement measurement with sub-pixel accuracy. Existing research lacks systematic validation of the feasibility of deep learning-based dense feature point detection and matching techniques for structural health monitoring, which provides new directions and possibilities for further research. Summary of the Invention

[0007] The purpose of the present invention is to provide a semi-dense sub-pixel displacement measurement method based on LoFTR, which uses dense feature point detection based on deep learning to realize sub-pixel displacement measurement, aiming to achieve higher precision and robust displacement measurement without relying on specific markers and can be directly applied to complex and changeable scenes.

[0008] The present invention adopts the following technical solution: a semi-dense sub-pixel displacement measurement method based on LoFTR, comprising the following steps:

[0009] S1. Use the improved camera to shoot the target structure from the front to obtain the corresponding image, and preprocess the image to obtain a matching pair.

[0010] S2. Use the LoFTR matching framework to process the matching pairs, establish a semi-dense key point matching relationship between the frame pairs, and obtain the feature point matching set.

[0011] S3. Based on the feature point matching set, the RANSAC algorithm and linear interpolation method are used to determine the sub-pixel relative displacement of the target structure and complete the measurement of sub-pixel displacement.

[0012] Furthermore, an improved camera housing is designed using shock-proof and vibration-reducing materials, and a sealed and dust-proof structure is designed to obtain an improved camera.

[0013] Furthermore, in step S1, the matching pair obtained includes the following:

[0014] The expression of the projection transformation relationship obtained by camera calibration is:

[0015]

[0016] Where SF is the conversion coefficient between the 2D image coordinate system and the 3D structure coordinate system, and its unit is mm / pixel; p is the unit length of the camera sensor, and its unit is mm / pixel; Z is the actual distance between the camera position and the target structure; and f is the focal length of the camera.

[0017] Multiplying the pixel displacement with the SF gives the actual displacement of the target structure.

[0018] Furthermore, in step S1, a region to be processed in the image is selected as a region of interest image.

[0019] The region of interest in the first frame of the image is divided into a set number of mask region images; the mask region images correspond to different local region images in the image; and a displacement image sequence with a displacement range of 1 to 10 pixels is constructed in the vertical and horizontal directions of the mask region images using a 16-neighborhood bicubic interpolation algorithm with a unit of 0.1 pixel.

[0020] The region of interest of the first frame in the image is paired with the regions of interest of all frames in sequence to generate matching pairs.

[0021] Furthermore, in step S3, determining the sub-pixel relative displacement of the target structure includes the following steps:

[0022] The true displacement value and the displacement value of the displacement image sequence are plotted as the horizontal and vertical coordinates respectively to form a curve graph. The curve graph shows a monotonically increasing trend, indicating that there is a consistent corresponding relationship between the true displacement value and the displacement value of the displacement image sequence.

[0023] Based on this consistent correspondence, the mask area image is divided into a set number of sub-areas Ω1, Ω2, ..., Ω along the main axis direction of the target structure. n , where Ω n Indicates the nth sub-region.

[0024] The RANSAC (Random Sample Consensus) algorithm is used to perform geometric consistency screening on the feature point matching subset in the sub-region, and the feature points in the feature point matching subset in the sub-region are set to satisfy the similarity transformation model. The specific expression is:

[0025] k k,i =Ak 1,i +t

[0026] Among them, k k,i represents the two-dimensional coordinates of the i-th feature point in the k-th frame image, A represents the scaling and rotation matrix, and t represents the translation vector.

[0027] Calculate the geometric error e of the i-th feature point i , the specific expression is:

[0028] e i =‖k k,i -(Ak 1,i +t)‖2.

[0029] When e i ≤d max When , the feature point is an interior point, and the interior point set is obtained; where d max Indicates the error threshold.

[0030] Based on the internal point set, calculate the initial relative displacement of the jth sub-region in the kth frame image The specific expression is:

[0031]

[0032] in, represents the set of interior points of the j-th subregion, x k,i Indicates the horizontal coordinate of the i-th feature point in the k-th frame image in the pixel coordinate system, x 1,i Indicates the horizontal coordinate of the i-th feature point in the first frame image in the pixel coordinate system, It represents the average vertical displacement of all internal points in the jth subregion of the kth frame image relative to the 1st frame image.

[0033] Perform deviation correction on the initial relative displacement to obtain the initial relative displacement after deviation correction The specific expression is:

[0034]

[0035] in, represents the horizontal deviation of the j-th sub-region.

[0036] The initial relative displacement after the deviation correction is corrected using the linear interpolation method to obtain the sub-pixel relative displacement of the target structure.

[0037] Furthermore, the present invention also proposes an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the semi-dense sub-pixel displacement measurement method based on LoFTR when executing the computer program.

[0038] Furthermore, the present invention also proposes a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is run by a processor, the semi-dense sub-pixel displacement measurement method based on LoFTR is executed.

[0039] Compared with the prior art, the present invention adopts the above technical solution and has the following technical effects:

[0040] 1. The present invention can achieve high measurement accuracy in both texture-rich areas and texture-poor surfaces, and its effectiveness in modal identification and cable force identification has also been verified.

[0041] 2. The LoFTR in the present invention can provide stable matching results in different regions, thereby improving the robustness of displacement measurement.

[0042] 3. By introducing error curve slope analysis, the present invention effectively improves the continuity and accuracy of the measurement results and enhances the stability of the method under different displacement change conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 It is an overall implementation flow chart of the present invention.

[0044] Figure 2 2 is a schematic structural diagram of a camera housing according to an embodiment of the present invention.

[0045] Figure 3 3 is a schematic diagram of selecting a region of interest for the first frame of image and a corresponding mask image in an embodiment of the present invention.

[0046] Figure 4 2 is a diagram showing the applicability analysis results of the present invention.

[0047] Figure 5 This is a visual comparison diagram of different feature point matching methods obtained by different methods in the embodiments of the present invention.

[0048] Figure 6 Schematic diagram of the modal parameter identification results of the cantilever beam structure in an embodiment of the present invention.

[0049] Figure 7Schematic diagram of the cable force test of the embodiment of the present invention.

[0050] Figure 8 : This is the Fourier spectrum of the vibration displacement of the cable tested by the UAV in the embodiment of the present invention.

[0051] Figure 9 1 is a diagram showing the cable force identification results in an embodiment of the present invention. DETAILED DESCRIPTION

[0052] In order to make the purpose, technical solutions and advantages disclosed in the embodiments of the present invention clearer, the embodiments of the present invention are further described in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the embodiments of the present invention and are not intended to limit the embodiments of the present invention. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application. Examples of the embodiments are shown in the accompanying drawings, where the same or similar numbers throughout represent the same or similar elements or elements with the same or similar functions.

[0053] It should be noted that the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or server that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or are inherent to these processes, methods, products or devices.

[0054] Like reference numerals and letters denote like items in the following drawings, and thus, once an item is defined in one drawing, it does not require further definition or explanation in subsequent drawings.

[0055] In the description of the present invention, it should be noted that the terms "upper", "lower", "inside", "outside", etc. indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, or are the orientations or positional relationships in which the inventive product is usually placed when in use. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, they should not be understood as limiting the present invention.

[0056] In the description of the present invention, it should also be noted that, unless otherwise expressly specified or limited, the terms "disposed," "installed," and "connected" should be understood broadly. For example, they may refer to fixed, detachable, or integral connections; mechanical or electrical connections; direct or indirect connections through an intermediate medium; and internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on the specific circumstances.

[0057] The present invention will be further described below in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and are not intended to limit the scope of protection of the present invention.

[0058] To achieve the above objectives, the present invention proposes a semi-dense sub-pixel displacement measurement method based on LoFTR, such as Figure 1 The specific steps are as follows:

[0059] The S1 test beam has a total length of 2 meters, with 0.2 meters of it serving as a fixed end. Boundary constraints are applied via supports. To enhance visual features, the cantilever beam surface is sprayed with a speckle pattern, and six circular markers are affixed to assist in motion tracking. The cantilever beam is struck with a hammer to generate an impact load, and a modified camera is used to record a dynamic image sequence of the impact response. The modified camera has an image resolution of 512 × 1024 pixels and a sampling frequency of 1000 Hz, capturing images from the front of the cantilever beam. The captured images undergo preprocessing, including camera calibration, region of interest selection, applicability analysis, and frame pair construction. This is done to complete the algorithm's applicability analysis, reduce computational complexity, and ultimately generate matching pairs.

[0060] To ensure the calibration accuracy, the present invention uses a specially designed camera housing, such as Figure 2 As shown, the housing is made of shock-proof and vibration-damping materials and is designed with a sealed and dust-proof structure to reduce the impact of environmental interference on the camera posture and intrinsic parameters, thereby improving the accuracy of calibration and subsequent displacement measurements.

[0061] Camera calibration: The expression of the projection transformation relationship obtained through camera calibration is:

[0062]

[0063] Where SF is the conversion coefficient between the 2D image coordinate system and the 3D structure coordinate system, and its unit is mm / pixel; p is the unit length of the camera sensor, and its unit is mm / pixel; Z is the actual distance between the camera position and the target structure; and f is the focal length of the camera.

[0064] Multiplying the pixel displacement with the SF gives the actual displacement of the target structure.

[0065] In actual operation, local pixel size measurements can be performed on areas such as concrete surfaces, and an average value can be read in combination with the average physical size of the structure to achieve the conversion from pixel displacement to actual displacement.

[0066] Region of Interest (ROI) selection: Select the area of ​​interest (ROI) in the image to be processed, rather than the entire image. ROI selection can help you avoid irrelevant background interference and further focus on key targets.

[0067] Two region of interest images are divided on the first frame image, such as Figure 3 As shown, the size of the first region of interest image is 496×96 pixels, the size of the second region of interest image is 528×96 pixels, and all subsequent images are cropped within the range of the region of interest image.

[0068] The improved camera underwent two sets of shock and vibration tests: the first set applied shock at circular marker No. 3, and the second set applied shock at circular marker No. 5, collecting 3034 and 5147 frames, respectively. Framepairs were constructed from the two region-of-interest images, consisting of the first frame and its corresponding time series image. The first set of tests generated 3034 pairs of framepairs, while the second set generated 5147 pairs. The LoFTR matching framework provides four pre-trained models for direct use: indoor_DS, indoor_OT, outdoor_DS, and outdoor_OT. "Indoor" and "Outdoor" represent models optimized for indoor and outdoor scenarios, respectively, to accommodate different application requirements. "OT" and "DS" represent different feature matching strategies. Optimal Transport (OT) emphasizes global consistency and high-precision matching, making it suitable for tasks requiring high precision, while Dual-Softmax (DS) offers higher computational efficiency and is more suitable for applications requiring high real-time performance.

[0069] Applicability analysis: The region of interest in the first frame of the image was divided into 15 masked region images along the beam length. The masked region images correspond to different local region images in the image, facilitating subsequent more refined local displacement analysis. The outdoor_DS and outdoor_OT models provided by LoFTR, trained on outdoor data, were selected for applicability analysis. Using a 16-neighborhood bicubic interpolation algorithm, a 101-frame displacement image sequence with a displacement range of 1 to 10 pixels (or appropriately expanded based on the estimated structural displacement range, such as simulating a range of 20 pixels or even larger) was constructed in the vertical and horizontal directions of the masked region images, using 0.1 pixel units.

[0070] Frame pair construction: The region of interest of the first frame in the image is paired with the regions of interest of all frames in sequence to generate matching pairs.

[0071] S2. Use the LoFTR (Local Feature Matching with Transformers) framework to process the matching pairs, establish a semi-dense key point matching relationship between the frame pairs, and obtain a feature point matching set.

[0072] S3. Based on the feature point matching set, combined with mask area segmentation, random sampling consistent feature point screening, displacement calculation and error correction, the sub-pixel relative displacement of the target structure is determined and the sub-pixel displacement measurement is completed; the specific content is as follows:

[0073] The true displacement value and the displacement value of the displacement image sequence are plotted as the horizontal and vertical coordinates respectively to form a curve graph. The curve graph shows a monotonically increasing trend, indicating that there is a consistent corresponding relationship between the true displacement value and the displacement value of the displacement image sequence.

[0074] Based on this consistent correspondence, the mask area image is divided into a set number of sub-areas Ω1, Ω2, ..., Ω along the main axis direction of the target structure. n , where Ω n represents the nth subregion. Ω1, Ω3, and Ω6 correspond to the area where the circular marker is attached, and Ω2, Ω4, and Ω5 correspond to the other beam segments of the cantilever beam, covering smooth areas without feature point markers, to evaluate the matching stability and displacement calculation accuracy of different regions.

[0075] Figure 4 (a) shows the comparison results of the calculated and simulated displacements of the three sub-mask regions Ω4, Ω5, and Ω6 under the outdoor_DS and outdoor_OT models. As can be seen from the figure, within the displacement range of 1-10 pixels, the degree of consistency between the predicted displacements and the true values ​​in different regions varies. Although Ω4 and Ω5 are both general beam sections of the cantilever beam, their matching accuracy varies due to different background images. Among them, in the Ω4 region, the outdoor_ds and outdoor_ot models both show better matching results than Ω5 and Ω6. However, within the displacement range of 0-4 pixels, all regions show good consistency. In the 4-6 pixel range, the predicted displacements of Ω5 and Ω6 fail to maintain good monotonicity, and the calculated values ​​grow slowly, resulting in reduced stability of the constructed interpolation function.

[0076] To further analyze the matching characteristics, the first ROI image is rotated 90°. At this time, the calculated horizontal displacement of the rotated image is equivalent to the vertical displacement of the original image. Figure 4(b) shows the comparison of the calculated and simulated displacements for Ω4, Ω5, and Ω6 after rotation. It can be observed that for the Ω4 region, both the original and rotated images maintain a high degree of fit with the true values. For the Ω5 and Ω6 regions, the deviation between the displacements calculated from the rotated images and the true values ​​is significantly smaller than that calculated from the original images. Furthermore, in the Ω6 region, the displacements calculated using the outdoor_DS model match the displacements calculated using the outdoor_OT model better. After comprehensive consideration, the outdoor_DS model was ultimately selected to calculate the horizontal displacements of the rotated images and use them as the vertical displacements for the corresponding regions in the original images. A corrected interpolation table was constructed for each region to further improve the calculation accuracy.

[0077] The RANSAC algorithm is used to perform geometric consistency screening on the feature point matching subset within the sub-region, and the feature points in the feature point matching subset within the sub-region are set to satisfy the similarity transformation model. The specific expression is:

[0078] k k,i =Ak 1,i +t

[0079] Among them, k k,i represents the two-dimensional coordinates of the i-th feature point in the k-th frame image, A represents the scaling and rotation matrix, and t represents the translation vector.

[0080] Calculate the geometric error e of the i-th feature point i , the specific expression is:

[0081] e i =‖k k,i -(Ak 1,i +t)‖2

[0082] When e i ≤d max When , the feature point is an interior point, where d max Represents the error threshold; obtains the set of internal points.

[0083] Based on the internal point set, calculate the initial relative displacement of the jth sub-region in the kth frame image The specific expression is:

[0084]

[0085] in, represents the set of interior points of the j-th subregion, x k,i Indicates the horizontal coordinate of the i-th feature point in the k-th frame image in the pixel coordinate system, x 1,i Indicates the horizontal coordinate of the i-th feature point in the first frame image in the pixel coordinate system, represents the average vertical displacement of all internal points in the jth subregion of the kth frame image relative to the first frame image, yk,i Indicates the vertical coordinate of the i-th feature point in the k-th frame image in the pixel coordinate system, y 1,i Indicates the vertical coordinate of the i-th feature point in the first frame image in the pixel coordinate system.

[0086] Perform deviation correction on the initial relative displacement to obtain the initial relative displacement after deviation correction The specific expression is:

[0087]

[0088] in, represents the horizontal deviation of the j-th sub-region, represents the vertical deviation of the j-th sub-region.

[0089] The initial relative displacement after the deviation correction is corrected using the linear interpolation method to obtain the sub-pixel relative displacement of the target structure.

[0090] The present invention compares the displacement measurement effects of the circle center detection algorithm and the DIC (Digital Image Correlation) algorithm. The circle center detection algorithm adopts the gradient-based Hough transform method, while the DIC algorithm adopts the open source DIC toolkit Ncorr. The present invention measures the structural responses of 15 sub-areas calculated under small displacement conditions and large displacement conditions. Since the high-speed camera shoots the cantilever beam head-on, all measurements are evaluated in pixels. In the small displacement condition, the maximum amplitude of all sub-areas is within 1 pixel, which can better evaluate the sub-pixel measurement accuracy of different algorithms. In the large displacement condition, the amplitude increases significantly, and the maximum amplitude is about 4.2 pixels.

[0091] In larger displacement measurements under larger displacement conditions, for areas with circular markers, the proposed method, the center detection algorithm, and the DIC algorithm all achieved relatively consistent measurement results, with a high degree of agreement among the three. However, in small displacement tests under smaller displacement conditions, the proposed method achieved significantly higher agreement with the DIC calculation results than the center detection algorithm. Furthermore, in areas without markers under both small and larger displacement conditions, the proposed method and DIC maintained good consistency in measurement results.

[0092] like Figure 5As shown in the figure, the present invention compares LoFTR with several traditional feature point detection and matching algorithms, including ORB (Oriented FAST and Rotated BRIEF, oriented fast feature points and rotation invariant BRIEF descriptor), SIFT (Scale-Invariant Feature Transform, scale invariant feature transform) and BRISK (Binary Robust Invariant Scalable Keypoints, accelerated robust feature key point detection), and analyzes their applicability in cantilever beam vibration testing. Since the feature point coordinates generated by ORB are integer pixel level by default, the method of the present invention adopts sub-pixel optimization method to refine the coordinates to improve the accuracy of feature point matching. Specifically, ORB is first used to extract feature points and descriptors, and then the detected key points are converted into floating-point coordinates, and cv2.cornerSubPix is ​​used to further optimize its sub-pixel accuracy to improve the matching effect.

[0093] In terms of feature point distribution, the matching points generated by LoFTR are relatively evenly distributed across the entire region of interest, while traditional methods generate far fewer matching points in areas with attached markers than in beam sections with sprayed speckles. This allows LoFTR to provide stable matching results across different regions, while traditional methods have limited applicability. Furthermore, in areas where the structure vibrates, the number of feature point matches using traditional matching methods decreases significantly, or even fails completely. This is especially true in the presence of large amplitude displacements, where SIFT and BRISK also struggle to maintain stable feature point matching.

[0094] In terms of measurement accuracy, the displacement measurements obtained using different feature matching methods show that the traditional method has high matching accuracy in the speckle area, but the matching accuracy drops significantly in the area where the markers are attached. In particular, the ORB method, even with sub-pixel optimization, still has significantly lower displacement calculation accuracy for matching points than other algorithms. In addition, the traditional feature point detection algorithm has many matching deficiencies during the matching process, especially in high-frequency vibration areas and beam sections without obvious textures. This makes it difficult to obtain sufficient feature points for displacement calculation during the measurement process. LoFTR can provide more stable matching results, thereby effectively improving measurement accuracy and robustness.

[0095] After obtaining the dynamic displacement time history curve of the semi-dense measurement points of the cantilever beam structure under impact load, the frequency response function matrix of the structure is estimated by combining the input impact load. By performing singular value decomposition on this frequency response function matrix, the modal parameters such as frequency and vibration mode of the cantilever beam structure can be obtained. The singular value spectrum of the frequency response function estimated by visual measurement displacement is as follows: Figure 6As shown in (a), it can be seen that the first two frequencies of the cantilever beam structure are 16.2Hz and 103Hz respectively. In order to verify the correctness of the identified frequencies, the DIC algorithm is used to calculate the dynamic displacement time history curves of the dense measurement points of the cantilever beam structure. The first two frequencies identified are 16.1Hz and 103Hz, which are basically consistent. In addition, the first two displacement vibration modes of the cantilever beam structure are identified, as shown in Figure 6 As shown in (b) (first-order vibration mode) and (c) (second-order vibration mode), it can be seen from the figure that the displacement vibration mode extracted by the proposed LoFTR algorithm for displacement identification is completely consistent with the displacement vibration mode extracted by the DIC algorithm, which further verifies the accuracy of the proposed method in sub-pixel displacement identification.

[0096] Since the displacement amplitude of the cantilever beam vibration experiment is small, in order to further verify the robustness of the proposed method in large displacement calculation, the present invention carried out the cable vibration displacement extraction and cable force calculation experiments in a laboratory environment. Figure 7 As shown in Figure 2 . The experiment used accelerometers, drone-mounted cameras, and fixed cameras to record the vibration response of the cable under excitation loads, and calculated the cable force based on the vibration method. During the experiment, a PCB393B04 accelerometer was used for data acquisition, coupled with a National Instruments (NI) Model PXIe-1082 data acquisition system. The sampling frequency was set to 1000 Hz, and the data acquisition time was 2 minutes. For visual measurement, the fixed camera had a resolution of 2048×2048 pixels and a sampling frequency of 25 Hz, while the drone-mounted camera had a resolution of 3840×2160 pixels and a sampling frequency of 30 Hz. During the experiment, the cable was gently tapped with a hammer to stimulate free vibration. The vibration responses recorded by the drone-mounted camera, fixed camera, and accelerometer were simultaneously collected to evaluate the accuracy of the cable force calculation using different sensing methods.

[0097] After obtaining the vibration displacement of the cable, Fourier transform is performed on it to identify its natural frequency parameters, and then combined with the physical parameters of the cable to estimate the internal force of the cable. The Fourier spectrum of the cable vibration displacement estimated by LoFTR and SIFT method is shown in the figure below. Figure 8 As shown in the figure, it can be seen that the first two vibration frequencies of the cable identified by the two methods are exactly the same, which verifies the correctness of the extracted displacement and can be further used to estimate the cable force of the cable. The cable force estimated by using the first two frequency differences and the mass per unit length of the cable and the cable length is as follows: Figure 9As shown in the figure, (UAV, Accelerometer, and Camera represent the estimation of cable force using drone, accelerometer, and standing camera, respectively; SIFT is the traditional feature point detection method) In order to verify the correctness of the visual method for estimating cable force, the cable force estimated using accelerometer and standing camera are also plotted in the figure. It can be seen from the figure that the cable force estimated by the displacement extracted by the two feature detection algorithms is consistent with the reference value, with an error of less than 2%, verifying the correctness of the proposed method in large displacement measurement.

[0098] An embodiment of the present invention further provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable by the processor. It should be noted that when the processor executes the computer program, it corresponds to the specific steps of the method provided in the embodiment of the present invention and has the corresponding functional modules and beneficial effects of the method. For technical details not fully described in this embodiment, please refer to the method provided in the embodiment of the present invention.

[0099] The present invention also provides a computer-readable storage medium storing a computer program. It should be noted that when executed by a processor, the computer program corresponds to the specific steps of the method provided in the present invention and has the corresponding functional modules and beneficial effects. For technical details not fully described in this embodiment, please refer to the method provided in the present invention.

[0100] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A semi-dense sub-pixel displacement measurement method based on LoFTR, characterized in that: include: S1. Use the improved camera to shoot the target structure from the front to obtain the corresponding image, and preprocess the image to obtain a matching pair; S2. Use the LoFTR matching framework to process the matching pairs, establish a semi-dense key point matching relationship between the frame pairs, and obtain the feature point matching set; S3. Based on the feature point matching set, the RANSAC algorithm and linear interpolation method are used to determine the sub-pixel relative displacement of the target structure and complete the measurement of sub-pixel displacement.

2. The semi-dense sub-pixel displacement measurement method based on LoFTR according to claim 1, characterized in that: An improved camera housing is designed using shock-proof and vibration-reducing materials, and a sealed and dust-proof structure is designed to obtain an improved camera.

3. The semi-dense sub-pixel displacement measurement method based on LoFTR according to claim 1, characterized in that: In step S1, the matching pair obtained includes the following: The expression of the projection transformation relationship obtained by camera calibration is: Where SF represents the conversion coefficient between the two-dimensional image coordinate system and the three-dimensional structure coordinate system, p represents the unit length of the camera sensor, Z represents the actual distance between the camera position and the target structure, and f represents the focal length of the camera. Multiplying the pixel displacement with the SF gives the actual displacement of the target structure.

4. The semi-dense sub-pixel displacement measurement method based on LoFTR according to claim 1, characterized in that: In step S1, a region to be processed in an image is selected as a region of interest image; The region of interest in the first frame of the image is divided into a set number of mask region images; the mask region images correspond to different local region images in the image; using the 16-neighborhood bicubic interpolation algorithm, a displacement image sequence with a displacement range of 1 to 10 pixels is constructed in the vertical and horizontal directions of the mask region images in units of 0.1 pixels; The region of interest of the first frame in the image is paired with the regions of interest of all frames in sequence to generate matching pairs.

5. The semi-dense sub-pixel displacement measurement method based on LoFTR according to claim 4, characterized in that: In step S3, determining the sub-pixel relative displacement of the target structure includes the following: The true displacement value and the displacement value of the displacement image sequence are plotted as the horizontal and vertical coordinates respectively to form a curve graph, which shows a monotonically increasing trend, indicating that there is a consistent corresponding relationship between the true displacement value and the displacement value of the displacement image sequence; Based on this consistent correspondence, the mask area image is divided into a set number of sub-areas Ω1, Ω2, ..., Ω along the main axis direction of the target structure. n , where Ω n represents the nth sub-region; The RANSAC algorithm is used to perform geometric consistency screening on the feature point matching subset within the sub-region, and the feature points in the feature point matching subset within the sub-region are set to satisfy the similarity transformation model. The specific expression is: k k,i =And 1,i +t Among them, k k,i represents the two-dimensional coordinates of the i-th feature point in the k-th frame image, A represents the scaling and rotation matrix, and t represents the translation vector; Calculate the geometric error e of the i-th feature point i , the specific expression is: and i =‖k k,i -(And 1,i +t)‖2; When e i ≤d max When , the feature point is an interior point, and the interior point set is obtained; where d max represents the error threshold; Based on the internal point set, calculate the initial relative displacement of the jth sub-region in the kth frame image The specific expression is: in, represents the set of interior points of the j-th subregion, x k,i Indicates the horizontal coordinate of the i-th feature point in the k-th frame image in the pixel coordinate system, x 1,i Indicates the horizontal coordinate of the i-th feature point in the first frame image in the pixel coordinate system, represents the average vertical displacement of all internal points in the jth subregion of the kth frame image relative to the first frame image; Perform deviation correction on the initial relative displacement to obtain the initial relative displacement after deviation correction The specific expression is: in, represents the horizontal deviation of the j-th sub-region; The initial relative displacement after the deviation correction is corrected using the linear interpolation method to obtain the sub-pixel relative displacement of the target structure.

6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the semi-dense sub-pixel displacement measurement method based on LoFTR according to any one of claims 1 to 5 are implemented.

7. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the semi-dense sub-pixel displacement measurement method based on LoFTR according to any one of claims 1 to 5 is executed.