Structural unmarked three-way displacement real-time measurement method based on deep learning stereoscopic vision

By applying deep learning and stereo vision technology in structural deformation measurement, markless three-dimensional displacement measurement is achieved, solving the problem of manually selecting areas and difficulty in obtaining three-dimensional deformation in the prior art, and achieving high-precision and automated structural monitoring.

CN120043448APending Publication Date: 2025-05-27SHANGHAI CHOYOIN CONSTR GRP CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510167359.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-15
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing structural deformation measurement methods mainly rely on manual selection of areas of interest, and cannot realize automated monitoring, and it is difficult to directly obtain the three-dimensional spatial deformation of the structure.

Method used

The stereoscopic vision technology based on deep learning is adopted, combining semantic segmentation and SuperPoint feature point detection network to realize label-free three-dimensional displacement measurement. Through the synchronous acquisition of binocular cameras and the application of deep learning technology, the area to be monitored by the structure is automatically selected and three-dimensional reconstruction is carried out to realize real-time measurement of the multi-measuring point three-way displacement of the structure.

Benefits of technology

It realizes automatic selection and adjustment of monitoring areas without manual intervention, significantly improving the intelligence and automation level of the monitoring process, and can perform high-precision three-dimensional displacement measurements, which are suitable for complex and dynamically changing monitoring scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120043448A_ABST
    Figure CN120043448A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of structure health monitoring, and provides a structure unmarked three-way displacement real-time measurement method based on deep learning stereoscopic vision, which comprises the following steps: erecting and calibrating a binocular camera; obtaining a structure monitoring image sequence, and automatically extracting a to-be-monitored region of the structure through the semantic segmentation model; tracking coordinates of feature points of the to-be-monitored region of the structure through a deep learning feature point detection model; obtaining feature point coordinates of sub-pixel precision through sub-pixel refinement; three-dimensional coordinates of the feature points at all moments are obtained through stereo matching and three-dimensional reconstruction; and obtaining three-direction displacement information of the structure according to the three-dimensional coordinates and the three-dimensional point displacement. According to the invention, three-dimensional displacement measurement of the structure can be carried out without manually selecting a to-be-monitored area of the structure and designing a mark, the cost is low, the precision is high, and the engineering practicability is high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of structural health monitoring, and in particular to a method for real-time measurement of three-dimensional displacement without markers of a structure based on deep learning stereo vision. Background Art

[0002] In the field of civil engineering, the structural deformation measurement technology based on stereo vision and deep learning processes image data and performs three-dimensional reconstruction on feature points on the structure in the image to obtain the actual deformation information of the structure.

[0003] In recent years, computer vision-based technologies have developed rapidly and have been gradually applied to the field of civil engineering. With the help of image acquisition devices and digital image processing technologies, computer vision-based structural displacement monitoring methods and systems have been developed. The computer vision-based civil engineering structure displacement measurement method has been gradually adopted due to advantages such as convenient installation, low system cost, and high measurement resolution.

[0004] In 2003, Wahbeh AM et al. obtained the displacement response of a bridge in "A vision-based approach for the direct measurement of displacements in vibrating systems" by installing red LED light sources on the bridge and using a color filtering method to track the light sources; in 2014, Busca et al. used template matching, edge detection, and digital image correlation to measure the response of a bridge when a train passed by in "Vibration Monitoring of Multiple Bridge Points by Means of a Unique Vision-Based Measuring System" and compared the results with displacement sensors; in 2015, Feng et al. developed a vision sensing system based on template matching for two-dimensional displacement measurement of civil structures in "A Vision-Based Sensor for Noncontact Structural Displacement Measurement" and tested high-contrast artificial targets and natural feature points; in 2016, Yoon et al. used an optical flow-based method and photos taken by a drone to track the natural features of a structure to obtain displacements in "Target-free approach for vision-based structural system identification using consumer-grade cameras".

[0005] However, previous methods mainly relied on prior knowledge to select regions of interest, and were not applicable to automated structural monitoring processing. In addition, previous monitoring mainly used a single camera for two-dimensional deformation monitoring, and could not directly obtain the three-dimensional spatial deformation of the structure under actual service conditions. Summary of the Invention

[0006] In order to automatically select the regions of the structure to be monitored and solve the limitations of manual marking in practical engineering applications, the present invention provides a method for real-time measurement of three-way displacement without markers of a structure based on deep learning stereo vision. By introducing semantic segmentation technology and the SuperPoint deep learning feature point detection network into structural deformation measurement, through synchronous acquisition by two cameras, and combining deep learning technology with stereo vision technology, it is possible to achieve real-time measurement of three-way displacement at multiple measurement points of the structure with low cost and high precision. Without the need to manually select the regions of the structure to be monitored and design markers, three-dimensional displacement measurement of the structure can be carried out, and the engineering practicability is relatively strong.

[0007] Technical solution of the present invention: A method for real-time measurement of three-way displacement without markers of a structure based on deep learning stereo vision, comprising the following steps: Step 1: Set up and calibrate a binocular camera; Step 2: Obtain a sequence of structural monitoring images, and automatically extract the regions of the structure to be monitored through a semantic segmentation model; Step 3: Track the coordinates of the feature points in the regions of the structure to be monitored through a deep learning feature point detection model; Step 4: Obtain the coordinates of the feature points with sub-pixel accuracy through sub-pixel refinement; Step 5: Obtain the three-dimensional coordinates of the feature points at each moment through stereo matching and three-dimensional reconstruction; Step 6: Obtain the three-way displacement information of the structure according to the three-dimensional coordinates and three-dimensional point displacements.

[0008] Compared with the prior art, the present invention has the following beneficial effects: (1) Intelligently and automatically select the regions of the structure to be monitored: This method uses a semantic segmentation model based on the K-net network to realize the automatic division of the pixel range of the regions of the structure to be monitored; by calculating the moving average distance of the feature pixels between adjacent images, the monitoring region range is dynamically updated; compared with the prior art, the present invention can automatically select and adjust the monitoring region without manual intervention, greatly improving the intelligence and automation level of the monitoring process, and is particularly suitable for complex and dynamically changing monitoring scenarios; (2) Unmarked natural feature point detection based on deep learning: This method uses an improved SuperPoint network to achieve automatic detection of natural feature points on the structure surface, completely eliminating the need for manual marking. Compared with existing methods, the present invention significantly reduces the complexity of monitoring preparation work, greatly improves the applicability and practical application value in various environments, and is especially suitable for structures and scenarios where manual marking is difficult. (3) High-precision sub-pixel displacement measurement: This method applies a technique based on quadratic surface fitting to the coordinates of feature points to achieve sub-pixel precision positioning of feature points, thereby performing high-precision three-way displacement measurement of feature points on the structure surface. Compared with existing technologies, the present invention significantly improves the measurement accuracy of the three-way displacement of the structure, can capture more subtle structural deformations, and provides more reliable and accurate data support for structural health monitoring. (4) Innovative integration of deep learning and stereo vision: This method innovatively combines deep learning technology with stereo vision technology to achieve high-precision three-dimensional displacement measurement. Compared with traditional methods, the present invention not only improves the measurement accuracy and robustness, but also can maintain efficient and stable performance in complex environments, providing a low-cost and high-precision comprehensive solution for structural monitoring. Brief Description of the Drawings

[0009] Figure 1 It is a schematic diagram of the scenario of the embodiment of the present invention. Figure 2 It is a structural diagram of the SuperPoint network improved using depthwise separable convolution.

[0010] Figure 3 It is a schematic flow diagram of a method for real-time measurement of unmarked three-way displacement of a structure based on deep learning stereo vision of the present invention. Detailed Embodiment

[0011] The technical solution provided by the present application will be further described below in conjunction with specific embodiments and their accompanying drawings. In combination with the following description, the advantages and features of the present application will be more clearly understood.

[0012] A method for real-time measurement of unmarked three-way displacement of a structure based on deep learning stereo vision includes the following steps: Step 1: Set up and calibrate the binocular camera; First, select a position that can capture the area to be monitored of the structure, set up the binocular camera so that the optical axes of the two cameras form a certain angle, use a tripod to keep the two cameras stable, adjust parameters such as the focal length of the binocular camera so that the area to be monitored of the structure can appear simultaneously and clearly in the fields of view of the two cameras, and finally use a black and white checkerboard pattern to perform stereo calibration on the binocular camera to obtain the parameters of the binocular camera. Step 2: Obtain the structural monitoring image sequence and automatically extract the area to be monitored of the structure through the semantic segmentation model; To solve the problems of low extraction efficiency and insufficient accuracy of traditional extraction methods in complex scenarios, a modular design concept is adopted. Based on the K-Net framework, a hierarchical semantic segmentation network, namely the K-Net deep learning semantic segmentation model, is constructed. This network includes three main modules: feature extraction, feature fusion, and output decoding, which are used for high-precision segmentation of complex scenarios in images; Step 3: Track the coordinates of the feature points in the area to be monitored of the structure through the deep learning feature point detection model; First, use the SuperPoint network improved by depthwise separable convolution to detect feature points in the left and right initial images, and then use the LightGlue network to perform stereo matching within the respective image sequences of the left and right cameras, calculate the pixel displacement of the feature points between adjacent sequences, and update the pixel range of the area to be monitored of the structure to achieve automatic adjustment of the monitoring area; Step 4: Obtain the feature point coordinates with sub-pixel accuracy through sub-pixel refinement; Use the sub-pixel refinement technology for all detected feature points to obtain the feature point coordinates with sub-pixel accuracy; Step 5: Obtain the three-dimensional coordinates of the feature points at each moment through stereo matching and three-dimensional reconstruction; First, perform stereo matching on the feature points in the initial image pair of the left and right cameras, eliminate the mismatched points, then calculate the disparity of the matching point pairs according to the matching results, reconstruct the three-dimensional coordinates of the feature points through the disparity principle, and finally perform stereo matching and three-dimensional reconstruction on the image sequences of the left and right cameras to obtain the three-dimensional coordinates of the feature points at each moment; Step 6: Obtain the three-way displacement information of the structure; According to the three-dimensional coordinates and three-dimensional point displacements, obtain the three-way displacement information of the structure.

[0013] Furthermore, the specific content of Step 1 is as follows: Step 1.1: According to the type and focal length requirements of the selected camera, adjust parameters such as the focal length, exposure time, and aperture of the camera to ensure that the binocular camera can obtain clear and stable images under different lighting conditions; Step 1.2: Use the camera calibration method (such as Zhang Zhengyou camera calibration method) to obtain the internal and external parameters of the binocular camera, including the internal parameter matrices of the left and right cameras respectively and , and the external parameter matrix is the spatial transformation matrix for converting the right camera coordinate system to the left camera coordinate system . The internal and external parameters are represented by the following formula: where and For the focal lengths of the image in the axis and the axis directions, and are the pixel coordinates of the principal point of the image., are respectively the axis of the right camera coordinate system and the axis of the left camera coordinate system, the cosine value of the included angle; are respectively the axis of the right camera coordinate system and the axis of the left camera coordinate system, the cosine value of the included angle; are respectively the axis of the right camera coordinate system and the axis of the left camera coordinate system, the cosine value of the included angle; is the rotation matrix, is the translation vector, are the components of the translation vector in the world coordinate system along the axis, axis, axis; The world coordinate system takes the direction from the left camera to the right as the positive direction of the axis, takes the vertically downward direction as the positive direction of the axis, and takes the direction perpendicular to the surface of the structure and inward as the positive direction of the axis, represented by ; The left and right camera coordinate systems are respectively represented by and ; The left and right image coordinate systems are respectively represented by and ; Step 1.3. Make the left camera coordinate system coincide with the world coordinate system. The three-dimensional coordinates of the feature point in the world coordinate system are calculated by the following formula: Among them, is the average focal length; are the components of the translation vector in the world coordinate system along the axis, axis, axis; is the baseline distance, representing the distance between the axes of the left and right cameras; is the parallax, representing the absolute value of the difference in the coordinates of the corresponding points of the spatial point in the left and right image coordinate systems; is the coordinate value of the feature point in the world coordinate system; is the coordinate value of the feature point in the left camera coordinate system; is the coordinate value of the feature point in the right camera coordinate system; Furthermore, the specific content of the said Step 2 is: Step 2.1: Semantically annotate the area to be monitored for the structure and generate a certain number of annotated datasets; As an example, during annotation, use the AI-assisted annotation in the Labelme software to ensure the integrity and accuracy of the annotated area; Step 2.2: Adopt the modular design concept and, based on the K-Net framework, construct a hierarchical semantic segmentation network, namely the K-Net deep learning semantic segmentation model; This network includes three main modules: feature extraction, feature fusion, and output decoding, which are used for high-precision segmentation of complex scenes in images; Step 2.3: Adopt a transfer learning mechanism to train the K-Net deep learning semantic segmentation model. Specifically: First, perform pre-training on a general dataset so that the K-Net deep learning semantic segmentation model obtains a good representation of general features; Then freeze the general feature extraction layer of the lower layer of the K-Net deep learning semantic segmentation model and only fine-tune the weights of the upper layer, and use data augmentation techniques to expand the dataset on the annotated dataset; This mechanism can achieve efficient learning and accurate prediction of new tasks under extremely few data conditions; Step 2.4: After extracting the area to be monitored for the structure through the K-Net deep learning semantic segmentation model, adopt a post-processing algorithm based on graph optimization to ensure the smoothness and continuity of the output area boundary; The specific steps of this algorithm are as follows: First, represent the boundary of the segmented area to be monitored for the structure as an undirected graph ; Among them, is the vertex set, representing the pixel points on the boundary; is the edge set, indicating the connection between adjacent pixel points; Each vertex in the vertex set is associated with a position label , representing the new position of this point after optimization; Then, define the global energy function, and the expression is as follows: Among them, is the total energy function, which is used to evaluate the quality of the current label assignment ; is the data term, which measures the degree to which the position label of each vertex deviates from the original position, and is used to ensure that the optimized position is still consistent with the original image data; is the weight factor, which is used to balance the influence between the data term and the smooth term. A higher value will pay more attention to the smooth effect, while a lower value will pay more attention to data fidelity; is the smooth term, which measures the adjacent vertices The change of the tags between promotes the continuity and consistency between adjacent pixel points, thus obtaining a smoother boundary; Finally, the α-expansion graph cut algorithm is used to minimize the energy function, obtaining a boundary contour considering global information, significantly improving the geometric accuracy of the structure to be monitored area; Furthermore, the specific steps of step 3 are as follows: Step 3.1: Replace the ordinary convolution in the shared encoder of the SuperPoint network with depthwise separable convolution; Step 3.2: Use image processing technology to generate a training dataset containing lines, polylines, and cubes by specifying the coordinates of feature points, and then use data augmentation technologies such as optical processing and homography transformation to obtain feature points under different lighting conditions and different perspectives; Step 3.3: Train the improved SuperPoint network. When calculating the loss, use the original image and the distorted image transformed by the homography matrix as a pair to calculate the loss simultaneously. The expression of the loss function is as follows: Among them, respectively represent the output features of the feature point detection end, the output features of the descriptor calculation end, and the feature point label value of the original image; respectively represent the output features of the feature point detection end, the output features of the descriptor calculation end, and the feature point label value of the distorted image; represents the corresponding relationship of all points before and after the transformation of the image pair; represents the feature point position loss function, using full convolutional cross-entropy loss, and the expression is as follows: Among them, represents the image size after being scaled by the shared encoder, which is 1 / 8 of the original image size; represents the feature point descriptor loss function, and the expression is as follows: Among them, represents the original image at and the distorted image at the descriptor vectors; is an indicator variable indicating whether this pair of descriptors matches; respectively represent the coordinates after applying the homography transformation to the original image points and the coordinates in the distorted image; is a weight parameter used to balance the matching and non-matching losses; and are the thresholds for positive and negative samples respectively; Step 3.4: In the initial images of the left and right cameras, apply the SuperPoint network improved by depthwise separable convolution to the area of the structure to be monitored for feature point extraction; Step 3.5: Use the deep learning stereo matching model based on the LightGlue network for stereo matching within the respective image sequences of the left and right cameras to obtain the coordinate change values of the same feature points between adjacent images, and take the average value of all the feature point coordinate change values as the moving distances of the pixel range of the structure to be monitored in the and directions respectively, and the moving distances in the and directions are calculated by the following formula: where, is the number of feature points, is the image serial number; Furthermore, the specific steps of step 4 are as follows: Step 4.1: For each detected feature point, use the quadratic surface fitting method to establish a neighborhood centered on this feature point, and use the gradient information of this neighborhood to calculate the parameters of the quadratic surface by the least squares method. The quadratic surface is calculated by the following formula: where, represents the gray value at the point in the image, represents the coefficient to be determined; Step 4.2: Establish the fitted quadratic surface and obtain the sub-pixel accurate coordinates of this feature point by the method of finding the extreme value.

[0014] Furthermore, the specific steps of step 5 are as follows: Step 5.1: Use the deep learning stereo matching model to match the feature points in the initial images of the left and right cameras and eliminate the mis-matched points, and reconstruct the initial three-dimensional coordinates of the feature points according to the matching results; Step 5.2: Perform stereo matching on the left and right image sequences, and reconstruct the three-dimensional coordinates of the feature points at any moment according to the matching results; Step 5.3: Obtain the three-dimensional coordinates of the feature points at each moment; Further, step 6 is specifically as follows: Taking the three-dimensional coordinates of each feature point at the initial moment as the reference coordinates, subtracting the reference coordinates from the three-dimensional coordinates of each feature point at each moment to obtain the three-dimensional displacements of each feature point at each moment.

[0015] The above description is only a description of the preferred embodiments of the present application, and is not any limitation on the scope of the present application. Any change or modification made by any person skilled in the art according to the technical content disclosed above shall be regarded as an equivalent effective embodiment, and all fall within the scope of protection of the technical solution of the present application.

Claims

1. A method for real-time measurement of structural markerless three-dimensional displacement based on deep learning stereo vision, characterized in that: The following steps are involved: Step 1: Set up and calibrate the binocular camera; Step 2: Obtain a sequence of structure monitoring images and automatically extract the structure to be monitored area through a semantic segmentation model; Step 3: Track the coordinates of the feature points in the structure to be monitored through the deep learning feature point detection model; Step 4: Obtain feature point coordinates with sub-pixel accuracy through sub-pixel refinement; Step 5: Obtain the three-dimensional coordinates of the feature points at each moment through stereo matching and three-dimensional reconstruction; Step 6: According to the three-dimensional coordinates and three-dimensional point displacement, the three-dimensional displacement information of the structure is obtained.

2. The method for real-time measurement of three-dimensional displacement without marking based on deep learning stereo vision according to claim 1, characterized in that: The step 1 is specifically as follows: Step 1.1, select a position where the structure to be monitored can be photographed, so that the optical axes of the two cameras form a certain angle to set up the binocular cameras, use a tripod to keep the two cameras stable, and adjust the camera's focal length, exposure time, and aperture parameters according to the type and focal length requirements of the selected camera to ensure that the binocular camera can obtain clear and stable images under different lighting conditions, so that the structure to be monitored can appear clearly in the field of view of the two cameras at the same time; Step 1.2: Use Zhang Zhengyou's camera calibration method to obtain the internal and external parameters of the binocular camera, including the internal parameter matrix of the left and right cameras. and , the external parameter matrix is ​​the spatial transformation matrix from the right camera coordinate system to the left camera coordinate system , the internal and external parameters are expressed by the following formula: in, and For the image Axis and The focal length along the axis, and is the pixel coordinate of the principal point of the image, They are respectively the right camera coordinate system The axes are respectively related to the left camera coordinate system The cosine of the angle between the axes, They are respectively the right camera coordinate system The axes are respectively related to the left camera coordinate system The cosine of the angle between the axes, They are respectively the right camera coordinate system The axes are respectively related to the left camera coordinate system The cosine of the angle between the axes, is the rotation matrix, is the translation vector, are the translation vectors in the world coordinate system axis, axis, The weight of the axis; The world coordinate system is from the left camera to the right. The positive direction of the axis is vertically downward. The positive direction of the axis is perpendicular to the structural surface and inward. Axis positive direction, use Represented by; the left and right camera coordinate systems are respectively and The left and right image coordinate systems are represented by and express; Step 1.3: Let the left camera coordinate system With the world coordinate system The three-dimensional coordinates of the feature points in the world coordinate system are calculated by the following formula: in, is the average focal length; are the translation vectors in the world coordinate system axis, axis, The weight of the axis; is the baseline distance, which represents the distance between the axes of the coordinates of the left and right cameras; is the disparity, representing the corresponding point of the spatial point in the left and right image coordinate systems The absolute value of the difference between the coordinates; is the coordinate value of the feature point in the world coordinate system; is the coordinate value of the feature point in the left camera coordinate system; is the coordinate value of the feature point in the right camera coordinate system.

3. The method for real-time measurement of three-dimensional displacement without marking based on deep learning stereo vision according to claim 1, characterized in that: The step 2 is specifically as follows: Step 2.1: semantically annotate the structure area to be monitored and generate a certain number of annotated data sets. When annotating, use the AI-assisted annotation in the Labelme software to ensure the integrity and accuracy of the annotated area. Step 2.2: Adopting modular design concept and based on K-Net framework, a hierarchical semantic segmentation network is constructed, namely K-Net deep learning semantic segmentation model; the network includes three main modules: feature extraction, feature fusion and output decoding, which is used for high-precision segmentation of complex scenes in images; Step 2.3: Use the transfer learning mechanism to train the K-Net deep learning semantic segmentation model. Specifically, pre-train on a general dataset to enable the K-Net deep learning semantic segmentation model to obtain a good representation of general features. Then, the low-level general feature extraction layer of the K-Net deep learning semantic segmentation model is frozen, and only the weights of the high-level layers are fine-tuned. The data set is expanded using data augmentation technology on the labeled data set. This mechanism can achieve efficient learning and accurate prediction of new tasks with very little data. Step 2.4: After extracting the structure area to be monitored, a post-processing algorithm based on graph optimization is used to ensure the smoothness and continuity of the output area boundary.

4. The method for real-time measurement of structural marker-free three-dimensional displacement based on deep learning stereo vision according to claim 3, characterized in that: The specific steps of the post-processing algorithm based on graph optimization in step 2.4 are as follows: First, the boundary of the segmented structure to be monitored is represented as an undirected graph ;in is a set of vertices, representing pixels on the boundary. is an edge set, indicating the connection between adjacent pixels; each vertex Associated with a location tag , represents the new position of the point after optimization; Then, define the global energy function, the expression is as follows: in, is the total energy function used to evaluate the current label assignment quality; is a data item that measures each vertex Location Tags The degree of deviation from the original position is used to ensure that the optimized position remains consistent with the original image data; is a weight factor used to balance the impact between the data term and the smoothing term. Values ​​place more emphasis on smoothing, while lower The value will pay more attention to data fidelity; is the smoothness term, which measures the relationship between adjacent vertices The change of labels between pixels promotes the continuity and consistency between adjacent pixels, thus obtaining a smoother boundary. Finally, the α-expansion graph cut algorithm is used to minimize the energy function and obtain the boundary contour that takes global information into account, which significantly improves the geometric accuracy of the structure to be monitored.

5. The method for real-time measurement of three-dimensional displacement without marking based on deep learning stereo vision according to claim 1, characterized in that: In step 3, the SuperPoint network improved by deep separable convolution is first used to detect feature points of the left and right initial images, and then the LightGlue network is used to perform stereo matching in the image sequences of the left and right cameras respectively, and the pixel displacement of the feature points is calculated between adjacent sequences, and the pixel range of the structure to be monitored is updated to realize automatic adjustment of the monitoring area; specifically, Step 3.1, replace the ordinary convolution in the shared encoder of the SuperPoint network with depth-wise separable convolution; Step 3.2: Use image processing technology to generate a training data set containing straight lines, polylines, and cubes by specifying the coordinates of feature points. Then use data enhancement techniques such as optical processing and homography transformation to obtain feature points under different lighting conditions and different viewing angles. Step 3.3: Train the improved SuperPoint network. When calculating the loss, use the original image and the distorted image transformed by the homography matrix as a pair to calculate the loss simultaneously. The expression of the loss function is as follows: in, They represent the output features of the feature point detection end, the output features of the descriptor calculation end, and the feature point label value of the original image respectively; They represent the output features of the feature point detection end, the output features of the descriptor calculation end, and the feature point label value of the distorted image respectively; Represents the correspondence between all points before and after the image transformation; Represents the feature point position loss function, using full convolution cross entropy loss, the expression is as follows: in, Represents the image size after being scaled by the shared encoder, which is 1 / 8 of the original image size; Represents the feature point descriptor loss function, and the expression is as follows: in, Represents the original image Distorting and distorting images The descriptor vector at ; is an indicator variable indicating whether the pair of descriptors matches; Represent the coordinates after applying the homography transformation to the original image points and the coordinates in the distorted image, respectively; is a weight parameter used to balance matching and non-matching losses; and are the thresholds for positive and negative samples respectively; Step 3.4: In the initial images of the left and right cameras, the SuperPoint network improved by deep separable convolution is applied to the structure to be monitored area to extract feature points; Step 3.5: Use the deep learning stereo matching model based on the LightGlue network to perform stereo matching in the image sequences of the left and right cameras, obtain the coordinate change value of the same feature point between adjacent images, and use the average value of the coordinate change value of all feature points as the pixel range of the structure to be monitored. and The distance moved in the direction, and Distance moved in a direction and Calculated by the following formula: in, is the number of feature points, Is the image sequence number.

6. The method for real-time measurement of three-dimensional displacement without marking based on deep learning stereo vision according to claim 1, characterized in that: The step 4 uses sub-pixel refinement technology on all detected feature points to obtain feature point coordinates with sub-pixel accuracy; specifically: Step 4.1: For each detected feature point, use the quadratic surface fitting method to establish a The neighborhood of , using the gradient information of the neighborhood to calculate the parameters of the quadratic surface by the least squares method, the quadratic surface is calculated by the following formula: in, Indicates the midpoint of the image The gray value at represents the coefficient to be determined; Step 4.2: Establish the fitted quadratic surface and obtain the sub-pixel precision coordinates of the feature point by the method of finding the extreme value.

7. The method for real-time measurement of three-dimensional displacement without marking based on deep learning stereo vision according to claim 1, characterized in that: The step 5 is specifically as follows: Step 5.1, use the deep learning stereo matching model to match the feature points of the left and right camera images at the initial moment and remove the mismatched points, and reconstruct the initial three-dimensional coordinates of the feature points based on the matching results; Step 5.2, stereo matching is performed on the left and right image sequences, and the three-dimensional coordinates of the feature points at any time are reconstructed according to the matching results; Step 5.3: Obtain the three-dimensional coordinates of the feature points at each moment.

8. The method for real-time measurement of three-dimensional displacement without marking based on deep learning stereo vision according to claim 1, characterized in that: The step 6 specifically includes: taking the three-dimensional coordinates of each feature point at the initial moment as the reference coordinates, and subtracting the three-dimensional coordinates of each feature point at each moment from the reference coordinates to obtain the three-dimensional displacement of the feature point at each moment.

Citation Information

Cited By

  • Structural deformation monitoring method and system based on secondary stereoscopic vision calibration

    CN120807441A

  • Mobile fire detection system and method

    CN121214628A

  • A mobile fire detection system and method

    CN121214628B