Building settlement monitoring method and system based on multi-datum point affine correction
By employing a multi-reference point affine correction method, combined with deep learning and digital image correlation algorithms, errors introduced by the environment and equipment are eliminated, achieving high-precision building settlement monitoring. This solves the problems of insufficient monitoring complexity and accuracy in existing technologies, and improves monitoring efficiency and reliability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- EAST CHINA JIAOTONG UNIVERSITY
- Filing Date
- 2026-03-19
- Publication Date
- 2026-04-17
AI Technical Summary
Existing methods for monitoring building settlement are complex to operate, require a high degree of human intervention, have low monitoring efficiency, and are susceptible to human factors. Furthermore, methods based on a single or limited reference point have limited accuracy when subjected to external environmental disturbances and overall structural deformation. Deep learning methods still have room for improvement in terms of multi-field coupling and multi-reference point constraints.
A multi-reference point affine correction method is adopted. By acquiring a target group containing an initial reference image and real-time monitoring images, the sub-pixel displacement of the target is extracted using a deep learning model and digital image correlation algorithm. Combined with affine transformation moment calculation, the error introduced by environmental vibration and equipment pose change is eliminated, thus achieving high-precision settlement monitoring.
It improves the signal-to-noise ratio and reliability of monitoring data, reduces systematic errors caused by human intervention, and enables efficient assessment of large-scale building deformation.
Smart Images

Figure CN121876907A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of engineering precision measurement and computer vision application technology, specifically to a method and system for monitoring building settlement based on multi-reference point affine correction. Background Technology
[0002] Building settlement monitoring is an important research area in geotechnical and structural engineering, and it is of great significance for ensuring the safety of building structures and extending their service life. Existing settlement monitoring methods mainly include leveling, total station measurement, and the installation of deformation sensors. Although these traditional methods have good measurement accuracy, they suffer from drawbacks in practical applications, such as complex operation procedures, high degree of manual intervention, low monitoring efficiency, long monitoring cycles, and susceptibility to human factors. These methods cannot meet the demands of modern engineering for automated, real-time, and multi-point synchronous monitoring.
[0003] In existing technologies, some image-based settlement monitoring methods typically use a single reference point or a limited number of reference points for calibration, resulting in limited accuracy when dealing with external environmental disturbances and overall structural deformation. Furthermore, while digital image correlation methods have achieved sub-pixel level displacement and deformation measurements, their ability to correct for rigid body motion disturbances and noise is insufficient. In recent years, machine learning methods such as deep learning have been increasingly applied to target detection and recognition, improving the intelligence level of target localization. However, there is still room for improvement in multi-field coupling and multi-reference point constraints, which somewhat affects the overall monitoring accuracy and robustness.
[0004] Therefore, developing a building settlement monitoring method and system that combines multi-reference point constraints, deep learning target detection, and digital image correlation methods is of practical significance and application value for achieving high-precision, automated, and interference-resistant settlement monitoring of building structures. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention provides a building settlement monitoring method and system based on multi-reference point affine correction, which can effectively improve the signal-to-noise ratio and reliability of monitoring data, reduce systematic errors caused by human intervention, has excellent engineering applicability, and can achieve efficient assessment of large-scale building deformation.
[0006] The first aspect of the present invention is to provide a method for monitoring building settlement based on multi-reference point affine correction, the method comprising: Acquire a monitoring image set including an initial reference image and real-time monitoring images. The monitoring area is equipped with a static reference target group located in a non-settlement area and a dynamic monitoring target group located on the surface of the building. The monitoring image set is detected to obtain the positioning areas of the static reference target group and the dynamic monitoring target group, respectively; Based on the positioning area, a reference template is determined and matching is performed to determine the initial pixel coordinates of the static reference target group and the dynamic monitoring target group in the initial reference image, and their sub-pixel level real-time pixel coordinates in the real-time monitoring image are tracked and obtained. Based on the initial pixel coordinates of the static reference target group and the corresponding sub-pixel level real-time pixel coordinates, fit an affine transformation matrix; The sub-pixel level real-time pixel coordinates of the dynamic monitoring target group are corrected by inverse transformation using the affine transformation matrix. The vector difference between the corrected pixel coordinates and the initial pixel coordinates is calculated to obtain the sub-pixel level corrected pixel displacement. Based on the determined pixel-to-physical size conversion coefficient, the sub-pixel level corrected pixel displacement is converted into the physical settlement of the building.
[0007] According to one aspect of the above technical solution, the steps of determining a reference template based on the positioning region to perform matching, determining the initial pixel coordinates of the static reference target group and the dynamic monitoring target group in the initial reference image, and tracking and obtaining their sub-pixel level real-time pixel coordinates in the real-time monitoring image include: In the initial reference image, the center coordinates of each target localization region output by the deep learning model are determined as the initial pixel coordinates of the target. Centered on the initial pixel coordinates, an image block is cropped according to a preset template size and used as a reference template area for subsequent tracking; In the real-time monitoring image, a search area is defined with the initial pixel coordinates as the center. The ZNCC algorithm is used to calculate the correlation coefficient between the reference template area and each sub-region of the real-time image within the search area, and the position with the largest correlation coefficient value is determined as the real-time coordinate of the target at the integer pixel level. Within the neighborhood of the real-time coordinates at the integer pixel level, a discrete correlation coefficient matrix consisting of multiple ZNCC correlation coefficient values is constructed; The discrete correlation coefficient matrix is fitted using a quadratic surface fitting method, and the coordinates of the mathematical extremum points of the fitted quadratic surface are used as the sub-pixel level real-time pixel coordinates of the target.
[0008] Using the discrete pixel coordinates within the neighborhood of the real-time integer-level coordinates as independent variables and the corresponding ZNCC correlation coefficients as dependent variables, a quadratic surface mathematical model is constructed:
[0009] In the formula, Indicates continuous local coordinates The fitted correlation coefficient value at the location, , , , , , The coefficients of the quadratic surface equation are to be determined by regression analysis of the discrete matrix data using the least squares method. The least squares method is used to perform regression analysis on the discrete correlation coefficient matrix to determine the undetermined coefficients of the quadratic surface equation. The equation of the quadratic surface is given by the variables respectively. and Find the first-order partial derivatives, set them to zero, and solve the system of equations to obtain the coordinates of the extreme points of the quadratic surface. , () is used as the sub-pixel level real-time pixel coordinates of the target.
[0010] According to one aspect of the above technical solution, the coordinates of the extreme points of the quadratic surface ( , The formula for solving ) is as follows:
[0011] In the formula, The x-coordinate of the target's sub-pixel-level real-time pixel coordinates. The ordinate is the sub-pixel level real-time pixel coordinate of the target. , , , , Let x, y, and x' be the variables in the equation of the quadratic surface, respectively. 2 variables x and y 2 The polynomial coefficients.
[0012] According to one aspect of the above technical solution, the step of fitting an affine transformation matrix based on the initial pixel coordinates of the static reference target group and the corresponding sub-pixel level real-time pixel coordinates includes: For any static reference target in the static reference target group, its initial pixel coordinates (x, y) and corresponding sub-pixel level real-time pixel coordinates (x, y) are... , The following affine transformation mapping relationship exists between them:
[0013]
[0014] Among them, a 11 a 12 Let a be the coefficients of variables x and y in the abscissa mapping equation. 21 a 22Let be the coefficients of variables x and y in the ordinate mapping equation, respectively, and t be the coefficients of these variables. x t y The constant terms of the horizontal and vertical coordinate mapping equations, respectively, together serve as six undetermined parameters characterizing the affine transformation; By substituting the initial pixel coordinates and subpixel-level real-time pixel coordinates of each target in the static benchmark target group into the mapping relationship, an overdetermined linear equation system for solving the six undetermined parameters is constructed. The least squares method is used to solve the overdetermined linear equation system to determine the optimal solution for the six undetermined parameters; The determined optimal solution parameters are combined to form an affine transformation matrix M in the following form:
[0015] Among them, the upper left corner of the matrix submatrix Linear rotation and non-uniform scaling transformation operators in a two-dimensional plane are defined, and the third column vector... The translation transformation operator for the coordinate system is defined in the third line. To maintain the homogeneous extension term of the dimensions of the projected geometry calculation.
[0016] According to one aspect of the above technical solution, the step of using the affine transformation matrix to perform inverse transformation correction on the sub-pixel level real-time pixel coordinates of the dynamic monitoring target group, calculating the vector difference between the corrected pixel coordinates and the initial pixel coordinates, and obtaining the sub-pixel level corrected pixel displacement includes: Calculate the inverse matrix M⁻¹ of the affine transformation matrix; Using the inverse matrix, the sub-pixel level real-time pixel coordinates of each target in the dynamic monitoring target group are subjected to inverse coordinate transformation to obtain their corresponding sub-pixel level corrected pixel coordinates. Calculate the vector difference between the subpixel-level corrected pixel coordinates and the corresponding initial pixel coordinates of each dynamic monitoring target, and determine the vector difference as the subpixel-level corrected pixel displacement of the target.
[0017] According to one aspect of the above technical solution, the static reference target group includes at least three non-collinearly spaced static reference targets. Each static reference target is fixedly set on a stable structure that is not affected by building settlement and maintains its relative position unchanged during the monitoring process. The static reference targets are used to provide multi-reference constraints for solving the affine transformation matrix.
[0018] A second aspect of the present invention is to provide a building settlement monitoring system based on multi-reference point affine correction, using the method described in the above technical solution, wherein the system includes: The image acquisition module is used to acquire a set of monitoring images including an initial reference image and real-time monitoring images. The monitoring area is equipped with a static reference target group located in a non-settlement area and a dynamic monitoring target group located on the surface of the building. The target detection module is used to detect the monitoring image set and obtain the positioning areas of the static reference target group and the dynamic monitoring target group respectively; The digital image correlation processing module is used to determine a reference template based on the positioning area to perform matching, determine the initial pixel coordinates of the static reference target group and the dynamic monitoring target group in the initial reference image, and track and obtain their sub-pixel level real-time pixel coordinates in the real-time monitoring image. The affine transformation matrix fitting module is used to fit the affine transformation matrix based on the initial pixel coordinates of the static reference target group and the corresponding sub-pixel level real-time pixel coordinates. The affine correction and displacement calculation module is used to perform inverse transformation correction on the sub-pixel level real-time pixel coordinates of the dynamic monitoring target group using the affine transformation matrix, calculate the vector difference between the corrected pixel coordinates and the initial pixel coordinates, and obtain the sub-pixel level corrected pixel displacement. The settlement calculation module is used to convert the sub-pixel level corrected pixel displacement into the physical settlement of the building based on the determined pixel-to-physical size conversion coefficient.
[0019] A third aspect of the present invention is to provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the multi-reference point affine correction method for monitoring building settlement described in the above-mentioned technical solutions.
[0020] A fourth aspect of the present invention is to provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the multi-reference point affine correction method for monitoring building settlement described in the above technical solutions.
[0021] Compared with existing technologies, the building settlement monitoring method and system based on multi-reference point affine correction shown in this invention has the following advantages: This invention constructs a dual-target system encompassing both static and dynamic monitoring within the building monitoring area. It acquires time-series monitoring images using image acquisition equipment, applies deep learning models and digital image correlation algorithms to extract sub-pixel-level displacements of the targets, and performs multi-reference point affine transformation operations based on the static reference target to eliminate rigid body motion errors introduced by environmental vibrations and equipment pose changes. Finally, it maps the corrected pixel displacements to actual physical settlement by combining camera calibration parameters, achieving synchronous and accurate perception of settlement changes at multiple measurement points. This invention effectively improves the signal-to-noise ratio and reliability of monitoring data, reduces systematic errors caused by human intervention, has excellent engineering applicability, and enables efficient assessment of large-scale building deformation. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 This is a schematic diagram of the overall process of a building settlement monitoring method based on multi-reference point affine correction according to an embodiment of the present invention; Figure 2 This is a schematic diagram of a building settlement monitoring system based on multi-reference point affine correction according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of a terminal according to an embodiment of the present invention; Figure label: Image acquisition module-10, target detection module-20, digital image correlation processing module-30, affine transformation matrix fitting module-40, affine correction and displacement calculation module-50, settlement calculation module-60. Detailed Implementation
[0024] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0025] The terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.
[0026] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a mutually exclusive, independent, or alternative embodiment. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described in this application can be combined with other embodiments.
[0027] Example 1 Please see Figure 1 One embodiment of the present invention provides a building settlement monitoring method based on multi-reference point affine correction, the method comprising the following steps: Step S1: Obtain a monitoring image set including an initial reference image and a real-time monitoring image. The monitoring area is equipped with a static reference target group located in a non-settlement area and a dynamic monitoring target group located on the surface of the building.
[0028] The monitoring area of the building can be acquired in a time sequence using image acquisition equipment to obtain a monitoring image set including initial reference images and real-time monitoring images. Static reference target groups located in non-settlement areas and dynamic monitoring target groups located on the surface of the building are deployed within the monitoring area.
[0029] First, the optical imaging monitoring field of view is defined based on the spatial topology and geometric boundaries of the building to be monitored. Within this field of view, the spatial installation pose and optical axis pointing parameters of the image acquisition equipment are determined to ensure that the effective light-sensing area of the imaging sensor completely covers the entire space including both the static reference and the dynamic monitoring object, and meets the spatial resolution requirements for sub-pixel processing.
[0030] Static benchmark target groups are deployed in the foundation-stable zone or at non-deformation control points within the monitoring area to construct an absolute spatial reference system that remains stationary relative to the geodetic coordinate system. Simultaneously, dynamic monitoring target groups are deployed at the settlement nodes to be measured or key load-bearing components on the building surface, achieving a rigid connection between them and the structural surface, thereby mapping the physical displacement of the building into the spatial position coordinate changes of the targets in real time.
[0031] Specifically, the static reference target and the dynamic monitoring target maintain the same physical geometric dimensions and adhere to a unified imaging scale constraint. The target surface features include coded characters, numbers, or specific geometric patterns for unique indexing, constituting a unique identifier for feature matching and target retrieval by computer vision algorithms. Optically, the target body is made of a high-absorbency material or coated with a matte finish, giving its surface diffuse reflection optical properties and creating a high-contrast grayscale gradient distribution.
[0032] After target deployment and imaging optical path calibration are completed, the time-series image acquisition process is initiated to acquire a zero-time image of the building in its initial state, which is defined as the initial reference image. Subsequently, within the monitoring period, continuous image sampling operations are performed according to a preset discrete time step or an external trigger signal to generate time-series data containing multiple frames of real-time monitoring images.
[0033] It should be noted that the final monitoring image set consists of an initial reference image and a sequence of real-time monitoring images arranged by timestamps. Each frame of image data synchronously records the grayscale information of both the static reference target group and the dynamic monitoring target group in the pixel matrix, forming a homogeneous observation data set containing both spatial reference information and structural deformation information.
[0034] Step S2: Use a deep learning model to detect the monitoring image set and obtain the positioning areas of the static benchmark target group and the dynamic monitoring target group respectively.
[0035] In implementation, a predetermined number of sample images are first selected from the monitoring image set according to statistical sampling principles. The selected samples cover discrete time series nodes, differentiated light radiation intensity, multi-dimensional shooting perspective geometric parameters, and non-uniform target spatial distribution, constructing a multimodal sample set that includes scene brightness gradient, background texture complexity, multi-scale imaging features, and local occlusion factors.
[0036] For static baseline targets and dynamic monitoring targets in the sample images, a training dataset with supervised information is constructed using a combination of manual and semi-automatic annotation. This process aims to establish a mapping relationship between image planar pixel coordinates and target category attributes, generating structured metadata containing target bounding box coordinate parameters and category labels.
[0037] Specifically, the manual annotation process is performed using the labelimg annotation tool. In the image coordinate system, the operator draws a rectangular region of interest closely following the edge of the target and assigns it corresponding category attribute values (static baseline target or dynamic monitoring target). The system ultimately generates a txt format annotation file corresponding to the original image filename index. This file records the center point coordinates (x, y) and width and height dimensions (w, h) of each target in numerical form.
[0038] It should be noted that the semi-automatic annotation process incorporates the X-anylabelimg tool for auxiliary processing. This tool utilizes a candidate region generation network or interactive recommendation algorithm from a pre-trained model to output high-confidence candidate bounding boxes for potential target regions in the image. Operators then perform secondary verification based on these candidate results, manually correcting and completing the accuracy of the bounding box regression coordinates and the correctness of the classification labels, ultimately outputting the final labeled data that meets the IoU (Intersection over Union) threshold requirements.
[0039] After completing the sample annotation, the sample images and their corresponding annotation files are sequentially organized according to the data interface specifications of the deep learning framework to form a training dataset. This dataset consists of the original pixel matrix data, one-hot encoded class labels, and normalized bounding box regression parameters, serving as the input tensor for the supervised learning algorithm.
[0040] During the model training phase, the training dataset is input into the deep learning object detection network architecture to perform forward propagation and backward update. The model constructs a composite loss function based on labeled data, which includes classification loss and localization regression loss. The gradient is calculated using stochastic gradient descent (SGD) or Adam optimization algorithm, and the network weight parameters are iteratively updated. This enables the convolutional neural network to extract high-dimensional feature representations of static benchmark targets and dynamic monitoring targets in terms of grayscale texture, edge contours, and morphological features.
[0041] After the model training converges, the monitored image set is input into the deep learning model with fixed parameters to perform forward inference operations. The model performs multi-layer convolutional feature extraction and feature pyramid fusion on the input tensor, generates dense anchor boxes on the feature map, and selects the optimal prediction result through a non-maximum suppression algorithm.
[0042] Specifically, in the inference output stage, the model determines the category attribute (static baseline target or dynamic monitoring target) of the candidate region based on the classification confidence threshold and outputs the corrected bounding box coordinate data. This process transforms unstructured image pixel information into structured target location information, outputting precise pixel coordinate region data of the static baseline target group and the dynamic monitoring target group in the time-series image.
[0043] Step S3: Based on the positioning area, a reference template is determined, and a digital image correlation algorithm is used to perform correlation matching to determine the initial pixel coordinates of the static reference target group and the dynamic monitoring target group in the initial reference image, and their sub-pixel level real-time pixel coordinates in the real-time monitoring image are tracked and obtained.
[0044] During implementation, geometric center extraction is performed based on the target localization region data output by the deep learning model.
[0045] Specifically, for each target detection box in the initial reference image, the arithmetic mean of its corner coordinates is calculated to determine the geometric centroid coordinates of the localization region. These centroid coordinates are defined as the initial pixel coordinates of the corresponding target in the image coordinate system, serving as the spatial reference origin for subsequent temporal tracking algorithms.
[0046] Furthermore, using the initial pixel coordinates as anchor points, a region of interest extraction operation is performed in the initial reference image. A sub-image block containing the target is extracted according to a preset convolution kernel size or template window size. This sub-image block constitutes a reference template region for subsequent digital image correlation operations, storing the grayscale intensity matrix and texture gradient distribution information of the target and its neighboring background.
[0047] It should be noted that the coordinate search strategy for real-time monitoring images employs a local search window mechanism based on prior location. A rectangular search area is delineated in the real-time monitoring image, centered on the initial pixel coordinates in the initial reference image. The physical size parameters of this search area depend on the estimated maximum displacement vector magnitude of the building and the temporal resolution of image acquisition, thus limiting the traversal range of the relevant matching algorithm.
[0048] Within the search area, the zero-mean normalized cross-correlation (ZNCC) algorithm is used to calculate the correlation between the reference template region and candidate sub-regions in the real-time monitoring image. ZNCC correlation coefficient. The calculation formula is: (1) In the formula, This is the zero-mean normalized cross-correlation coefficient between the template image and the image to be matched at the reference position (x, y) of the matching window. This represents the grayscale value of the template image at pixel coordinates (u,v); For the template image in the relevant computational region The average grayscale value of all pixels within the range; To find the matching window within the image to be matched, using (x,y) as the reference position, in relative coordinates... The grayscale value at that location; In order to The mean grayscale value of all pixels within the matching window cropped for the reference position; The set of pixel coordinates involved in the relevant calculations; correlation coefficient ZNCC The range of values for ) (or R(x,y)) is The closer the value is to 1, the higher the degree of matching between the template and the window to be matched at position (x,y).
[0049] Specifically, the system iterates through every integer pixel coordinate within the search area. and calculate the corresponding The values are used to generate a correlation coefficient response surface in the discrete domain. By comparing the magnitudes of the values on the response surface, the correlation coefficient is extracted. The discrete coordinate position corresponding to the global maximum value is determined as the real-time coordinate of the target in the real-time monitoring image at the integer pixel level.
[0050] After obtaining the real-time coordinates at the integer pixel level, sub-pixel level interpolation operations are performed to overcome the physical resolution limitations of the imaging sensor.
[0051] Specifically, taking the pixel-level coordinate as the center, extract the 8 adjacent discrete positions in its 3×3 neighborhood, calculate the ZNCC correlation coefficient values at these positions, and construct a local discrete correlation coefficient matrix composed of discrete grid coordinates and their corresponding correlation coefficients.
[0052] In the sub-pixel solution stage, a quadratic surface fitting method is used to fit the local discrete correlation coefficient matrix. The discrete pixel coordinates within the neighborhood are taken as the independent variable, and the local coordinates are set as follows: , The corresponding correlation coefficient value is The mathematical model of the quadratic surface is then constructed as follows: (2) In the formula, C(x,y) represents the fitting correlation coefficient value at the continuous local coordinates (x,y), a1, a2, a3, a4, and a5 are the coefficients of variables x, y, x2, xy, and y2 in the quadratic surface equation, respectively, and f is the constant term of the quadratic surface equation. Together, they serve as the undetermined coefficients of the quadratic surface equation, which are obtained by regression solution of discrete matrix data using the least squares method.
[0053] To determine the theoretical extreme points of the fitted surface, the functions were analyzed respectively. Regarding variables and Find the first-order partial derivatives. The system of partial derivative equations is as follows: (3) In the formula, Let x be the first-order partial derivative of the continuous local surface function with respect to the variable x. Let y be the first-order partial derivative of the continuous local surface function with respect to the variable y, where x and y are the abscissa and ordinate of the continuous local coordinates, respectively, and a1, a2, a3, a4, and a5 are the polynomial coefficients of the quadratic surface equation.
[0054] It should be noted that by setting the partial derivatives to zero to locate the stationary points of the continuous surface, and by solving the above system of linear equations simultaneously, the coordinates of the sub-pixel matching points can be analytically obtained. , Its analytical solution takes the following form: (4) In the formula, x' is the abscissa of the extreme point of the quadratic surface, y' is the ordinate of the extreme point of the quadratic surface, and a1, a2, a3, a4, and a5 are the polynomial coefficients of the equation of the quadratic surface.
[0055] The obtained ( , The coordinates of the extreme point of the local correlation coefficient quadratic surface correspond to the position of the maximum value of the correlation coefficient in this neighborhood, in pixels. By using the position of this extreme point in the image coordinate system as the sub-pixel real-time pixel coordinates of the target in the current real-time monitoring image, sub-pixel level interpolation optimization and fine correction can be achieved based on the integer pixel level matching results.
[0056] It should be noted that, through the joint operation process of deep learning region localization, ZNCC statistical correlation analysis and quadratic surface extremum interpolation algorithm, a complete data processing flow from the initial coordinate extraction of static benchmark target group and dynamic monitoring target group to real-time sub-pixel coordinate calculation is realized. Finally, the output is a sub-pixel level accurate image coordinate sequence containing temporal information, which serves as the numerical input basis for subsequent spatial coordinate transformation and physical displacement calculation.
[0057] Step S4: Fit an affine transformation matrix based on the initial pixel coordinates of the static reference target group and the corresponding sub-pixel level real-time pixel coordinates.
[0058] During implementation, several targets that remain stable relative to the building structure throughout the entire monitoring period are selected to form a static benchmark target group.
[0059] Specifically, for any static reference target i in this set, its pixel coordinates in the initial reference image are defined as follows: The sub-pixel level real-time pixel coordinates obtained by the digital image correlation and sub-pixel interpolation algorithm in the real-time monitored image at time k are defined as... Based on the rigid body kinematics assumption, the static reference target has zero displacement in physical space. Its coordinate differences on the image plane at different times are entirely composed of linear geometric transformation components such as in-plane translation, rotation, and anisotropic scaling caused by optical axis drift, focal length variation, or mechanical vibration of the imaging system.
[0060] It should be noted that the initial pixel coordinates With sub-pixel level real-time pixel coordinates ( , The mapping relationship between them can be expressed by the following affine transformation formula: (5) In the formula, a 11 a 12 Let a be the coefficients of variables x and y in the abscissa mapping equation. 21 a 22 Let be the coefficients of variables x and y in the ordinate mapping equation, respectively, and t be the coefficients of these variables. x t y The constant terms of the horizontal and vertical coordinate mapping equations, respectively, together serve as six undetermined parameters characterizing the affine transformation. Specifically, by constraining the spatial topology of the static reference target group, it is ensured that the number N of targets involved in the calculation satisfies N≥3 and that all targets are not collinear on the image plane. Based on this condition, an overdetermined linear equation system with respect to six undetermined transformation parameters is constructed. The observation equations for all N reference targets are then uniformly encapsulated into matrix operation form. .
[0061] Where A is the initial pixel coordinates of each target. Composition The design matrix is 6×1, where p is the 6×1 dimensional parameter vector to be solved. b represents the sub-pixel level real-time pixel coordinates corresponding to each target. Composition Dimensional observation vector.
[0062] In the parameter solving stage, the least squares method is used to numerically solve the overdetermined linear equation system. This is achieved by constructing an objective function. That is, minimizing the sum of squared Euclidean distance residuals between the coordinates predicted by the affine model and the actual observed subpixel coordinates of all static benchmark targets, taking the partial derivative of the objective function with respect to the parameter vector p and setting it to zero, thereby obtaining the optimal estimated solution of parameter p under the minimum mean square error criterion.
[0063] Specifically, after obtaining the optimal solutions for the above six parameters, they are reorganized according to the homogeneous coordinate rule to construct... The affine transformation matrix M of dimension 1 has the following matrix form: (6) The 2×2 submatrix in the upper left corner of the matrix. Linear rotation and non-uniform scaling transformation operators in a two-dimensional plane are defined, and the third column vector... The translation transformation operator for the coordinate system is defined in the third line. To maintain the homogeneous extension term of the dimensions of the projected geometry calculation.
[0064] It should be noted that for any initial pixel coordinate point in the image space Expand it into a homogeneous coordinate vector Then, by performing a linear multiplication operation with matrix M, its mapped coordinates in the real-time monitoring image coordinate system can be calculated. This process mathematically establishes a global geometric mapping rule from the initial reference state space to the real-time monitoring state space, and quantifies the global rigid body motion components generated between images at each time step due to the external pose changes of the imaging system.
[0065] Step S5: Use the affine transformation matrix to perform inverse transformation correction on the sub-pixel level real-time pixel coordinates of the dynamic monitoring target group, calculate the vector difference between the corrected pixel coordinates and the initial pixel coordinates, and obtain the sub-pixel level corrected pixel displacement. During implementation, the pixel coordinates of the dynamically monitored target in the initial reference image are set as follows: And obtain the sub-pixel level real-time pixel coordinates in the real-time monitoring image at the corresponding monitoring time, calculated by correlation matching and sub-pixel interpolation algorithms. , To establish a unified measurement benchmark for pixel coordinates at different monitoring times relative to the initial reference image coordinate system, the affine transformation matrix M obtained in the previous steps is used to measure the sub-pixel level real-time pixel coordinates of the dynamically monitored target group. , Perform the inverse affine transformation correction operation.
[0066] Specifically, firstly, the real-time sub-pixel level pixel coordinates on the two-dimensional plane ( , Projecting onto homogeneous coordinate space, we construct a three-dimensional column vector form in homogeneous coordinates:
[0067] Furthermore, regarding the affine transformation matrix... Perform a matrix inversion operation to obtain its inverse matrix. , as the inverse linear mapping operator. This inverse matrix Mathematically, this represents the inverse mapping relationship from the real-time monitoring coordinate system to the initial reference coordinate system. Its internal elements consist of inverse rotation, scaling, and translation parameters corresponding to the forward transformation. Its matrix form is shown below: (7) Among them, the first two rows and first two columns of elements (i.e. This forms a two-dimensional linear transformation submatrix, used to describe inverse rotation and scaling transformations in the plane. The third column elements are translation vectors (i.e., ...). ), representing the inverse translation of the coordinate system; third line This is a homogeneous coordinate extension term that preserves the operational dimension.
[0068] It should be noted that, in order to eliminate the global rigid body displacement components caused by device pose changes or environmental interference during image acquisition, this inverse matrix is used. The inverse transformation multiplication operation on the real-time coordinate vector is calculated using the following formula: (8) In the formula, This refers to the corrected sub-pixel level pixel coordinates that are traced back and unified to the initial reference image coordinate system. These coordinate values represent the theoretical spatial position of the dynamically monitored target after removing global background motion.
[0069] Based on the corrected coordinate data, vector difference and modulus calculations are performed to obtain the true sub-pixel level corrected pixel displacement. .
[0070] Specifically, the corrected sub-pixel level pixel coordinates With the stored initial pixel coordinates Perform vector subtraction and use the two-dimensional Euclidean distance formula to calculate the displacement amplitude. The calculation formula is as follows: (9) In the formula, This is for the final sub-pixel level correction of pixel displacement.
[0071] The displacement value achieves spatial alignment of the coordinate reference through inverse transformation in the algorithm logic, and mathematically represents the projection modulus of the deformation displacement vector of the structure relative to the initial state on the imaging plane.
[0072] Step S6: Based on the determined pixel-to-physical size conversion coefficient, the sub-pixel level corrected pixel displacement is converted into the physical settlement of the building.
[0073] Retrieve the pixel-to-physical size conversion coefficients pre-determined based on Zhang Zhengyou's camera calibration method. Subpixel-level correction of pixel displacement Implement spatial scale reconstruction and value tracing.
[0074] It should be noted that the transformation coefficient K is defined as a linear scaling factor between the image discretization sampling space and the real physical measurement space. Its physical properties characterize the projection resolution and actual physical coverage length of a unit pixel on the target surface of the imaging sensor onto the tangent plane of the measured object, with units of millimeters per pixel (mm / pixel). From a geometric optics perspective, the value of this coefficient is constrained by internal parameter matrices such as the camera's effective focal length and pixel physical size, as well as external pose parameters such as the object distance along the optical axis and the angle between the optical axis and the measured surface. This parameter establishes a proportional mapping relationship between the pixel coordinate system and the world coordinate system, mapping the sub-pixel displacement vector based on the image grayscale gradient to an absolute physical dimension conforming to the International System of Units (SI).
[0075] Furthermore, following the linear projection imaging model and the geometric principles of similar triangles, a numerical mapping calculation is performed from the image pixel domain to the objective physical domain. This calculation process corrects pixel displacement at the sub-pixel level. As input variables, and the conversion coefficients By performing scalar multiplication, the actual settlement displacement of the building in three-dimensional Euclidean space can be calculated.
[0076] Specifically, physical settlement The calculation model is as follows: (10) Where, in the formula Defined as the final calculated physical settlement of the building, its unit of measurement is millimeters (mm). Mathematically, this value is equal to the linear product of the conversion coefficient and the pixel displacement vector. Physically and geometrically, it represents the actual Euclidean distance displacement of the monitored target point relative to the initial reference state along the vertical direction of gravity during the monitoring period, constituting an absolute quantitative index describing the changes in the spatial geometric position of the structure.
[0077] Compared with existing technologies, the building settlement monitoring method and system based on multi-reference point affine correction shown in this invention has the following advantages: This example constructs a dual-target system encompassing both static and dynamic benchmarks within the building monitoring area. It acquires time-series monitoring images using image acquisition equipment, applies deep learning models and digital image correlation algorithms to extract sub-pixel-level displacements of the targets, and performs multi-benchmark affine transformation operations based on the static benchmark target to eliminate rigid body motion errors introduced by environmental vibrations and equipment pose changes. Finally, by combining camera calibration parameters, the corrected pixel displacements are mapped to actual physical settlement, achieving synchronous and accurate perception of settlement changes at multiple measurement points. This example effectively improves the signal-to-noise ratio and reliability of monitoring data, reduces systematic errors caused by human intervention, and possesses excellent engineering applicability, enabling efficient assessment of large-scale building deformation.
[0078] Example 2 Please see Figure 2 A second embodiment of the present invention provides a building settlement monitoring system based on multi-reference point affine correction, applied to the method described in any of the above embodiments, the system comprising: Image acquisition module 10 is used to acquire data in a time sequence of the building monitoring area, and obtain a monitoring image set including an initial reference image and a real-time monitoring image. The monitoring area is equipped with a static reference target group and a dynamic monitoring target group. The target detection module 20 is used to detect the monitoring image set using a deep learning model, and to obtain the positioning areas of the static reference target group and the dynamic monitoring target group respectively. The digital image correlation processing module 30 is used to, based on the positioning area, employ a digital image correlation algorithm to extract image blocks centered on the initial pixel coordinates of each target in the initial reference image as a reference template area for subsequent tracking, and perform correlation matching on the real-time monitoring image according to the reference template area, thereby determining the sub-pixel level real-time pixel coordinates of the target in the real-time monitoring image.
[0079] The affine transformation matrix fitting module 40 is used to fit the affine transformation matrix based on the initial integer pixel coordinates of the static reference target group and the corresponding sub-pixel level real-time pixel coordinates. The affine correction and displacement calculation module 50 is used to perform inverse transformation correction on the sub-pixel level real-time pixel coordinates of the dynamic monitoring target group using the affine transformation matrix, calculate the vector difference between the corrected pixel coordinates and the initial pixel coordinates, and obtain the sub-pixel level corrected pixel displacement. The settlement calculation module 60 is used to convert the sub-pixel level corrected pixel displacement into the physical settlement of the building based on the determined pixel-to-physical size conversion coefficient.
[0080] Compared with existing technologies, the building settlement monitoring method and system based on multi-reference point affine correction shown in this example have the following advantages: This example constructs a dual-target system encompassing both static and dynamic benchmarks within the building monitoring area. It acquires time-series monitoring images using image acquisition equipment, applies deep learning models and digital image correlation algorithms to extract sub-pixel-level displacements of the targets, and performs multi-benchmark affine transformation operations based on the static benchmark target to eliminate rigid body motion errors introduced by environmental vibrations and equipment pose changes. Finally, by combining camera calibration parameters, the corrected pixel displacements are mapped to actual physical settlement, achieving synchronous and accurate perception of settlement changes at multiple measurement points. This example effectively improves the signal-to-noise ratio and reliability of monitoring data, reduces systematic errors caused by human intervention, and possesses excellent engineering applicability, enabling efficient assessment of large-scale building deformation.
[0081] Example 3 A third embodiment of the present invention provides a readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the building settlement monitoring method based on multi-reference point affine correction described in the above embodiments.
[0082] Example 4 The fourth embodiment of the present invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. When the processor executes the computer program, it implements the building settlement monitoring method based on multi-reference point affine correction described in the above embodiments.
[0083] Example 5 For examples consistent with the above embodiments, please refer to... Figure 3 , Figure 3 A schematic diagram of a terminal structure provided in an embodiment of this application is shown in the figure. It includes a processor, an input device, an output device, and a memory, which are interconnected. The memory stores a computer program, which includes program instructions. The processor is configured to invoke the program instructions. The program includes instructions for performing the following steps: Acquire a monitoring image set including an initial reference image and real-time monitoring images. The monitoring area is equipped with a static reference target group located in a non-settlement area and a dynamic monitoring target group located on the surface of the building. The monitoring image set is detected to obtain the positioning areas of the static reference target group and the dynamic monitoring target group, respectively; Based on the positioning area, a reference template is determined and matching is performed to determine the initial pixel coordinates of the static reference target group and the dynamic monitoring target group in the initial reference image, and their sub-pixel level real-time pixel coordinates in the real-time monitoring image are tracked and obtained. Based on the initial pixel coordinates of the static reference target group and the corresponding sub-pixel level real-time pixel coordinates, fit an affine transformation matrix; The sub-pixel level real-time pixel coordinates of the dynamic monitoring target group are corrected by inverse transformation using the affine transformation matrix. The vector difference between the corrected pixel coordinates and the initial pixel coordinates is calculated to obtain the sub-pixel level corrected pixel displacement. Based on the determined pixel-to-physical size conversion coefficient, the sub-pixel level corrected pixel displacement is converted into the physical settlement of the building.
[0084] Throughout this specification, the terms "one embodiment," "some embodiments," "illustrative example," or "specific example," etc., are intended to indicate that a particular feature, structure, material, or characteristic described in connection with that embodiment or example is at least included in one or more technical solutions of the present invention. In this document, the illustrative statements of the above terms do not necessarily refer to identical embodiments or examples. Furthermore, the specific features, structures, materials, or characteristics described can be combined or incorporated in any suitable manner in one or more embodiments or examples.
[0085] The above description is merely a preferred embodiment of the present invention, intended to aid in understanding the core technical ideas and implementation path of the invention. However, the scope of protection of the present invention is not limited thereto. It should be understood that those skilled in the art can make various equivalent substitutions, adaptive changes, or modifications to the above embodiments without departing from the concept and principles of the present invention, based on the needs of actual application scenarios. For example, functionally equivalent substitutions can be made to some technical features, or logically reasonable adjustments can be made to the order of steps. Any modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention should be covered within the scope of protection of the claims of the present invention. Therefore, the actual scope of protection of the present invention should be strictly determined by the content defined in the appended claims.
Claims
1. A method for monitoring building settlement based on multi-reference point affine correction, characterized in that, The method includes: Acquire a monitoring image set including initial reference images and real-time monitoring images. Static reference target groups located in non-settlement areas and dynamic monitoring target groups located on building surfaces are deployed within the monitoring area. The monitoring image set is detected to obtain the positioning areas of the static reference target group and the dynamic monitoring target group, respectively; Based on the positioning area, a reference template is determined and matching is performed to determine the initial pixel coordinates of the static reference target group and the dynamic monitoring target group in the initial reference image, and their sub-pixel level real-time pixel coordinates in the real-time monitoring image are tracked and obtained. Based on the initial pixel coordinates of the static reference target group and the corresponding sub-pixel level real-time pixel coordinates, fit an affine transformation matrix; The sub-pixel level real-time pixel coordinates of the dynamic monitoring target group are corrected by inverse transformation using the affine transformation matrix. The vector difference between the corrected pixel coordinates and the initial pixel coordinates is calculated to obtain the sub-pixel level corrected pixel displacement. Based on the determined pixel-to-physical size conversion coefficient, the sub-pixel level corrected pixel displacement is converted into the physical settlement of the building.
2. The building settlement monitoring method based on multi-reference point affine correction according to claim 1, characterized in that, The steps of determining a reference template based on the positioning region, performing matching, determining the initial pixel coordinates of the static reference target group and the dynamic monitoring target group in the initial reference image, and tracking and obtaining their sub-pixel level real-time pixel coordinates in the real-time monitoring image include: In the initial reference image, the center coordinates of each target localization region output by the deep learning model are determined as the initial pixel coordinates of the target. Centered on the initial pixel coordinates, an image block is cropped according to a preset template size and used as a reference template area for subsequent tracking; In the real-time monitoring image, a search area is defined with the initial pixel coordinates as the center. The ZNCC algorithm is used to calculate the correlation coefficient between the reference template area and each sub-region of the real-time image within the search area, and the position with the largest correlation coefficient value is determined as the real-time coordinate of the target at the integer pixel level. Within the neighborhood of the real-time coordinates at the integer pixel level, a discrete correlation coefficient matrix consisting of multiple ZNCC correlation coefficient values is constructed; The discrete correlation coefficient matrix is fitted using a quadratic surface fitting method, and the coordinates of the mathematical extremum points of the fitted quadratic surface are used as the sub-pixel level real-time pixel coordinates of the target.
3. The building settlement monitoring method based on multi-reference point affine correction according to claim 2, characterized in that, The step of fitting the discrete correlation coefficient matrix using a quadratic surface fitting method and using the coordinates of the mathematical extrema points of the fitted quadratic surface as the sub-pixel level real-time pixel coordinates of the target includes: Using the discrete pixel coordinates within the neighborhood of the real-time integer-level coordinates as independent variables and the corresponding ZNCC correlation coefficients as dependent variables, a quadratic surface mathematical model is constructed: In the formula, Indicates continuous local coordinates The fitted correlation coefficient value at the location, , , , , , The coefficients of the quadratic surface equation are to be determined by regression analysis of the discrete matrix data using the least squares method. The least squares method is used to perform regression analysis on the discrete correlation coefficient matrix to determine the undetermined coefficients of the quadratic surface equation. The equation of the quadratic surface is given by the variables respectively. and Find the first-order partial derivatives, set them to zero, and solve the system of equations to obtain the coordinates of the extreme points of the quadratic surface. , () is used as the sub-pixel level real-time pixel coordinates of the target.
4. The building settlement monitoring method based on multi-reference point affine correction according to claim 3, characterized in that, The coordinates of the extreme points of the quadratic surface ( , The formula for solving ) is as follows: In the formula, The x-coordinate of the target's sub-pixel-level real-time pixel coordinates. The ordinate is the sub-pixel level real-time pixel coordinate of the target. , , , , Let x, y, and x' be the variables in the equation of the quadratic surface, respectively. 2 variables x and y 2 The polynomial coefficients.
5. The building settlement monitoring method based on multi-reference point affine correction according to claim 1, characterized in that, The step of fitting an affine transformation matrix based on the initial pixel coordinates and corresponding sub-pixel level real-time pixel coordinates of the static reference target group includes: For any static reference target in the static reference target group, its initial pixel coordinates With the corresponding sub-pixel level real-time pixel coordinates The following affine transformation mapping relationship exists between them: wherein a 11 , a 12 are coefficients of variable x and variable y in the horizontal coordinate mapping equation, respectively, a 21 , a 22 are coefficients of variable x and variable y in the vertical coordinate mapping equation, respectively, t x , t y are constant terms of the horizontal coordinate mapping equation and the vertical coordinate mapping equation, respectively, which together serve as six undetermined parameters representing the affine transformation; By substituting the initial pixel coordinates and subpixel-level real-time pixel coordinates of each target in the static benchmark target group into the mapping relationship, an overdetermined linear equation system for solving the six undetermined parameters is constructed. The least squares method is used to solve the overdetermined linear equation system to determine the optimal solution for the six undetermined parameters; The determined optimal solution parameters are combined to form an affine transformation matrix M in the following form: Among them, the upper left corner of the matrix submatrix Linear rotation and non-uniform scaling transformation operators in a two-dimensional plane are defined, and the third column vector... The translation transformation operator for the coordinate system is defined in the third line. To maintain the homogeneous extension term of the dimensions of the projected geometry calculation.
6. The building settlement monitoring method based on multi-reference point affine correction according to claim 1, characterized in that, The steps of using the affine transformation matrix to perform inverse transformation correction on the sub-pixel level real-time pixel coordinates of the dynamic monitoring target group, calculating the vector difference between the corrected pixel coordinates and the initial pixel coordinates, and obtaining the sub-pixel level corrected pixel displacement include: Calculate the inverse matrix M of the affine transformation matrix. -1 ; Using the inverse matrix, the sub-pixel level real-time pixel coordinates of each target in the dynamic monitoring target group are subjected to inverse coordinate transformation to obtain their corresponding sub-pixel level corrected pixel coordinates. Calculate the vector difference between the subpixel-level corrected pixel coordinates and the corresponding initial pixel coordinates of each dynamic monitoring target, and determine the vector difference as the subpixel-level corrected pixel displacement of the target.
7. The building settlement monitoring method based on multi-reference point affine correction according to claim 1, characterized in that, The static reference target group includes at least three non-collinearly spaced static reference targets. Each static reference target is fixed on a stable structure unaffected by building settlement and maintains its relative position during monitoring. The static reference targets are used to provide multi-reference constraints for solving the affine transformation matrix.
8. A building settlement monitoring system based on multi-reference point affine correction, characterized in that, The system includes: The image acquisition module is used to acquire a set of monitoring images including an initial reference image and real-time monitoring images. The monitoring area is equipped with a static reference target group located in a non-settlement area and a dynamic monitoring target group located on the surface of the building. The target detection module is used to detect the monitoring image set and obtain the positioning areas of the static reference target group and the dynamic monitoring target group respectively; The digital image correlation processing module is used to determine a reference template based on the positioning area to perform matching, determine the initial pixel coordinates of the static reference target group and the dynamic monitoring target group in the initial reference image, and track and obtain their sub-pixel level real-time pixel coordinates in the real-time monitoring image. The affine transformation matrix fitting module is used to fit the affine transformation matrix based on the initial pixel coordinates of the static reference target group and the corresponding sub-pixel level real-time pixel coordinates. The affine correction and displacement calculation module is used to perform inverse transformation correction on the sub-pixel level real-time pixel coordinates of the dynamic monitoring target group using the affine transformation matrix, calculate the vector difference between the corrected pixel coordinates and the initial pixel coordinates, and obtain the sub-pixel level corrected pixel displacement. The settlement calculation module is used to convert the sub-pixel level corrected pixel displacement into the physical settlement of the building based on the determined pixel-to-physical size conversion coefficient.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the building settlement monitoring method based on multi-reference point affine correction as described in any one of claims 1-7.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the building settlement monitoring method based on multi-reference point affine correction as described in any one of claims 1-7.
Citation Information
Patent Citations
Three-dimensional deformation monitoring method and device, computer equipment and storage medium
CN115049718A
Bridge displacement change target-free monitoring method based on digital image correlation
CN116147498A
Steel rail longitudinal displacement monitoring method and system based on visual monitoring
CN116295040A
Settlement detection method and device based on machine vision
CN120538476A
Side slope slippage monitoring and early warning method based on image recognition technology
CN120612785A