Image deformation monitoring method, device, equipment and program product
Images are collected by a binocular camera and polar line constraint correction are performed. Combined with multi-scale census transformation and weighted Hamming distance calculation, the cost matrix is generated and multi-directional cost aggregation optimization is performed, which solves the problems of insufficient robustness, limited matching accuracy and low real-time performance of the deformation monitoring method in the existing technology, and achieves more efficient and accurate deformation monitoring.
Patent Information
- Application Number
- CN202510278559.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-27
AI Technical Summary
In the prior art, deformation monitoring methods have problems such as insufficient robustness, limited matching accuracy and low real-time performance.
Images are collected by a binocular camera, image correction of polar line constraints is carried out, images with different resolutions are constructed, census features and local grayscale variance are obtained, weighted coefficients are determined, dynamic weights and weighted Hamming distances are calculated, cost matrix is generated, and optimal parallax and depth deformation information is finally determined through multi-directional cost aggregation optimization.
It improves the calculation range and efficiency of binocular matching, improves matching accuracy and robustness, enhances adaptability to textured complex areas and low-texture areas, and obtains more accurate and reliable parallax and depth deformation information.
Smart Images

Figure CN120220053A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of deformation monitoring, and in particular, to an image deformation monitoring method, device, equipment and program product. Background Art
[0002] With the rapid development of infrastructure such as buildings and public transportation, the scale and volume of its construction are constantly expanding. However, during the long-term service of infrastructure, affected by internal and external factors such as natural disasters, human disturbances, and material aging, it will inevitably have an irreversible negative impact on the structural performance of the infrastructure, endangering the safety and service life of the infrastructure. If the damage to the infrastructure cannot be effectively monitored and timely maintained, serious safety accidents are extremely likely to occur. Therefore, accurately measuring the deformation characteristics of the infrastructure under service conditions and taking timely intervention measures are of great significance for ensuring the long-term stability and safety of the structure.
[0003] Traditional deformation monitoring methods usually rely on manual measurement, and direct observations are made manually with the help of tools such as total stations and theodolites. However, such manual measurement methods have problems such as low measurement efficiency and high labor costs, and it is difficult to meet the accuracy and efficiency requirements of modern engineering monitoring. In recent years, new monitoring means such as three-dimensional laser scanning technology and satellite positioning technology (such as GPS and GNSS) have been gradually introduced. Although they can achieve high-precision monitoring, their costs are high, data processing is complex, and there are limitations in some specific environments (such as satellite signal failure areas). In addition, the application of digital image correlation technology has gradually emerged, capturing structural displacement information through high-resolution cameras and combining image processing algorithms for deformation measurement. Although the accuracy has been improved, the contradiction between algorithm complexity and real-time performance restricts its wide application in engineering.
[0004] Binocular vision technology has become an important direction in deformation monitoring research in recent years due to its characteristics such as low cost, high precision, and strong applicability. The monitoring technology based on binocular vision can calculate the three-dimensional spatial displacement of the target through camera calibration and stereo matching algorithms, and has the advantage of not requiring the pasting of target marks. However, the current deformation monitoring methods based on binocular vision still have problems such as insufficient robustness, limited matching accuracy, and low real-time performance. Summary of the Invention
[0005] In view of this, the embodiments of the present application provide an image deformation monitoring method, device, equipment and program product to solve the problems of insufficient robustness, limited matching accuracy, and low real-time performance existing in the existing deformation monitoring methods.
[0006] The first aspect of the embodiments of the present application provides an image deformation monitoring method, and the method includes:
[0007] Collect the first binocular image through a binocular camera, and perform image rectification through epipolar constraint according to the pre-calibrated camera parameters to obtain the second binocular image;
[0008] Construct images with different resolutions based on the second binocular image, obtain the census features of the images with different resolutions, and determine the local gray variance of the images with different resolutions. Determine the first weight coefficient according to the local gray variance;
[0009] Perform weighted fusion on the census features of different resolutions according to the first weight coefficient to determine the weighted census feature of each pixel;
[0010] Determine the dynamic weight according to the gray difference of the corresponding pixels of the second binocular image and the distance of the pixels from the central pixel. According to the dynamic weight and the weighted census feature, traverse the pixel window corresponding to the candidate disparity in the epipolar direction to determine the weighted Hamming distance corresponding to different disparities. Determine the cost matrix corresponding to the second binocular image according to the weighted Hamming distance corresponding to different disparities;
[0011] Optimize the cost matrix through multi-directional cost aggregation, determine the optimal disparity according to the optimized cost matrix, determine the depth of the pixel according to the optimal disparity, and determine the deformation information of the pixel in the depth direction according to the change value of the depth.
[0012] Combined with the first aspect, in the first possible implementation manner of the first aspect, after obtaining the second binocular image, the method further includes:
[0013] Determine the sub-region calculation path of the planar deformation between the second binocular image and the predetermined reference image according to the similarity of the feature points and sub-regions of the second binocular image or the predetermined reference image;
[0014] Determine the deformation vector of each sub-region through an iterative optimization method according to the sub-region calculation path.
[0015] Combined with the first possible implementation manner of the first aspect, in the second possible implementation manner of the first aspect, determining the deformation vector of each sub-region through an iterative optimization method includes:
[0016] Determine the initial displacement value of the sub-region between any current image in the second binocular image and the predetermined reference image through the normalized cross-correlation method of the correlation function;
[0017] Perform interpolation processing on the gray values of the current image and the reference image according to the initial displacement value to obtain the interpolated current image and the interpolated reference image;
[0018] Through the inverse compositional Gauss-Newton method, perform iterative optimization on the sub-region matching of the interpolated current image and the parameter image to determine the deformation vector of the sub-region.
[0019] Combined with the second possible implementation manner of the first aspect, in the third possible implementation manner of the first aspect, through the inverse compositional Gauss-Newton method, perform iterative optimization on the sub-region matching of the interpolated current image and the parameter image to determine the deformation vector of the sub-region, including:
[0020] Determine the objective function of the gray difference between the reference image sub-region and the deformed sub-region of the current image through the normalized least squares criterion;
[0021] Approximate the objective function by second-order Taylor expansion, solve the deformation parameter increment through the Gauss-Newton method and update the parameters inversely;
[0022] Stop the iteration when the deformation parameter increment is less than a predetermined parameter threshold to obtain the deformation vector of the sub-region.
[0023] Combined with the first possible implementation manner of the first aspect, in the fourth possible implementation manner of the first aspect, determine the sub-region calculation path of the planar deformation between the second binocular image and the predetermined reference image according to the similarity of the feature points and sub-regions of the second binocular image or the predetermined reference image, including:
[0024] Detect the feature points in the second binocular image or the predetermined reference image, and determine the first sub-region as the calculation starting point according to the matching result of the feature points of the second binocular image and the reference image;
[0025] According to the similarity between the adjacent sub-region of the first sub-region and the first sub-region, find the second sub-region with the highest similarity as the propagation direction of the first sub-region, and according to the similarity between the adjacent sub-region of the second sub-region and the second sub-region, find the third sub-region with the highest similarity as the propagation direction of the second sub-region. By repeatedly determining the propagation direction according to the similarity of the adjacent sub-regions until the second binocular image or the reference image is traversed, obtain the sub-region calculation path.
[0026] Combined with the fourth possible implementation manner of the first aspect, in the fifth possible implementation manner of the first aspect, detecting the feature points in the second binocular image or the predetermined reference image includes:
[0027] Perform multi-scale Gaussian blur processing on the second binocular image or the predetermined reference image to generate Gaussian blur images of different scales;
[0028] Subtract the adjacent-scale Gaussian blur images pixel by pixel to generate a Gaussian difference pyramid;
[0029] Detect the extreme points of the Difference of Gaussian (DoG) pyramid and filter out the edge extreme points;
[0030] Use the filtered extreme points as the feature points in the second binocular image or a predetermined reference image.
[0031] Combined with the first aspect, in the sixth possible implementation manner of the first aspect, optimizing the cost matrix through multi-directional cost aggregation and determining the optimal disparity according to the optimized cost matrix includes:
[0032] Determine the dynamic weight according to the gray level difference of adjacent pixels;
[0033] Adaptively adjust the smoothness according to the dynamic weight to determine the energy function corresponding to the cost matrix;
[0034] Determine the path cost values in multiple predetermined directions respectively, and superimpose the path cost values in multiple directions to obtain the total path cost value;
[0035] Determine the optimal disparity according to the minimum value of the total path cost value.
[0036] Combined with the first aspect, in the seventh possible implementation manner of the first aspect, according to the pre-calibrated camera parameters, perform epipolar constraint-based image rectification to obtain the second binocular image, including:
[0037] Determine the internal parameters and distortion parameters of each camera in the binocular camera through Zhang's calibration method;
[0038] Perform binocular calibration according to the internal parameters of the camera to determine the external parameters of the binocular camera;
[0039] Perform epipolar rectification on the first binocular image according to the external parameters to minimize the horizontal alignment of the epipolar lines in the first binocular image and obtain the second binocular image.
[0040] The second aspect of the embodiments of the present application provides an image deformation monitoring device, and the device includes:
[0041] An image acquisition and preprocessing unit, configured to acquire a first binocular image through a binocular camera, and perform image rectification through epipolar constraint according to pre-calibrated camera parameters to obtain a second binocular image;
[0042] A multi-scale image construction unit, configured to construct images with different resolutions according to the second binocular image, obtain the census features of the images with different resolutions, and determine the local gray variance of the images with different resolutions, and determine the first weight coefficient according to the local gray variance;
[0043] A weighted fusion unit, configured to perform weighted fusion on census features of different resolutions according to the first weight coefficient to determine the weighted census feature of each pixel;
[0044] A cost matrix determination unit, configured to determine a dynamic weight according to the gray difference of corresponding pixels of the second binocular image and the distance of the pixel from the central pixel, traverse a pixel window corresponding to a candidate disparity in the epipolar direction according to the dynamic weight and the weighted census feature, determine the weighted Hamming distance corresponding to different disparities, and determine the cost matrix corresponding to the second binocular image according to the weighted Hamming distance corresponding to different disparities;
[0045] A depth deformation determination unit, configured to optimize the cost matrix through multi-directional cost aggregation, determine the optimal disparity according to the optimized cost matrix, determine the depth of the pixel according to the optimal disparity, and determine the deformation information of the pixel in the depth direction according to the change value of the depth.
[0046] A third aspect of the embodiments of the present application provides an image deformation monitoring device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the image deformation monitoring device implements the method according to any one of the first aspect.
[0047] A fourth aspect of the embodiments of the present application provides a computer program product, which, when running on a computer, causes the computer to execute the method according to the first aspect or its various implementation manners.
[0048] A fifth aspect of the embodiments of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of the method according to any one of the first aspect are implemented.
[0049] A sixth aspect of the embodiments of the present application provides a chip, configured to implement the method according to each implementation manner in the first aspect. Specifically, the above chip includes: a processor, configured to call and run a computer program from a memory, so that a device installed with the above chip executes the method according to the first aspect or its various implementation manners.
[0050] The beneficial effects of the embodiments of the present application compared with the prior art are as follows: In the embodiments of the present application, a first binocular image is collected by a binocular camera. Through the pre-calibrated camera parameters, epipolar constraint is used for image correction to obtain a second binocular image. Based on the corrected second binocular image, the weighted Hamming distance corresponding to different disparities is determined, effectively optimizing the calculation range of binocular matching, which is beneficial to improving the matching accuracy and efficiency. By determining the weighted Hamming distance through census transforms with different resolutions, the local information and global information can be effectively combined to dynamically adjust the weight distribution, significantly improving the quality of the cost matrix. For areas with complex textures, local detail features can be more accurately extracted. For areas with low textures, noise interference can be more effectively reduced and matching instability problems can be reduced, which is beneficial to enhancing the robustness to illumination changes and local noise, and obtaining more accurate and reliable disparity and depth deformation information. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0052] Figure 1 is a schematic implementation flowchart of an image deformation monitoring method provided by an embodiment of the present application;
[0053] Figure 2 is a schematic diagram of the relationship between disparity and cost provided by an embodiment of the present application;
[0054] Figure 3 is a schematic diagram of disparity calculation provided by an embodiment of the present application;
[0055] Figure 4 is a schematic diagram of the disparity effect provided by an embodiment of the present application;
[0056] Figure 5 is a schematic implementation flowchart of normalized cross-correlation calculation provided by an embodiment of the present application;
[0057] Figure 6 is a schematic implementation flowchart of the inverse compositional Gauss-Newton method provided by an embodiment of the present application;
[0058] Figure 7 is a schematic diagram of an image deformation monitoring device provided by an embodiment of the present application;
[0059] Figure 8 is a schematic diagram of an image deformation monitoring device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0060] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0061] In order to illustrate the technical solutions described in the present application, specific embodiments will be used for illustration below.
[0062] With the rapid expansion of the scale of infrastructure such as buildings and public transportation, the long-term service safety faces severe challenges. Under the action of internal and external factors such as natural disasters, human disturbances, and material aging, the infrastructure is prone to irreversible degradation of structural performance, seriously threatening its safety and service life. If the structural damage cannot be effectively monitored and timely maintained, major safety accidents may be triggered. Therefore, it is urgent to develop high-precision deformation monitoring technologies to provide support for structural safety assessment and maintenance decision-making.
[0063] Traditional manual measurement relies on equipment such as total stations and theodolites, and has problems such as low efficiency (time-consuming for single-point measurement) and high cost (continuous human input is required), making it difficult to meet the all-weather and high-density monitoring requirements of modern engineering. Although new technologies such as three-dimensional laser scanning and satellite positioning (GPS / GNSS) have improved the accuracy, the equipment procurement and data processing costs are high, the satellite signals are restricted by occluded areas, and the point cloud data processing process is complex, making it difficult to achieve dynamic monitoring.
[0064] Digital image correlation technology captures displacement information through a high-resolution camera. Although it breaks through the limitations of contact measurement, it is restricted by the contradiction between algorithm complexity and real-time performance, and the promotion of engineering applications is blocked.
[0065] Binocular vision technology has become a research hotspot in the field of deformation monitoring due to its characteristics such as low cost, non-contact, and high adaptability. However, in scenarios such as low texture, sudden illumination changes, and occlusion, the matching failure rate of binocular vision technology is still high, and error accumulation is easy, and the complexity of the global optimization calculation process is high, making it difficult to meet the dynamic monitoring requirements.
[0066] To solve the above problems, an image deformation monitoring method is proposed in the embodiments of the present application. Figure 1 The schematic diagram of the implementation process of this method is described in detail as follows:
[0067] In S101, a first binocular image is collected by a binocular camera, and according to the pre-calibrated camera parameters, image correction is performed through epipolar constraint to obtain a second binocular image.
[0068] Among them, the first binocular image includes the left image and the right image collected by the binocular camera.
[0069] Before using the binocular camera to collect the first binocular image in the embodiment of the present application, it further includes the step of calibrating the binocular camera, including calibrating the internal parameters and distortion parameters of the camera, as well as the external parameters of the binocular camera, etc.
[0070] For example, the Zhang calibration method can be used to determine the internal parameters (such as focal length, principal point coordinates, pixel size, etc.) and distortion parameters (such as radial distortion, tangential distortion) of each camera in the binocular camera. Based on the calibrated internal parameters, binocular calibration can be performed outside the binocular machine to determine the external parameters of the binocular camera (such as rotation matrix and translation vector). Based on the determined distortion parameters, distortion correction can be performed on the collected first binocular image.
[0071] Based on the calibrated external parameters, epipolar rectification can be performed on the first binocular image collected by the binocular camera to minimize the horizontal alignment of the epipolar lines of the first binocular image and obtain the second binocular image. The epipolar rectification process can include:
[0072] Epipolar transformation: Calculate the epipolar transformation matrix of the left and right images according to the external parameters.
[0073] Image remapping: Use the epipolar transformation matrix to remap the left and right images so that the corresponding points are on the same epipolar line.
[0074] Error evaluation: Evaluate the effect of epipolar rectification by calculating the reprojection error, and adjust the epipolar transformation matrix to minimize the error.
[0075] By minimizing the reprojection error of the left and right images through epipolar rectification, the calculation range of binocular matching can be effectively optimized and the calculation efficiency can be improved. For example Figure 2 As shown, when calculating the pixel cost of a certain pixel point in the left image and the right image, it only needs to search and calculate within the parallax range (x to x + dmax) in the same epipolar line direction, which can effectively reduce the calculation range of binocular matching and improve the calculation efficiency.
[0076] In the embodiment of the present application, for the problem of image quality degradation in a complex shooting environment, an image preprocessing scheme of collaborative noise reduction of median filtering and wavelet transform combined with histogram equalization enhancement is proposed.
[0077] Median filtering (such as a 5×5 window) can remove salt-and-pepper noise and protect the edge information of the image. Wavelet transform performs soft threshold processing on the high-frequency coefficients to effectively suppress Gaussian noise. Through histogram equalization processing, the gray-scale distribution can be dynamically adjusted, effectively enhancing the overall contrast of the image, providing a high-quality and highly reliable image data basis for subsequent disparity calculation and displacement monitoring, and effectively improving the matching accuracy.
[0078] In S102, different-resolution images are constructed based on the second binocular image, census features of the different-resolution images are obtained, and the local gray-scale variance of the different-resolution images is determined. The first weight coefficient is determined according to the local gray-scale variance.
[0079] The above different-resolution images are also images of different scales.
[0080] In the traditional Census transform, the calculation of features is limited to a single-scale window, making it difficult to simultaneously consider local texture information and global structural features. This makes the matching effect of the algorithm less than ideal in texture-complex regions and low-texture sparse regions. At the same time, the traditional Hamming distance performs poorly in scenarios of noise interference and illumination changes and is easily affected by local outliers, thus reducing the matching accuracy.
[0081] To solve these problems, the embodiments of the present application adopt a multi-scale Census transform. By constructing a multi-scale space of the image, the Census features of the images in each scale space are calculated separately, and combined with a weight-based Hamming distance calculation method, the robustness and accuracy of the cost matrix can be greatly improved.
[0082] The multi-scale Census transform extracts local information and global information by constructing a multi-scale space of the image and performing hierarchical processing on the image at different scales.
[0083] When constructing different-resolution images based on the second binocular image, the Gaussian pyramid method can be used to perform Gaussian blur and downsampling processing on the image layer by layer to obtain different-resolution images.
[0084] On images of different resolutions (i.e., different scales), the census features of each pixel can be calculated separately. The census features can be determined by comparing the gray-scale values of the pixels within the window with the central pixel. The determined census features can be used to describe the local structural relationship. The comparison window is a window determined by the central pixel and the pixels around the central pixel. The census feature determined based on the comparison window is the census feature of the central pixel.
[0085] The calculation formula of the census feature can be expressed as:
[0086]
[0087] Among them, C s (u, v) is the census feature of the pixel (u, v) at scale s, W represents the window area centered on the pixel (u, v), and I s (x, y) is the gray value of the pixel (x, y) at scale s, and Θ(x) is the step function, and its definition is as follows:
[0088]
[0089] By calculating the census features of images with different resolutions scale by scale, the local texture characteristics and global structure characteristics at different scales can be captured.
[0090] Determine the local gray variance of images with different resolutions. The first weight coefficient can be determined according to the ratio of the local gray variance at different scales to the sum of the local gray variances at all scales. The calculation formula can be expressed as:
[0091]
[0092] Among them, is the local gray variance at scale s, S is the number of multi-scales, and α s is the weight coefficient of scale s, which is used to control the contribution ratio of the census features at different scales to the weighted census feature.
[0093] In S103, according to the first weight coefficient, the census features with different resolutions are weighted and fused to determine the weighted census feature of each pixel.
[0094] In order to effectively fuse the feature information at different scales, the embodiments of the present application can weight the census features at each scale. The calculation formula of the weighted census feature after fusing the final multi-scale census features can be expressed as:
[0095]
[0096] Among them, C(u, v) represents the weighted census feature. By weighting and fusing the census features, the fused Census feature not only retains local details but also can reflect the global structure, improving the adaptability of the algorithm in high-texture regions and low-texture regions.
[0097] In S104, according to the gray difference between the corresponding pixels of the second binocular image and the distance of the pixel from the central pixel, the dynamic weight is determined. According to the dynamic weight and the weighted census feature, the pixel window corresponding to the candidate disparity is traversed in the epipolar direction to determine the weighted Hamming distance corresponding to different disparities. According to the weighted Hamming distance corresponding to different disparities, the cost matrix corresponding to the second binocular image is determined.
[0098] In view of the problem that the traditional Hamming distance is sensitive to noise and illumination changes, the embodiment of the present application adopts an improved weighted Hamming distance calculation method. During the generation process of the cost matrix, different weights are assigned to the Hamming distances of the pixels within the matching window. The weight assignment not only considers the second binocular image, including the gray difference between the pixels of the left image and the right image, but also considers the distance from the pixel to the center of the window. The improved weighted Hamming distance calculation formula can be expressed as:
[0099]
[0100] Where H is the weighted Hamming distance, d(p) represents the Hamming distance value of pixel p within the window, and w(p) is the dynamic weight. The definition of the dynamic weight can be expressed as:
[0101]
[0102] Where I L (p), I R (p) respectively represent the gray values of the corresponding pixels p in the left and right images, |I L (p) - I R (p)| is the gray difference, ‖p c - p‖ is the Euclidean distance from the center pixel p c of the window to pixel p, and γ1 and γ2 are parameters for controlling the weight distribution. By introducing the dual constraints of the gray difference and the pixel distance, the dynamic weight can assign a larger weight to the center of the window while reducing the influence of abnormal noise on the matching, thereby enhancing the robustness of the weighted Hamming distance.
[0103] It can be seen that the method for determining the cost matrix based on the multi-scale census transform and the improved weighted Hamming distance proposed in the embodiment of the present application can effectively combine local and global information and dynamically adjust the weight distribution, thus significantly improving the quality of the cost matrix. In the area with complex texture, the method can accurately extract local detail features; in the low-texture area, the method can effectively avoid noise interference and unstable matching problems, providing a more reliable input for subsequent disparity estimation.
[0104] In S105, the cost matrix is optimized through multi-directional cost aggregation, the optimal disparity is determined according to the optimized cost matrix, the depth of the pixel is determined according to the optimal disparity, and the deformation information of the pixel in the depth direction is determined according to the change value of the depth.
[0105] After the cost calculation is completed, cost matrices corresponding to the left and right images are generated. However, since the cost calculation only considers the local correlation of pixels, it is very sensitive to noise and prone to matching errors in texture smooth areas and object edge areas. To solve this problem, in the embodiments of the present application, the cost value is optimized through cost aggregation, so that the optimized cost value can more accurately reflect the correlation between pixels. The goal of cost aggregation is to find the optimal cost value for each pixel, so as to minimize the global energy function of the entire image. To achieve this goal, the embodiments of the present application can introduce a dynamic weight mechanism based on gray gradient in cost aggregation and combine the eight-path aggregation method for efficient calculation.
[0106] The dynamic weight can adaptively adjust the smooth term in the cost aggregation process by measuring the gray difference of pixels. The dynamic weight can be expressed as follows:
[0107]
[0108] where, ‖I p -I q ‖ represents the gray difference between pixels p and q, and σ represents the gray gradient, which is used to control the degree of gradient sensitivity of the dynamic weight. In the edge area, due to the large gray gradient, the dynamic weight λ p,q is small, which can effectively weaken the effect of disparity smoothing, so as to retain the true disparity boundary; while in the smooth area, due to the small gray difference, the dynamic weight λ p,q is large, which can enhance the continuity of the disparity and reduce the jumping phenomenon. By introducing the dynamic weight, the cost aggregation can adapt to the local characteristics of the image, so as to more accurately optimize the disparity estimation globally.
[0109] The energy function optimized based on the dynamic weight can be expressed as follows:
[0110]
[0111] where C(p, D(p)) is the matching cost of pixel p at the disparity value D(p), λ p,q is the dynamic weight, N is the neighborhood range of pixel p, and |D(p) - D(q)| represents the discontinuity of the disparity value. The first term on the right is the data term, which is used to measure the accumulated value of the disparity and the original matching cost; the second term on the right is the smooth term, which is used to punish the change of the disparity within the neighborhood, so that the disparity map maintains a certain continuity globally.
[0112] To efficiently solve this energy function, the embodiments of the present application may adopt a dynamic programming method based on multi-path aggregation, such as a dynamic programming method using eight-path aggregation. The eight-path aggregation calculates the path cost value pixel by pixel along eight directions in the image (including the horizontal direction, the vertical direction, and two sets of diagonal directions), and superimposes the calculation results of the path cost values in the eight directions to obtain the total path cost value, and optimizes the cost matrix according to the total path cost value. The calculation formula of the path cost value can be expressed as follows:
[0113]
[0114] Wherein, L r (p, d) represents the path cost value of pixel p on path r with a disparity value of d, and N r (p) is the neighboring pixel of pixel p on path r, and λ p,q is the dynamic weight. The core idea of this calculation formula is to cumulatively calculate the matching cost of the current pixel and the path cost value of the previous pixel pixel by pixel on each path, and at the same time combine the smooth term to penalize the jump of the disparity value.
[0115] The final total path cost value is obtained by superimposing and calculating the cost values of the eight paths:
[0116]
[0117] Wherein, S(p, d) is the total path cost value of pixel p at the disparity value d. The optimization of the disparity value is completed by selecting the disparity d with the minimum total path cost value, that is:
[0118] D(p) = argmin d S(p, d)
[0119] By decomposing the two-dimensional global cost aggregation problem into a one-dimensional path aggregation problem through eight-path aggregation, the computational complexity can be effectively and significantly reduced, and the processing speed can be significantly improved. Experiments show that the traditional method usually takes more than one minute to process an image, while the method based on eight-path aggregation proposed in the embodiments of the present application only takes a few seconds to complete the processing. At the same time, the introduction of the dynamic weight further enhances the robustness of the disparity estimation, making the method show higher accuracy and stability in both the edge region and the texture smooth region.
[0120] After cost aggregation, an optimized cost matrix can be obtained. According to the principle of cost calculation, the smaller the cost value between the pixels of the left and right images, the higher the correlation between the two pixels. In the embodiments of the present application, the WTA (Winner-Takes-All) algorithm can be adopted, that is, among the cost values corresponding to the disparities of the pixels in the matching range of the right image for a certain pixel in the left image, the disparity corresponding to the minimum cost value is selected as the optimal disparity. In order to make the disparity calculation reach sub-pixel accuracy, the embodiments of the present application can adopt the method of quadratic curve interpolation. As Figure 3 shown, in the embodiments of the present application, quadratic curve fitting can be performed on the cost value of the optimal disparity and the cost values of the two adjacent disparities before and after. The disparity value corresponding to the extreme point of the curve is the final disparity value. As Figure 3 shown, quadratic curve fitting is performed on the optimal disparity 7 and the two adjacent disparities 6 and 8 before and after, and the extreme point 6.5 of the curve is obtained as the final disparity value.
[0121] In addition, before calculating the depth, the embodiments of the present application can also eliminate the wrong matching points. The left-right consistency method (L-R Check), the RemovePeaks algorithm, and the UniquenessCheck algorithm can be used successively to eliminate the gross errors, so as to improve the calculation accuracy. After calculating the disparity, the depth information corresponding to each pixel can be determined according to the determined disparity. The calculation formula for calculating the depth through the disparity can be expressed as follows:
[0122]
[0123] According to the principle of camera imaging, the depth Z can be obtained based on X R -X T (disparity), f (camera focal length), and B (distance between the centers of the two cameras). After obtaining the depth map, the average depth of the target area can be calculated within the region of interest (ROI), and the difference between the average depths before and after the displacement is compared as the displacement of the target vertical photography plane. For example Figure 4 shown, for the region of interest ROI in the original image ( Figure 4 the left image in, corresponding to the left image or the right image in the second binocular image collected), the optimal disparity map ( Figure 4 the middle image in) is determined by minimizing the energy function. According to the optimal disparity map, the depth map ( Figure 4 the right image in) is generated using the above depth calculation formula.
[0124] In summary, in the embodiments of the present application, by introducing the method of multi-scale census transform and weighted Hamming distance calculation, and combining dynamic programming for cost aggregation, not only the matching accuracy is significantly improved, but also there is an obvious advantage in terms of computational efficiency compared with traditional methods. In addition, the embodiments of the present application further optimize the calculation strategies of disparity and depth, use the quadratic curve interpolation method to achieve sub-pixel level disparity calculation, and combine a variety of false matching point elimination techniques, effectively improving the accuracy and robustness of the final result.
[0125] In a possible implementation manner, after obtaining the depth displacement of the second binocular image, the embodiments of the present application can further determine the plane displacement of the second binocular image through an iterative optimization method. The overall goal of calculating the plane displacement is to obtain the displacement within the region of interest (ROI) of the picture. The overall calculation idea is to iteratively obtain the corresponding relationship by comparing the position changes of the subsets of the reference image and the current image, and finally obtain the continuous displacement. The transformation between the initial reference image and the current image is represented by a linear, first-order transformation. The target vector for calculating the plane displacement in the embodiments of the present application is:
[0126]
[0127] In the formula, x refi , y refj represent the x and y coordinates in the sub-region of the initial reference image, x refc and y refc are the x coordinate and y coordinate of the center of the sub-region of the initial reference image, and are the x coordinate and y coordinate in the current image. The deformation vector p includes displacements in two directions, shear deformations in two directions, and volume deformations in two directions.
[0128] The embodiments of the present application can determine the sub-region calculation path according to the feature points of any image in the predetermined reference image or the second binocular image and the similarity of each sub-region, and determine the deformation vector of each sub-region according to the determined calculation path.
[0129] Among them, when determining the deformation vector of the sub-region through an iterative method, the initial displacement value of the sub-region between any current image in the second binocular image and the predetermined reference image can be determined by the normalized cross-correlation method of the relevant function; then, according to the initial displacement value, gray value interpolation processing is performed on the current image and the reference image to obtain the interpolated current image and the interpolated reference image; through the inverse compositional Gauss-Newton method, iterative optimization of the sub-region matching of the interpolated current image and the parameter image is performed to determine the deformation vector of the sub-region.
[0130] During iterative calculation, an initial value needs to be obtained, so the initial displacement value calculation is required. In the embodiments of the present application, the initial displacement value of integer pixels is obtained based on a correlation function, and the correlation function is used to compare the correlation of pixel gray values between the two. In the embodiments of the present application, the normalized cross-correlation criterion in the correlation function can be used, and the formula is as follows:
[0131]
[0132] Wherein, in the formula, f and g are respectively the function representations of the reference image and the current image, representing the gray value of each coordinate, and f m and g m are respectively the average gray values of the final reference image and the current image sub-region, and are defined as:
[0133]
[0134] Wherein, n(S) is the number of pixels in the sub-region S, and the formula indicates good matching when approaching 1 in C cc When approaching 1, it indicates good matching.
[0135] Perform normalized cross-correlation calculation on all sub-regions in the reference image and the current image, and calculate the pixel displacement of the sub-region with the highest matching degree, and use this as the initial displacement value of integer pixels. As Figure 5 shown in the flow schematic diagram of normalized cross-correlation calculation, for any subset in the reference image, the sub-region with the highest matching degree with the subset of the reference image in the current image can be determined through the extreme point of the correlation function.
[0136] After calculating the initial displacement value of integer pixels, in order to obtain sub-pixel accuracy in the Gauss-Newton iteration, the embodiments of the present application can perform interpolation processing on the reference image and the current image, for example, five-degree B-spline interpolation can be performed. Approximate the gray value of the image through the linear combination of B-spline and basis spline, and the B-spline interpolation can be expressed as:
[0137] g(x) = Σ k∈Z c(k)β n (x - k)
[0138] Wherein, c(k) is the B-spline coefficient value at integer k, β n (x - k) represents the B-spline kernel value at x - k, and g(x) represents the interpolation signal value at x. n is the B-spline kernel, and in the embodiments of the present application, b can be set to 5 (five-degree kernel), and Z is the set of integers. The equation of the B-spline kernel can be expressed as:
[0139]
[0140] The present application can determine the B-spline coefficients by performing discrete Fourier transform on the formula, and can interpolate the gray value using the one-dimensional B-spline coefficient value to obtain the interpolated gray image.
[0141] After calculating the initial displacement value and performing interpolation, the embodiments of the present application can iteratively obtain an accurate deformation vector p based on the sub-region matching algorithm. The algorithm in the embodiments of the present application is different from the single-pixel matching of traditional image displacement detection algorithms, which only considers planar rigid body displacement and may cause inaccurate displacement calculation in engineering applications. In view of the irregularity and complexity of object displacement, and considering that the measurement region will be sheared and magnified after deformation, the embodiments of the present application comprehensively consider the displacements of multiple pixels in the image sub-region through iterative matching of the deformed image sub-regions, and the calculated displacement is more accurate than single-pixel matching. The embodiments of the present application perform iterative matching of the Inverse compositional Gauss-Newton method on the image sub-regions to accurately match the sub-regions before and after deformation, thereby obtaining the accurate displacement of the sub-regions. The iteration uses the function C LS ,C LS is the normalized least squares criterion, expressed as:
[0142]
[0143] Equation C LS is defined as a function that accepts a single parameter p. The iterative equation used in the embodiments of the present application is obtained by performing a second-order Taylor series expansion of C LS near p0:
[0144]
[0145] where p0 contains the initial deformation parameters, and the value of Δp is calculated by determining where the derivative of Δp is a zero vector in the equation. Among them, is the gradient of C Ls at p0, is the hessian matrix of C Ls at p0. By taking the derivative of the equation to obtain an explicit solution for Δp, and using p0 + Δp to approach the required solution, through continuous iteration until a sufficiently close solution is found. The iterative optimization equation used in the embodiments of the present application can be the Newton-Raphson iteration formula, expressed as:
[0146]
[0147] After simplification, it is obtained:
[0148] It can be approximated by the hessian matrix to obtain the Gauss-Newton iteration formula.
[0149] Such as Figure 6The figure shows a schematic diagram of the inverse compositional Gauss-Newton method in an embodiment of the present application. Different from the forward Gauss-Newton method, in the inverse compositional Gauss-Newton iteration, the solution of Δp is performed for the reference image, and the deformation of the current image is calculated by taking the inverse matrix of Δp. Therefore, the Hessian matrix is calculated at the same position in the reference image sub-region in each iteration, which means that the hessian matrix is the same in each iteration, so it only needs to be calculated once, and it has a faster response speed in the iteration. When Δp is small enough, such as less than a preset deformation threshold, it is considered that the exact solution is obtained by iteration, and the iteration is stopped.
[0150] The traditional displacement calculation method obtains the displacement value based on pixel matching, without considering the influence of surrounding pixels, and the calculation robustness is poor. Moreover, most of the algorithms based on pixel matching require a target for assistance, and the calculation accuracy is poor when there is no monitoring target. The algorithm in the embodiment of the present application is based on pixel sub-region matching, with stronger robustness, and can match the corresponding sub-regions before and after displacement without the participation of a target. The embodiment of the present application performs planar displacement calculation based on the Gauss-Newton iteration method, performs pixel matching on small displacement sub-regions, and has higher calculation accuracy.
[0151] In addition, since the propagation direction of the sub-region and the selection of the initial point are closely related to the calculation accuracy, in order to efficiently and accurately solve the displacement calculation of the entire image, the embodiment of the present application proposes a reliability-guided algorithm based on SIFT feature matching. The algorithm can use the SIFT feature matching algorithm to select the initial sub-region to calculate the position, and then determine the calculation direction of the sub-region in combination with the calculation results of the relevant functions of the surrounding sub-regions. Among them, the SIFT algorithm can detect the key points in the image, and these points will not change their characteristics due to external factors such as movement, rotation, scaling, affine transformation, and brightness. Therefore, the method proposed in this paper can capture the key points of the calculation and further improve the calculation accuracy when performing the displacement calculation of the entire image. The calculation steps of the SIFT feature matching algorithm can be as follows:
[0152] 1. Construct a scale space by using Gaussian kernel filtering. Define the original image as L(x, y, σ), and define the convolution operation of the original image I(x, y) and a two-dimensional Gaussian function G(x, y, σ) with a variable scale. The Gaussian function is shown in the following formula.
[0153]
[0154] L(x, y, σ) = G(x, y, σ) * I(x, y)
[0155] The scale space calculation equation is as shown in the above formula. In the above formula, (x, y) represents the position of the image pixel, * represents the convolution operation, and σ represents the image scale parameter.
[0156] Through continuous convolution operations, the image is continuously scaled by a factor, resulting in images of different resolutions, and Gaussian blur with different parameters is applied to each scaled image.
[0157] 2. Subtract each pair of adjacent images after Gaussian blur and reconstruct to obtain a Difference of Gaussian (DOG) pyramid, as Figure 7 shown in the schematic diagram of a DOG pyramid provided by an embodiment of this application. In the schematic diagram of the DOG pyramid, through continuous convolution operations, the image is continuously scaled by a multiple to obtain images of different resolutions, and Gaussian blur with different parameters is applied to each scaled image. Subtract each pair of adjacent images after Gaussian blur processing and reconstruct to obtain the Figure 6 DOG pyramid as shown. Perform extreme point detection based on the obtained DOG pyramid, traverse the image to find the pixel points whose pixel values are greater than or less than all surrounding adjacent points as extreme points.
[0158] Then, judge the points with large gray-scale changes in the edge contour part of the matrix and eliminate them to improve the accuracy of key point positioning. Use the points after feature matching as the starting points for iterative calculation.
[0159] After determining the seed points, select the sub-region calculation direction. To reduce the accumulation of errors in the propagation direction, this paper selects the sub-region propagation direction based on the correlation function. Use the feature points obtained by the SIFT feature matching algorithm as the initial points, calculate the correlation function for the sub-regions in the four adjacent directions, and take the direction with the highest calculated value of the correlation function as the calculation path. The schematic diagram is shown in the figure.
[0160] This paper calculates the initial value of sub-region iteration through displacement initial value calculation, obtains the sub-pixel displacement of the sub-region through multi-pixel displacement calculation based on sub-region matching, and finally completes the calculation of the planar displacement by combining the full-image displacement detection of the SIFT feature matching algorithm. After measuring the true distance corresponding to each pixel, the planar displacement value of the target or the photographed object within the ROI area can be obtained. Finally, by combining with the determined image depth displacement, image deformation monitoring based on binocular vision can be realized.
[0161] In the embodiments of the present application, aiming at the problem of binocular deformation monitoring in complex environments, a high-precision monitoring algorithm integrating camera calibration, image preprocessing, disparity calculation, depth analysis, and planar displacement measurement is proposed. In image preprocessing, the Zhang's calibration method combined with epipolar rectification significantly reduces the matching error, and a median + wavelet filtering and histogram equalization method is designed to enhance the signal-to-noise ratio and contrast of the image. In disparity calculation, the multi-scale census transform and weighted Hamming distance method proposed in this paper can effectively combine local and global information, enhancing the robustness of the cost matrix; the eight-path aggregation optimization algorithm based on dynamic programming further improves the accuracy and calculation efficiency of disparity calculation. In addition, the full-image displacement calculation method based on SIFT feature guidance proposed in this paper, through sub-region matching and Gauss-Newton iterative optimization, not only solves the problem of cumulative displacement propagation error in traditional methods, but also significantly improves the displacement measurement accuracy in low-texture regions. Experimental results show that the image deformation monitoring method in the embodiments of the present application exhibits good robustness and accuracy in complex scenarios, and can be widely applied to fields such as engineering deformation monitoring and structural health assessment, providing new ideas and practical support for the development of image monitoring technology.
[0162] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution is prior or posterior. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0163] Figure 7 The following is a schematic diagram of an image deformation monitoring device provided by the embodiments of the present application. The device includes:
[0164] An image acquisition and preprocessing unit 701, configured to acquire a first binocular image through a binocular camera, and perform image rectification through epipolar constraint according to pre-calibrated camera parameters to obtain a second binocular image;
[0165] A multi-scale image construction unit 702, configured to construct images with different resolutions according to the second binocular image, obtain the census features of the images with different resolutions, and determine the local gray variance of the images with different resolutions, and determine a first weight coefficient according to the local gray variance;
[0166] A weighted fusion unit 703, configured to perform weighted fusion on the census features with different resolutions according to the first weight coefficient to determine the weighted census feature of each pixel;
[0167] A cost matrix determination unit 704 is configured to determine a dynamic weight according to the gray difference of corresponding pixels of the second binocular image and the distance of the pixel from the central pixel, traverse a pixel window corresponding to a candidate disparity in the epipolar direction according to the dynamic weight and the weighted census feature, determine a weighted Hamming distance corresponding to different disparities, and determine a cost matrix corresponding to the second binocular image according to the weighted Hamming distance corresponding to different disparities;
[0168] A depth deformation determination unit 705 is configured to optimize the cost matrix through multi-directional cost aggregation, determine an optimal disparity according to the optimized cost matrix, determine the depth of a pixel according to the optimal disparity, and determine deformation information of the pixel in the depth direction according to a change value of the depth.
[0169] Figure 7 The image deformation monitoring device shown is associated with Figure 1 the image deformation monitoring method shown.
[0170] Figure 8 is a schematic diagram of an image deformation monitoring device provided by an embodiment of the present application. As Figure 8 shown, the image deformation monitoring device 8 of this embodiment includes: a processor 80, a memory 81, and a computer program 82 stored in the memory 81 and executable on the processor 80, such as a deformation monitoring program. When the processor 80 executes the computer program 82, the steps in each embodiment of the above-mentioned various image deformation monitoring methods are implemented. Alternatively, when the processor 80 executes the computer program 82, the functions of each module / unit in each of the above-mentioned device embodiments are implemented.
[0171] Exemplarily, the computer program 82 may be divided into one or more modules / units. The one or more modules / units are stored in the memory 81 and executed by the processor 80 to complete the present application. The one or more modules / units may be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program 82 in the image deformation monitoring device 8.
[0172] The image deformation monitoring device may include, but is not limited to, a processor 80 and a memory 81. Those skilled in the art can understand that Figure 8 this is only an example of an image deformation monitoring device 8, and does not constitute a limitation on the image deformation monitoring device 8. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the image deformation monitoring device may further include an input / output device, a network access device, a bus, etc.
[0173] The so-called processor 80 may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0174] The memory 81 may be an internal storage unit of the image deformation monitoring device 8, such as the hard disk or memory of the image deformation monitoring device 8. The memory 81 may also be an external storage device of the image deformation monitoring device 8, such as a plug-in hard disk equipped on the image deformation monitoring device 8, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 81 may also include both the internal storage unit of the image deformation monitoring device 8 and the external storage device. The memory 81 is used to store the computer program and other programs and data required by the image deformation monitoring device. The memory 81 may also be used to temporarily store the data that has been output or will be output.
[0175] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit. The above integrated unit may be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.
[0176] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0177] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.
[0178] In the embodiments provided in this application, it should be understood that the disclosed device / terminal device and method can be implemented in other ways. For example, the device / terminal device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be electrical, mechanical or other forms.
[0179] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0180] In addition, the functional units in each embodiment of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0181] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of this application, it can also be completed by hardware related to computer program instructions. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0182] In addition, an embodiment of this application also provides a computer program product. When it runs on a computer, it causes the computer to execute the methods in the above-described various implementation manners.
[0183] The above-described embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A method for monitoring image deformation, characterized in that: The method comprises: A first binocular image is acquired by a binocular camera, and image correction is performed by epipolar constraints according to pre-calibrated camera parameters to obtain a second binocular image; constructing images with different resolutions according to the second binocular image, obtaining survey features of the images with different resolutions, determining local grayscale variances of the images with different resolutions, and determining a first weight coefficient according to the local grayscale variance; Performing weighted fusion on survey features of different resolutions according to the first weight coefficient to determine a weighted survey feature of each pixel; Determine a dynamic weight according to the grayscale difference of the pixels corresponding to the second binocular image and the distance between the pixels and the center pixel, traverse the pixel window corresponding to the candidate disparity in the epipolar direction according to the dynamic weight and the weighted census feature, determine the weighted Hamming distances corresponding to different disparities, and determine the cost matrix corresponding to the second binocular image according to the weighted Hamming distances corresponding to the different disparities; The cost matrix is optimized through multi-directional cost aggregation, an optimal disparity is determined according to the optimized cost matrix, a depth of a pixel is determined according to the optimal disparity, and deformation information of the pixel in a depth direction is determined according to a change value of the depth.
2. The method according to claim 1, characterized in that After obtaining the second binocular image, the method further includes: Determining a sub-region calculation path of a plane deformation between the second binocular image and a predetermined reference image according to the similarity between the feature points and the sub-region of the second binocular image or the predetermined reference image; According to the sub-region calculation path, the deformation vector of each sub-region is determined by an iterative optimization method.
3. The method according to claim 2, characterized in that The deformation vector of each sub-area is determined by an iterative optimization method, including: Determine an initial displacement value of a sub-area between any current image in the second binocular image and a predetermined reference image by a normalized cross-correlation method of a correlation function; Performing grayscale interpolation processing on the current image and the reference image according to the initial displacement value to obtain an interpolated current image and an interpolated reference image; By using the inverse synthetic Gauss-Newton method, the sub-region matching of the interpolated current image and the parameter image is iteratively optimized to determine the deformation vector of the sub-region.
4. The method according to claim 3, characterized in that By using the inverse synthetic Gauss-Newton method, iterative optimization of sub-region matching is performed on the interpolated current image and the parameter image to determine the deformation vector of the sub-region, including: Determine the objective function of the grayscale difference between the reference image subregion and the deformed subregion of the current image by using the normalized least squares criterion; The objective function is approximated by the second-order Taylor expansion, and the deformation parameter increment is solved by the Gauss-Newton method and the parameters are updated inversely; When the deformation parameter increment is less than a predetermined parameter threshold, the iteration is stopped to obtain the deformation vector of the sub-region.
5. The method according to claim 2, characterized in that: Determining a sub-region calculation path of a plane deformation between the second binocular image and a predetermined reference image according to the similarity between the feature point and the sub-region of the second binocular image or the predetermined reference image, comprising: Detecting feature points in the second binocular image or a predetermined reference image, and determining a first sub-area as a calculation starting point according to a matching result of the feature points between the second binocular image and the reference image; According to the similarity between the adjacent sub-region of the first sub-region and the first sub-region, the second sub-region with the highest similarity is found as the propagation direction of the first sub-region, and according to the similarity between the adjacent sub-region of the second sub-region and the second sub-region, the third sub-region with the highest similarity is found as the propagation direction of the second sub-region, and the propagation direction is repeatedly determined according to the similarity of the adjacent sub-regions until the second binocular image or the reference image is traversed to obtain the sub-region calculation path.
6. The method according to claim 5, characterized in that Detecting feature points in the second binocular image or a predetermined reference image includes: Performing multi-scale Gaussian blur processing on the second binocular image or the predetermined reference image to generate Gaussian blurred images of different scales; Subtract the Gaussian blurred images of adjacent scales pixel by pixel to generate a Gaussian difference pyramid; Detecting extreme points of the Gaussian difference pyramid and filtering out edge extreme points; The screened extreme value points are used as feature points in the second binocular image or a predetermined reference image.
7. The method according to claim 1, characterized in that The cost matrix is optimized by multi-directional cost aggregation, and an optimal disparity is determined according to the optimized cost matrix, including: Determine dynamic weights based on grayscale differences of adjacent pixels; Adaptively adjusting the smoothness according to the dynamic weights to determine an energy function corresponding to the cost matrix; Determine path cost values respectively according to a plurality of predetermined directions, and superimpose the path cost values of the plurality of directions to obtain a total path cost value; The optimal disparity is determined according to the minimum value of the total path cost.
8. The method according to claim 1, characterized in that According to the pre-calibrated camera parameters, image correction is performed through epipolar constraints to obtain the second binocular image, including: Determine the internal parameters and distortion parameters of each camera in the binocular camera by Zhang calibration method; Performing binocular calibration according to the internal parameters of the camera to determine the external parameters of the binocular camera; The first binocular image is subjected to epipolar correction according to the external parameters to minimize epipolar horizontal alignment of the first binocular image, thereby obtaining a second binocular image.
9. An image deformation monitoring device, characterized in that: The device comprises: An image acquisition preprocessing unit, used to acquire a first binocular image through a binocular camera, and perform image correction through epipolar constraints according to pre-calibrated camera parameters to obtain a second binocular image; a multi-scale image construction unit, configured to construct images of different resolutions according to the second binocular image, obtain survey features of the images of different resolutions, determine local grayscale variances of the images of different resolutions, and determine a first weight coefficient according to the local grayscale variance; A weighted fusion unit, configured to perform weighted fusion on the census features of different resolutions according to the first weight coefficient to determine a weighted census feature of each pixel; a cost matrix determination unit, configured to determine a dynamic weight according to a grayscale difference of pixels corresponding to the second binocular image and a distance between the pixels and a central pixel, traverse a pixel window corresponding to a candidate disparity in an epipolar direction according to the dynamic weight and the weighted census feature, determine a weighted Hamming distance corresponding to different disparities, and determine a cost matrix corresponding to the second binocular image according to the weighted Hamming distance corresponding to the different disparities; A depth deformation determination unit is used to optimize the cost matrix through multi-directional cost aggregation, determine the optimal disparity according to the optimized cost matrix, determine the depth of the pixel according to the optimal disparity, and determine the deformation information of the pixel in the depth direction according to the change value of the depth.
10. An image deformation monitoring device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the image deformation monitoring device implements the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Semi-global stereo matching method based on binocular vision
CN116934874A
Image deformation monitoring method and device, equipment and storage medium
CN117593349A
Cited By
Parameter adaptive matching binocular ranging method and system based on scene features
CN120876574A