Displacement detection method, device, equipment and program product
By introducing a self-attention mechanism and dynamic path adaptive aggregation method in the stereo matching algorithm, the problem of low matching accuracy in complex scenarios in the prior art is solved, and more efficient displacement detection is achieved.
Patent Information
- Application Number
- CN202510064806.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-15
Smart Images

Figure CN119991804A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision, and in particular to displacement detection methods, devices, equipment and program products. Background Art
[0002] In the field of computer vision, depth calculation, displacement calculation and stereo matching are key technical applications, which are widely used in fields such as autonomous driving, robot navigation and 3D reconstruction. However, existing stereo matching algorithms have obvious shortcomings when dealing with complex scenes. Traditional cost calculation usually relies on the grayscale value, color value or gradient information of the local window. This local method shows great limitations when facing long-distance dependencies, especially in complex scenes such as texture similarity, illumination changes and occlusion. The matching accuracy is not high. In addition, the traditional semi-global matching (SGM) algorithm has poor aggregation effect when aggregating costs in scenes of different complexities, especially when dealing with drastic parallax changes or continuous areas, resulting in low matching accuracy, which is not conducive to improving displacement detection accuracy and improving computational efficiency. Summary of the invention
[0003] In view of this, the embodiments of the present application provide a displacement detection method, device, equipment and program product to solve the problem in the prior art that the matching accuracy is not high, which is not conducive to improving the displacement detection accuracy.
[0004] A first aspect of an embodiment of the present application provides a displacement detection method, the method comprising:
[0005] Acquire a current image, where the current image includes a first image and a second image captured by a binocular camera;
[0006] capturing self-attention scores between pixels in the first image and the second image and other pixels in the image through a self-attention mechanism, and determining cost values of pixels in the first image and the second image according to the attention scores;
[0007] Determining the number of aggregation paths for each pixel according to the local structural complexity of the first image or the second image, performing cost aggregation according to the number of aggregation paths, and determining an optimal cost;
[0008] A depth displacement of the current image relative to a predetermined reference image is determined according to the optimal cost.
[0009] In combination with the first aspect, in a first possible implementation manner of the first aspect, after acquiring the current image, the method further includes:
[0010] Downsampling the current image and a predetermined reference image to obtain a plurality of sampled image pairs with different resolutions;
[0011] In order from low to high resolution, the plane displacement of the sampled image pair is estimated by a pre-trained displacement prediction model, and the plane displacement of the sampled image pair of the next resolution is corrected to obtain the plane displacement result after layer-by-layer iteration.
[0012] In combination with the first possible implementation manner of the first aspect, in a second possible implementation manner of the first aspect, in order of resolution from low to high, the plane displacement of the sampled image pair is estimated by a pre-trained displacement prediction model, and the plane displacement of the sampled image pair of the next resolution is corrected to obtain a plane displacement result after layer-by-layer iteration, including:
[0013] When the number of sampled images with different resolutions is L, the first plane displacement vector of the first sampled image pair with the lowest resolution is estimated by using a pre-trained displacement prediction model;
[0014] estimating the plane displacement of the second sampling image pair by using the pre-trained displacement prediction model, and correcting the plane displacement of the second sampling image pair by using the first plane displacement vector to obtain a first plane displacement vector;
[0015] The plane displacement of the third sampling image pair is estimated by the pre-trained displacement prediction model, and the plane displacement of the third sampling image pair is corrected by the first plane displacement vector to obtain the third plane displacement vector, until the Lth plane displacement vector is obtained by iterating layer by layer, and the Lth plane displacement vector is the plane displacement result.
[0016] In combination with the first possible implementation manner of the first aspect, in a third possible implementation manner of the first aspect, estimating the plane displacement of the sampled image pair by using a pre-trained displacement prediction model includes:
[0017] estimating an initial displacement value of each pixel in the sampled image pair by using a pre-trained displacement prediction model;
[0018] The plane displacement of the sampled image pair is determined according to the initial displacement value of each pixel in combination with the least square method.
[0019] In combination with the first aspect, in a fourth possible implementation manner of the first aspect, capturing self-attention scores between pixels in the first image and the second image and other pixels in the image through a self-attention mechanism, and determining cost values of the pixels in the first image and the second image according to the attention scores includes:
[0020] Performing feature extraction on the first image and the second image to determine first feature matrices of the first image and the second image;
[0021] Introducing a self-attention mechanism into the first feature matrix to obtain a self-attention score between pixels of the first image and the second image;
[0022] Update the first feature matrix according to the attention score to obtain a second feature matrix;
[0023] Cost values of pixels in the first image and the second image are determined based on the second feature matrix.
[0024] In combination with the fourth possible implementation manner of the first aspect, in a fifth possible implementation manner of the first aspect, introducing a self-attention mechanism in the first feature matrix to obtain a self-attention score between pixels of the first image and the second image includes:
[0025] Determine the query matrix, key matrix and value matrix through the pre-trained weight matrix;
[0026] determining a self-attention score between pixels of the first image and the second image based on the query matrix and the key matrix;
[0027] Updating the first feature matrix according to the attention score to obtain a second feature matrix includes:
[0028] The second feature matrix is determined according to the attention score and the value matrix.
[0029] In combination with the first aspect, in a sixth possible implementation manner of the first aspect, determining the number of aggregation paths for each pixel according to the local structural complexity of the first image or the second image, performing cost aggregation according to the number of aggregation paths, and determining the optimal cost includes:
[0030] determining a local structural complexity of each pixel in the first image or the second image;
[0031] According to a preset correspondence between local structure complexity and the number of paths, the number of paths corresponding to the local structure complexity of each pixel is determined;
[0032] Costs are accumulated according to the determined number of paths to determine the optimal cost.
[0033] A second aspect of an embodiment of the present application provides a displacement detection device, the device comprising:
[0034] An image acquisition unit, used to acquire a current image, wherein the current image includes a first image and a second image acquired by a binocular camera;
[0035] a cost value determining unit, configured to capture self-attention scores between pixels in the first image and the second image and other pixels in the image through a self-attention mechanism, and determine cost values of pixels in the first image and the second image according to the attention scores;
[0036] An optimal cost determination unit, configured to determine the number of aggregation paths for each pixel according to the local structural complexity of the first image or the second image, perform cost aggregation according to the number of aggregation paths, and determine an optimal cost;
[0037] A depth displacement determining unit is used to determine the depth displacement of the current image relative to a predetermined reference image according to the optimal cost.
[0038] A third aspect of an embodiment of the present application provides a position detection device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the position detection device implements a method as described in any one of the first aspects.
[0039] A fourth aspect of the embodiments of the present application provides a computer program product, which, when executed on a computer, enables the computer to execute the method in the first aspect or its various implementations.
[0040] A fifth aspect of an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method described in any one of the first aspects are implemented.
[0041] The sixth aspect of the embodiment of the present application provides a chip for implementing the methods in each implementation of the first aspect. Specifically, the chip includes: a processor for calling and running a computer program from a memory, so that a device equipped with the chip executes the method in the first aspect or its implementation.
[0042] Compared with the prior art, the embodiments of the present application have the following beneficial effects: the embodiments of the present application capture the self-attention score between the pixels of the first image and the second image captured by the binocular camera through the self-attention mechanism, determine the cost value of the pixel according to the attention score, and determine the number of aggregated paths for each pixel according to the complexity of the local structure, and perform cost aggregation according to the number of aggregated paths to determine the optimal cost, so that the depth displacement can be determined according to the optimal cost. Since this method can effectively capture global correlation, improve the matching accuracy in complex scenes, and dynamically adjust the number of paths according to the local structure of the image, it can significantly reduce computational redundancy and improve computational efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0044] Figure 1 It is a schematic diagram of an implementation flow of a position detection method provided in an embodiment of the present application;
[0045] Figure 2 This is a schematic diagram of an implementation process of determining a cost value provided by an embodiment of the present application;
[0046] Figure 3 It is a schematic diagram of an implementation flow of a displacement prediction model provided in an embodiment of the present application;
[0047] Figure 4 This is a schematic diagram of a feature extraction implementation process provided by an embodiment of the present application;
[0048] Figure 5 is a schematic diagram of a position detection device provided in an embodiment of the present application;
[0049] Figure 6 It is a schematic diagram of a position detection device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0050] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.
[0051] In order to illustrate the technical solution described in this application, a specific embodiment is provided below for illustration.
[0052] In the field of computer vision, deep displacement calculation and stereo matching are key technical applications, widely used in fields such as autonomous driving, robot navigation and 3D reconstruction. However, existing stereo matching algorithms have obvious shortcomings when dealing with complex scenes. For example, traditional cost calculation usually relies on the grayscale value, color value or gradient information of the local window. This local method will show great limitations when facing long-distance dependencies, especially in complex scenes such as texture similarity, illumination changes and occlusion. The matching accuracy is often unsatisfactory. In addition, the traditional semi-global matching (SGM) algorithm uses a fixed number of paths when aggregating costs, and cannot be adaptively adjusted according to the complexity of the scene, resulting in low computational efficiency and poor aggregation effect, especially when dealing with drastic parallax changes or continuous areas.
[0053] In terms of plane displacement calculation, existing traditional methods mainly rely on manual features and local optimization, which are computationally complex and inefficient, and are easily affected by image noise and complex deformation. Although deep learning technology is gradually applied to displacement prediction, existing deep learning methods usually predict local areas, ignoring the multi-scale and global information of image features, resulting in the accuracy of displacement estimation still being difficult to meet high-demand application scenarios.
[0054] Based on the above problems, the embodiment of the present application enhances the global correlation between pixels in the cost calculation by introducing the self-attention mechanism, thereby improving the matching accuracy in complex scenes. By combining the dynamic path adaptive aggregation method, the aggregation path is adaptively adjusted according to the local structure of the image, which can effectively reduce redundant calculations and improve computing efficiency. In addition, the embodiment of the present application integrates the displacement prediction model of deep learning and the multi-scale adaptive matching algorithm in the plane displacement calculation, and realizes high-precision and efficient displacement calculation by combining multi-level features and global optimization iterations. These improvements not only make up for the shortcomings of the above-mentioned technologies, but also provide a new solution for deep calculations in complex scenes, which has high practical value and innovative significance.
[0055] Figure 1 A schematic diagram of a displacement detection method for implementing the present application is provided in an embodiment, as detailed below:
[0056] In S101, a current image is acquired, where the current image includes a first image and a second image captured by a binocular camera.
[0057] In the displacement detection process, the embodiment of the present application can collect an image of a target to be monitored in real time as a current image. The current image can include a first image and a second image collected by a binocular camera, and the depth of a pixel in the current image can be determined based on the pixel cost between the first image and the second image. The current image is not limited to an image collected in real time, and can also be any image that needs to be estimated for displacement.
[0058] The current image in the embodiment of the present application is used for comparative analysis with the reference image to determine the changes in the image, thereby detecting displacement, including information such as depth displacement and plane displacement. The reference image can be a pre-set image, such as an image acquired at the beginning of monitoring, or an image acquired at other time points.
[0059] In S102, self-attention scores between pixels in the first image and the second image and other pixels in the image are captured through a self-attention mechanism, and cost values of pixels in the first image and the second image are determined according to the attention scores.
[0060] Among them, the self-attention mechanism is used to enable each pixel to refer to global information when calculating the cost, thereby improving the accuracy of matching. The process of determining the cost value can be as follows Figure 2 As shown, including:
[0061] In S201, feature extraction is performed on the first image and the second image to determine first feature matrices of the first image and the second image.
[0062] In the embodiment of the present application, when performing feature extraction, a census transformation can be used to generate a binary descriptor by comparing the grayscale values of a pixel with its neighboring pixels. The calculation formula of the census transformation can be expressed as:
[0063]
[0064] Where p is the current pixel, N(p) is the neighborhood of pixel p in the first image, I(p) represents the grayscale value of pixel p, and q is a pixel index of the second image, which is associated with pixel p in the first image. Calculating I[·] is the indicator function. In this way, we get the first feature matrix of the first image and the second image, including F L and F R .
[0065] In S202, a self-attention mechanism is introduced into the first feature matrix to obtain a self-attention score between pixels of the first image and the second image.
[0066] Applying the self-attention mechanism to the feature matrix can be used to capture the global correlation between pixels. The embodiment of the present application can construct a query matrix, a key matrix, and a value matrix, such as a weight matrix that can be pre-trained and learned, including a query weight matrix W Q , value weight matrix W V and key weight matrix W K , generate the query matrix QL , key matrix K R Sum value matrix V R :
[0067] Q L =F L W Q ,K R =F R W K ,V R =F R W V
[0068] In order to avoid numerical instability caused by too large an inner product result, the embodiment of the present application can scale the inner product result, and the scaling factor can be the square root of the dimension of the key vector. By calculating the dot product of the query matrix and the key matrix, the self-attention score between pixels can be obtained, which is expressed as:
[0069]
[0070] Among them, d k is the dimension of the key vector, A pq is the self-attention score between pixel p and pixel q.
[0071] In a possible implementation, the embodiment of the present application can also normalize the attention score through the Softmax function, which is expressed as:
[0072]
[0073] Among them, A pq ' represents the similarity between pixel p in the first image and pixel q in the second image, that is, the normalized self-attention score.
[0074] In S203, the first feature matrix is updated according to the attention score to obtain a second feature matrix.
[0075] The second feature matrix of the second image can be obtained by R After attention weighting, we get:
[0076]
[0077] Among them, Z R (q) is the feature representation of pixel q in the second image, which can be obtained by the value matrix V of the second image R Similarly, each pixel p in the first image is weighted by the value matrix of the second image to obtain its updated second feature matrix, expressed as Z L (p).
[0078]
[0079] Through the self-attention mechanism, the pixels in the second image not only contain local information, but also combine the pixel features related to it in the global range.
[0080] In S204, cost values of pixels in the first image and the second image are determined according to the second feature matrix.
[0081] Finally, based on the updated second feature matrix Z L and Z R , calculate the cost value:
[0082] C(p,d)=||Z L (p)-Z R (pd)||
[0083] Where C(p,d) is the cost value of pixel p at disparity d.
[0084] By supplementing the global information of the self-attention mechanism, the cost calculation becomes more robust and can effectively match similar areas in complex scenes.
[0085] like Figure 3 The figure shows a schematic diagram of the process of generating a feature matrix. After the first image and the second image are transformed by Census, a first feature submatrix is generated. The self-attention score between the pixels in the first image and the second image is determined by the self-attention mechanism. The first feature matrix is updated by the self-attention mechanism to obtain the second feature matrix. The cost value of the pixel at different disparities can be calculated by the second feature matrix determined by the first image and the second image respectively.
[0086] In S103, the number of aggregation paths for each pixel is determined according to the local structural complexity of the first image or the second image, and cost aggregation is performed according to the number of aggregation paths to determine an optimal cost.
[0087] After obtaining the cost values of the pixels of the first image and the second image at different parallaxes, the embodiment of the present application may further perform cost aggregation through dynamic path selection to improve the efficiency of cost aggregation.
[0088] The traditional semi-global matching (SGM) algorithm accumulates the cost between pixels through a fixed number of paths (usually 8) during cost aggregation. However, fixed path aggregation is not very accurate when dealing with areas with drastic parallax changes or continuous areas, and redundant calculation paths increase the computational overhead. In order to improve the efficiency of cost aggregation, the embodiment of the present application proposes a dynamic path adaptive aggregation method based on the local structural characteristics of the image.
[0089] The basic idea of the dynamic path adaptive aggregation in the embodiment of the present application is to dynamically adjust the number of aggregation paths for each pixel according to the local structural complexity of the image to reduce redundant calculations and improve efficiency. Specifically, for each pixel p in the image (the first image or the second image), first calculate its local structural complexity S(p):
[0090]
[0091] in, represents the gradient of pixel p, and N(p) is the neighborhood of pixel p. The structural complexity S(p) is used to reflect the complexity of the area where the pixel is located. In one implementation, the number of paths dynamically adjusted according to the structural complexity value can be expressed as:
[0092]
[0093] Among them, T1 and T2 are preset structural complexity thresholds, and the number of paths r(p) increases with the increase of complexity. In complex areas, the number of paths increases to enhance the aggregation effect; while in simple areas, the number of paths is reduced to improve computational efficiency.
[0094] The cumulative formula of path cost can be expressed as:
[0095] L r (p,d)=C(p,d)+min(L r (pr,d),L r (pr,d±1)+P1,L r (pr,d±2)+P2)
[0096] Among them, L r (p, d) represents the accumulated cost of the path, C(p, d) is the cost value of pixel p at disparity d, P1 and P2 are penalty terms for small (disparity change amplitude is less than the preset amplitude threshold) and large (disparity change amplitude is greater than the preset amplitude threshold) disparity changes.
[0097] Compared with the aggregation method with a fixed number of paths, the selection of dynamic paths and adaptive control of the number of paths can avoid unnecessary redundant calculations and improve the aggregation effect. In areas with drastic parallax changes, more paths are used to enhance cost accumulation; in flat areas, fewer paths are used to reduce calculation time. Finally, the optimal cost value of the pixel is obtained by accumulating the costs of multiple paths:
[0098]
[0099] In S104, a depth displacement of the current image relative to a predetermined reference image is determined according to the optimal cost.
[0100] After determining the optimal cost value, the depth of each pixel in the first image or the second image can be determined according to the relationship between the optimal cost value and the depth. Based on the change in the depth of the current image and the depth of the reference image, the depth displacement of the current image can be determined.
[0101] The displacement detection method of the present application significantly improves the accuracy of cost calculation and the efficiency of cost aggregation by combining the self-attention mechanism and dynamic path adaptive aggregation. In cost calculation, the self-attention mechanism enables each pixel to use global information for matching, improving the matching accuracy in complex scenes; in cost aggregation, the dynamic path adaptive aggregation flexibly adjusts the number of paths according to the local structural characteristics of the image, reducing redundant calculations and enhancing the real-time performance of the algorithm. It has better performance when processing complex and non-continuous parallax scenes, and can significantly reduce the calculation time while maintaining high accuracy.
[0102] The displacement detection in the embodiment of the present application may also include plane displacement detection. The displacement prediction model of deep learning and multi-scale adaptive sub-area matching can be introduced to improve the traditional plane displacement calculation method. The image features are automatically learned through the convolutional neural network (CNN), the initial displacement value is predicted quickly and accurately, and the displacement estimation is optimized layer by layer using a multi-scale pyramid structure, thereby improving the overall calculation accuracy and speed.
[0103] In the training phase of the displacement prediction model, a large amount of image data needs to be prepared first. High-quality original images can be selected as the basic material. The basic material can be a public dataset (such as COCO, Image Net) or a self-owned image resource. Some known transformations can be applied to the original image, including translation, rotation, scaling, shearing, and adding noise. These transformations are used to simulate the displacement and deformation in real scenes. Finally, sub-regions are extracted from the original image and the transformed image, and the reference image sub-region S is extracted from them. ref and the current image subarea S cur These sub-regions are local areas extracted from the original reference image and the current image, and are usually N×N matrices containing grayscale or RGB values. These sub-regions reflect the displacement information between the current frame and the reference frame in the image.
[0104] The output of the displacement prediction model is the displacement vector p pre d , its specific form can be:
[0105] p pred =[Δx,Δy,θ,γ x ,γ y ,λ]
[0106] Among them, the predicted displacement vector p pred Can include translation (Δx, Δy), rotation (θ), shear (γ x ,γ y ) and scaling (λ), which are used as initial values for subsequent fine iterative optimization and can significantly improve iterative efficiency and accuracy.
[0107] like Figure 4 As shown in the figure, when the displacement prediction model is calculated, the input image sub-area S is firstly processed through the convolution layer. ref and S cur Perform feature extraction. The convolution operation extracts local features by sliding the convolution kernel (filter) on the input image.
[0108] The convolution operation formula can be expressed as follows:
[0109]
[0110] in: is the value of the feature map at position (i, j) after the l-th convolution. is the convolution kernel (weight matrix) of the lth layer. is the pixel value of the previous layer input, b (l) is the bias term. Through this formula, the convolutional layer extracts local edge, texture, and shape features.
[0111] Next is the pooling layer, which is used to reduce the dimension and downsample the feature map of the convolution output, retaining the main features while reducing the amount of calculation and preventing overfitting. The commonly used pooling method is maximum pooling, and the maximum pooling formula can be:
[0112] in: is the output value of the position (i, j) after pooling. K is the pooling window size (such as 2×2). The pooling layer retains the most significant features by filtering out the maximum value of the local area.
[0113] The pooled feature map is flattened into a one-dimensional vector and further processed by a fully connected layer. These vectors are combined through linear transformations and mapped to the final displacement output. The fully connected layer formula can be expressed as follows:
[0114] z=W*x+b
[0115] Where: x is the flattened feature vector, W is the weight matrix of the fully connected layer, and b is the bias term. The fully connected layer performs comprehensive processing on the features to generate feature representations for displacement estimation. After the convolutional layer and the fully connected layer, an activation function (such as ReLU, Sigmoid) is applied to the output to introduce nonlinearity and improve the model's ability to express complex relationships. The ReLU (RectifiedLinearUnit) formula can be expressed as follows:
[0116] f(x)=max(0,x)
[0117] The ReLU activation function sets negative values to zero and only retains positive values, which enhances the nonlinear mapping ability of the model. CNN optimizes weights through the back-propagation algorithm to minimize the loss function. The loss function measures the error between the displacement vector predicted by the model and the true displacement vector. Commonly used loss functions include normalized cross correlation (NCC), which can be expressed as follows:
[0118]
[0119] in, and are the average grayscale values of the reference image and the current image sub-area, respectively. By continuously iteratively adjusting the weights of the network, the displacement prediction model learns how to effectively predict plane displacement. In the plane displacement calculation, the displacement prediction model achieves efficient prediction from the input image sub-area to the plane displacement vector through deep feature extraction and nonlinear mapping of convolution, pooling, and fully connected layers. The convolution layer extracts local features, the pooling layer reduces the dimension and simplifies the features, and the fully connected layer integrates and outputs the displacement estimate. The loss function guides the model learning, so that the displacement prediction model can quickly respond to complex scene changes and provide accurate initial value estimation for subsequent multi-scale adaptive sub-area matching.
[0120] After the training of the displacement prediction model is completed, displacement prediction can be performed based on the trained displacement prediction model to improve the efficiency of plane displacement calculation.
[0121] In a further optimization method, the embodiment of the present application can downsample the current image and the reference image to obtain multiple sampled image pairs with different resolutions. Each sampled image pair includes a reference image and a current image, and the resolution of the reference image is the same as the resolution of the current image. For example, among multiple sampled image pairs with different resolutions obtained by downsampling, if the sampling rate of one sampled image pair is 1 / 2 of the original image, then the resolution of the reference image and the current image in the sampled image pair is 1 / 2 of the original image resolution.
[0122] The embodiment of the present application can perform downsampling at different resolutions to construct a pyramid-shaped image pair sequence. The reference image and the current image are downsampled to generate a multi-layer image pyramid. From bottom to top, the resolution of each layer of the image gradually decreases. The bottom layer of the pyramid is the original image, and the resolution decreases as it goes up. The pyramid construction formula can be:
[0123]
[0124] in, and The pyramid images of the reference image and the current image at the lth layer can be respectively. l=0,1,...,L, L is the number of pyramid layers, the bottom layer (l=0) is the original image, and the resolution is lower as it goes up. Matching starts at the highest layer of the pyramid (the lowest resolution layer), and the initial displacement value output by the displacement prediction model is used for preliminary displacement calculation. The least squares method is used to estimate the displacement, and the first plane displacement vector of the highest layer (the first sampled image pair) is obtained:
[0125]
[0126] where p L is the preliminary displacement estimate of the highest layer L, and (i′, j′) is the coordinate after transformation based on the current estimated displacement p. In the matching process of each layer, the displacement estimate of the previous layer can be used to compensate and adjust the current layer image, that is, to correct the current image in order to better match the sub-area of the next layer:
[0127]
[0128] Among them, p (l+1) is the displacement estimate from the previous layer (lower resolution layer), and the Warp function uses the estimated value p of the previous layer (l+1) For the current image Perform geometric transformations (such as translation, rotation, etc.). Then perform sub-area matching on the compensated image, update the displacement estimate of the current layer, and achieve higher-precision displacement estimation by refining the matching layer by layer:
[0129]
[0130] Among them, p lis the displacement estimate of the current layer l. This formula means that on the compensated image, the sum of squared pixel differences is minimized and the matching optimization is performed in combination with the estimate of the previous layer. By refining layer by layer from low resolution to high resolution, the plane displacement of the second sampled image pair can be obtained in turn, and the first plane displacement vector is corrected, and the plane displacement of the third sampled image pair is corrected to obtain the third plane displacement vector, until the Lth plane displacement vector is corrected. This optimization process can continuously adjust and optimize the displacement estimate, making the final displacement vector more accurate and effectively reducing noise and error accumulation. This method uses the optimization results of the previous layer at each layer, forming a process of gradually approaching the true displacement, which can significantly improve matching accuracy and computational robustness.
[0131] By gradually refining the displacement from high to low layers, the accurate plane displacement vector P0 of the original resolution layer is finally obtained. The displacement prediction model and multi-scale adaptive sub-area matching form a synergistic effect in the calculation of plane displacement, combining the fast prediction of deep learning and the precise adjustment of traditional optimization. The displacement prediction model quickly provides the initial displacement value by learning the features of the image sub-area. The output displacement vector contains information such as translation, rotation, and shear, but the initial estimation may have errors under complex deformation and noise interference. Multi-scale adaptive sub-area matching optimizes the initial value provided by the displacement prediction model in a multi-level and step-by-step manner by constructing an image pyramid and matching layer by layer from low resolution to high resolution, approximating the true displacement value, thereby significantly improving the matching accuracy. The displacement prediction model captures high-dimensional features such as edges, textures, and local patterns, but does not process global information well; multi-scale matching effectively supplements the adaptability of the displacement prediction model to global displacement changes by integrating global and local features at different scales, making the displacement estimation more comprehensive. The initial value provided by the displacement prediction model significantly reduces the search range and calculation time of multi-scale matching, speeds up the convergence of matching, and enables the calculation process to have both the rapid response capability of deep learning and the precise adjustment advantage of traditional optimization methods. The combination of the two reflects the fusion and innovation of deep learning and traditional optimization methods, and provides new ideas for displacement calculation in complex scenarios, which not only maintains the accuracy of calculation, but also greatly improves efficiency.
[0132] In a possible implementation, the embodiment of the present application can use the displacement obtained by multi-scale adaptive matching as the initial value, use the inverse synthetic Gauss-Newton method to perform optimization iteration, and further fine-tune the displacement. The basic formula for optimization iteration is:
[0133]
[0134] Among them, E(p) is the error function, E 2 (p) is the Hessian matrix, The gradient vector is used to gradually adjust the displacement parameters output by the model, and finally obtain a high-precision displacement estimate. This algorithm combines deep learning fast prediction and multi-scale layer-by-layer optimization, making the plane displacement calculation more efficient and accurate in complex scenarios.
[0135] The embodiment of the present application proposes an improved algorithm combining self-attention mechanism and dynamic path adaptive aggregation for the problem of deep displacement calculation and stereo matching in complex scenes, and introduces deep learning displacement prediction model and multi-scale adaptive matching technology in plane displacement calculation. Through the self-attention mechanism, the global correlation is effectively captured in the cost calculation, which improves the matching accuracy in complex scenes such as texture similarity and lighting changes; dynamic path adaptive aggregation dynamically adjusts the number of paths according to the local structure of the image, significantly reducing computational redundancy and improving computational efficiency.
[0136] In the plane displacement calculation, the displacement prediction model quickly predicts the initial displacement value, and multi-scale matching optimizes the displacement estimation layer by layer, realizing efficient and accurate displacement calculation. The improved method of this application not only overcomes the shortcomings of the existing technology, but also provides a more robust and efficient solution for displacement calculation in complex and dynamic scenes, which has important theoretical significance and practical application value.
[0137] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0138] Figure 5 A schematic diagram of a displacement prediction device provided in an embodiment of the present application, the device comprising:
[0139] An image acquisition unit 501 is used to acquire a current image, where the current image includes a first image and a second image captured by a binocular camera;
[0140] A cost value determining unit 502 is used to capture self-attention scores between pixels in the first image and the second image and other pixels in the image through a self-attention mechanism, and determine cost values of pixels in the first image and the second image according to the attention scores;
[0141] An optimal cost determination unit 503 is used to determine the number of aggregation paths for each pixel according to the local structural complexity of the first image or the second image, perform cost aggregation according to the number of aggregation paths, and determine the optimal cost;
[0142] The depth displacement determining unit 504 is configured to determine the depth displacement of the current image relative to a predetermined reference image according to the optimal cost.
[0143] The displacement prediction device corresponds to the above-mentioned displacement prediction method.
[0144] Figure 6 Schematic diagram of a position detection device provided in an embodiment of the present application. Figure 6 As shown, the position detection device 6 of this embodiment includes: a processor 60, a memory 61, and a computer program 62 stored in the memory 61 and executable on the processor 60, such as a displacement prediction program. When the processor 60 executes the computer program 62, the steps in the above-mentioned displacement prediction method embodiments are implemented. Alternatively, when the processor 60 executes the computer program 62, the functions of the modules / units in the above-mentioned device embodiments are implemented.
[0145] Exemplarily, the computer program 62 may be divided into one or more modules / units, which are stored in the memory 61 and executed by the processor 60 to complete the present application. The one or more modules / units may be a series of computer program instruction segments capable of completing specific functions, which are used to describe the execution process of the computer program 62 in the position detection device 6.
[0146] The position detection device 6 can be a computing device such as a desktop computer, a notebook, a PDA, or a cloud server. The position detection device can include, but is not limited to, a processor 60 and a memory 61. Those skilled in the art can understand that Figure 6 It is only an example of a position detection device 6 and does not constitute a limitation of the position detection device 6. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the position detection device may also include input and output devices, network access devices, buses, etc.
[0147] The processor 60 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0148] The memory 61 may be an internal storage unit of the position detection device 6, such as a hard disk or memory of the position detection device 6. The memory 61 may also be an external storage device of the position detection device 6, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (SecureDigital, SD) card, a flash card (Flash Card), etc. equipped on the position detection device 6. Further, the memory 61 may also include both an internal storage unit and an external storage device of the position detection device 6. The memory 61 is used to store the computer program and other programs and data required by the position detection device. The memory 61 may also be used to temporarily store data that has been output or is to be output.
[0149] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.
[0150] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0151] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0152] In the embodiments provided in the present application, it should be understood that the disclosed devices / terminal equipment and methods can be implemented in other ways. For example, the device / terminal equipment embodiments described above are only schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0153] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0154] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0155] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by hardware related to computer program instructions. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.
[0156] In addition, an embodiment of the present application also provides a computer program product, which, when executed on a computer, enables the computer to execute the methods in the above-mentioned implementation modes.
[0157] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A displacement detection method, characterized in that: The method comprises: Acquire a current image, where the current image includes a first image and a second image captured by a binocular camera; capturing self-attention scores between pixels in the first image and the second image and other pixels in the image through a self-attention mechanism, and determining cost values of pixels in the first image and the second image according to the attention scores; Determining the number of aggregation paths for each pixel according to the local structural complexity of the first image or the second image, performing cost aggregation according to the number of aggregation paths, and determining an optimal cost; A depth displacement of the current image relative to a predetermined reference image is determined according to the optimal cost.
2. The method according to claim 1, characterized in that After acquiring the current image, the method further includes: Downsampling the current image and a predetermined reference image to obtain a plurality of sampled image pairs with different resolutions; In order from low to high resolution, the plane displacement of the sampled image pair is estimated by a pre-trained displacement prediction model, and the plane displacement of the sampled image pair of the next resolution is corrected to obtain the plane displacement result after layer-by-layer iteration.
3. The method according to claim 2, characterized in that In order from low to high resolution, the plane displacement of the sampled image pair is estimated by using a pre-trained displacement prediction model, and the plane displacement of the sampled image pair of the next resolution is corrected to obtain the plane displacement result after layer-by-layer iteration, including: When the number of sampled images with different resolutions is L, the first plane displacement vector of the first sampled image pair with the lowest resolution is estimated by using a pre-trained displacement prediction model; estimating the plane displacement of the second sampling image pair by using the pre-trained displacement prediction model, and correcting the plane displacement of the second sampling image pair by using the first plane displacement vector to obtain a first plane displacement vector; The plane displacement of the third sampling image pair is estimated by the pre-trained displacement prediction model, and the plane displacement of the third sampling image pair is corrected by the first plane displacement vector to obtain the third plane displacement vector, until the Lth plane displacement vector is obtained by iterating layer by layer, and the Lth plane displacement vector is the plane displacement result.
4. The method according to claim 2, characterized in that: The planar displacement of the sampled image pair is estimated by using a pre-trained displacement prediction model, including: estimating an initial displacement value of each pixel in the sampled image pair by using a pre-trained displacement prediction model; The plane displacement of the sampled image pair is determined according to the initial displacement value of each pixel in combination with the least square method.
5. The method according to claim 1, characterized in that Capturing self-attention scores between pixels in the first image and the second image and other pixels in the image through a self-attention mechanism, and determining cost values of pixels in the first image and the second image according to the attention scores, including: Performing feature extraction on the first image and the second image to determine first feature matrices of the first image and the second image; Introducing a self-attention mechanism into the first feature matrix to obtain a self-attention score between pixels of the first image and the second image; Update the first feature matrix according to the attention score to obtain a second feature matrix; Cost values of pixels in the first image and the second image are determined based on the second feature matrix.
6. The method according to claim 5, characterized in that Introducing a self-attention mechanism into the first feature matrix to obtain a self-attention score between pixels of the first image and the second image includes: Determine the query matrix, key matrix and value matrix through the pre-trained weight matrix; determining a self-attention score between pixels of the first image and the second image based on the query matrix and the key matrix; Updating the first feature matrix according to the attention score to obtain a second feature matrix includes: The second feature matrix is determined according to the attention score and the value matrix.
7. The method according to claim 1, characterized in that Determining the number of aggregation paths for each pixel according to the local structural complexity of the first image or the second image, performing cost aggregation according to the number of aggregation paths, and determining an optimal cost, including: determining a local structural complexity of each pixel in the first image or the second image; According to a preset correspondence between local structure complexity and the number of paths, the number of paths corresponding to the local structure complexity of each pixel is determined; Costs are accumulated according to the determined number of paths to determine the optimal cost.
8. A displacement detection device, characterized in that: The device comprises: An image acquisition unit, used to acquire a current image, wherein the current image includes a first image and a second image acquired by a binocular camera; a cost value determining unit, configured to capture self-attention scores between pixels in the first image and the second image and other pixels in the image through a self-attention mechanism, and determine cost values of pixels in the first image and the second image according to the attention scores; An optimal cost determination unit, configured to determine the number of aggregation paths for each pixel according to the local structural complexity of the first image or the second image, perform cost aggregation according to the number of aggregation paths, and determine an optimal cost; A depth displacement determining unit is used to determine the depth displacement of the current image relative to a predetermined reference image according to the optimal cost.
9. A position detection device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the position detection device implements the method according to any one of claims 1 to 7.
10. A computer program product comprising computer program instructions, characterized in that When the computer program is executed, the method according to any one of claims 1 to 7 is performed.
Citation Information
Patent Citations
Cost aggregation method and device and storage medium
CN113348483A
Binocular stereo matching method based on multi-scale feature extraction and adaptive aggregation
CN114742875A
Binocular depth estimation method and system based on attention mechanism and multilevel cost body
CN116258758A