Deep learning dense matching method and system based on probabilistic error elimination
By introducing a parallax soft regression module based on probability error removal in the intensive matching method, the shortcomings of the traditional intensive matching method in terms of high precision and real-time requirements are solved, and higher matching accuracy and automation are achieved.
Patent Information
- Application Number
- CN202410982076.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-22
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-07-22
AI Technical Summary
The existing intensive matching methods are difficult to meet the needs of high precision and real-time in automatic measurement of surveying and mapping digital surface model (DSM), and the traditional methods are poorly robust, especially in repetitive textures, weak textures, non-textures and occlusion areas with low accuracy.
Using a deep learning intensive matching method based on probability error culling, a parallax soft regression module based on probability error culling is designed. By sorting and eliminating the parallax probabilities, the impact of small probability errors on the final parallax regression result is reduced.
It improves the accuracy of intensive matching, reduces the amount of manual modification, and improves the efficiency and accuracy of DSM automatic testing.
Smart Images

Figure CN119206265B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of dense matching technology, and in particular to a deep learning dense matching method and system based on probabilistic error elimination. Background Art
[0002] As the basic information of modern geospatial analysis, Digital Surface Model (DSM) is essential for understanding and interpreting the surface of the earth and the structure of objects on it. It provides indispensable three-dimensional spatial data support for urban planning, architectural design, environmental monitoring, agricultural optimization, disaster management and other fields by accurately reflecting the height information of natural terrain and artificial structures. The comprehensive application of DSM not only greatly improves the accuracy and efficiency of spatial data analysis, but also provides practical solutions for technical challenges such as wireless network design, new energy resource assessment, and aviation navigation safety. DSM can be calculated by distance information obtained by remote sensing and the position and attitude of the sensor. At present, there are airborne laser radar, dense matching and other methods to obtain distance information. Compared with the acquisition method of airborne laser radar, the method based on dense matching has the advantages of low cost and easy acquisition, so it is more widely used. The dense matching process constructs a disparity map by searching for the same-name points pixel by pixel in the optical image, and then calculates the distance information according to the sensor position and attitude. However, the dense matching process is computationally intensive, so the two-dimensional search is usually converted into a one-dimensional search using the kernel line constraint. The specific process is: first find the kernel line of the overlapping area image, then find the same-name point for each pixel on the left and right images along the kernel line direction, and count the coordinate difference of the same-name point as the disparity value of the pixel point. The disparity values of all pixels form a disparity map, and finally convert the disparity map into the required digital product based on the internal and external orientation elements. In the process of obtaining distance information, it is necessary to find the same-name point for each pixel in the left and right images, that is, the dense matching process is the most critical step, and its accuracy will directly affect the accuracy of subsequent applications. However, the existing dense matching methods are difficult to meet the needs of automatic measurement of surveying and mapping DSM to a certain extent, mainly including the following three aspects:
[0003] (1) The degree of automation is low and cannot meet the growing demand for surveying and mapping accuracy and real-time performance. The existing operation mode still relies on computer-assisted manual collection. Especially for high-precision scenes, manual collection is still required, which is inefficient, time-consuming and labor-intensive.
[0004] (2) Traditional dense matching methods use manually set feature descriptors, which are more dependent on professional domain knowledge and have poor robustness. Traditional methods use the same feature descriptors in the matching process, but the feature sets between different data sets may be different, and the same feature descriptors are likely to lead to poor robustness;
[0005] (3) Affected by various factors, the accuracy is low. In areas with repeated textures, weak textures, no textures, occlusions, and strong lighting, the matching accuracy of traditional methods is low, so a lot of post-processing is required, and these processes may reduce the final matching accuracy.
[0006] In the past five years, with the gradual advancement of the practical application of artificial intelligence, deep learning algorithms, as an important part of artificial intelligence, have achieved remarkable achievements in indoor and autonomous driving scenarios. Compared with traditional algorithms, deep learning methods are data-driven and have strong autonomous learning capabilities. This self-learning ability enables the algorithm to construct corresponding feature descriptors according to the tasks completed, and no longer requires manual design. Therefore, it can improve the degree of automation while making the algorithm more accurate and more robust. At present, deep learning dense matching networks are represented by two typical structures. One is a dense matching general network structure represented by DispNet. This type of network structure is similar to "U-Net" and is divided into two parts: feature extraction and resolution recovery. The reason why it is called a general type network is that the learning process relies on a large number of network parameters to fit the dense matching process, and there are fewer special designs for dense matching. In addition to dense matching tasks, it can also be used for other tasks such as classification, segmentation, and repair. Its specific application is determined by the type of data set and the corresponding label type. Although more parameters are used, the network structure is relatively simple and the calculation speed is fast; the other is a dense matching dedicated network structure represented by GCNet, such as PSMNet, GwcNet and other network structures. This type of network draws on the traditional dense matching idea to design the network structure. Its characteristics are that it has a special matching cost construction module, forms a probability map of each disparity through three-dimensional convolution, and then uses the disparity soft regression method to estimate the final disparity value. Its matching accuracy is much higher than that of the general network. However, this type of network has certain problems. The probability error at a distance far from the true disparity value will be amplified, and this error is positively correlated with the disparity regression distance, which leads to the deviation of the final disparity value. Summary of the invention
[0007] As the disparity regression distance increases, disturbances in tiny probability values may affect the final disparity regression results. Therefore, these tiny disturbances should be eliminated. To address this problem, the present invention proposes a deep learning intensive matching method and system based on probability error elimination, and designs a disparity soft regression module based on probability error elimination. By eliminating tiny probability errors, the interference of tiny probability errors on the final disparity regression value can be effectively reduced, thereby improving the matching accuracy.
[0008] In order to achieve the above object, the present invention adopts the following technical solutions:
[0009] A deep learning dense matching method based on probability error elimination, including:
[0010] Constructing a deep learning dense matching network;
[0011] In the above network structure, the disparity soft regression adopts a disparity soft regression module based on probability error elimination. First, the disparity probability of each pixel is sorted from large to small, the first k values of the disparity probability are retained, and the remaining probability values are set to zero to eliminate small probability errors.
[0012] The smooth L1 loss function is used as the target constraint function for network training and learning, and the trained network is used to perform dense matching on the target scene data.
[0013] According to the deep learning dense matching method based on probability error elimination of the present invention, further, the deep learning dense matching network includes feature extraction, pyramid pooling, multi-scale feature fusion, matching cost construction, disparity calculation and a disparity soft regression module based on probability error elimination; the entire network input is a left image and a right image of the epipolar line image, and feature extraction is performed respectively, and then multi-scale features are obtained through pyramid pooling, and then multi-scale feature fusion is performed on the left and right feature maps respectively; after extracting features respectively, a matching cost is constructed by translating along the epipolar line, and then disparity calculation is performed, and finally a left disparity map is constructed using a disparity soft regression module based on probability error elimination.
[0014] According to the deep learning dense matching method based on probabilistic error elimination of the present invention, further, feature extraction is used to extract features of each level of the image, and at the same time, multiple layers of dilated convolution are added; pyramid pooling and multi-scale feature fusion first perform multi-scale pooling on the feature extraction results respectively, and then perform 1×1 convolution respectively, and then upsample to the size before pooling and splice with the preliminary feature extraction results in the channel dimension, and finally pass through the multi-scale feature fusion layer and enter the matching cost construction.
[0015] According to the deep learning dense matching method based on probabilistic error elimination of the present invention, further, the matching cost construction process is: by keeping the left feature map unchanged, the right feature map is translated along the baseline direction according to the step size d, and then spliced with the left feature map and superimposed into the cost matrix, the excess part after translation is cut off, and the missing part is filled with 0, finally forming a four-dimensional tensor (channel, disparity, height, width) of the matching cost.
[0016] According to the deep learning dense matching method based on probability error elimination of the present invention, further, the disparity calculation is to calculate the constructed matching cost using three-dimensional convolution to form a matching cost probability tensor, that is, the probability of each pixel at each disparity value.
[0017] According to the deep learning dense matching method based on probability error elimination of the present invention, further, a disparity soft regression module based on probability error elimination is adopted, and the disparity soft regression process of a pixel (x, y) at any position is described as follows:
[0018]
[0019] Where d represents the current disparity value, P d Indicates the probability that the current pixel disparity value is d, maxdisp is the maximum disparity range of the match, < d is the effective indicator value of the current pixel at the disparity value d, and M is obtained by matching the cost probability tensor.
[0020] According to the deep learning dense matching method based on probability error elimination of the present invention, further, the calculation process of M is: first, a tensor M with a shape completely equal to the matching cost probability tensor is constructed, and all are initialized to 0, and then the matching cost probability tensors are sorted in order from large to small, and the indexes are recorded, and then the index values corresponding to the first k probability values are selected, and the tensor M is set to 1, and the effective indicator tensor M can be obtained; the construction process of the tensor M of the pixel (x, y) at any position is described as follows:
[0021] P′, Idx=F sort (P)
[0022]
[0023] Where d represents the current disparity value, F sort is a sorting function that sorts the disparity probability of each pixel in the disparity dimension, Idx d is the probability ranking value of the current disparity value d, P is the matching cost probability tensor, P′ is the sorted matching cost probability tensor, and k represents the ranking value.
[0024] According to the deep learning dense matching method based on probability error elimination of the present invention, further, the smooth L1 loss function is a smooth L1 loss between the true disparity value and the predicted disparity value.
[0025] A deep learning dense matching system based on probability error elimination is used to implement the deep learning dense matching method based on probability error elimination as mentioned above. The system comprises a model construction module, an improved parallax soft regression module and a loss function construction module, wherein:
[0026] Model building module for building deep learning dense matching networks;
[0027] An improved disparity soft regression module is used for disparity soft regression in the above network structure. A disparity soft regression module based on probability error elimination is adopted. First, the disparity probability of each pixel is sorted from large to small, the first k values of the disparity probability are retained, and the remaining probability values are set to zero to eliminate small probability errors.
[0028] The loss function construction module is used to adopt the smooth L1 loss function as the target constraint function for network training and learning, and use the trained network to perform dense matching on the target scene data.
[0029] Compared with the prior art, the present invention has the following advantages:
[0030] The core step of constructing a digital elevation model (DSM) using homologous remote sensing satellite images with overlapping areas is dense matching. The dense matching process actually searches for points of the same name pixel by pixel, calculates disparity values, and forms a disparity map. The DSM is then calculated using the position and attitude of the remote sensing satellite image when it was taken. The accuracy of the disparity value will directly affect the final accuracy of the DSM automatic measurement, and thus affect the amount of manual modification. The higher the accuracy, the fewer DSM grid points that do not meet the accuracy, and the less manual modification is required. Conversely, the more manual modification is required.
[0031] In order to improve the accuracy of dense matching and reduce the amount of manual modification, there have been many studies on dense matching methods, mainly including traditional methods and deep learning methods. Previous studies have shown that deep learning dense matching methods, starting from GCNet, to the later PSMNet and GwcNet, have surpassed the traditional SGBM matching method and reduced the amount of manual modification to a certain extent. These methods all use a disparity soft regression module at the end. This module multiplies the disparity probability and the corresponding disparity value to take the disparity soft regression value as the final disparity value. This method can accurately calculate the disparity value without probability error. However, the actual situation is that there is a small probability error at each disparity position. As the distance from the true disparity value increases, the error increases, thereby reducing the matching accuracy, such as Figure 1 (a) shows that the number of grid points that meet the accuracy requirements is reduced, the amount of final manual verification work is increased, and the production efficiency of DSM products is reduced. In the parallax soft regression operation, the present invention sorts the parallax probability of each pixel from large to small, retains the first k values of the parallax probability, and sets the remaining probability values to zero to eliminate small probability errors, thereby avoiding the influence of small probability values on the parallax regression accuracy and achieving the purpose of improving the matching accuracy. Figure 1(b) The present invention can also convert the obtained disparity map into distance information according to the sensor position and posture, and then convert it into elevation information for constructing a digital surface model, providing indispensable three-dimensional spatial data support for urban planning, architectural design, environmental monitoring, agricultural optimization, disaster management and other fields, and has good application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0033] Figure 1 1 is a schematic diagram of the principles of a traditional parallax soft regression and a parallax soft regression module based on probability error elimination according to an embodiment of the present invention;
[0034] Figure 2 It is a flowchart of a deep learning dense matching method based on probability error elimination according to an embodiment of the present invention;
[0035] Figure 3 is a schematic diagram of the deep learning dense matching network structure of an embodiment of the present invention;
[0036] Figure 4 Schematic diagram of the pyramid pooling network structure of an embodiment of the present invention;
[0037] Figure 5 is a schematic diagram of constructing a matching cost according to an embodiment of the present invention;
[0038] Figure 6 is a schematic diagram of a disparity calculation network structure according to an embodiment of the present invention;
[0039] Figure 7 is an image of the KITTI2012 dataset and a corresponding disparity map according to an embodiment of the present invention. DETAILED DESCRIPTION
[0040] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0041] The deep learning dense matching method based on probability error elimination in this embodiment is as follows: Figure 2 As shown, the following steps are included:
[0042] Step S101: construct a deep learning dense matching network.
[0043] Step S102: In the above network structure, the disparity soft regression adopts a disparity soft regression module based on probability error elimination. First, the disparity probability of each pixel is sorted from large to small, the first k values of the disparity probability are retained, and the remaining probability values are set to zero to eliminate small probability errors.
[0044] Step S103: adopt a smooth L1 loss function as the target constraint function for network training and learning, and use the trained network to perform dense matching on the target scene data.
[0045] The deep learning dense matching network includes feature extraction, pyramid pooling, multi-scale feature fusion, matching cost construction, disparity calculation and disparity soft regression module based on probability error elimination. The entire network input is the left image of the epipolar line image i L and right figure I R First, feature extraction is performed separately, and then multi-scale features are obtained through pyramid pooling. Then, multi-scale feature fusion is performed on the left and right feature maps respectively. The convolution weights used for feature extraction and multi-scale feature fusion are shared. After extracting features separately, the matching cost is constructed by translating along the kernel line, and then the disparity is calculated by three-dimensional convolution. Finally, the disparity soft regression module based on probability error elimination is used to construct the left disparity map. The entire network structure is as follows: Figure 3 As shown, each module is described below.
[0046] (1) Feature extraction module
[0047] Feature extraction is used to extract features at each level of the image. In order to expand the "field of view" of convolution, multiple layers of dilated convolution are added to ensure that the convolution kernel has sufficient "field of view" to increase the matching accuracy. Its network structure is shown in Table 1.
[0048] Table 1 Feature extraction module structure
[0049]
[0050]
[0051] This module is used as the input of the network. L and right figure I R As input, feature extraction is performed respectively, and the initial extraction results of left and right features are output Extraction results The process can be described as:
[0052]
[0053] In the formula, F Net-fe (·) is the feature extraction network, θ fe is the weight of the feature extraction network.
[0054] (2) Pyramid pooling and multi-scale feature fusion
[0055] The pyramid pooling module uses multi-scale information to give segmentation and positioning higher accuracy, which can solve the matching ambiguity caused by smooth areas, weak texture areas and repeated texture areas. When local information is difficult to determine the best matching position, it expands the field of view and adds global information as an aid to find the best matching point for the pixels in the area, ultimately achieving the purpose of improving accuracy.
[0056] This module first performs 64×64, 32×32, 16×16, and 8×8 pooling on the feature extraction results, and then performs 1×1 convolution to fully integrate the information between channels. It then upsamples to the size before pooling and concatenates it with the preliminary feature extraction results in the channel dimension. Finally, it passes through the multi-scale feature fusion layer and enters the matching cost construction. Its specific structure is as follows: Figure 4 shown.
[0057] (3) Matching cost construction
[0058] Matching cost construction is an important mark to distinguish general dense matching network from special dense matching network. This process is based on the tensor of the superposition of the two-pair feature maps in the disparity dimension from zero to the maximum disparity maxdisp of the baseline direction offset, which lays the foundation for the subsequent disparity calculation and the final disparity. The specific construction process is as follows Figure 5 As shown in the figure, by keeping the left feature map unchanged, the right feature map is translated along the baseline direction according to a certain step length d, and then concatenated with the left feature map and superimposed into the cost matrix. The extra part after translation is cut off, and the missing part is filled with 0, finally forming a four-dimensional tensor (channel, disparity, height, width) of the matching cost.
[0059] The matching cost construction increases the disparity dimension of the feature map, changing it from a three-dimensional tensor to a four-dimensional tensor. As the search range (maximum disparity maxdisp) increases, the required memory also increases exponentially.
[0060] (4) Parallax calculation
[0061] The essence of disparity calculation is to use three-dimensional convolution to calculate the constructed matching cost, forming a matching cost probability tensor, that is, the probability of each pixel at each disparity value. This process can not only utilize the relevant information in the two-dimensional direction of the plane, but also utilize the correlation between the disparity dimension information, so the overall accuracy is higher than the general network of dense matching. In the specific implementation, 6 convolution groups are connected in series, and each convolution block consists of two layers of 3×3×3 three-dimensional convolution layers. A skip-layer structure is set between blocks to increase multi-level information fusion. Its specific structure is as follows Figure 6 shown.
[0062] The matching cost tensor C is constructed as input and passed through the disparity calculation network F C-Net and the softmax function F softmax The matching cost probability tensor P is formed by processing, and the process can be described as:
[0063] P=F softmax (-F C-Net (C,θ c ))
[0064] In the formula, θ c To calculate the network F C-Net The convolution kernel weight takes a negative sign because the input is the cost and the output is the probability.
[0065] (5) Parallax soft regression module based on probability error elimination
[0066] After forming the probability of the pixel disparity value, the traditional method selects the position with the highest probability as the disparity value of the pixel, that is, selects the subscript of the maximum probability. This operation usually uses the argmax function, but because the process is not differentiable, it cannot be applied to the deep learning end-to-end network that requires backpropagation. Therefore, the disparity soft regression method is often used to convert the classification problem into a regression problem. This will amplify the probability error at a position far away from the true value of the disparity. Therefore, a disparity soft regression module based on probability error elimination is used to reduce the influence of the probability error on the distance between the disparity value and the true value. The disparity soft regression process of a pixel (x, y) at any position is described as follows:
[0067]
[0068] Where d represents the current disparity value, P d Indicates the probability that the current pixel disparity value is d, maxdisp is the maximum disparity range, M dis the effective indicator value of the current pixel at the disparity value d, and M is obtained by matching the cost probability tensor. The calculation process of M is as follows: first, a tensor M with the same shape as the matching cost probability tensor P is constructed, and all are initialized to 0, then the matching cost probability tensor P is sorted from large to small, and the index is recorded, and then the index value corresponding to the first k probability values is selected, and the tensor M is set to 1 to obtain the effective indicator tensor M; the construction process of the tensor M of the pixel (x, y) at any position is described as follows:
[0069] P′, Idx=F sort (P)
[0070]
[0071] Where d represents the current disparity value, F sort is a sorting function that sorts the disparity probability of each pixel in the disparity dimension, Idx d is the probability ranking value of the current disparity value d, P is the matching cost probability tensor, P′ is the sorted matching cost probability tensor, and k represents the ranking value.
[0072] (6) Loss Function
[0073] During the matching process, the choice of loss function has a direct impact on the model performance. The traditional L1 loss (also known as the absolute value loss) is sensitive to outliers, while the L2 loss (also known as the mean squared error loss) is more sensitive to outliers because it gives higher weights to larger errors. However, both loss functions have potential flaws in different situations. The L1 loss is not smooth near the zero point, which may lead to unstable updates when using gradient descent methods, while the L2 loss may cause the model to overreact to slight deviations due to excessive punishment of outliers. To overcome these limitations, this paper adopts the Smooth L1 Loss, which combines the advantages of both. It can reduce the overall impact of the mismatched area while maintaining the continuity of the loss function, specifically defined as the true disparity value D True The smooth L1 loss between the predicted disparity value D can be described by the following formula:
[0074] Loss = F smoothL1 (|DD True |)#()
[0075] In the formula, F smoothL1 is the smooth L1 loss, which is defined as follows:
[0076]
[0077] Based on the above method, this embodiment also proposes a deep learning dense matching system based on probability error elimination, which includes a model construction module, an improved parallax soft regression module and a loss function construction module, wherein:
[0078] Model building module for building deep learning dense matching networks.
[0079] An improved disparity soft regression module is used for disparity soft regression in the above network structure. A disparity soft regression module based on probability error elimination is adopted. First, the disparity probability of each pixel is sorted from large to small, the first k values of the disparity probability are retained, and the remaining probability values are set to zero to eliminate small probability errors.
[0080] The loss function construction module is used to adopt the smooth L1 loss function as the target constraint function for network training and learning, and use the trained network to perform dense matching on the target scene data.
[0081] In order to verify the effectiveness of this solution, the following is a further explanation based on the test data:
[0082] ①Dataset
[0083] The KITTI2012 dataset is a car driving scene, such as Figure 7 As shown in the figure, the image contains a large number of objects such as roads, cars, and houses. The data labels are calculated from the vehicle-mounted LiDAR data. Limited by the resolution of the LiDAR and the reflection reception, the reflection area is distributed in a point cloud shape. The area at infinity (sky) and outside the field of view is the default value, so the disparity label is semi-dense. The dataset contains 194 pairs of training images and 195 pairs of test images, with a size of 1226×370 pixels. It should be noted that the test set does not provide disparity annotations and requires online testing, so this article only uses 194 pairs of training images.
[0084] ②Software and hardware environment
[0085] In terms of software, the experiment was carried out under the Windows 10 operating system, and the virtual environment was created using Anaconda. The deep learning framework was Pytorch 1.13. In terms of hardware, the CPU was an Intel Xeon processor E5-2680v2 with a main frequency of 2.8 GHz, which could be turbo-accelerated to 3.4 GHZ. The memory was four 16 GB DDR3 generation memory sticks, totaling 64 GB with a frequency of 1333 MHz. The graphics card was two NVDIARTXA6000s with 10752 stream processors and 48 GB of video memory.
[0086] ③Evaluation indicators
[0087] The evaluation index uses two indicators commonly used in dense matching: absolute end point error (EPE) and p point error (pPE). Among them, p is 1, 2, 3, namely 1PE, 2PE and 3PE. These two indicators are complementary. The EPE indicator focuses more on the overall and is more sensitive to gross errors; 3PE pays more attention to whether the pixels meet the standards and is more sensitive to the number of points with errors greater than 3 pixels. Taking into account the actual production, this paper determines the optimal model with the minimum 3PE.
[0088] ④Algorithm selection and parameter setting
[0089] In terms of algorithm parameters, the learning rate is set to 1×10 -4 , the training rounds are set to 5000, the batchsize parameter is implemented by gradient accumulation simulation and is set to 8, the optimizer is Adam, β1 = 0.9, β2 = 0.999. Since the dense matching network has a large memory requirement and gradients need to be stored during training, the original image is randomly cropped into a 512 pixel × 256 pixel image during training. This operation can save memory on the one hand and enhance the dataset on the other.
[0090] ⑤ Experimental results
[0091] In the actual calculation process, since four indicators, EPE, 1PE, 2PE and 3PE, are used, and each indicator reaches the optimal value at different times during the training process, for a comprehensive comparison, the optimal results of each evaluation indicator are taken at the round where they are compared. The experimental results are shown in Table 2.
[0092] Table 2 Experimental results
[0093]
[0094] In summary, deep learning has developed rapidly in recent years and has surpassed traditional dense matching methods in the field of dense matching, forming two major types of networks: general and special dense matching networks. The special network structure design is more in line with dense matching tasks. However, when calculating the disparity value, the disparity probability value error of its disparity soft regression module will increase as the distance from the disparity to the true value increases, resulting in a deviation in the disparity value regression value. In response to this problem, the present invention designs a disparity soft regression module based on probability error, which can effectively eliminate the disturbance of small probability and thus improve the matching accuracy. Experiments were carried out on the KITTI2012 data set, and the results showed that the module can effectively reduce the EPE, 1PE, 2PE and 3PE indicators at the same time.
[0095] Unless otherwise specifically stated, the relative steps, numerical expressions and values of the components and steps set forth in these embodiments do not limit the scope of the present invention.
[0096] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0097] The units and method steps of each example described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. A person of ordinary skill in the art may use different methods to implement the described functions for each specific application, but such implementation is not considered to be beyond the scope of the present invention.
[0098] Those skilled in the art will appreciate that all or part of the steps in the above method can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium, such as a read-only memory, a disk or an optical disk. Optionally, all or part of the steps in the above embodiment can also be implemented using one or more integrated circuits, and accordingly, each module / unit in the above embodiment can be implemented in the form of hardware or in the form of software function modules. The present invention is not limited to any specific form of combination of hardware and software.
[0099] Finally, it should be noted that the above-described embodiments are only specific implementations of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The protection scope of the present invention is not limited thereto. Although the present invention is described in detail with reference to the above-described embodiments, ordinary technicians in the field should understand that any technician familiar with the technical field can still modify the technical solutions recorded in the above-described embodiments within the technical scope disclosed by the present invention, or can easily think of changes, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A deep learning dense matching method based on probability error elimination, characterized in that: Include: Step 1: Build a deep learning dense matching network; Step 2: In the above network structure, the disparity soft regression adopts a disparity soft regression module based on probability error elimination. First, the disparity probability of each pixel is sorted from large to small, the first k values of the disparity probability are retained, and the remaining probability values are set to zero to eliminate small probability errors; specifically, it includes: The disparity soft regression process of any pixel (x, y) is described as follows: Where d represents the current disparity value, P d Indicates the probability that the current pixel disparity value is d, maxdisp is the maximum disparity range, M d is the effective indicator value of the current pixel at the disparity value d, and M is obtained by matching the cost probability tensor; the calculation process of M is: first construct a tensor M with the same shape as the matching cost probability tensor, and initialize all to 0, then sort the matching cost probability tensor from large to small, and record the index, then select the index value corresponding to the first k probability values, set the tensor M to 1, and you can get the effective indicator tensor M; the construction process of the tensor M of the pixel (x, y) at any position is described as follows: P′,Idx=F sort (P) Where d represents the current disparity value, F sort is a sorting function that sorts the disparity probability of each pixel in the disparity dimension, Idx d is the probability ranking value of the current disparity value d, P is the matching cost probability tensor, P′ is the sorted matching cost probability tensor, and k represents the ranking value; Step 3: Use the smooth L1 loss function as the target constraint function for network training and use the trained network to perform dense matching on the target scene data.
2. The deep learning dense matching method based on probability error elimination according to claim 1 is characterized in that: The deep learning dense matching network includes feature extraction, pyramid pooling, multi-scale feature fusion, matching cost construction, disparity calculation and disparity soft regression module based on probability error elimination; the entire network input is the left and right images of the epipolar line image, and feature extraction is performed separately, and then multi-scale features are obtained through pyramid pooling, and then multi-scale feature fusion is performed on the left and right feature maps respectively; after extracting the features separately, the matching cost is constructed by translating along the epipolar line, and then the disparity is calculated, and finally the left disparity map is constructed using the disparity soft regression module based on probability error elimination.
3. The deep learning dense matching method based on probability error elimination according to claim 2 is characterized in that: Feature extraction is used to extract features at all levels of the image, while adding multiple layers of dilated convolutions; pyramid pooling and multi-scale feature fusion first perform multi-scale pooling on the feature extraction results, then perform 1×1 convolution on them, and then upsample them to the size before pooling and concatenate them with the preliminary feature extraction results in the channel dimension, and finally pass through the multi-scale feature fusion layer and enter the matching cost construction.
4. The deep learning dense matching method based on probability error elimination according to claim 2 is characterized in that: The matching cost construction process is as follows: by keeping the left feature map unchanged, the right feature map is translated along the baseline direction according to the step size d, and then it is spliced with the left feature map and superimposed into the cost matrix. The excess part after translation is cut off and the missing part is filled with 0, finally forming a four-dimensional tensor (channel, disparity, height, width) of the matching cost.
5. The deep learning dense matching method based on probability error elimination according to claim 4 is characterized in that: Disparity calculation uses three-dimensional convolution to calculate the constructed matching cost to form a matching cost probability tensor, that is, the probability of each pixel at each disparity value.
6. The deep learning dense matching method based on probability error elimination according to claim 1 is characterized in that: The smooth L1 loss function is a smooth L1 loss between the true disparity value and the predicted disparity value.
7. A deep learning dense matching system based on probabilistic error elimination, characterized in that: The system is used to implement the deep learning dense matching method based on probability error elimination as described in any one of claims 1 to 6, comprising a model construction module, an improved parallax soft regression module and a loss function construction module, wherein: Model building module for building deep learning dense matching networks; The improved disparity soft regression module is used for disparity soft regression in the above network structure. The disparity soft regression module based on probability error elimination is adopted. First, the disparity probability of each pixel is sorted from large to small, the first k values of the disparity probability are retained, and the remaining probability values are set to zero to eliminate small probability errors. Specifically, it includes: The disparity soft regression process of any pixel (x, y) is described as follows: Where d represents the current disparity value, P d Indicates the probability that the current pixel disparity value is d, maxdisp is the maximum disparity range, M d is the effective indicator value of the current pixel at the disparity value d, and M is obtained by matching the cost probability tensor; the calculation process of M is as follows: first, a tensor M with the same shape as the matching cost probability tensor is constructed, and all are initialized to 0, and then the matching cost probability tensor is sorted from large to small, and the index is recorded, and then the index value corresponding to the first k probability values is selected, and the tensor M is set to 1 to obtain the effective indicator tensor M; the construction process of the tensor M of the pixel (x, y) at any position is described as follows: P′,Idx=F sort (P) Where d represents the current disparity value, F sort is a sorting function that sorts the disparity probability of each pixel in the disparity dimension, Idx d is the probability ranking value of the current disparity value d, P is the matching cost probability tensor, P′ is the sorted matching cost probability tensor, and k represents the ranking value; The loss function construction module is used to adopt the smooth L1 loss function as the target constraint function for network training and learning, and use the trained network to perform dense matching on the target scene data.
8. A computer device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Deep learning semi-supervised dense matching method and system based on consistency constraint
CN113780389A
Intensive matching optimization method and device based on frame-level twin structure
CN118297867A