A stereo matching method and network based on residual cost volume
By constructing multi-scale residual cost volumes and performing residual heterogeneous aggregation, the problem of information redundancy in the binocular stereo matching algorithm is solved, a balance between accuracy and speed is achieved, and the quality of stereo matching is improved.
Patent Information
- Application Number
- CN202310553305.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-16
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2043-05-16
AI Technical Summary
In existing binocular stereo matching algorithms, the 4D cost volume has a lot of redundant information and is slow, making it difficult to achieve a good balance between accuracy and speed.
Construct multiple residual cost volumes of different scales and dimensions, and use residual heterogeneous aggregation to perform cost aggregation, realize information interaction of polymorphic cost representation, and optimize the disparity map through feature pyramid and disparity regression.
It improves the accuracy and inference speed of the binocular stereo matching network, solves the information redundancy problem of the multi-scale cost volume network, and improves the quality of stereo matching.
Smart Images

Figure CN116681655B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of stereo matching technology, and in particular to a stereo matching method and network based on residual cost volume. Background Art
[0002] As the most popular depth estimation scheme, binocular stereo matching is indispensable in many real-world applications, including autonomous driving, robot navigation, 3D reconstruction, and augmented reality. The performance of the binocular stereo matching algorithm directly affects the final performance of the above technologies. Therefore, optimizing the stereo matching algorithm so that it can accurately estimate the target depth in a relatively short time is of great research significance.
[0003] With the rise of deep learning and neural networks, significant progress has been made in many computer vision problems that were difficult to solve with traditional methods. The classic stereo matching network is based on an end-to-end convolutional neural network framework and is divided into four modules: feature extraction, cost volume construction, cost aggregation, and disparity regression. Cost volume construction is the most significant step that distinguishes the stereo matching network from other computer vision tasks, and it also plays a decisive role in the speed and accuracy of the stereo matching network. However, the commonly used 4D cost volume has a lot of redundant information and is slow. It is difficult to achieve a good balance between accuracy and speed in subsequent processing based on this cost volume. Summary of the Invention
[0004] The present invention constructs the first residual cost volume of different scales and dimensions based on the extracted feature pyramid, and adopts the residual heterogeneous aggregation method to perform cost aggregation on the first residual cost volume, which can efficiently aggregate heterogeneous cost representations and realize information interaction of polymorphic cost representations, thereby solving the information redundancy problem of multi-scale cost volume network and enabling the binocular stereo matching network to achieve a better balance between accuracy and inference speed.
[0005] In a first aspect, an embodiment provides a stereo matching method based on residual cost volumes, characterized in that it includes: obtaining a left view and a right view to be matched; performing feature extraction on the left view and the right view to obtain a feature pyramid; constructing a plurality of first residual cost volumes of different scales based on the feature pyramid, each of the first residual cost volumes having a different dimension; performing cost aggregation on all the first residual cost volumes by residual isomerization aggregation to obtain a plurality of second cost volumes of different scales, the second cost volumes corresponding one-to-one to the first residual cost volumes and the corresponding first residual cost volumes and the second cost volumes having the same scale and dimension; performing disparity regression on each of the second cost volumes respectively to obtain a plurality of first disparity maps of different sizes; performing error correction based on all the first disparity maps to obtain a second disparity map having the same size as the left view and the right view.
[0006] In some embodiments, the residual heterogeneous aggregation method includes: performing inner-scale cost aggregation on each of the first residual cost volumes to obtain an inner-scale aggregated cost volume; and performing information fusion between cost volumes on all the inner-scale aggregated cost volumes to obtain multiple second cost volumes of different scales.
[0007] In some embodiments, information fusion is performed between cost volumes on all the inner-scale aggregated cost volumes, including: cross-scale cost aggregation operations at multiple scales, wherein: the cross-scale cost aggregation operation at each scale includes: sampling the inner-scale aggregated cost volumes of different scales to the same scale; performing information fusion between cost volumes on the inner-scale aggregated cost volumes of the same scale to obtain the second cost volume of one scale.
[0008] In some embodiments, the first residual cost volume includes: a 1 / 3 scale residual 3D cost volume, a 1 / 6 scale residual 4D cost volume, and a 1 / 12 scale regular 4D cost volume.
[0009] In some embodiments, the residual isomerization mode is represented by the following method: Where I is the identity transformation; The second cost volume is 1 / 3 of the scale; The second price volume is 1 / 6 scale, is the second cost volume of 1 / 12 scale, T is the forward residual cost volume transformation; is the reverse residual cost volume transformation; S q Refers to the squeeze operation, which is used to convert the 4D cost volume into a 3D cost volume; US q Refers to the unsqueeze operation, which is used to convert the 3D cost volume into a 4D cost volume; It is the 1 / 3 scale residual 3D cost volume; It is the 1 / 6 scale residual 4D cost volume; It is a 1 / 12 scale regular 4D value volume.
[0010] In some embodiments, an inner-scale cost aggregation is performed on each of the first residual cost volumes through two layers of convolution.
[0011] In some embodiments, performing disparity regression on each of the second cost volumes to obtain a plurality of first disparity maps of different sizes includes: performing disparity regression on the second cost volume of 1 / 3 scale to obtain the first disparity maps of the left view and the right view. Figure 1 / 3 of the first disparity map; perform disparity regression on the second cost volume of 1 / 6 scale to obtain the size of the left view and the right view Figure 1 / 6 first disparity map; perform disparity regression on the second cost volume of 1 / 12 scale to obtain the size of the left view and the right view Figure 1 / 12 first disparity map.
[0012] In some embodiments, performing error correction on all the first disparity maps to obtain a second disparity map having the same size as the left view and the right view includes: performing error correction on all the first disparity maps to obtain a second disparity map having the same size as the left view and the right view. Figure 1 / 12 is upsampled to obtain the left view and the right view with a size of Figure 1 / 6 of the first up-sampled disparity map; adding the first up-sampled disparity map to the first disparity map of the same size to obtain a first error correction map of a size of 1 / 6; up-sampling the first error correction map to obtain a first error correction map of a size of the left view and the right view Figure 1 / 3; adding the second up-sampled error map to the first disparity map of the same size to obtain a second error correction map with a size of 1 / 3; upsampling the second error correction map to obtain a second disparity map with the same size as the left view and the right view.
[0013] In the second aspect, the present invention provides a stereo matching network based on residual cost volume, comprising: a feature extraction module, a cost volume construction module, a cost aggregation module, a disparity regression module and a disparity optimization module; the feature extraction module is used to perform feature extraction on the acquired left view and right view to be matched to obtain a feature pyramid; the cost volume construction module is used to construct a plurality of first residual cost volumes of different scales according to the feature pyramid, and each first residual cost volume has a different dimension; the cost aggregation module is used to perform cost aggregation on all the first residual cost volumes by residual heterogeneous aggregation to obtain a plurality of second cost volumes of different scales, and the second cost volumes correspond one-to-one to the first residual cost volumes and the corresponding first residual cost volumes and the second cost volumes have the same scale and dimension; the disparity regression module is used to perform disparity regression on each of the second cost volumes respectively to obtain a plurality of first disparity maps of different sizes; the disparity optimization module is used to perform error correction based on all the first disparity maps to obtain a second disparity map with the same size as the left view and the right view.
[0014] In some embodiments, the cost aggregation module includes an intra-scale cost aggregation sub-module and a cross-modality cost aggregation sub-module; the intra-scale cost aggregation sub-module is used to perform intra-scale cost aggregation on each of the first residual cost volumes to obtain an intra-scale aggregated cost volume; the cross-modality cost aggregation sub-module is used to perform information fusion between cost volumes on all the intra-scale aggregated cost volumes to obtain multiple second cost volumes of different scales.
[0015] According to the method of the above embodiment, multiple residual cost volumes of different scales and dimensions are constructed based on the extracted feature pyramid, and information fusion of these residual cost volumes is performed using residual heterogeneous aggregation. This can efficiently aggregate heterogeneous cost representations and realize information interaction of polymorphic cost representations, thereby solving the information redundancy problem of the multi-scale cost volume network and enabling the binocular stereo matching network to achieve a better balance between accuracy and inference speed; error correction is performed based on multiple first disparity maps of different scales, which can effectively improve the quality of stereo matching. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 Flowchart of the stereo matching method based on residual cost volume provided by the present invention;
[0017] Figure 2 This is a flow chart of a residual isomerization method according to an embodiment;
[0018] Figure 3 A flowchart of a cross-scale cost aggregation operation for each scale according to an embodiment;
[0019] Figure 4 A flowchart of performing disparity regression on each second cost volume according to an embodiment;
[0020] Figure 5 A flowchart of performing disparity regression on each second cost volume respectively according to an embodiment;
[0021] Figure 6 A schematic diagram of the structure of a stereo matching network based on residual cost volume provided by the present invention;
[0022] Figure 7 A schematic diagram of the structure of a cost aggregation module according to an embodiment. DETAILED DESCRIPTION
[0023] The present invention will be further described in detail below by means of specific embodiments in conjunction with the accompanying drawings. Similar elements in different embodiments are numbered with associated similar elements. In the following embodiments, many detailed descriptions are provided to enable the present application to be better understood. However, those skilled in the art will readily appreciate that some of the features may be omitted in different circumstances, or may be replaced by other elements, materials, or methods. In some cases, some operations related to the present application are not shown or described in the specification. This is to avoid the core portion of the present application being overwhelmed by excessive descriptions, and for those skilled in the art, it is not necessary to describe these related operations in detail. They will fully understand the related operations based on the description in the specification and the general technical knowledge in the art.
[0024] In addition, the features, operations, or characteristics described in the specification may be combined in any appropriate manner to form various embodiments. Furthermore, the steps or actions in the method description may be reordered or adjusted in a manner readily apparent to those skilled in the art. Therefore, the various sequences in the specification and drawings are provided solely for the purpose of clearly describing a particular embodiment and are not intended to be mandatory, unless otherwise specified.
[0025] The serial numbers assigned to components herein, such as "first," "second," etc., are used solely to distinguish the objects being described and do not convey any sequential or technical meaning. References to "connection" and "coupling" herein, unless otherwise specified, include both direct and indirect connections (couplings).
[0026] In the stereo matching process, the construction of the cost volume is the most significant step that distinguishes the stereo matching network from other computer vision tasks. It also plays a decisive role in the speed and accuracy of the stereo matching network. However, the commonly used 4D cost volume has a lot of redundant information and is slow. Subsequent cost aggregation and other processing based on the cost volume can achieve a good balance between accuracy and speed.
[0027] In an embodiment of the present invention, multiple cost volumes of different scales and dimensions are constructed based on the extracted feature pyramid, and information fusion of these cost volumes is performed using residual heterogeneous aggregation. This can efficiently aggregate heterogeneous cost representations and realize information interaction of polymorphic cost representations, thereby solving the information redundancy problem of the multi-scale cost volume network and enabling the binocular stereo matching network to achieve a better balance between accuracy and inference speed.
[0028] Please refer to Figure 1 In one embodiment of the present invention, a stereo matching method based on residual cost volume is provided, comprising:
[0029] S10: Obtain the left view and right view to be matched.
[0030] S20: Extract features from the left view and the right view to obtain a feature pyramid.
[0031] S30: Construct multiple first residual cost volumes of different scales according to the feature pyramid, where each first residual cost volume has a different dimension.
[0032] In some embodiments, the first residual cost volume includes: a 1 / 3 scale residual 3D cost volume, a 1 / 6 scale residual 4D cost volume, and a 1 / 12 scale conventional 4D cost volume. The construction method of this first residual cost volume is to balance the parameter quantity of the residual cost volume and the information quality of the cost representation. Among them, the 1 / 3 scale often occupies a large proportion of parameters in the overall cost volume due to its high resolution. Selecting the 3D residual cost volume can effectively reduce its computational cost; the 1 / 6 scale and 1 / 12 scale consider the final disparity estimation performance of the network, and need to provide more accurate initial disparity and upper-level semantic features, thereby constructing a 1 / 6 scale residual 4D cost volume and a 1 / 12 scale conventional 4D cost volume.
[0033] S40: Cost aggregation is performed on all the first residual cost volumes by residual heterogeneous aggregation to obtain a plurality of second cost volumes of different scales, where the second cost volumes correspond one-to-one to the first residual cost volumes and the corresponding first residual cost volumes and the second cost volumes have the same scale and dimension.
[0034] In some embodiments, such as Figure 2 As shown, the isomerization modes of the residues include:
[0035] S41: Performing an inner-scale cost aggregation on each first residual cost volume to obtain an inner-scale aggregated cost volume.
[0036] In some embodiments, an inner-scale cost aggregation is performed on each of the first residual cost volumes through two layers of convolution.
[0037] S42: Perform cross-scale cost aggregation on all intra-scale aggregated cost volumes to obtain multiple second cost volumes of different scales.
[0038] In some embodiments, information fusion between cost volumes is performed on all intra-scale aggregated cost volumes, including: a cross-scale cost aggregation operation of multiple scales, wherein:
[0039] The cross-scale cost aggregation operation at each scale is as follows: Figure 3 Shown, including:
[0040] S420: Sampling the inner-scale aggregated cost volumes of different scales to the same scale;
[0041] S421: performing information fusion between the cost volumes of the inner-scale aggregated cost volume of the same scale to obtain the second cost volume of the same scale.
[0042] In some embodiments, the information fusion between cost volumes for all inner-scale aggregated cost volumes includes cross-scale cost aggregation operations of N scales. In each cross-scale cost aggregation operation of each scale, the inner-scale cost aggregation results of N scales are aggregated. In this embodiment, N is 3, and the 3 scales correspond to the three scales (1 / 3, 1 / 6, and 1 / 12) of the first residual cost volume, respectively.
[0043] In some embodiments, the residual isomerization mode is represented by the following:
[0044]
[0045] Where I is the identity transformation; The second cost volume is 1 / 3 of the size; The second price volume is 1 / 6 scale. is the second cost volume at 1 / 12 scale, and T is the positive residual cost volume transformation; is the reverse residual cost volume transformation; S q Refers to the squeeze operation, which is used to convert the 4D cost volume into a 3D cost volume, that is, to achieve dimensionality reduction. Specifically, the number of channels is first adjusted to 1 using the 3D convolution layer, and then the channel dimension is removed by the squeeze operation; US q Refers to the unsqueeze operation, which is used to convert the 3D cost volume into a 4D cost volume, that is, to achieve dimensionality increase. Correspondingly, the unsqueeze operation is first used to add channel dimensions to the 3D cost volume, and then the convolution layer is used to increase the number of channels to keep consistent with the corresponding 4D cost volume; It is the 1 / 3 scale residual 3D cost volume; It is the 1 / 6 scale residual 4D cost volume; It is a 1 / 12 scale regular 4D value volume.
[0046] By performing cost aggregation on all the constructed first residual cost volumes through residual heterogeneous aggregation, heterogeneous cost representations can be efficiently aggregated and information interaction of polymorphic cost representations can be realized, thereby solving the information redundancy problem of multi-scale cost volume networks and enabling binocular stereo matching networks to achieve a better balance between accuracy and inference speed.
[0047] S50: Perform disparity regression on each second cost volume to obtain multiple first disparity maps of different sizes.
[0048] In some embodiments, disparity regression is performed on each second cost volume to obtain multiple first disparity maps of different sizes, such as Figure 4 As shown, including:
[0049] S51: Perform disparity regression on the second cost volume of 1 / 3 scale to obtain the left view and right view sizes. Figure 1 / 3 first disparity map.
[0050] S52: Perform disparity regression on the second cost volume at 1 / 6 scale to obtain the left view and right view sizes. Figure 1 / 6 first disparity map.
[0051] S53: Perform disparity regression on the second cost volume at 1 / 12 scale to obtain the left view and right view sizes. Figure 1 / 12 first disparity map.
[0052] S60: Perform error correction according to all the first disparity maps to obtain a second disparity map with the same size as the left view and the right view.
[0053] In some embodiments, error correction is performed based on all first disparity maps to obtain a second disparity map with the same size as the left view and the right view, such as Figure 5 Shown, including:
[0054] S61: Dimensions for left and right views Figure 1 The first disparity map of / 12 is upsampled to obtain the size of the left view and the right view Figure 1 The first upsampled disparity map of / 6.
[0055] S62: Add the first up-sampled disparity map to the first disparity map of the same size to obtain a first error correction map of 1 / 6 size.
[0056] S63: Upsample the first error correction image to obtain left and right view sizes. Figure 1 The second upsampling error graph of / 3.
[0057] S64: Add the second up-sampled error map to the first disparity map of the same size to obtain a second error correction map of 1 / 3 the size.
[0058] S65: Up-sampling the second error correction image to obtain a second disparity image with the same size as the left view and the right view.
[0059] In some embodiments, the addition operation in step S62 and step S64 refers to adding the values of each corresponding pixel point in the two images, thereby achieving error correction of the first disparity map and obtaining a second disparity map of the same size as the left view and the right view. Error correction is performed based on multiple first disparity maps of different scales, which can effectively improve the quality of stereo matching.
[0060] Another embodiment of the present invention provides a stereo matching network based on residual cost volume, such as Figure 6As shown, it includes: a feature extraction module 10, a cost volume construction module 20, a cost aggregation module 30, a disparity regression module 40 and a disparity optimization module 50; the feature extraction module 10 is used to extract features from the acquired left view and right view to obtain a feature pyramid; the cost volume construction module 20 is used to construct a plurality of first residual cost volumes of different scales according to the feature pyramid, and each first residual cost volume has a different dimension; the cost aggregation module 30 is used to perform cost aggregation on all first residual cost volumes by residual heterogeneous aggregation to obtain a plurality of second cost volumes of different scales, and the second cost volumes correspond one-to-one to the first cost volumes and the corresponding first residual cost volumes and second cost volumes have the same scale and dimension; the disparity regression module 40 is used to perform disparity regression on each second cost volume respectively to obtain a plurality of first disparity maps of different sizes; the disparity optimization module 50 is used to perform error correction based on all first disparity maps to obtain a second disparity map with the same size as the left view and the right view.
[0061] In some embodiments, the residual isomerization mode is represented by the following:
[0062]
[0063] Where I is the identity transformation; The second cost volume is 1 / 3 of the size; The second price volume is 1 / 6 scale. is the second cost volume at 1 / 12 scale, and T is the positive residual cost volume transformation; is the reverse residual cost volume transformation; S q Refers to the squeeze operation, which is used to convert the 4D cost volume into a 3D cost volume; US q Refers to the unsqueeze operation, which is used to convert the 3D cost volume into a 4D cost volume; It is the 1 / 3 scale residual 3D cost volume; It is the 1 / 6 scale residual 4D cost volume; It is a 1 / 12 scale regular 4D value volume.
[0064] By performing cost aggregation on all constructed first residual cost volumes through the cost aggregation module 30, heterogeneous cost representations can be efficiently aggregated, and information interaction of polymorphic cost representations can be realized, thereby solving the information redundancy problem of the multi-scale cost volume network and enabling the binocular stereo matching network to achieve a better balance between accuracy and inference speed.
[0065] In some embodiments, such as Figure 7As shown, the cost aggregation module 30 includes an intra-scale cost aggregation sub-module 31 and a cross-modality cost aggregation sub-module 32; the intra-scale cost aggregation sub-module 31 is used to perform intra-scale cost aggregation on each first residual cost volume to obtain an intra-scale aggregated cost volume; the cross-modality cost aggregation sub-module 32 is used to perform information fusion between cost volumes on all intra-scale aggregated cost volumes to obtain multiple second cost volumes of different scales.
[0066] The implementation of the stereo matching network provided in this embodiment is the same as that of the aforementioned method and will not be described in detail here.
[0067] The above examples are used to illustrate the present invention, which are only used to help understand the present invention and are not intended to limit the present invention. Those skilled in the art can make several simple deductions, modifications or substitutions based on the concept of the present invention.
Claims
1. A stereo matching method based on residual cost volume, characterized in that: include: Get the left view and right view to be matched; Performing feature extraction on the left view and the right view to obtain a feature pyramid; According to the feature pyramid, a plurality of first residual cost volumes of different scales are constructed, each of which has a different dimension; wherein the first residual cost volume includes: a 1 / 3 scale residual 3D cost volume, a 1 / 6 scale residual 4D cost volume, and a 1 / 12 scale conventional 4D cost volume; Aggregating all the first residual cost volumes by residual heterogeneous aggregation to obtain a plurality of second cost volumes of different scales, where the second cost volumes correspond one-to-one to the first residual cost volumes and the corresponding first residual cost volumes and the second cost volumes have the same scale and dimension; Performing disparity regression on each of the second cost volumes to obtain a plurality of first disparity maps of different sizes; performing error correction based on all the first disparity maps to obtain a second disparity map having the same size as the left view and the right view; Wherein, the residual isomerization mode is represented by the following method: Where I is the identity transformation; The second cost volume is 1 / 3 of the scale; The second price volume is 1 / 6 scale, is the second cost volume of 1 / 12 scale, T is the forward residual cost volume transformation; is the reverse residual cost volume transformation; S q Refers to the squeeze operation, which is used to convert the 4D cost volume into a 3D cost volume; US q Refers to the unsqueeze operation, which is used to convert the 3D cost volume into a 4D cost volume; It is the 1 / 3 scale residual 3D cost volume; It is the 1 / 6 scale residual 4D cost volume; It is a 1 / 12 scale regular 4D value volume.
2. The method according to claim 1, wherein The residual isomerization polymerization method includes: Performing an inner-scale cost aggregation on each of the first residual cost volumes to obtain an inner-scale aggregated cost volume; Information fusion is performed between cost volumes on all the inner-scale aggregated cost volumes to obtain the second cost volumes of multiple different scales.
3. The method according to claim 2, wherein Performing information fusion between cost volumes on all the intra-scale aggregated cost volumes, including: cross-scale cost aggregation operations of multiple scales, wherein: The cross-scale cost aggregation operation at each scale includes: Sampling the inner-scale aggregated cost volumes of different scales to the same scale; The intra-scale aggregated cost volumes of the same scale are subjected to information fusion between the cost volumes to obtain the second cost volume of the same scale.
4. The method according to claim 3, wherein The inner scale cost is aggregated for each of the first residual cost volumes through two layers of convolution.
5. The method according to claim 1, wherein The performing disparity regression on each of the second cost volumes to obtain a plurality of first disparity maps of different sizes includes: Performing disparity regression on the second cost volume at a scale of 1 / 3 to obtain a first disparity map having a size of 1 / 3 of the left view and the right view; Performing disparity regression on the second cost volume at a scale of 1 / 6 to obtain a first disparity map having a size of 1 / 6 of the left view and the right view; Disparity regression is performed on the second cost volume at a scale of 1 / 12 to obtain a first disparity map with a size of 1 / 12 of the left view and the right view.
6. The method according to claim 5, wherein The performing error correction according to all the first disparity maps to obtain a second disparity map having the same size as the left view and the right view includes: Upsampling the first disparity map having a size of 1 / 12 of the left view and the right view to obtain a first upsampled disparity map having a size of 1 / 6 of the left view and the right view; Adding the first upsampled disparity map to the first disparity map of the same size to obtain a first error correction map with a size of 1 / 6; Upsampling the first error correction image to obtain a second upsampled error image with a size of 1 / 3 of the left view and the right view; Adding the second upsampled error map to the first disparity map of the same size to obtain a second error correction map of 1 / 3 the size; The second error correction map is upsampled to obtain a second disparity map with the same size as the left view and the right view.
7. A stereo matching network based on residual cost volume, characterized by: include: Feature extraction module, cost volume construction module, cost aggregation module, disparity regression module and disparity optimization module; The feature extraction module is used to extract features from the acquired left view and right view to be matched to obtain a feature pyramid; The cost volume construction module is used to construct a plurality of first residual cost volumes of different scales according to the feature pyramid, each of which has a different dimension; wherein the first residual cost volume includes: a 1 / 3 scale residual 3D cost volume, a 1 / 6 scale residual 4D cost volume, and a 1 / 12 scale conventional 4D cost volume; The cost aggregation module is used to perform cost aggregation on all the first residual cost volumes by residual heterogeneous aggregation to obtain a plurality of second cost volumes of different scales, where the second cost volumes correspond to the first residual cost volumes one-to-one and the corresponding first residual cost volumes and the second cost volumes have the same scale and dimension; The disparity regression module is used to perform disparity regression on each of the second cost volumes respectively to obtain a plurality of first disparity maps of different sizes; The disparity optimization module is configured to perform error correction based on all the first disparity maps to obtain a second disparity map with the same size as the left view and the right view; Wherein, the residual isomerization mode is represented by the following method: Where I is the identity transformation; The second cost volume is 1 / 3 of the scale; The second price volume is 1 / 6 scale, is the second cost volume of 1 / 12 scale, T is the forward residual cost volume transformation; is the reverse residual cost volume transformation; S q Refers to the squeeze operation, which is used to convert the 4D cost volume into a 3D cost volume; US q Refers to the unsqueeze operation, which is used to convert the 3D cost volume into a 4D cost volume; It is the 1 / 3 scale residual 3D cost volume; It is the 1 / 6 scale residual 4D cost volume; It is a 1 / 12 scale regular 4D value volume.
8. The stereo matching network according to claim 7, wherein: The cost aggregation module includes an intra-scale cost aggregation sub-module and a cross-modality cost aggregation sub-module; the intra-scale cost aggregation sub-module is used to perform intra-scale cost aggregation on each of the first residual cost volumes to obtain an intra-scale aggregated cost volume; the cross-modality cost aggregation sub-module is used to perform information fusion between cost volumes on all the intra-scale aggregated cost volumes to obtain multiple second cost volumes of different scales.
Citation Information
Patent Citations
Stereo matching method
CN111508013A
End-to-end stereo matching method based on convolutional neural network
CN111696148A