Binocular stereo matching method and system
By adopting grouped convolution backbone and redesigned SPPF network in binocular stereo matching technology, the problem of inefficiency in feature extraction and loop matching in the prior art is solved, and a more efficient matching effect and a simplified deployment process is achieved.
Patent Information
- Application Number
- CN202510226632.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-10
AI Technical Summary
The existing binocular stereo matching technology has inefficiency and computational complexity problems in the process of extracting features and loop matching, resulting in increased difficulty in network deployment on the end side and reduced operational efficiency.
Grouped convolution backbone is used to extract features, obtain higher semantic features through multiple downsampling, avoid false positive problems caused by repeated textures, and use the redesigned SPPF network to replace common loop matching modules, improving the matching effect and operation efficiency of the model.
It improves the matching effect and operation efficiency of binocular stereo matching, avoids false positive problems caused by repeated textures, and simplifies the difficulty of network deployment on the end side.
Smart Images

Figure CN120125858A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and specifically to a binocular stereo matching method and system. Background Art
[0002] In the prior art, binocular stereo matching has very wide applications in the fields of unmanned aerial vehicles and robots. Currently, when performing binocular matching, two backbone networks are mostly used to separately extract the features of the left-eye image and the right-eye image, and then the stereo matching similarity is calculated through multiple training loops. Then, a neural network is used to generate a disparity map corresponding to the left-eye image and the right-eye image according to the stereo matching similarity, thereby completing depth estimation.
[0003] However, the above process has the following problems:
[0004] 1. When extracting the features of the left-eye image and the right-eye image through two backbones respectively, the two extraction processes cannot be parallel, thereby reducing the running efficiency of the network at the edge side;
[0005] 2. Multiple loop matching increases the difficulty of network deployment at the edge side and also reduces the calculation efficiency. Summary of the Invention
[0006] Aiming at the deficiencies of the prior art, the present invention provides a binocular stereo matching method and system. It extracts features through a grouped convolution backbone, and obtains higher semantic features through multiple downsamplings, which can avoid false alarm problems caused by repeated textures. Finally, a redesigned SPPF network is used to replace the common loop matching module, improving the matching effect and running efficiency of the model.
[0007] To achieve the above object, the present invention provides the following technical solutions:
[0008] On the one hand, a binocular stereo matching method is provided, which includes the following steps:
[0009] Input both the left-eye image and the right-eye image into the backbone module, and the backbone module fuses the left-eye image and the right-eye image and extracts the fused image features;
[0010] Process the fused image features through the SPPF module to output a feature map;
[0011] Perform upsampling processing and feature aggregation processing on the feature map through the upsampling + concat + feature aggregation module to output a feature aggregation image;
[0012] Input the feature aggregation image into the head output layer head, and output a disparity map after convolution calculation to complete binocular stereo matching.
[0013] Preferably, before fusing the left-eye image and the right-eye image, preprocessing is performed on both the left-eye image and the right-eye image.
[0014] Preferably, the backbone module uses the RepVGG network structure to complete image feature extraction. The RepVGG network structure has a total of 5 layers, and the RepVGG network structure is used to complete the fusion of the left-eye image and the right-eye image and the extraction of the fused image features.
[0015] Preferably, using the RepVGG network structure to complete the fusion of the left-eye image and the right-eye image and the extraction of the fused image features includes the following steps:
[0016] Fuse the left-eye image and the right-eye image by the concat method to obtain a fused image;
[0017] Input the fused image into the first layer of the RepVGG network structure, and perform the first downsampling and convolution operation with the parameter settings of conv3×3, group = 2, stride = 2, and input the result of the first convolution operation into the second layer of the RepVGG network structure, and perform the second downsampling and convolution operation again with the parameter settings of conv3×3, group = 2, stride = 2 to obtain the convolution-processed image features;
[0018] Input the convolution-processed image features into the third layer, and after downsampling and convolution calculation, obtain the first fused image features;
[0019] Input the first fused image features into the fourth layer, and after downsampling and convolution calculation, obtain the second fused image features;
[0020] Input the second fused image features into the fifth layer, and after downsampling and convolution calculation, obtain the third fused image features.
[0021] Preferably, the fused image features are processed by the SPPF module to output feature maps, including the following steps:
[0022] Input the first fused image features, the second fused image features, and the third fused image features into an SPPF module correspondingly to output the first feature map, the second feature map, and the third feature map correspondingly.
[0023] Preferably, each SPPF module includes a first convolutional layer, 3 parallel MaxPool modules, a concat module, and a second convolutional layer;
[0024] The fused image features are input into the first convolutional layer. The convolution calculation results pass through 3 parallel MaxPool modules to obtain pooling results. Further, the 3 pooling results are concatenated through a concat module, and the concatenated result is subjected to feature mapping through the second convolutional layer to obtain a feature map.
[0025] Preferably, the second convolutional layer completes the convolution calculation through a CrossConv module.
[0026] Preferably, the feature map is subjected to upsampling processing and feature aggregation processing through an upsampling + concat + feature aggregation module to output a feature aggregation image. Specifically, it includes the following steps:
[0027] The third feature map is upsampled through a first upsampling + concat + feature aggregation module, and the upsampling result is concatenated with the second feature map, and the concatenated result is subjected to convolution calculation, and the convolution calculation result is sequentially processed through m repvggblock modules to obtain a first processing result;
[0028] The first processing result is upsampled through a second upsampling + concat + feature aggregation module, and the upsampling result is concatenated with the first feature map, and the concatenated result is subjected to convolution calculation, and the convolution calculation result is sequentially processed through m repvggblock modules to obtain a second processing result;
[0029] The second processing result is upsampled through a third upsampling + concat + feature aggregation module, and the upsampling result is concatenated with the convolutional processed image features, and the concatenated result is subjected to convolution calculation, and the convolution calculation result is sequentially processed through m repvggblock modules to obtain a third processing result as the feature aggregation image.
[0030] Preferably, when each repvggblock module performs calculations, group = 1 is taken.
[0031] On the other hand, a binocular stereo matching system is also provided, which includes:
[0032] A backbone module, which is used to fuse the left-eye image and the right-eye image and extract the fused image features;
[0033] Three SPPF modules, which are respectively used to input the first fused image features, the second fused image features, and the third fused image features into one SPPF module to correspondingly output the first feature map, the second feature map, and the third feature map;
[0034] The first upsampling + concat + feature aggregation module is used to upsample the third feature map, perform a concat operation on the upsampling result and the second feature map, perform convolutional calculations on the concat result, and sequentially process the convolutional calculation result through m RepVGG block modules to obtain a first processing result;
[0035] The second upsampling + concat + feature aggregation module is used to upsample the first processing result, perform a concat operation on the upsampling result and the first feature map, perform convolutional calculations on the concat result, and sequentially process the convolutional calculation result through m RepVGG block modules to obtain a second processing result;
[0036] The third upsampling + concat + feature aggregation module is used to upsample the second processing result, perform a concat operation on the upsampling result and the convolutional processed image features, perform convolutional calculations on the concat result, and sequentially process the convolutional calculation result through m RepVGG block modules to obtain a third processing result as the feature aggregation image;
[0037] The head output layer head is used to receive the feature aggregation image and output a disparity map after convolutional calculations to complete binocular stereo matching.
[0038] Compared with the prior art, the present application has the following beneficial effects:
[0039] The present application can extract features through a grouped convolutional backbone, which obtains higher semantic features through multiple downsamplings, can avoid false alarm problems caused by repeated textures, and finally redesigned the SPPF network using cross convolutions and replaced the common loop matching module with this network, improving the matching effect and running efficiency of the model. Description of the Drawings
[0040] Figure 1 It is a flowchart of the steps of the binocular stereo matching method in the present invention;
[0041] Figure 2 It is a structural schematic diagram of the backbone module in the present invention;
[0042] Figure 3 It is a structural schematic diagram of the 3rd - 5th layers of the backbone module in the present invention;
[0043] Figure 4 It is a structural schematic diagram of the SPPF module in the present invention;
[0044] Figure 5Schematic diagram of the upsampling + concat + feature aggregation module in the present invention;
[0045] Figure 6 Schematic diagram of the binocular stereo matching system in the present invention. Detailed implementation manners
[0046] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0047] Embodiment 1:
[0048] As Figure 1 shown, this embodiment provides a binocular stereo matching method, which includes the following steps:
[0049] S1. Obtain the left-eye image ImgL and the right-eye image ImgR; in this embodiment, the left-eye image and the right-eye image can be obtained by a binocular camera mounted on a device such as a drone, and both the left-eye image and the right-eye image come from the same scene;
[0050] And preprocess both the left-eye image ImgL and the right-eye image ImgR. The preprocessing includes flipping, cropping, size scaling, etc., so that the resolutions of the left-eye image ImgL and the right-eye image ImgR both meet the preset conditions. For example, in this embodiment, the resolutions of the preprocessed left-eye image ImgL and right-eye image ImgR can both be 800×576;
[0051] S2. Input both the left-eye image ImgL and the right-eye image ImgR into the backbone module, and fuse the left-eye image ImgL and the right-eye image ImgR through the backbone module, and extract the fused image features;
[0052] For example, in this embodiment, as Figure 2 shown, the backbone module uses the repvgg network structure to complete image feature extraction. The repvgg network structure has a total of 5 layers. Based on this, the repvgg network structure is used to complete the fusion of the left-eye image ImgL and the right-eye image ImgR and the extraction of the fused image features, which includes the following steps:
[0053] Fuse the left-eye image ImgL and the right-eye image ImgR in the concat manner to obtain a fused image. For example, in this embodiment, the preprocessed left-eye image ImgL and right-eye image ImgR with a resolution of 800×576 can be subjected to feature fusion through concat to obtain a 1×6×800×576 fused image;
[0054] Input the fused image into the first layer of the repvgg network structure, and perform the first downsampling and convolution operation with the parameter settings of conv3×3, group = 2, and stride = 2. Then input the result of the first convolution operation into the second layer of the repvgg network structure, and perform the second downsampling and convolution operation again with the parameter settings of conv3×3, group = 2, and stride = 2 to obtain the convolutional processed image features. For example, after the 1×6×800×576 fused image passes through the first layer and the second layer of the repvgg network structure in sequence, convolutional processed image features of 1×128×200×144 can be obtained;
[0055] Input the convolutional processed image features into the third layer, and after downsampling and convolution calculation, obtain the first fused image features. For example, fused image features with a size of 1×256×100×72;
[0056] Input the first fused image features into the fourth layer, and after downsampling and convolution calculation, obtain the second fused image features. For example, fused image features with a size of 1×512×50×36;
[0057] Input the second fused image features into the fifth layer, and after downsampling and convolution calculation, obtain the third fused image features. For example, fused image features with a size of 1×1024×25×18;
[0058] Among them, as Figure 3 shown, the structures of the third layer, the fourth layer, and the fifth layer are the same, and each includes a convolutional layer with stride = 2 and group = 2 and n repvggblock modules (the value range of n is 2 - 5). Moreover, the previous fused image features first pass through the convolutional layer to complete the convolution calculation, and the convolution calculation results are sequentially processed by n repvggblock modules to obtain the next fused image features;
[0059] S3. Process the fused image features through the SPPF module to output the feature map;
[0060] In this embodiment, different from other networks that explicitly match the left and right image features, as Figure 2 shown, in this embodiment, 3 SPPF modules are directly used to complete the operation to increase the receptive field while reducing the computational amount;
[0061] Further, as Figure 4 shown, each SPPF module includes a first convolutional layer 101, three parallel MaxPool modules 102, a concat module 103, and a second convolutional layer 104;
[0062] On this basis, the fused image features are processed by the SPPF module to output a feature map, including the following steps:
[0063] Input the first fused image feature, the second fused image feature, and the third fused image feature into an SPPF module correspondingly to output the first feature map, the second feature map, and the third feature map correspondingly.
[0064] Specifically, a fused image feature (i.e., one of the first fused image feature, the second fused image feature, and the third fused image feature) is input into the first convolutional layer 101, and the number of channels is reduced by half through a 1×1 convolution operation with group = 2. The convolution calculation result passes through three parallel MaxPool modules 102 to obtain corresponding pooling results. Further, the three pooling results are concatenated by the concat module 103, and the concatenated result is subjected to feature mapping through the second convolutional layer 104 (also with group = 2) to finally obtain the corresponding feature map;
[0065] Meanwhile, the second convolutional layer 104 in this embodiment completes the convolution calculation through the CrossConv module. Thus, by adding three SPPF modules to the 3rd, 4th, and 5th layers of the backbone module respectively, feature matching can be completed in different receptive fields, expanding the feature matching range. At the same time, in order to preserve more detailed information, the output of the 2nd layer network of the backbone module does not pass through the SPPF module because using convolution to replace the explicit feature matching module will cause the problem of too short matching distance;
[0066] In addition, different from the operation of only downsampling the image features twice in a general stereo matching network, this embodiment extracts semantic features at different levels through 5 times of downsampling, enabling the model to pay more attention to the overall features of the image and avoiding the problem of incorrect matching caused by the influence of repeated textures.
[0067] S4. Upsample and perform feature aggregation processing on the feature map through the upsampling + concat + feature aggregation module to output a feature aggregation image. Specifically, it includes the following steps:
[0068] The third feature map is upsampled by the first upsampling + concat + feature aggregation module 201, and the upsampling result is concatenated with the second feature map, and 1×1 convolution calculation is performed on the concatenation result, and the convolution calculation result is sequentially processed through m repvggblock modules (the value range of m is 2-5) to obtain a first processing result;
[0069] For example, as Figure 5 shown, the third feature map of 1×512×25×18 from the 5th layer is upsampled to become a feature map of 1×512×50×36, and then the 1×512×50×36 feature map is concatenated with the second feature map of 1×256×50×36 from the 4th layer. The concatenation result is then subjected to 1×1 convolution calculation to obtain a feature of 1×256×50×36, and then the 1×256×50×36 feature is sequentially processed through 3 repvggblock modules to obtain a first processing result;
[0070] The first processing result is upsampled by the second upsampling + concat + feature aggregation module 202, and the upsampling result is concatenated with the first feature map, and 1×1 convolution calculation is performed on the concatenation result, and the convolution calculation result is sequentially processed through m repvggblock modules (the value range of m is 2-5) to obtain a second processing result, and the obtaining method of the second processing result is the same as the obtaining method of the above first processing result;
[0071] The second processing result is upsampled by the third upsampling + concat + feature aggregation module 203, and the upsampling result is concatenated with the convolution-processed image feature, and 1×1 convolution calculation is performed on the concatenation result, and the convolution calculation result is sequentially processed through m repvggblock modules (the value range of m is 2-5) to obtain a third processing result as the feature aggregation image, and the obtaining method of the third processing result is the same as the obtaining method of the above first processing result;
[0072] In this embodiment, when each repvggblock module performs calculations, group = 1 is taken, that is, grouped convolution is no longer performed;
[0073] S5. The feature aggregation image is input into the head output layer head, and after convolution calculation, a disparity map is output to complete binocular stereo matching.
[0074] Therefore, in this embodiment, a carefully designed grouped convolution backbone is used to extract features. The backbone obtains higher semantic features through multiple downsamplings, which can avoid false alarms caused by repeated textures. Finally, the SPPF network is redesigned using cross convolution, and this network is used to replace the common loop matching module, improving the matching effect and running efficiency of the model.
[0075] Embodiment 2:
[0076] This embodiment provides a binocular stereo matching system, which can implement the binocular stereo matching method described in Embodiment 1. As Figure 6 shown, the binocular stereo matching system includes:
[0077] The backbone module 100 is used to fuse the left-eye image ImgL and the right-eye image ImgR, and extract the fused image features. Among them, the backbone module 1 uses the repvgg network structure to complete image feature extraction, and its structure and image feature extraction process are the same as those in Embodiment 1;
[0078] Three SPPF modules are respectively used to input the first fused image feature, the second fused image feature, and the third fused image feature into an SPPF module to correspondingly output the first feature map, the second feature map, and the third feature map. Among them, each SPPF module includes a first convolutional layer 101, 3 parallel MaxPool modules 102, a concat module 103, and a second convolutional layer 104;
[0079] The first upsampling + concat + feature aggregation module 201 is used to upsample the third feature map, perform concat processing on the upsampling result and the second feature map, perform 1×1 convolutional calculation on the concat processing result, and sequentially process the convolutional calculation result through m repvggblock modules (the value range of m is 2-5) to obtain the first processing result;
[0080] The second upsampling + concat + feature aggregation module 202 is used to upsample the first processing result, perform concat processing on the upsampling result and the first feature map, perform 1×1 convolutional calculation on the concat processing result, and sequentially process the convolutional calculation result through m repvggblock modules (the value range of m is 2-5) to obtain the second processing result;
[0081] The third upsampling + concat + feature aggregation module 203 is used to perform upsampling on the second processing result, perform concat processing on the upsampling result and the convolutional processed image features, perform 1×1 convolutional calculation on the concat processing result, and sequentially process the convolutional calculation result through m repvggblock modules (the value range of m is 2-5) to obtain the third processing result as the feature aggregation image; wherein, the processing processes of the first upsampling + concat + feature aggregation module 201, the second upsampling + concat + feature aggregation module 202, and the third upsampling + concat + feature aggregation module 203 on the image are the same as those in Embodiment 1;
[0082] The head output layer head 300 is used to receive the feature aggregation image and output a disparity map after convolutional calculation to complete binocular stereo matching.
[0083] In summary, the present application can extract features through a grouped convolutional backbone, which obtains higher semantic features through multiple downsamplings, can avoid false alarm problems caused by repeated textures, and finally redesigns the SPPF network using cross convolutions and replaces the common loop matching module with this network, improving the matching effect and running efficiency of the model.
[0084] It should be noted that the technical features in the above Embodiments 1-2 can be combined arbitrarily, and the combined technical solutions all fall within the protection scope of the present application. In this text, terms such as "including", "comprising" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. Without more limitations, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.
[0085] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principle and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A binocular stereo matching method, characterized in that: The steps include: Input the left eye image and the right eye image into the backbone module, fuse the left eye image and the right eye image through the backbone module, and extract the fused image features; The fused image features are processed through the SPPF module to output a feature map; The feature map is upsampled and feature aggregated through the upsampling + concat + feature aggregation module to output a feature aggregated image; The feature aggregation image is input into the head output layer head, and a disparity map is output after convolution calculation to complete binocular stereo matching.
2. The binocular stereo matching method according to claim 1, wherein: Before fusing the left image and the right image, both the left image and the right image are preprocessed.
3. The binocular stereo matching method according to claim 1, wherein: The backbone module uses a repvgg network structure to complete image feature extraction. The repvgg network structure is divided into 5 layers, and the repvgg network structure is used to complete the fusion of the left eye image and the right eye image and the extraction of the fused image features.
4. The binocular stereo matching method according to claim 3, characterized in that: The repvgg network structure is used to complete the fusion of the left and right images and the extraction of fused image features, which includes the following steps: The left image and the right image are fused by a concat method to obtain a fused image; Input the fused image into the first layer of the repvgg network structure, and perform the first downsampling and convolution operation by using the parameter setting of conv3×3, group=2, and stride=2, and input the result of the first convolution operation into the second layer of the repvgg network structure, and perform the second downsampling and convolution operation by using the parameter setting of conv3×3, group=2, and stride=2 again, so as to obtain the convolution processed image features; Inputting the convolution processed image features into the third layer, and obtaining the first fused image features after downsampling and convolution calculation; Inputting the first fused image feature into the fourth layer, and obtaining the second fused image feature after downsampling and convolution calculation; The second fused image feature is input into the fifth layer, and after downsampling and convolution calculation, the third fused image feature is obtained.
5. The binocular stereo matching method according to claim 4, characterized in that: The fused image features are processed by the SPPF module to output a feature map, including the following steps: The first fused image feature, the second fused image feature, and the third fused image feature are correspondingly input into an SPPF module to output a first feature map, a second feature map, and a third feature map respectively.
6. The binocular stereo matching method according to claim 5, characterized in that: Each SPPF module includes the first convolutional layer, three parallel MaxPool modules, a concat module, and the second convolutional layer; The fused image features are input into the first convolutional layer, and the convolution calculation results are passed through three parallel MaxPool modules to obtain the pooling results. The three pooling results are further spliced through the concat module, and the spliced results are feature mapped through the second convolutional layer to obtain the feature map.
7. The binocular stereo matching method according to claim 6, characterized in that: The second convolutional layer completes the convolution calculation through the CrossConv module.
8. The binocular stereo matching method according to claim 5, characterized in that: The feature map is upsampled and feature aggregated by the upsampling + concat + feature aggregation module to output a feature aggregated image. Specifically, it includes the following steps: The third feature map is upsampled by the first upsampling+concat+feature aggregation module, and the upsampling result is concat-processed with the second feature map, and a convolution calculation is performed on the concat processing result, and the convolution calculation result is processed by m repvggblock modules in sequence to obtain a first processing result; The first processing result is upsampled by a second upsampling+concat+feature aggregation module, and the upsampling result is concat-processed with the first feature map, and a convolution calculation is performed on the concat processing result, and the convolution calculation result is processed by m repvggblock modules in sequence to obtain a second processing result; The second processing result is upsampled by the third upsampling+concat+feature aggregation module, and the upsampling result is concat processed with the convolution processing image feature, and the concat processing result is convolution calculated, and the convolution calculation result is processed by m repvggblock modules in sequence to obtain the third processing result as the feature aggregation image.
9. The binocular stereo matching method according to claim 8, characterized in that: Each repvggblock module takes group=1 when performing calculations.
10. A binocular stereo matching system, characterized in that: include: The backbone module is used to fuse the left image and the right image and extract the features of the fused image; Three SPPF modules, which are respectively used to input the first fused image feature, the second fused image feature, and the third fused image feature into one SPPF module to output a first feature map, a second feature map, and a third feature map respectively; A first upsampling+concat+feature aggregation module is used to upsample the third feature map, concat the upsampling result with the second feature map, perform convolution calculation on the concat processing result, and sequentially process the convolution calculation result through m repvggblock modules to obtain a first processing result; A second upsampling+concat+feature aggregation module is used to upsample the first processing result, concat the upsampling result with the first feature map, perform convolution calculation on the concat processing result, and sequentially process the convolution calculation result through m repvggblock modules to obtain a second processing result; a third upsampling+concat+feature aggregation module, which is used to upsample the second processing result, concat the upsampling result with the convolution processing image feature, perform convolution calculation on the concat processing result, and sequentially process the convolution calculation result through m repvggblock modules to obtain a third processing result as a feature aggregation image; The head output layer is used to receive the feature aggregation image and output the disparity map after convolution calculation to complete binocular stereo matching.