Stereo matching method and device, electronic equipment and storage medium

By extracting feature maps of different resolutions on resource-constrained devices, building the grouping distance cost volume, and using mixed cost aggregation and parallax regression techniques, the trade-off between parallax accuracy and running speed in stereo matching is solved, and efficient and accurate parallax estimation is achieved.

CN119967142AActive Publication Date: 2025-05-09SUN YAT SEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510028134.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-09
Estimated Expiration
2045-01-08

AI Technical Summary

Technical Problem

On devices with resource-constrained, it is difficult to achieve stereoscopic matching between high parallax accuracy and high operating speed at the same time.

Method used

A stereo matching method is proposed, by extracting feature maps of different resolutions, constructing grouping distance cost volumes, and using mixed cost aggregation and disparity regression techniques to gradually refine the disparity map.

Benefits of technology

It realizes efficient extraction of multi-scale features on resource-constrained devices, improves computing efficiency, and obtains accurate parallax maps to meet the needs of real-time and high precision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119967142A_ABST
    Figure CN119967142A_ABST
Patent Text Reader

Abstract

The invention discloses a stereo matching method and device, electronic equipment and a storage medium, and relates to the technical field of computer vision, and the method comprises the steps: extracting three feature maps of different resolutions from a stereo image; constructing a grouping distance cost volume by using the feature map, carrying out mixed cost aggregation on the grouping distance cost volume, and then carrying out parallax regression to obtain a parallax map; and constructing a grouping distance cost volume according to the feature map, carrying out mixed cost aggregation on the grouping distance cost volume and then carrying out parallax regression to obtain a residual image, and adding the residual image and the up-sampled parallax image to obtain a parallax image with higher precision. According to the method, multi-scale feature context information can be efficiently extracted by extracting the feature maps of different resolutions of the three-dimensional image, more spatial details are reserved, and the final disparity map is obtained through the grouping distance cost volume and disparity regression of three stages, so that the calculation efficiency is improved, and the calculation cost is reduced. And the accurate disparity map can be obtained by fully utilizing the depth information of the three-dimensional image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer vision technology, and in particular to a stereo matching method, device, electronic device and storage medium. Background Art

[0002] With the rapid development of technologies such as robotics, the Internet of Things, and autonomous driving, perception technology has become particularly important in these systems. Stereo matching, as a key perception technology, can generate depth information by predicting the pixel correspondence between the reference image and the target image using the principle of triangulation, providing these systems with reliable real-time environmental perception. However, it is difficult to achieve both high parallax accuracy and fast running speed when implementing stereo matching on resource-constrained devices. Summary of the invention

[0003] The main purpose of the embodiments of the present application is to propose a stereo matching method, apparatus, electronic device and storage medium to achieve stereo matching with high parallax accuracy and high operating speed on a resource-constrained device.

[0004] To achieve the above object, an embodiment of the present application provides a stereo matching method, which includes the following steps:

[0005] Extracting a first feature map of a unit resolution, a second feature map of twice the unit resolution, and a third feature map of four times the unit resolution from a stereoscopic image;

[0006] constructing a first group distance cost volume of the unit resolution by using the first feature map, performing mixed cost aggregation on the first group distance cost volume and then performing disparity regression to obtain a first residual map as a first disparity map;

[0007] Constructing a second group distance cost volume with twice the unit resolution according to the second feature map, performing mixed cost aggregation on the second group distance cost volume and then performing disparity regression to obtain a second residual map, and adding the second residual map to the upsampled first disparity map to obtain a second disparity map;

[0008] constructing a third grouping distance cost volume with four times the unit resolution according to the third feature map, performing mixed cost aggregation on the third grouping distance cost volume and then performing disparity regression to obtain a third residual map, and adding the third residual map to the upsampled second disparity map to obtain a third disparity map;

[0009] The third disparity map is determined as a target disparity map of the stereoscopic image.

[0010] In some embodiments, extracting a first feature map of a unit resolution, a second feature map of twice the unit resolution, and a third feature map of four times the unit resolution from a stereoscopic image comprises the following steps:

[0011] Input the stereoscopic image into two block convolutional layers respectively;

[0012] Using one of the block convolution layers to perform downsampling with a first step length to obtain a first initial feature map of the unit resolution; using another block convolution layer to perform downsampling with a second step length to obtain a second initial feature map of half the unit resolution;

[0013] Upsampling the first initial feature map and the second initial feature map to twice the unit resolution and then adding them to obtain a third initial feature map;

[0014] Downsampling the third initial feature map three times in sequence with a third step size to obtain a third candidate feature map with four times the unit resolution, a second candidate feature map with twice the unit resolution, and a first candidate feature map with the unit resolution, respectively;

[0015] The first candidate feature map is determined as the first feature map, the first feature map is upsampled to twice the unit resolution and then added to the second candidate feature map to obtain the second feature map, and the second feature map is upsampled to four times the unit resolution and then added to the third candidate feature map to obtain the third feature map.

[0016] In some embodiments, the step of constructing each of the group distance cost volumes comprises the following steps:

[0017] Splitting the feature map into a plurality of left features of a left channel and a plurality of right features of a right channel according to channels;

[0018] Calculate the L1 distance of each set of features as a distance cost volume; wherein each set of features includes a pair of the left feature and the right feature corresponding to each other;

[0019] The distance cost volumes are packed into a four-dimensional distance cost volume as the corresponding group distance cost volume.

[0020] In some embodiments, calculating the L1 distance of each set of features as the distance cost volume comprises the following steps:

[0021] According to the first expression, the L1 distance of each set of features is calculated as the distance cost volume;

[0022] The first expression is as follows:

[0023]

[0024] Among them, C gwd is the distance cost volume, g is the number of groups, d is the disparity, x and y are the horizontal and vertical coordinates of the pixel respectively, F L and F R are the left feature and the right feature respectively;

[0025] The step of packing the distance cost volumes into a four-dimensional distance cost volume as the corresponding group distance cost volume comprises the following steps:

[0026] According to the second expression, each of the distance cost volumes is packed into a four-dimensional distance cost volume as the corresponding group distance cost volume;

[0027] The second expression is:

[0028] C final (g,d,x,y)=C g (d,x,y);

[0029] Among them, C final is the group distance cost volume, g, d, x, y are the group number, disparity, pixel horizontal coordinate, and pixel vertical coordinate, respectively.

[0030] In some embodiments, the step of performing mixed cost aggregation on each of the grouped distance cost volumes comprises the following steps:

[0031] The grouped distance cost volume is sequentially passed through eight three-dimensional convolutional layers for mixed cost aggregation; wherein the first three-dimensional convolutional layer and the last three-dimensional convolutional layer of the eight three-dimensional convolutional layers are respectively used for dimension expansion and dimension recovery, and the middle six three-dimensional convolutional layers of the eight three-dimensional convolutional layers are used for local information learning of the grouped distance cost volume for regularization.

[0032] In some embodiments, performing mixed cost aggregation on the first grouped distance cost volume and then performing disparity regression to obtain a first residual map as the first disparity map includes the following steps:

[0033] Performing mixed cost aggregation on the first grouped distance cost volume and then performing a full-range disparity search to obtain the first residual map as the first disparity map; wherein the full range is obtained by multiplying the maximum value of the disparity search range at the original resolution of the stereoscopic image by the unit resolution;

[0034] The step of performing mixed cost aggregation on the second grouped distance cost volume and then performing disparity regression to obtain a second residual map comprises the following steps:

[0035] According to the third expression, the second group distance cost volume is subjected to mixed cost aggregation, and then a disparity search within a preset range is performed to obtain the second residual map; wherein the preset range is obtained by limiting the residual calculation;

[0036] The third expression is:

[0037] d Actual (x,y)=d1(x,y)+Δd(x,y);

[0038] Among them, d Actual is the corrected disparity value at the pixel position (x, y), where x and y are the horizontal coordinate and vertical coordinate of the pixel respectively, d1 is the disparity estimate obtained by performing the full range disparity search, and Δd represents the disparity offset;

[0039] The step of performing mixed cost aggregation on the third grouped distance cost volume and then performing disparity regression to obtain a third residual map comprises the following steps:

[0040] The third residual map is obtained by performing mixed cost aggregation on the third grouped distance cost volume according to the third expression and then performing a disparity search in the preset range.

[0041] In some embodiments, the disparity regression step comprises the following steps:

[0042] Performing mixed cost aggregation on each of the group distance cost volumes and then using a soft minimum regression function to convert each of the group distance cost volumes into a probability distribution;

[0043] Calculating a corresponding disparity value according to a weighted sum of the probability distributions of each discrete disparity level within each of the grouped distance cost volumes;

[0044] The expression for calculating the disparity value is as follows:

[0045]

[0046] Among them, d i represents the disparity value of the i-th pixel; M is the maximum disparity; d is a disparity candidate value, and the value range of d is 0 to M; d′ is used to represent another disparity candidate value traversed and summed, which is the same group of candidate values ​​as d; exp() is an exponential function used to convert the cost value into a non-negative number; C i (d) and C i (d′) is the cost value, which represents the matching error corresponding to a certain disparity value d or d′, generated by the grouped distance cost volume. The smaller the cost value, the better the match.

[0047] To achieve the above object, another aspect of the embodiment of the present application provides a stereo matching device, the device comprising:

[0048] A multi-scale feature extraction unit, configured to extract a first feature map of a unit resolution, a second feature map of twice the unit resolution, and a third feature map of four times the unit resolution from a stereoscopic image;

[0049] A first matching unit is configured to construct a first grouping distance cost volume of the unit resolution by using the first feature map, perform mixed cost aggregation on the first grouping distance cost volume, and then perform disparity regression to obtain a first residual map as a first disparity map;

[0050] a second matching unit, configured to construct a second grouping distance cost volume having twice the unit resolution according to the second feature map, perform mixed cost aggregation on the second grouping distance cost volume and then perform disparity regression to obtain a second residual map, and add the second residual map to the upsampled first disparity map to obtain a second disparity map;

[0051] a third matching unit, configured to construct a third grouping distance cost volume of four times the unit resolution according to the third feature map, perform mixed cost aggregation on the third grouping distance cost volume and then perform disparity regression to obtain a third residual map, and add the third residual map to the upsampled second disparity map to obtain a third disparity map;

[0052] The disparity determination unit is configured to determine the third disparity map as a target disparity map of the stereoscopic image.

[0053] To achieve the above objective, another aspect of an embodiment of the present application provides an electronic device, the electronic device comprising a memory and a processor, the memory storing a computer program, and the processor implementing the above method when executing the computer program.

[0054] To achieve the above objective, another aspect of an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program implements the above method when executed by a processor.

[0055] The embodiments of the present application include at least the following beneficial effects:

[0056] The present application can extract a first feature map of unit resolution, a second feature map of twice the unit resolution, and a third feature map of four times the unit resolution for a stereoscopic image; use the first feature map to construct a first group distance cost volume of unit resolution, perform mixed cost aggregation on the first group distance cost volume and then perform disparity regression to obtain a first residual map as a first disparity map; construct a second group distance cost volume of twice the unit resolution according to the second feature map, perform mixed cost aggregation on the second group distance cost volume and then perform disparity regression to obtain a second residual map, and add the second residual map to the upsampled first disparity map to obtain a second disparity map; construct a third group distance cost volume of four times the unit resolution according to the third feature map, perform mixed cost aggregation on the third group distance cost volume and then perform disparity regression to obtain a third residual map, and add the third residual map to the upsampled second disparity map to obtain a third disparity map; determine the third disparity map as the target disparity map of the stereoscopic image. By extracting feature maps of different resolutions of stereo images, the present application can efficiently extract multi-scale feature context information and retain more spatial details, and then obtain the final disparity map through three-stage grouping distance cost volume and disparity regression, which not only improves the computational efficiency, but also makes full use of the depth information of the stereo image to obtain an accurate disparity map. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0058] Figure 1 A schematic diagram of a flow chart of a stereo matching method provided in an embodiment of the present application;

[0059] Figure 2 An example flow chart of a stereo matching method based on multi-scale block feature extraction and mixed cost aggregation provided in an embodiment of the present application;

[0060] Figure 3 An example structural diagram of a multi-scale slicing feature extraction module provided in an embodiment of the present application;

[0061] Figure 4 An example flow chart for constructing a group distance cost volume provided in an embodiment of the present application;

[0062] Figure 5 An example structural diagram of a hybrid three-dimensional convolution cost aggregation module provided in an embodiment of the present application;

[0063] Figure 6A schematic diagram of the structure of a stereo matching device provided in an embodiment of the present application;

[0064] Figure 7 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0065] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is further described in detail below in conjunction with the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present application. They are only examples of devices and methods consistent with some aspects of the embodiments of the present application as detailed in the attached claims.

[0066] It is understood that the terms "first", "second", etc. used in this application can be used to describe various concepts in this article, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another concept. For example, without departing from the scope of the embodiment of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "if" as used herein can be interpreted as "at the time of" or "when" or "in response to determination".

[0067] The terms "at least one", "multiple", "each", "any", etc. used in this application, at least one includes one, two or more, multiple includes two or more, each refers to each of the corresponding multiple, and any refers to any one of the multiple.

[0068] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0069] Before describing the embodiments of the present application in detail, some related technologies that may be involved in the embodiments of the present application are first described as follows:

[0070] Most of the current precision-oriented stereo matching methods are based on deep convolutional neural networks, which have achieved remarkable success in generating disparity maps. For example, models based on highly complex 3D convolutional neural networks can make full use of depth information in the cost volume aggregation step and achieve very accurate disparity estimation results. However, such precision-oriented methods require a lot of computing resources and are difficult to meet the real-time task requirements on resource-constrained edge devices.

[0071] In order to solve the real-time problem on resource-constrained edge devices, some studies have designed a coarse-to-fine disparity estimation strategy, that is, a coarse disparity estimation is performed by constructing a full-range cost volume on a low-resolution feature map, and then the disparity result is gradually refined through bilinear interpolation upsampling and residual calculation. This coarse-to-fine strategy reduces computing resources to a certain extent, thereby increasing the running speed of the model. However, the use of low-resolution and low-quality feature maps may lead to the loss of detail information, thereby affecting the final disparity estimation accuracy.

[0072] In addition, the feature extraction module often has an important impact on the overall performance of the stereo matching method. Currently, mainstream stereo matching methods mostly use a symmetrical encoding-decoding network with a U-Net architecture to extract feature maps of different scales through multiple layers of standard two-dimensional convolution. However, the U-Net architecture is prone to losing key feature detail information during the downsampling process, resulting in a decrease in the accuracy of disparity estimation. The full distance and full correlation methods widely used in cost volume construction only generate single-channel distance or correlation maps, which limits the ability to express feature information.

[0073] In the cost aggregation step, precision-oriented methods often use a stacked hourglass structured 3D convolutional cost aggregation module, combined with intermediate supervision for high-precision disparity estimation. Although this type of method improves the aggregation quality, it is not suitable for various real-time task requirements on resource-constrained edge devices because the stacked hourglass structured 3D convolution requires a lot of computing resources.

[0074] In order to achieve real-time, high-precision stereo matching on resource-constrained devices, existing research has also attempted to design lightweight method architectures. For example, by adopting a coarse-to-fine strategy for feature extraction and a 2D convolutional cost volume aggregation module, the computational complexity can be reduced to a certain extent. However, these methods usually lack sufficient expression of high-quality disparity candidates, resulting in insufficient accuracy of disparity estimation results.

[0075] In summary, current stereo matching algorithms still face the challenge of information loss in the feature extraction module, which affects the accuracy of disparity estimation. At the same time, the cost aggregation module has insufficient regularization capability or high computational complexity, which makes it difficult to deploy in real time on resource-constrained platforms.

[0076] To solve the above problems, the embodiments of the present application provide a stereo matching method, device, electronic device and storage medium. The technical solution of the present application includes: extracting a first feature map of unit resolution, a second feature map of twice the unit resolution and a third feature map of four times the unit resolution from a stereo image; constructing a first group distance cost volume of unit resolution using the first feature map, performing mixed cost aggregation on the first group distance cost volume and then performing disparity regression to obtain a first residual map as a first disparity map; constructing a second group distance cost volume of twice the unit resolution according to the second feature map, performing mixed cost aggregation on the second group distance cost volume and then performing disparity regression to obtain a second residual map, adding the second residual map to the upsampled first disparity map to obtain a second disparity map; constructing a third group distance cost volume of four times the unit resolution according to the third feature map, performing mixed cost aggregation on the third group distance cost volume and then performing disparity regression to obtain a third residual map, adding the third residual map to the upsampled second disparity map to obtain a third disparity map; determining the third disparity map as the target disparity map of the stereo image. By extracting feature maps of different resolutions of stereo images, the present application can efficiently extract multi-scale feature context information and retain more spatial details, and then obtain the final disparity map through three-stage grouping distance cost volume and disparity regression, which not only improves the computational efficiency, but also makes full use of the depth information of the stereo image to obtain an accurate disparity map.

[0077] The embodiments of the present application provide a stereo matching method, device, electronic device and storage medium, which relate to the field of computer vision technology. The stereo matching method, device, electronic device and storage medium provided in the embodiments of the present application can be applied to a terminal, can also be applied to a server, and can also be software running in a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, and a car terminal, etc., but is not limited to this; the server side can be configured as an independent physical server, or a server cluster or distributed system composed of multiple physical servers, and can also be configured as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network; the software can be an application that implements a knowledge extraction method, etc., but is not limited to the above forms.

[0078] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments, in which tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0079] Reference Figure 1 The present application embodiment provides a stereo matching method, which may include but is not limited to S100 to S140, as follows:

[0080] S100: extracting a first feature map of a unit resolution, a second feature map of twice the unit resolution, and a third feature map of four times the unit resolution from a stereoscopic image.

[0081] Further, S100 may include the following steps S101 to S105:

[0082] S101: Input the stereoscopic image into two block convolution layers respectively;

[0083] S102: using one of the block convolution layers to perform downsampling with a first step size to obtain a first initial feature map of the unit resolution; using another block convolution layer to perform downsampling with a second step size to obtain a second initial feature map of half the unit resolution;

[0084] S103: upsampling the first initial feature map and the second initial feature map to twice the unit resolution and then adding them to obtain a third initial feature map;

[0085] S104: downsampling the third initial feature map three times in sequence with a third step size to obtain a third candidate feature map with four times the unit resolution, a second candidate feature map with two times the unit resolution, and a first candidate feature map with the unit resolution respectively;

[0086] S105: Determine the first candidate feature map as the first feature map, upsample the first feature map to twice the unit resolution and then add it to the second candidate feature map to obtain the second feature map, upsample the second feature map to four times the unit resolution and then add it to the third candidate feature map to obtain the third feature map.

[0087] S110: constructing the first grouping distance cost volume of the unit resolution by using the first feature map, performing mixed cost aggregation on the first grouping distance cost volume and then performing disparity regression to obtain a first residual map as a first disparity map.

[0088] Further, performing mixed cost aggregation on the first grouped distance cost volume and then performing disparity regression to obtain a first residual map as a first disparity map includes the following steps:

[0089] The first grouped distance cost volume is subjected to mixed cost aggregation and then a full-range disparity search is performed to obtain the first residual map as the first disparity map; wherein the full range is obtained by multiplying the maximum value of the disparity search range at the original resolution of the stereoscopic image by the unit resolution.

[0090] S120: constructing a second grouped distance cost volume with twice the unit resolution according to the second feature map, performing mixed cost aggregation on the second grouped distance cost volume and then performing disparity regression to obtain a second residual map, and adding the second residual map to the upsampled first disparity map to obtain a second disparity map.

[0091] Further, performing mixed cost aggregation on the second group distance cost volume and then performing disparity regression to obtain a second residual map includes the following steps:

[0092] According to the third expression, the second group distance cost volume is subjected to mixed cost aggregation, and then a disparity search within a preset range is performed to obtain the second residual map; wherein the preset range is obtained by limiting the residual calculation;

[0093] The third expression is:

[0094] d Actual (x,y)=d1(x,y)++Δd(x,y);

[0095] Among them, d Actual is the corrected disparity value at the pixel position (x, y), where x and y are the horizontal coordinate and vertical coordinate of the pixel respectively, d1 is the disparity estimate obtained by performing the full-range disparity search, and Δd represents the disparity offset.

[0096] S130: constructing a third grouped distance cost volume with four times the unit resolution according to the third feature map, performing mixed cost aggregation on the third grouped distance cost volume and then performing disparity regression to obtain a third residual map, and adding the third residual map to the upsampled second disparity map to obtain a third disparity map.

[0097] Further, performing mixed cost aggregation on the third group distance cost volume and then performing disparity regression to obtain a third residual map includes the following steps:

[0098] The third residual map is obtained by performing mixed cost aggregation on the third grouped distance cost volume according to the third expression and then performing a disparity search in the preset range.

[0099] S140: Determine the third disparity map as a target disparity map of the stereoscopic image.

[0100] As an optional implementation manner, the step of constructing each of the group distance cost volumes includes the following steps:

[0101] Splitting the feature map into a plurality of left features of a left channel and a plurality of right features of a right channel according to channels;

[0102] Calculate the L1 distance of each set of features as a distance cost volume; wherein each set of features includes a pair of the left feature and the right feature corresponding to each other;

[0103] The distance cost volumes are packed into a four-dimensional distance cost volume as the corresponding group distance cost volume.

[0104] More specifically, the step of calculating the L1 distance of each set of features as the distance cost volume comprises the following steps:

[0105] According to the first expression, the L1 distance of each set of features is calculated as the distance cost volume;

[0106] The first expression is as follows:

[0107]

[0108] Among them, C gwd is the distance cost volume, g is the number of groups, d is the disparity, x and y are the horizontal and vertical coordinates of the pixel respectively, F L and F R are the left feature and the right feature respectively;

[0109] The step of packing the distance cost volumes into a four-dimensional distance cost volume as the corresponding group distance cost volume comprises the following steps:

[0110] According to the second expression, each of the distance cost volumes is packed into a four-dimensional distance cost volume as the corresponding group distance cost volume;

[0111] The second expression is:

[0112] C final (g,d,x,y)=C g (d,x,y);

[0113] Among them, C final is the group distance cost volume, g, d, x, y are the group number, disparity, pixel horizontal coordinate, and pixel vertical coordinate, respectively.

[0114] As an optional implementation, the step of performing mixed cost aggregation on each of the group distance cost volumes includes the following steps:

[0115] The grouped distance cost volume is sequentially passed through eight three-dimensional convolutional layers for mixed cost aggregation; wherein the first three-dimensional convolutional layer and the last three-dimensional convolutional layer of the eight three-dimensional convolutional layers are respectively used for dimension expansion and dimension recovery, and the middle six three-dimensional convolutional layers of the eight three-dimensional convolutional layers are used for local information learning of the grouped distance cost volume for regularization.

[0116] As an optional implementation, the disparity regression step includes the following steps:

[0117] Performing mixed cost aggregation on each of the group distance cost volumes and then using a soft minimum regression function to convert each of the group distance cost volumes into a probability distribution;

[0118] Calculating a corresponding disparity value according to a weighted sum of the probability distributions of each discrete disparity level within each of the grouped distance cost volumes;

[0119] The expression for calculating the disparity value is as follows:

[0120]

[0121] Among them, d i represents the disparity value of the i-th pixel; M is the maximum disparity; d is a disparity candidate value, and the value range of d is 0 to M; d′ is used to represent another disparity candidate value traversed and summed, which is the same group of candidate values ​​as d; exp() is an exponential function used to convert the cost value into a non-negative number; C i (d) and C i (d′) is the cost value, which represents the matching error corresponding to a certain disparity value d or d′, generated by the grouped distance cost volume. The smaller the cost value, the better the match.

[0122] Next, the solution of the embodiment of the present application will be introduced and explained in detail with reference to specific application examples.

[0123] The purpose of this embodiment is to provide a real-time, high-precision stereo matching method to address the trade-off between running speed and disparity accuracy in stereo matching in the prior art, especially in various real-time application scenarios on resource-constrained edge devices. By designing a multi-scale block feature extraction module, this embodiment can effectively capture multi-scale contextual information in stereo images and improve the accuracy of feature extraction. At the same time, a group distance-based method is used to construct a high-quality full-range and compact cost volume, providing higher-precision disparity candidates for subsequent depth prediction. This embodiment also develops a lightweight hybrid three-dimensional convolutional cost aggregation module that can efficiently regularize the group distance cost volume, taking into account both computational efficiency and disparity prediction accuracy.

[0124] Through the three-stage coarse-to-fine residual calculation strategy for disparity estimation, this embodiment achieves high-quality feature extraction at low resolution, and gradually refines the disparity map to generate the final high-resolution disparity result. The optimized method can meet the real-time task requirements on resource-constrained edge devices and significantly improve the disparity prediction accuracy, overcoming the limitations of existing precision-oriented or speed-oriented stereo matching methods that are difficult to balance between real-time and high precision.

[0125] The purpose of this embodiment is achieved through the following technical solutions:

[0126] This embodiment designs a real-time and high-precision stereo matching method to solve the trade-off problem between real-time and precision in the prior art, while meeting the demand for efficient perception on resource-constrained edge devices. To achieve the goal of real-time and high precision, this embodiment adopts the following technical means and measures.

[0127] First, a novel multi-scale tile feature extraction module is proposed, which combines the advantages of tile layers and convolution layers to efficiently capture the multi-scale contextual information of left and right features from stereo images. Unlike the widely used maximum pooling downsampling method, this module achieves a step-by-step change in feature resolution through tile layers and convolution layers, which can retain more spatial information of features at low resolution, especially in complex scenes that require fine parallax details. This module also fuses multi-scale features through upsampling and convolution operations with a specific structure, enabling it to quickly provide high-quality feature maps to construct cost volumes.

[0128] Secondly, this embodiment adopts a group distance-based method to construct full-range and compact cost volumes, by dividing the left and right features into multiple groups along the channel dimension and calculating the distance map of each group, thereby generating high-quality disparity candidates that fuse multi-channel feature information. The full-range cost volume is generated at a low resolution (1 / 8), and the ability to capture global features is improved by introducing contextual information. The compact cost volume is generated at medium (1 / 4) and high resolutions (1 / 2), but focuses on specific candidate areas (such as disparity candidates between -2 and 2), thereby greatly reducing computing resources. It should be noted that this embodiment can use a resolution of 1 / 8 of the original resolution of the stereoscopic image as a unit resolution.

[0129] In order to efficiently regularize the group distance cost volume, this embodiment develops a lightweight hybrid 3D convolutional cost aggregation module. This module significantly improves the regularization effect of the cost volume at a lower computational cost by aggregating the global context information of the cost volume step by step. In the full-range cost volume, global features are captured through a larger receptive field, while in the compact cost volume, small convolutional layers are used to efficiently integrate local features, thereby achieving a balance between global and local information.

[0130] Finally, this embodiment proposes a three-stage stereo method based on a coarse-to-fine strategy. In the first stage, a novel multi-scale block feature extraction module is used to generate low-resolution (1 / 8) left and right feature maps, and then the left and right feature compositions are used to construct the full-range cost volume of the grouped distance. Finally, the cost volume is regularized by the hybrid three-dimensional convolution cost aggregation module and the coarse disparity map is generated after disparity regression. In the second stage, after upsampling the coarse disparity map, the disparity residual is calculated, and a compact cost volume of medium resolution (1 / 4) is constructed to reduce the computational complexity. In the third stage, the disparity map of the second stage is further upsampled and the new disparity residual is calculated to construct a compact cost volume of high resolution (1 / 2) to generate a high-resolution, high-precision disparity map. Through stage-by-stage optimization and residual calculation methods, this method greatly improves real-time performance while ensuring high precision, meeting various real-time requirements of resource-limited edge devices.

[0131] Next, this embodiment will be described in more detail.

[0132] Embodiment 1: Overall implementation of the solution of this embodiment.

[0133] This embodiment proposes a stereo matching method based on multi-scale block feature extraction and mixed cost aggregation. The overall process of the method is as follows: Figure 2 As shown, it includes four steps: feature extraction, cost volume construction, cost aggregation, and disparity regression. The inference time is accelerated by residual calculation in the second and third stages.

[0134] The specific steps are as follows:

[0135] 1. Feature extraction stage: An innovative multi-scale slicing feature extraction module is used to extract multi-scale left and right features from stereo images. The input stereo image size is [3, H, W]. After processing, feature maps of three resolutions (1 / 8, 1 / 4 and 1 / 2) are generated, namely [8c, H / 8, W / 8], [4c, H / 4, W / 4] and [2c, H / 2, W / 2], where H, W and c represent the image height, image width and hyperparameters affecting the channel, respectively.

[0136] 2. Grouped distance cost volume construction stage: The first stage uses low-resolution features to construct a full-range grouped distance cost volume with a resolution of 1 / 8. The second stage constructs a compact grouped distance cost volume with a resolution of 1 / 4 based on the disparity results of the first stage. The third stage constructs a compact grouped distance cost volume with a resolution of 1 / 2 based on the disparity results of the second stage.

[0137] 3. Cost aggregation stage: Use the hybrid 3D convolution cost aggregation module to normalize the group distance cost volume. In the first stage, large receptive field convolution is used to perform global information aggregation on the full range group distance cost volume, and in the subsequent stage, lightweight 3D convolution is used to optimize the compact group distance cost volume.

[0138] 4. Disparity regression stage: Each stage uses a soft minimum regression (soft argmin) operation to regress the disparity map of the first stage or the residual map of the second and third stages from the regularized grouped distance cost volume. After the soft minimum regression operation, the first stage can directly obtain a coarse disparity of 1 / 8 resolution. Figure 1 .

[0139] 5. Residual calculation method: In order to achieve fast disparity estimation in the second stage, the disparity search range is limited to between -2 and 2. Figure 1 To reconstruct the right feature, and construct a 1 / 4 resolution group distance cost volume with the left feature to generate the residual Figure 2 The second stage requires the residual Figure 2 With parallax Figure 1 Adding for fine parallax at 1 / 4 resolution Figure 2 , and then use bilinear interpolation to convert the disparity Figure 2 Upsample to the original image size. The third stage is similar to the disparity estimation process in the second stage. Figure 3 With parallax Figure 2 Added to get the final high precision parallax at 1 / 2 resolution Figure 3 , and bilinear interpolation is used to upsample the disparity by 3 to the original image size.

[0140] Embodiment 2: Implementation of multi-scale slicing feature extraction module.

[0141] The structure of the multi-scale block feature extraction module proposed in this embodiment is as follows: Figure 3 shown.

[0142] 1. The input stereo image first passes through two block convolutional layers, and is downsampled to 1 / 8 and 1 / 16 resolutions with stride 8 and stride 16 respectively to obtain two features.

[0143] 2. The above features are upsampled to 1 / 2 resolution by bilinear interpolation and added (Concat), and then combined with two standard 3×3 2D convolutional layers to further extract multi-scale features. The standard 2D convolutional layer is followed by batch normalization (Batch Normalization 2d) and activation function (ReLU) operations by default.

[0144] 3. Finally, three block convolutions with a step size of 2 are used to gradually downsample the above multi-scale features to 1 / 2, 1 / 4 and 1 / 8 resolutions. Similarly, two standard 3×3 two-dimensional convolutional layers are combined to further extract multi-scale features. Then, double-line interpolation and two standard 3×3 two-dimensional convolutional layers are used to upsample and generate feature maps with three resolutions of 1 / 2, 1 / 4 and 1 / 8.

[0145] Optimization effect: This module uses block convolution to extract 1 / 8 and 1 / 16 resolution features from the input stereo image and fuses multi-scale context information. In addition, this module replaces the traditional pooling operation with block convolution to retain more spatial information. Its efficient feature extraction capability significantly improves the matching quality of low-resolution feature maps and provides high-precision input for the subsequent construction of group distance cost volume.

[0146] Embodiment 3: Constructing a group distance cost volume.

[0147] The method of constructing the group distance cost volume in this embodiment is as follows: Figure 4 shown.

[0148] 1. Divide the left and right image features into 48 groups along the channel dimension.

[0149] 2. Calculate the L1 distance of each group one by one. The formula is as follows:

[0150]

[0151] Among them C gwd is the group distance cost volume, g is the number of groups, d is the disparity, x and y are the horizontal and vertical coordinates of the pixel respectively, and F L and F R Left and right features respectively.

[0152] 3. Pack the cost volumes of all groups into a final 4D cost volume with resolutions of [N,M / 8,H / 8,W / 8], [N,M / 4,H / 4,W / 4], and [N,M / 2,H / 2,W / 2], respectively. The formula is as follows:

[0153] C final (g,d,x,y)=C g (d, x, y);

[0154] Among them C final is the final group distance cost volume, g, d, x, y are the number of groups, disparity, pixel horizontal coordinate, and pixel vertical coordinate, respectively.

[0155] Optimization effect: By grouping distance calculation, the speed of cost volume construction is improved while retaining high-quality disparity candidates of multi-channel features.

[0156] Example 4: Implementation of a hybrid three-dimensional convolutional cost aggregation module.

[0157] The structure of the hybrid 3D convolution cost aggregation module proposed in this embodiment is as follows: Figure 5 shown.

[0158] 1. In the first stage, a 5×5×5 3D convolutional layer is used to aggregate global context information for the full range of grouped distance cost volumes. The 3D convolutional layer is preceded by batch normalization (Batch Normalization 3d) and activation function (ReLU) operations by default.

[0159] 2. In the second and third stages, a 3×3×3 lightweight 3D convolutional layer is used to regularize the group distance compact cost volume.

[0160] 3. The grouped distance cost volume of each stage is output after being normalized by eight layers of 3D convolution. The first and last layers are used for dimension expansion and restoration, respectively, and the middle six layers are used for local information learning for regularization. The overall structure is divided into three stages, so the whole method contains a total of 24 3D convolution layers.

[0161] Optimization effect: The module design fully considers the characteristics of group distance cost volume at different stages, and realizes efficient group distance cost volume regularization with lightweight three-dimensional convolution.

[0162] Example 5: Implementation of residual calculation method

[0163] This embodiment adopts the residual calculation method to perform disparity estimation, which is specifically implemented as follows:

[0164] 1. In the first stage, a full-range grouping distance cost volume with a resolution of 1 / 8 is constructed, and the disparity search range is M=24 (ie, 192 / 8, 192 refers to the maximum value of the disparity search range at the original resolution).

[0165] 2. In the second and third stages, the disparity search range is limited to [-2, 2] through residual calculation, which helps to reduce unnecessary calculations. The formula is as follows:

[0166] d Actual (x,y)=d1(x,y)+Δd(x,y)

[0167] where d Actual is the disparity value at the pixel position (x, y) after the second or third stage of refinement and correction, where x and y are the horizontal and vertical coordinates of the pixel, respectively, d1 is the disparity estimate obtained in the first stage, and △d represents a small range disparity offset.

[0168] 3. Use the reconstructed right and left features to construct a compact group distance cost volume, which greatly speeds up the calculation. Since the first stage uses a low resolution (1 / 8) to construct the group distance cost volume, some errors may be introduced, so the residual of the second stage needs to be Figure 2 Rough parallax with the first stage Figure 1 Add together to get the second stage fine disparity Figure 2 The calculation method of the residual in the third stage is similar to that in the second stage. Figure 3 Parallax with the second stage Figure 2 Add together to get the final high-precision parallax Figure 3 .

[0169] Optimization effect: Through coarse-to-fine disparity prediction, the accuracy of high-resolution disparity maps is significantly improved while reducing computational overhead. In the first stage, the full disparity search range is used to capture the global structure, while in the second and third stages, the disparity estimation is refined by limiting the search range and using compact grouped distance cost volumes, thus achieving fast and high-precision disparity estimation.

[0170] Example 6: Implementation of parallax regression

[0171] In the design of this embodiment, the disparity regression step further processes the full range and compact grouped distance cost volumes to obtain the final disparity map. The specific implementation is as follows:

[0172] After the grouped distance cost volume is processed by the hybrid 3D convolutional cost aggregation module, the grouped distance cost volume is converted into a probability distribution using the soft minimum regression (soft argmin) function. The final calculated disparity value is calculated as the weighted sum of the probability distributions of each discrete disparity level within the grouped distance cost volume. The formula is as follows:

[0173] where d i is the expected disparity value of the i-th pixel, which is the final calculated disparity value. M is the maximum disparity. d is a possible disparity candidate value, ranging from 0 to M. d′ is used to represent another disparity candidate value to be traversed and summed, which is the same group of candidate values ​​as d. exp() is an exponential function used to convert the cost value into a non-negative number. C i (d) and C i (d′) is the cost value, which represents the matching error corresponding to a certain disparity value d or d′, generated by the grouped distance cost volume. The smaller it is, the better the match.

[0174] Optimization effect: This method effectively utilizes the multi-disparity candidate information of the grouped distance cost volume, thereby improving the accuracy of disparity prediction.

[0175] Example 7: Implementation of loss function.

[0176] In order to improve the accuracy of disparity estimation, this embodiment uses a smooth L1 loss function to supervise the training of the proposed method. This loss function is insensitive to outliers and has high robustness. The loss function of each stage is defined as:

[0177]

[0178] Where L K is the loss function value of the Kth (first, second and third) stage, d is the disparity prediction, is the true disparity, D is the number of disparity candidates, d i and They represent the i-th disparity prediction value and the true disparity value respectively. The formula of the smooth L1 function is as follows:

[0179]

[0180] Smooth L1 is a smooth L1 function.

[0181] In the three-stage method of this embodiment, the loss functions of different stages are weighted with different weights, and the final total loss function is:

[0182]

[0183] Where Lt is the total loss function of the final three stages, d is the disparity prediction, is the true disparity, the weight coefficients are set to λ1=0.5, λ2=0.7 and λ3=1, respectively. L1, L2 and L3 represent the loss functions of the first, second and third stages, respectively.

[0184] Optimization effect: By introducing a multi-scale supervised learning strategy, the proposed method can be optimized at 1 / 8, 1 / 4 and 1 / 2 resolutions, further improving the robustness and accuracy of disparity estimation.

[0185] Example 8: Dataset, evaluation metrics, method optimization and performance verification.

[0186] 1. Dataset: The method proposed in this embodiment is pre-trained on the Scene Flow dataset for 90 epochs, and then fine-tuned on the KITTI 2012 and KITTI 2015 datasets for 1800 epochs. Scene Flow contains 35454 pairs of stereo images for training and 4370 pairs of stereo images for testing, with a resolution of 540×960 and high-quality dense disparity labels. KITTI 2012 contains 160 pairs of training images and 34 pairs of test images, and KITTI 2015 contains 200 pairs of training images and 200 pairs of test images, with a resolution of 376×1240.

[0187] 2. Evaluation indicators: Three-pixel error (3px-all) refers to the proportion of pixels with an error exceeding 3 pixels. All-area error (D1-all) refers to the proportion of pixels with an error exceeding 3 pixels and 5% of the ground truth disparity. The smaller the two error indicators of 3px-all and D1-all, the higher the disparity accuracy.

[0188] 3. Method optimization and performance verification: On the KITTI 2012 and KITTI 2015 datasets, the results are reported through online benchmarks and compared with current mainstream methods. The 3px-all of the proposed three-stage method is 3.21%, 2.38%, and 2.02%, respectively, while the D1-all is 3.76%, 2.90%, and 2.50%, respectively, which is better than all current speed-oriented stereo matching methods. On resource-constrained edge devices such as NVIDIA Jetson AGX Orin, after TensorRT optimization, the inference speed of the proposed three-stage method reaches 28-40 frames per second, far exceeding the real-time requirements. Compared with existing stereo matching methods, this embodiment achieves an efficient balance between disparity accuracy and inference speed.

[0189] In summary, the beneficial effects of this embodiment include:

[0190] 1. This embodiment designs a multi-scale slicing feature extraction module, which can efficiently extract the multi-scale feature context information of the input image and retain more spatial details by combining the slicing layer and the convolution layer. This structure significantly improves the disparity accuracy of depth estimation. In addition, the multi-scale slicing feature extraction module uses lightweight slicing and convolution operations to quickly generate left and right feature maps of different resolutions, providing high-quality input for subsequent cost volume construction and disparity calculation, meeting real-time performance requirements.

[0191] 2. Compared with the traditional cost volume construction method based on single distance calculation, this embodiment divides the left and right features into multiple groups along the channel dimension by adopting the group distance method, and generates an independent distance map for each group. This structural design can integrate more comprehensive feature information to obtain high-quality disparity candidates, improve the initial quality of the cost volume, and thus provide a more reliable basis for disparity prediction.

[0192] 3. This embodiment innovatively develops a lightweight hybrid 3D convolution aggregation module, which effectively aggregates the global context information of the grouped distance cost volume by adopting a progressive DC pipeline structure. At each stage of the grouped cost volume, the cost aggregation module flexibly adjusts the receptive domain of the convolution operation according to the different characteristics of the full range and compact cost volumes, which not only ensures computational efficiency, but also makes full use of depth information for accurate regularization.

[0193] 4. The three-stage disparity estimation method from coarse to fine designed in this embodiment effectively reduces the amount of calculation by performing residual calculation in the refinement stage. The disparity map generated in the previous stage is gradually refined through the residual calculation strategy, and in the subsequent stage, only a small range of candidates near the coarse disparity map need to be focused on. This structure greatly reduces the computational complexity of the subsequent high-resolution stage and improves the detail performance of disparity prediction.

[0194] 5. The multi-scale slicing feature extraction module and hybrid 3D convolution cost aggregation module designed in this embodiment are suitable for expansion and adjustment in different computing environments, especially resource-constrained edge devices (such as NVIDIA Jetson AGX Orin). These modules are lightweight and can efficiently process multi-resolution data, ensuring the adaptability and reliability of the stereo matching method in embedded scenarios.

[0195] 6. This embodiment takes full consideration of the optimization of computing resources during the design process, and the proposed method can exceed the real-time performance requirements after being optimized using TensorRT. The entire method can achieve high-precision disparity prediction while significantly improving the inference time, meeting the real-time requirements of resource-constrained edge devices.

[0196] 7. By improving the key links of feature extraction, cost volume construction and cost aggregation, this embodiment not only improves the disparity accuracy, but also enhances the robustness to flat and interference areas. Compared with the prior art, the stereo matching method of this embodiment achieves the most advanced balance between inference speed and disparity accuracy.

[0197] Reference Figure 6 The present application also provides a stereo matching device, which can implement the above-mentioned stereo matching method. The device includes:

[0198] A multi-scale feature extraction unit, configured to extract a first feature map of a unit resolution, a second feature map of twice the unit resolution, and a third feature map of four times the unit resolution from a stereoscopic image;

[0199] A first matching unit is configured to construct a first grouping distance cost volume of the unit resolution by using the first feature map, perform mixed cost aggregation on the first grouping distance cost volume, and then perform disparity regression to obtain a first residual map as a first disparity map;

[0200] a second matching unit, configured to construct a second grouping distance cost volume having twice the unit resolution according to the second feature map, perform mixed cost aggregation on the second grouping distance cost volume and then perform disparity regression to obtain a second residual map, and add the second residual map to the upsampled first disparity map to obtain a second disparity map;

[0201] a third matching unit, configured to construct a third grouping distance cost volume of four times the unit resolution according to the third feature map, perform mixed cost aggregation on the third grouping distance cost volume and then perform disparity regression to obtain a third residual map, and add the third residual map to the upsampled second disparity map to obtain a third disparity map;

[0202] The disparity determination unit is configured to determine the third disparity map as a target disparity map of the stereoscopic image.

[0203] It can be understood that the contents of the above method embodiments are all applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0204] The embodiment of the present application also provides an electronic device, the electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the above-mentioned stereo matching method when executing the computer program. The electronic device can be any smart terminal including a tablet computer, a car computer, etc.

[0205] It can be understood that the contents of the above method embodiments are all applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0206] See also Figure 7 , Figure 7 The hardware structure of an electronic device of another embodiment is illustrated, and the electronic device includes:

[0207] The processor 701 may be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;

[0208] The memory 702 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 702 can store an operating system and other application programs. When the technical solution provided in the embodiment of this specification is implemented by software or firmware, the relevant program code is stored in the memory 702, and the processor 701 calls and executes a stereo matching method of the embodiment of this application;

[0209] Input / output interface 703, used to implement information input and output;

[0210] Communication interface 704, used to realize communication interaction between the device and other devices, which can be realized by wired mode (such as USB, network cable, etc.) or wireless mode (such as mobile network, WIFI, Bluetooth, etc.);

[0211] A bus 705 that transmits information between the various components of the device (e.g., the processor 701, the memory 702, the input / output interface 703, and the communication interface 704);

[0212] The processor 701 , the memory 702 , the input / output interface 703 and the communication interface 704 are connected to each other in communication within the device via a bus 705 .

[0213] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned stereo matching method is implemented.

[0214] It can be understood that the contents of the above method embodiments are all applicable to the present storage medium embodiments, the functions specifically implemented by the present storage medium embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0215] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0216] The embodiments described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0217] Those skilled in the art will appreciate that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.

[0218] The device embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0219] Those skilled in the art will appreciate that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices may be implemented as software, firmware, hardware, or a suitable combination thereof.

[0220] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0221] It should be understood that in the present application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0222] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the above units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0223] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0224] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0225] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including multiple instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, referred to as ROM), random access memory (Random Access Memory, referred to as RAM), disk or optical disk and other media that can store programs.

[0226] The preferred embodiments of the present application are described above with reference to the accompanying drawings, but the scope of the rights of the present application is not limited thereto. Any modification, equivalent substitution and improvement made by a person skilled in the art without departing from the scope and essence of the present application should be within the scope of the rights of the present application.

Claims

1. A stereo matching method, characterized in that: The method comprises the following steps: Extracting a first feature map of a unit resolution, a second feature map of twice the unit resolution, and a third feature map of four times the unit resolution from a stereoscopic image; constructing a first group distance cost volume of the unit resolution by using the first feature map, performing mixed cost aggregation on the first group distance cost volume and then performing disparity regression to obtain a first residual map as a first disparity map; Constructing a second group distance cost volume with twice the unit resolution according to the second feature map, performing mixed cost aggregation on the second group distance cost volume and then performing disparity regression to obtain a second residual map, and adding the second residual map to the upsampled first disparity map to obtain a second disparity map; constructing a third grouping distance cost volume with four times the unit resolution according to the third feature map, performing mixed cost aggregation on the third grouping distance cost volume and then performing disparity regression to obtain a third residual map, and adding the third residual map to the upsampled second disparity map to obtain a third disparity map; The third disparity map is determined as a target disparity map of the stereoscopic image.

2. A stereo matching method according to claim 1, characterized in that: The step of extracting a first feature map of a unit resolution, a second feature map of twice the unit resolution, and a third feature map of four times the unit resolution from a stereoscopic image comprises the following steps: Input the stereoscopic image into two block convolutional layers respectively; Using one of the block convolution layers to perform downsampling with a first step length to obtain a first initial feature map of the unit resolution; using another block convolution layer to perform downsampling with a second step length to obtain a second initial feature map of half the unit resolution; Upsampling the first initial feature map and the second initial feature map to twice the unit resolution and then adding them to obtain a third initial feature map; Downsampling the third initial feature map three times in sequence with a third step size to obtain a third candidate feature map with four times the unit resolution, a second candidate feature map with twice the unit resolution, and a first candidate feature map with the unit resolution, respectively; The first candidate feature map is determined as the first feature map, the first feature map is upsampled to twice the unit resolution and then added to the second candidate feature map to obtain the second feature map, and the second feature map is upsampled to four times the unit resolution and then added to the third candidate feature map to obtain the third feature map.

3. A stereo matching method according to claim 1, characterized in that: The step of constructing each of the group distance cost volumes comprises the following steps: Splitting the feature map into a plurality of left features of a left channel and a plurality of right features of a right channel according to channels; Calculate the L1 distance of each set of features as a distance cost volume; wherein each set of features includes a pair of the left feature and the right feature corresponding to each other; The distance cost volumes are packed into a four-dimensional distance cost volume as the corresponding group distance cost volume.

4. A stereo matching method according to claim 3, characterized in that: The step of calculating the L1 distance of each set of features as the distance cost volume includes the following steps: According to the first expression, the L1 distance of each group of features is calculated as the distance cost volume; The first expression is as follows: Among them, C gwd is the distance cost volume, g is the number of groups, d is the disparity, x and y are the horizontal and vertical coordinates of the pixel respectively, F L and F R are the left feature and the right feature respectively; The step of packing the distance cost volumes into a four-dimensional distance cost volume as the corresponding group distance cost volume comprises the following steps: According to the second expression, each of the distance cost volumes is packed into a four-dimensional distance cost volume as the corresponding group distance cost volume; The second expression is: C final (g,d,x,y)=C g (d,x,y); Among them, C final is the group distance cost volume, g, d, x, y are the group number, disparity, pixel horizontal coordinate, and pixel vertical coordinate, respectively.

5. A stereo matching method according to claim 1, characterized in that: The step of performing mixed cost aggregation on each of the grouped distance cost volumes comprises the following steps: The grouped distance cost volume is sequentially passed through eight three-dimensional convolutional layers for mixed cost aggregation; wherein the first three-dimensional convolutional layer and the last three-dimensional convolutional layer of the eight three-dimensional convolutional layers are respectively used for dimension expansion and dimension recovery, and the middle six three-dimensional convolutional layers of the eight three-dimensional convolutional layers are used for local information learning of the grouped distance cost volume for regularization.

6. A stereo matching method according to claim 1, characterized in that: The step of performing mixed cost aggregation on the first grouped distance cost volume and then performing disparity regression to obtain a first residual map as a first disparity map includes the following steps: Performing mixed cost aggregation on the first grouped distance cost volume and then performing a full-range disparity search to obtain the first residual map as the first disparity map; wherein the full range is obtained by multiplying the maximum value of the disparity search range at the original resolution of the stereoscopic image by the unit resolution; The step of performing mixed cost aggregation on the second grouped distance cost volume and then performing disparity regression to obtain a second residual map comprises the following steps: According to the third expression, the second group distance cost volume is subjected to mixed cost aggregation, and then a disparity search within a preset range is performed to obtain the second residual map; wherein the preset range is obtained by limiting the residual calculation; The third expression is: d Actual (x,y)=d1(x,y)+Δd(x,y); Among them, d Actual is the corrected disparity value at the pixel position (x, y), where x and y are the horizontal coordinate and vertical coordinate of the pixel respectively, d1 is the disparity estimate obtained by performing the full range disparity search, and Δd represents the disparity offset; The step of performing mixed cost aggregation on the third grouped distance cost volume and then performing disparity regression to obtain a third residual map comprises the following steps: The third residual map is obtained by performing mixed cost aggregation on the third grouped distance cost volume according to the third expression and then performing a disparity search in the preset range.

7. A stereo matching method according to claim 1, characterized in that: The disparity regression step comprises the following steps: Performing mixed cost aggregation on each of the group distance cost volumes and then using a soft minimum regression function to convert each of the group distance cost volumes into a probability distribution; Calculating a corresponding disparity value according to a weighted sum of the probability distributions of each discrete disparity level within each of the grouped distance cost volumes; The expression for calculating the disparity value is as follows: Among them, d i represents the disparity value of the i-th pixel; M is the maximum disparity; d is a disparity candidate value, and the value range of d is 0 to M; d′ is used to represent another disparity candidate value traversed and summed, which is the same group of candidate values ​​as d; exp() is an exponential function used to convert the cost value into a non-negative number; C i (d) and C i (d′) is the cost value, which represents the matching error corresponding to a certain disparity value d or d′, generated by the grouped distance cost volume. The smaller the cost value, the better the match.

8. A stereo matching device, characterized in that: The device comprises: A multi-scale feature extraction unit, configured to extract a first feature map of a unit resolution, a second feature map of twice the unit resolution, and a third feature map of four times the unit resolution from a stereoscopic image; A first matching unit is configured to construct a first grouping distance cost volume of the unit resolution by using the first feature map, perform mixed cost aggregation on the first grouping distance cost volume, and then perform disparity regression to obtain a first residual map as a first disparity map; a second matching unit, configured to construct a second grouping distance cost volume having twice the unit resolution according to the second feature map, perform mixed cost aggregation on the second grouping distance cost volume and then perform disparity regression to obtain a second residual map, and add the second residual map to the upsampled first disparity map to obtain a second disparity map; a third matching unit, configured to construct a third grouping distance cost volume of four times the unit resolution according to the third feature map, perform mixed cost aggregation on the third grouping distance cost volume and then perform disparity regression to obtain a third residual map, and add the third residual map to the upsampled second disparity map to obtain a third disparity map; The disparity determination unit is configured to determine the third disparity map as a target disparity map of the stereoscopic image.

9. An electronic device, characterized in that: The electronic device comprises a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Two-stage real-time binocular depth estimation method and device based on grouping mixing

    CN115546279A

  • Hybrid cost body binocular stereo matching method, device and storage medium

    WO2023240764A1