Parallax estimation method and device, image processing apparatus, and storage medium

CN115239783BActive Publication Date: 2026-08-28ZTE CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110442253.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-04-23
Publication Date
2026-08-28
Estimated Expiration
2041-04-23

AI Technical Summary

Technical Problem

[0003]目前用于视差估计的神经网络大多采用较小的卷积核以及下采样层以扩大网络的感受野,但实际的感受野比该方法理论上的感受野要小得多,并不能获得足够多的上下文信息和不同场景中特征的依赖关系,大多数卷积神经网络用于视差估计往往会丢失许多相关联信息,影响视差估计的精度,进而影响匹配精度

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115239783B_ABST
    Figure CN115239783B_ABST
Patent Text Reader

Abstract

The application provides a disparity estimation method and device, an image processing apparatus and a storage medium. The method comprises processing an input image based on a first network model to obtain a direct cost volume of the input image, the input image comprising a first image and a second image, the first network model comprising a convolutional neural network, a pyramid convolutional network and a spatial pyramid pooling layer; processing the input image based on a second network model to obtain an associated cost volume of the input image, the second network model comprising a residual network; determining an estimated cost of the input image according to the associated cost volume and the direct cost volume; and calculating an estimated disparity corresponding to the first image and the second image according to the estimated cost.
Need to check novelty before this filing date? Find Prior Art

Claims

1. A disparity estimation method, characterized in that, include: The input image is processed based on the first network model to obtain the direct cost volume of the input image. The input image includes a first image and a second image. The first network model includes a convolutional neural network, a pyramid convolutional network, and a spatial pyramid pooling layer. The direct cost volume is obtained by concatenating the matching cost of the first image and the matching cost of the second image. The input image is processed based on the second network model to obtain the associated cost volume of the input image. The second network model includes a residual network. The associated cost volume is obtained by concatenating the low-level feature matching cost of the first image and the low-level feature matching cost of the second image. The estimated cost of the input image is determined based on the associated cost volume and the direct cost volume; The estimated disparity corresponding to the first image and the second image is calculated based on the estimated cost.

2. The method according to claim 1, characterized in that, The process of processing the input image based on the first network model to obtain the direct cost volume of the input image includes: Feature information of the first image and the second image is extracted based on the convolutional neural network; Based on the pyramid convolutional network, convolution operations are performed on the feature information of the first image and the second image respectively to obtain the multi-scale feature information of the first image and the multi-scale feature information of the second image. Based on the spatial pyramid pooling layer, the multi-scale feature information of the first image and the multi-scale feature information of the second image are aggregated respectively to obtain the matching cost of the first image and the matching cost of the second image; The matching cost of the first image is concatenated with the matching cost of the second image to obtain the direct cost volume of the input image.

3. The method according to claim 1, characterized in that, The process of processing the input image based on the second network model to obtain the associated cost volume of the input image includes: The low-level feature matching costs of the first image and the second image are obtained based on the residual network. The low-level feature matching cost of the first image and the low-level feature matching cost of the second image are grouped to obtain at least one associated cost information group, and each associated cost information group includes at least one low-level feature matching cost. Update the underlying feature matching cost for each associated cost information group; The updated low-level feature matching cost of the first image is concatenated with the updated low-level feature matching cost of the second image to obtain the associated cost volume; After repeating the grouping, updating, and stitching operations a preset number of times, the resulting associated cost volume is used as the associated cost volume of the input image.

4. The method according to claim 3, characterized in that, The updating of the underlying feature matching cost for each associated cost information group includes: Calculate the mean of the underlying feature matching cost for each associated cost information group; Replace the underlying feature matching cost of each associated cost information group with the mean of the underlying feature matching cost of that associated cost information group.

5. The method according to claim 3, characterized in that, The number of associated cost information groups is: ,in, group Indicates the number of related cost information groups. This represents the number of underlying feature matching costs. epoch Indicates the current iteration number. epoch The maximum value is equal to the preset number of times.

6. The method according to claim 1, characterized in that, Determining the estimated cost of the input image based on the associated cost volume and the direct cost volume includes: The average of the associated cost volume and the direct cost volume is used as the cost volume of the input image; The cost of the input image is estimated by aggregating the cost volume using a 3D convolutional neural network.

7. The method according to claim 1, characterized in that, The disparity between the first image and the second image is: , This indicates an estimate of parallax. Indicates the maximum parallax. Indicates parallax. This represents the Softmax function. This indicates the estimated cost.

8. A parallax estimation device, characterized in that, include: The first volume calculation module is configured to process the input image based on the first network model to obtain the direct cost volume of the input image. The input image includes a first image and a second image. The first network model includes a convolutional neural network, a pyramid convolutional network, and a spatial pyramid pooling layer. The direct cost volume is obtained by concatenating the matching cost of the first image and the matching cost of the second image. The second volume calculation module is configured to process the input image based on a second network model to obtain the associated cost volume of the input image. The second network model includes a residual network, and the associated cost volume is obtained by concatenating the low-level feature matching cost of the first image and the low-level feature matching cost of the second image. The cost estimation module is configured to determine the estimated cost of the input image based on the associated cost volume and the direct cost volume; The disparity estimation module is configured to calculate the disparity between the first image and the second image based on the estimated cost.

9. An image processing apparatus, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the disparity estimation method as described in any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the disparity estimation method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Census transformation-based binocular stereo matching method

    CN110473217A

  • Stereo matching cost volume construction method for binocular ranging

    CN111462212A