Depth estimation method and depth estimation device

Through the self-attention and cross-attention modules, the internal and external information interaction of the image is generated, and the depth estimation accuracy problem of texture poor areas in multi-view stereo vision is solved, and high-precision depth map generation is achieved.

CN114913215BActive Publication Date: 2025-08-15YUANLI JUHE (CHONGQING) ROBOTICS TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210323890.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-11-25
Filing Date
2022-03-29
Publication Date
2025-08-15
Estimated Expiration
2042-03-29

AI Technical Summary

Technical Problem

Existing multi-view stereo vision technology is difficult to achieve high-precision depth map estimation in areas such as poor texture, repeated textures and non-Lambertian surface areas, and the potential correspondence between images is not fully utilized to affect the depth map estimation accuracy.

Method used

The self-attention and cross-attention modules are used to interact with the internal and external information of the image. Through the feature matching deformation module and the adaptive receptive field module, the global features of the image are generated, and the depth map is generated based on the external parameter information.

Benefits of technology

Improved the accuracy of depth estimation, especially in areas such as weak textures, repeated textures and non-Lambertian surfaces.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114913215B_ABST
    Figure CN114913215B_ABST
Patent Text Reader

Abstract

A depth estimation method and apparatus, comprising: obtaining a reference image and a source image, wherein the reference image and the source image are images captured from different perspectives for the same scene or the same object; performing feature extraction on the reference image and the source image to obtain local features of the reference image and the source image; performing information interaction processing on the local features of the reference image and the source image to obtain global features of the reference image and the source image; obtaining external parameters of the reference image and the source image, and obtaining a depth map of the reference image based on the global features and external parameters of the reference image and the source image. The depth estimation method and apparatus can utilize internal attention and external attention to aggregate contextual information within and between images, so that the depth estimation method can improve the accuracy of depth estimation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of multi-view stereo (MVS) vision technology, and more specifically to a depth estimation method and a depth estimation device. Background Art

[0002] Multiple View Stereo (MVS) has long been a hot topic in computer vision research. Its goal is to establish dense correspondences from multiple images with known camera poses, thereby producing a dense 3D point cloud reconstruction. MVS typically uses a two-step process to reconstruct a dense 3D model of a scene: first, a depth map is estimated for each image, and then these depth maps are fused to form a unified point cloud representation. Depth map estimation is the key to this process.

[0003] Currently, depth map estimation is based on local features. This makes it difficult to achieve high-precision depth map estimation results for challenging areas in MVS, such as those with poor texture, repetitive textures, and non-Lambertian surfaces. Furthermore, when calculating the matching cost, the features to be compared are simply extracted from each image, failing to consider the underlying correspondence between images. This also affects the accuracy of depth map estimation. Summary of the Invention

[0004] According to one aspect of the present application, a method for depth estimation is provided, the method comprising: acquiring a reference image and a source image, wherein the reference image and the source image are images captured from different perspectives for the same scene or the same object; performing feature extraction on the reference image and the source image respectively to obtain local features of the reference image and the source image; performing information interaction processing on the local features of the reference image and the source image respectively to obtain global features of the reference image and the source image respectively; acquiring external parameters of the reference image and the source image respectively, and obtaining a depth map of the reference image based on the global features and external parameters of the reference image and the source image respectively.

[0005] In one embodiment of the present application, the performing information interaction processing on the local features of the reference image and the source image to obtain the global features of the reference image and the source image includes: performing global information interaction processing on the local features of the reference image and the source image to obtain the global features of the reference image and the source image; or performing semi-global information interaction processing on the local features of the reference image and the source image to obtain the semi-global features of the reference image and the source image, and performing global information interaction processing on the semi-global features of the reference image and the source image to obtain the global features of the reference image and the source image.

[0006] In one embodiment of the present application, the global information interaction processing is performed by a feature matching deformation module, and the feature matching deformation module includes a self-attention module for performing intra-image information interaction processing and a cross-attention module for performing inter-image information interaction processing; the global information interaction processing of the local features of the reference image and the source image to obtain the global features of the reference image and the source image includes: inputting the local features of the reference image into the self-attention module to obtain the global features of the reference image; inputting the local features of the source image into the self-attention module, and the result obtained is used as the first input of the cross-attention module; using the processing result of the self-attention module on the local features of the reference image as the second input of the cross-attention module, and the output of the cross-attention module is the global feature of the source image.

[0007] In one embodiment of the present application, the feature matching deformation module includes a plurality of cascade modules, each of the cascade modules includes the self-attention module and the cross-attention module; the global features of the reference image are obtained by: inputting the local features of the reference image into the self-attention module in the first cascade module of the plurality of cascade modules to obtain the first-level processing results of the local features of the reference image; inputting the first-level processing results into the self-attention module in the second cascade module of the plurality of cascade modules to obtain the second-level processing results of the local features of the reference image; and so on, until the self-attention module in the last cascade module of the plurality of cascade modules outputs the global features of the reference image; the source image The global features are obtained in the following manner: the local features of the source image are input into the self-attention module in the first cascade module, and the obtained results are input together with the first-level processing results of the local features of the reference image into the cross-attention module in the first cascade module to obtain the first-level processing results of the source image; the first-level processing results of the source image are input into the self-attention module in the second cascade module, and the obtained results are input together with the second-level processing results of the local features of the reference image into the cross-attention module in the second cascade module to obtain the second-level processing results of the source image; and so on, until the cross-attention module in the last cascade module in the multiple cascade modules outputs the global features of the source image.

[0008] In one embodiment of the present application, the method further includes: position encoding each pixel in the local features of the reference image and the source image, obtaining the position-encoded local features and inputting them into the feature matching and deformation module for obtaining the global features of the reference image and the source image.

[0009] In one embodiment of the present application, the semi-global information interaction processing is performed by an adaptive receptive field module, and the global information interaction processing is performed by a feature matching deformation module. The adaptive receptive field module includes a deformable convolution block for performing deformable convolution processing, and the feature matching deformation module includes a self-attention module for performing intra-image information interaction processing and a cross-attention module for performing inter-image information interaction processing; the semi-global information interaction processing is performed on the local features of the reference image and the source image to obtain the semi-global features of the reference image and the source image, including: inputting the local features of the reference image and the source image into the deformable shaped convolution block to obtain the semi-global features of the reference image and the source image respectively; the global information interaction processing of the semi-global features of the reference image and the source image respectively to obtain the global features of the reference image and the source image respectively, including: inputting the semi-global features of the reference image into the self-attention module to obtain the global features of the reference image; inputting the semi-global features of the source image into the self-attention module to obtain the first input of the cross-attention module; using the processing result of the self-attention module on the semi-global features of the reference image as the second input of the cross-attention module, and the output of the cross-attention module is the global features of the source image.

[0010] In one embodiment of the present application, the feature matching deformation module includes a plurality of cascade modules, each of the cascade modules includes the self-attention module and the cross-attention module; the global features of the reference image are obtained in the following manner: the semi-global features of the reference image are input into the self-attention module in the first cascade module of the plurality of cascade modules to obtain the first-level processing result of the semi-global features of the reference image; the first-level processing result is input into the self-attention module in the second cascade module of the plurality of cascade modules to obtain the second-level processing result of the semi-global features of the reference image; and so on, until the self-attention module in the last cascade module of the plurality of cascade modules outputs the global features of the reference image; the global features of the source image are obtained by: inputting the semi-global features of the reference image into the self-attention module in the first cascade module of the plurality of cascade modules to obtain the second-level processing result of the semi-global features of the reference image; and so on. The local features are obtained in the following manner: the semi-global features of the source image are input into the self-attention module in the first cascade module, and the obtained results are input together with the first-level processing results of the semi-global features of the reference image into the cross-attention module in the first cascade module to obtain the first-level processing results of the source image; the first-level processing results of the source image are input into the self-attention module in the second cascade module, and the obtained results are input together with the second-level processing results of the semi-global features of the reference image into the cross-attention module in the second cascade module to obtain the second-level processing results of the source image; and so on, until the cross-attention module in the last cascade module in the multiple cascade modules outputs the global features of the source image.

[0011] In one embodiment of the present application, the method further includes: position encoding each pixel in the semi-global features of the reference image and the source image, obtaining the position-encoded semi-global features and inputting them into the feature matching and deformation module to obtain the global features of the reference image and the source image.

[0012] In one embodiment of the present application, the depth map of the reference image is obtained based on the global features and external parameters of the reference image and the source image, including: based on the external parameters of the reference image and the source image, transforming the global features of the reference image and the source image to a reference feature plane via a microbending operation to obtain feature bodies of the reference image and the source image; correlating the feature body of the reference image with the feature body of the source image to obtain a correlation body, regularizing the correlation body to obtain a probability body, and obtaining the depth map of the reference image based on the probability body.

[0013] In one embodiment of the present application, the feature extraction is performed by a feature pyramid module, which outputs multiple local features of different resolutions for the reference image and multiple local features of different resolutions for the source image, so that the feature matching deformation module outputs multiple global features of different resolutions for the reference image and multiple global features of different resolutions for the source image; the global features of the minimum resolution of the reference image and the source image are used to generate a depth map of the minimum resolution of the reference image, and the depth map of the minimum resolution is used as an initial stage depth map to act on the microbending operation to obtain a depth map of the next smallest resolution of the reference image, and so on, until a depth map of the maximum resolution of the reference image is obtained.

[0014] In one embodiment of the present application, the multiple local features of different resolutions are all input into the feature matching and deformation module to obtain multiple global features of different resolutions; or only the local feature of the minimum resolution among the multiple local features of different resolutions is input into the feature matching and deformation module to obtain the global feature of the minimum resolution, and the global features of other resolutions are obtained based on the fusion of the global feature of the minimum resolution and the local features of other resolutions.

[0015] According to another aspect of the present application, a depth estimation device is also provided, which includes a feature extraction module, a feature matching deformation module and a depth estimation module, wherein: the feature extraction module is used to obtain a reference image and a source image, and perform feature extraction on the reference image and the source image respectively to obtain local features of the reference image and the source image, wherein the reference image and the source image are images collected from different perspectives for the same scene or the same object; the feature matching deformation module is used to perform information interaction processing on the local features of the reference image and the source image respectively to obtain global features of the reference image and the source image respectively; the depth estimation module is used to obtain external parameters of the reference image and the source image respectively, and output a depth map of the reference image based on the global features and external parameters of the reference image and the source image respectively.

[0016] In one embodiment of the present application, the device also includes an adaptive receptive field module, which is used to perform semi-global information interaction processing on the local features of the reference image and the source image to obtain the semi-global features of the reference image and the source image; the feature matching deformation module is also used to perform global information interaction processing on the semi-global features of the reference image and the source image to obtain the global features of the reference image and the source image.

[0017] According to another aspect of the present application, a depth estimation device is provided, which includes a memory and a processor, wherein the memory stores a computer program executed by the processor, and when the computer program is executed by the processor, the processor executes the above-mentioned depth estimation method.

[0018] According to another aspect of the present application, a storage medium is provided, on which a computer program is stored. When the computer program is executed, the depth estimation method described above is executed.

[0019] According to another aspect of the present application, a computer program product is provided, comprising a computer program / instruction, wherein the computer program / instruction implements the above-mentioned depth estimation method when executed by a processor.

[0020] According to the embodiments of the present application, the depth estimation method and depth estimation device can utilize internal attention and external attention to aggregate contextual information within and between images, so that the depth estimation method can improve the accuracy of depth estimation and obtain high-precision depth estimation results for areas such as weak textures, repeated textures, and non-Lambertian surfaces. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The above and other purposes, features, and advantages of the present application will become more apparent through a more detailed description of the embodiments of the present application in conjunction with the accompanying drawings. The accompanying drawings are intended to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the drawings, the same reference numerals generally represent the same components or steps.

[0022] Figure 1 A schematic block diagram of an example electronic device for implementing a depth estimation method and a depth estimation apparatus according to an embodiment of the present invention is shown.

[0023] Figure 2 A schematic flowchart of a depth estimation method according to an embodiment of the present application is shown.

[0024] Figure 3 A schematic diagram of a feature matching deformation module in a depth estimation method according to an embodiment of the present application and a processing flow chart thereof are shown.

[0025] Figure 4 A flow chart of model processing used in a depth estimation method according to an embodiment of the present application is shown.

[0026] Figure 5 A schematic block diagram of a depth estimation device according to an embodiment of the present application is shown.

[0027] Figure 6A schematic block diagram of a depth estimation device according to another embodiment of the present application is shown. DETAILED DESCRIPTION

[0028] In order to make the purpose, technical solutions and advantages of the present application more apparent, the following is a detailed description of example embodiments of the present application with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the example embodiments described herein. Based on the embodiments of the present application described in this application, all other embodiments obtained by those skilled in the art without creative work should fall within the scope of protection of this application.

[0029] In recent years, significant progress has been made in AI-based research on computer vision, deep learning, machine learning, image processing, and image recognition. Artificial Intelligence (AI) is an emerging science and technology that studies and develops theories, methods, technologies, and application systems for simulating and extending human intelligence. AI is a comprehensive discipline encompassing numerous technologies, including chips, big data, cloud computing, the Internet of Things, distributed storage, deep learning, machine learning, and neural networks. Computer vision, a key branch of AI, specifically enables machines to understand the world. Computer vision technologies typically include face recognition, liveness detection, fingerprint recognition and anti-counterfeiting verification, biometric recognition, face detection, pedestrian detection, object detection, pedestrian recognition, image processing, image recognition, image semantic understanding, image retrieval, text recognition, video processing, video content recognition, behavior recognition, 3D reconstruction, virtual reality, augmented reality, simultaneous localization and mapping (SLAM), computational photography, and robotic navigation and positioning. With the research and advancement of artificial intelligence technology, this technology has been applied in many fields, such as security, urban management, traffic management, building management, park management, facial access, facial attendance, logistics management, warehouse management, robots, intelligent marketing, computational photography, mobile phone imaging, cloud services, smart homes, wearable devices, unmanned driving, autonomous driving, smart medical care, facial payment, facial unlocking, fingerprint unlocking, identity verification, smart screens, smart TVs, cameras, mobile Internet, live streaming, beauty, makeup, medical beauty, smart temperature measurement and other fields.

[0030] Below, refer to Figure 1 An example electronic device 100 for implementing the depth estimation method and apparatus according to the embodiments of the present invention will be described.

[0031] like Figure 1As shown, the electronic device 100 includes one or more processors 102, one or more storage devices 104, an input device 106, and an output device 108, which are interconnected via a bus system 110 and / or other forms of connection mechanisms (not shown). It should be noted that Figure 1 The components and structure of the electronic device 100 shown are merely exemplary and non-limiting. The electronic device may also have other components and structures as needed.

[0032] The processor 102 may be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 100 to perform desired functions.

[0033] The storage device 104 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), a hard disk, a flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 102 may execute the program instructions to implement the client functions and / or other desired functions in the embodiments of the present invention (implemented by the processor) described below. Various applications and various data may also be stored in the computer-readable storage medium, such as various data used and / or generated by the application.

[0034] The input device 106 may be a device used by a user to input instructions, and may include one or more of a keyboard, a mouse, a microphone, a touch screen, etc. In addition, the input device 106 may also be any interface for receiving information.

[0035] The output device 108 can output various information (such as images or sounds) to the outside (such as a user), and can include one or more of a display, a speaker, etc. In addition, the output device 108 can also be any other device with an output function.

[0036] Exemplarily, an example electronic device for implementing the method and apparatus for training an image processing model according to an embodiment of the present invention may be implemented as a terminal such as a smart phone, a tablet computer, a camera, or the like.

[0037] Below, we will refer to Figure 2 Describe the depth estimation method 200 according to an embodiment of the present application. Figure 2 As shown, the depth estimation method 200 may include the following steps:

[0038] In step S210 , a reference image and a source image are acquired, wherein the reference image and the source image are images captured from different perspectives for the same scene or the same object.

[0039] In step S220 , feature extraction is performed on the reference image and the source image respectively to obtain local features of the reference image and the source image.

[0040] In step S230, information interaction processing is performed on local features of the reference image and the source image to obtain global features of the reference image and the source image.

[0041] In step S240 , extrinsic parameters of the reference image and the source image are obtained, and a depth map of the reference image is obtained based on the global features and extrinsic parameters of the reference image and the source image.

[0042] In an embodiment of the present application, the images acquired in step S210 are multiple images from different perspectives captured for the same scene or the same object. For any one of these images, it can be used as a reference image, and the remaining images can be used as source images to calculate the depth map of the reference image. In this way, each image can be used as a reference image, and the depth map of the reference image can be calculated in combination with other source images. Then, the depth maps of the reference images are fused to obtain a three-dimensional model of the scene or object. In particular, when calculating the depth map of each reference image, the reference image and other source images are obtained. First, the local features of the reference image and the source image are obtained by feature extraction (or local feature maps). Here, it should be understood that the local features (maps) can only reflect the relationship between a pixel and its adjacent pixels in the image (reference image or source image). Then, the local features of the reference image and the source image are subjected to information interaction processing to obtain the global features of the reference image and the source image (or global feature maps). Here, information interaction processing of the local features of the reference image and the source image can refer to: processing the local features of the reference image and the source image based on the self-attention mechanism, and processing their processing results based on the cross-attention mechanism to obtain the global features of the reference image and the source image. It should be understood that the global feature (map) can reflect the relationship between a pixel and any other pixel in the image (reference image or source image). Finally, based on the global features of the reference image and the source image and the external parameters of the reference image and the source image, the depth map of the reference image is calculated.

[0043] Therefore, the depth estimation method according to the embodiment of the present application can utilize internal attention (self-attention) and external attention (cross-attention) to aggregate contextual information (global features) within the image (within the reference image) and between images (between the reference image and the source image), so that the depth estimation method can improve the accuracy of depth estimation, and can also obtain high-precision depth estimation results for areas such as weak textures, repeated textures and non-Lambertian surfaces.

[0044] In an embodiment of the present application, the feature extraction module used in step S220 can be a feature pyramid network (FPN), which (for example, includes three convolution blocks, each of which contains three layers of ordinary convolution) can output multiple local features of different resolutions for the reference image, and output multiple local features of different resolutions for the source image, so that in step S230, multiple global features of different resolutions are output for the reference image, and multiple global features of different resolutions are output for the source image. Based on this, the global features of the minimum resolution of the reference image and the source image in step S240 are used to generate a depth map of the minimum resolution of the reference image, and the depth map of the minimum resolution is used as the initial stage depth map to act on the microbending operation to obtain a depth map of the second smallest resolution of the reference image, and so on, until a depth map of the maximum resolution of the reference image is obtained. In this embodiment, the features of the previous stage can be used as reference prior information to guide the next stage to output a more refined depth map.

[0045] The following describes how to generate global features in conjunction with different embodiments. For the sake of brevity, the following description does not mention local features (and semi-global features to be described later) and global features of different resolutions, because the processing of features of different resolutions is similar.

[0046] In one embodiment of the present application, performing information exchange processing on local features of the reference image and the source image to obtain global features of the reference image and the source image in step S230 may include performing global information exchange processing on the local features of the reference image and the source image to obtain global features of the reference image and the source image. In this embodiment, the information exchange processing on the local features of the reference image and the source image is global, and thus the global features of the reference image and the source image can be directly obtained.

[0047] In an embodiment of the present application, global information interaction processing can be performed by a feature matching transformer module (FeatureMatching Transformer, referred to as FMT), which can include a self-attention module for performing intra-image information interaction processing and a cross-attention module for performing inter-image information interaction processing.

[0048] Based on this, the aforementioned global information interaction processing of the local features of the reference image and the source image to obtain the global features of the reference image and the source image can include: inputting the local features of the reference image into the self-attention module to obtain the global features of the reference image; inputting the local features of the source image into the self-attention module, and using the result obtained as the first input of the cross-attention module; using the processing result of the self-attention module on the local features of the reference image as the second input of the cross-attention module, and the output of the cross-attention module is the global features of the source image. In this embodiment, the global features of the reference image are obtained through the self-attention module based on the local features of the reference image; and the global features of the source image are obtained through the cross-attention module based on the local features of the source image and the processing result of the local features of the reference image by the self-attention module.

[0049] In a further embodiment of the present application, the above-mentioned feature matching deformation module may include multiple cascade modules, each of which includes the self-attention module and the cross-attention module. Based on this, the global features of the reference image can be obtained in the following manner: the local features of the reference image are input into the self-attention module in the first cascade module among the multiple cascade modules to obtain the first-level processing results of the local features of the reference image; the first-level processing results are input into the self-attention module in the second cascade module among the multiple cascade modules to obtain the second-level processing results of the local features of the reference image; and so on, until the self-attention module in the last cascade module among the multiple cascade modules outputs the global features of the reference image. The global features of the source image can be obtained in the following manner: the local features of the source image are input into the self-attention module in the first cascade module, and the obtained results are input together with the first-level processing results of the local features of the reference image into the cross-attention module in the first cascade module to obtain the first-level processing results of the source image; the first-level processing results of the source image are input into the self-attention module in the second cascade module, and the obtained results are input together with the second-level processing results of the local features of the reference image into the cross-attention module in the second cascade module to obtain the second-level processing results of the source image; and so on, until the cross-attention module in the last cascade module in the multiple cascade modules outputs the global features of the source image. In this embodiment, the local features of the reference image and the source image are processed sequentially by multiple attention blocks, so that high-precision global features of the reference image and the source image can be obtained.

[0050] In a further embodiment of the present application, method 200 may further include: position encoding each pixel in the local features of each of the reference image and the source image, obtaining the position-encoded local features and inputting them into the feature matching and deformation module for obtaining the global features of each of the reference image and the source image. In this embodiment, when obtaining the global features of each of the reference image and the source image, in addition to the local features of the reference image and the source image themselves, the global features are also obtained based on the position encoding information of each pixel in their local features. This can enhance position consistency and make the feature matching and deformation module robust to local features of different resolutions.

[0051] In another embodiment of the present application, the information interaction processing of the local features of the reference image and the source image to obtain the global features of the reference image and the source image in step S230 may include: performing semi-global information interaction processing on the local features of the reference image and the source image to obtain the semi-global features (or semi-global feature maps) of the reference image and the source image, and performing global information interaction processing on the semi-global features of the reference image and the source image to obtain the global features of the reference image and the source image. Here, it should be understood that the local feature (map) can refer to the information of each pixel in the feature map that can reflect the relationship between a pixel and its neighboring pixels in the image and the relationship with pixels farther away. In this embodiment, the local features of the reference image and the source image are first processed into semi-global features through semi-global information interaction processing, and then the semi-global features of the reference image and the source image are processed into global features through global information interaction processing. This processing flow from local to semi-global to global makes the final global features more accurate because of the addition of a transition stage (semi-global features).

[0052] In an embodiment of the present application, the semi-global information interaction processing can be performed by an adaptive receptive field module (Adaptive Receptive Field, abbreviated as ARF), and the global information interaction processing can be performed by a feature matching deformation module. The adaptive receptive field module may include a deformable convolution block for performing deformable convolution processing. The feature matching deformation module may include a self-attention module for performing intra-image information interaction processing and a cross-attention module for performing inter-image information interaction processing. In one example, the adaptive receptive field module may include two deformable convolution blocks, and each deformable convolution block may include three layers of deformable convolution. Since the adaptive receptive field module includes a deformable convolution block, the adaptive receptive field module can learn additional offsets of the sampling position and can adaptively enlarge the receptive field according to the local context, thereby obtaining semi-global features. The applicant has experimentally proved that the semi-global features processed by the adaptive receptive field module are more friendly to the feature matching deformation module, so that the feature matching deformation module can output more accurate global features, thereby further improving the depth estimation accuracy of the reference image.

[0053] Based on this, the aforementioned semi-global information interaction processing of the local features of the reference image and the source image to obtain the semi-global features of the reference image and the source image can include: inputting the local features of the reference image and the source image into a deformable convolution block to obtain the semi-global features of the reference image and the source image. The aforementioned global information interaction processing of the semi-global features of the reference image and the source image to obtain the global features of the reference image and the source image can include: inputting the semi-global features of the reference image into a self-attention module to obtain the global features of the reference image; inputting the semi-global features of the source image into the self-attention module, and using the obtained result as the first input of the cross-attention module; using the processing result of the self-attention module on the semi-global features of the reference image as the second input of the cross-attention module, and the output of the cross-attention module is the global features of the source image. In this embodiment, the global features of the reference image are obtained by the self-attention module based on the semi-global features of the reference image; and the global features of the source image are obtained by the cross-attention module based on the semi-global features of the source image and the processing result of the self-attention module on the semi-global features of the reference image.

[0054] In a further embodiment of the present application, the above-mentioned feature matching deformation module may include multiple cascade modules, each of which includes the self-attention module and the cross-attention module. Based on this, the global features of the reference image can be obtained in the following manner: the semi-global features of the reference image are input into the self-attention module in the first cascade module among the multiple cascade modules to obtain the first-level processing result of the semi-global features of the reference image; the first-level processing result is input into the self-attention module in the second cascade module among the multiple cascade modules to obtain the second-level processing result of the semi-global features of the reference image; and so on, until the self-attention module in the last cascade module among the multiple cascade modules outputs the global features of the reference image. The global features of the source image can be obtained in the following manner: the semi-global features of the source image are input into the self-attention module in the first cascade module, and the obtained results are input together with the first-level processing results of the semi-global features of the reference image into the cross-attention module in the first cascade module to obtain the first-level processing results of the source image; the first-level processing results of the source image are input into the self-attention module in the second cascade module, and the obtained results are input together with the second-level processing results of the semi-global features of the reference image into the cross-attention module in the second cascade module to obtain the second-level processing results of the source image; and so on, until the cross-attention module in the last cascade module in the multiple cascade modules outputs the global features of the source image. In this embodiment, the semi-global features of the reference image and the source image are processed sequentially by multiple attention blocks, so that high-precision global features of the reference image and the source image can be obtained.

[0055] In a further embodiment of the present application, method 200 may further include: position encoding each pixel in the semi-global features of each of the reference image and the source image, and inputting the position-encoded semi-global features into the feature matching and deformation module to obtain the global features of each of the reference image and the source image. In this embodiment, when obtaining the global features of each of the reference image and the source image, in addition to the semi-global features themselves, the position encoding information of each pixel in these semi-global features is also used. This can enhance position consistency and make the feature matching and deformation module robust to semi-global features of different resolutions.

[0056] The following combination Figure 3 To understand the operation process of the feature matching deformation module. Figure 3 As shown, F0 is the local feature (or semi-global feature) of the reference image, {F i} Local features (or semi-global features) of the source image, where i ranges from 1 to N-1, and N is a natural number greater than 1. Figure 3 In , it is shown as the local features (or semi-global features) of the two source images, that is, N is 3. Figure 3 As shown, the local feature (or semi-global feature) F0 of the reference image and the local feature (or semi-global feature) F i After positional encoding, multiple feature vectors are obtained (such as Figure 3 The long strip shape shown in the figure is input into the FMT (also called attention block) including the self-attention module (Intra-Attention) and the cross-attention module (Inter-Attention). The feature vector of the reference image is input into the self-attention module in the first cascade module of Na cascade modules to obtain the first-level processing result of the feature vector of the reference image; the first-level processing result is input into the self-attention module in the second cascade module of Na cascade modules to obtain the second-level processing result of the feature vector of the reference image; and so on, until the self-attention module in the last cascade module of Na cascade modules outputs the global features of the reference image. The feature vector of the source image is input into the self-attention module in the first cascade module, and the result obtained together with the first-level processing result of the feature vector of the reference image is input into the cross-attention module in the first cascade module to obtain the first-level processing result of the source image; the first-level processing result of the source image is input into the self-attention module in the second cascade module, and the result obtained together with the second-level processing result of the feature vector of the reference image is input into the cross-attention module in the second cascade module to obtain the second-level processing result of the source image; and so on, until the cross-attention module in the last cascade module in Na cascade modules outputs the global features of the source image.

[0057] The above describes how to generate global features from the perspective of features of one resolution. In one embodiment of the present application, multiple local features (or semi-global features) of different resolutions can be input into the feature matching deformation module to obtain multiple global features of different resolutions. In this embodiment, the feature matching deformation module processes the local features (or semi-global features) of different resolutions to obtain accurate calculation results. In another embodiment of the present application, only the local features (or semi-global features) of the minimum resolution among the multiple local features of different resolutions can be input into the feature matching deformation module to obtain the global features of the minimum resolution, and the global features of other resolutions are obtained based on the fusion of the global features of the minimum resolution and the local features (or semi-global features) of other resolutions. In this embodiment, the feature matching deformation module only processes the local features (or semi-global features) of the minimum resolution, which can reduce the amount of calculation, thereby reducing the required computing resources.

[0058] In an embodiment of the present application, step S240 obtains a depth map of the reference image based on the respective global features and respective extrinsic parameters of the reference image and the source image (i.e., the posture of the camera that captured the reference image and the source image), which may include: based on the respective extrinsic parameters of the reference image and the source image, transforming the respective global features of the reference image and the source image to the reference feature plane via a differentiable warping operation to obtain a feature volume of the reference image and the source image; correlating the feature volume of the reference image with the feature volume of the source image to obtain a correlation volume, regularizing the correlation volume to obtain a probability volume, and obtaining a depth map of the reference image based on the probability volume.

[0059] In this embodiment, the global features of the reference image and the source image obtained in step S230 can both be three-dimensional features (C×H×W, where C is the number of feature channels and W and H are the width and height of the feature). The three-dimensional features are transformed onto the reference feature plane through a microbending operation to form a four-dimensional feature tensor (C×H×W×D, where D is the interval data divided by the depth range), i.e., a feature volume. The feature volume of the reference image and the feature volume of the source image are correlated to obtain a correlation volume. The correlation volume is then regularized using a three-dimensional convolutional neural network to obtain a probability volume. The probability volume is a 1×H×W three-dimensional feature, and each position in the H×W plane of the feature corresponds to a probability vector, which is the probability value of various depths corresponding to that position. Finally, a depth map of the reference image is obtained based on the probability volume. That is, a winner-take-all operation is performed on each of the aforementioned positions. That is, for each position, the depth value with the maximum probability value in the probability vector is used as the predicted depth value at that position, thereby obtaining a depth map. For example, the probability vector at a certain location is [0.2, 0.5, 0.3], and the corresponding depth interval vector is [1-2m, 2-3m, 3-4m]. Therefore, the probability vector contains three elements: 0, 2, 0.5, and 0.3, each corresponding to a depth interval. The largest element, 0.5, corresponds to the depth interval of 2-3m, which is the predicted depth value at the current location.

[0060] The above describes in detail the depth estimation method according to the embodiment of the present application. Assuming that the above method can be implemented by a model, the model can be called TransMVSNet model. Figure 4 To summarize the model's processing flow, Figure 4 In the example shown, the process may include a combination of the above embodiments. Figure 4 As shown, obtain the reference image I0 and the source image {I i}, where i ranges from 1 to N-1, and N is a natural number greater than 1. Figure 4 In the figure, two source images are shown, that is, N is 3. When obtaining the reference image I0 and the source image {I i}, they are input into FPN, and local features of three resolutions (as an example) are obtained respectively. Each local feature is input into ARF to obtain semi-global features of three resolutions. Among them, only the semi-global features of the minimum resolution are input into FMT including the intra-attention module and the inter-attention module, and the FMT outputs the global features. The global features of other resolutions are obtained by fusing the global features of the minimum resolution with the semi-global features of the resolution (such as Figure 4 Then, for the global feature with the minimum resolution, it is subjected to the W operation (i.e., the microbending operation) to obtain the feature volume, Figure 4 The three small cubes on the right side of W in the figure are the feature bodies of the reference image (which can be called the reference feature body) and the feature bodies of the two source images (which can be called the source feature bodies). Then, the feature body of the reference image is correlated with the feature bodies of the two source images respectively to obtain two correlation bodies (two small cubes on the right side of the three small cubes). The two correlation bodies are then fused to obtain a fused correlation body (a small cube on the right side of the two small cubes). The fused correlation body is regularized to obtain a probability body (the small cube on the far right). Finally, the probability body is subjected to a winner-takes-all operation to obtain a depth map of the minimum resolution (i.e., the depth map of the initial stage). The calculation of depth maps of other resolutions combines the depth maps of the previous stage, and finally obtains a depth map (Depth Map) with the same size as the reference image (i.e., the same resolution as the reference image).

[0061] Based on the above description, the depth estimation method according to the embodiment of the present application can utilize internal attention and external attention to aggregate contextual information within and between images, so that the depth estimation method can improve the accuracy of depth estimation and obtain high-precision depth estimation results for areas with weak textures, repeated textures, and non-Lambertian surfaces. After experimental verification by the applicant, the method of the present application has achieved state-of-the-art performance on the DTU dataset, the Tanks and Temples online test set, and the BlendedMVS dataset, fully demonstrating the advantages of TransMVSNet, such as good results and strong generalization performance.

[0062] The following combination Figures 5 to 6According to another aspect of the present application, a depth estimation device is provided, which can be used to perform the depth estimation method according to the embodiment of the present application described above. Those skilled in the art can understand the structure and specific operation of the depth estimation device according to the embodiment of the present application in combination with the above content. For the sake of brevity, the specific details are not repeated here, and only some main operations are described.

[0063] Figure 5 FIG. 5 shows a schematic block diagram of a depth estimation device 500 according to an embodiment of the present application. Figure 5 As shown, the depth estimation device 500 includes a feature extraction module 510, a feature matching deformation module 520, and a depth estimation module 530. The feature extraction module 510 is used to obtain a reference image and a source image, perform feature extraction on the reference image and the source image respectively, and obtain local features of the reference image and the source image respectively, wherein the reference image and the source image are images captured from different perspectives of the same scene or the same object. The feature matching deformation module 520 is used to perform information exchange processing on the local features of the reference image and the source image respectively, and obtain global features of the reference image and the source image respectively. The depth estimation module 530 is used to obtain external parameters of the reference image and the source image respectively, and output a depth map of the reference image based on the global features and external parameters of the reference image and the source image respectively.

[0064] In an embodiment of the present application, the depth estimation device 500 may include an adaptive receptive field module (not shown), which is used to perform semi-global information interaction processing on the local features of the reference image and the source image to obtain the semi-global features of the reference image and the source image; the feature matching deformation module 520 is also used to perform global information interaction processing on the semi-global features of the reference image and the source image to obtain the global features of the reference image and the source image.

[0065] Figure 6 FIG. 5 shows a schematic block diagram of a depth estimation device 600 according to another embodiment of the present application. Figure 6 As shown, the depth estimation device 600 according to an embodiment of the present application may include a memory 610 and a processor 620. The memory 610 stores a computer program executed by the processor 620. When the computer program is executed by the processor 620, the processor 620 performs the depth estimation method according to the embodiment of the present application described above. Those skilled in the art can understand the specific operation of the depth estimation device 600 according to the embodiment of the present application in combination with the above content. For the sake of brevity, the specific details are not repeated here.

[0066] In addition, according to an embodiment of the present application, a storage medium is further provided, on which program instructions are stored, and when the program instructions are run by a computer or a processor, the corresponding steps of the depth estimation method of the embodiment of the present application are used to execute. The storage medium may include, for example, a memory card of a smart phone, a storage component of a tablet computer, a hard disk of a personal computer, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a portable compact disk read-only memory (CD-ROM), a USB memory, or any combination of the above storage media. The computer-readable storage medium may be any combination of one or more computer-readable storage media.

[0067] According to an embodiment of the present application, a computer program product is also provided, including a computer program / instruction, which implements the above-mentioned depth estimation method when executed by a processor.

[0068] Based on the above description, the depth estimation method and depth estimation device according to the embodiments of the present application can utilize internal attention and external attention to aggregate contextual information within and between images, so that the depth estimation method can improve the accuracy of depth estimation, and can also obtain high-precision depth estimation results for areas such as weak textures, repeated textures, and non-Lambertian surfaces.

[0069] Although example embodiments have been described herein with reference to the accompanying drawings, it should be understood that the above example embodiments are merely illustrative and are not intended to limit the scope of the present application. Various changes and modifications may be made therein by those skilled in the art without departing from the scope and spirit of the present application. All such changes and modifications are intended to be included within the scope of the present application as required by the appended claims.

[0070] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0071] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units described is merely a logical function division. In actual implementation, other division methods may be used, such as combining or integrating multiple units or components into another device, or ignoring or not performing some features.

[0072] In the description provided herein, a large number of specific details are described. However, it is understood that the embodiments of the present application can be practiced without these specific details. In some instances, well-known methods, structures, and techniques are not shown in detail so as not to obscure the understanding of this description.

[0073] Similarly, it should be understood that in order to streamline the present application and aid in understanding one or more of the various inventive aspects, in the description of the exemplary embodiments of the present application, the various features of the present application are sometimes grouped together into a single embodiment, figure, or description thereof. However, this approach of the present application should not be interpreted as reflecting the intention that the application claimed for protection requires more features than those explicitly recited in each claim. More precisely, as reflected in the corresponding claims, the inventive point is that the corresponding technical problem can be solved with fewer features than all the features of a single disclosed embodiment. Therefore, the claims following the detailed description are hereby expressly incorporated into the detailed description, with each claim itself serving as a separate embodiment of the present application.

[0074] It will be understood by those skilled in the art that, except where mutually exclusive, all features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all processes or units of any method or apparatus disclosed herein may be combined in any combination. Unless expressly stated otherwise, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) may be replaced by an alternative feature providing the same, equivalent, or similar purpose.

[0075] Furthermore, those skilled in the art will appreciate that although some embodiments described herein include certain features included in other embodiments but not other features, combinations of features from different embodiments are intended to be within the scope of this application and to form different embodiments. For example, in the claims, any of the claimed embodiments may be used in any combination.

[0076] The various component embodiments of the present application can be implemented in hardware, or in a software module running on one or more processors, or in a combination thereof. Those skilled in the art will appreciate that a microprocessor or digital signal processor (DSP) can be used in practice to implement some or all of the functions of some modules according to the embodiments of the present application. The application can also be implemented as a part or all of a device program (e.g., a computer program and a computer program product) for performing the method described herein. Such a program implementing the present application can be stored on a computer-readable medium, or can have the form of one or more signals. Such a signal can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.

[0077] It should be noted that the above embodiments illustrate rather than limit the present application, and that a person skilled in the art may devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference symbols placed between brackets should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present application may be implemented by means of hardware comprising several different elements and by means of appropriately programmed computers. In a unit claim enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third etc. does not indicate any order. These words may be interpreted as names.

[0078] The above description is merely a specific embodiment or illustration of a specific embodiment of the present application, and the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present application should be included in the scope of protection of the present application. The scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A depth estimation method, characterized in that: The method comprises: Acquire a reference image and a source image, wherein the reference image and the source image are images captured from different perspectives with respect to the same scene or the same object; Performing feature extraction on the reference image and the source image respectively to obtain local features of the reference image and the source image; Performing information interaction processing on local features of the reference image and the source image to obtain global features of the reference image and the source image; Obtaining extrinsic parameters of the reference image and the source image, and obtaining a depth map of the reference image based on the global features and extrinsic parameters of the reference image and the source image; The performing information interaction processing on the local features of the reference image and the source image to obtain the global features of the reference image and the source image, includes: Based on the local features of the reference image, the global features of the reference image are obtained through a self-attention module; and based on the local features of the source image and the global features of the reference image obtained through the self-attention module, the global features of the source image are obtained through a cross-attention module.

2. The method according to claim 1, characterized in that The performing information interaction processing on the local features of the reference image and the source image to obtain the global features of the reference image and the source image, includes: Performing global information interaction processing on local features of the reference image and the source image to obtain global features of the reference image and the source image; or Performing semi-global information interaction processing on the local features of the reference image and the source image to obtain the semi-global features of the reference image and the source image, and performing semi-global information interaction processing on the local features of the reference image and the source image to obtain the semi-global features of the reference image and the source image. The global information interaction processing is performed on the features to obtain the global features of the reference image and the source image respectively.

3. The method according to claim 2, characterized in that The global information interaction processing is performed by a feature matching and deformation module, which includes a self-attention module for performing intra-image information interaction processing and a cross-attention module for performing inter-image information interaction processing; The performing global information interaction processing on the local features of the reference image and the source image to obtain the global features of the reference image and the source image includes: Inputting the local features of the reference image into the self-attention module to obtain the global features of the reference image; Inputting the local features of the source image into the self-attention module, and the obtained result is used as the first input of the cross-attention module; The processing result of the self-attention module on the local features of the reference image is used as the second input of the cross-attention module, and the output of the cross-attention module is the global features of the source image.

4. The method according to claim 3, characterized in that The feature matching deformation module includes a plurality of cascade modules, each of the cascade modules includes the self-attention module and the cross-attention module; The global features of the reference image are obtained by: inputting the local features of the reference image into a self-attention module in a first cascade module among the multiple cascade modules to obtain a first-level processing result of the local features of the reference image; inputting the first-level processing result into a self-attention module in a second cascade module among the multiple cascade modules to obtain a second-level processing result of the local features of the reference image; and so on, until the self-attention module in the last cascade module among the multiple cascade modules outputs the global features of the reference image; The global features of the source image are obtained in the following manner: the local features of the source image are input into the self-attention module in the first cascade module, and the obtained result is input together with the first-level processing result of the local features of the reference image into the cross-attention module in the first cascade module to obtain the first-level processing result of the source image; the first-level processing result of the source image is input into the self-attention module in the second cascade module, and the obtained result is input together with the second-level processing result of the local features of the reference image into the cross-attention module in the second cascade module to obtain the second-level processing result of the source image; and so on, until the cross-attention module in the last cascade module among the multiple cascade modules outputs the global features of the source image.

5. The method according to claim 3 or 4, characterized in that The method further comprises: Each pixel in the local features of the reference image and the source image is position-encoded, and the local features obtained after position encoding are input into the feature matching and deformation module to obtain the global features of the reference image and the source image.

6. The method according to claim 2, characterized in that The semi-global information interaction processing is performed by an adaptive receptive field module, and the global information interaction processing is performed by a feature matching deformation module, wherein the adaptive receptive field module includes a deformable convolution block for performing deformable convolution processing, and the feature matching deformation module includes a self-attention module for performing intra-image information interaction processing and a cross-attention module for performing inter-image information interaction processing; The performing semi-global information interaction processing on the local features of the reference image and the source image to obtain the semi-global features of the reference image and the source image, comprising: inputting the local features of the reference image and the source image into the deformable convolution block to obtain the semi-global features of the reference image and the source image; The global information interaction processing is performed on the semi-global features of each of the reference image and the source image to obtain the global features of each of the reference image and the source image, including: inputting the semi-global features of the reference image into the self-attention module to obtain the global features of the reference image; inputting the semi-global features of the source image into the self-attention module, and the obtained result is the first input of the cross-attention module; and using the processing result of the semi-global features of the reference image by the self-attention module as the second input of the cross-attention module, and the output of the cross-attention module is the global features of the source image.

7. The method according to claim 6, characterized in that The feature matching deformation module includes a plurality of cascade modules, each of the cascade modules includes the self-attention module and the cross-attention module; The global features of the reference image are obtained by: inputting the semi-global features of the reference image into a self-attention module in a first cascade module among the multiple cascade modules to obtain a first-level processing result of the semi-global features of the reference image; inputting the first-level processing result into a self-attention module in a second cascade module among the multiple cascade modules to obtain a second-level processing result of the semi-global features of the reference image; and so on, until the self-attention module in the last cascade module among the multiple cascade modules outputs the global features of the reference image; The global features of the source image are obtained in the following manner: the semi-global features of the source image are input into the self-attention module in the first cascade module, and the obtained results are input together with the first-level processing results of the semi-global features of the reference image into the cross-attention module in the first cascade module to obtain the first-level processing results of the source image; the first-level processing results of the source image are input into the self-attention module in the second cascade module, and the obtained results are input together with the second-level processing results of the semi-global features of the reference image into the cross-attention module in the second cascade module to obtain the second-level processing results of the source image; and so on, until the cross-attention module in the last cascade module among the multiple cascade modules outputs the global features of the source image.

8. The method according to claim 6 or 7, characterized in that The method further comprises: Position encoding is performed on each pixel in the semi-global features of the reference image and the source image, and the obtained position-encoded semi-global features are input into the feature matching and deformation module to obtain the global features of the reference image and the source image.

9. The method according to claim 3 or 6, characterized in that The obtaining of the depth map of the reference image based on the global features and extrinsic parameters of the reference image and the source image includes: Based on the extrinsic parameters of the reference image and the source image, the global features of the reference image and the source image are transformed onto a reference feature plane via a differentiable bending operation to obtain feature volumes of the reference image and the source image; The feature volume of the reference image and the feature volume of the source image are correlated to obtain a correlation volume, the correlation volume is regularized to obtain a probability volume, and a depth map of the reference image is obtained based on the probability volume.

10. The method according to claim 9, characterized in that The feature extraction is performed by a feature pyramid module, which outputs a plurality of local features of different resolutions for the reference image and a plurality of local features of different resolutions for the source image, so that the feature matching and deformation module outputs a plurality of global features of different resolutions for the reference image and a plurality of global features of different resolutions for the source image; The global features of the minimum resolution of each of the reference image and the source image are used to generate a depth map of the minimum resolution of the reference image. The depth map of the minimum resolution is used as an initial-stage depth map to act on the microbending operation to obtain a depth map of the next-smallest resolution of the reference image, and so on, until a depth map of the maximum resolution of the reference image is obtained.

11. The method according to claim 10, characterized in that The multiple local features of different resolutions are input into the feature matching and deformation module to obtain multiple global features of different resolutions; or Among the multiple local features of different resolutions, only the local feature with the minimum resolution is input into the feature matching deformation module to obtain the global feature with the minimum resolution. The global features of other resolutions are obtained based on the fusion of the global feature with the minimum resolution and the local features of other resolutions.

12. A depth estimation device, characterized in that: The depth estimation device includes a memory and a processor, wherein the memory stores a computer program to be executed by the processor, and when the computer program is executed by the processor, the processor executes the depth estimation method according to any one of claims 1 to 11.

13. A storage medium, characterized in that: The storage medium stores a computer program, which, when executed, executes the depth estimation method according to any one of claims 1 to 11.

14. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the depth estimation method according to any one of claims 1 to 11 is implemented.

Citation Information

Patent Citations

  • Model generation method and device based on multi-view panoramic image

    CN111402345A

  • Feature pyramid multi-view three-dimensional reconstruction method and system

    CN113345082A