Depth estimation method and apparatus, electronic device, and computer-readable storage medium
Patent Information
- Application Number
- CN202011440325.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-07
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2040-12-07
AI Technical Summary
但是,现有的深度估计方法所能估计的深度范围较小,无法满足智能设备对近距离或远距离目标的自动对焦需求
[0091]通过将待处理图像映射到预设平面获取待处理图像各像素点在预设成像平面上的位置信息,并在深度估计过程中使用待处理图像中各像素点在预设平面上的位置信息,消除了相机参数对深度估计范围的影响,使得同一个网络模型可以对不同相机参数对应的待处理图像进行深度估计,在保证宽范围深度估计的同时,节省了计算资源和存储空间。
Smart Images

Figure CN114596349B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and more specifically, to a depth estimation method, apparatus, electronic device, and computer-readable storage medium. Background Technology
[0002] Autofocus is a core function of many smart devices when capturing images and videos. Regardless of the distance between the object and the camera, users want the object of interest to be in focus, and depth estimation is the foundation for achieving fast autofocus. However, existing depth estimation methods have a limited depth range, which cannot meet the autofocus requirements of smart devices for both close-up and distant targets. Therefore, it is necessary to improve existing depth estimation methods. Summary of the Invention
[0003] The purpose of this application is to at least solve one of the aforementioned technical defects. The technical solution provided by the embodiments of this application is as follows:
[0004] In a first aspect, embodiments of this application provide a depth estimation method, including:
[0005] The image to be processed is mapped onto a preset plane, and the first position information of the pixels in the image to be processed on the preset plane is obtained;
[0006] Depth estimation is performed on the image to be processed based on the first location information.
[0007] In one optional embodiment of this application, obtaining the first position information of pixels in the image to be processed on a preset plane includes:
[0008] Based on the second camera parameters corresponding to the current imaging plane, obtain the second position information of the pixels in the image to be processed on the current imaging plane;
[0009] The first position information is obtained based on the second position information, the first camera parameters corresponding to the preset plane, and the second camera parameters.
[0010] In one alternative embodiment of this application, the camera parameters include at least one of the camera's focal length, principal point position, and sensor size.
[0011] In one optional embodiment of this application, obtaining the first position information based on the second position information, the first camera parameters corresponding to the preset plane, and the second camera parameters includes:
[0012] Based on the first camera parameters and the second camera parameters, obtain the mapping relationship between the first position information and the second position information;
[0013] Based on the second location information and the mapping relationship, the first location information is obtained.
[0014] In one optional embodiment of this application, depth estimation of the image to be processed based on the first location information includes:
[0015] The image to be processed is downsampled at least once using an encoding network to obtain the corresponding feature map.
[0016] The depth estimation result is obtained by performing at least one feature upsampling based on the feature map and the first location information through the first decoding network.
[0017] In one optional embodiment of this application, the encoding network includes a plurality of feature downsampling units, and at least one feature downsampling unit includes a first feature downsampling module and a second feature downsampling module, wherein,
[0018] The first feature downsampling module performs feature downsampling on the input image or feature map to obtain a first feature map. The second feature downsampling module performs feature downsampling on the input image or feature map to obtain a second feature map. The first feature map and the second feature map are then fused and output.
[0019] In one optional embodiment of this application, the convolution kernel used by the first feature downsampling module includes a first convolution kernel of a first dimension and a second convolution kernel of a second dimension, wherein the value of the second convolution kernel is zero.
[0020] The first dimension is determined based on the number of convolutional kernels used by the first feature downsampling module in the previous feature downsampling unit, and the second dimension is determined based on the number of convolutional kernels used by the second feature downsampling module in the previous feature downsampling unit.
[0021] In one optional embodiment of this application, the convolution kernels used by the first feature downsampling module and the second feature downsampling module are standard convolution kernels or pointwise convolution kernels.
[0022] In one optional embodiment of this application, the number of convolution kernels used by the first feature downsampling module is the same as the number of convolution kernels used by the first feature downsampling module in the previous feature downsampling unit, and the number of convolution kernels used by the second feature downsampling module is the same as the number of convolution kernels used by the second feature downsampling module in the previous feature downsampling unit.
[0023] In one optional embodiment of this application, the convolution kernels used by the first feature downsampling module and the second feature downsampling module are depthwise convolution kernels.
[0024] In one optional embodiment of this application, it further includes:
[0025] The second decoding network performs at least one feature upsampling based on the first feature map output by at least one first downsampling module in the encoding network to obtain the corresponding processing result.
[0026] In one optional embodiment of this application, the corresponding processing result obtained by the second decoding network is the semantic parsing result of the image to be processed.
[0027] In one optional embodiment of this application, at least one feature upsampling is performed based on the feature map and the first location information to obtain a depth estimation result, including:
[0028] The first position information and the feature map output from at least one feature downsampling are fused to obtain at least one first fused feature map;
[0029] Feature upsampling is performed based on at least one first fused feature map, corresponding to at least one feature downsampling.
[0030] In one optional embodiment of this application, feature upsampling corresponding to at least one feature downsampling is performed based on at least one first fused feature map, including:
[0031] The first fused feature map is fused with the input feature map corresponding to the corresponding feature upsampling to obtain the corresponding second fused feature map;
[0032] Feature upsampling is performed based on the second fused feature map to output the resulting feature map.
[0033] In one optional embodiment of this application, the method further includes:
[0034] Acquire at least two consecutive frames containing the image to be processed;
[0035] Based on at least two consecutive frames of images, obtain the first disparity information corresponding to the image to be processed;
[0036] Based on the feature map and the first location information, at least one feature upsampling is performed to obtain the depth estimation result, including:
[0037] Based on the feature map, the first location information, and the first disparity information, at least one feature upsampling is performed to obtain the depth estimation result.
[0038] In one optional embodiment of this application, obtaining first disparity information corresponding to the image to be processed based on at least two consecutive images includes:
[0039] Obtain the second disparity information between two adjacent frames in at least two consecutive images;
[0040] First disparity information is obtained based on second disparity information.
[0041] In one optional embodiment of this application, obtaining the first disparity information based on the second disparity information includes:
[0042] Based on the second disparity information, obtain the corresponding average disparity information or cumulative disparity information, and use the average disparity information or cumulative disparity information as the first disparity information.
[0043] In one optional embodiment of this application, at least one feature upsampling is performed based on the feature map, first location information, and first disparity information to obtain a depth estimation result, including:
[0044] The first position information, the first disparity information, and the feature map output from at least one feature downsampling are fused to obtain at least one third fused feature map.
[0045] Feature upsampling is performed based on at least one third fusion feature map, corresponding to at least one feature downsampling.
[0046] In one optional embodiment of this application, feature upsampling corresponding to at least one feature downsampling is performed based on at least one third fused feature map, including:
[0047] The third fused feature map is fused with the corresponding input feature map that is upsampled to obtain the corresponding fourth fused feature map.
[0048] Feature upsampling is performed based on the fourth fused feature map to output the resulting feature map.
[0049] Secondly, embodiments of this application provide an image processing method, including:
[0050] The encoding network performs feature downsampling on the image to be processed at least once to obtain the corresponding feature map. The encoding network includes several feature downsampling units, and at least one feature downsampling unit includes a first feature downsampling module and a second feature downsampling module. The first feature downsampling module and the second feature downsampling module perform feature downsampling on the input image to be processed or the feature map respectively to obtain a first feature map and a second feature map. The first feature map and the second feature map are then fused and output.
[0051] The first processing result is obtained by upsampling the features at least once based on the feature map output by the first decoding network and the encoding network.
[0052] The second processing result is obtained by upsampling the features at least once based on the first feature map output by at least one first downsampling module in the encoding network through the second decoding network.
[0053] In one optional embodiment of this application, the convolution kernel used by the first feature downsampling module includes a first convolution kernel of a first dimension and a second convolution kernel of a second dimension, wherein the value of the second convolution kernel is zero.
[0054] The first dimension is determined based on the number of convolutional kernels used by the first feature downsampling module in the previous feature downsampling unit, and the second dimension is determined based on the number of convolutional kernels used by the second feature downsampling module in the previous feature downsampling unit.
[0055] In one optional embodiment of this application, the number of convolution kernels used by the first feature downsampling module is the same as the number of convolution kernels used by the first feature downsampling module in the previous feature downsampling unit, and the number of convolution kernels used by the second feature downsampling module is the same as the number of convolution kernels used by the second feature downsampling module in the previous feature downsampling unit.
[0056] In one optional embodiment of this application, the first processing result is the depth estimation result of the image to be processed, and the second processing result is the semantic parsing result of the image to be processed.
[0057] Thirdly, embodiments of this application provide a depth estimation method, including:
[0058] Acquire at least two consecutive frames containing the image to be processed;
[0059] Based on at least two consecutive frames of images, obtain the first disparity information corresponding to the image to be processed;
[0060] Depth estimation is performed on the image to be processed based on the first disparity information.
[0061] In one optional embodiment of this application, obtaining first disparity information corresponding to the image to be processed based on at least two consecutive images includes:
[0062] Obtain the second disparity information between two adjacent frames in at least two consecutive images;
[0063] First disparity information is obtained based on second disparity information.
[0064] In one optional embodiment of this application, obtaining the first disparity information based on the second disparity information includes:
[0065] Based on the second disparity information, obtain the corresponding average disparity information or cumulative disparity information, and use the average disparity information or cumulative disparity information as the first disparity information.
[0066] In one optional embodiment of this application, depth estimation of the image to be processed based on first disparity information includes:
[0067] The image to be processed is downsampled at least once using an encoding network to obtain the corresponding feature map.
[0068] The depth estimation result is obtained by performing at least one feature upsampling based on the feature map and the first disparity information through the first decoding network.
[0069] In one optional embodiment of this application, feature upsampling is performed at least once based on the feature map and first disparity information, including:
[0070] The first disparity information and the feature map output from at least one feature downsampling are fused to obtain at least one fifth fused feature map.
[0071] Feature upsampling is performed based on at least one fifth fusion feature map, corresponding to at least one feature downsampling.
[0072] In one optional embodiment of this application, feature upsampling corresponding to at least one feature downsampling is performed based on at least one fifth fusion feature map, including:
[0073] The fifth fused feature map is fused with the input feature map corresponding to the feature upsampling to obtain the corresponding sixth fused feature map;
[0074] Feature upsampling is performed based on the sixth fusion feature map to output the resulting feature map.
[0075] Fourthly, embodiments of this application provide a depth estimation apparatus, including:
[0076] The position information acquisition module is used to map the image to be processed onto a preset plane and acquire the first position information of the pixels in the image to be processed on the preset plane;
[0077] The depth estimation module is used to perform depth estimation on the image to be processed based on the first location information.
[0078] Fifthly, embodiments of this application provide an image processing apparatus, including:
[0079] An encoding module is used to perform feature downsampling on the image to be processed at least once through an encoding network to obtain a corresponding feature map. The encoding network includes several feature downsampling units, and at least one feature downsampling unit includes a first feature downsampling module and a second feature downsampling module. The first feature downsampling module and the second feature downsampling module respectively perform feature downsampling on the input image to be processed or the feature map to obtain a first feature map and a second feature map. The first feature map and the second feature map are then fused and output.
[0080] The first decoding module is used to perform at least one feature upsampling based on the feature map output by the encoding network through the first decoding network to obtain the first processing result.
[0081] The second decoding module is used to perform at least one feature upsampling based on the first feature map output by at least one first downsampling module in the encoding network through the second decoding network to obtain the second processing result.
[0082] Sixthly, embodiments of this application provide a depth estimation apparatus, including:
[0083] A continuous image acquisition module is used to acquire at least two consecutive frames of images containing the image to be processed;
[0084] The disparity information acquisition module is used to acquire the first disparity information corresponding to the image to be processed based on at least two consecutive frames of images.
[0085] The depth estimation module is used to perform depth estimation on the image to be processed based on the first disparity information.
[0086] In a seventh aspect, embodiments of this application provide an electronic device, including a memory and a processor;
[0087] The memory contains computer programs;
[0088] A processor is configured to execute a computer program to implement the methods provided in the first aspect embodiment or any alternative embodiment of the first aspect, the second aspect embodiment or any alternative embodiment of the second aspect, or the third aspect embodiment or any alternative embodiment of the third aspect.
[0089] Eighthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the methods provided in the first aspect embodiment or any optional embodiment of the first aspect, the second aspect embodiment or any optional embodiment of the second aspect, or the third aspect embodiment or any optional embodiment of the third aspect.
[0090] The beneficial effects of the technical solution provided in this application are:
[0091] By mapping the image to be processed onto a preset plane, the position information of each pixel in the image to be processed on the preset imaging plane is obtained. The position information of each pixel in the image to be processed on the preset plane is used in the depth estimation process, eliminating the influence of camera parameters on the depth estimation range. This allows the same network model to perform depth estimation on images to be processed corresponding to different camera parameters, saving computational resources and storage space while ensuring a wide range of depth estimation. Attached Figure Description
[0092] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below.
[0093] Figure 1 A schematic diagram illustrating the autofocus process for smart devices;
[0094] Figure 2 A diagram showing the size comparison of an object on different imaging planes;
[0095] Figure 3 This is a schematic diagram of the network structure of a depth estimation network model in the prior art;
[0096] Figure 4 A flowchart illustrating a depth estimation method provided in an embodiment of this application;
[0097] Figure 5 This is a schematic diagram of coordinate transformation from the current imaging plane to the preset imaging plane in an embodiment of this application;
[0098] Figure 6 This is a schematic diagram illustrating IPM acquiring first location information in one example of an embodiment of this application;
[0099] Figure 7 This is a schematic diagram illustrating depth estimation using a preset network model combined with IPM in one example of an embodiment of this application.
[0100] Figure 8 This is a comparative diagram of several indicators of autofocus and auto exposure functions in existing technologies;
[0101] Figure 9a This is a schematic diagram illustrating human body parsing using a human body parsing network model, provided as an example of an embodiment of this application.
[0102] Figure 9b for Figure 9a A detailed schematic diagram of the network structure of the human body parsing network model shown in the figure;
[0103] Figure 10a This is a schematic diagram illustrating depth estimation based on a preset network model obtained by extending a human parsing network model, as provided in an example of an embodiment of this application.
[0104] Figure 10b for Figure 10a A detailed schematic diagram of the network structure of the preset network model shown;
[0105] Figure 11a This is a schematic diagram showing zero padding of the convolution kernel corresponding to the human body parsing and encoding module in one example of an embodiment of this application;
[0106] Figure 11b This is a schematic diagram showing zero padding of the convolutional kernel corresponding to the depth estimation coding module in one example of an embodiment of this application.
[0107] Figure 12 A schematic diagram of the angle offset and histogram of the angle offset introduced by the movement of the hand holding the shooting device in one example of an embodiment of this application;
[0108] Figure 13 This is an example of a parallax map between two adjacent frames in a series of k consecutive frames in this application embodiment;
[0109] Figure 14 This is a schematic diagram comparing depth images obtained from a single frame image and from k consecutive frames, as an example of an embodiment of this application.
[0110] Figure 15 This is a schematic diagram illustrating how MFbD acquires first disparity information in an example of an embodiment of this application;
[0111] Figure 16 This is a schematic diagram illustrating depth estimation using a preset network model combined with IPM and MFbD in one example of an embodiment of this application.
[0112] Figure 17 for Figure 16 A detailed schematic diagram of the combined network structure shown;
[0113] Figure 18 This is a schematic diagram illustrating the process of generating an image from a target camera based on an image from the NYUDv2 dataset, as an example of an embodiment of this application.
[0114] Figure 19 Histogram statistics for depth values in the NYUDv2 and KITTI datasets;
[0115] Figure 20 A schematic flowchart of an image processing method provided in an embodiment of this application;
[0116] Figure 21 This is a schematic diagram illustrating depth estimation using another preset network model derived from an extension of the human parsing network model, provided as an example of an embodiment of this application.
[0117] Figure 22 A flowchart illustrating another depth estimation method provided in an embodiment of this application;
[0118] Figure 23 This is a schematic diagram illustrating depth estimation based on a preset network model derived from an extension of the human parsing network model, as provided in an example of an embodiment of this application.
[0119] Figure 24A structural block diagram of a depth estimation device provided in an embodiment of this application;
[0120] Figure 25 A structural block diagram of an image processing apparatus provided in an embodiment of this application;
[0121] Figure 26 A structural block diagram of another depth estimation device provided in the embodiments of this application;
[0122] Figure 27 This is a schematic diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0123] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this application, and should not be construed as limiting the invention.
[0124] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this application means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we say an element is “connected” or “coupled” to another element, it can be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein can include wireless connections or wireless coupling. The term “and / or” as used herein includes all or any units and all combinations of one or more associated listed items.
[0125] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0126] Autofocus is a core function of many smart devices when shooting images and videos. Regardless of the distance between the object and the camera, users want the subject to be in focus. However, some smart devices (such as smartphones) still have autofocus issues, especially when shooting objects that are too close, sometimes failing to focus. Figure 1The diagram illustrates the autofocus process of a smart device (or camera) and the acquisition of the autofocused image. The left image is the input image before focusing; the middle image is the depth estimation result (depth image) corresponding to the input image; and the right image is the image after autofocus processing based on the estimated depth image. Many smart devices, such as mobile phones, require a wide depth estimation range for autofocus, often covering a range from 0.07m to 300m. This means that users need autofocus for both near and far targets when taking photos with smart devices, requiring the device to have a wide-range depth estimation capability.
[0127] The depth value of each pixel in an input image refers to the distance between the object corresponding to that pixel and the camera's optical center (origin) in the camera coordinate system, from the Z-axis point of the object. Existing depth estimation schemes work by estimating the depth based on the camera's imaging model and the size of the object region in the image on the imaging plane; that is, estimating the Z-axis distance between the object being photographed and the imaging device. For example... Figure 2 As shown, for the same real object at a distance d from the imaging device (i.e., a depth value of d), different camera focal lengths f1 and f2 result in different image sizes on the imaging plane. The size of the object on the imaging plane depends on the distance between the object and the imaging device, as well as the parameters of the imaging device (such as camera focal length, principal point position (also known as principal point coordinates), sensor size, etc.).
[0128] Most existing depth estimation schemes use an encoder-decoder network model based on the above depth estimation principle to obtain the depth image corresponding to the input image to complete the depth estimation, such as... Figure 3As shown, this depth estimation network model consists of two parts: an encoder and a decoder. It is typically trained and tested using public datasets such as NYUDv2 and KITTI, where image samples from the same dataset correspond to images captured by the same camera with identical camera parameters. Since the same depth value corresponds to different sizes on the imaging plane depending on the camera parameters (i.e., the relationship between object depth values and object sizes on the imaging plane differs depending on the camera parameters), it is necessary to use image samples from the same dataset for both training and testing. Clearly, the depth estimation network model trained using the above method can only be used to estimate the depth of images with the same camera parameters as the training dataset; that is, the trained depth estimation network model can only cover the depth range that the camera parameters corresponding to that specific dataset can cover. For example, NYUDv2 is an indoor scene dataset, with a depth range of 0.5m to 10m when captured. Therefore, a depth estimation network model trained on NYUDv2 can only cover a depth range of 0.5m to 10m. KITTI is an outdoor scene dataset, with a depth range of 1m to 100m when captured. Thus, a depth estimation network model trained on KITTI can only cover a depth range of 1m to 100m. Each dataset can only cover a portion of the "wide range" of depth. Therefore, models trained on different training sets can only estimate a partial depth range.
[0129] Considering the principles and network structures of existing depth estimation schemes, achieving wide-range depth estimation requires storing multiple depth estimation network models trained on different datasets in the smart device, which consumes significant computational resources and storage space. To address this issue, this application provides a depth estimation method, which will be further described below.
[0130] Figure 4 A flowchart illustrating a depth estimation method provided in an embodiment of this application is shown below. Figure 4 As shown, the method may include:
[0131] Step S401: Map the image to be processed onto a preset plane and obtain the first position information of the pixels in the image to be processed on the preset plane;
[0132] Step 402: Perform depth estimation on the image to be processed based on the first location information.
[0133] The first location information can be the coordinates of each pixel.
[0134] Specifically, for multiple images to be processed corresponding to different camera parameters, they are each mapped onto the same preset plane (also called a preset imaging plane). First position information corresponding to each image is obtained. Then, based on this first position information, the size of the object in each image on the preset imaging plane can be obtained, and the object's depth value can be estimated based on this size. Since different images to be processed are mapped onto the same preset imaging plane, the calculation scale of the object size in each image is the same, ensuring a consistent correspondence between the object size and the object's depth value. In other words, mapping the images to be processed onto the preset imaging plane and then using the object's size on that plane for depth estimation eliminates the influence of camera parameters on the depth estimation range of the same depth estimation network model. This means that the same depth estimation network model can be used to estimate the depth of images to be processed corresponding to different camera parameters.
[0135] The solution provided in this application obtains the position information of each pixel in the image to be processed on the preset imaging plane by mapping the image to be processed to the preset imaging plane, and uses the position information of each pixel in the image to be processed on the preset imaging plane during the depth estimation process. This eliminates the influence of camera parameters on the depth estimation range, so that the same network model can perform depth estimation for images to be processed corresponding to different camera parameters. While ensuring a wide range of depth estimation, it saves computing resources and storage space.
[0136] In one optional embodiment of this application, obtaining the first position information of pixels in the image to be processed on a preset imaging plane includes:
[0137] Obtain the second position information of the pixels in the image to be processed on the current imaging plane, and obtain the second camera parameters corresponding to the current imaging plane;
[0138] The first position information is obtained based on the second position information, the first camera parameters corresponding to the preset imaging plane, and the second camera parameters.
[0139] Further, based on the second position information, the first camera parameters corresponding to the preset imaging plane, and the second camera parameters, the first position information is obtained, including:
[0140] Based on the first camera parameters and the second camera parameters, obtain the mapping relationship between the first position information and the second position information;
[0141] Based on the second location information and the mapping relationship, the first location information is obtained.
[0142] Specifically, the process of mapping the image to be processed from the current imaging plane to the preset imaging plane is the process of transforming the coordinates of each pixel in the image to be processed, that is, converting the second position information of each pixel in the current imaging plane into the first position information in the preset imaging plane.
[0143] Specifically, such as Figure 5 As shown, the first camera parameters corresponding to the preset imaging plane U include: focal length. Sensor size The second camera parameters corresponding to the current imaging plane P include: focal length f(f x f y ), sensor size (d x d y ) and principal point coordinates (I cx I cy The coordinate transformation process described above may include:
[0144] First, determine the coordinates (i.e., the second position information) of each pixel in the image to be processed in the current imaging plane P, using the following formula:
[0145] P x (i, j) = I cx -i
[0146] P y (i, j) = I cy -j
[0147] Among them, P x (i, j) represents the x-axis coordinates of pixel (i, j) in the current imaging plane P, where P y (i, j) is the y-coordinate of pixel (i, j) in the current imaging plane P.
[0148] Then, based on the principle of similar triangles, the coordinates of each pixel in the image to be processed on the current imaging plane are converted into coordinates on a preset imaging plane (i.e., the first position information). That is, the mapping relationship between the current imaging plane and the preset imaging plane is obtained based on the principle of similar triangles, and the first position information is obtained based on this mapping relationship and the second position information. The formula is as follows:
[0149]
[0150]
[0151] Among them, U x (i, j) represents the x-axis coordinates of pixel (i, j) in the preset imaging plane U. y (i, j) is the y-axis coordinate of pixel (i, j) in the preset imaging plane U.
[0152] Understandably, an Image Mapping Module (IPM) can be pre-defined based on the above processing procedure to perform coordinate transformation on each pixel in the image to be processed. The input of the IPM is the image to be processed and the corresponding camera parameters (i.e., the second camera parameters), and the output is the first position information corresponding to the image to be processed. Figure 6 This is a schematic diagram of the input and output of the IPM module. The input is the x-axis camera parameters and y-axis camera parameters corresponding to the second camera parameters. After passing through the IPM, it outputs two coordinate matrices of size W*H (which can be denoted as W*H*2). Each element in these two coordinate matrices is the x-axis coordinate and y-axis coordinate of the corresponding pixel in the preset imaging plane.
[0153] In one optional embodiment of this application, depth estimation of the image to be processed based on the first location information includes:
[0154] The depth of the image to be processed is estimated based on the first location information by using a pre-set network model.
[0155] Specifically, the preset network model is used to perform depth estimation on the image to be processed by combining the first location information and output the corresponding depth estimation result. More specifically, the preset network model can be combined with the IPM (Integrated Peripheral Model). During the depth estimation process of the image to be processed, the preset network model uses the first location information output by the IPM to output the corresponding depth estimation result.
[0156] In one optional embodiment of this application, depth estimation of the image to be processed based on the first location information includes:
[0157] The image to be processed is downsampled at least once using an encoding network to obtain the corresponding feature map.
[0158] The depth estimation result is obtained by performing at least one feature upsampling based on the feature map and the first location information through the first decoding network.
[0159] The resulting depth can be the corresponding depth image.
[0160] Specifically, in the encoding network, the image to be processed undergoes multiple feature downsampling operations. Each feature downsampling generates a corresponding feature map, which can be understood as the feature map output by the encoding network. On the other hand, in the first decoding network, combining the feature map output by the encoding network and the first position information, multiple feature upsampling operations are performed on the feature map output by the encoding network (i.e., the output feature map of the encoding network) to obtain the depth image corresponding to the image to be processed. Each feature upsampling operation corresponds to a feature downsampling operation in the encoding network. For example... Figure 7As shown, the preset network model is combined with IPM, and the first position information output by IPM and the feature downsampled in the coding unit (i.e., the coding network) correspond to the features. Figure 1 It is used in the feature upsampling corresponding to the decoding unit (i.e., the first decoding network).
[0161] In one optional embodiment of this application, at least one feature upsampling is performed based on the feature map and the first location information to obtain a depth estimation result, including:
[0162] The first position information and the feature map output from at least one feature downsampling are fused to obtain at least one first fused feature map;
[0163] Feature upsampling is performed based on at least one first fused feature map, corresponding to at least one feature downsampling.
[0164] Further, feature upsampling corresponding to at least one feature downsampling is performed based on at least one first fused feature map, including:
[0165] The first fused feature map is fused with the input feature map corresponding to the corresponding feature upsampling to obtain the corresponding second fused feature map;
[0166] Feature upsampling is performed based on the second fused feature map to output the resulting feature map.
[0167] Specifically, firstly, the first position information and the feature map corresponding to at least one feature downsampling are fused to obtain at least one first fused feature map; then, each first fused feature map is used for feature upsampling corresponding to the feature downsampling. More specifically, each first fused feature map is fused with the input feature map of the corresponding feature upsampling to obtain a second fused feature map, and then feature upsampling is performed on the second fused feature map.
[0168] It should be noted that during the process of obtaining the first fusion feature map, the feature map that is fused with the first position information can be selected according to the actual needs. For example, the feature map output from a certain feature downsampling can be selected, or the feature map output from several feature downsamplings can be selected. In the subsequent process of obtaining the second fusion feature map, the input feature map of the corresponding feature upsampling needs to be selected and the corresponding first fusion feature map is fused together. That is, the first fusion feature map needs to be applied to the corresponding feature upsampling.
[0169] For many shooting devices, automatic exposure and automatic focus are often required to be present simultaneously in preview mode. Current technology necessitates different network models to implement these two functions, such as... Figure 8As shown, for the automatic exposure function, the network model needs to be able to perform human body analysis and exposure parameter settings, which takes approximately 25 milliseconds; for the autofocus function, the model needs to be able to perform depth estimation, which takes approximately 7 to 10 milliseconds. If the network models for both functions are run separately, it will consume a lot of computing resources and take too long, affecting real-time performance.
[0170] Considering that the features extracted for the human body parsing task and the features required for depth estimation are very similar in preview mode, the encoding network mentioned above can be shared to some extent by the human body parsing task and the depth estimation task.
[0171] It is understood that, besides the human body parsing task, if the features extracted by other tasks can be utilized by the depth estimation task, then these other tasks can also share the encoding network with the depth estimation task. For example, other tasks could also be the more widely applicable semantic parsing task. This application embodiment only uses the human body parsing task as an example to describe the solution in detail, but it is not limited thereto. Similarly, it is understood that the solution provided by this application embodiment is applicable not only in preview mode but also in non-preview mode.
[0172] To ensure depth estimation while maintaining human body parsing performance, the encoding network includes several feature downsampling units. At least one feature downsampling unit includes a first feature downsampling module (such as a semantic parsing encoding module, or more specifically, a human body parsing encoding module) and a second feature downsampling module (i.e., a depth estimation encoding module). The first feature downsampling module downsamples the input image or feature map to obtain a first feature map, and the second feature downsampling module downsamples the input image or feature map to obtain a second feature map. The first and second feature maps are then fused and output.
[0173] In one optional embodiment of this application, if the convolution kernels used by the first feature downsampling module and the second feature downsampling module are standard convolution or pointwise convolution kernels, then the convolution kernels used by the first feature downsampling module include a first convolution kernel with a first dimension and a second convolution kernel with a second dimension, wherein the value of the second convolution kernel is zero. The first dimension is determined based on the number of convolution kernels used by the first feature downsampling module in the previous feature downsampling unit, and the second dimension is determined based on the number of convolution kernels used by the second feature downsampling module in the previous feature downsampling unit.
[0174] Specifically, the kernels of standard convolution or pointwise convolution are generally three-dimensional convolution kernels. Each three-dimensional convolution kernel can be seen as a superposition of multiple two-dimensional convolution kernels. Therefore, this three-dimensional convolution kernel can be divided into two parts: a first convolution kernel in the first dimension and a second convolution kernel in the second dimension. For example, a three-dimensional convolution kernel with dimensions a*b*c can be divided into a first convolution kernel of a1*b*c (first dimension) and a second convolution kernel of a2*b*c (second dimension), where a = a1 + a2. a1 and a2 can be referred to as the heights of the first and second convolution kernels, respectively. These heights can be determined based on the number of convolution kernels in the previous feature downsampling unit. Specifically, the height of the first convolution kernel is equal to the number of convolution kernels in the first feature downsampling module of the previous feature downsampling unit, and the height of the second convolution kernel is equal to the number of convolution kernels in the second feature downsampling module of the previous feature downsampling unit. Furthermore, the second convolution kernel is zero (i.e., all weights in the second convolution kernel are zero). By setting the convolution kernels as described above, it can be ensured that the first feature downsampling module extracts only the feature information from the feature map output by the first feature downsampling module in the previous feature downsampling unit. This means the human body analysis module extracts only the human body analysis feature information from the feature map output by the human body analysis module in the previous feature downsampling unit. Furthermore, each two-dimensional convolution kernel superimposed in the three-dimensional convolution kernel can be referred to as a slice. In other words, each three-dimensional convolution kernel is composed of multiple slices. The number of slices in the first convolution kernel is equal to the number of convolution kernels in the first feature downsampling module in the previous feature downsampling unit, and the number of slices in the second convolution kernel is equal to the number of convolution kernels in the second feature downsampling module in the previous feature downsampling unit.
[0175] It is understandable that the height of the convolution kernel in the second feature downsampling module is the same as the height of the convolution kernel in the first feature downsampling module (the convolution kernel dimensions are the same).
[0176] In one optional embodiment of this application, if the convolution kernels used by the first feature downsampling module and the second feature downsampling module are depthwise convolution kernels, then the number of convolution kernels used by the first feature downsampling module is the same as the number of convolution kernels used by the first feature downsampling module in the previous feature downsampling unit, and the number of convolution kernels used by the second feature downsampling module is the same as the number of convolution kernels used by the second feature downsampling module in the previous feature downsampling unit.
[0177] Specifically, the convolution kernel for depthwise convolution is generally a two-dimensional convolution kernel. Each convolution kernel performs a convolution operation on the feature map of one channel output by the previous feature downsampling unit. Therefore, the number of channels in the previous feature downsampling unit is equal to the number of channels in the depthwise convolution kernel. In other words, the number of convolution kernels used by the first feature downsampling module is the same as the number of convolution kernels used by the first feature downsampling module in the previous feature downsampling unit, and the number of convolution kernels used by the second feature downsampling module is the same as the number of convolution kernels used by the second feature downsampling module in the previous feature downsampling unit.
[0178] In one optional embodiment of this application, the method may further include:
[0179] The second decoding network performs at least one feature upsampling based on the first feature map output by at least one first downsampling module in the encoding network to obtain the corresponding processing result.
[0180] The first decoding network performs at least one feature downsampling based on the second feature map output by at least one second downsampling module in the encoding network to obtain the corresponding processing result.
[0181] The processing result obtained by the second decoding network can be the semantic parsing result of the image to be processed. The processing result obtained by the first decoding network can be the depth estimation result of the image to be processed.
[0182] Specifically, semantic parsing is similar to semantic segmentation; their basic task can be considered as assigning a category to each pixel. Semantic parsing typically involves more detailed categories than semantic segmentation. For example, semantic segmentation might categorize pixels into human body, blue sky, and grass, while semantic parsing might categorize them into eyebrows, nose, and mouth.
[0183] Next, we will continue to use the first feature downsampling module as the human body parsing encoding module as an example to explain the above scheme in detail. The preset network model of this application can be understood as an extension of the existing human body parsing network model, such as... Figure 9a The diagram shows the network structure of an existing human body parsing network model, which includes a human body parsing encoding module and a human body parsing decoding module.
[0184] Specifically, such as Figure 9bAs shown, the human body parsing encoding module of the human body parsing network model consists of multiple convolutional layers, including standard convolutions and depth-wise convolutions. The input image passes through multiple convolutional layers to obtain shallow, mid-level, and deep feature maps. For example, after passing through C 1*k*k convolutional kernels, the input image yields C shallow feature maps; then, it passes through D standard convolutional kernels to obtain D mid-level feature maps; and finally, it passes through D depth-wise convolutional kernels (also called channel-wise convolutional kernels) to obtain D deep feature maps. The next step is the human body parsing decoding module (i.e., the second decoding network), which also consists of multiple convolutional layers. This module upsamples the features of the output image from the human body parsing encoding module and finally outputs the human body parsing result corresponding to the image to be processed.
[0185] Specifically, a depth estimation coding module (i.e., a second feature downsampling module) is added to the human body parsing coding module (i.e., the first feature downsampling module) to obtain the coding network in the preset network model of this application embodiment. On the other hand, a depth estimation decoding unit (i.e., the first decoding network) corresponding to the depth estimation coding unit is added. Then, the depth estimation decoding unit and the human body parsing decoding unit (i.e., the second decoding network) constitute the decoding network in the preset network model of this application embodiment. The network structure diagram of the preset network model is shown below. Figure 10a As shown.
[0186] Specifically, such as Figure 10bAs shown, the method for expanding the human body parsing encoding module is as follows: In the human body parsing network model, when processing the image to be processed to the shallow feature map, the human body parsing task requires C 1*k*k convolutional kernels, and then an additional C' 1*k*k convolutional kernels are added for the depth estimation task. Therefore, due to the (C+C') convolutional kernels, a total of (C+C') channels of shallow feature maps are obtained (where the human body parsing task corresponds to C channels of shallow feature maps, and the depth estimation task corresponds to C' channels of shallow feature maps). Then, after passing through the additional D' (C+C')*k*k standard convolutional kernels, a total of (D+D') channels of mid-level feature maps are obtained (where the human body parsing task corresponds to D channels of mid-level feature maps, and the depth estimation task corresponds to D' channels of mid-level feature maps). Then, after passing through (D... (D+D') 1*k*k channel-wise convolutional kernels are used to obtain (D+D') deep feature maps (where D channels correspond to the human body parsing task and D' channels correspond to the depth estimation task). These are then processed by (M+M') (D+D')*1*1 point-wise convolutional kernels to obtain (M+M') feature maps (where M channels correspond to the human body parsing task and M' channels correspond to the depth estimation task). The subsequent steps are the human body parsing decoding module and the depth estimation decoding module. The human body parsing decoding module upsamples the output feature map corresponding to the human body parsing encoding module, and the depth estimation decoding module upsamples the output feature map corresponding to the depth estimation encoding module. The final output is the human body parsing result and the depth estimation result (i.e., the depth image) corresponding to the image to be processed. As can be seen from the above feature downsampling process, in the standard convolution operation and the pointwise convolution operation, the human body parsing encoding module only extracts the features of the part of the feature map corresponding to the human body parsing task in the input feature map, while the depth estimation encoding module extracts the features of the part of the feature map corresponding to the human body parsing task in the input feature map and the features of the part of the feature map corresponding to the depth estimation task in the input feature map.
[0187] It should be noted that, for the encoding network provided in this application embodiment, when utilizing the features extracted in the human body parsing task, only the above-described standard convolution operation method can be used, or only the above-described depthwise convolution operation and pointwise convolution operation combined can be used, or both methods can be used simultaneously. Similarly, the above-described "shallow feature map", "middle feature map" and "deep feature map" are only illustrative examples and are not limited thereto. That is, the specific feature maps output by which convolutional layers are "shallow feature maps", "middle feature maps" or "deep feature maps" can be defined according to actual needs.
[0188] Furthermore, in the process of expanding the human body parsing encoding module to obtain the depth estimation encoding module, firstly, the newly added convolutional kernel corresponding to the depth estimation is determined; then, since human body parsing does not need to use the features extracted by depth estimation, and to avoid using computationally expensive merging operations, zero-padding is applied to the convolutional kernel corresponding to the human body parsing encoding module, and the dimension of the zero-padding is equal to the dimension of the newly added convolutional kernel; finally, since depth estimation can use the features extracted by human body parsing, the newly added convolutional kernel and the original convolutional kernel corresponding to the human body parsing module are superimposed to obtain the convolutional kernel corresponding to the depth estimation encoding module. (See also...) Figure 9b and Figure 10b Taking the process of shallow feature maps undergoing D standard convolutions to obtain mid-level feature maps as an example, the human body parsing network model uses D convolutional kernels with a dimension of C*k*k. However, in the preset network model of this application embodiment, D convolutional kernels with a dimension of (C+C')*k*k are used in the human body parsing encoding module. Figure 11a As shown, its first C-dimensional convolutional kernel (the part with height C) is the same as the convolutional kernel in the human body parsing network model, while the newly added C'-dimensional convolutional kernel (the part with height C') has a value of 0. This avoids the computationally expensive merging operation without affecting human body parsing. The convolutional kernel dimension used in the depth estimation encoding module of the preset network model is also (C+C')*k*k.
[0189] It should be noted that the above is a method for expanding convolutional kernels (i.e., expanding the encoding module), and different expansion methods can be used in different network branches. For example, for depthwise convolutional kernels from the human body parsing encoding module, if they are used in depthwise convolutional layers, they do not need to be padded with zeros. For standard convolutional kernels from the depth estimation encoding module, if the current layer does not need to utilize the features extracted from human body parsing, then the convolutional kernel of the current layer needs to be padded with zeros, such as... Figure 11b As shown, the convolution kernels of the dimension corresponding to C are all 0, and the convolution kernels of the dimension corresponding to C' are depth estimation convolution kernels.
[0190] In the above-mentioned scheme of sharing the encoding network for human body parsing and depth estimation tasks, the performance of depth estimation is improved on the one hand, and the size of the preset network model is controlled while improving performance on the other hand. The specific analysis is as follows:
[0191] (1) Performance: Human body analysis and depth estimation tasks share many similarities in their features. Reusing human body analysis features in depth estimation can significantly improve the performance. Compared to networks used in standalone depth estimation tasks, the encoding part of the depth estimation network in this scheme can obtain more semantic features, not only in terms of quantity but also in terms of richer meaning. More semantic features can improve the performance of depth estimation. For example, if the category of each pixel in an image is known, such as a part being the human body, then the depth values of those pixels should be similar.
[0192] (2) Model size: Separate human body parsing and separate depth estimation require two models, which occupy a relatively large amount of storage space. However, for the multi-task network of this application, while ensuring the performance of depth estimation, the number of convolutional kernels required for depth estimation is much smaller than that of a separate depth estimation network, so our multi-task network occupies less storage space.
[0193] (3) Real-time performance: By reusing the human body parsing coding network, the processing efficiency of depth estimation can be greatly improved. If the embodiments of this application are applied to the mobile phone photo preview mode, then the real-time performance requirement in the preview mode is guaranteed.
[0194] In preview mode, hand movements while holding the shooting device introduce random angular shifts. For example... Figure 12 As shown, Figure (a) illustrates the horizontal and vertical offsets caused by hand movements, with the offset distribution being almost symmetrical; Figure (b) shows the histogram statistics of angular offsets, with most offsets being relatively small. In preview mode, the hand movement response on the shooting device manifests as: objects closer to the image will have a larger positional offset compared to objects farther away. These offsets reflect the differences in depth values between objects. For example, objects closer to the shooting device's depth value have a larger parallax value, while objects farther away from the shooting device's depth value have a smaller parallax value in the same image. Figure 13 As shown in the figure, the upper part of the figure is a continuous k-frame image, and the lower part is the disparity map between every two adjacent frames and the depth map of the k-th frame image.
[0195] In one optional embodiment of this application, the method may further include:
[0196] Obtain consecutive images containing at least two adjacent frames of the image to be processed;
[0197] Based on at least two consecutive frames of images, obtain the first disparity information corresponding to the image to be processed;
[0198] Then, based on the first location information, depth estimation is performed on the image to be processed to obtain the depth image corresponding to the image to be processed, including:
[0199] Depth estimation is performed on the image to be processed based on the first disparity information and the first position information to obtain the depth image corresponding to the image to be processed.
[0200] Furthermore, based on at least two consecutive frames of images, the first disparity information corresponding to the image to be processed is obtained, including:
[0201] Obtain the second disparity information between two adjacent frames in at least two consecutive images;
[0202] First disparity information is obtained based on second disparity information.
[0203] Further, obtaining the first disparity information based on the second disparity information includes:
[0204] Based on the second disparity information, obtain the corresponding average disparity information or cumulative disparity information, and use the average disparity information or cumulative disparity information as the first disparity information.
[0205] Specifically, refer again Figure 13 For k consecutive frames, a disparity map (i.e., second disparity information) can be calculated for every two adjacent frames. The depth map is then calculated based on (k-1) consecutively accumulated or averaged absolute disparity maps (i.e., first disparity information). For example, the disparity value between frame 2 and frame 1 is |i2-i1| in the horizontal direction and |j2-j1| in the vertical direction; the disparity value between frame 3 and frame 2 is |i3-i2| in the horizontal direction and |j3-j2| in the vertical direction; similarly, the disparity value between frame k and frame (k-1) is |i2-i1| in the horizontal direction. k -i k-1 |, with the vertical direction being |j k -j k-1 |
[0206] Specifically, depth estimation using only a single image is prone to being unrobust, such as... Figure 14 As shown, when estimating the depth map of the k-th frame image alone, the depth values of distant and near objects are mixed together, resulting in inaccurate depth estimation. However, after fusing disparity information from multiple frames, the depth values of distant and near objects can be better distinguished. This is because the disparity value of distant objects is smaller than that of near objects. This additional information helps the pre-defined network model to better estimate the depth values, improving the robustness of depth estimation.
[0207] Understandably, a multi-frame based disparity (MFbD) module can be pre-defined based on the above processing procedure to obtain the first disparity information of the image to be processed. The input of the MFbD is at least two consecutive frames of images containing the image to be processed (i.e., k consecutive frames), and the output is the first disparity information corresponding to the image to be processed. Figure 15 This is a schematic diagram of the input and output of MFbD. The input is a k-frame image in preview mode. After passing through the MFbD convolution kernel, it outputs two disparity matrices of size W*H (which can be denoted as W*H*2). Each element in these two disparity matrices represents the horizontal and vertical disparity information of the corresponding pixel.
[0208] In one optional embodiment of this application, at least one feature upsampling is performed based on the feature map, first location information, and first disparity information to obtain a depth estimation result, including:
[0209] The first position information, the first disparity information, and the feature map output from at least one feature downsampling are fused to obtain at least one third fused feature map.
[0210] Feature upsampling is performed based on at least one third fusion feature map, corresponding to at least one feature downsampling.
[0211] In one optional embodiment of this application, feature upsampling corresponding to at least one feature downsampling is performed based on at least one third fused feature map, including:
[0212] The third fused feature map is fused with the corresponding input feature map that is upsampled to obtain the corresponding fourth fused feature map.
[0213] Feature upsampling is performed based on the fourth fused feature map to output the resulting feature map.
[0214] It should be noted that, as Figure 16-17 As shown, MFbD can be used alone or in combination with the preset network model described above, just like IPM, to perform depth estimation on the image to be processed. In the combined structure, the connection method between MFbD and the preset network is the same as that between IPM and the preset network. The first disparity information output by MFbD is used in the same way as the first position information in the depth estimation of the preset network model, and will not be described again here.
[0215] In summary, among the user feedback issues, the autofocus function of some smart devices is not good enough, especially for close-up shooting. The technical solution provided in this application can provide a wide range of depth map estimation (covering micro-depth) to improve the subsequent autofocus function of the mobile phone. Since the preset network model in the solution provided in this application can be adapted to different cameras, it can be used for the main camera and telephoto lens in the rear camera system, as well as for different cameras in the front camera series. To save computing resources, the solution provided in this application can use a single model to complete the functions of automatic exposure and automatic focus. Applying the technical solution of the embodiments of this application, regardless of the distance of the object, accurate depth estimation can achieve accurate focus, thereby obtaining a clear image of the object.
[0216] This application proposes a deep learning-based multi-task wide-range depth estimation method. Wide-range depth estimation has broad application prospects in fields such as autofocus, AI (Artificial Intelligence) cameras, AR (Augmented Reality) / VR (Virtual Reality), and can be used in various smartphones, robots, and AR / VR devices. This application achieves wide-range depth support by incorporating image plane coordinate transformation into convolutional layers, achieves more robust near-far distance depth estimation by fusing multi-frame image information, and integrates depth estimation and other tasks into the same network in a multi-task manner, achieving efficient wide-range depth estimation that meets the real-time requirements of preview mode.
[0217] The training process of the preset network model used in the above process is described below. This part includes the acquisition of training data, the design of loss function and the design of rating criteria, which are described in detail below.
[0218] 1. Obtaining training data:
[0219] (1) Based on images captured by other cameras (such as images from datasets like NYUDV2, KITTI, and DIODE), the training data for the target camera (i.e., the camera corresponding to the preset imaging plane, whose camera parameters are the first camera parameters mentioned above) is calculated using the following formula:
[0220]
[0221] Among them, [x I y I [1] represents the coordinates of each pixel in the input image, and K represents the coordinates of that pixel. -1 It is the inverse matrix of the camera intrinsic parameter matrix used to capture the current image (including camera focal length, image principal point coordinates, sensor size, i.e., the second camera parameter mentioned earlier), Ks This is the intrinsic parameter matrix of the target camera. z is the depth value of each pixel in the captured image, and t... z R is the set translation depth, and R is the set rotation matrix.
[0222] (2) Change t Z To obtain the required micro-depth data, modify R to expand the training data.
[0223] (3) There will be many holes in the image generated in step (2). Use bilinear interpolation to fill the holes.
[0224] (4) Cropping the image obtained in the above steps, removing low-resolution and poor-quality areas, thus obtaining the training data.
[0225] like Figure 18 As shown, the left image is an image captured in the NYUDV2 dataset, and the right image is the target image (i.e., training data) obtained by transforming the original image.
[0226] 2. Design of the loss function:
[0227] Acceptable error cost: More penalty is applied to the more important depth range; the specific loss function is shown in the following formula:
[0228]
[0229]
[0230]
[0231] Where D(i) is the cost function for the i-th depth value interval, which can be L1, L2, or other functions that measure the error between the true depth value and the estimated depth value. The value representing the acceptable error, d i This represents the interval containing the true values of the i-th depth value. This indicates that in the t-th batch of data, d i The average acceptable error for this interval, where α and β represent hyperparameters, and L represents the moving average error across all intervals. For example... Figure 19 The image shows the histogram statistics of the depth values in the NYUDv2 and KITTI datasets.
[0232] As shown in Table 1, the acceptable error values are for different depth ranges. For example, for Level 1, the acceptable error range for the corresponding depth range [0.07m, 0.08m) is 0.005m, and the acceptable error range for the depth range [1.2m, 1.6m) is 0.2m.
[0233] Table 1
[0234] 1 [0.07,0.08) 0.005 … … … 21 [1.2,1.6) 0.2 … … … 24 [5,10) 2.5 … … …
[0235] 3. Evaluation Criteria Design:
[0236] If the estimated depth value Difference from the true depth value If the value is less than AE, then it is considered to meet the condition. The number of pixels that meet the condition is denoted as N. R The total number of pixels in the test set is denoted as N. all Then, the acceptable error for the defined evaluation criterion is expressed as:
[0237] Figure 20 This application also provides a flowchart illustrating an image processing method, as shown in the embodiments below. Figure 20 As shown, the method may include:
[0238] Step S2001: The image to be processed is downsampled at least once through the encoding network to obtain the corresponding feature map. The encoding network includes several feature downsampling units. At least one feature downsampling unit includes a first feature downsampling module and a second feature downsampling module. The first feature downsampling module and the second feature downsampling module respectively downsample the input image to be processed or the feature map to obtain a first feature map and a second feature map. The first feature map and the second feature map are fused and then output.
[0239] Step S2002: Through the first decoding network, at least one feature upsampling is performed based on the feature map output by the encoding network to obtain the first processing result;
[0240] Step S2003: Through the second decoding network, at least one feature upsampling is performed on the first feature map output by at least one first downsampling module in the encoding network to obtain the second processing result.
[0241] The solution provided in this application performs multiple feature downsampling operations on the image to be processed in the encoding module through the first feature downsampling module and the second feature downsampling module, respectively. Then, the corresponding decoding network decodes the feature maps output by the two feature downsampling modules to obtain the corresponding processing results. While obtaining two processing results in one model, the second downsampling module in the encoding module can reuse the features extracted by the first downsampling module, thereby making the processing results obtained after decoding more accurate.
[0242] In one optional embodiment of this application, the convolution kernel used by the first feature downsampling module includes a first convolution kernel of a first dimension and a second convolution kernel of a second dimension, wherein the value of the second convolution kernel is zero.
[0243] The first dimension is determined based on the number of convolutional kernels used by the first feature downsampling module in the previous feature downsampling unit, and the second dimension is determined based on the number of convolutional kernels used by the second feature downsampling module in the previous feature downsampling unit.
[0244] In one optional embodiment of this application, the number of convolution kernels used by the first feature downsampling module is the same as the number of convolution kernels used by the first feature downsampling module in the previous feature downsampling unit, and the number of convolution kernels used by the second feature downsampling module is the same as the number of convolution kernels used by the second feature downsampling module in the previous feature downsampling unit.
[0245] In one optional embodiment of this application, the first processing result is the depth estimation result of the image to be processed, and the second processing result is the semantic parsing result of the image to be processed.
[0246] It should be noted that in this embodiment, the depth estimation task and other related tasks share the encoding network, and the depth estimation feature downsampling module in the encoding network can reuse the features extracted by the feature downsampling modules of related tasks. The specific implementation process of the above scheme is described above and will not be repeated here. Figure 21 As shown, the depth estimation task and the human body parsing task share the same encoding network, which includes a human body parsing encoding module and a depth estimation module.
[0247] Figure 22 A flowchart illustrating a depth estimation method provided in an embodiment of this application is shown below. Figure 22 As shown, the method may include:
[0248] Step S2201: Obtain at least two consecutive frames of images containing the image to be processed;
[0249] Step S2202: Based on at least two consecutive frames of images, obtain the first disparity information corresponding to the image to be processed;
[0250] Step S2203: Depth estimation is performed on the image to be processed based on the first disparity information.
[0251] The solution provided in this application obtains the disparity information of the image to be processed by using at least two consecutive images containing the image to be processed, and performs depth estimation by combining the disparity information of the image to be processed. This eliminates the influence of disparity introduced during the acquisition of the image to be processed on the depth estimation, and improves the accuracy and robustness of the depth estimation.
[0252] In one optional embodiment of this application, obtaining first disparity information corresponding to the image to be processed based on at least two consecutive images includes:
[0253] Obtain the second disparity information between two adjacent frames in at least two consecutive images;
[0254] First disparity information is obtained based on second disparity information.
[0255] In one optional embodiment of this application, obtaining the first disparity information based on the second disparity information includes:
[0256] Based on the second disparity information, obtain the corresponding average disparity information or cumulative disparity information, and use the average disparity information or cumulative disparity information as the first disparity information.
[0257] In one optional embodiment of this application, depth estimation of the image to be processed based on first disparity information includes:
[0258] The image to be processed is downsampled at least once using an encoding network to obtain the corresponding feature map.
[0259] The depth estimation result is obtained by performing at least one feature upsampling based on the feature map and the first disparity information through the first decoding network.
[0260] In one optional embodiment of this application, feature upsampling is performed at least once based on the feature map and first disparity information, including:
[0261] The first disparity information and the feature map output from at least one feature downsampling are fused to obtain at least one fifth fused feature map.
[0262] Feature upsampling is performed based on at least one fifth fusion feature map, corresponding to at least one feature downsampling.
[0263] In one optional embodiment of this application, feature upsampling corresponding to at least one feature downsampling is performed based on at least one fifth fusion feature map, including:
[0264] The fifth fused feature map is fused with the input feature map corresponding to the feature upsampling to obtain the corresponding sixth fused feature map;
[0265] Feature upsampling is performed based on the sixth fusion feature map to output the resulting feature map.
[0266] It should be noted that, as Figure 23 As shown, MFbD is combined with the previously mentioned preset network model to perform depth estimation on the image to be processed. In the combined structure, the first disparity information output by MFbD is used in the same way as the first position information in the depth estimation of the preset network model, and will not be repeated here.
[0267] Figure 24A structural block diagram of a depth estimation device provided in an embodiment of this application is shown below. Figure 24 As shown, the device 2400 may include: a location information acquisition module 2401 and a depth image acquisition module 2402, wherein:
[0268] The position information acquisition module 2401 is used to map the image to be processed onto a preset plane and acquire the first position information of the pixels in the image to be processed on the preset plane;
[0269] The depth image acquisition module 2402 is used to perform depth estimation on the image to be processed based on the first location information.
[0270] The solution provided in this application obtains the position information of each pixel in the image to be processed on the preset imaging plane by mapping the image to be processed to the preset imaging plane, and uses the position information of each pixel in the image to be processed on the preset imaging plane during the depth estimation process. This eliminates the influence of camera parameters on the depth estimation range, so that the same network model can perform depth estimation for images to be processed corresponding to different camera parameters. While ensuring a wide range of depth estimation, it saves computing resources and storage space.
[0271] In one optional embodiment of this application, the first location information acquisition module is specifically used for:
[0272] Based on the second camera parameters corresponding to the current imaging plane, obtain the second position information of the pixels in the image to be processed on the current imaging plane;
[0273] The first position information is obtained based on the second position information, the first camera parameters corresponding to the preset plane, and the second camera parameters.
[0274] In one alternative embodiment of this application, the camera parameters include at least one of the camera's focal length, principal point position, and sensor size.
[0275] In an optional embodiment of this application, the first location information acquisition module is further configured to:
[0276] Based on the first camera parameters and the second camera parameters, obtain the mapping relationship between the first position information and the second position information;
[0277] Based on the second location information and the mapping relationship, the first location information is obtained.
[0278] In one optional embodiment of this application, the depth image acquisition module includes a feature downsampling submodule and a feature upsampling submodule, wherein:
[0279] The feature downsampling submodule is used to perform feature downsampling on the image to be processed at least once through the encoding network to obtain the corresponding feature map;
[0280] The first feature upsampling submodule is used to perform at least one feature upsampling based on the feature map and the first location information through the first decoding network to obtain a depth estimation result.
[0281] In one optional embodiment of this application, the coding network includes a plurality of feature downsampling units, and at least one feature downsampling unit includes a first feature downsampling module and a second feature downsampling module, wherein,
[0282] The first feature downsampling module performs feature downsampling on the input image or feature map to obtain a first feature map, and the second feature downsampling module performs feature downsampling on the input image or feature map to obtain a second feature map. The first feature map and the second feature map are then fused and output.
[0283] In one optional embodiment of this application, the convolution kernel used by the first feature downsampling module includes a first convolution kernel of a first dimension and a second convolution kernel of a second dimension, wherein the value of the second convolution kernel is zero.
[0284] The first dimension is determined based on the number of convolutional kernels used by the first feature downsampling module in the previous feature downsampling unit, and the second dimension is determined based on the number of convolutional kernels used by the second feature downsampling module in the previous feature downsampling unit.
[0285] In one optional embodiment of this application, the convolution kernels used by the first feature downsampling module and the second feature downsampling module are standard convolution kernels or pointwise convolution kernels.
[0286] In one optional embodiment of this application, the number of convolution kernels used by the first feature downsampling module is the same as the number of convolution kernels used by the first feature downsampling module in the previous feature downsampling unit, and the number of convolution kernels used by the second feature downsampling module is the same as the number of convolution kernels used by the second feature downsampling module in the previous feature downsampling unit.
[0287] In one optional embodiment of this application, the convolution kernels used by the first feature downsampling module and the second feature downsampling module are depthwise convolution kernels.
[0288] In an optional embodiment of this application, the device may further include a second feature upsampling submodule, used for:
[0289] The second decoding network performs at least one feature upsampling based on the first feature map output by at least one first downsampling module in the encoding network to obtain the corresponding processing result.
[0290] In one optional embodiment of this application, the corresponding processing result obtained by the second decoding network is the semantic parsing result of the image to be processed.
[0291] In one optional embodiment of this application, the first feature upsampling submodule is specifically used for:
[0292] The first location information and the feature map output from at least one feature downsampling are fused to obtain at least one first fused feature map;
[0293] Based on the at least one first fused feature map, feature upsampling is performed corresponding to the at least one feature downsampling.
[0294] In an optional embodiment of this application, the first feature upsampling submodule is further configured to:
[0295] The first fused feature map is fused with the input feature map corresponding to the corresponding feature upsampling to obtain the corresponding second fused feature map;
[0296] Based on the second fused feature map, feature upsampling is performed to output the resulting feature map.
[0297] In an optional embodiment of this application, the device may further include a parallax information acquisition module, used for:
[0298] Acquire at least two consecutive frames containing the image to be processed;
[0299] Based on at least two consecutive images, obtain the first disparity information corresponding to the image to be processed;
[0300] Accordingly, the depth estimation module is further used for:
[0301] Based on the feature map, the first position information, and the first disparity information, the output feature map of the encoding network is upsampled to correspond to at least one feature downsampling to obtain the depth estimation result.
[0302] In one optional embodiment of this application, the disparity information acquisition module is specifically used for:
[0303] Obtain the second disparity information between two adjacent frames in at least two consecutive images;
[0304] First disparity information is obtained based on second disparity information.
[0305] In an optional embodiment of this application, the parallax information acquisition module is further configured to:
[0306] Based on the second disparity information, obtain the corresponding average disparity information or cumulative disparity information, and use the average disparity information or cumulative disparity information as the first disparity information.
[0307] In an optional embodiment of this application, the first feature upsampling submodule is further configured to:
[0308] The first location information, the first disparity information, and the feature map output from at least one feature downsampling are fused to obtain at least one third fused feature map.
[0309] Based on the at least one third fusion feature map, feature upsampling is performed corresponding to the at least one feature downsampling.
[0310] In an optional embodiment of this application, the first feature upsampling submodule is further configured to:
[0311] The third fused feature map is fused with the corresponding input feature map corresponding to the feature upsampling to obtain the corresponding fourth fused feature map;
[0312] Based on the fourth fused feature map, feature upsampling is performed to output the resulting feature map.
[0313] Figure 25 This is a structural block diagram of an image processing apparatus provided in an embodiment of this application, such as... Figure 25 As shown, the device 2500 may include: an encoding module 2501, a first decoding module 2502, and a second decoding module 2503, wherein:
[0314] The encoding module 2501 is used to perform feature downsampling on the image to be processed at least once through the encoding network to obtain the corresponding feature map. The encoding network includes a plurality of feature downsampling units, and at least one feature downsampling unit includes a first feature downsampling module and a second feature downsampling module. The first feature downsampling module and the second feature downsampling module respectively perform feature downsampling on the input image to be processed or the feature map to obtain a first feature map and a second feature map. The first feature map and the second feature map are then fused and output.
[0315] The first decoding module 2502 is used to perform at least one feature upsampling based on the feature map output by the encoding network through the first decoding network to obtain a first processing result;
[0316] The second decoding module 2503 is used to perform at least one feature upsampling based on the first feature map output by at least one first downsampling module in the encoding network through the second decoding network to obtain a second processing result. The solution provided in this application obtains the position information of each pixel in the image to be processed on the preset imaging plane by mapping the image to be processed onto the preset imaging plane, and uses the position information of each pixel in the image to be processed on the preset imaging plane during depth estimation. This eliminates the influence of camera parameters on the depth estimation range, allowing the same network model to perform depth estimation for images to be processed corresponding to different camera parameters. While ensuring a wide range of depth estimation, it saves computational resources and storage space.
[0317] In one optional embodiment of this application, the convolution kernel used by the first feature downsampling module includes a first convolution kernel of a first dimension and a second convolution kernel of a second dimension, wherein the value of the second convolution kernel is zero.
[0318] The first dimension is determined based on the number of convolutional kernels used by the first feature downsampling module in the previous feature downsampling unit, and the second dimension is determined based on the number of convolutional kernels used by the second feature downsampling module in the previous feature downsampling unit.
[0319] In one optional embodiment of this application, the number of convolution kernels used by the first feature downsampling module is the same as the number of convolution kernels used by the first feature downsampling module in the previous feature downsampling unit, and the number of convolution kernels used by the second feature downsampling module is the same as the number of convolution kernels used by the second feature downsampling module in the previous feature downsampling unit.
[0320] In one optional embodiment of this application, the first processing result is the depth estimation result of the image to be processed, and the second processing result is the semantic parsing result of the image to be processed.
[0321] Figure 26 A structural block diagram of a depth estimation device provided in an embodiment of this application is shown below. Figure 26 As shown, the device 2600 may include: a continuous image acquisition module 2601, a disparity information acquisition module 2602, and a depth estimation module 2603, wherein:
[0322] The continuous image acquisition module 2601 is used to acquire at least two consecutive frames of images containing the image to be processed;
[0323] The disparity information acquisition module 2602 is used to acquire first disparity information corresponding to the image to be processed based on at least two consecutive frames of images;
[0324] The depth estimation module 2603 is used to perform depth estimation on the image to be processed based on the first disparity information.
[0325] In one optional embodiment of this application, the first disparity information acquisition module is specifically used for:
[0326] Obtain the second disparity information between two adjacent frames in at least two consecutive images;
[0327] First disparity information is obtained based on second disparity information.
[0328] In an optional embodiment of this application, the first disparity information acquisition module is further configured to:
[0329] Based on the second disparity information, obtain the corresponding average disparity information or cumulative disparity information, and use the average disparity information or cumulative disparity information as the first disparity information.
[0330] In one optional embodiment of this application, the depth estimation module includes: a feature downsampling submodule and a feature upsampling submodule, wherein:
[0331] The feature downsampling submodule is used to perform feature downsampling on the image to be processed at least once through the encoding network to obtain the corresponding feature map;
[0332] The feature upsampling submodule is used to perform at least one feature upsampling based on the feature map and the first disparity information through the first decoding network to obtain the depth estimation result.
[0333] In one optional embodiment of this application, the feature upsampling submodule is specifically used for:
[0334] The first disparity information and the feature map output from at least one feature downsampling are fused to obtain at least one fifth fused feature map.
[0335] Based on the at least one fifth fusion feature map, feature upsampling is performed corresponding to the at least one feature downsampling.
[0336] In an optional embodiment of this application, the feature upsampling submodule is further configured to:
[0337] The fifth fused feature map is fused with the corresponding input feature map corresponding to the feature upsampling to obtain the corresponding sixth fused feature map;
[0338] Based on the sixth fused feature map, feature upsampling is performed to output the resulting feature map.
[0339] The following is for reference. Figure 27 It illustrates an electronic device suitable for implementing embodiments of this application (e.g., performing...). Figure 4 , Figure 20 or Figure 22 The diagram illustrates the structure of the terminal device or server 1800 of the method shown. The electronic devices in the embodiments of this application may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), wearable devices, etc., as well as fixed terminals such as digital TVs, desktop computers, etc. Figure 27 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0340] The electronic device includes a memory and a processor. The memory stores a program for performing the methods described in the above-described method embodiments; the processor is configured to execute the program stored in the memory. The processor may be referred to as processing device 1801 as described below, and the memory may include at least one of read-only memory (ROM) 1802, random access memory (RAM) 1803, and storage device 1808 as described below, as follows:
[0341] like Figure 27 As shown, electronic device 1800 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 1801, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 1802 or a program loaded from storage device 1808 into random access memory (RAM) 1803. RAM 1803 also stores various programs and data required for the operation of electronic device 1800. Processing device 1801, ROM 1802, and RAM 1803 are interconnected via bus 1804. Input / output (I / O) interface 1805 is also connected to bus 1804.
[0342] Typically, the following devices can be connected to the I / O interface 1805: input devices 1806 including, for example, a touchscreen, touchpad, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc.; output devices 1807 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1808 including, for example, magnetic tape, hard disk, etc.; and communication devices 1809. Communication device 1809 allows electronic device 1800 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 27 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively.
[0343] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 1809, or installed from storage device 1808, or installed from ROM 1802. When the computer program is executed by processing device 1801, it performs the functions defined in the methods of embodiments of this application.
[0344] It should be noted that the computer-readable storage medium described above in this application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0345] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0346] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0347] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to:
[0348] The image to be processed is mapped onto a preset plane, and the first position information of the pixels in the image to be processed on the preset plane is obtained; depth estimation of the image to be processed is performed based on the first position information.
[0349] Alternatively, the image to be processed is downsampled at least once using an encoding network to obtain a corresponding feature map. The encoding network includes several feature downsampling units, and at least one feature downsampling unit includes a first feature downsampling module and a second feature downsampling module. The first feature downsampling module and the second feature downsampling module respectively downsample the input image to be processed or the feature map to obtain a first feature map and a second feature map. The first feature map and the second feature map are then fused and output. The first decoding network performs feature upsampling at least once based on the feature map output by the encoding network to obtain a first processing result. The second decoding network performs feature upsampling at least once based on the first feature map output by at least one first downsampling module in the encoding network to obtain a second processing result.
[0350] Alternatively, acquire at least two consecutive frames containing the image to be processed; acquire first disparity information corresponding to the image to be processed based on the at least two consecutive frames; and perform depth estimation on the image to be processed based on the first disparity information.
[0351] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof. These programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0352] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0353] The modules or units described in the embodiments of this application can be implemented in software or hardware. The names of modules or units do not necessarily limit the specific unit; for example, a first location information acquisition module can also be described as a "module for acquiring first location information".
[0354] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0355] In the context of this application, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0356] The apparatus provided in this application embodiment can implement at least one of multiple modules through an AI model. AI-related functions can be executed through non-volatile memory, volatile memory, and a processor.
[0357] The processor may include one or more processors. In this case, the one or more processors may be general-purpose processors, such as central processing unit (CPU), application processor (AP), etc., or pure graphics processing unit, such as graphics processing unit (GPU), vision processing unit (VPU), and / or AI-specific processors, such as neural processing unit (NPU).
[0358] The one or more processors control the processing of input data based on predefined operating rules or artificial intelligence (AI) models stored in non-volatile and volatile memory. These predefined operating rules or AI models are provided through training or learning.
[0359] Here, "providing through learning" refers to obtaining predefined operating rules or an AI model with desired characteristics by applying a learning algorithm to multiple learning datasets. This learning can be performed within the device itself, in which the AI is executed according to the embodiment, and / or can be implemented via a separate server / system.
[0360] This AI model can contain multiple neural network layers. Each layer has multiple weight values, and the computation of a layer is performed using the results of the previous layer and the multiple weights of the current layer. Examples of neural networks include, but are not limited to, Convolutional Neural Networks (CNNs), Deep Neural Networks (DNNs), Recurrent Neural Networks (RNNs), Restricted Boltzmann Machines (RBMs), Deep Belief Networks (DBNs), Bidirectional Recurrent Deep Neural Networks (BRDNNs), Generative Adversarial Networks (GANs), and Deep Q-Networks.
[0361] A learning algorithm is a method of training a predetermined target device (e.g., a robot) using multiple learning data sets to enable, allow, or control the target device to make determinations or predictions. Examples of such learning algorithms include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.
[0362] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific methods implemented by the computer-readable medium described above when executed by an electronic device can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
Claims
1. A depth estimation method, characterized in that, include: The image to be processed is mapped onto a preset plane, and the first position information of the pixels in the image to be processed on the preset plane is obtained; The image to be processed is downsampled at least once using an encoding network to obtain the corresponding feature map; The depth estimation result is obtained by performing at least one feature upsampling based on the feature map and the first location information through the first decoding network. The step of obtaining the first position information of the pixels in the image to be processed on the preset plane includes: Based on the second camera parameters corresponding to the current imaging plane, the second position information of the pixels in the image to be processed on the current imaging plane is obtained; The first position information is obtained based on the second position information, the first camera parameters corresponding to the preset plane, and the second camera parameters.
2. The method according to claim 1, characterized in that, Camera parameters include at least one of the following: camera focal length, principal point position, and sensor size.
3. The method according to claim 1 or 2, characterized in that, The step of obtaining the first position information based on the second position information, the first camera parameters corresponding to the preset plane, and the second camera parameters includes: Based on the first camera parameters and the second camera parameters, obtain the mapping relationship between the first location information and the second location information; Based on the second location information and the mapping relationship, the first location information is obtained.
4. The method according to claim 1, characterized in that, The encoding network includes several feature downsampling units, and at least one feature downsampling unit includes a first feature downsampling module and a second feature downsampling module, wherein... The first feature downsampling module performs feature downsampling on the input image or feature map to obtain a first feature map, and the second feature downsampling module performs feature downsampling on the input image or feature map to obtain a second feature map. The first feature map and the second feature map are then fused and output.
5. The method according to claim 4, characterized in that, The convolution kernel used by the first feature downsampling module includes a first convolution kernel in a first dimension and a second convolution kernel in a second dimension, wherein the value of the second convolution kernel is zero. The first dimension is determined based on the number of convolutional kernels used by the first feature downsampling module in the previous feature downsampling unit, and the second dimension is determined based on the number of convolutional kernels used by the second feature downsampling module in the previous feature downsampling unit.
6. The method according to claim 4, characterized in that, The number of convolutional kernels used by the first feature downsampling module is the same as the number of convolutional kernels used by the first feature downsampling module in the previous feature downsampling unit, and the number of convolutional kernels used by the second feature downsampling module is the same as the number of convolutional kernels used by the second feature downsampling module in the previous feature downsampling unit.
7. The method according to any one of claims 4-6, characterized in that, Also includes: The second decoding network performs at least one feature upsampling based on the first feature map output by at least one first downsampling module in the encoding network to obtain the corresponding processing result.
8. The method according to claim 7, characterized in that, The corresponding processing result obtained by the second decoding network is the semantic parsing result of the image to be processed.
9. The method according to claim 1, characterized in that, The step of performing at least one feature upsampling based on the feature map and the first location information to obtain a depth estimation result includes: The first location information and the feature map output from at least one feature downsampling are fused to obtain at least one first fused feature map; Based on the at least one first fused feature map, feature upsampling is performed corresponding to the at least one feature downsampling.
10. The method according to claim 1, characterized in that, The feature upsampling based on the at least one first fused feature map, corresponding to the at least one feature downsampling, includes: The first fused feature map is fused with the input feature map corresponding to the corresponding feature upsampling to obtain the corresponding second fused feature map; Based on the second fused feature map, feature upsampling is performed to output the resulting feature map.
11. The method according to claim 1, characterized in that, The method further includes: Acquire at least two consecutive frames containing the image to be processed; Based on the at least two consecutive images, obtain the first disparity information corresponding to the image to be processed; The step of performing at least one feature upsampling based on the feature map and the first location information to obtain a depth estimation result includes: Based on the feature map, the first location information, and the first disparity information, at least one feature upsampling is performed to obtain the depth estimation result.
12. The method according to claim 11, characterized in that, obtaining the first disparity information corresponding to the image to be processed based on the at least two consecutive images includes: Obtain the second disparity information between two adjacent frames in the at least two consecutive images; The first disparity information is obtained based on the second disparity information.
13. The method according to claim 12, characterized in that, The step of obtaining the first disparity information based on the second disparity information includes: Based on the second disparity information, obtain the corresponding average disparity information or cumulative disparity information, and use the average disparity information or the cumulative disparity information as the first disparity information.
14. The method according to claim 11, characterized in that, The process of performing at least one feature upsampling based on the feature map, the first location information, and the first disparity information to obtain a depth estimation result includes: The first location information, the first disparity information, and the feature map output from at least one feature downsampling are fused to obtain at least one third fused feature map. Based on the at least one third fusion feature map, feature upsampling is performed corresponding to the at least one feature downsampling.
15. The method according to claim 14, characterized in that, The feature upsampling based on the at least one third fused feature map corresponding to the at least one feature downsampling includes: The third fused feature map is fused with the corresponding input feature map corresponding to the feature upsampling to obtain the corresponding fourth fused feature map; Based on the fourth fused feature map, feature upsampling is performed to output the resulting feature map.
16. An electronic device, characterized in that, Including memory and processor; The memory stores computer programs; The processor is configured to execute the computer program to implement the method of any one of claims 1 to 15.
17. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method of any one of claims 1 to 15.
Citation Information
Patent Citations
METHOD AND DEVICE FOR DETECTING PLANE AREA USING DEPTH IMAGE, and Non-Transitory COMPUTER READABLE RECORDING MEDIUM
KR102074929B1