Depth estimation method, storage medium and computer device
By regional fusion of depth images with high dynamic range and low dynamic range, the problem that the prior art cannot meet the high-scale environmental depth information and high-precision face depth information at the same time is solved, and higher depth estimation accuracy is achieved.
Patent Information
- Application Number
- CN202110732659.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-06-29
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2041-06-29
AI Technical Summary
The existing depth estimation methods cannot meet the needs of high-scale environmental depth information and high-precision face depth information at the same time, especially in the face heavy lighting technology.
The target depth image is generated by acquiring a first depth image of a high dynamic range and a second depth image of a low dynamic range, and image fusing the two according to the region where the target of interest is located.
Adaptively obtaining the target depth image where the area where the target of interest is located is in a high dynamic range and the area where the target of interest is located is in a low dynamic range, thereby improving the accuracy of depth estimation.
Smart Images

Figure CN113343973B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing, and in particular, to a depth estimation method, a storage medium and a computer device. Background Art
[0002] There are two existing depth estimation methods. One is to perform depth estimation on the received original image through a depth estimation model to obtain a depth map with a high dynamic range (the scale range of the depth map is large, that is, the difference between the maximum depth value and the minimum depth value is large, such as tens of meters); the other is to perform depth estimation on the received original depth map through a three-dimensional reconstruction model to obtain a high-precision depth map (the scale range of the depth map is small, that is, the difference between the maximum depth value and the minimum depth value is small, such as a few centimeters).
[0003] In some specific application scenarios, such as face re-illumination technology, high-scale environmental depth information and high-precision face depth information are required at the same time, and the existing depth estimation methods cannot meet these requirements at the same time. Therefore, how to adaptively obtain the depth information of the image to be recognized is a technical problem that needs to be solved urgently. Summary of the invention
[0004] The present application provides a depth estimation method, a storage medium, and a computer device, which can solve the technical problem of how to adaptively obtain the depth information of an image to be identified.
[0005] In a first aspect, an embodiment of the present application provides a depth estimation method, the method comprising:
[0006] Acquire a first depth image of the target image based on the depth estimation model, and determine a first area where the target of interest is located in the first depth image;
[0007] Acquire a second depth image of the target image based on the three-dimensional reconstruction model, and determine a second area where the target of interest is located in the second depth image;
[0008] Fusing the depth image of the first area and the depth image of the second area to obtain a third depth image of the object of interest;
[0009] A target depth image is generated based on the third depth image and the first depth image.
[0010] In a second aspect, an embodiment of the present application provides a storage medium, wherein the storage medium stores a computer program, and the computer program is suitable for being loaded by a processor and executing the steps of the above method.
[0011] In a third aspect, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the steps of the above method are implemented when the processor executes the program.
[0012] In an embodiment of the present application, by acquiring a first depth image with a high dynamic range and a second depth image with a low dynamic range, and then performing image fusion on the first depth image and the second depth image according to the area where the target of interest is located, a target depth image in which the area where the target of interest is located is adaptively acquired as a high dynamic range and the area where the target of non-interest is located is a low dynamic range, thereby obtaining a target depth image with higher accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without paying any creative work.
[0014] Figure 1 A flowchart of a depth estimation method provided in an embodiment of the present application;
[0015] Figure 2 A flowchart of a depth estimation method provided in an embodiment of the present application;
[0016] Figure 3 A schematic diagram of a process for obtaining a normalization factor provided in an embodiment of the present application;
[0017] Figure 4 A flowchart of a depth estimation method provided in an embodiment of the present application;
[0018] Figure 5 A flowchart of a depth estimation method provided in an embodiment of the present application;
[0019] Figure 6 A schematic diagram of the structure of a depth estimation device provided in an embodiment of the present application;
[0020] Figure 7 It is a structural diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0021] In order to make the features and advantages of the present application more obvious and easy to understand, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present application.
[0022] When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are only examples of devices and methods consistent with some aspects of the present application as detailed in the attached claims. The flowcharts shown in the accompanying drawings are only exemplary and do not have to be executed according to the steps shown. For example, some steps are parallel and there is no strict logical order, so the actual execution order is variable. In addition, the terms "first", "second", "third", "fourth", "fifth", "sixth", "seventh", and "eighth" are only for the purpose of distinction and should not be used as limitations of the present disclosure.
[0023] The depth estimation method and depth estimation device disclosed in the embodiments of the present application can be applied to the field of image processing, such as face recognition, 3D modeling, etc. The depth estimation device can include but is not limited to smart terminals such as smart interactive tablets, mobile phones, personal computers, and laptops.
[0024] In an embodiment of the present application, the depth estimation device can obtain a first depth image with a high dynamic range and a second depth image with a low dynamic range, and then fuse the first depth image with the second depth image according to the area where the target of interest is located, so as to adaptively obtain a target depth image in which the area where the target of interest is located is a high dynamic range and the area where the non-target of interest is located is a low dynamic range, thereby obtaining a target depth image with higher accuracy.
[0025] The following will be combined Figure 1~Figure 5 , the depth estimation method provided in the embodiments of the present application is introduced in detail.
[0026] See also Figure 1 , which is a flow chart of a depth estimation method provided in an embodiment of the present application. Figure 1 As shown, the method may include the following steps S101 to S104.
[0027] S101, acquiring a first depth image of a target image based on a depth estimation model, and determining a first area where an object of interest is located in the first depth image.
[0028] Specifically, the pixel value of each pixel in the first depth image represents the depth information of the pixel at the original image position corresponding to the pixel in the real world. The depth information can be understood as the distance between the real object corresponding to the pixel and the camera or reference plane. It should be noted that the depth estimation model uses a macro algorithm, that is, a high dynamic range depth estimation algorithm, and the depth image output by the depth estimation model is also a high dynamic range depth image. The target of interest can be a small object such as a face, a figurine, a doll, a desktop plant, etc., which requires high precision.
[0029] After receiving the target image for depth estimation, the processor inputs the received target image into the depth estimation model so that the depth estimation model performs depth estimation on the received target image, generates a first depth image of the target image, and outputs the obtained first depth image.
[0030] When the processor receives the first depth image output by the depth estimation model, it determines a first area in the first depth image where the object of interest is located.
[0031] S102, acquiring a second depth image of the target image based on the three-dimensional reconstruction model, and determining a second area where the target of interest is located in the second depth image.
[0032] Specifically, the pixel value of each pixel in the second depth image also represents the depth information of the pixel at the original image position corresponding to the pixel in the real world, and the depth information can also be understood as the distance between the real object corresponding to the pixel and the camera or reference plane. It should be noted that the 3D reconstruction model uses a microscopic algorithm, that is, a low dynamic range depth estimation algorithm, and the depth image output by the 3D reconstruction model is also a low dynamic range depth image.
[0033] After receiving the target image for depth estimation, the processor also inputs the received target image into the 3D reconstruction model so that the 3D reconstruction model performs depth estimation on the received target image, generates a second depth image of the target image, and outputs the obtained second depth image.
[0034] When the processor receives the second depth image output by the three-dimensional reconstruction model, it determines a second area in the second depth image where the object of interest is located.
[0035] S103: Fusing the depth image of the first area and the depth image of the second area to obtain a third depth image of the object of interest.
[0036] Specifically, a sub-depth image corresponding to the first area in the first depth image is obtained, and a sub-depth image corresponding to the second area in the second depth image is obtained, and then the depth image corresponding to the first area is fused with the depth image corresponding to the second area to obtain a third depth image of the target of interest. It can be understood that the third depth image only includes the area where the target of interest is located.
[0037] S104: Generate a target depth image based on the third depth image and the first depth image.
[0038] Specifically, the third depth image is merged with the first depth image, that is, firstly, images to be merged corresponding to other areas of the first depth image except the first area where the target of interest is located are obtained, and then the third depth image and the image to be merged are merged according to their corresponding areas to obtain the target depth image.
[0039] Optionally, when generating the third depth image, the first area corresponding to the target of interest is gradient-guided according to the second area corresponding to the target of interest, and the depth image corresponding to the first area in the first depth image is directly replaced by the third depth image without affecting the integrity of the first depth image.
[0040] Optionally, if there is no object of interest in the target image, the first depth image is directly used as the target depth image.
[0041] In an embodiment of the present application, by acquiring a first depth image with a high dynamic range and a second depth image with a low dynamic range, and then performing image fusion on the first depth image and the second depth image according to the area where the target of interest is located, a target depth image in which the area where the target of interest is located is adaptively acquired as a high dynamic range and the area where the target of non-interest is located is a low dynamic range, thereby obtaining a target depth image with higher accuracy.
[0042] See also Figure 2 , which is a flow chart of a depth estimation method provided in an embodiment of the present application. Figure 2 As shown, the method may include the following steps S201 to S210.
[0043] S201, obtaining a target image mask where an object of interest is located in a target image.
[0044] Specifically, the processor first inputs the target image into the target detection model, and then identifies whether there is an object of interest in the target image based on the target detection model. When the target detection model detects the object of interest, the region information of the object of interest in the target image is obtained, that is, the target region where the object of interest is located in the target image is determined, and then the target image mask is generated based on.
[0045] The target image mask is an image representing the area where the target of interest is located in the target image. The image size of the target image mask is consistent with the image size of the target image. The pixel value of each pixel in the target image mask only includes a first preset value and a second preset value, wherein the first preset value represents the area where the target of interest is located, and the second preset value represents the area where the target of non-interest is located.
[0046] S202: Acquire a first depth image of the target image based on a depth estimation model.
[0047] Specifically, after receiving the target image for depth estimation, the processor inputs the received target image into the depth estimation model so that the depth estimation model performs depth estimation on the received target image, generates a first depth image of the target image, and outputs the obtained first depth image.
[0048] S203: If the image size of the first depth image is inconsistent with the image size of the target image, perform image sampling on the first depth image to obtain a sampled image that is consistent with the image size of the target image.
[0049] Specifically, the image size of the first depth image and the image size of the target image are obtained, and the image size of the first depth image and the image size of the target image are compared. If the image size of the first depth image is inconsistent with the image size of the target image, the first depth image is sampled to obtain a sampled image that is consistent with the image size of the target image.
[0050] Optionally, if the image size of the first depth image is smaller than the image size of the target image, the first depth image is upsampled so as to enlarge the first depth image to obtain a sampled image having the same image size as the target image; if the image size of the first depth image is larger than the image size of the target image, the first depth image is downsampled so as to reduce the first depth image to obtain a sampled image having the same image size as the target image.
[0051] S204: Replace the first depth image with the sampling image.
[0052] Since the first depth image only represents the depth information corresponding to each pixel, it is impossible to directly determine whether there is an object of interest in the first depth image based on the depth information, and it is even more impossible to determine the area where the object of interest is located in the first depth image. Therefore, by sampling the first depth image, the first depth image is stretched / contracted to ensure that the image size of the first depth image is consistent with the image size of the target image, so that the first area where the object of interest is located in the first depth image can be directly determined based on the area where the object of interest is located in the target image, avoiding the situation where the first area where the object of interest is located cannot be obtained through the first depth image.
[0053] S205 : Determine a first region where the object of interest is located in the first depth image based on the object image mask.
[0054] Specifically, determine the first positions of each pixel point whose pixel value is the first preset value in the target image mask, and then determine the second positions corresponding to each first position in the first depth image. According to each second position, obtain the first area where the target of interest is located. Specifically, the pixel point corresponding to each second position in the first depth image can be used as the first pixel point corresponding to the first area.
[0055] Exemplarily, the first area where the object of interest is located in the first depth image may be determined by the following formula:
[0056]
[0057] in, is a first region where the object of interest is located in the first depth image; is the pixel at the i-th row and j-th column in the first depth image; is the pixel at the i-th row and j-th column in the target image mask, and 1 is the first preset value; is the first depth image.
[0058] S206: Acquire a second depth image of the target image based on the three-dimensional reconstruction model, and determine a second region where the target of interest is located in the second depth image based on the target image mask.
[0059] Specifically, determine the first positions of each pixel point whose pixel value is a first preset value in the target image mask, and then determine the third positions corresponding to each first position in the second depth image. According to each third position, obtain the second area where the target of interest is located. Specifically, the pixel point corresponding to each third position in the first depth image can be used as the second pixel point corresponding to the second area.
[0060] Exemplarily, the second area where the object of interest is located in the second depth image may be determined by the following formula:
[0061]
[0062] in, is a second region where the object of interest is located in the second depth image; is the pixel at the i-th row and j-th column in the second depth image; is the pixel at the i-th row and j-th column in the target image mask, and 1 is the first preset value; is the second depth image.
[0063] Since the depth image only represents the depth information corresponding to each pixel, it is impossible to directly determine whether there is an object of interest in the depth image based on the depth information, and it is even more impossible to determine the area where the object of interest is located in the depth image. Therefore, by first obtaining the target area where the object of interest is located in the target image, and then generating a target image mask based on the target area, the area where the object of interest is located in the first depth image and the second depth image is obtained according to the target image mask.
[0064] S207 , obtaining first depth information corresponding to each pixel in the first area and second depth information corresponding to each pixel in the second area.
[0065] Specifically, the depth information corresponding to a pixel point refers to a pixel value corresponding to the pixel point.
[0066] S208: Acquire a normalization factor based on first depth information corresponding to each pixel in the first area and second depth information corresponding to each pixel in the second area.
[0067] Specifically, the normalization factor is used to normalize the first depth information and the second depth information to the same scale range. It should be noted that since the first depth image and the second depth image are depth images obtained by performing depth estimation on the same target image based on different models, and the depth information obtained by different models is different relative depths, that is, the scales of the depth information obtained by different models are different, it can be understood that the scale of the depth information is a standard for describing the depth, such as centimeters, meters, feet, inches, etc.
[0068] S209: Fusing the depth image of the first area and the depth image of the second area based on the normalization factor to obtain a third depth image of the object of interest.
[0069] Specifically, the depth image of the first area and / or the depth image of the second area is normalized according to the normalization factor, and then the normalized depth image of the first area and the depth image of the second area are fused to obtain a third image of the target of interest.
[0070] S210: Generate a target depth image based on the third depth image and the first depth image.
[0071] For details, please refer to step S104, which will not be described again here.
[0072] In an embodiment of the present application, a normalization factor between the first depth image and the second depth image is calculated and then normalized to obtain a first depth image and a second depth image of the same scale to ensure the accuracy of the third depth image obtained by image fusion, that is, to avoid image fusion errors due to the different scales of the first depth image and the second depth image, or to avoid the situation where the third depth image obtained by image fusion cannot correctly represent the depth information.
[0073] Please refer to Figure 3 , provides a schematic diagram of a process for obtaining a normalization factor for an embodiment of the present application. Figure 3 As shown, the method may include the following steps S301 to S302.
[0074] S301 : Calculate first average depth information corresponding to each piece of first depth information and calculate second average depth information corresponding to each piece of second depth information.
[0075] Specifically, first depth information corresponding to the first area is obtained, that is, the first pixel value corresponding to each pixel point in the first area is obtained, and then the average pixel value of each first pixel value is calculated to obtain the first average depth information. Similarly, second depth information corresponding to the second area is also obtained, that is, the second pixel value corresponding to each pixel point in the second area is obtained, and then the average pixel value of each second pixel value is calculated to obtain the second average depth information.
[0076] Exemplarily, the first average depth information may be obtained by the following formula:
[0077]
[0078] in, is the first average depth information; is the number of pixels in the first area; refers to the pixel in the first area; is the pixel value corresponding to the pixel point.
[0079] Exemplarily, the first average depth information may be obtained by the following formula:
[0080]
[0081] Among them, among them, is the second average depth information; is the number of pixels in the second area; refers to the pixel in the second area; is the pixel value corresponding to the pixel point.
[0082] S302: Generate a normalization factor based on the first average depth information and the second average depth information.
[0083] Specifically, a quotient between the first average depth information and the second average depth information is obtained, and the obtained quotient is used as a normalization factor.
[0084] In the embodiment of the present application, by acquiring the first average depth information and the second average depth information, the accuracy of the relevant calculation basis is improved, thereby improving the accuracy of the normalization factor.
[0085] See also Figure 4 , which is a flow chart of a depth estimation method provided in an embodiment of the present application. Figure 4 As shown, the method may include the following steps S401 to S408.
[0086] S401, acquiring a first depth image of a target image based on a depth estimation model, and determining a first area where an object of interest is located in the first depth image.
[0087] Please refer to step S101 for details, which will not be described again here.
[0088] S402: Acquire a second depth image of the target image based on the three-dimensional reconstruction model, and determine a second area where the target of interest is located in the second depth image.
[0089] For details, please refer to step S102, which will not be described again here.
[0090] S403: Obtain first depth information corresponding to each pixel in the first area and second depth information corresponding to each pixel in the second area.
[0091] S404: Calculate first average depth information corresponding to each piece of first depth information and calculate second average depth information corresponding to each piece of second depth information.
[0092] Please refer to step S301 for details, which will not be described again here.
[0093] S405 , obtaining a sum of the second average depth information and a preset value, and obtaining a quotient of the first average depth information and the sum, and using the quotient as a normalization factor.
[0094] Specifically, since the second average depth information may be zero, it is necessary to first perform relevant preprocessing on the second average depth information, that is, add the second average depth information to a preset value to obtain a sum between the two. It can be understood that the preset value is a very small positive number. Then, based on the obtained sum, the quotient between the first average depth and the obtained sum is obtained, and the obtained quotient is used as a normalization factor.
[0095] Exemplarily, the normalization factor can be obtained by the following formula:
[0096]
[0097] in, is the normalization factor; is the first average depth information; is the second average depth information; is the default value.
[0098] S406: Normalize the second depth image based on the normalization factor to obtain a normalized second depth image.
[0099] Specifically, each pixel in the second depth image is normalized based on the normalization factor to obtain a normalized second depth image.
[0100] Exemplarily, the second depth image may be normalized by the following formula:
[0101]
[0102] in, is the second depth image; is the normalization factor.
[0103] S407 , fusing the depth image of the first area and the depth image of the second area in the normalized second depth image based on an image fusion algorithm to obtain a third depth image of the object of interest.
[0104] S408: Generate a target depth image based on the third depth image and the first depth image.
[0105] Optionally, the image fusion algorithm can be a Poisson fusion algorithm. After obtaining the normalized second depth image, the first gradient field of the first depth image is calculated, the second gradient field of the depth image of the second area in the second depth image is calculated, and then the third gradient field of the target depth image is calculated. The fusion divergence of the target depth image is calculated based on the first gradient field, the second gradient field and the third gradient field, and then the coefficient matrix is calculated based on the fusion divergence. Finally, the target depth image is reconstructed based on the obtained coefficient matrix. The gradient field. It should be noted that the gradient field includes the gradient value of each pixel point in the corresponding area of the gradient field, and the target depth image includes the depth image in other areas except the first area in the first depth image, and the depth image in the second area in the second depth image. It can be understood that the third depth image is a depth image obtained by reconstructing the gradient field of the depth image in the second area based on the coefficient matrix. Further, the foregoing description is only an image fusion process based on the Poisson fusion algorithm, and does not mean that the present application can only perform image fusion through this image fusion algorithm.
[0106] In an embodiment of the present application, a preset value that does not affect the calculation process, that is, a very small positive number, is set, and then the preset value is added to the second average depth information, and a normalization factor is calculated based on this to avoid calculation errors caused by the second average depth information being zero.
[0107] See also Figure 5 , which is a flow chart of a depth estimation method provided in an embodiment of the present application. Figure 5 As shown, the method may include the following steps S501 to S507.
[0108] S501, acquiring a first depth image of a target image based on a depth estimation model, and determining a first sub-region where each target of interest is located in the first depth image;
[0109] Specifically, after the processor receives the target image for depth estimation, the received target image is input into the depth estimation model so that the depth estimation model performs depth estimation on the received target image, generates a first depth image of the target image, and outputs the obtained first depth image, and then determines the first sub-area where each target of interest is located in the first depth image. It should be noted that each target of interest corresponds to a separate first sub-area.
[0110] Optionally, a target image mask corresponding to each target of interest may be obtained, and then the first sub-region corresponding to each target of interest may be determined based on each target image mask.
[0111] S502, acquiring a second depth image of the target image based on the three-dimensional reconstruction model, and determining a second sub-region where each target of interest is located in the second depth image;
[0112] Specifically, after the processor receives the target image for depth estimation, the received target image is input into the three-dimensional reconstruction model so that the three-dimensional reconstruction model performs depth estimation on the received target image, generates a second depth image of the target image, and outputs the obtained second depth image, and then determines the second sub-area where each target of interest in the second depth image is located. It should be noted that each target of interest corresponds to a separate second sub-area.
[0113] S503, traverse the target of interest;
[0114] Specifically, the objects of interest are traversed, and the first sub-region and the second sub-region corresponding to the currently traversed object of interest are obtained.
[0115] S504, fusing the depth image of the first sub-region corresponding to the currently traversed target of interest and the depth image of the second sub-region corresponding to the currently traversed target of interest to obtain a third depth image of the currently traversed target of interest;
[0116] Optionally, the first pixel value of each pixel in the depth image of the first sub-area can be obtained first, and then the first average depth information of the first sub-area can be calculated based on the first pixel value of each pixel; at the same time, the second pixel value of each pixel in the depth image of the second sub-area can be obtained, and then the second average depth information of the second sub-area can be calculated based on the second pixel value of each pixel; then the normalization factor is calculated based on the first average depth information and the second average depth information, and then the second depth image is normalized based on the normalization factor to obtain the second depth information corresponding to the currently traversed target of interest. It should be noted that each time the second depth image is normalized, it is normalized based on the initial second depth image, that is, during the traversal of the target of interest, the normalization process of the second depth image does not change the original second depth image. Finally, the image fusion is performed based on the depth image of the first sub-area and the normalized depth image of the second sub-area to obtain the third depth image of the currently traversed target of interest.
[0117] S505, if it is the first traversal, generating a target depth image based on the third depth image of the current target of interest and the first depth image;
[0118] Specifically, if it is the first traversal, the third depth image is merged with the first depth image, that is, firstly obtain the image to be merged corresponding to other areas of the first depth image except the first sub-area where the target of interest is located, and then merge the third depth image with the image to be merged according to their corresponding areas to obtain the target depth image.
[0119] S506, if it is not the first traversal, generating a new target depth image based on the third depth image of the target of interest currently traversed and the target depth image generated last time;
[0120] Specifically, if it is not the first traversal, the third depth image is merged with the target depth image generated by the previous traversal process, that is, first obtain the images to be merged corresponding to other areas of the target depth image generated by the previous traversal process except the first sub-area where the target of interest is located, and then merge the third depth image with the image to be merged according to their corresponding areas to obtain a new target depth image generated by the current traversal process.
[0121] S507: If all the interested targets are traversed, a new target depth image is used as the target depth image.
[0122] Specifically, if all the objects of interest are traversed, the target depth image obtained after the last traversal is used as the final target depth image, that is, the image fusion result of the second sub-regions corresponding to all the objects of interest and the first depth image.
[0123] In the embodiment of the present application, the area where each target of interest is located is first determined, and then image fusion is performed on each target of interest respectively, so as to calculate a more matching fusion parameter for each target of interest, thereby improving the image fusion effect.
[0124] The following will be combined with the attached Figure 6 The depth estimation device provided in the embodiment of the present application is described in detail. Figure 6 Depth estimation device, used to execute this application Figure 1~Figure 5 For the convenience of explanation, only the part related to the embodiment of the present application is shown. For the specific technical details not disclosed, please refer to the present application. Figure 1~Figure 5 The embodiment shown.
[0125] See also Figure 6 , is a schematic diagram of the structure of a depth estimation device provided in an embodiment of the present application. Figure 6 As shown, the depth estimation device 1 of the embodiment of the present application may include: a first acquisition module 10 , a second acquisition module 20 , an image fusion module 30 , and an image generation module 40 .
[0126] A first acquisition module 10 is used to acquire a first depth image of the target image based on the depth estimation model, and determine a first area where the target of interest is located in the first depth image;
[0127] A second acquisition module 20, configured to acquire a second depth image of the target image based on the three-dimensional reconstruction model, and determine a second region where the target of interest is located in the second depth image;
[0128] An image fusion module 30, configured to fuse the depth image of the first area and the depth image of the second area to obtain a third depth image of the object of interest;
[0129] The image generating module 40 is configured to generate a target depth image based on the third depth image and the first depth image.
[0130] In an embodiment of the present application, by acquiring a first depth image with a high dynamic range and a second depth image with a low dynamic range, and then performing image fusion on the first depth image and the second depth image according to the area where the target of interest is located, a target depth image in which the area where the target of interest is located is adaptively acquired as a high dynamic range and the area where the target of non-interest is located is a low dynamic range, thereby obtaining a target depth image with higher accuracy.
[0131] Optionally, the image fusion module 30 is specifically used for:
[0132] Acquire first depth information corresponding to each pixel in the first area and second depth information corresponding to each pixel in the second area;
[0133] Acquire a normalization factor based on first depth information corresponding to each pixel in the first area and second depth information corresponding to each pixel in the second area;
[0134] The depth image of the first area and the depth image of the second area are fused based on the normalization factor to obtain a third depth image of the object of interest.
[0135] Optionally, the image fusion module 30 is specifically used for:
[0136] Calculating first average depth information corresponding to each piece of first depth information and calculating second average depth information corresponding to each piece of second depth information;
[0137] A normalization factor is generated based on the first average depth information and the second average depth information.
[0138] Optionally, the image fusion module 30 is specifically used for:
[0139] Obtaining a sum of the second average depth information and a preset value, and obtaining a quotient of the first average depth information and the sum, and using the quotient as a normalization factor;
[0140] The depth image of the first area and the depth image of the second area are fused based on the normalization factor to obtain a third depth image of the object of interest, including:
[0141] Normalizing the second depth image based on the normalization factor to obtain a normalized second depth image;
[0142] The depth image of the first area and the depth image of the second area in the normalized second depth image are fused based on an image fusion algorithm to obtain a third depth image of the object of interest.
[0143] Optionally, the first acquisition module 10 is further used for:
[0144] Acquire a first depth image of the target image based on the depth estimation model, and determine a first sub-region where each target of interest is located in the first depth image;
[0145] The second acquisition module 20 is further used for:
[0146] Acquire a second depth image of the target image based on the three-dimensional reconstruction model, and determine a second sub-region where each target of interest is located in the second depth image;
[0147] A traversal module 50, used to traverse the objects of interest;
[0148] The image fusion module 30 is further used for:
[0149] Fusing the depth image of the first sub-region corresponding to the currently traversed target of interest and the depth image of the second sub-region corresponding to the currently traversed target of interest to obtain a third depth image of the currently traversed target of interest;
[0150] The image generation module 40 is further used for:
[0151] If it is the first traversal, a target depth image is generated based on the third depth image of the current target of interest and the first depth image;
[0152] If it is not the first traversal, a new target depth image is generated based on the third depth image of the target of interest currently traversed and the target depth image generated last time;
[0153] If all the targets of interest are traversed, the new target depth image is used as the target depth image.
[0154] Optionally, the depth estimation device 1 further includes: a mask acquisition module 60 .
[0155] The mask acquisition module 60 is used to acquire a target image mask where the target of interest is located in the target image;
[0156] The first acquisition module 10 is further used to determine a first area where the object of interest is located in the first depth image based on the target image mask;
[0157] The second acquisition module 20 is further configured to determine a second region where the object of interest is located in the second depth image based on the object image mask.
[0158] Optionally, the depth estimation device 1 further includes: an image correction module 70 .
[0159] The image correction module 70 is specifically used for:
[0160] If the image size of the first depth image is inconsistent with the image size of the target image, performing image sampling on the first depth image to obtain a sampled image that is consistent with the image size of the target image;
[0161] Replace the first depth image with the sampled image.
[0162] The present application also provides a storage medium that can store multiple program instructions, which are suitable for being loaded and executed by a processor as described above. Figure 1 to Figure 5 The method steps of the embodiment shown in the figure can be found in the specific implementation process. Figure 1 to Figure 5 The specific description of the illustrated embodiment will not be repeated here.
[0163] See also Figure 7 , which is a schematic diagram of the structure of a computer device according to an embodiment of the present application. Figure 7 As shown, the computer device 1000 may include: at least one processor 1001, at least one memory 1002, at least one network interface 1003, at least one input and output interface 1004, at least one communication bus 1005 and at least one display unit 1006. Among them, the processor 1001 may include one or more processing cores. The processor 1001 uses various interfaces and lines to connect the various parts of the entire computer device 1000, and executes various functions and processes data of the terminal 1000 by running or executing instructions, programs, code sets or instruction sets stored in the memory 1002, and calling data stored in the memory 1002. The memory 1002 can be a high-speed RAM memory or a non-volatile memory (non-volatile memory), such as at least one disk storage. The memory 1002 can optionally be at least one storage device located away from the aforementioned processor 1001. Among them, the network interface 1003 can optionally include a standard wired interface, a wireless interface (such as a WI-FI interface). The communication bus 1005 is used to realize the connection and communication between these components. As shown in FIG. Figure 7 As shown, the memory 1002 as a storage medium of a terminal device may include an operating system, a network communication module, an input and output interface module, and a depth estimation program.
[0164] exist Figure 7 In the computer device 1000 shown, the input and output interface 1004 is mainly used to provide an input interface for the user and the access device, and to obtain data input by the user and the access device.
[0165] In one embodiment.
[0166] The processor 1001 may be used to call the depth estimation program stored in the memory 1002, and specifically perform the following operations:
[0167] Acquire a first depth image of the target image based on the depth estimation model, and determine a first area where the target of interest is located in the first depth image;
[0168] Acquire a second depth image of the target image based on the three-dimensional reconstruction model, and determine a second area where the target of interest is located in the second depth image;
[0169] Fusing the depth image of the first area and the depth image of the second area to obtain a third depth image of the object of interest;
[0170] A target depth image is generated based on the third depth image and the first depth image.
[0171] Optionally, when the processor 1001 performs fusion of the depth image of the first area and the depth image of the second area to obtain a third depth image of the object of interest, the processor 1001 specifically performs the following operations:
[0172] Acquire first depth information corresponding to each pixel in the first area and second depth information corresponding to each pixel in the second area;
[0173] Acquire a normalization factor based on first depth information corresponding to each pixel in the first area and second depth information corresponding to each pixel in the second area;
[0174] The depth image of the first area and the depth image of the second area are fused based on the normalization factor to obtain a third depth image of the object of interest.
[0175] Optionally, when the processor 1001 acquires the normalization factor based on the first depth information corresponding to each pixel in the first area and the second depth information corresponding to each pixel in the second area, the processor 1001 specifically performs the following operations:
[0176] Calculating first average depth information corresponding to each piece of first depth information and calculating second average depth information corresponding to each piece of second depth information;
[0177] A normalization factor is generated based on the first average depth information and the second average depth information.
[0178] Optionally, when the processor 1001 generates a normalization factor based on the first average depth information and the second average depth information, the processor 1001 specifically performs the following operations:
[0179] Obtaining a sum of the second average depth information and a preset value, and obtaining a quotient of the first average depth information and the sum, and using the quotient as a normalization factor;
[0180] When the processor 1001 performs fusion of the depth image of the first area and the depth image of the second area based on the normalization factor to obtain the third depth image of the object of interest, the following operations are specifically performed:
[0181] Normalizing the second depth image based on the normalization factor to obtain a normalized second depth image;
[0182] The depth image of the first area and the depth image of the second area in the normalized second depth image are fused based on an image fusion algorithm to obtain a third depth image of the object of interest.
[0183] Optionally, the processor 1001 may also be used to call a depth estimation program stored in the memory 1002, and specifically perform the following operations:
[0184] Acquire a first depth image of the target image based on the depth estimation model, and determine a first sub-region where each object of interest is located in the first depth image;
[0185] Acquire a second depth image of the target image based on the three-dimensional reconstruction model, and determine a second sub-region where each target of interest is located in the second depth image;
[0186] Traverse the objects of interest;
[0187] Fusing the depth image of the first sub-region corresponding to the currently traversed target of interest and the depth image of the second sub-region corresponding to the currently traversed target of interest to obtain a third depth image of the currently traversed target of interest;
[0188] If it is the first traversal, a target depth image is generated based on the third depth image of the current target of interest and the first depth image;
[0189] If it is not the first traversal, a new target depth image is generated based on the third depth image of the target of interest currently traversed and the target depth image generated last time;
[0190] If all the targets of interest are traversed, the new target depth image is used as the target depth image.
[0191] Optionally, before determining the first area where the object of interest is located in the first depth image, the processor 1001 further performs the following operations:
[0192] Obtain a target image mask where the target of interest is located in the target image;
[0193] Determining a first area in the first depth image where the object of interest is located includes:
[0194] Determine a first region where the target of interest is located in the first depth image based on the target image mask;
[0195] Determining a second area in the second depth image where the object of interest is located includes:
[0196] A second region in the second depth image where the object of interest is located is determined based on the object image mask.
[0197] Optionally, after executing the process of acquiring the first depth image of the target image based on the depth estimation model, the processor 1001 further performs the following operations:
[0198] If the image size of the first depth image is inconsistent with the image size of the target image, performing image sampling on the first depth image to obtain a sampled image that is consistent with the image size of the target image;
[0199] Replace the first depth image with the sampled image.
[0200] In an embodiment of the present application, by acquiring a first depth image with a high dynamic range and a second depth image with a low dynamic range, and then performing image fusion on the first depth image and the second depth image according to the area where the target of interest is located, a target depth image in which the area where the target of interest is located is adaptively acquired as a high dynamic range and the area where the target of non-interest is located is a low dynamic range, thereby obtaining a target depth image with higher accuracy.
[0201] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.
[0202] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0203] The above is a description of a depth estimation method, a depth estimation device, a storage medium and a device provided in the present application. For technicians in this field, according to the ideas of the embodiments of the present application, there may be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A depth estimation method, characterized in that: The method comprises: Acquire a first depth image of the target image based on the depth estimation model, and determine a first area where the target of interest is located in the first depth image; Acquire a second depth image of the target image based on the three-dimensional reconstruction model, and determine a second area where the target of interest is located in the second depth image; Acquire first depth information corresponding to each pixel in the first area and second depth information corresponding to each pixel in the second area; Acquire a normalization factor based on first depth information corresponding to each pixel in the first area and second depth information corresponding to each pixel in the second area; fusing the depth image of the first area and the depth image of the second area based on the normalization factor to obtain a third depth image of the object of interest; A target depth image is generated based on the third depth image and the first depth image.
2. The method according to claim 1, characterized in that The obtaining a normalization factor based on first depth information corresponding to each pixel in the first area and second depth information corresponding to each pixel in the second area includes: Calculating first average depth information corresponding to each piece of the first depth information and calculating second average depth information corresponding to each piece of the second depth information; The normalization factor is generated based on the first average depth information and the second average depth information.
3. The method according to claim 2, characterized in that The generating the normalization factor based on the first average depth information and the second average depth information includes: Acquire a sum of the second average depth information and a preset value, and acquire a quotient of the first average depth information and the sum, and use the quotient as the normalization factor; The fusing the depth image of the first area and the depth image of the second area based on the normalization factor to obtain a third depth image of the object of interest includes: Normalizing the second depth image based on the normalization factor to obtain a normalized second depth image; The depth image of the first area and the depth image of the second area in the normalized second depth image are fused based on an image fusion algorithm to obtain a third depth image of the object of interest.
4. The method according to claim 3, characterized in that: The image fusion algorithm includes Poisson fusion.
5. The method according to claim 1, characterized in that: If there are at least two objects of interest in the target image, the method includes: Acquire a first depth image of the target image based on the depth estimation model, and determine a first sub-region where each of the objects of interest are located in the first depth image; Acquire a second depth image of the target image based on the three-dimensional reconstruction model, and determine a second sub-region where each of the objects of interest are located in the second depth image; Traversing the objects of interest; Fusing the depth image of the first sub-region corresponding to the currently traversed target of interest and the depth image of the second sub-region corresponding to the currently traversed target of interest to obtain a third depth image of the currently traversed target of interest; If it is the first traversal, generating a target depth image based on the third depth image of the current target of interest and the first depth image; If it is not the first traversal, generating a new target depth image based on the third depth image of the target of interest currently traversed and the target depth image generated last time; If all the objects of interest are traversed, the new object depth image is used as the target depth image.
6. The method according to claim 1, characterized in that Before determining the first area where the object of interest is located in the first depth image, the method further includes: Acquire a target image mask where the target of interest is located in the target image; The determining a first area where the object of interest is located in the first depth image includes: Determine a first region where the target of interest is located in the first depth image based on the target image mask; The determining a second area where the object of interest is located in the second depth image includes: A second region in the second depth image where the object of interest is located is determined based on the object image mask.
7. The method according to claim 1, characterized in that After acquiring the first depth image of the target image based on the depth estimation model, the method further includes: If the image size of the first depth image is inconsistent with the image size of the target image, performing image sampling on the first depth image to obtain a sampled image that is consistent with the image size of the target image; The first depth image is replaced by the sampled image.
8. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the depth estimation method according to any one of claims 1 to 7 is implemented.
9. A computer device, characterized in that: include: A processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the steps of the depth estimation method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Depth estimation method and device, equipment and storage medium
CN109472821A
Monitoring video target detection method based on background elimination
CN109993091A