Method and electronic device for acquiring depth image

By combining the first depth image captured by the electronic device and the second depth image obtained by processing the RGB image, the first point cloud data of the target scene is generated, which solves the incompleteness problem when the active light depth camera obtains the depth image, realizes the integrity and accuracy of the target depth image, and reduces cost and power consumption.

CN114445478BActive Publication Date: 2025-09-26HUAWEI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202011211807.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-11-03
Publication Date
2025-09-26
Estimated Expiration
2040-11-03

AI Technical Summary

Technical Problem

When existing active optical depth cameras acquire depth images, the depth in some images cannot be determined due to the presence of light-absorbing and highly reflective materials in the real scene, resulting in incomplete depth images and incomplete content of the target scene.

Method used

By combining the first depth image captured by the electronic device and the second depth image obtained by processing the RGB image, and utilizing the accuracy of the first depth image and the completeness of the second depth image, first point cloud data of the target scene is generated, thereby obtaining a complete and accurate target depth image.

Benefits of technology

The image integrity and depth value accuracy of the target scene in the acquired depth image are improved, production costs are reduced, and power consumption of electronic devices is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114445478B_ABST
    Figure CN114445478B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a method for obtaining a depth image and an electronic device, which relates to the technical field of image processing and can solve the problems that the obtained depth image is incomplete and the content of the target scene in the depth image is incomplete. The specific solution includes: the electronic device can obtain a first depth image and a second depth image; wherein, the first depth image is a depth image of the target scene collected by the electronic device, the first depth image includes N pixel points, the second depth image is a depth image obtained by processing the RGB image of the target scene, the second depth image includes M pixel points, both M and N are integers, and 1 < N < M. After that, the electronic device can obtain the first point cloud data of the target scene according to the first depth image and the second depth image. Finally, the electronic device can obtain the target depth image of the target scene according to the first point cloud data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the technical field of electronic devices, and in particular to a method for acquiring a depth image and an electronic device. Background Art

[0002] With the development of science and technology, augmented reality (AR) devices (such as mobile phones or glasses, etc.) have gradually entered people's lives and are widely used in industry, education, entertainment and other fields. Specifically, users can see the interaction effects between virtual content and real scenes (such as occlusion and collision between objects) through AR devices. The realization of the above-mentioned interactive effects depends on the depth image provided by the AR device to the AR application. Among them, the depth image refers to an image that uses the distance (depth) from the image collector to each point in the scene as the pixel value, which directly reflects the geometric shape of the visible surface of the object.

[0003] Currently, active light depth cameras (such as infrared structured light cameras) can be used to acquire depth images. When acquiring depth images, active light depth cameras can calculate the depth of each point in the image by analyzing the time it takes for light to reach each point.

[0004] However, when active optical depth cameras acquire depth images, the presence of a large number of light-absorbing and highly reflective materials in real scenes prevents the light emitted by the active optical depth camera from returning. This results in uncertainty about the depth of some points in the image. This inability to determine the depth of some points in the image results in incomplete depth images, and the content of the target scene in the depth image is incomplete. Summary of the Invention

[0005] The present application provides a method for acquiring a depth image, which can solve the problem that the acquired depth image is incomplete and the content of the target scene in the depth image is incomplete.

[0006] In a first aspect, the present application provides a method for acquiring a depth image, which can be applied to an electronic device.

[0007] The electronic device can obtain a first depth image and a second depth image respectively. The first depth image is a depth image of the target scene captured by the electronic device (such as a camera of the electronic device), and the first depth image includes N pixels. The second depth image is a depth image obtained by processing the RGB image of the target scene, that is, the second depth image is not a depth image directly captured by the electronic device, and the second depth image includes M pixels, where M and N are both integers.

[0008] Generally speaking, the depth values of the depth images directly captured by an electronic device (such as the above-mentioned first depth image) are relatively accurate. For example, the depth values of the depth images captured by an electronic device through a Time-of-Flight (ToF) camera are relatively accurate. However, the number of pixel points in such depth images is less, so the image integrity of the target scene in this depth image is relatively low. The above-mentioned second depth image is a depth image obtained by processing the RGB image of the target scene; therefore, the number of pixel points in the second depth image is more (such as 1 < N < M), and the image integrity of the target scene in this second depth image is relatively high; however, the accuracy of the depth values of this second depth image is relatively low.

[0009] Based on this, the electronic device can obtain the first point cloud data of the target scene according to the first depth image and the second depth image. The first point cloud data includes the two-dimensional features and depth values of M first feature points, and the M first feature points correspond to the M pixel points one by one. Among them, the depth values of the M first feature points are obtained by the electronic device according to the depth values of the M pixel points and the depth values of the N pixel points. Then, the electronic device can obtain the target depth image of the target scene according to the first point cloud data.

[0010] That is to say, the electronic device can combine the respective advantages of the above-mentioned first depth image and the second depth image (such as the depth values of the first depth image being relatively accurate and the image integrity of the target scene in the second depth image being relatively high), and obtain the first point cloud data of the target scene according to the first depth image and the second depth image. Therefore, the first point cloud data can present a 3D point cloud of the above-mentioned target scene that is relatively complete and has relatively accurate depth values. Among them, the target depth image of the target scene is obtained according to the first point cloud data. The image integrity of the target scene in this target depth image is relatively high, and the depth values are relatively accurate.

[0011] In summary, by using the method of this application, the image integrity of the target scene in the obtained target depth image can be improved, and the accuracy of the depth values can be improved.

[0012] In combination with the first aspect, in a possible design method, the above-mentioned method of "the electronic device can obtain the first point cloud data of the target scene based on the first depth image and the second depth image" includes: the electronic device can obtain the second point cloud data based on the first depth image; wherein the second point cloud data includes the two-dimensional features and depth values ​​of N feature points, the N feature points correspond one-to-one to N pixels, and the depth values ​​of the N feature points are obtained based on the depth values ​​of the N pixels. In addition, the electronic device can obtain the third point cloud data based on the second depth image; wherein the third point cloud data includes the two-dimensional features and depth values ​​of M second feature points, the M second feature points correspond one-to-one to M pixels, and the depth values ​​of the M second feature points are obtained based on the depth values ​​of the M pixels. Afterwards, the electronic device can obtain the first point cloud data based on the second point cloud data and the third point cloud data.

[0013] It can be understood that since the depth values ​​of the pixels in the first depth image are relatively accurate, and the N feature points included in the second point cloud data correspond one-to-one to the N pixels in the second depth image (the displayed content is the target scene). Therefore, the N feature points in the second point cloud data can accurately record the target scene. Since the number of pixels in the second depth image is relatively large, and the M feature points included in the third point cloud data correspond one-to-one to the M pixels in the second depth image. Therefore, the M feature points in the third point cloud data can completely record the target scene. Furthermore, since the second point cloud data can accurately record the target scene, the third point cloud data can completely record the target scene. Therefore, the electronic device can obtain the first point cloud data that can accurately and completely record the target scene based on the second point cloud data and the third point cloud data.

[0014] In conjunction with the first aspect, in another possible design, the method of "an electronic device may obtain third point cloud data based on a second depth image" includes: the electronic device scales the depth values ​​of pixels in the second depth image based on the depth values ​​of the pixels in the first depth image to obtain a third depth image; wherein the difference between the depth values ​​of the pixels in the third depth image and the depth values ​​of the corresponding pixels in the first depth image is less than the difference between the depth values ​​of the corresponding pixels in the second depth image and the depth values ​​of the corresponding pixels in the first depth image. Thereafter, the electronic device may obtain third point cloud data based on the third depth image.

[0015] It is understandable that the electronic device scales the depth values ​​of the pixels in the second depth image based on the depth values ​​of the pixels in the first depth image to obtain a third depth image. Since the depth values ​​of the pixels in the first depth image are more accurate, the depth values ​​of the third depth image obtained by the electronic device are more accurate than those of the second depth image. That is, the difference between the depth value of the pixel in the third depth image and the depth value of the corresponding pixel in the first depth image is less than the difference between the depth value of the corresponding pixel in the second depth image and the depth value of the corresponding pixel in the first depth image.

[0016] In addition, since the second depth image is a depth image obtained by processing the RGB image of the target scene, the depth value of each pixel in the second depth image has no length unit (such as meters, centimeters, etc.). Therefore, the electronic device can scale the depth value of the pixel in the second depth image according to the depth value of the pixel in the first depth image to obtain the depth value of each pixel in the third depth image and its length unit. In this way, compared with the point cloud data obtained by the electronic device based on the second depth image, the third point cloud data obtained by the electronic device based on the third depth image can record the target scene more accurately.

[0017] In conjunction with the first aspect, in another possible design, the method of "the electronic device may scale the depth values ​​of pixels in the second depth image based on the depth values ​​of the pixels in the first depth image to obtain a third depth image" includes: the electronic device may determine a first median value and a second median value; wherein the first median value is the median value of the depth values ​​of N pixels; and the second median value is the median value of the depth values ​​of M pixels. The electronic device may then scale the depth values ​​of the M pixels in the second depth image according to a ratio between the first median value and the second median value to obtain the third depth image.

[0018] It can be understood that since the median value can represent the central tendency of a set of data, and the first median value is the median value of the depth values ​​of N pixels, and the second median value is the median value of the depth values ​​of M pixels. Therefore, the electronic device can scale the depth values ​​of the M pixels in the second depth image according to the ratio of the first median value to the second median value to obtain a third depth image. In other words, the electronic device adjusts the depth values ​​of the pixels in the second depth image according to the depth values ​​of the pixels in the first depth image to obtain the depth values ​​of the third depth image. And since the depth values ​​of the pixels in the first depth image are more accurate. Therefore, compared with the depth values ​​of the pixels in the second depth image, the depth values ​​of the third depth image obtained by the electronic device are more accurate.

[0019] For example, the first depth image includes depth values ​​of 3 pixels, which are 1 meter, 5 meters, and 8 meters respectively, and the first median value is 5 meters. The second depth image includes depth values ​​of 5 pixels, which are 2, 6, 10, 12, and 20 respectively, and the second median value is 10. Since the ratio between the first median value and the second median value is 0.5, the electronic device can reduce the depth values ​​of the 5 pixels in the second depth image to 1 / 2 of the original. That is, the depth values ​​of the 5 pixels in the reduced second depth image are 1 meter, 3 meters, 5 meters, 6 meters, and 10 meters respectively.

[0020] In combination with the first aspect, in another possible design, the electronic device may process the depth image to obtain point cloud data using the following formula:

[0021] z i =d i ;

[0022] Among them, (x i ,y i , z i ) represents the three-dimensional features of the feature point in the point cloud data corresponding to the i-th pixel in the depth image, (u i , v i ) represents the two-dimensional feature of the i-th pixel in the depth image, d i represents the depth value of the i-th pixel in the depth image, (u0, v0) represents the two-dimensional feature of the center point of the depth image, and f x Represents the first focal length of the depth image, f y Indicates the second focal length of the depth image.

[0023] In combination with the first aspect, in another possible design method, the above-mentioned method of "the electronic device can obtain the first point cloud data based on the second point cloud data and the third point cloud data" includes: the electronic device can determine N target points corresponding to the N feature points of the second point cloud data from the M feature points of the third point cloud data; wherein the two-dimensional features of the N feature points are the same as the two-dimensional features of the N target points. Afterwards, the electronic device can adjust the three-dimensional features of the N target points based on the three-dimensional features of the N feature points, and adjust the three-dimensional features of MN traction points in the second point cloud data, where the MN traction points are points other than the N target points in the third point cloud data, and the three-dimensional features include two-dimensional features and depth values. Then, the electronic device can obtain the first point cloud data based on the adjusted three-dimensional features of the N target points and the three-dimensional features of the MN traction points.

[0024] It can be understood that since the second point cloud data is obtained from the first depth image, and the depth value of the first depth image is relatively accurate. Therefore, the three-dimensional features of the second point cloud data are also relatively accurate. Therefore, the electronic device can adjust the three-dimensional features of the N target points according to the three-dimensional features of the N feature points, and adjust the three-dimensional features of the MN traction points in the second point cloud data (the MN traction points are points other than the N target points in the third point cloud data) to obtain the first point cloud data, which includes the three-dimensional features of the adjusted N target points and the three-dimensional features of the MN traction points. Since the three-dimensional features of the N feature points in the second point cloud data are relatively accurate. Therefore, the first point cloud data obtained by the electronic device after adjusting the three-dimensional features of the N feature points is also relatively accurate.

[0025] In combination with the first aspect, in another possible design method, the electronic device can obtain the first point cloud data by using the following formula:

[0026]

[0027] Among them, v i ′ represents the three-dimensional feature of the i-th target point among the N target points after adjustment, Represents the three-dimensional features of the feature point corresponding to the i-th target point in the second point cloud data, v′ i,2 represents the three-dimensional features of the i-th traction point among the MN traction points after adjustment, represents the desired three-dimensional feature of the i-th traction point, λ is the coefficient, 0<λ<1.

[0028] In combination with the first aspect, in another possible design, the electronic device may process the point cloud data to obtain a depth image using the following formula:

[0029] d i =z i .

[0030] In combination with the first aspect, in another possible design method, the above-mentioned method for obtaining a depth image also includes: the electronic device can perform image semantic segmentation on the RGB image, and use an artificial intelligence AI algorithm to calculate the segmented RGB image to obtain a second depth image.

[0031] It should be noted that since there may be two objects (also referred to as objects) connected or partially overlapped in the RGB image, and the electronic device cannot distinguish which object the edge where the two objects contact belongs to. Therefore, in the depth image obtained by the electronic device using the artificial intelligence AI algorithm, the depth value at the edge where the two objects contact is inaccurate. And image semantic segmentation can effectively distinguish two different objects. Therefore, after the electronic device performs image semantic segmentation on the RGB image, the edge contours of the two overlapping objects in the RGB image become clearer. In this way, for the second depth image obtained by the electronic device using the artificial intelligence AI algorithm to calculate the segmented RGB image, its depth value can be more accurate.

[0032] In a second aspect, an embodiment of the present application provides an electronic device, which includes: a memory and a processor, the memory is coupled to the processor; the memory is used to store computer program code, and the computer program code includes computer instructions; when the computer instructions are executed by the processor, the electronic device is caused to execute the method as described in the first aspect and any of its possible design manners.

[0033] In a third aspect, an embodiment of the present application provides a device for obtaining a depth image, the device includes: an AI information prediction module, a 2D-to-3D projection module, and a 3D-to-2D projection module.

[0034] The AI information prediction module is used to obtain a first depth image and a second depth image. The first depth image is the depth image of the target scene collected by the electronic device. The first depth image includes N pixel points, and the second depth image is the depth image obtained by processing the RGB image of the target scene. The second depth image includes M pixel points, and both M and N are integers, 1 < N < M; the depth image includes the pixel information of the target scene, and the pixel information includes the two-dimensional features and depth values of each pixel point.

[0035] The 2D-to-3D projection module is used to obtain the first point cloud data of the target scene according to the first depth image and the second depth image; wherein, the first point cloud data includes the two-dimensional features and depth values of M first feature points, and the M first feature points correspond to the M pixel points one by one; the depth values of the M first feature points are obtained according to the depth values of the M pixel points and the depth values of the N pixel points.

[0036] The 3D-to-2D projection module is used to obtain the target depth image of the target scene according to the first point cloud data.

[0037] In combination with the third aspect, in a possible design, the 3D to 2D projection module is specifically used to obtain second point cloud data based on the first depth image; wherein, the second point cloud data includes two-dimensional features and depth values ​​of N feature points, the N feature points correspond one-to-one to N pixels, and the depth values ​​of the N feature points are obtained based on the depth values ​​of the N pixels; third point cloud data is obtained based on the second depth image; wherein, the third point cloud data includes two-dimensional features and depth values ​​of M second feature points, the M second feature points correspond one-to-one to M pixels, and the depth values ​​of the M second feature points are obtained based on the depth values ​​of the M pixels; the first point cloud data is obtained based on the second point cloud data and the third point cloud data.

[0038] In conjunction with the third aspect, in another possible design, the apparatus further includes a scaling module. The scaling module is configured to scale the depth values ​​of pixels in the second depth image based on the depth values ​​of the pixels in the first depth image to obtain a third depth image; wherein the difference between the depth values ​​of the pixels in the third depth image and the depth values ​​of the corresponding pixels in the first depth image is less than the difference between the depth values ​​of the corresponding pixels in the second depth image and the depth values ​​of the corresponding pixels in the first depth image. The 2D-to-3D projection module is specifically configured to obtain third point cloud data based on the third depth image.

[0039] In combination with the third aspect, in another possible design, the scale scaling module is specifically used to determine a first median value and a second median value; wherein the first median value is the median value of the depth values ​​of N pixels; the second median value is the median value of the depth values ​​of M pixels; according to the ratio of the first median value to the second median value, the depth values ​​of the M pixels in the second depth image are scaled to obtain a third depth image.

[0040] In conjunction with the third aspect, in another possible design, the 2D to 3D projection module is specifically used to obtain point cloud data based on the depth image using the following formula:

[0041] z i =d i ;

[0042] Among them, (x i ,y i , z i ) represents the three-dimensional features of the feature point in the point cloud data corresponding to the i-th pixel in the depth image, (u i , v i ) represents the two-dimensional feature of the i-th pixel in the depth image, d i represents the depth value of the i-th pixel in the depth image, (u0, v0) represents the two-dimensional feature of the center point of the depth image, and f x and represents the first focal length of the depth image, fy Indicates the second focal length of the depth image.

[0043] In conjunction with the third aspect, in another possible design, the apparatus further includes a point cloud fusion module. The point cloud fusion module is configured to determine, from the M feature points in the third point cloud data, N target points corresponding to the N feature points in the second point cloud data; adjust the three-dimensional features of the N target points based on the three-dimensional features of the N feature points, and adjust the three-dimensional features of MN traction points in the second point cloud data, where the MN traction points are points other than the N target points in the third point cloud data, and the three-dimensional features include two-dimensional features and depth values; and obtain the first point cloud data based on the adjusted three-dimensional features of the N target points and the three-dimensional features of the MN traction points.

[0044] In conjunction with the third aspect, in another possible design, the point cloud fusion module is specifically configured to obtain the adjusted three-dimensional features of the N target points and the three-dimensional features of the MN traction points using the following formula:

[0045]

[0046] Among them, v i ′ represents the three-dimensional feature of the i-th target point among the N target points after adjustment, Represents the three-dimensional features of the feature point corresponding to the i-th target point in the second point cloud data, v′ i,2 represents the three-dimensional features of the i-th traction point among the MN traction points after adjustment, represents the desired three-dimensional feature of the i-th traction point, λ is the coefficient, 0<λ<1.

[0047] In conjunction with the third aspect, in another possible design, the 3D to 2D projection module is specifically used to obtain a depth image based on the point cloud data using the following formula:

[0048] d i =z i .

[0049] Optionally, the AI ​​information prediction module 1301 is further used to perform image semantic segmentation on the RGB image, and use an artificial intelligence AI algorithm to calculate the segmented RGB image to obtain a second depth image.

[0050] In a fourth aspect, an embodiment of the present application provides a chip system, which is applied to an electronic device. The chip system includes one or more interface circuits and one or more processors. The interface circuits and processors are interconnected via a circuit. The interface circuit is configured to receive a signal from a memory of the electronic device and send the signal to the processor, the signal including computer instructions stored in the memory. When the processor executes the computer instructions, the electronic device executes the method described in the first aspect and any possible design thereof.

[0051] In a fifth aspect, an embodiment of the present application provides a computer storage medium, which includes computer instructions. When the computer instructions are executed on an electronic device, the electronic device executes the method described in the first aspect and any possible design method thereof.

[0052] In a sixth aspect, an embodiment of the present application provides a computer program product, which, when executed on a computer, enables the computer to execute the method described in the first aspect and any possible design thereof.

[0053] It can be understood that the beneficial effects that can be achieved by the electronic device described in the second aspect and any possible design thereof provided above, the depth image acquisition device described in the third aspect and any possible design thereof, the chip system described in the fourth aspect, the computer storage medium described in the fifth aspect, and the computer program product described in the sixth aspect can refer to the beneficial effects in the first aspect and any possible design thereof, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 A schematic diagram of a depth image example provided in an embodiment of the present application;

[0055] Figure 2A A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application;

[0056] Figure 2B A schematic diagram of the software structure of an electronic device provided in an embodiment of the present application;

[0057] Figure 3 A schematic diagram of a flow chart of a method for acquiring a depth image provided in an embodiment of the present application;

[0058] Figure 4A A schematic diagram of an RGB image example provided in an embodiment of the present application;

[0059] Figure 4B A schematic diagram of another depth image example provided in an embodiment of the present application;

[0060] Figure 4CA schematic diagram of another depth image example provided in an embodiment of the present application;

[0061] Figure 4D A schematic diagram of another depth image example provided in an embodiment of the present application;

[0062] Figure 4E A schematic diagram of another depth image example provided in an embodiment of the present application;

[0063] Figure 5 A schematic diagram of another depth image example provided in an embodiment of the present application;

[0064] Figure 6 A schematic diagram of another depth image example provided in an embodiment of the present application;

[0065] Figure 7A A schematic diagram of a flow chart of another method for acquiring a depth image provided in an embodiment of the present application;

[0066] Figure 7B A schematic diagram of a flow chart of another method for acquiring a depth image provided in an embodiment of the present application;

[0067] Figure 8 A schematic diagram of a point cloud data example provided in an embodiment of the present application;

[0068] Figure 9 A schematic diagram of another point cloud data example provided in an embodiment of the present application;

[0069] Figure 10 A schematic diagram of another point cloud data example provided in an embodiment of the present application;

[0070] Figure 11 A schematic diagram of another point cloud data example provided in an embodiment of the present application;

[0071] Figure 12 A schematic diagram of a 3D point cloud space example provided in an embodiment of the present application;

[0072] Figure 13 A schematic diagram of the composition of a depth image acquisition device provided in an embodiment of the present application;

[0073] Figure 14 A schematic diagram of a flow chart of a method for acquiring a depth image provided in an embodiment of the present application;

[0074] Figure 15 A schematic diagram of the structural composition of a chip system provided in an embodiment of the present application. DETAILED DESCRIPTION

[0075] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0076] In this application, the character " / " generally indicates that the preceding and following objects are in an "or" relationship. For example, A / B can be understood as A or B.

[0077] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature identified as "first" or "second" may explicitly or implicitly include one or more of such features. In the description of this embodiment, unless otherwise specified, "plurality" means two or more.

[0078] Furthermore, the terms "including," "having," and any variations thereof, as used in the description of this application are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or modules is not limited to the listed steps or modules, but may optionally include other steps or modules not listed, or may optionally include other steps or modules inherent to the process, method, product, or apparatus.

[0079] Additionally, in the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described in this application as "exemplary" or "for example" should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present concepts in a concrete manner.

[0080] In order to facilitate understanding of the technical solution of the present application, before giving a detailed introduction to the method for obtaining a depth image in the embodiment of the present application, the professional terms mentioned in the embodiment of the present application are first introduced.

[0081] 1. Depth Image

[0082] A depth image, also known as a range image, refers to an image that uses the distance (depth) from the image collector to each point in the scene as a pixel value. It directly reflects the geometric shape of the visible surface of the scene. In other words, the depth value of a pixel in the depth image of the target scene represents the distance between a point in the target scene and the image collector (usually in millimeters). Generally, the farther a point in the scene is from the image collector, that is, the larger the depth value of the pixel, the lighter the color of the pixel in the depth image; the closer a point in the scene is to the image collector, that is, the smaller the depth value of the pixel, the darker the color of the pixel in the depth image.

[0083] 2. Point cloud data

[0084] Point cloud data can record the target scene in the form of points. Each point in the point cloud data contains three-dimensional coordinates. Point cloud data can also include color information (RGB). Color information is usually obtained by acquiring a color image through a camera, and then assigning the color information of the pixel point at the corresponding position to the corresponding point in the point cloud. Point cloud data can also include intensity information. Intensity information is obtained by the echo intensity collected by the laser scanner receiving device. This intensity information is related to the surface material, roughness, incident angle direction of the target, as well as the instrument's emission energy and laser wavelength.

[0085] In addition, the depth image can be calculated into point cloud data after coordinate transformation, and the point cloud data with rules and necessary information can also be inversely calculated into depth image data.

[0086] 3. Two focal lengths f in the camera intrinsic parameters x and f y

[0087] In the pinhole model of an everyday camera, a lens typically has only one focal length. However, the focal length of the lens in the pinhole model differs from the two focal lengths in the intrinsic parameters. The following explains these two focal lengths based on the law of perspective.

[0088] Images captured by a camera follow the laws of linear perspective. This means that the width and height of an object decrease proportionally as the distance from the camera increases. For a rectangular image, the width and height of an object decrease in different proportions depending on the distance from the camera. These proportions are determined by the camera's focal length.

[0089] When the image captured by the camera is a rectangular image, the ratio between the width and height of the object in the image and the actual width and height of the object is:

[0090]

[0091] Among them, f x is the first focal length of the camera, f y is the first focal length of the camera, d is the distance from the object to the camera (i.e., the lens), x is the width of the object in the image, w is the actual width of the object, y is the height of the object in the image, and h is the actual height of the object.

[0092] When the image captured by the camera is a square image, the ratio between the width and height of the object in the image and the actual width and height of the object is:

[0093]

[0094] After introducing the professional terms mentioned in the embodiments of this application, the conventional technologies are introduced below.

[0095] With the development of technology, users can now use AR devices to see the interaction between virtual content and real scenes (such as occlusion and collision between objects). The realization of these interactions depends on the depth image provided by the AR device to the AR application. Currently, depth images can be obtained through the following three conventional technologies.

[0096] Conventional technology 1: Use a depth camera to obtain a depth image.

[0097] Among them, depth cameras include active light depth cameras and passive light cameras.

[0098] When acquiring a depth image, an active light depth camera calculates the depth of each point by analyzing the time it takes for light to reach each point in the image. A passive light camera uses a camera array consisting of multiple ordinary cameras (cameras that capture RGB images) and triangulates the positioning relationship to calculate the depth value of each pixel in the RGB image, thereby generating a depth image.

[0099] However, when the active light depth camera acquires the depth image, due to the presence of a large number of light-absorbing and highly reflective materials in the real scene, the light emitted by the active light depth camera cannot be returned, resulting in the inability to determine the depth of some points in some images. Since the depth of some points in the image cannot be known, the acquired depth image appears incomplete, and the content of the target scene in the depth image is incomplete. When the passive light camera acquires the depth image, if the texture of the photographed object is low (such as a white wall), the passive light camera may not be able to obtain the point representing the feature position (this point can determine the relative position of other points with respect to this point), resulting in the inability to determine the depth of some points in some images. Since the depth of some points in the image cannot be known, the acquired depth image appears incomplete, and the content of the target scene in the depth image is incomplete. In addition, if hardware devices such as depth cameras are added to electronic devices, the cost of the electronic devices will increase, and the power consumption of the electronic devices will increase.

[0100] Conventional technology 2: Use artificial intelligence (AI) algorithms to obtain depth images.

[0101] While using AI algorithms can reduce the cost of electronic devices and produce a continuous, complete depth image without holes, the depth values ​​obtained using AI algorithms lack units and can differ significantly from the actual depth values, resulting in low accuracy. Furthermore, depth noise may be introduced when calculating depth values ​​using AI algorithms, further reducing the accuracy of the calculated depth values.

[0102] Conventional technique three: optimizing the depth image obtained by conventional technique one through RGB image.

[0103] Based on the information of RGB images, assuming that the depth of areas with consistent textures is continuous, bilateral guided filtering is used to optimize the depth image obtained by the depth camera to improve the accuracy of the depth value at the edge. However, in flat scenes with rich textures, using bilateral guided filtering for optimization will result in incorrect depth values. For example, Figure 1 As shown, the depth values ​​of a person's hair and face should be the same. However, due to the different colors of the face and hair, the optimized depth values ​​of the hair and face are different. Similarly, the depth values ​​of leaves at different locations vary. However, because the leaves have the same color, the optimized depth values ​​of the leaves are the same.

[0104] In summary, conventional technical solutions all require the addition of hardware devices (such as ToF cameras) to improve the completeness and accuracy of the content displayed by the acquired depth image. However, adding hardware devices to electronic devices will increase production costs and increase the power consumption of electronic devices.

[0105] To this end, embodiments of the present application provide a method for acquiring a depth image. This method can be applied to the process of acquiring a depth image. Specifically, the method of embodiments of the present application can be applied to multiple scenarios, such as face unlocking scenarios, face payment scenarios, AR scenarios, 3D modeling scenarios, or large aperture scenarios.

[0106] Specifically, in this method, the electronic device can obtain a target depth image of the target scene based on the depth image of the target scene collected and the depth image obtained by processing the RGB image of the target scene. Since the depth value of the depth image of the target scene collected is relatively accurate, the depth image obtained by processing the RGB image of the target scene is relatively complete. Therefore, based on the depth image of the target scene collected and the depth image obtained by processing the RGB image of the target scene, the target depth image obtained is accurate and complete. In addition, the method provided in the embodiment of the present application does not require additional hardware equipment, thereby reducing production costs and reducing the power consumption of electronic devices.

[0107] For example, the electronic device in the embodiments of the present application may be a tablet computer, a mobile phone, a desktop computer, a laptop computer, a handheld computer, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a cellular phone, a personal digital assistant (PDA), an augmented reality (AR) device, a virtual reality (VR) device, an in-vehicle device, or the like. The embodiments of the present application do not impose any special restrictions on the specific form of the electronic device.

[0108] The execution subject of the depth image acquisition method provided in this application may be a depth image acquisition device, which may be Figure 2A The electronic device shown. At the same time, the execution device can also be a central processing unit (CPU) of the electronic device, or a control module for acquiring a depth image in the electronic device. In the embodiment of the present application, the method for acquiring a depth image executed by an electronic device is taken as an example to illustrate the method for acquiring a depth image provided in the embodiment of the present application.

[0109] Please refer to Figure 2A , this application here takes electronic equipment as Figure 2A Taking the mobile phone 200 shown as an example, the electronic device provided by this application is introduced. Figure 2AThe illustrated cell phone 200 is merely one example of an electronic device, and the cell phone 200 may have more or fewer components than shown, may combine two or more components, or may have a different configuration of components. Figure 2A The various components shown in the drawings may be implemented in hardware, software, or a combination of hardware and software, including one or more signal processing and / or application specific integrated circuits.

[0110] like Figure 2A As shown, the mobile phone 200 may include: a processor 210, an external memory interface 220, an internal memory 221, a universal serial bus (USB) interface 230, a charging management module 240, a power management module 241, a battery 242, an antenna 1, an antenna 2, a mobile communication module 250, a wireless communication module 260, an audio module 270, a speaker 270A, a receiver 270B, a microphone 270C, an earphone interface 270D, a sensor module 280, a button 290, a motor 291, an indicator 292, a camera 293, a display screen 294, and a subscriber identification module (SIM) card interface 295, etc.

[0111] The sensor module 280 may include a depth sensor, a pressure sensor, a gyroscope sensor, an air pressure sensor, a magnetic sensor, an acceleration sensor, a distance sensor, a proximity sensor, a fingerprint sensor, a temperature sensor, a touch sensor, an ambient light sensor, a bone conduction sensor, and other sensors. The depth sensor may capture a depth image of a target scene.

[0112] It should be understood that the structure illustrated in this embodiment does not constitute a specific limitation on the mobile phone 200. In other embodiments, the mobile phone 200 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The components shown in the illustrations may be implemented in hardware, software, or a combination of software and hardware.

[0113] The processor 210 may include one or more processing units. For example, the processor 210 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent devices or integrated into one or more processors.

[0114] The controller can be the nerve center and command center of the mobile phone 200. The controller can generate operation control signals according to the instruction operation code and timing signal to complete the control of instruction fetching and execution.

[0115] Processor 210 may also include a memory for storing instructions and data. In some embodiments, the memory in processor 210 is a cache memory. This memory can store instructions or data that have just been used or are being recycled by processor 210. If processor 210 needs to use the same instruction or data again, it can directly access the memory. This avoids duplicate accesses, reduces processor 210 latency, and thus improves system efficiency.

[0116] In some embodiments, the processor 210 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface.

[0117] It is understood that the interface connection relationship between the modules illustrated in this embodiment is merely illustrative and does not limit the structure of the mobile phone 200. In other embodiments, the mobile phone 200 may also adopt different interface connection methods from the above embodiments, or a combination of multiple interface connection methods.

[0118] The charging management module 240 is used to receive charging input from a charger. The charger can be a wireless charger or a wired charger. While charging the battery 242, the charging management module 240 can also power the electronic device through the power management module 241.

[0119] The power management module 241 is used to connect the battery 242, the charging management module 240, and the processor 210. The power management module 241 receives input from the battery 242 and / or the charging management module 240 and provides power to the processor 210, the internal memory 221, the external memory, the display 294, the camera 293, and the wireless communication module 260. In some embodiments, the power management module 241 and the charging management module 240 can also be provided in the same device.

[0120] The wireless communication function of mobile phone 200 can be implemented through antenna 1, antenna 2, mobile communication module 250, wireless communication module 260, modem processor, and baseband processor. In some embodiments, antenna 1 of mobile phone 200 is coupled to mobile communication module 250, and antenna 2 is coupled to wireless communication module 260, so that mobile phone 200 can communicate with the network and other devices through wireless communication technology. For example, in this embodiment of the present application, mobile phone 200 can send the above-mentioned target depth image to other devices through wireless communication technology.

[0121] Antenna 1 and Antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in mobile phone 200 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In other embodiments, the antennas can be used in conjunction with a tuning switch.

[0122] Mobile communication module 250 can provide wireless communication solutions for mobile phone 200, including 2G / 3G / 4G / 5G. Mobile communication module 250 can include at least one filter, switch, power amplifier, low-noise amplifier (LNA), etc. Mobile communication module 250 can receive electromagnetic waves from antenna 1, filter and amplify the received electromagnetic waves, and transmit them to the modem processor for demodulation.

[0123] The mobile communication module 250 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via the antenna 1. In some embodiments, at least some functional modules of the mobile communication module 250 can be set in the processor 210. In some embodiments, at least some functional modules of the mobile communication module 250 can be set in the same device as at least some modules of the processor 210.

[0124] The wireless communication module 260 can provide wireless communication solutions for the mobile phone 200, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR), etc. For example, in the embodiment of the present application, the mobile phone 200 can access the Wi-Fi network through the wireless communication module 260.

[0125] Wireless communication module 260 can be one or more devices that integrate at least one communication processing module. Wireless communication module 260 receives electromagnetic waves via antenna 2, frequency-modulates and filters the electromagnetic wave signals, and transmits the processed signals to processor 210. Wireless communication module 260 can also receive signals to be transmitted from processor 210, frequency-modulate and amplify them, and then convert them into electromagnetic waves for radiation via antenna 2.

[0126] Mobile phone 200 implements display functionality through a GPU, display screen 294, and an application processor. The GPU is a microprocessor for image processing that connects display screen 294 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 210 may include one or more GPUs that execute program instructions to generate or modify display information.

[0127] The display screen 294 is used to display images, videos, etc. The display screen 294 includes a display panel. For example, in the embodiment of the present application, the display screen 294 can be used to display the above-mentioned target depth image, etc.

[0128] The mobile phone 200 can implement a shooting function through an ISP, a camera 293, a video codec, a GPU, a display screen 294, and an application processor. The ISP is used to process data fed back by the camera 293. The camera 293 is used to capture static images or videos. In some embodiments, the mobile phone 200 may include 1 or N cameras 293, where N is a positive integer greater than 1. Exemplarily, the camera 293 may be an active light depth camera (e.g., an infrared structured light camera, an infrared ToF camera, a lidar camera, etc.), or a passive light depth camera (e.g., a binocular vision camera, a multi-camera array, etc.).

[0129] The external memory interface 220 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the mobile phone 200. The external memory card communicates with the processor 210 via the external memory interface 220 to implement data storage functions. For example, files such as music and videos can be stored on the external memory card.

[0130] The internal memory 221 can be used to store computer executable program code, which includes instructions. The processor 210 executes various functional applications and data processing of the mobile phone 200 by running the instructions stored in the internal memory 221. For example, in an embodiment of the present application, the processor 210 can execute instructions stored in the internal memory 221, and the internal memory 221 can include a program storage area and a data storage area.

[0131] The program storage area may store an operating system, at least one application required for a function (such as a sound playback function, an image playback function, etc.), etc. The data storage area may store data created during the use of the mobile phone 200 (such as audio data, a phone book, etc.), etc. In addition, the internal memory 221 may include a high-speed random access memory and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.

[0132] The mobile phone 200 can implement audio functions such as music playback and recording through the audio module 270, the speaker 270A, the receiver 270B, the microphone 270C, the headphone jack 270D, and the application processor.

[0133] The buttons 290 include a power button, a volume button, etc. The button 290 can be a mechanical button. It can also be a touch button. The motor 291 can generate a vibration prompt. The motor 291 can be used for an incoming call vibration prompt, or for touch vibration feedback. The indicator 292 can be an indicator light, which can be used to indicate the charging status, power changes, messages, missed calls, notifications, etc. The SIM card interface 295 is used to connect the SIM card. The SIM card can be connected to and separated from the mobile phone 200 by inserting it into the SIM card interface 295 or pulling it out from the SIM card interface 295. The mobile phone 200 can support 1 or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 295 can support Nano SIM cards, Micro SIM cards, SIM cards, etc.

[0134] although Figure 2A Not shown, the mobile phone 200 may also include a flash, a micro-projection device, a near field communication (NFC) device, etc., which will not be described in detail here.

[0135] After introducing the hardware structure of the electronic device, this application takes the electronic device as a mobile phone 200 as an example to introduce the system architecture of the electronic device provided by this application. The system architecture of the mobile phone 200 can adopt a layered architecture, an event-driven architecture, a micro-kernel architecture, a micro-service architecture, or a cloud architecture. Taking the system as an example, the software structure of the mobile phone 200 is illustrated.

[0136] Figure 2B It is a software structure block diagram of the mobile phone 200 according to an embodiment of the present invention.

[0137] A layered architecture divides software into several layers, each with distinct roles and responsibilities. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers: the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer.

[0138] The application layer can include a series of application packages. Figure 2B As shown, the application package may include applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, short message, etc.

[0139] The application framework layer provides application programming interface (API) and programming framework for the applications in the application layer. The application framework layer includes some predefined functions. Figure 2BAs shown, the application framework layer may include a window manager, a content provider, a view system, a phone manager, a resource manager, a notification manager, and the like.

[0140] The window manager is used to manage window programs. It can obtain the display size, determine whether there is a status bar, lock the screen, take screenshots, etc.

[0141] Content providers are used to store and retrieve data and make it accessible to applications. The data may include videos, images, audio, calls made and received, browsing history and bookmarks, phone books, etc.

[0142] The view system includes visual controls, such as those for displaying text and images. The view system is used to build applications. A display interface can consist of one or more views. For example, a display interface containing a text notification icon might include a view for displaying text and a view for displaying images.

[0143] The phone manager is used to provide communication functions of the mobile phone 200, such as management of call status (including answering, hanging up, etc.).

[0144] The resource manager provides various resources for applications, such as localized strings, icons, images, layout files, video files, and so on.

[0145] The Notification Manager allows applications to display notifications in the status bar. These messages can be displayed briefly and then disappear automatically without user interaction. For example, the Notification Manager is used to notify users of completed downloads and message reminders. The Notification Manager can also display notifications in the top status bar of the system as icons or scrolling text, such as notifications from background applications, or as dialog windows on the screen. Examples include text messages in the status bar, beeps, vibrations on electronic devices, and flashing indicator lights.

[0146] The Android Runtime consists of core libraries and a virtual machine (VM). The Android runtime is responsible for scheduling and management of the Android system. The core library consists of two parts: one for Java-based functions and the other for the Android core library. The application layer and application framework layer run in the VM. The VM executes Java files from the application layer and application framework layer as binary files. The VM is responsible for performing functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.

[0147] The system library can include multiple functional modules, such as surface manager, media library, 3D graphics processing library (such as OpenGL ES), 2D graphics engine (such as SGL), etc.

[0148] The surface manager is used to manage the display subsystem and provide fusion of 2D and 3D layers for multiple applications.

[0149] The media library supports playback and recording of a variety of common audio and video formats, as well as static image files. The media library can support a variety of audio and video encoding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.

[0150] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.

[0151] A 2D graphics engine is a drawing engine for 2D drawings.

[0152] The kernel layer is the layer between hardware and software. The kernel layer includes at least display driver, camera driver, audio driver, and sensor driver.

[0153] The following describes the workflow of the software and hardware of the mobile phone 200 in conjunction with the capture and photo shooting scene.

[0154] When the touch sensor 180K receives a touch operation, the corresponding hardware interrupt is sent to the kernel layer. The kernel layer processes the touch operation into a raw input event (including touch coordinates, touch operation timestamp, and other information). The raw input event is stored in the kernel layer. The application framework layer obtains the raw input event from the kernel layer and identifies the control corresponding to the input event. For example, if the touch operation is a touch single-click operation and the control corresponding to the single-click operation is the control of the camera application icon, the camera application calls the interface of the application framework layer to start the camera application, and then starts the camera driver by calling the kernel layer to capture a still image or video through the camera 193.

[0155] The methods in the following embodiments can all be implemented in an electronic device having the above hardware structure and the above system architecture. In the following embodiments, the methods of the embodiments of the present application are described by taking the electronic device being a mobile phone 200 as an example.

[0156] The present invention provides a method for obtaining a depth image. Figure 3 As shown, the method for acquiring the depth image may include S301-S303.

[0157] S301 : The mobile phone 200 obtains a first depth image and a second depth image.

[0158] Among them, the first depth image (hereinafter referred to as depth image 1) can be a depth image of the target scene collected by the mobile phone 200. The depth image 1 can include N pixels, where N is an integer and N>1.

[0159] In one possible implementation, the mobile phone 200 may collect a depth image 1 of the target scene through an active optical depth camera. For example, Figure 4A is an example of an RGB image, the target scene is Figure 4A The mobile phone 200 can collect the scene as shown by the ToF camera. Figure 4A The depth image 1 of the scene shown (as Figure 4B ).

[0160] It should be noted that the depth value of the depth image 1 collected by the active light depth camera is relatively accurate. However, the depth image 1 collected by the active light depth camera has fewer pixels, which causes the content displayed in the depth image 1 to be incomplete and discontinuous. (For the specific reasons, please refer to the above introduction to conventional technology 1, which will not be repeated here.) For example, Figure 4B As shown, the 401 and 402 areas do not have corresponding content in the target scene, resulting in Figure 4B The image of the target scene in the shown depth image is incomplete.

[0161] In the embodiment of the present application, the second depth image (hereinafter referred to as depth image 2) may be a depth image obtained by processing the RGB image of the target scene. The depth image 2 may include M pixels, where M is an integer, 1 <N<M。

[0162] In one possible implementation, the mobile phone 200 may first shoot the target scene with an ordinary camera to obtain an RGB image of the target scene (e.g. Figure 4A ). Afterwards, the mobile phone 200 inputs the RGB image into the neural network model to obtain the depth image 2. The neural network model can be constructed by a preset algorithm. For example, the mobile phone 200 can use the deeplabv3+ algorithm to Figure 4A The RGB image shown is processed to predict the depth image 2 (such as Figure 4C ).

[0163] In another possible implementation, mobile phone 200 can first capture the target scene using a standard camera to obtain an RGB image of the target scene. Then, mobile phone 200 can perform semantic segmentation on the RGB image. Finally, mobile phone 200 inputs the segmented RGB image into a neural network model to obtain a second depth image.

[0164] It should be noted that, since there may be two objects (also called objects) connected or partially overlapping in the RGB image, the electronic device cannot distinguish which object the edge where the two objects touch belongs to. Therefore, in the depth image calculated by the mobile phone 200 using the artificial intelligence AI algorithm, the depth value at the edge where the two objects touch is inaccurate. Image semantic segmentation can effectively distinguish two different objects. Therefore, after the mobile phone 200 performs image semantic segmentation on the RGB image, the edge contours of the two overlapping objects in the RGB image are clearer. In this way, the second depth image obtained by the mobile phone 200 using the artificial intelligence AI algorithm to calculate the segmented RGB image can have a more accurate depth value. For example, Figure 4D It is an object image recognized by the mobile phone 200 when the mobile phone 200 does not perform image semantic segmentation on the RGB image. Figure 4E The object image is recognized by the mobile phone 200 when the mobile phone 200 performs image semantic segmentation on the RGB image.

[0165] It should be noted that the RGB image captured by an ordinary camera is relatively complete, that is, each object and each feature in the target scene will be displayed in the RGB image. In this way, the depth image obtained by the mobile phone 200 processing the RGB image of the target scene is also relatively complete, that is, the content displayed in the depth image 2 will not have problems such as incomplete content and discontinuous content. However, since the depth value of the depth image obtained based on the neural network model usually does not include a length unit (such as millimeters, centimeters, etc.). Therefore, the depth value of the depth image 2 is not accurate. For example, assume that the depth image A is the depth image obtained by the mobile phone 200 processing the RGB image of the target scene. The depth value of pixel A in the depth image A is 0.7. However, the mobile phone 200 cannot obtain the length unit of the depth value of pixel A, that is, the mobile phone 200 cannot determine whether the depth value of pixel A is 0.7 mm, 0.7 cm, or 0.7 m, etc.

[0166] In an embodiment of the present application, the depth image may include pixel information of the target scene, and the pixel information includes two-dimensional features and depth values ​​of each pixel. Among them, the two-dimensional features may include the two-dimensional coordinates of the pixel points. For example, depth image A is a depth image collected by mobile phone 200. The two-dimensional coordinates of pixel A in depth image A are (1,5), and the depth value of pixel A is 5000 mm. For another example, depth image B is a depth image obtained by processing the RGB image of the target scene. The two-dimensional coordinates of pixel B in image B are (3,6), and the depth value of pixel B is 7.

[0167] In some embodiments, the pixel information may further include a color value for each pixel. The color value of a pixel is determined by an R value, a G value, and a B value. The R value, the G value, and the B value all have a value range of [0, 255]. For example, if the R value, the G value, and the B value of pixel A are 0, 0, and 0, the color value of the pixel is #000000, and the color of pixel A is black. For another example, if the R value, the G value, and the B value of pixel B are 255, 255, and 255, the color value of the pixel is #FFFFFF, and the color of pixel A is white.

[0168] S302 : The mobile phone 200 obtains first point cloud data of the target scene according to the first depth image and the second depth image.

[0169] The first point cloud data (hereinafter referred to as point cloud data 1) may include two-dimensional features and depth values ​​of M first feature points, and the M first feature points correspond one-to-one to M pixel points. For example, Figure 5 For a portrait's depth image and 3D point cloud, there's a one-to-one correspondence between pixels in the depth image and feature points in the 3D point cloud data. For example, Table 1 shows the correspondence between three feature points in the point cloud data and three pixels in the depth image. The three feature points in the point cloud data are: Feature Point A, Feature Point B, and Feature Point C. The three pixels in the depth image are: Pixel A, Pixel B, and Pixel C.

[0170] Table 1

[0171] Feature Points Pixels Feature point A Pixel A Feature point B Pixel B Feature point C Pixel C

[0172] That is, feature point A corresponds to pixel point A, feature point B corresponds to pixel point B, and feature point C corresponds to pixel point C.

[0173] In some embodiments, mobile phone 200 can adjust the pixel information of pixels in depth image 2 based on the pixel information of pixels in depth image 1, and ultimately obtain the first point cloud data of the target scene. For example, mobile phone 200 can adjust the depth values ​​of M pixels in depth image 2 based on the depth values ​​of N pixels in depth image 1, and obtain the adjusted depth values ​​of the M pixels. Subsequently, mobile phone 200 can obtain the depth values ​​of the M first feature points in point cloud data 1 based on the adjusted depth values ​​of the M pixels.

[0174] It should be noted that mobile phone 200 merely adjusts the pixel information of the pixels in depth image 2 based on the pixel information of the pixels in depth image 1 to obtain point cloud data 1 of the target scene. That is, for feature point 1 (any feature point in point cloud data 1) in point cloud data 1, the two-dimensional features of feature point 1 are associated with the two-dimensional features of pixel A (the pixel corresponding to feature point 1 in depth image 1) and pixel B (the pixel corresponding to feature point 1 in depth image 2), and the depth value of feature point 1 is associated with the depth values ​​of pixel A (the pixel corresponding to feature point 1 in depth image 1) and pixel B (the pixel corresponding to feature point 1 in depth image 2). Therefore, the two-dimensional features of feature point 1 may not be the same as the two-dimensional features of pixel A and pixel B, and the depth value of feature point 1 may not be the same as the depth values ​​of pixel A and pixel B.

[0175] For example, the 2D feature of feature point A in point cloud data 1 is (3, 7), with a depth of 5 meters. The 2D feature of the pixel corresponding to feature point A in depth image 1 is (2, 5), with a depth of 4.8 meters. The 2D feature of the pixel corresponding to feature point A in depth image 2 is (3, 6), with a depth of 10.

[0176] It is understandable that since the depth values ​​of the N pixels of depth image 1 are relatively accurate, and the number of pixels of depth image 2 is larger (including M pixels, M>N), the first point cloud data obtained by mobile phone 200 based on depth image 1 and depth image 2 can record the complete and accurate target scene.

[0177] S303 : The mobile phone 200 obtains a target depth image of the target scene according to the first point cloud data.

[0178] For example, Figure 5 As shown, the mobile phone 200 can convert each feature point in the point cloud data into a pixel point in the depth image, and finally obtain a depth image of the portrait.

[0179] It is understandable that since the M first feature points in the point cloud data 1 can record the target scene completely and accurately, the electronic device can obtain a target depth image of the target scene based on the first point cloud data, and the target depth image is complete and accurate.

[0180] In one possible implementation, the mobile phone 200 may process the point cloud data 1 using Formula 1 to obtain a target depth image of the target scene.

[0181]

[0182] Among them, (x i ,y i , z i) represents the three-dimensional features of the feature point in the point cloud data corresponding to the i-th pixel in the depth image, (u i ,v i ) represents the two-dimensional feature of the i-th pixel in the depth image, d i represents the depth value of the i-th pixel in the depth image, (u0, v0) represents the two-dimensional feature of the center point of the depth image, and f x Represents the first focal length of the depth image, f y Indicates the second focal length of the depth image.

[0183] It should be noted that the center point of the depth image refers to the point at the center position of the depth image. For example, Figure 6 is a schematic diagram of a depth image A, where the four vertices of the depth image A are A, B, C, and D. The center point of the depth image A is the intersection E between AD and BC, that is, point E is the center point of the depth image A.

[0184] For example, the following uses the correspondence between feature point A in point cloud data 1 and the first pixel in the target depth image as an example to illustrate Formula 1. For example, the three-dimensional feature of feature point A in point cloud data 1 is (2, 4, 8), the two-dimensional feature of the center point in the target depth image is (8, 6), the first focal length of the target depth image is 4, and the second focal length of the target depth image is 2. Then the two-dimensional feature (u1, v1) and depth value d1 of the first pixel in the target depth image are:

[0185]

[0186]

[0187] d1=8.

[0188] That is to say, the two-dimensional feature of the first pixel in the target depth image is (9, 7) and the depth value is 8.

[0189] Based on the above technical solution, an electronic device can obtain a first depth image and a second depth image. Since the first depth image is a depth image of the target scene captured by the electronic device and includes N pixels, the depth values ​​of the N pixels in the first depth image are relatively accurate. Furthermore, since the second depth image is a depth image obtained by processing an RGB image of the target scene and includes M pixels, the second depth image has a greater number of pixels and is relatively complete. Subsequently, the electronic device can obtain first point cloud data of the target scene based on the first and second depth images. The first point cloud data includes two-dimensional features and depth values ​​of M first feature points, with the M first feature points corresponding one-to-one to the M pixels. Since the second depth image has a larger number of pixels, the M first feature points in the first point cloud data can fully record the target scene. Furthermore, since the depth values ​​of the pixels in the first depth image are relatively accurate, and the depth values ​​of the M first feature points are derived based on the depth values ​​of the M pixels and the depth values ​​of the N pixels, the depth values ​​of the M first feature points in the first point cloud data are relatively accurate. Furthermore, since the M first feature points in the first point cloud data can record a complete and accurate target scene, the electronic device can obtain a target depth image of the target scene based on the first point cloud data, and the target depth image is complete and accurate. Therefore, compared with conventional technical solutions, the embodiments of the present application can solve the problem of incomplete depth images and incomplete content of the target scene displayed by the depth image.

[0190] It should be noted that the above S302 is only a general description of the process of the mobile phone 200 obtaining the point cloud data 1. For a detailed description of the mobile phone 200 obtaining the first point cloud data of the target scene based on the depth image 1 and the depth image 2, please refer to S701-S703.

[0191] The following describes a method (i.e., S302) for the mobile phone 200 to obtain point cloud data 1 in an embodiment of the present application.

[0192] In some embodiments, the mobile phone 200 may execute S701-S703 to obtain point cloud data 1. Specifically, Figure 7A As shown, S302 includes S701-S703.

[0193] S701. The mobile phone 200 obtains second point cloud data according to the first depth image.

[0194] The second point cloud data (also referred to as point cloud data 2) includes two-dimensional features and depth values ​​of N feature points, and the N feature points correspond to N pixel points one by one. For a detailed description of the correspondence between feature points and pixel points, please refer to Table 1 in S302. Figure 5 The description is not repeated here.

[0195] In one possible implementation, mobile phone 200 obtains the depth values ​​of N feature points of point cloud data 2 based on the depth values ​​of N pixels in depth image 1. Mobile phone 200 obtains the two-dimensional features of N feature points of point cloud data 2 based on the two-dimensional features of N pixels in depth image 1, the two-dimensional features of the center point of depth image 1, the first focal length of depth image 1, and the second focal length.

[0196] In one possible design, the mobile phone 200 processes the depth image using Formula 2 to obtain point cloud data.

[0197]

[0198] Among them, (x i ,y i , z i ) represents the three-dimensional features of the feature point in the point cloud data corresponding to the i-th pixel in the depth image, (u i ,v i ) represents the two-dimensional feature of the i-th pixel in the depth image, d i represents the depth value of the i-th pixel in the depth image, (u0, v0) represents the two-dimensional feature of the center point of the depth image, and f x Represents the first focal length of the depth image, f y Represents the second focal length of the depth image. For details, please refer to the description of the parameters in Formula 1 in S303, which will not be repeated here.

[0199] For example, the following uses the first pixel A in the depth image 1 and the feature point A in the point cloud data 2 as an example to illustrate Formula 2. For example, the depth value of pixel A is 8, the two-dimensional feature of pixel A is (4, 8), the two-dimensional feature of the center point of the depth image 1 is (2, 6), the first focal length of the depth image 1 is 2, and the second focal length of the depth image 2 is 4. Then the three-dimensional feature of feature point A (x i ,y i ,z i )for:

[0200]

[0201]

[0202] z i =8.

[0203] That is, the three-dimensional feature of feature point A is (8, 4, 8).

[0204] S702 : The mobile phone 200 obtains third point cloud data according to the second depth image.

[0205] The third point cloud data (also referred to as point cloud data 3) includes two-dimensional features and depth values ​​of M second feature points, and the M second feature points correspond to M pixels one by one. For a detailed description of the correspondence between feature points and pixels, please refer to Table 1 in S302. Figure 5 The description is not repeated here.

[0206] In the embodiment of the present application, the mobile phone 200 can obtain the depth values ​​of the M feature points in the point cloud data 3 based on the depth values ​​of the M pixels in the depth image 2. Specifically, the mobile phone 200 can scale the depth values ​​of the pixels in the depth image 2 based on the depth values ​​of the pixels in the depth image 1 to obtain a third depth image (also referred to as depth image 3). Thereafter, the mobile phone 200 obtains point cloud data 3 based on the depth image 3.

[0207] It should be noted that, for the process in which the mobile phone 200 obtains the point cloud data 3 based on the depth image 3 , reference may be made to the description in S701 , which will not be repeated here.

[0208] It is understandable that, since the depth values ​​of the pixels in depth image 1 are more accurate, the depth values ​​of depth image 3 obtained by mobile phone 200 are more accurate than those of depth image 2. That is, the difference between the depth value of a pixel in depth image 3 and the depth value of the corresponding pixel in the first depth image is smaller than the difference between the depth value of the corresponding pixel in depth image 2 and the depth value of the corresponding pixel in depth image 1.

[0209] The following describes in detail the process in which the mobile phone 200 scales the depth values ​​of the pixels in the depth image 2 according to the depth values ​​of the pixels in the depth image 1 to obtain the depth image 3 .

[0210] In some embodiments, mobile phone 200 may determine a first median value and a second median value. The first median value is the median value of the depth values ​​of N pixels, and the second median value is the median value of the depth values ​​of M pixels. Mobile phone 200 may then scale the depth values ​​of the M pixels in depth image 2 according to a ratio of the first median value to the second median value to obtain depth image 3.

[0211] It should be noted that the median value is to arrange a group of data in order from small to large or from large to small. If the number of data in this group is an odd number, the middle data is the median value. For example, depth image 1 includes the depth values ​​of 3 pixels, which are 1 meter, 5 meters, and 8 meters respectively, then the first median value is 5 meters. Depth image 2 includes the depth values ​​of 5 pixels, which are 2, 6, 10, 12, and 20 respectively, then the second median value is 10. Since the ratio between the first median value and the second median value is 0.5. Therefore, mobile phone 200 can reduce the depth values ​​of the 5 pixels in depth image 2 to 1 / 2 of the original. That is, the depth values ​​of the 5 pixels of the reduced depth image 2 are 1 meter, 3 meters, 5 meters, 6 meters, and 10 meters respectively. Therefore, the depth values ​​of the 5 pixels of depth image 3 are 1 meter, 3 meters, 5 meters, 6 meters, and 10 meters respectively.

[0212] If the number of data in this group is an even number, the average of the two middle numbers is the median. For example, depth image 1 includes the depth values ​​of 4 pixels, which are 1 meter, 4 meters, 6 meters, and 8 meters respectively, so the first median is 5 meters. Depth image 2 includes the depth values ​​of 5 pixels, which are 2, 6, 10, 12, and 20 respectively, so the second median is 10. Since the ratio between the first median and the second median is 0.5. Therefore, mobile phone 200 can reduce the depth values ​​of the 5 pixels in depth image 2 to 1 / 2 of the original. That is, the depth values ​​of the 5 pixels of the reduced depth image 2 are 1 meter, 3 meters, 5 meters, 6 meters, and 10 meters respectively. Therefore, the depth values ​​of the 5 pixels of depth image 3 are 1 meter, 3 meters, 5 meters, 6 meters, and 10 meters respectively.

[0213] It can be understood that since the median value can represent the central tendency of a set of data, and the first median value is the median value of the depth values ​​of N pixels, and the second median value is the median value of the depth values ​​of M pixels. Therefore, the mobile phone 200 can scale the depth values ​​of the M pixels in the depth image 3 according to the ratio of the first median value to the second median value to obtain the depth image 3. In other words, the mobile phone 200 adjusts the depth values ​​of the pixels in the depth image 2 according to the depth values ​​of the pixels in the depth image 1 to obtain the depth values ​​of the depth image 3. And since the depth values ​​of the pixels in the depth image 1 are more accurate. Therefore, compared with the depth values ​​of the pixels in the depth image, the depth value of the depth image 3 obtained by the mobile phone 200 is more accurate.

[0214] In other embodiments, mobile phone 200 may also determine a first average value and a second average value. The first average value is the average of the depth values ​​of N pixels, and the second average value is the average of the depth values ​​of M pixels. Mobile phone 200 may then scale the depth values ​​of the M pixels in depth image 2 according to the ratio of the first average value to the second average value to obtain depth image 3.

[0215] It should be noted that since the median value can represent the central tendency of a set of data, the ratio between the first median value and the second median value is more representative, that is, it can represent the ratio between the depth values ​​of the pixels in depth image 1 and their corresponding pixels in depth image 2. Therefore, under normal circumstances, mobile phone 200 scales the depth values ​​of M pixels in depth image 2 by the ratio of the first median value to the second median value to obtain depth image 3.

[0216] In addition, since the depth image 2 is a depth image obtained by processing the RGB image of the target scene, the depth value of each pixel in the depth image 2 has no length unit (such as meters, centimeters, etc.). Therefore, the mobile phone 200 can scale the depth value of the pixel in the depth image 2 according to the depth value of the pixel in the depth image 1 to obtain the depth value of each pixel in the depth image and its length unit. In this way, compared with the point cloud data obtained by the mobile phone 200 based on the depth image 2, the point cloud data obtained by the mobile phone 200 based on the depth image 3 can record the target scene more accurately.

[0217] It should be noted that the embodiment of the present application does not limit the order in which the mobile phone 200 executes S701 and S702. In other words, the mobile phone 200 may execute S701 first and then S702; or the mobile phone 200 may execute S702 first and then S701.

[0218] S703 : The mobile phone 200 obtains the first point cloud data according to the second point cloud data and the third point cloud data.

[0219] It is understandable that since point cloud data 2 can accurately record the target scene and point cloud data 3 can completely record the target scene, mobile phone 200 can obtain point cloud data 1 that can accurately and completely record the target scene based on point cloud data 2 and point cloud data 3.

[0220] In other embodiments, the mobile phone 200 may execute S7001-S7003 to obtain point cloud data 1. Specifically, Figure 7B As shown, S703 includes S7001-S7003.

[0221] S7001. The mobile phone 200 determines N target points corresponding to the N feature points of the second point cloud data from the M feature points of the third point cloud data.

[0222] The points in the target scene corresponding to the N target points in the point cloud data 3 are the same as the points in the target scene corresponding to the N feature points in the point cloud data 2.

[0223] For example, Figure 8As shown in (a), point cloud data 3 includes 36 feature points, of which 36 feature points include 10 target points (i.e. Figure 8 The black dots in (a). Figure 8 As shown in (b), point cloud data 2 includes 10 feature points. Figure 8 The 10 target points in (a) and Figure 8 The 10 feature points in (b) correspond one to one. For example, if feature point A in point cloud data 2 corresponds to feature point a in point cloud data 1, then feature point A is the target point.

[0224] In some embodiments, the two-dimensional features of the N pixels in the depth image 2 corresponding to the N target points in the point cloud data 3 are the same as the two-dimensional features of the N pixels in the depth image 1 corresponding to the N feature points in the point cloud data 2. In other words, the mobile phone 200 can obtain the target points in the point cloud data 2 corresponding to the feature points in the point cloud data 1 based on the two-dimensional features of the pixels in the depth images 1 and 3.

[0225] In other embodiments, the mobile phone 200 may also obtain multiple pixel groups by analyzing the depth image 1 and the depth image 3. Each pixel group includes two pixel points, which are pixel points with the same object features represented in the depth image 1 and the depth image 3. For example, pixel A in the depth image 1 represents the eye of a person in the image, and pixel B in the depth image 3 also represents the eye of a person in the image, then pixel A and pixel B can form a pixel group. Then the feature point b corresponding to pixel B in the point cloud data 3 corresponds to the feature point a corresponding to pixel A in the point cloud data 1, and feature point b is the target point.

[0226] S7002. The mobile phone 200 may adjust the three-dimensional features of the N target points according to the three-dimensional features of the N feature points, and adjust the three-dimensional features of the MN traction points in the second point cloud data.

[0227] In the embodiment of the present application, the three-dimensional feature includes a two-dimensional feature and a depth value. For example, the two-dimensional feature of feature point A is (1, 5) and the depth value is 7. Then the three-dimensional feature of feature point A is (1, 5, 7). MN traction points are points other than the N target points in the third point cloud data. For example, Figure 9 The black points are target points, and the white points are traction points. For example, feature points B, C, and D are all traction points.

[0228] In some embodiments, the mobile phone 200 can adjust the three-dimensional features of the N target points and the three-dimensional features of the MN traction points using a grid-optimized traction fusion method. Specifically, the mobile phone 200 can adjust the three-dimensional features of the N target points and the three-dimensional features of the MN traction points using Formula 3.

[0229]

[0230] Among them, v i ′ represents the three-dimensional feature of the i-th target point among the N target points after adjustment, Represents the three-dimensional features of the feature point corresponding to the i-th target point in the second point cloud data, v′ i,2 represents the three-dimensional features of the i-th traction point among the MN traction points after adjustment, represents the desired three-dimensional feature of the i-th traction point, λ is the coefficient, 0<λ<1.

[0231] It should be noted that the mobile phone 200 can be used according to E data and E reg Get E.

[0232] in, That is to say, E data It is determined by the difference between the adjusted three-dimensional features of the target point and the three-dimensional features of the feature points in the point cloud data 1 corresponding to the target point. data The smaller it is, the closer the three-dimensional features of the adjusted target point are to the three-dimensional features of the feature points in the point cloud data 1 corresponding to the target point, and the more accurate the three-dimensional features of the adjusted target point are.

[0233] E reg It represents the difference between the adjusted 3D feature of the traction point and the expected 3D feature of the traction point. reg The smaller the value, the closer the three-dimensional features of the adjusted traction point are to the three-dimensional features of the expected traction point, and the more accurate the three-dimensional features of the adjusted traction point are. Specifically, the mobile phone 200 can determine E based on the similarity between the triangle formed by the traction point and other feature points (target point or traction point) in the point cloud data 2 (i.e., the triangle formed by the expected traction point) and the triangle formed by the adjusted traction point and the target point (or traction point). reg .

[0234] For example, Figure 10 As shown in (a), V2 is the traction point. ΔV0V2V1 is a triangle formed by the traction point V2, feature point V0, and feature point V1 in point cloud data 2. Figure 10 As shown in (b), V2' is the adjusted traction point. ΔV0′V2′V1′ is a triangle formed by the adjusted traction point V2', feature point V0', and feature point V1'. When the similarity between ΔV0′V2′V1′ and ΔV0V2V1 is greater, the 3D features of V2' are closer to the 3D features of V2, and E reg The smaller.

[0235] In the embodiment of the present application, when E is minimum, the adjusted N target points and MN traction points can accurately record the target scene.

[0236] In other embodiments, the mobile phone 200 may also adjust the N target points and MN traction points based on a neural network model.

[0237] S7003. The mobile phone 200 may obtain first point cloud data based on the adjusted three-dimensional features of the N target points and the three-dimensional features of the MN traction points.

[0238] The point cloud data 1 includes: when E is minimum, the adjusted three-dimensional features of N target points and the three-dimensional features of MN traction points. For example, Figure 11 The 3D point cloud shown in (a) is the point cloud corresponding to point cloud data 2 (10 target points and 26 traction points). Figure 11 The 3D point cloud image shown in (b) is the point cloud image corresponding to point cloud data 1 obtained by mobile phone 200 based on the adjusted 3D features of the 10 target points and the 3D features of the 26 traction points. For example, the 3D features of target point A' are the adjusted 3D features of target point A.

[0239] It is understandable that since point cloud data 2 is derived from depth image 1, and the depth values ​​of depth image 1 are relatively accurate, the 3D features of point cloud data 1 are also relatively accurate. Therefore, mobile phone 200 can adjust the 3D features of N target points based on the 3D features of N feature points, and adjust the 3D features of MN traction points in point cloud data 2 to obtain point cloud data 1. Since the 3D features of the N feature points in point cloud data 2 are relatively accurate, the point cloud data 1 obtained by mobile phone 200 after adjusting the 3D features of the N feature points is also relatively accurate.

[0240] Based on the above technical solution, since the depth values ​​of the pixels in the depth image 1 are relatively accurate, and the N feature points included in the point cloud data 2 correspond one-to-one to the N pixels in the depth image 2. Therefore, the N feature points in the point cloud data 2 can accurately record the target scene. Since the number of pixels in the depth image 2 is relatively large, and the M feature points included in the point cloud data 3 correspond one-to-one to the M pixels in the depth image 2. Therefore, the M feature points in the point cloud data 3 can completely record the target scene. Furthermore, since the point cloud data 2 can accurately record the target scene, the point cloud data 3 can completely record the target scene. Therefore, the mobile phone 200 can obtain the point cloud data 1 that can accurately and completely record the target scene based on the point cloud data 2 and the point cloud data 3.

[0241] For example, combined Figure 4A The target scene in the RGB image shown, Figure 12(a) is the 3D space point cloud 1 (i.e., the 3D space point cloud corresponding to point cloud data 2) converted from the depth image (i.e., depth image 1) captured by the ToF camera. Figure 12 (b) is the 3D space point cloud 2 (i.e., the 3D space point cloud corresponding to the point cloud data 3) converted from the depth image (i.e., depth image 2) obtained by processing the RGB image of the target scene by the mobile phone 200. Figure 12 (c) is the 3D space point cloud 3 (i.e., the 3D space point cloud corresponding to the point cloud data 1) converted from the depth image (i.e., the target depth image) obtained by the mobile phone 200 according to the method provided in the embodiment of the present application.

[0242] The above mainly introduces the solution provided by the embodiment of the present application from the perspective of an electronic device. It is understandable that, in order to realize the above functions, the electronic device includes a hardware structure and / or software module corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the steps of a method for acquiring a depth image of each example described in the embodiment disclosed in this application, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or electronic device software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0243] In the embodiment of the present application, the depth image acquisition device can be divided into functional modules or functional units according to the above method example. For example, each functional module or functional unit can be divided according to each function, or two or more functions can be integrated into one processing module. The above-mentioned integrated module can be implemented in the form of hardware or in the form of software functional modules or functional units. Among them, the division of modules or units in the embodiment of the present application is schematic and is only a logical functional division. There may be other division methods in actual implementation.

[0244] Please refer to Figure 13 , which shows a schematic diagram of a depth image acquisition device provided by an embodiment of the present application. The depth image acquisition device can be a functional module in the above-mentioned electronic device (such as mobile phone 200) for implementing the method of the embodiment of the present application. Figure 13 As shown, the depth image acquisition device may include: an AI information prediction module 1301, a scale scaling module 1302, a 2D to 3D projection module 1303, a point cloud fusion module 1304, and a 3D to 2D projection module 1305.

[0245] Among them, the AI information prediction module 1301 is used to obtain a first depth image and a second depth image. The first depth image is the depth image of the target scene collected by the electronic device. The first depth image includes N pixel points. The second depth image is the depth image obtained by processing the RGB image of the target scene. The second depth image includes M pixel points. Both M and N are integers, and 1 < N < M. The depth image includes the pixel information of the target scene. The pixel information includes the two-dimensional features and depth values of each pixel point.

[0246] The 2D to 3D projection module 1303 is used to obtain the first point cloud data of the target scene according to the first depth image and the second depth image. Among them, the first point cloud data includes the two-dimensional features and depth values of M first feature points. The M first feature points correspond to the M pixel points one by one. The depth values of the M first feature points are obtained according to the depth values of the M pixel points and the depth values of the N pixel points.

[0247] The 3D to 2D projection module 1305 is used to obtain the target depth image of the target scene according to the first point cloud data.

[0248] Optionally, the 2D to 3D projection module 1303 is specifically used to obtain second point cloud data according to the first depth image. Among them, the second point cloud data includes the two-dimensional features and depth values of N feature points. The N feature points correspond to the N pixel points one by one. The depth values of the N feature points are obtained according to the depth values of the N pixel points. Obtain third point cloud data according to the second depth image. Among them, the third point cloud data includes the two-dimensional features and depth values of M second feature points. The M second feature points correspond to the M pixel points one by one. The depth values of the M second feature points are obtained according to the depth values of the M pixel points. The first point cloud data is obtained according to the second point cloud data and the third point cloud data.

[0249] Optionally, the scale scaling module 1302 is used to scale the depth values of the pixel points in the second depth image according to the depth values of the pixel points in the first depth image to obtain a third depth image. Among them, the difference between the depth value of the pixel point in the third depth image and the depth value of the corresponding pixel point in the first depth image is smaller than the difference between the depth value of the corresponding pixel point in the second depth image and the depth value of the corresponding pixel point in the first depth image. The 2D to 3D projection module 1303 is specifically used to obtain third point cloud data according to the third depth image.

[0250] Optionally, the scale scaling module 1302 is specifically used to determine a first median value and a second median value. Among them, the first median value is the median value of the depth values of N pixel points. The second median value is the median value of the depth values of M pixel points. Scale the depth values of the M pixel points in the second depth image according to the ratio of the first median value to the second median value to obtain a third depth image.

[0251] Optionally, the 2D to 3D projection module 1303 is specifically configured to obtain point cloud data according to the depth image using the following formula:

[0252] z i =d i ;

[0253] Among them, (x i ,y i ,z i ) represents the three-dimensional features of the feature point in the point cloud data corresponding to the i-th pixel in the depth image, (u i ,v i ) represents the two-dimensional feature of the i-th pixel in the depth image, d i represents the depth value of the i-th pixel in the depth image, (u0, v0) represents the two-dimensional feature of the center point of the depth image, and f x and represents the first focal length of the depth image, f y Indicates the second focal length of the depth image.

[0254] Optionally, the point cloud fusion module 1304 is used to determine N target points corresponding to the N feature points of the second point cloud data from the M feature points of the third point cloud data; adjust the three-dimensional features of the N target points according to the three-dimensional features of the N feature points, and adjust the three-dimensional features of MN traction points in the second point cloud data, where the MN traction points are points other than the N target points in the third point cloud data, and the three-dimensional features include two-dimensional features and depth values; obtain the first point cloud data based on the adjusted three-dimensional features of the N target points and the three-dimensional features of the MN traction points.

[0255] Optionally, the point cloud fusion module 1304 is specifically configured to obtain the adjusted three-dimensional features of the N target points and the three-dimensional features of the MN traction points using the following formula:

[0256]

[0257] Among them, v i ′ represents the three-dimensional feature of the i-th target point among the N target points after adjustment, Represents the three-dimensional features of the feature point corresponding to the i-th target point in the second point cloud data, v′ i,2 represents the three-dimensional features of the i-th traction point among the MN traction points after adjustment, represents the desired three-dimensional feature of the i-th traction point, λ is the coefficient, 0<λ<1.

[0258] Optionally, the 3D to 2D projection module 1305 is configured to obtain a depth image based on the point cloud data using the following formula:

[0259] d i =z i .

[0260] Optionally, the AI ​​information prediction module 1301 is further used to perform image semantic segmentation on the RGB image, and use an artificial intelligence AI algorithm to calculate the segmented RGB image to obtain a second depth image.

[0261] For example, combined Figure 4A 、 Figure 4B ,as well as Figure 4C ,like Figure 14 As shown, the AI ​​information prediction module 1301 can predict the RGB image (i.e. Figure 4A ) to obtain depth image 2 (i.e. Figure 4C ). Afterwards, the point cloud fusion module 1304 can perform the depth image 2 and the depth image 1 (ie Figure 4B ) to obtain the target depth image.

[0262] Some other embodiments of the present application provide an electronic device (such as Figure 2A The mobile phone 200 shown in the figure) has multiple preset applications installed in the electronic device. The electronic device may include: a memory and one or more processors. The memory and the processor are coupled. The electronic device may also include a camera. Alternatively, the electronic device may be connected to an external camera. The memory is used to store computer program code, which includes computer instructions. When the processor executes the computer instructions, the electronic device may execute the various functions or steps executed by the mobile phone in the above method embodiment. The structure of the electronic device can refer to Figure 2A The structure of the mobile phone 200 is shown.

[0263] The present application also provides a chip system. Figure 15 As shown, the chip system includes at least one processor 1501 and at least one interface circuit 1502. The processor 1501 and the interface circuit 1502 can be interconnected via a line. For example, the interface circuit 1502 can be used to receive signals from other devices (such as a memory of an electronic device). For another example, the interface circuit 1502 can be used to send signals to other devices (such as the processor 1501). Exemplarily, the interface circuit 1502 can read instructions stored in the memory and send the instructions to the processor 1501. When the instructions are executed by the processor 1501, the electronic device (such as Figure 2A The mobile phone 200 shown in FIG2 executes each step in the above embodiment. Of course, the chip system may also include other discrete components, which are not specifically limited in the embodiment of the present application.

[0264] The embodiment of the present application further provides a computer storage medium, which includes computer instructions. When the computer instructions are in the above electronic device (such as Figure 2A When the method is executed on the mobile phone 200 shown in the figure, the electronic device executes the functions or steps executed by the mobile phone in the above method embodiment.

[0265] The embodiment of the present application further provides a computer program product, which, when executed on a computer, enables the computer to execute the functions or steps executed by the mobile phone in the above method embodiment.

[0266] Through the description of the above implementation methods, technical personnel in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0267] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0268] The units described as separate components may or may not be physically separate, and the components shown as units may be one physical unit or multiple physical units, that is, they may be located in one place or distributed in multiple places. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0269] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0270] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a device (which can be a single-chip microcomputer, chip, etc.) or a processor (processor) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0271] The above content is only a specific embodiment of this application, but the scope of protection of this application is not limited to this. Any changes or replacements within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A method for acquiring a depth image, characterized in that: Applied to an electronic device, the method includes: Obtain a first depth image and a second depth image. The first depth image is a depth image of a target scene collected by the electronic device. The first depth image includes N pixel points. The second depth image is a depth image obtained by processing the RGB image of the target scene. The second depth image includes M pixel points. Both M and N are integers, and 1 < N < M. The depth image includes pixel information of the target scene. The pixel information includes two-dimensional features and depth values of each pixel point. Obtain second point cloud data according to the first depth image. Among them, the second point cloud data includes two-dimensional features and depth values of N feature points. The N feature points correspond to the N pixel points one by one. The depth values of the N feature points are obtained according to the depth values of the N pixel points. Scale the depth values of the pixel points in the second depth image according to the depth values of the pixel points in the first depth image to obtain a third depth image. Among them, the difference between the depth value of a pixel point in the third depth image and the depth value of the corresponding pixel point in the first depth image is less than the difference between the depth value of the corresponding pixel point in the second depth image and the depth value of the corresponding pixel point in the first depth image. Obtain third point cloud data according to the third depth image. Among them, the third point cloud data includes two-dimensional features and depth values of M second feature points. The M second feature points correspond to the M pixel points one by one. The depth values of the M second feature points are obtained according to the depth values of the M pixel points. Obtain first point cloud data of the target scene according to the second point cloud data and the third point cloud data. Among them, the first point cloud data includes two-dimensional features and depth values of M first feature points. The M first feature points correspond to the M pixel points one by one. The depth values of the M first feature points are obtained according to the depth values of the M pixel points and the depth values of the N pixel points. Obtain a target depth image of the target scene according to the first point cloud data.

2. The method according to claim 1, characterized in that The step of scaling the depth values of the pixel points in the second depth image according to the depth values of the pixel points in the first depth image to obtain a third depth image includes: Determine a first median value and a second median value. Among them, the first median value is the median value of the depth values of the N pixel points. The second median value is the median value of the depth values of the M pixel points. Scale the depth values of the M pixel points in the second depth image according to the ratio of the first median value to the second median value to obtain the third depth image.

3. The method according to claim 1 or 2, characterized in that Process the depth image through the following formula to obtain point cloud data: , , ; in, represents the three-dimensional features of the feature point in the point cloud data corresponding to the i-th pixel in the depth image, represents the two-dimensional feature of the i-th pixel in the depth image, Indicates the depth image The depth value of the i-th pixel, A two-dimensional feature representing the center point of the depth image, represents the first focal length of the depth image, Represents the second focal length of the depth image.

4. The method according to claim 1 or 2, characterized in that The step of obtaining the first point cloud data according to the second point cloud data and the third point cloud data includes: Determine N target points corresponding to the N feature points of the second point cloud data from the M second feature points of the third point cloud data. Adjust the three-dimensional features of the N target points according to the three-dimensional features of the N feature points, and adjust the three-dimensional features of the M - N traction points in the second point cloud data, where the M - N traction points are points other than the N target points in the third point cloud data, and the three-dimensional features include two-dimensional features and depth values; Obtain the first point cloud data according to the adjusted three-dimensional features of the N target points and the three-dimensional features of the M - N traction points.

5. The method according to claim 4, characterized in that Obtain the first point cloud data through the following formula: ; in, represents the three-dimensional features of the i-th target point among the N target points after adjustment, represents the three-dimensional features of the feature points corresponding to the i-th target point in the second point cloud data, represents the three-dimensional feature of the i-th traction point among the MN traction points after adjustment, represents the desired three-dimensional features of the i-th traction point, is the coefficient, 0< <1.

6. The method according to claim 3, characterized in that Process the point cloud data through the following formula to obtain the depth image: , , 。 7. The method according to claim 1 or 2, characterized in that The method further includes: Perform image semantic segmentation on the RGB image, and calculate the segmented RGB image using an artificial intelligence (AI) algorithm to obtain the second depth image.

8. An electronic device, characterized in that: The electronic device includes: a memory and one or more processors; the memory is coupled to the processors, and the memory is used to store computer program code, and the computer program code includes computer instructions. When the computer instructions are executed by the one or more processors, the electronic device performs the following operations: Obtain a first depth image and a second depth image, where the first depth image is the depth image of the acquired target scene, the first depth image includes N pixel points, the second depth image is the depth image obtained by processing the RGB image of the target scene, the second depth image includes M pixel points, M and N are both integers, and 1 < N < M; the depth image includes the pixel information of the target scene, and the pixel information includes the two-dimensional features and depth values of each pixel point; Obtain a second point cloud data according to the first depth image; where the second point cloud data includes the two-dimensional features and depth values of N feature points, the N feature points correspond to the N pixel points one by one, and the depth values of the N feature points are obtained according to the depth values of the N pixel points; Scale the depth values of the pixel points in the second depth image according to the depth values of the pixel points in the first depth image to obtain a third depth image; where the difference between the depth value of a pixel point in the third depth image and the depth value of the corresponding pixel point in the first depth image is less than the difference between the depth value of the corresponding pixel point in the second depth image and the depth value of the corresponding pixel point in the first depth image; Obtain a third point cloud data according to the third depth image; where the third point cloud data includes the two-dimensional features and depth values of M second feature points, the M second feature points correspond to the M pixel points one by one, and the depth values of the M second feature points are obtained according to the depth values of the M pixel points; ​ ​ 9. The electronic device according to claim 8, wherein: When the computer instructions are executed by the one or more processors, the electronic device further performs the following steps: Determine a first median value and a second median value; wherein the first median value is the median value of the depth values ​​of the N pixels; and the second median value is the median value of the depth values ​​of the M pixels; The depth values ​​of the M pixels in the second depth image are scaled according to a ratio of the first median value to the second median value to obtain the third depth image.

10. The electronic device according to claim 8 or 9, characterized in that: The electronic device processes the depth image to obtain point cloud data using the following formula: , , ; in, represents the three-dimensional features of the feature point in the point cloud data corresponding to the i-th pixel in the depth image, represents the two-dimensional feature of the i-th pixel in the depth image, represents the depth value of the i-th pixel in the depth image, A two-dimensional feature representing the center point of the depth image, Indicates the depth The first focal length of the image, Represents the second focal length of the depth image.

11. The electronic device according to claim 8 or 9, characterized in that: When the computer instructions are executed by the one or more processors, the electronic device further performs the following steps: Determine N target points corresponding to the N feature points of the second point cloud data from the M second feature points of the third point cloud data; Adjusting the three-dimensional features of the N target points based on the three-dimensional features of the N feature points, and adjusting the three-dimensional features of MN traction points in the second point cloud data, where the MN traction points are points other than the N target points in the third point cloud data, and the three-dimensional features include two-dimensional features and depth values; The first point cloud data is obtained according to the adjusted three-dimensional features of the N target points and the three-dimensional features of the MN traction points.

12. The electronic device according to claim 11, wherein: The electronic device obtains the first point cloud data by the following formula: ; in, represents the three-dimensional features of the i-th target point among the N target points after adjustment, represents the three-dimensional features of the feature points corresponding to the i-th target point in the second point cloud data, represents the three-dimensional feature of the i-th traction point among the MN traction points after adjustment, represents the desired three-dimensional features of the i-th traction point, is the coefficient, 0< <1.

13. The electronic device according to claim 10, wherein: The depth image is obtained by processing the point cloud data using the following formula: , , 。 14. The electronic device according to claim 8 or 9, characterized in that: When the computer instructions are executed by the one or more processors, the electronic device further performs the following steps: The RGB image is subjected to image semantic segmentation, and an artificial intelligence (AI) algorithm is used to calculate the segmented RGB image to obtain the second depth image.

15. A chip system, characterized in that: The chip system is applied to an electronic device; the chip system includes one or more interface circuits and one or more processors; the interface circuit and the processor are interconnected through lines; the interface circuit is used to receive a signal from the memory of the electronic device and send the signal to the processor, the signal including computer instructions stored in the memory; when the processor executes the computer instructions, the electronic device executes the method as described in any one of claims 1 to 7.

16. A computer storage medium, characterized in that The method comprises computer instructions, which, when executed on an electronic device, enable the electronic device to execute the method according to any one of claims 1 to 7.

17. A computer program product, characterized in that When the computer program product is run on a computer, the computer is caused to perform the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Camera motion and image brightness-based Kinect depth reconstruction algorithm

    CN106780592A