Image processing method, electronic device, storage medium and computer program product

By generating a reference depth image using a color camera to complete the missing depth information from the depth camera, and performing coordinate system transformation and image restoration, the problem of facial image distortion when shooting at close range is solved, thus improving image quality.

CN121397384APending Publication Date: 2026-01-23HONOR DEVICE CO LTD
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
CN202410941689.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-07-12
Publication Date
2026-01-23

AI Technical Summary

Technical Problem

When shooting at close range, the depth information of the hair edges and facial edges of the user's face is missing, resulting in reduced image quality and distortion.

Method used

A reference depth image is generated by using image features acquired by a color camera to fill in the missing depth information in the depth image captured by the depth camera. Coordinate system transformation and depth offset processing are then performed, and image inpainting algorithms are combined to improve image quality.

Benefits of technology

It effectively reduces image distortion and improves image quality, especially the clarity and integrity of facial details in user selfie scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121397384A_ABST
    Figure CN121397384A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, in particular to an image processing method, electronic equipment, a storage medium and a computer program product. In the method, the electronic equipment shoots the same object through a depth camera and a color camera to obtain a depth image and a color image, and the electronic equipment can obtain the depth image and the color image through image features, such as textures, patterns, shapes, illumination, shadows, spatial relationships, semantic information and the like, in the color image shot by the color camera. And obtaining a reference depth image comprising depth difference values or depth relationships among the pixel points in the corresponding color image. And then, the electronic equipment corrects the depth image shot by the depth camera based on the reference depth image, that is, depth information of pixel points with missing depth information in the depth image is supplemented. Therefore, the electronic equipment can complement the depth information of the pixel points with the missing depth information in the depth image, so that image distortion can be avoided, and the image quality is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to an image processing method, electronic device, storage medium, and computer program product. Background Technology

[0002] For reference Figure 1 The illustration of a selfie scenario assumes that when the distance between the phone 110 and the user's face 120 is relatively short, for example, within 35 to 40 centimeters, the user's shooting posture is more natural and comfortable. However, because the dimensions of areas such as the hairline and facial edges of the user's face 120 are relatively small, they may be below the resolution threshold of the sensor in the depth camera of the phone 110, causing the sensor to be unable to acquire depth information in these tiny areas. Therefore, close-range shooting leads to the loss of some depth information in the image, resulting in a decrease in image quality. Summary of the Invention

[0003] To address the problem of image distortion caused by close-up photography, resulting in reduced image quality, this application provides an image processing method, electronic device, storage medium, and computer program product, including:

[0004] In a first aspect, embodiments of this application provide an image processing method applied to an electronic device, comprising acquiring a first depth image and a color image of an object captured by the electronic device; obtaining a reference depth image based on image features of the color image, wherein the reference depth image includes the depth relationship between each pixel in the color image; and determining the depth information of a first pixel in the first depth image that is missing depth information based on the reference depth image to obtain a second depth image.

[0005] Based on the above scheme, by using a reference depth image obtained from image features in a color image within the same coordinate system as the depth image, the depth image captured by the depth camera can be corrected. This can reduce distortion in certain parts of the object and improve the shooting effect. In other words, by supplementing the depth information of pixels lacking depth information in the first depth image, image distortion can be avoided, and image quality can be improved.

[0006] Furthermore, in this embodiment, only one color image is needed to complete the depth information of pixels missing depth information in the first depth image, which can save the processing resources of electronic devices.

[0007] It can be understood that the first depth image can be an image captured by a time-of-flight camera in an electronic device, and the color image can be an image captured by a color camera in an electronic device. Image features can include texture, pattern, shape, lighting, shadow, spatial relationships, and semantic information, etc.

[0008] It is understood that multiple pixels in a reference depth image may lack depth information. This embodiment of the application uses the first pixel lacking depth information as an example for illustration.

[0009] In some optional implementations of the first aspect, determining the depth information of the first pixel based on a reference depth image includes: determining a plurality of associated pixels in the reference depth image that are associated with the first pixel, wherein the plurality of associated pixels include a first associated pixel corresponding to the first pixel and at least one second associated pixel, and at least some of the second associated pixels have pixels in the first depth image that do not lack depth information; and determining the depth information of the first pixel based on the depth information of the plurality of pixels corresponding to the plurality of associated pixels in the first depth image.

[0010] In some optional implementations of the first aspect, determining multiple associated pixels in the reference depth image that are associated with the first pixel includes: determining at least one second pixel in the first depth image, wherein the second pixel is a pixel that does not lack depth information; and taking the pixel corresponding to the second pixel in the reference depth image as the second associated pixel.

[0011] In some optional implementations of the first aspect, the depth information of the first pixel is determined based on the depth information of multiple pixels corresponding to multiple associated pixels in the first depth image, including: obtaining the depth information of the first pixel based on the depth relationship between the first and second associated pixels in the reference depth image, and the depth information of the second pixel in the first depth image.

[0012] For example, such as Figure 2 As shown, assuming the depth information of the first pixel p1 in the first depth image P1 is missing, while the depth information of pixel p2 is not missing, and the pixel corresponding to the first pixel p1 in the reference depth image P2 is the first associated pixel p1', and the pixel corresponding to the first pixel p2 is the second associated pixel p2'. If the depth ratio between the first associated pixel p1' and the second associated pixel p2' in the reference depth image P2 is 5:1, then 5 times the depth value of the second pixel p2 in the depth image P1 can be determined as the depth value of the first pixel p1.

[0013] In some alternative implementations of the first aspect, multiple associated pixels of the first pixel are determined from the reference depth image, and at least one pixel in the first depth image whose similarity to the depth information of the first associated pixel is higher than a similarity threshold and whose corresponding pixel in the first depth image does not lack depth information is determined as the second associated pixel.

[0014] In some optional implementation forms of the first aspect, the depth information of the first pixel point is determined based on the depth information of the plurality of pixel points corresponding to the plurality of associated pixel points in the first depth image, including: determining the plurality of pixel points corresponding to the plurality of associated pixel points from the first depth image; and determining the depth information of the first pixel point based on the depth information of the plurality of pixel points.

[0015] For example, as shown in FIG. 1, it is assumed that the depth information of a first pixel point p1 in a depth image P1 is missing, and a first associated pixel point p1' corresponding to the first pixel point p1 in a reference depth image P2. If the depth ratio between a second associated pixel point p2' and the first associated pixel point p1' is 1:1, the depth ratio between a third associated pixel point p3' and the first associated pixel point p1' is 1:1.2, and the pixel point corresponding to the second associated pixel point p2' in the depth image P1 is a second pixel point p2, and the second pixel point corresponding to the third associated pixel point p3' is a pixel point p3, the depth information of the first pixel point p1 can be determined according to the depth values of the second pixel point p2 and the second pixel point p3. For example, the average of the depth values of the second pixel point p2 and the second pixel point p3 is determined as the depth value of the first pixel point p1. Figure 3

[0016] In some optional implementation forms of the first aspect, the method further includes: performing bias processing on the depth information of each pixel point in the second depth image based on a depth offset to obtain a third depth image, wherein the depth offset is a user input offset, or is determined based on a shooting distance between the electronic device and the object and a first relationship, wherein the first relationship is a one-to-one correspondence relationship between the shooting distance and the depth offset.

[0017] It can be understood that after obtaining the second depth image, the electronic device can first perform overall translation on the second depth image based on the depth offset, project the second depth image to a position closer to the depth camera, or project the second depth image to a position farther away from the depth camera, to obtain a third depth image. In this way, the distortion of the part of the object can be further reduced, and the shooting effect can be further improved.

[0018] For example, if the electronic device detects that the distance between the object and the electronic device is 30 centimeters, the depth offset can be determined as 30 centimeters based on the one-to-one correspondence relationship between the pre-stored shooting distance and the depth offset.

[0019] ​In some optional implementation forms of the first aspect, the method further includes: performing conversion processing on the third depth image based on the scaling factor and a coordinate conversion relationship to obtain a fourth depth image, wherein a third pixel point in the fourth depth image representing the first position of the object has a same position in the fourth depth image as a fourth pixel point in the color image representing the first position, and the coordinate conversion relationship is a conversion relationship between a coordinate system of a depth camera capturing the first depth image and a coordinate system of a color camera capturing the color image; and fusing the fourth depth image and the color image to obtain the target image.

[0020] In this way, the third pixel point in the fourth depth image representing the first position of the object has a one-to-one correspondence with the fourth pixel point in the color image representing the first position, so that the distortion of the partial part of the object can be further reduced, and the shooting effect can be further improved.

[0021] In some optional implementation forms of the first aspect, the scaling factor is determined by: determining at least one key point from the second depth image; determining an overall depth of the second depth image according to depth information of the at least one key point; and taking a ratio of a total depth to the overall depth as the scaling factor, wherein the total depth is a sum of the overall depth and a depth offset of the second depth image.

[0022] In some optional implementation forms of the first aspect, the object is a face, and the key point includes a pixel point corresponding to at least one of the following features in the face: left corner of eye, right corner of eye, left corner of mouth, right corner of mouth, and tip of nose.

[0023] In some specific implementation forms, coordinates of multiple face key points in the second depth image in the color image coordinate system can be obtained, and a barycentric coordinate of the multiple face key points can be determined based on the coordinates of the multiple face key points in the color image coordinate system. Then, depth information corresponding to the barycentric coordinate in the second depth image can be determined as the overall depth d of the second depth image. Subsequently, a scaling factor can be determined based on a ratio of a total depth of the overall depth of the complete depth image and the depth offset offset to the overall depth of the complete depth image.

[0024] For example, the scaling factor can be represented by the following formula (1):

[0025] α=(d+offset) / d Formula (1)

[0026] Wherein, α represents the scaling factor, d represents the overall depth of the third depth image, and offset represents the depth offset.

[0027] In the embodiments of the present application, noise is introduced in the process of overall translation of the second depth image and re-projection of the third depth image, thereby affecting the definition of the finally generated target image. Therefore, after obtaining the target image, an adaptive feature fusion algorithm can be used to improve the details of the target image, for example, to improve the definition of the target image, and obtain a target image with high definition.

[0028] Further, since the distance between the face and the camera (for example, the depth camera and the color camera mentioned above) is less than the distance between the ear and the camera, the face observed by the camera will become wider than the actual face of the user, and the ear observed by the camera will become smaller than the actual ear of the user, so the face will block part of the ear. Therefore, when the camera captures the face of the user, it will not be able to collect image information such as color information and depth information of part of the ear, thereby causing part of the ear to be incomplete in the captured image.

[0029] Therefore, an image inpainting algorithm can be used to complete the image after detail improvement. The image inpainting algorithm can include a mask-aware transformer for large hole in image inpainting (MAT), a generative landmark guided face inpaintor (LaFIn), and a second version of deep fill (DeepFill v2) algorithm.

[0030] In a second aspect, the present application provides an electronic device, comprising: a memory for storing instructions executed by one or more processors of the electronic device, and a processor, which is one of the one or more processors of the electronic device, for executing the image processing method of the first aspect of the present application or any of the image processing methods mentioned in the first aspect.

[0031] In a third aspect, the present application provides a readable storage medium, and the readable storage medium stores instructions, which, when executed on an electronic device, cause the electronic device to execute the image processing method of the first aspect of the present application or any of the image processing methods mentioned in the first aspect.

[0032] In a fourth aspect, the embodiments of the present application provide a computer program product, which includes computer instructions, and when executed by an electronic device, the electronic device executes the computer program code of the image processing method of the first aspect of the present application or any of the image processing methods mentioned in the first aspect. BRIEF DESCRIPTION OF DRAWINGS

[0033] Figure 1 An application scenario schematic diagram is shown;

[0034] Figure 2 According to some embodiments of the present application, a schematic diagram corresponding to a depth image P1 and a reference depth image P2 is shown;

[0035] Figure 3 According to some embodiments of the present application, another schematic diagram corresponding to a depth image P1 and a reference depth image P2 is shown;

[0036] Figure 4A According to some embodiments of the present application, a schematic diagram of a parallax is shown;

[0037] Figure 4B According to some embodiments of the present application, a schematic diagram corresponding to a depth image and a color image is shown;

[0038] Figure 4C According to some embodiments of the present application, a schematic diagram of a distorted image is shown;

[0039] Figure 5 According to some embodiments of the present application, a schematic diagram of a flow of an image processing method is shown;

[0040] Figure 6 According to some embodiments of the present application, a schematic diagram of re-projection and canvas scaling is shown;

[0041] Figure 7 According to some embodiments of the present application, another schematic diagram of a flow of an image processing method is shown;

[0042] Figure 8 According to some embodiments of the present application, a schematic diagram of a calibration image is shown;

[0043] Figure 9 According to some embodiments of the present application, a schematic diagram of a hardware structure of an electronic device is shown;

[0044] Figure 10 According to some embodiments of the present application, a schematic diagram of a software structure of an electronic device is shown. DETAILED DESCRIPTION

[0045] Embodiments of the present application include but are not limited to an image processing method, an electronic device, a storage medium and a computer program product.

[0046] It can be understood that the image processing method mentioned in the embodiments of the present application can be applied to an electronic device. The electronic device can be referred to as a terminal, a user equipment (UE), a mobile station (MS), a mobile terminal (MT), etc. In some specific implementation manners, the electronic device can be a mobile phone, a tablet computer, a wearable device with wireless communication function, a computer with wireless transceiver function, a virtual reality (VR) device, an augmented reality (AR) device, etc., which are not limited in the present application.

[0047] For ease of illustration, the electronic device is taken as a mobile phone in the following description. In some specific implementation manners, the mobile phone can have a depth camera and a color camera. The depth camera can capture a depth image and an infrared image corresponding to each pixel point at the same position. The image parameters in the depth image include depth information, which can represent the distance from the object to the depth camera. The image parameters in the infrared image include infrared light information, which can represent the reflection of infrared light by the object. The color camera can capture a color image, which is used to reflect the color of the object.

[0048] As mentioned above, as shown in FIG. 1, the size of the hair edge region, the face edge region and other regions of the face 120 of the user can be smaller than the resolution threshold of the sensor of the depth camera in the mobile phone 110, so that the sensor cannot obtain the depth information of these small regions. Therefore, close-range shooting can cause the partial depth information of the image to be distorted, resulting in reduced image quality. Figure 1

[0049] Therefore, in order to solve the above problems, the embodiments of the present application provide an image processing method. In the method, the electronic device captures a depth image and a color image of the same object by using a depth camera and a color camera at the same time. The electronic device can obtain a reference depth image including the depth difference or depth relationship between each pixel point in the corresponding color image by using the image features in the color image captured by the color camera, such as texture, pattern, shape, illumination, shadow, spatial relationship and semantic information, etc. Then, the electronic device corrects the depth image captured by the depth camera based on the reference depth image, that is, supplements the depth information of the pixel points with missing depth information in the depth image.

[0050] In this way, the electronic device can complete the depth information of the pixel points with missing depth information in the depth image, so as to avoid image distortion and improve image quality. ​

[0051] Specifically, such as Figure 2 As shown, in some embodiments, it is assumed that the depth information of pixel p1 in depth image P1 is missing, while the depth information of pixel p2 is not missing. In the reference depth image P2, the pixel corresponding to pixel p1 is pixel p1', and the pixel corresponding to pixel p2 is pixel p2'. If the depth ratio between pixel p1' and pixel p2' in the reference depth image P2 is 5:1, then 5 times the depth value of pixel p2 in depth image P1 can be determined as the depth value of pixel p1.

[0052] like Figure 3 As shown, in some embodiments, it is assumed that the depth information of pixel p1 in depth image P1 is missing, while the corresponding pixel in reference depth image P2 is p1'. If in reference depth image P2, the depth ratio between pixel p2' and pixel p1' is 1:1, and the depth ratio between pixel p3' and pixel p1' is 1:1.2, and in depth image P1, the corresponding pixel for pixel p2' is pixel p2, and the corresponding pixel for pixel p3' is pixel p3, then the depth information of pixel p1 can be determined based on the depth values ​​of pixel p2 and pixel p3. For example, the average of the depth values ​​of pixel p2 and pixel p3 can be used as the depth value of pixel p1.

[0053] For example, in a user selfie scenario, for each pixel in the hair edge region that lacks depth information, the electronic device can determine the depth information of each pixel in the hair edge region based on the depth information of multiple pixels in the hair middle region that has depth information.

[0054] In some embodiments, as described above, the electronic device may include a depth camera and a color camera, which are typically located in different positions within the electronic device. Because when the same object is photographed from different positions, the position of the object in the depth image captured by the depth camera and the position of the object in the color image captured by the color camera differ, as... Figure 4A The discrepancy shown (hereinafter referred to as parallax) means that when fusing depth and color images, pixels at the same location in the depth and color images may not correspond one-to-one, resulting in image distortion.

[0055] For example, in a user selfie scenario, such as Figure 4BAs shown, the position of region A in the face in the depth image is A1, and the position of region A in the color image is A2. When the depth image and the color image are fused, the pixel point at position A1 in the depth image will be fused with the pixel point at position B2 in the color image, and the pixel point at position A2 in the color image will be fused with the pixel point at position B1 in the depth image. For example, a certain pixel point of the face in the depth image is fused with a certain pixel point of the ear in the color image, which will cause the nose to be too large, the face to be too long and wide, and thus the image to be distorted.

[0056] It needs to be understood that, Figure 4B In actual applications, the position of the same point on the object in the depth image and the position of the same point on the object in the color image can also be less than Figure 4B the deviation shown.

[0057] Therefore, before correcting the depth image captured by the depth camera, the electronic device can convert the depth image and the color image obtained by simultaneously capturing the same object through the depth camera and the color camera into the same coordinate system, and in the coordinate system, adjust the conversion relationship between different pixel points representing the same point on the object in the depth image and the color image, so that the pixel points at the same position in the depth image and the color image correspond one by one. In this way, the electronic device corrects the depth image captured by the depth camera through the reference depth image obtained by the image features in the color image in the same coordinate system as the depth image, which can reduce the distortion of part of the object and improve the shooting effect.

[0058] Moreover, after correcting the depth image captured by the depth camera, the electronic device can first perform overall translation on the depth image with the depth information supplemented (hereinafter referred to as the second depth image) based on the depth offset, project the second depth image to a position closer to the depth camera, or project the second depth image to a position farther from the depth camera, to obtain a third depth image. Then, the scaling factor can be determined based on the depth information of the key points in the third depth image and the depth offset. The electronic device can perform conversion processing on the third depth image through the scaling factor and the coordinate conversion relationship, for example, enlarge or reduce the third depth image, to obtain a fourth depth image. The position of the third pixel point representing the first position of the object in the fourth depth image is the same as the position of the fourth pixel point representing the first position in the color image, and the coordinate conversion relationship is the conversion relationship between the coordinate system of the depth camera and the coordinate system of the color camera. The fourth depth image and the color image are fused to obtain a target image. The quality of the target image is higher than that of the fused image obtained by fusing the first depth image and the color image.

[0059] It can be understood that in some optional implementations, the depth offset can be a depth offset offset input by the user in real time to the electronic device, or can be determined according to a one-to-one correspondence relationship between a pre-stored shooting distance and the depth offset of the electronic device. For example, if the electronic device detects that the distance between the object and the electronic device is 30 centimeters, the depth offset can be determined to be 30 centimeters based on the pre-stored one-to-one correspondence relationship between the shooting distance and the depth offset.

[0060] It can be understood that in some optional implementations, the scaling factor can be determined in the following manner:

[0061] The electronic device can obtain depth information of a plurality of key points in the second depth image, and determine an overall depth d corresponding to the second depth image based on the depth information of the plurality of key points. The key points can be points whose relative positional relationship does not change when the depth camera takes a close-up shot of the object at different angles. For example, in a user selfie scenario, the key points of the user's face can include the left corner of the eye, the right corner of the eye, the left corner of the mouth, the right corner of the mouth, the tip of the nose, etc., also known as face landmarks. The electronic device can determine the scaling factor based on the ratio of the total depth of the depth offset to the overall depth corresponding to the second depth image. For example, the scaling factor can be represented by the following formula (1):

[0062] α = (d + offset) / d Formula (1)

[0063] Wherein, α represents the scaling factor, d represents the overall depth corresponding to the second depth image, and offset represents the depth offset.

[0064] In some specific implementations, the electronic device can obtain the coordinates of a plurality of face landmarks in the color image coordinate system, and determine the barycentric coordinates of the plurality of face landmarks based on the coordinates of the plurality of face landmarks in the color image coordinate system. Then the electronic device can determine the depth information corresponding to the barycentric coordinates in the second depth image as the overall depth d corresponding to the second depth image.

[0065] In addition, due to the influence of "near large and far small", the size of different parts observed by the camera will change compared to the actual size of the part, so the part observed by the camera that is larger than the actual size will block the part that is smaller than the actual size.

[0066] For example, in a user selfie scenario, for the face of the user, since the distance between the face and the camera (e.g., the aforementioned depth camera and color camera) is less than the distance between the ear and the camera, the face observed by the camera will become wider than the actual face of the user, and the ear observed by the camera will become narrower than the actual ear of the user, so the face will block part of the ear. Thus, when the camera captures the face of the user, the camera cannot collect image information (e.g., color information and depth information) of part of the ear, which will cause the captured image to present a partial ear defect as shown in Figure 4C

[0067] Therefore, after the fourth depth image and the color image are fused to obtain the target image, an image enhancement algorithm and an image completion algorithm can be used to enhance the details of the partial defect in the target image and complete the partial defect. In this way, the quality of the image can be further improved.

[0068] For example, in a user selfie scenario, an adaptively spatial feature fusion (ASFFNet) algorithm can be used to enhance the details of the target image, and an image inpainting algorithm can be used to complete the image after the details are enhanced.

[0069] The image processing method mentioned in the embodiments of the present application will be described in detail below.

[0070] It can be understood that the image processing method mentioned in the embodiments of the present application can be applied to image beautification, photographing, video recording, video call, and the like. For example, in an image beautification scenario, the image processing method can be used to process images in a gallery; for another example, in a photographing scenario, the image processing method can be used to process images captured by a camera in real time; for another example, in a video recording or video call scenario, the image processing method can be used to process each frame of image or part of the image. The embodiments of the present application do not limit the specific application scenarios of the image processing method.

[0071] For ease of illustration, the image processing method is applied to a user selfie scenario in the following description. In a user selfie scenario, the object can include the face of the user.

[0072] As shown in Figure 5 , a flowchart of an image processing method is shown. The image processing method can be executed by an electronic device, for example, a mobile phone 100. Specifically, the image processing method can include:

[0073] First, an RGB image, a partial depth image, and camera parameters are obtained.

[0074] ​It can be understood that the RGB image can be a color image obtained by a color camera in the mobile phone 100 photographing a user's face, also known as a red green blue (RGB) image.

[0075] It can be understood that the mobile phone 100 can obtain the depth image by, for example, a time of flight (TOF) camera. Since the light reflected by the user's face is reflected back to the TOF camera through different paths, the flight time measured by the TOF camera can be inaccurate, which can cause some pixels in the depth image obtained by the TOF camera to lack depth information. In addition, in close-range photography, the size of the edge region of the user's hair, the edge region of the face, and other regions can be smaller than the resolution threshold of the sensor in the TOF camera, so that the sensor cannot obtain the depth information of these small regions. Thus, the depth image obtained by the TOF camera and having some missing depth information is a defective depth image.

[0076] It can be understood that the conversion relationship between the depth camera coordinate system and the color camera coordinate system can include a translation relationship and a rotation relationship, which can specifically be represented as a translation matrix and a rotation matrix between the depth camera coordinate system and the color camera coordinate system.

[0077] In some specific implementations, the translation matrix and the rotation matrix that convert the depth camera coordinate system to coincide with the color camera coordinate system can be collectively referred to as the extrinsic parameters of the depth camera. The conversion relationship between the depth camera coordinate system and the depth image coordinate system can be referred to as the intrinsic parameters of the depth camera, which can include the focal length, the principal point coordinates, and other parameters of the depth camera. The conversion relationship between the color camera coordinate system and the color image coordinate system can be referred to as the intrinsic parameters of the color camera, which can include the focal length, the principal point coordinates, and other parameters of the color camera. The extrinsic parameters of the depth camera, the intrinsic parameters of the depth camera, and the intrinsic parameters of the color camera can be collectively referred to as camera parameters.

[0078] Then, the RGB image and the defective depth image are aligned in RGBD.

[0079] It can be understood that the RGB image and the defective depth image can be aligned in RGBD to obtain an aligned depth image. Here, RGB can refer to three primary colors, and D can refer to depth information.

[0080] It can be understood that, as described above, the depth camera and the color camera in the electronic device are located at different positions. When photographing the user's face from different positions, there is a parallax between the position of the user's face in the depth image photographed by the depth camera and the position of the user's face in the color image photographed by the color camera, which can cause the pixels at the same position in the depth image and the color image to not correspond one by one, thereby causing distortion in the images.

[0081] In some optional implementations, the RGBD alignment can be performed on the RGB image and the incomplete depth map based on the color image and the camera parameters, so that the pixels at the same positions in the incomplete depth image and the color image correspond to each other. In this way, the distortion of the partial part of the object can be reduced, and the shooting effect can be improved.

[0082] Specifically, the depth image and the color image captured by the depth camera and the color camera can be converted to the same coordinate system, and in the coordinate system, the conversion relationship between the different pixels representing the same point on the object in the depth image and the color image can be adjusted, so that the pixels at the same positions in the depth image and the color image correspond to each other.

[0083] Then, the depth estimation is performed on the RGB image.

[0084] It can be understood that the monocular depth estimation algorithm, such as the DepthAnything algorithm, can be used to perform the depth estimation on the RGB image to obtain the reference depth image including the depth difference or depth relationship between the pixels.

[0085] In some specific estimation processes of the DepthAnything algorithm, the depth estimation can be performed on the RGB image based on the image features in the RGB image, such as texture, pattern, shape, illumination, shadow, spatial relationship, and semantic information.

[0086] Further, the depth completion is performed on the aligned depth map.

[0087] It can be understood that since the reference depth image is obtained by performing the depth estimation on the RGB image, the pixels at the same positions in the reference depth image and the RGB image correspond to each other, and therefore, the depth completion can be performed on the aligned depth map based on the conversion relationship between the depth camera coordinate system and the color camera coordinate system and based on the reference depth image to obtain the complete depth map.

[0088] The transformation matrix can include rotation parameters, translation parameters, and scaling parameters. Based on the transformation matrix, the pixels in the reference depth image can be mapped to the corresponding pixels with depth information in the aligned depth map in the color camera coordinate system. The corresponding pixels in the reference depth image and the aligned depth map can represent the same point on the user's face. In this way, the coordinates of the corresponding pixels in the aligned depth map in the color camera coordinate system can be calculated based on the transformation matrix and the pixels in the reference depth image, so that the depth information of the pixels with missing depth information in the aligned depth map can be obtained.

[0089] In some optional implementations, for the pixel point with missing depth information in the aligned depth map, a first associated pixel point corresponding to the first pixel point with missing depth information can be determined from the reference depth image, and then a second associated pixel point with a depth ratio to the first associated pixel point within a ratio interval, for example, [1:1~1.5], such as 1:1 or 1:1.2, can be determined from the reference depth image based on the depth relationship, for example, the depth ratio, between the pixel points in the reference depth image. Then a second pixel point corresponding to the second associated pixel point can be determined from the aligned depth map, and then the depth information of the first pixel point with missing depth information can be determined based on the depth information of the second pixel point.

[0090] Further, the complete depth map is translated as a whole.

[0091] It can be understood that after obtaining the complete depth map, the complete depth map can be translated as a whole based on the depth offset to obtain the depth map at the target distance.

[0092] In some optional implementations, the depth offset can be a depth offset offset input by the user in real time to the electronic device, or can be determined according to a one-to-one correspondence relationship between the shooting distance and the depth offset pre-stored by the electronic device. For example, if the electronic device detects that the distance between the object and the electronic device is 30 centimeters, the depth offset of 30 centimeters can be determined based on the one-to-one correspondence relationship between the pre-stored shooting distance and the depth offset.

[0093] In some specific implementations, as shown in Figure 6 The depth value of each pixel point in the complete depth map can be added to the depth offset offset to translate the complete depth map as a whole to obtain the depth map at the target distance.

[0094] Further, the depth map at the target distance is re-projected and canvas scaled.

[0095] It can be understood that after obtaining the depth map at the target distance, the depth map at the target distance can be re-projected to the color image coordinate system to integrate the depth information and the color information. Since the complete depth map is translated as a whole, the positions of the pixel points in the depth map at the target distance in the depth camera coordinate system change relative to the positions of the pixel points in the complete depth map in the depth camera coordinate system. Therefore, the canvas of the depth map at the target distance can be scaled based on the scaling factor. Then the depth map after canvas scaling and the color image are integrated to obtain the RGB image at the target distance.

[0096] It can be understood that in some optional implementations, the scaling factor can be determined as follows:

[0097] The coordinates of the plurality of facial landmarks in the color image coordinate system in the complete depth map are obtained, and the barycentric coordinates of the plurality of facial landmarks are determined based on the coordinates of the plurality of facial landmarks in the color image coordinate system. Then, the depth information corresponding to the barycentric coordinates in the second depth image can be determined as the overall depth d corresponding to the second depth image. Then, the scaling factor can be determined based on the ratio of the total depth of the depth offset to the overall depth corresponding to the complete depth map. For example, the scaling factor can be represented by the above formula (1).

[0098] For example, the depth image at the position corresponding to the before-adjustment can be re-projected and canvas-scaled to adjust the depth image at the position corresponding to the before-adjustment to the position corresponding to the after-adjustment.

[0099] Further, the RGB image at the target distance is enhanced in detail.

[0100] It can be understood that, in the process of overall translation of the complete depth map and re-projection of the depth map at the target distance, noise will be introduced in the corresponding depth map, thereby affecting the clarity of the finally generated RGB image at the target distance. Therefore, after obtaining the RGB image at the target distance, an adaptive feature fusion algorithm can be used to enhance the details of the RGB image at the target distance, for example, to enhance the clarity of the RGB image at the target distance, to obtain a high-clarity RGB image. The high-clarity RGB image does not contain ears.

[0101] Further, the high-clarity RGB image is face-completed to obtain an ear-containing RGB.

[0102] It can be understood that, as mentioned above, because the distance between the face and the camera (such as the depth camera and the color camera mentioned above) is less than the distance between the ear and the camera, the face observed by the camera will become larger than the actual face of the user, and the ear observed by the camera will become smaller than the actual ear of the user, so the face will block part of the ear. Therefore, when the camera captures the face of the user, it will not be able to collect image information such as color information and depth information of part of the ear, thereby causing part of the ear to be incomplete in the captured image.

[0103] Therefore, an image inpainting algorithm can be used to complete the image after detail enhancement. The image inpainting algorithm can include the MAT algorithm, the LaFIn algorithm, and the second version of the DeepFill (DeepFill v2) algorithm.

[0104] The following will be further described in combination with some steps in Figure 7 Figure 5 For example, the depth image at the position corresponding to the before-adjustment can be re-projected and canvas-scaled to adjust the depth image at the position corresponding to the before-adjustment to the position corresponding to the after-adjustment.​Figure 7 As shown, a flowchart of another image processing method is shown. The image processing method can be performed by an electronic device, for example, the mobile phone 100. Specifically, the image processing method can include:

[0105] 701: obtaining a first depth image and a color image;

[0106] It can be understood that the color image can be an image obtained by a color camera in the mobile phone 100 for photographing a user's face, also known as a trichromatic image. The first depth image can be a partial depth information missing incomplete depth map photographed by a TOF camera in the mobile phone 100.

[0107] 702: obtaining a reference depth image based on image features of the color image, wherein the reference depth image includes depth relationships between pixel points in the color image.

[0108] It can be understood that monocular depth estimation algorithms, such as the DepthAnything algorithm, can be used to estimate the depth of the color image to obtain a reference depth image including depth differences or depth relationships between pixel points.

[0109] In some specific estimation processes of the DepthAnything algorithm, the depth of the color image can be estimated based on image features in the color image, such as texture, pattern, shape, illumination, shadow, spatial relationship, and semantic information.

[0110] 703: determining the depth information of the pixel points based on the reference depth image corresponding to the pixel points with missing depth information in the first depth image to obtain a second depth image.

[0111] It can be understood that the way to determine the depth information of the pixel points based on the reference depth image can include mapping the reference depth image and the depth image to the same coordinate system, such as a spatial coordinate system or a plane coordinate system. The spatial coordinate system can be a color camera coordinate system and a depth camera coordinate system, and the plane coordinate system can be a color image coordinate system and a depth image coordinate system. And based on the coordinates of the pixel points in the reference depth image in the coordinate system, the depth information of the pixel points with missing depth information in the depth image is supplemented.

[0112] It should be understood that the color camera coordinate system, the depth camera coordinate system, the color image coordinate system, and the depth image coordinate system can be converted to each other through the conversion relationship between the coordinate systems. In some specific implementations, the conversion relationship between the coordinate systems can be obtained by calibrating the cameras. The specific calibration method will be described in detail in Figure 8 , and to avoid repetition, it will not be described here.

[0113] The following takes the mapping of the reference depth image and the depth image to the color camera coordinate system and the color image coordinate system as an example to describe the specific implementation of the depth information of the pixel point with missing depth information in the supplementary depth image.

[0114] In some specific implementations, the depth image can be mapped to the color camera coordinate system based on the conversion relationship between the depth image coordinate system and the depth camera coordinate system and based on the conversion relationship between the depth camera coordinate system and the color camera coordinate system. And the reference depth image can also be mapped to the color camera coordinate system based on the conversion relationship between the color image coordinate system and the color camera coordinate system.

[0115] Then, the coordinates of the plurality of key points in the depth image in the color camera coordinate system can be obtained, and the coordinates of the corresponding plurality of key points in the reference depth image in the color camera coordinate system can be obtained, to determine the transformation matrix between the depth image and the reference depth image.

[0116] The transformation matrix can include rotation parameters, translation parameters and scaling parameters. Based on the transformation matrix, the pixel points in the reference depth image can be mapped to the color camera coordinate system to coincide with the corresponding pixel points with depth information in the depth image. The corresponding pixel points in the reference depth image and the depth image can represent the same point on the object. In this way, the coordinates of the corresponding pixel points in the depth image in the color camera coordinate system can be calculated based on the transformation matrix and the pixel points in the reference depth image, so that the depth information of the pixel points with missing depth information in the depth image can be obtained.

[0117] In other optional implementations, the depth image can be mapped to the color image coordinate system based on the conversion relationship between the depth image coordinate system and the depth camera coordinate system, the conversion relationship between the depth camera coordinate system and the color camera coordinate system, and the conversion relationship between the color camera coordinate system and the color image coordinate system. Then the pixel points corresponding to each pixel point in the depth image can be determined from the reference depth image.

[0118] For the pixel point with missing depth information in the depth image, the first associated pixel point corresponding to the pixel point with missing depth information can be determined from the reference depth image, and then the second associated pixel point with a depth ratio within a ratio interval between the first associated pixel point can be determined from the reference depth image based on the depth relationship, such as the depth ratio, between each pixel point in the reference depth image, where the ratio interval can be [1:1~1.5], and specific examples include 1:1 and 1:1.2. The third associated pixel point corresponding to the second associated pixel point can be determined from the depth image, and then the depth information of the pixel point with missing depth information can be determined based on the depth information of the third associated pixel point.

[0119] The following describes a specific embodiment of camera calibration.

[0120] It can be understood that, as described above, the conversion relationship between the depth camera coordinate system and the color camera coordinate system can include a translation relationship and a rotation relationship, and specifically can be represented by a translation matrix and a rotation matrix between the depth camera coordinate system and the color camera coordinate system. In some specific implementations, the translation matrix and the rotation matrix that make the depth camera coordinate system coincide with the color camera coordinate system after conversion can be collectively referred to as the extrinsic parameters of the depth camera. The conversion relationship between the depth camera coordinate system and the depth image coordinate system can be referred to as the intrinsic parameters of the depth camera, and the intrinsic parameters of the depth camera can include the focal length, the principal point coordinates, and the like of the depth camera. The conversion relationship between the color camera coordinate system and the color image coordinate system can be referred to as the intrinsic parameters of the color camera, and the intrinsic parameters of the color camera can include the focal length, the principal point coordinates, and the like of the color camera.

[0121] In some specific implementations, the conversion relationship between the depth camera coordinate system and the depth image coordinate system can refer to the following formula (2):

[0122]

[0123] wherein, represents a coordinate in the depth image coordinate system, K d represents an intrinsic matrix of the depth camera coordinate system, represents a coordinate in the depth camera coordinate system.

[0124] The conversion relationship between the color camera coordinate system and the color image coordinate system can refer to the following formula (3):

[0125]

[0126] wherein, represents a coordinate in the color image coordinate system, K r represents an intrinsic matrix of the color camera coordinate system, represents a coordinate in the color camera coordinate system.

[0127] The conversion relationship between the depth camera coordinate system and the color camera coordinate system can refer to the following formula (4):

[0128]

[0129] wherein, represents a coordinate in the color camera coordinate system, represents a coordinate in the depth camera coordinate system, R can represent a rotation matrix of the depth camera coordinate system transformed to the color camera coordinate system, and T represents a translation matrix of the depth camera coordinate system transformed to the color camera coordinate system, K represents the coordinates in the color image coordinate system. r The intrinsic parameter matrix representing the coordinate system of the color camera. K represents the coordinates in the depth image coordinate system. d The intrinsic parameter matrix represents the coordinate system of the depth camera.

[0130] The following section will take the calibration of the intrinsic parameters of a color camera as an example to provide a detailed introduction to the calibration process of depth cameras and color cameras.

[0131] like Figure 8 The diagram shows a schematic of a calibration image. The calibration image may include multiple ellipses and has four corner points a, b, c, and d.

[0132] It's understandable that a spatial coordinate system can be established with the top left corner of the calibrated image as the origin, the horizontal direction to the right from the origin as the first direction X, the vertical downward direction along the direction of gravity as the second direction Y, and the direction perpendicular to the paper and outward as the third direction Z. It's understood that the first direction X, the second direction Y, and the third direction Z are mutually perpendicular. Furthermore, by using a shape detection algorithm to determine the center point of each ellipse, the coordinates of the center point of each ellipse in the spatial coordinate system can be determined.

[0133] In the specific calibration process, a color camera can be used to take calibration images from different positions and angles, and a shape detection algorithm can be used to determine the center point of each ellipse in the color calibration image taken by the color camera.

[0134] In some optional implementations, the coordinates of the center points of all ellipses in the color calibration image within the color image coordinate system can be input into a convex hull calculation function. This function can then be used to determine the outermost ellipse. The convex hull calculation function can be any OpenCV function that calculates the convex hull. The color image coordinate system can be a coordinate system established with the top-left corner of the color calibration image as the origin, the horizontal direction to the right as the first direction (X), and the vertical direction downwards as the second direction (Y).

[0135] Then, the four corner points of the color calibration image can be determined from the outermost ellipse. For example, the four corner points of the color calibration image can be determined from multiple ellipses in the color calibration image captured by the color camera based on the angles formed by the lines between the center point of the outermost ellipse and adjacent points. For example, the four corner points of the calibration image can be determined based on the four smallest angles formed by the lines between the center point of the outermost ellipse and adjacent points.

[0136] In some optional implementations, any one of the outermost circles of the color calibration image captured by the color camera can be selected, and the distances between the center points of the remaining circles of the outermost circles and the center point of the selected circle can be determined by traversing the remaining circles of the outermost circles. The center points of the two circles with the smallest distances can be determined as the adjacent points of the center point of the selected circle. The four corner points of the color calibration image can be determined from the plurality of circles in the color calibration image according to the angles formed by the straight lines between the center points of each circle and the adjacent points.

[0137] For example, the center point of the circle whose angle formed by the straight line between the center point and the adjacent point is less than 180° can be determined as the corner point of the color calibration image. As shown in FIG. 6, the angle formed by the straight line between the center point a of the circle and the adjacent points e and f is 90°, and the angle formed by the straight line between the center point e of the circle and the adjacent points a and g is 180°. Therefore, the center point a of the circle can be determined as the corner point of the color calibration image. Figure 8

[0138] Then, the affine transformation can be used to perform row sorting and column sorting on all the circles in the calibration image to obtain the correct order of all the circles in the calibration image.

[0139] Then, the intrinsic parameters of the color camera coordinate system can be determined based on the opencv calibration function, the coordinates of the center points of all the circles in the calibration image in the spatial coordinate system, the coordinates of the center points of all the circles in the color calibration image in the color image coordinate system, and the position of the color camera coordinate system in the spatial coordinate system. The color camera coordinate system can be a coordinate system with the optical center of the color camera as the origin, the horizontal direction of the imaging plane of the color camera as the first direction X, and the vertical direction of the imaging plane of the color camera as the second direction Y.

[0140] Similarly, the intrinsic parameters of the depth camera and the extrinsic parameters of the depth camera can be determined based on the above method, and details are not repeated here to avoid repetition.

[0141] In some implementations of determining the extrinsic parameters of the color camera, the conversion relationship between the color image coordinate system and the world coordinate system can be determined. The derivation process of the conversion relationship between the color image coordinate system and the world coordinate system can refer to the following calculation formula:

[0142]

[0143] wherein, represents the coordinates in the color image coordinate system, d x represents the pixel increment in the x-axis direction of the color image coordinate system, d y ​represents a pixel increment along the y-axis direction in the color image coordinate system, u0 and v0 are offsets of the color camera, f represents a focal length, R can represent a rotation matrix, and T can represent a translation matrix, represents a coordinate in the world coordinate system, f x represents a focal length along the x-axis direction of the image, f y represents a focal length along the y-axis direction of the image.

[0144] It can be understood that the image processing method provided by the embodiments of the present application can be applied to an electronic device. The hardware structure of the electronic device to which the image processing method provided by the embodiments of the present application is applicable is exemplarily introduced below.

[0145] As shown in Figure 9 The electronic device 900 can include a processor 910, an external memory interface 920, an internal memory 921, a universal serial bus (USB) interface 930, a charge management module 940, a power management module 941, a battery 942, an antenna 1, an antenna 2, a mobile communication module 950, a wireless communication module 960, an audio module 970, a loudspeaker 970A, a receiver 970B, a microphone 970C, a sensor module 980, a key 990, a motor 991, an indicator 992, a camera 993, a display screen 994, and a subscriber identification module (SIM) card interface 995, and the like.

[0146] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the electronic device. In other embodiments of the present application, the electronic device 900 can include more or fewer components than illustrated, or combine certain components, or split certain components, or different arrangement of components. The illustrated components can be implemented in hardware, software, or a combination of software and hardware.

[0147] The processor 910 can include one or more processing units, for example: the processor 910 can include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units can be independent devices, or can be integrated in one or more processors.

[0148] In some optional implementations, the processor 910 can execute the image processing method mentioned in the embodiments of the present application.

[0149] Specifically, if the electronic device 900 obtains a reference depth image including depth differences or depth relationships between pixels according to image features in a color image, such as texture, pattern, shape, illumination, shadow, spatial relationship, semantic information, etc. Then based on the reference depth image, the depth image captured by the depth camera is corrected, that is, the depth information of the pixel points with missing depth information in the depth image is supplemented. In this way, the depth information of the pixel points with missing depth information in the image can be completed, so that image distortion can be avoided, and image quality can be improved.

[0150] The controller can generate operation control signals according to the instruction operation code and the timing signal, and complete the control of fetching and executing instructions.

[0151] The processor 910 can also be provided with a memory for storing instructions and data. In some embodiments, the memory in the processor 910 is a cache memory. The memory can save instructions or data that the processor 910 has just used or repeatedly uses. If the processor 910 needs to use the instructions or data again, it can be directly called from the memory. This avoids repeated access and reduces the waiting time of the processor 910, thereby improving the efficiency of the system.

[0152] In some optional implementations, the memory can store instructions or data of the image processing method mentioned in the embodiments of the present application.

[0153] The electronic device realizes the display function through the GPU, the display screen 994, and the application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 994 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 910 can include one or more GPUs that execute program instructions to generate or change display information.

[0154] The display screen 994 is configured to display images, videos, and the like. The display screen 994 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flex light-emitting diode (FLED), a Mini-LED, a MicroLED, a Micro-OLED, a quantum dot light emitting diode (QLED), or the like. In some embodiments, the electronic device can include one or N display screens 994, where N is a positive integer greater than 1.

[0155] The electronic device can implement the photographing function through the ISP, the camera 993, the video codec, the GPU, the display screen 994, and the application processor.

[0156] The ISP is configured to process the data fed back by the camera 993. For example, when taking a photo, the shutter is opened, the light is transmitted to the camera photosensitive element through the lens, the light signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to the ISP for processing to convert it into an image visible to the naked eye. The ISP can also optimize the algorithm for the noise, brightness, and skin color of the image. The ISP can also optimize the exposure, color temperature, and other parameters of the shooting scene. In some embodiments, the ISP can be arranged in the camera 993.

[0157] In some optional implementations, the electronic device can have a depth camera and a color camera. The depth camera can capture a depth image, and the image parameters in the depth image include depth information and infrared light information. The depth information can represent the distance from the object to the depth camera, and the infrared light information can represent the reflection of the infrared light by the object. The color camera can capture a color image, and the color image is used to reflect the color of the object. In addition, the depth image and the color image have the same size.

[0158] The camera 993 is configured to capture still images or videos. An object projects an optical image through a lens to a photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, which is then transmitted to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into a standard image signal in RGB, YUV, or the like. In some embodiments, the electronic device can include one or N cameras 993, where N is a positive integer greater than 1.

[0159] The digital signal processor is configured to process digital signals, including digital image signals. For example, when the electronic device selects a frequency point, the digital signal processor is configured to perform Fourier transform on the frequency point energy, and the like.

[0160] The video codec is configured to compress or decompress digital videos. The electronic device can support one or more video codecs. In this way, the electronic device can play or record videos in multiple encoding formats, such as moving picture experts group (MPEG) 1, MPEG 2, MPEG 3, MPEG 4, and the like.

[0161] The NPU is a neural-network (NN) computing processor that is configured to process input information quickly by referring to the structure of a biological neural network, such as the transmission mode between human brain neurons, and is also configured to continuously self-learn. Through the NPU, the electronic device can implement intelligent cognition applications, such as image recognition, face recognition, voice recognition, text understanding, and the like.

[0162] The following provides an exemplary description of the software structure of the electronic device to which the image processing method provided in the embodiments of the present application is applicable.

[0163] The software system of the electronic device can adopt a layered architecture, an event-driven architecture, a micro-kernel architecture, a micro-service architecture, or a cloud architecture, and the like. The embodiments of the present application exemplarily illustrate the software structure of the electronic device by taking the layered architecture as an example. In some embodiments, the software structure of the electronic device can be the same.

[0164] Figure 10 is a software structure block diagram of the electronic device according to the embodiments of the present application.

[0165] A layered architecture divides software into several layers, each with a clear role and division of labor. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom, the application layer, the framework layer, the system runtime library, and the kernel layer.

[0166] As shown in Figure 10 , the application layer can include a series of application packages. The application packages can include image processing, gallery, navigation, WLAN, Bluetooth, music, video, short message, etc.

[0167] Among them, the image processing application can implement the image processing method provided in the embodiments of the present application.

[0168] In some specific implementations, the image processing application can obtain a reference depth image including depth differences or depth relationships between each pixel point according to image features in the color image, such as texture, pattern, shape, illumination, shadow, spatial relationship, and semantic information, etc. Then, based on the reference depth image, the depth image captured by the depth camera is corrected, that is, the depth information of the pixel points with missing depth information in the depth image is supplemented. In this way, the depth information of the pixel points with missing depth information in the image can be completed, so that image distortion can be avoided, and image quality can be improved.

[0169] In this way, the depth information of the pixel points with missing depth information in the image can be completed, so that image distortion can be avoided, and image quality can be improved.

[0170] The application framework layer provides application programming interfaces (application programming interface, API) and programming frameworks for the applications of the application layer. The application framework layer includes some pre-defined functions.

[0171] As shown in Figure 10 , the application framework layer can include a window manager, a content provider, a phone manager, a resource manager, a notification manager, a view system, etc.

[0172] The window manager is used to manage window programs. The window manager can obtain the size of the display screen, determine whether there is a status bar, lock the screen, and intercept the screen, etc.

[0173] The content provider is used to store and obtain data, and make the data accessible to the application. The data can include video, image, audio, dialed and received calls, browsing history and bookmarks, phonebook, etc.

[0174] The phone manager is used to provide communication functions of the electronic device. For example, management of call state (including call connection, call hang-up, etc.).

[0175] The resource manager provides various resources for the application, such as localized strings, icons, pictures, layout files, video files, etc.

[0176] The notification manager enables the application to display notification information in the status bar, which can be used to convey a message of the notification type, and can automatically disappear after a short stay without user interaction. For example, the notification manager is used to notify the completion of the download, message reminders, etc. The notification manager can also be a notification that appears in the form of a chart or a scroll bar text in the top status bar of the system, such as a notification of an application running in the background, and can also be a notification that appears in the form of a dialog window on the screen. For example, a text information prompt in the status bar, a prompt sound, an electronic device vibration, a flashing indicator light, etc.

[0177] The view system includes visual controls, such as controls for displaying text, controls for displaying pictures, etc. The view system can be used to build an application. A display interface can be composed of one or more views. For example, a display interface including a short message notification icon can include a view for displaying text and a view for displaying pictures.

[0178] The Android runtime includes a core library and a virtual machine. The Android runtime is responsible for scheduling and managing the Android system.

[0179] The core library includes two parts: one part is the function function that the java language needs to call, and the other part is the core library of Android.

[0180] The application layer and the application framework layer run in the virtual machine. The virtual machine executes the java files of the application layer and the application framework layer into binary files. The virtual machine is used to perform the management of the object life cycle, the management of the stack, the management of the thread, the management of the security and the exception, and the garbage collection, etc.

[0181] The system library can include multiple functional modules. For example: surface manager, three-dimensional graphics processing library (such as: OpenGL ES), two-dimensional graphics engine (such as: SGL), media library (Media Libraries), etc.

[0182] The surface manager is used to manage the display subsystem, and provides 2D and 3D layer fusion for multiple applications.

[0183] The three-dimensional graphics processing library is used to realize three-dimensional graphics drawing, image rendering, synthesis, and layer processing, etc.

[0184] The two-dimensional graphics engine is a two-dimensional drawing drawing engine.

[0185] The media library supports a variety of commonly used audio, video format playback and recording, as well as static image files, etc. The media library can support a variety of audio and video encoding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.

[0186] The kernel layer is the layer between hardware and software. The kernel layer at least includes display driver, camera driver, audio driver, sensor driver.

[0187] In some cases, the embodiments disclosed in the present application can be implemented in hardware, firmware, software or any combination thereof.

[0188] The embodiments disclosed in the present application can also be implemented as instructions carried by one or more transitory or non-transitory machine-readable (e.g., computer-readable) storage media, which can be read and executed by one or more processors. For example, the instructions can be distributed over the network or by other computer readable media. Thus, the machine-readable media can include any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computer), including without limitation, a floppy disk, an optical disc, an optical compact disc, a magneto-optical disk, a read-only memory (ROM), a random access memory (RAM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic or optical card, a flash memory, or a tangible, machine-readable storage used in the transmission of information over the Internet using electrical, optical, acoustical or other forms of propagated signals (e.g., carrier waves, infrared signals, digital signals, etc.). Accordingly, the machine-readable media includes any type of machine-readable medium suitable for storing or transmitting electronic instructions or information in a form readable by a machine (e.g., a computer).

[0189] Embodiments of the present application can be implemented as computer programs or program codes executing on programmable system including at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.

[0190] Program code can be applied to input instructions to perform the functions described in the present application and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of this application, a processing system includes any system that has a processor, such as a digital signal processor (DSP), a microcontroller, an application specific integrated circuit, or a microprocessor.

[0191] The program code can be implemented in a high level of programming language or a object-oriented programming language to communicate with processing system. In the case of desiring, the program code can be implemented in assembly or machine language. In fact, the mechanism described in the present application is not limited to the scope of any particular programming language. In any case, the language can be a compiled or interpreted language.

[0192] In some embodiments, the embodiments of the present application further provide a computer readable medium, which stores program codes, when the computer program codes are run on a computer, the computer is enabled to execute the method in the above aspects.

[0193] In some embodiments, the embodiments of the present application further provide a computer program product, which comprises: computer program codes, when the computer program codes are run on a computer, the computer is enabled to execute the method in the above aspects.

[0194] The above introduces the hardware structure that the electronic device can have. It can be understood that the structure illustrated by the embodiments of the present application does not constitute a specific limitation on the electronic device. In other embodiments of the present application, the electronic device can include more or fewer components than the illustration, or combine certain components, or split certain components, or different component arrangement. The components illustrated can be implemented in hardware, software or a combination of software and hardware.

[0195] In the drawings, some structural or methodical features can be shown in a specific arrangement and / or order. However, it should be understood that such specific arrangement and / or order can not be required. Rather, in some embodiments, these features can be arranged in a manner different from that shown in the illustrative drawings. In addition, inclusion of structural or methodical features in a particular figure does not imply that such features are required in all embodiments, and in some embodiments, these features can not be included or can be combined with other features.

[0196] It should be noted that in the examples and descriptions of the present patent, the relationship terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between the entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without more limitation, the element defined by the statement "including one" does not exclude the presence of another identical element in the process, method, article or device including the element.

[0197] While the application has been illustrated and described in connection with certain embodiments thereof, it will be readily apparent to those of ordinary skill in the art that various changes in form and detail can be made therein without departing from the spirit and scope of the application.

Claims

1. An image processing method, characterized in that, Applied to electronic devices, including: Acquire the first depth image and color image of the object captured by the electronic device; Based on the image features of the color image, a reference depth image is obtained, wherein the reference depth image includes the depth relationship between each pixel in the color image; For a first pixel in the first depth image that lacks depth information, the depth information of the first pixel is determined based on the reference depth image to obtain a second depth image.

2. The method according to claim 1, characterized in that, Determining the depth information of the first pixel based on the reference depth image includes: Multiple associated pixels in the reference depth image are identified that are associated with the first pixel. The multiple associated pixels include a first associated pixel corresponding to the first pixel and at least one second associated pixel. At least some of the at least one second associated pixels have corresponding pixels in the first depth image that do not lack depth information. The depth information of the first pixel is determined based on the depth information of the pixels corresponding to the multiple associated pixels in the first depth image.

3. The method according to claim 2, characterized in that, The determination of multiple associated pixels in the reference depth image that are associated with the first pixel includes: At least one second pixel is determined from the first depth image, wherein the second pixel is a pixel that does not lack depth information; The pixel corresponding to the second pixel in the reference depth image is taken as the second associated pixel.

4. The method according to claim 3, characterized in that, The step of determining the depth information of the first pixel based on the depth information of the multiple associated pixels corresponding to the multiple pixels in the first depth image includes: The depth information of the first pixel is obtained based on the depth relationship between the first associated pixel and the second associated pixel in the reference depth image, and the depth information of the second pixel in the first depth image.

5. The method according to claim 2, characterized in that, The determination of multiple associated pixels in the reference depth image that are associated with the first pixel includes: At least one pixel in the reference depth image whose similarity to the depth information of the first associated pixel is higher than a similarity threshold and whose corresponding pixel in the first depth image does not lack depth information is identified as the second associated pixel.

6. The method according to claim 5, characterized in that, The step of determining the depth information of the first pixel based on the depth information of the multiple associated pixels corresponding to the multiple pixels in the first depth image includes: Multiple pixels corresponding to the multiple associated pixels are determined from the first depth image; The depth information of the first pixel is determined based on the depth information of the plurality of pixels.

7. The method according to claim 1, characterized in that, The method further includes: The depth information of each pixel in the second depth image is offset based on the depth offset to obtain a third depth image. The depth offset is either an offset input by the user or determined based on the shooting distance between the electronic device and the object and a first relationship, wherein the first relationship is a one-to-one correspondence between the shooting distance and the depth offset.

8. The method according to claim 7, characterized in that, The method further includes: Based on the scaling factor and coordinate transformation relationship, the third depth image is transformed to obtain the fourth depth image. Wherein, the position of the third pixel representing the first position of the object in the fourth depth image is the same as the position of the fourth pixel representing the first position in the color image, and the coordinate transformation relationship is the transformation relationship between the coordinate system of the depth camera that captures the first depth image and the coordinate system of the color camera that captures the color image. The target image is obtained by fusing the fourth depth image and the color image.

9. The method according to claim 8, characterized in that, The scaling factor is determined in the following way: Determine at least one key point from the second depth image; The overall depth of the second depth image is determined based on the depth information of the at least one key point; The ratio of the total depth to the overall depth is used as the scaling factor, wherein the total depth is the sum of the overall depth of the second depth image and the depth offset.

10. The method according to claim 9, characterized in that, Corresponding to the object being a face, the key points include pixels corresponding to at least one of the following features in the face: Left corner of the eye, right corner of the eye, left corner of the mouth, right corner of the mouth, tip of the nose.

11. An electronic device, characterized in that, include: A memory for storing instructions executed by one or more processors of the electronic device, and a processor, being one of one or more processors of the electronic device, for performing the image processing method according to any one of claims 1-10.

12. A readable storage medium, characterized in that, The readable storage medium stores instructions that, when executed on an electronic device, cause the electronic device to perform the image processing method according to any one of claims 1-10.

13. A computer program product, characterized in that, The computer program product includes computer instructions, which, when executed by an electronic device, execute the computer program code of the image processing method as described in any one of claims 1-10.

Citation Information

Patent Citations

  • Depth image data processing method and mobile terminal

    CN107222737A

  • Method for capturing facial expression features, terminal device and computer readable storage medium

    CN111582121A

  • Depth information acquisition method, binocular camera module, storage medium and electronic equipment

    CN112615993A

  • Image processing method and device and electronic equipment

    CN112837219A

  • Face image restoration method and device, electronic equipment and storage medium

    CN114626993A