Image processing method, electronic device, storage medium and computer program product

CN121397384BActive Publication Date: 2026-09-11HONOR DEVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410941689.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-12
Publication Date
2026-09-11
Estimated Expiration
2044-07-12

AI Technical Summary

Technical Problem

[0003]为解决近距离拍摄导致所拍摄的图像发生畸变,从而导致图像质量降低的问题,本申请实施例提供一种图像处理方法、电子设备、存储介质及计算机程序产品,包括:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121397384B_ABST
    Figure CN121397384B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of image processing, in particular to an image processing method, an electronic device, a storage medium and a computer program product. In the method, an electronic device obtains a depth image and a color image by simultaneously photographing a same object through a depth camera and a color camera. The electronic device can obtain a reference depth image including depth difference values or depth relationships between each pixel point in the corresponding color image by image features such as texture, pattern, shape, illumination, shadow, spatial relationship and semantic information in the color image photographed by the color camera. Then, the electronic device corrects the depth image photographed by the depth camera based on the reference depth image, that is, supplements the depth information of the pixel points with missing depth information in the depth image. In this way, the electronic device can complete the depth information of the pixel points with missing depth information in the depth image, thereby avoiding image distortion and improving image quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to an image processing method, electronic device, storage medium, and computer program product. Background Technology

[0002] For reference Figure 1 The illustration of a selfie scenario assumes that when the distance between the phone 110 and the user's face 120 is relatively short, for example, within 35 to 40 centimeters, the user's shooting posture is more natural and comfortable. However, because the dimensions of areas such as the hairline and facial edges of the user's face 120 are relatively small, they may be below the resolution threshold of the sensor in the depth camera of the phone 110, causing the sensor to be unable to acquire depth information in these tiny areas. Therefore, close-range shooting leads to the loss of some depth information in the image, resulting in a decrease in image quality. Summary of the Invention

[0003] To address the problem of image distortion caused by close-up photography, resulting in reduced image quality, this application provides an image processing method, electronic device, storage medium, and computer program product, including:

[0004] In a first aspect, embodiments of this application provide an image processing method applied to an electronic device, comprising acquiring a first depth image and a color image of an object captured by the electronic device; obtaining a reference depth image based on image features of the color image, wherein the reference depth image includes the depth relationship between each pixel in the color image; and determining the depth information of a first pixel in the first depth image that is missing depth information based on the reference depth image to obtain a second depth image.

[0005] Based on the above scheme, by using a reference depth image obtained from image features in a color image within the same coordinate system as the depth image, the depth image captured by the depth camera can be corrected. This can reduce distortion in certain parts of the object and improve the shooting effect. In other words, by supplementing the depth information of pixels lacking depth information in the first depth image, image distortion can be avoided, and image quality can be improved.

[0006] Furthermore, in this embodiment, only one color image is needed to complete the depth information of pixels missing depth information in the first depth image, which can save the processing resources of electronic devices.

[0007] It can be understood that the first depth image can be an image captured by a time-of-flight camera in an electronic device, and the color image can be an image captured by a color camera in an electronic device. Image features can include texture, pattern, shape, lighting, shadow, spatial relationships, and semantic information, etc.

[0008] It is understood that multiple pixels in a reference depth image may lack depth information. This embodiment of the application uses the first pixel lacking depth information as an example for illustration.

[0009] In some optional implementations of the first aspect, determining the depth information of the first pixel based on a reference depth image includes: determining a plurality of associated pixels in the reference depth image that are associated with the first pixel, wherein the plurality of associated pixels include a first associated pixel corresponding to the first pixel and at least one second associated pixel, and at least some of the second associated pixels have pixels in the first depth image that do not lack depth information; and determining the depth information of the first pixel based on the depth information of the plurality of pixels corresponding to the plurality of associated pixels in the first depth image.

[0010] In some optional implementations of the first aspect, determining multiple associated pixels in the reference depth image that are associated with the first pixel includes: determining at least one second pixel in the first depth image, wherein the second pixel is a pixel that does not lack depth information; and taking the pixel corresponding to the second pixel in the reference depth image as the second associated pixel.

[0011] In some optional implementations of the first aspect, the depth information of the first pixel is determined based on the depth information of multiple pixels corresponding to multiple associated pixels in the first depth image, including: obtaining the depth information of the first pixel based on the depth relationship between the first and second associated pixels in the reference depth image, and the depth information of the second pixel in the first depth image.

[0012] For example, such as Figure 2 As shown, assuming the depth information of the first pixel p1 in the first depth image P1 is missing, while the depth information of pixel p2 is not missing, and the pixel corresponding to the first pixel p1 in the reference depth image P2 is the first associated pixel p1', and the pixel corresponding to the first pixel p2 is the second associated pixel p2'. If the depth ratio between the first associated pixel p1' and the second associated pixel p2' in the reference depth image P2 is 5:1, then 5 times the depth value of the second pixel p2 in the depth image P1 can be determined as the depth value of the first pixel p1.

[0013] In some alternative implementations of the first aspect, multiple associated pixels of the first pixel are determined from the reference depth image, and at least one pixel in the first depth image whose similarity to the depth information of the first associated pixel is higher than a similarity threshold and whose corresponding pixel in the first depth image does not lack depth information is determined as the second associated pixel.

[0014] In some optional implementations of the first aspect, the depth information of the first pixel is determined based on the depth information of multiple pixels corresponding to multiple associated pixels in the first depth image, including: determining multiple pixels corresponding to multiple associated pixels from the first depth image; and determining the depth information of the first pixel based on the depth information of multiple pixels.

[0015] For example, such as Figure 3 As shown, suppose the depth information of the first pixel p1 in depth image P1 is missing, and the first associated pixel of the first pixel p1 in reference depth image P2 is p1'. If in reference depth image P2, the depth ratio between the second associated pixel p2' and the second associated pixel p1' is 1:1, and the depth ratio between the second associated pixel p3' and the second associated pixel p1' is 1:1.2, and the pixel corresponding to p2' in depth image P1 is the second pixel p2, and the second pixel corresponding to the second associated pixel p3' is pixel p3, then the depth information of the first pixel p1 can be determined based on the depth values ​​of the second pixel p2 and the second pixel p3. For example, the average of the depth values ​​of the second pixel p2 and the second pixel p3 can be used as the depth value of the first pixel p1.

[0016] In some optional implementations of the first aspect, the method further includes: biasing the depth information of each pixel in the second depth image based on the depth offset to obtain a third depth image, wherein the depth offset is an offset input by the user, or determined based on the shooting distance between the electronic device and the object and a first relationship, wherein the first relationship is a one-to-one correspondence between the shooting distance and the depth offset.

[0017] It's understandable that after obtaining the second depth image, the electronic device can first perform an overall translation of the second depth image based on the depth offset, projecting the second depth image to a position closer to the depth camera, or projecting it to a position farther away from the depth camera, to obtain the third depth image. This can further reduce distortion in certain parts of the object, further improving the shooting effect.

[0018] For example, if an electronic device detects that the distance between the object and the electronic device is 30 centimeters, it can determine that the depth offset is 30 centimeters based on the pre-stored one-to-one correspondence between shooting distance and depth offset.

[0019] In some optional implementations of the first aspect, the method further includes: transforming the third depth image based on a scaling factor and a coordinate transformation relationship to obtain a fourth depth image, wherein the position of the third pixel representing the first position of the object in the fourth depth image is the same as the position of the fourth pixel representing the first position in the color image, and the coordinate transformation relationship is the transformation relationship between the coordinate system of the depth camera that captured the first depth image and the coordinate system of the color camera that captured the color image; fusing the fourth depth image and the color image to obtain the target image.

[0020] In this way, the position of the third pixel representing the first position of the object in the fourth depth image can correspond one-to-one with the position of the fourth pixel representing the first position in the color image, thereby further reducing distortion in parts of the object and further improving the shooting effect.

[0021] In some alternative implementations of the first aspect, the scaling factor is determined by: determining at least one key point from the second depth image; determining the overall depth of the second depth image based on the depth information of the at least one key point; and using the ratio of the total depth to the overall depth as the scaling factor, wherein the total depth is the sum of the overall depth and the depth offset of the second depth image.

[0022] In some alternative implementations of the first aspect, the object is a face, and the key points include pixels corresponding to at least one of the following features of the face: left corner of the eye, right corner of the eye, left corner of the mouth, right corner of the mouth, and tip of the nose.

[0023] In some specific implementations, the coordinates of multiple facial key points in the second depth image can be obtained in the color image coordinate system. Based on these coordinates, the centroid coordinates of the multiple facial key points can be determined. Then, the depth information corresponding to the centroid coordinates in the second depth image can be determined as the overall depth *d* of the second depth image. Next, the scaling factor can be determined based on the ratio of the total depth of the complete depth map and the depth offset *offset* to the overall depth of the complete depth map.

[0024] For example, the scaling factor can be expressed using the following formula (1):

[0025] α=(d+offset) / d Formula (1)

[0026] Where α represents the scaling factor, d represents the overall depth corresponding to the third depth image, and offset represents the depth offset.

[0027] In this embodiment, noise is introduced during the overall translation of the second depth image and the reprojection of the third depth image, affecting the clarity of the final generated target image. Therefore, after obtaining the target image, an adaptive feature fusion algorithm can be used to enhance the details of the target image, for example, to improve the clarity of the target image, resulting in a high-clarity target image.

[0028] Furthermore, because the distance between the face and the camera (such as the depth camera and color camera mentioned earlier) is smaller than the distance between the ear and the camera, the face observed by the camera will appear wider than the user's actual face, and the ear will appear narrower than the user's actual ear. Therefore, the face will partially obscure the ear. Consequently, when the camera captures the user's face, it may fail to capture image information about the ear, such as color and depth information, resulting in a partially obscured ear in the captured image.

[0029] Therefore, image inpainting algorithms can be used to complete images after detail enhancement. These algorithms can include mask-aware transformers for large-hole image inpainting (MAT), generative landmark-guided face inpainttors (LaFIn), and DeepFill v2.

[0030] In a second aspect, this application provides an electronic device, comprising: a memory for storing instructions executed by one or more processors of the electronic device, and a processor, which is one of the one or more processors of the electronic device, for executing the image processing method mentioned in the first aspect or any one of the first aspects of this application.

[0031] Thirdly, this application provides a readable storage medium storing instructions that, when executed on an electronic device, cause the electronic device to perform the image processing method mentioned in the first aspect or any one of the first aspects of this application.

[0032] Fourthly, embodiments of this application provide a computer program product, which includes computer instructions. When executed by an electronic device, the electronic device executes the computer program code of the image processing method mentioned in the first aspect or any one of the first aspects of this application. Attached Figure Description

[0033] Figure 1 A schematic diagram of an application scenario is shown;

[0034] Figure 2 According to some embodiments of this application, a schematic diagram is shown of a depth image P1 and a reference depth image P2 corresponding to each other.

[0035] Figure 3 According to some embodiments of this application, a schematic diagram is shown corresponding to another depth image P1 and a reference depth image P2;

[0036] Figure 4A A schematic diagram of parallax is shown according to some embodiments of this application;

[0037] Figure 4B According to some embodiments of this application, a schematic diagram showing the correspondence between a depth image and a color image is illustrated;

[0038] Figure 4C A schematic diagram of a distorted image is shown according to some embodiments of this application;

[0039] Figure 5 According to some embodiments of this application, a schematic flowchart of an image processing method is shown;

[0040] Figure 6 According to some embodiments of this application, a schematic diagram of reprojection and canvas scaling is shown;

[0041] Figure 7 According to some embodiments of this application, a flowchart of another image processing method is shown;

[0042] Figure 8 A schematic diagram of a calibration image is shown according to some embodiments of this application;

[0043] Figure 9 According to some embodiments of this application, a schematic diagram of the hardware structure of an electronic device is shown;

[0044] Figure 10 According to some embodiments of this application, a schematic diagram of the software structure of an electronic device is shown. Detailed Implementation

[0045] The embodiments of this application include, but are not limited to, an image processing method, an electronic device, a storage medium, and a computer program product.

[0046] It is understood that the image processing methods mentioned in the embodiments of this application can be applied to electronic devices. These electronic devices can be referred to as terminals, user equipment (UE), mobile stations (MS), mobile terminals (MT), etc. In some specific implementations, the electronic device can be a mobile phone, a tablet computer, a wearable device with wireless communication capabilities, a computer with wireless transceiver capabilities, a virtual reality (VR) device, an augmented reality (AR) device, etc., and this application does not limit the scope of the application to these devices.

[0047] For ease of explanation, the following description uses a mobile phone as an example. In some specific implementations, a mobile phone may have a depth camera and a color camera. The depth camera captures depth images and infrared images corresponding to pixels at the same location. The image parameters in the depth image include depth information, which represents the distance from the object to the depth camera. The image parameters in the infrared image include infrared light information, which represents the object's reflection of infrared light. The color camera captures color images, which reflect the color of the object.

[0048] As mentioned above, Figure 1 As shown, because the dimensions of areas such as the hairline and facial edges on the user's face (120) are relatively small, they may be below the resolution threshold of the depth camera sensor in the phone (110), causing the sensor to be unable to acquire depth information from these tiny areas. Therefore, close-up shooting can lead to distortion due to the loss of some depth information in the image, resulting in reduced image quality.

[0049] Therefore, to solve the above problems, this application provides an image processing method. In this method, an electronic device simultaneously captures images of the same object using a depth camera and a color camera to obtain a depth image and a color image. The electronic device can obtain a reference depth image, including the depth difference or depth relationship between pixels in the corresponding color image, based on image features in the color image captured by the color camera, such as texture, pattern, shape, lighting, shadow, spatial relationship, and semantic information. Then, the electronic device corrects the depth image captured by the depth camera based on the reference depth image, that is, it supplements the depth information of pixels missing depth information in the depth image.

[0050] In this way, electronic devices can complete the depth information of pixels that are missing depth information in a depth image, thereby avoiding image distortion and improving image quality.

[0051] Specifically, such as Figure 2 As shown, in some embodiments, it is assumed that the depth information of pixel p1 in depth image P1 is missing, while the depth information of pixel p2 is not missing. In the reference depth image P2, the pixel corresponding to pixel p1 is pixel p1', and the pixel corresponding to pixel p2 is pixel p2'. If the depth ratio between pixel p1' and pixel p2' in the reference depth image P2 is 5:1, then 5 times the depth value of pixel p2 in depth image P1 can be determined as the depth value of pixel p1.

[0052] like Figure 3 As shown, in some embodiments, it is assumed that the depth information of pixel p1 in depth image P1 is missing, while the corresponding pixel in reference depth image P2 is p1'. If in reference depth image P2, the depth ratio between pixel p2' and pixel p1' is 1:1, and the depth ratio between pixel p3' and pixel p1' is 1:1.2, and in depth image P1, the corresponding pixel for pixel p2' is pixel p2, and the corresponding pixel for pixel p3' is pixel p3, then the depth information of pixel p1 can be determined based on the depth values ​​of pixel p2 and pixel p3. For example, the average of the depth values ​​of pixel p2 and pixel p3 can be used as the depth value of pixel p1.

[0053] For example, in a user selfie scenario, for each pixel in the hair edge region that lacks depth information, the electronic device can determine the depth information of each pixel in the hair edge region based on the depth information of multiple pixels in the hair middle region that has depth information.

[0054] In some embodiments, as described above, the electronic device may include a depth camera and a color camera, which are typically located at different locations within the electronic device. Because when the same object is photographed from different locations, the position of the object in the depth image captured by the depth camera and the position of the object in the color image captured by the color camera differ, as... Figure 4A The discrepancy shown (hereinafter referred to as parallax) means that when fusing depth and color images, pixels at the same location in the depth and color images may not correspond one-to-one, resulting in image distortion.

[0055] For example, in a user selfie scenario, such as Figure 4BAs shown, region A in the face is located at position A1 in the depth image and position A2 in the color image. When fusing the depth and color images, the pixel at position A1 in the depth image will be fused with the pixel at position B2 in the color image, and the pixel at position A2 in the color image will be fused with the pixel at position B1 in the depth image. For example, fusing a pixel of the face in the depth image with a pixel of the ear in the color image will result in an enlarged nose and a stretched and widened face, causing image distortion.

[0056] It is necessary to understand that Figure 4B This embodiment is only intended to highlight situations where pixels at the same location in the depth image and color image cannot be matched one-to-one. In practical applications, the position of the same point on an object in the depth image and its position in the color image may be smaller than each other. Figure 4B deviation shown.

[0057] Therefore, before correcting the depth image captured by the depth camera, the electronic device can transform the depth image and color image, obtained simultaneously by the depth camera and color camera capturing the same object, into the same coordinate system. Within this coordinate system, the conversion relationship between different pixels representing the same point on the object in the depth image and color image is adjusted, ensuring a one-to-one correspondence between pixels at the same location in the depth image and color image. In this way, the electronic device can correct the depth image captured by the depth camera using a reference depth image obtained from the image features of the color image in the same coordinate system as the depth image. This reduces distortion in parts of the object and improves the shooting effect.

[0058] Furthermore, after correcting the depth image captured by the depth camera, the electronic device can first perform a global translation of the depth image supplemented with depth information (hereinafter referred to as the second depth image) based on the depth offset, projecting the second depth image to a position closer to the depth camera, or projecting it to a position farther away from the depth camera, to obtain the third depth image. Next, based on the depth information of key points in the third depth image and the depth offset, a scaling factor can be determined. The electronic device can then perform transformation processing on the third depth image using the scaling factor and coordinate transformation relationships, such as enlarging or reducing the third depth image, to obtain the fourth depth image. In the fourth depth image, the position of the third pixel representing the first position of the object is the same as the position of the fourth pixel representing the first position in the color image, and the coordinate transformation relationship is between the coordinate system of the depth camera and the coordinate system of the color camera. The fourth depth image and the color image are then fused to obtain the target image. The quality of the target image is higher than the quality of the fused image obtained by fusing the first depth image and the color image.

[0059] Understandably, in some optional implementations, the depth offset can be a depth offset input by the user to the electronic device in real time, or it can be determined based on a one-to-one correspondence between the shooting distance and the depth offset pre-stored by the electronic device. For example, if the electronic device detects that the distance between the object and the electronic device is 30 centimeters, the depth offset can be determined to be 30 centimeters based on the pre-stored one-to-one correspondence between the shooting distance and the depth offset.

[0060] It is understandable that, in some alternative implementations, the scaling factor can be determined in the following way:

[0061] The electronic device can acquire depth information of multiple key points in the second depth image and determine the overall depth d corresponding to the second depth image based on the depth information of multiple key points. Key points can be points whose relative positions do not change when the depth camera takes close-up shots of an object from different angles. For example, in a user selfie scenario, key points on the user's face can include the left corner of the eye, right corner of the eye, left corner of the mouth, right corner of the mouth, and tip of the nose, also known as facial landmarks. The electronic device can determine the scaling factor based on the ratio of the total depth of the overall depth and depth offset corresponding to the second depth image to the overall depth of the second depth image. For example, the scaling factor can be expressed using the following formula (1):

[0062] α=(d+offset) / d Formula (1)

[0063] Where α represents the scaling factor, d represents the overall depth corresponding to the second depth image, and offset represents the depth offset.

[0064] In some specific implementations, the electronic device can acquire the coordinates of multiple face landmarks in the color image coordinate system, and determine the centroid coordinates of the multiple face landmarks based on their coordinates. Then, the electronic device can determine the depth information corresponding to the centroid coordinates in the second depth image as the overall depth d of the second depth image.

[0065] In addition, due to the effect of "near is larger and far is smaller", the size of different parts observed by the camera will change relative to the actual size of the part. Therefore, the parts that appear larger than their actual size observed by the camera will block the parts that appear smaller than their actual size.

[0066] For example, in a selfie scenario, because the distance between the user's face and the camera (such as the depth camera and color camera mentioned earlier) is smaller than the distance between the ear and the camera, the face observed by the camera will appear wider than the user's actual face, and the ears will appear narrower than the user's actual ears. Therefore, the face will partially obscure the ears. Consequently, when the camera captures the user's face, it may fail to capture some image information about the ears, such as color and depth information, resulting in an image that appears as shown below. Figure 4C The ear shown is partially missing.

[0067] Therefore, after fusing the fourth depth image and the color image to obtain the target image, image enhancement and image completion algorithms can be used to enhance and complete the details of the missing parts in the target image. This can further improve the image quality.

[0068] For example, in a user selfie scenario, the adaptive spatial feature fusion (ASFFNet) algorithm can be used to enhance the details of the target image, and the image inpainting algorithm can be used to complete the image after the details have been enhanced.

[0069] The image processing methods mentioned in the embodiments of this application will be described in detail below.

[0070] It is understood that the image processing methods mentioned in the embodiments of this application can be applied to scenarios such as image enhancement, taking photos, recording videos, and video calls. For example, in the image enhancement scenario, the image processing method can be used to process images in a gallery; as another example, in the photo taking scenario, the image processing method can be used to process images captured by the camera in real time; as yet another example, in the video recording or video call scenario, the image processing method can be used to process each frame or a portion of the images. The embodiments of this application do not limit the specific application scenarios of the image processing methods.

[0071] For ease of explanation, the following description uses the application of this image processing method to a user selfie scenario. In a user selfie scenario, the object may include the user's face.

[0072] like Figure 5 The diagram illustrates a flowchart of an image processing method. This image processing method can be executed by an electronic device, such as a mobile phone 100. Specifically, the image processing method may include:

[0073] First, acquire the RGB image, incomplete depth map, and camera parameters.

[0074] It is understandable that an RGB image is a color image obtained by the color camera in a mobile phone 100 from taking a picture of the user's face, also known as a three-primary-color (red, green, blue, RGB) image.

[0075] It's understandable that a mobile phone can capture depth images using, for example, a time-of-flight (TOF) camera. However, because light reflected from the user's face can travel through different paths back to the TOF camera, the TOF camera's time-of-flight measurement may be inaccurate, potentially resulting in missing depth information for some pixels in the captured depth image. Furthermore, in close-up shots, areas such as the hairline and facial edges are relatively small and may be below the resolution threshold of the TOF camera's sensor, preventing the sensor from acquiring depth information from these tiny areas. Therefore, the depth image captured by the TOF camera, lacking some depth information, is an incomplete depth map.

[0076] It can be understood that the transformation relationship between the depth camera coordinate system and the color camera coordinate system can include translation and rotation relationships, which can be specifically represented by the translation matrix and rotation matrix between the depth camera coordinate system and the color camera coordinate system.

[0077] In some specific implementations, the translation and rotation matrices that transform the depth camera coordinate system to coincide with the color camera coordinate system can be collectively referred to as the extrinsic parameters of the depth camera. The transformation relationship between the depth camera coordinate system and the depth image coordinate system can be called the intrinsic parameters of the depth camera, which may include parameters such as the depth camera's focal length and principal point coordinates. Similarly, the transformation relationship between the color camera coordinate system and the color image coordinate system can be called the intrinsic parameters of the color camera, which may include parameters such as the color camera's focal length and principal point coordinates. In total, the extrinsic parameters, intrinsic parameters, and intrinsic parameters of the depth camera can be collectively referred to as camera parameters.

[0078] Then, RGBD alignment is performed on the RGB image and the incomplete depth map.

[0079] It's understandable that an RGB image and a missing depth map can be aligned using RGBD to obtain an aligned depth map. Here, RGB can refer to the three primary colors, and D can refer to depth information.

[0080] As mentioned earlier, depth cameras and color cameras are located in different positions in electronic devices. Because there is a parallax between the position of a user's face in the depth image captured by the depth camera and the position of the user's face in the color image captured by the color camera when the user's face is photographed from different positions, pixels at the same location in the depth and color images may not correspond one-to-one, resulting in image distortion.

[0081] In some alternative implementations, RGBD alignment can be performed on the RGB image and the missing depth map based on the color image and camera parameters, ensuring a one-to-one correspondence between pixels at the same location in the missing depth image and the color image. This can reduce distortion in parts of the object and improve the shooting effect.

[0082] Specifically, the depth images and color images captured by the depth camera and color camera can be converted to the same coordinate system. In this coordinate system, the conversion relationship between different pixels representing the same point on the object in the depth image and color image can be adjusted so that the pixels at the same position in the depth image and color image correspond one-to-one.

[0083] Next, depth estimation is performed on the RGB image.

[0084] Understandably, monocular depth estimation algorithms, such as the DepthAnything algorithm, can be used to estimate the depth of RGB images and obtain a reference depth image that includes the depth differences or depth relationships between each pixel.

[0085] In the specific estimation process of some DepthAnything algorithms, depth estimation of RGB images can be performed based on image features in the RGB image, such as texture, pattern, shape, lighting, shadow, spatial relationship and semantic information.

[0086] Furthermore, depth completion is performed on the aligned depth map.

[0087] It is understandable that since the reference depth image is obtained by depth estimation of the RGB image, the pixels at the same position in the reference depth image and the RGB image correspond one-to-one. Therefore, a complete depth map can be obtained by performing depth completion on the aligned depth map based on the transformation relationship between the depth camera coordinate system and the color camera coordinate system, and based on the reference depth image.

[0088] The transformation matrix can include rotation, translation, and scaling parameters. Based on this transformation matrix, pixels in the reference depth image can be mapped to the color camera coordinate system to coincide with the corresponding pixels with depth information in the aligned depth image. The corresponding pixels in the reference depth image and the aligned depth image can represent the same point on the user's face. Thus, based on this transformation matrix and the pixels in the reference depth image, the coordinates of the corresponding pixels in the aligned depth image in the color camera coordinate system can be calculated, thereby obtaining the depth information of pixels lacking depth information in the aligned depth image.

[0089] In some alternative implementations, for pixels lacking depth information in the aligned depth map, a first associated pixel corresponding to the first pixel lacking depth information can be determined from the reference depth image. Then, based on the depth relationships between pixels in the reference depth image, such as depth ratios, a second associated pixel whose depth ratio to the first associated pixel falls within a certain range (e.g., 1:1 to 1.5), specifically 1:1 or 1:1.2, can be determined from the reference depth image. Next, a second pixel corresponding to the second associated pixel is determined from the aligned depth map. Finally, based on the depth information of the second pixel, the depth information of the first pixel lacking depth information can be determined.

[0090] Furthermore, the entire depth map is translated.

[0091] It is understandable that after obtaining the complete depth map, the complete depth map can be translated as a whole based on the depth offset to obtain the depth map at the target distance.

[0092] In some alternative implementations, the depth offset can be a depth offset input by the user to the electronic device in real time, or it can be determined based on a one-to-one correspondence between the shooting distance and the depth offset pre-stored by the electronic device. For example, if the electronic device detects that the distance between the object and the electronic device is 30 centimeters, the depth offset can be determined to be 30 centimeters based on the pre-stored one-to-one correspondence between the shooting distance and the depth offset.

[0093] In some specific implementation methods, such as Figure 6 As shown, a depth offset can be added to the depth value of each pixel in the complete depth map to achieve overall translation of the complete depth map and obtain the depth map at the target distance.

[0094] Furthermore, the depth map at the target distance is reprojected and the canvas is scaled.

[0095] It's understandable that after obtaining the depth map at the target distance, it can be reprojected onto the color image coordinate system to fuse the depth and color information. Since the entire depth map has been translated, the positions of each pixel in the depth map at the target distance within the depth camera coordinate system change relative to their positions in the complete depth image. Therefore, the depth map at the target distance can be scaled based on a scaling factor. Then, the scaled depth map and the color image are fused to obtain the RGB image at the target distance.

[0096] It is understandable that, in some alternative implementations, the scaling factor can be determined in the following way:

[0097] Obtain the coordinates of multiple face landmarks in the color image coordinate system from the complete depth map, and determine the centroid coordinates of the multiple face landmarks based on their coordinates in the color image coordinate system. Then, the depth information corresponding to the centroid coordinates in the second depth image can be determined as the overall depth d of the second depth image. Next, the scaling factor can be determined based on the ratio of the total depth of the overall depth and the depth offset corresponding to the complete depth map to the overall depth corresponding to the complete depth map. For example, the scaling factor can be expressed using the above formula (1).

[0098] For example, the depth image at the corresponding position before adjustment can be reprojected and the canvas scaled to adjust the depth image at the corresponding position before adjustment to the corresponding position after adjustment.

[0099] Furthermore, the RGB image at the target distance is enhanced in detail.

[0100] It's understandable that noise is introduced into the depth map during the overall translation of the complete depth map and the reprojection of the depth map at the target distance, thus affecting the clarity of the final generated RGB image at the target distance. Therefore, after obtaining the RGB image at the target distance, an adaptive feature fusion algorithm can be used to enhance its details, for example, by improving its clarity to obtain a high-resolution RGB image. This high-resolution RGB image does not include the ears.

[0101] Furthermore, face completion is performed on the high-resolution RGB image to obtain an RGB image including the ears.

[0102] As mentioned earlier, it's understandable that because the distance between the face and the camera (such as depth and color cameras) is smaller than the distance between the ear and the camera, the face observed by the camera appears wider than the user's actual face, and the ear appears narrower than the user's actual ear. Therefore, the face partially obscures the ear. Consequently, when the camera captures the user's face, it may fail to capture image information about the ear, such as color and depth information, resulting in a partially obscured ear in the captured image.

[0103] Therefore, image inpainting algorithms can be used to complete the image after detail enhancement. These algorithms can include the MAT algorithm, the LaFIn algorithm, and the DeepFill v2 algorithm.

[0104] The following is combined with Figure 7 right Figure 5 Further explanation will be provided for some of the steps in the process. For example... Figure 7 The diagram illustrates a flowchart of another image processing method. This image processing method can be executed by an electronic device, such as a mobile phone 100. Specifically, the image processing method may include:

[0105] 701: Acquire the first depth image and color image;

[0106] It can be understood that the color image can be the image obtained by the color camera in the mobile phone 100 capturing the user's face, also known as a three-primary-color image. The first depth image can be a partial depth map with missing depth information captured by the TOF camera in the mobile phone 100.

[0107] 702: Based on the image features of the color image, a reference depth image is obtained, wherein the reference depth image includes the depth relationship between each pixel in the color image.

[0108] Understandably, monocular depth estimation algorithms, such as the DepthAnything algorithm, can be used to estimate the depth of a color image, resulting in a reference depth image that includes the depth differences or depth relationships between each pixel.

[0109] In the specific estimation process of some DepthAnything algorithms, depth estimation of color images can be performed based on image features in the color image, such as texture, pattern, shape, lighting, shadow, spatial relationship and semantic information.

[0110] 703: For pixels in the first depth image that lack depth information, the depth information of the pixels is determined based on the reference depth image to obtain the second depth image.

[0111] It is understandable that methods for determining pixel depth information based on a reference depth image can include mapping the reference depth image and the depth image to the same coordinate system, such as a spatial coordinate system or a planar coordinate system. The spatial coordinate system can be the color camera coordinate system and the depth camera coordinate system, while the planar coordinate system can be the color image coordinate system and the depth image coordinate system. Then, based on the coordinates of each pixel in the reference depth image within this coordinate system, the depth information of pixels missing from the depth image is supplemented.

[0112] It's important to understand that the color camera coordinate system, depth camera coordinate system, color image coordinate system, and depth image coordinate system can be converted to each other using transformation relationships between these systems. In some specific implementations, these transformation relationships can be obtained by calibrating the camera; the specific calibration method will be discussed later. Figure 8 The details are elaborated in the previous section, and will not be repeated here to avoid repetition.

[0113] The following example illustrates the specific implementation method for supplementing the depth information of pixels missing in the depth image, using the mapping of the reference depth image and the depth image to the color camera coordinate system and the color image coordinate system, respectively.

[0114] In some specific implementations, the depth image can be mapped to the color camera coordinate system based on the transformation relationship between the depth image coordinate system and the depth camera coordinate system, as well as the transformation relationship between the depth camera coordinate system and the color camera coordinate system. Furthermore, a reference depth image can also be mapped to the color camera coordinate system based on the transformation relationship between the color image coordinate system and the color camera coordinate system.

[0115] Then, the coordinates of multiple key points in the depth image in the color camera coordinate system can be obtained, as well as the coordinates of the corresponding multiple key points in the reference depth image in the color camera coordinate system, and the transformation matrix between the depth image and the reference depth image can be determined.

[0116] The transformation matrix can include rotation, translation, and scaling parameters. Based on this transformation matrix, pixels in the reference depth image can be mapped to the color camera coordinate system so that they coincide with the corresponding pixels in the depth image that contain depth information. The reference depth image and the corresponding pixels in the depth image can represent the same point on the object. Thus, based on this transformation matrix and the pixels in the reference depth image, the coordinates of the corresponding pixels in the depth image in the color camera coordinate system can be calculated, thereby obtaining the depth information of pixels in the depth image that lack depth information.

[0117] In other alternative implementations, the depth image can be mapped to the color image coordinate system based on the transformation relationships between the depth image coordinate system and the depth camera coordinate system, between the depth camera coordinate system and the color camera coordinate system, and between the color camera coordinate system and the color image coordinate system. This allows the pixels corresponding to each pixel in the depth image to be determined from the reference depth image.

[0118] For pixels in a depth image that lack depth information, a first associated pixel corresponding to the pixel lacking depth information can be determined from a reference depth image. Then, based on the depth relationships between pixels in the reference depth image, such as depth ratios, a second associated pixel whose depth ratio to the first associated pixel falls within a certain range (e.g., 1:1 to 1.5), specifically 1:1 or 1:1.2, can be identified. Next, a third associated pixel corresponding to the second associated pixel is determined from the depth image. Finally, based on the depth information of the third associated pixel, the depth information of the pixel lacking depth information can be determined.

[0119] The specific implementation method for camera calibration is described below.

[0120] As mentioned earlier, the transformation relationship between the depth camera coordinate system and the color camera coordinate system can include translation and rotation relationships, specifically represented by translation and rotation matrices between the two coordinate systems. In some specific implementations, the translation and rotation matrices that transform the depth camera coordinate system to coincide with the color camera coordinate system can be collectively referred to as the extrinsic parameters of the depth camera. The transformation relationship between the depth camera coordinate system and the depth image coordinate system can be referred to as the intrinsic parameters of the depth camera, which can include parameters such as the focal length and principal point coordinates. Similarly, the transformation relationship between the color camera coordinate system and the color image coordinate system can be referred to as the intrinsic parameters of the color camera, which can include parameters such as the focal length and principal point coordinates.

[0121] In some specific implementations, the transformation relationship between the depth camera coordinate system and the depth image coordinate system can be referred to the following formula (2):

[0122]

[0123] in, K represents the coordinates in the depth image coordinate system. d The intrinsic parameter matrix representing the depth camera coordinate system. This represents the coordinates in the depth camera coordinate system.

[0124] The transformation relationship between the color camera coordinate system and the color image coordinate system can be referred to the following formula (3):

[0125]

[0126] in, K represents the coordinates in the color image coordinate system. r The intrinsic parameter matrix representing the coordinate system of the color camera. This represents the coordinates in the color camera coordinate system.

[0127] The transformation relationship between the depth camera coordinate system and the color camera coordinate system can be referred to the following formula (4):

[0128]

[0129] in, Represents the coordinates in the color camera coordinate system. Let R represent the coordinates in the depth camera coordinate system, R represent the rotation matrix for transforming the depth camera coordinate system to the color camera coordinate system, and T represent the translation matrix for transforming the depth camera coordinate system to the color camera coordinate system. K represents the coordinates in the color image coordinate system. r The intrinsic parameter matrix representing the coordinate system of the color camera. K represents the coordinates in the depth image coordinate system. d The intrinsic parameter matrix represents the coordinate system of the depth camera.

[0130] The following section will take the calibration of the intrinsic parameters of a color camera as an example to provide a detailed introduction to the calibration process of depth cameras and color cameras.

[0131] like Figure 8 The diagram shows a schematic of a calibration image. The calibration image may include multiple ellipses and has four corner points a, b, c, and d.

[0132] It's understandable that a spatial coordinate system can be established with the top left corner of the calibrated image as the origin, the horizontal direction to the right from the origin as the first direction X, the vertical downward direction along the direction of gravity as the second direction Y, and the direction perpendicular to the paper and outward as the third direction Z. It's understood that the first direction X, the second direction Y, and the third direction Z are mutually perpendicular. Furthermore, by using a shape detection algorithm to determine the center point of each ellipse, the coordinates of the center point of each ellipse in the spatial coordinate system can be determined.

[0133] In the specific calibration process, a color camera can be used to take calibration images from different positions and angles, and a shape detection algorithm can be used to determine the center point of each ellipse in the color calibration image taken by the color camera.

[0134] In some optional implementations, the coordinates of the center points of all ellipses in the color calibration image within the color image coordinate system can be input into a convex hull calculation function. This function can then be used to determine the outermost ellipse. The convex hull calculation function can be any OpenCV function that calculates the convex hull. The color image coordinate system can be a coordinate system established with the top-left corner of the color calibration image as the origin, the horizontal direction to the right as the first direction (X), and the vertical direction downwards as the second direction (Y).

[0135] Then, the four corner points of the color calibration image can be determined from the outermost ellipse. For example, the four corner points of the color calibration image can be determined from multiple ellipses in the color calibration image captured by the color camera based on the angles formed by the lines between the center point of the outermost ellipse and adjacent points. For example, the four corner points of the calibration image can be determined based on the four smallest angles formed by the lines between the center point of the outermost ellipse and adjacent points.

[0136] In other alternative implementations, one can select any one of the multiple outermost ellipses in the color calibration image captured by the color camera, traverse the remaining outermost ellipses in the color calibration image, and determine the distance between the center point of the remaining ellipses and the center point of the selected ellipse. The center points of the ellipses with the two largest distances are then identified as the adjacent points of the selected ellipse's center point. Based on the angles formed by the lines between the center point of each ellipse and its adjacent points, the four corner points of the calibration image are determined from the multiple ellipses in the color calibration image.

[0137] For example, the center point of an ellipse whose center point is connected to the line between itself and its adjacent points at an angle less than 180° can be designated as a corner point of the color calibration image. Specifically... Figure 8 As shown, the angle formed by the lines between the center point a of the ellipse and the adjacent points e and f is 90°, and the angle formed by the lines between the center point e of the ellipse and the adjacent points a and g is 180°. Therefore, the center point a of the ellipse can be determined as the corner point of the color calibration image.

[0138] Next, affine transformations can be used to sort all the ellipses in the calibration image by row and column to obtain the correct order of all the ellipses in the calibration image.

[0139] Then, based on the OpenCV calibration function, the coordinates of the center points of all ellipses in the calibration image in the spatial coordinate system, the coordinates of the center points of all ellipses in the color calibration image in the color image coordinate system, and the position of the color camera coordinate system in the spatial coordinate system, the intrinsic parameters of the color camera coordinate system can be determined. The color camera coordinate system can be a coordinate system established with the optical center of the color camera as the origin, the direction horizontal to the color camera's imaging plane as the first direction X, and the direction perpendicular to the color camera's imaging plane as the second direction Y.

[0140] Similarly, the intrinsic and extrinsic parameters of the depth camera can be determined based on the above methods. To avoid repetition, they will not be elaborated here.

[0141] In some implementations that determine the extrinsic parameters of a color camera, the transformation relationship between the color image coordinate system and the world coordinate system can be determined. The derivation of this transformation relationship can be referenced in the following calculation formula:

[0142]

[0143] in, d represents the coordinates in the color image coordinate system. x d represents the pixel increment along the x-axis in the color image coordinate system. yThis represents the pixel increment along the y-axis in the color image coordinate system. u0 and v0 are the offsets of the color camera. f represents the focal length. R can represent the rotation matrix, and T can represent the translation matrix. f represents the coordinates in the world coordinate system. x f represents the focal length along the x-axis of the image. y This represents the focal length along the y-axis of the image.

[0144] It is understood that the image processing method provided in this application embodiment can be applied to electronic devices. The hardware structure of the electronic device to which the image processing method provided in this application embodiment is applicable will be described exemplarily below.

[0145] like Figure 9 As shown, the electronic device 900 may include a processor 910, an external memory interface 920, an internal memory 921, a universal serial bus (USB) interface 930, a charging management module 940, a power management module 941, a battery 942, antenna 1, antenna 2, a mobile communication module 950, a wireless communication module 960, an audio module 970, a speaker 970A, a receiver 970B, a microphone 970C, a sensor module 980, buttons 990, a motor 991, an indicator 992, a camera 993, a display screen 994, and a subscriber identification module (SIM) card interface 995, etc.

[0146] It is understood that the structures illustrated in the embodiments of the present invention do not constitute a specific limitation on the electronic device. In other embodiments of this application, the electronic device 900 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0147] Processor 910 may include one or more processing units, such as application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU). These different processing units may be independent devices or integrated into one or more processors.

[0148] In some alternative implementations, the processor 910 may execute the image processing methods mentioned in the embodiments of this application.

[0149] Specifically, the electronic device 900 obtains a reference depth image, including depth differences or depth relationships between pixels, based on image features in the color image, such as texture, pattern, shape, lighting, shadow, spatial relationships, and semantic information. Then, based on the reference depth image, the depth image captured by the depth camera is corrected, that is, the depth information of pixels missing depth information in the depth image is supplemented. In this way, the depth information of pixels lacking depth information in the image can be completed, thereby avoiding image distortion and improving image quality.

[0150] The controller can generate operation control signals based on the instruction opcode and timing signals to complete the control of instruction fetching and execution.

[0151] The processor 910 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 910 is a cache memory. This memory can store instructions or data that the processor 910 has just used or that are used repeatedly. If the processor 910 needs to use the instruction or data again, it can directly retrieve it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 910, and thus improves the efficiency of the system.

[0152] In some alternative implementations, the memory may store instructions or data of the image processing method mentioned in the embodiments of this application.

[0153] Electronic devices implement display functions through a GPU, a display screen 994, and an application processor. The GPU is a microprocessor for image processing, connecting the display screen 994 and the application processor. The GPU performs mathematical and geometric calculations and is used for graphics rendering. The processor 910 may include one or more GPUs, which execute program instructions to generate or modify display information.

[0154] Display screen 994 is used to display images, videos, etc. Display screen 994 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), Mini-LED, MicroLED, Micro-OLED, quantum dot light-emitting diodes (QLED), etc. In some embodiments, the electronic device may include one or N displays 994, where N is a positive integer greater than 1.

[0155] Electronic devices can achieve shooting functions through ISP, camera 993, video codec, GPU, display 994 and application processor.

[0156] The ISP (Image Signal Processor) is used to process data fed back from the camera 993. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, transforming it into an image visible to the naked eye. The ISP can also perform algorithmic optimization of image noise, brightness, and skin tone. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set within the camera 993.

[0157] In some alternative implementations, the electronic device may have a depth camera and a color camera. The depth camera captures a depth image containing image parameters including depth information and infrared light information. The depth information represents the distance of the object to the depth camera, and the infrared light information represents the object's reflection of infrared light. The color camera captures a color image reflecting the color of the object. Furthermore, the depth image and the color image are of the same size.

[0158] Camera 993 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard formats such as RGB and YUV. In some embodiments, the electronic device may include one or N cameras 993, where N is a positive integer greater than 1.

[0159] Digital signal processors (DSPs) are used to process digital signals. Besides digital image signals, they can also process other digital signals. For example, when an electronic device is selecting a frequency, a DSP can perform a Fourier transform on the frequency energy.

[0160] Video codecs are used to compress or decompress digital video. Electronic devices can support one or more video codecs. This allows the electronic device to play or record video in various encoded formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.

[0161] An NPU (Neural Processing Unit) is a computational processor for neural networks (NNs). By borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can rapidly process input information and continuously learn on its own. NPUs enable intelligent cognitive applications in electronic devices, such as image recognition, facial recognition, speech recognition, and text understanding.

[0162] The software structure of the electronic device to which the image processing method provided in the embodiments of this application is applicable will be described exemplarily below.

[0163] The software system of an electronic device can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture, etc. This embodiment of the invention uses a layered architecture as an example to illustrate the software structure of an electronic device. In some embodiments, the software structure of the electronic device can be the same.

[0164] Figure 10 This is a software structure block diagram of an electronic device according to an embodiment of this application.

[0165] A layered architecture divides software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom: the application layer, the framework layer, the system runtime library, and the kernel layer.

[0166] like Figure 10 As shown, the application layer can include a series of application packages. These application packages can include applications such as image processing, gallery, navigation, WLAN, Bluetooth, music, video, and SMS.

[0167] The image processing application can implement the image processing method provided in the embodiments of this application.

[0168] In some specific implementations, image processing applications can obtain a reference depth image based on image features in a color image, such as texture, pattern, shape, lighting, shadows, spatial relationships, and semantic information. This reference depth image includes depth differences or depth relationships between pixels. Then, based on this reference depth image, the depth image captured by the depth camera is corrected, that is, the depth information of pixels missing from the depth image is supplemented. In this way, the depth information of pixels lacking depth information in the image can be completed, thereby avoiding image distortion and improving image quality.

[0169] In this way, the depth information of pixels that are missing depth information in the image can be completed, thereby avoiding image distortion and improving image quality.

[0170] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer includes some predefined functions.

[0171] like Figure 10 As shown, the application framework layer may include a window manager, content provider, phone manager, resource manager, notification manager, view system, etc.

[0172] The window manager is used to manage windowed applications. It can retrieve screen size, determine the presence of a status bar, lock the screen, and capture screenshots, among other things.

[0173] Content providers store and retrieve data, making that data accessible to applications. This data may include videos, images, audio, made and received phone calls, browsing history and bookmarks, phone books, etc.

[0174] A phone manager is used to provide communication functions for electronic devices. For example, it manages call status (including connection and disconnection).

[0175] The file explorer provides applications with various resources, such as localized strings, icons, images, layout files, video files, and more.

[0176] The notification manager allows applications to display notifications in the status bar. These notifications can be used to deliver informational messages and can disappear automatically after a short pause, requiring no user interaction. For example, the notification manager can be used to notify users of download completion or message alerts. The notification manager can also display notifications as icons or scrolling text in the top status bar, such as notifications from background applications, or as dialog boxes on the screen. Examples include displaying text messages in the status bar, emitting sounds, vibrating electronic devices, and flashing indicator lights.

[0177] A view system includes visual controls, such as controls for displaying text and controls for displaying images. View systems can be used to build applications. A display interface can consist of one or more views. For example, a display interface including a text notification icon could include views for displaying text and views for displaying images.

[0178] The Android Runtime consists of core libraries and a virtual machine. The Android runtime is responsible for the scheduling and management of the Android system.

[0179] The core library consists of two parts: one part is the functionalities that need to be called by the Java language, and the other part is the Android core library.

[0180] The application layer and application framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.

[0181] System libraries can include multiple functional modules. For example: surface manager, 3D graphics processing library (e.g., OpenGL ES), 2D graphics engine (e.g., SGL), media libraries, etc.

[0182] The Surface Manager is used to manage the display subsystem and provides the blending of 2D and 3D layers for multiple applications.

[0183] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.

[0184] A 2D graphics engine is a drawing engine for 2D drawing.

[0185] The media library supports playback and recording of various common audio and video formats, as well as still image files. It supports multiple audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, and PNG.

[0186] The kernel layer is the layer between hardware and software. The kernel layer contains at least the display driver, camera driver, audio driver, and sensor driver.

[0187] In some cases, the embodiments disclosed in this application may be implemented in hardware, firmware, software, or any combination thereof.

[0188] The embodiments disclosed in this application can also be implemented as instructions carried or stored thereon on one or more temporary or non-temporary machine-readable (e.g., computer-readable) storage media, which can be read and executed by one or more processors. For example, the instructions can be distributed via a network or via other computer-readable media. Therefore, machine-readable media can include any mechanism for storing or transmitting information in a machine-readable (e.g., computer-readable) form, including but not limited to floppy disks, optical disks, optical discs, magneto-optical disks, read-only memory (ROM), random access memory (RAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic cards or optical cards, flash memory, or tangible machine-readable storage for transmitting information (e.g., carrier waves, infrared signals, digital signals, etc.) using the Internet in the form of electrical, optical, acoustic, or other propagation signals. Therefore, machine-readable media includes any type of machine-readable medium suitable for storing or transmitting electronic instructions or information in a machine-readable (e.g., computer-readable) form.

[0189] Embodiments of this application can be implemented as computer programs or program code that execute on a programmable system, the programmable system including at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.

[0190] Program code can be applied to input instructions to execute the functions described in this application and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of this application, the processing system includes any system having a processor such as, for example, a digital signal processor (DSP), a microcontroller, an application-specific integrated circuit (ASIC), or a microprocessor.

[0191] The program code can be implemented using a high-level procedural language or an object-oriented programming language to communicate with the processing system. Assembly language or machine language can also be used when needed. In fact, the mechanisms described in this application are not limited to any particular programming language. In either case, the language can be a compiled language or an interpreted language.

[0192] In some embodiments, this application also provides a computer-readable medium storing program code that, when run on a computer, causes the computer to perform the methods described above.

[0193] In some embodiments, this application also provides a computer program product comprising: computer program code that, when run on a computer, causes the computer to perform the methods described above.

[0194] The above describes the possible hardware structures of electronic devices. It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device. In other embodiments of this application, the electronic device may include more or fewer components than illustrated, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of both.

[0195] In the accompanying drawings, some structural or methodological features may be shown in a specific arrangement and / or order. However, it should be understood that such a specific arrangement and / or order may not be necessary. Rather, in some embodiments, these features may be arranged in a manner and / or order different from that shown in the illustrative drawings. Furthermore, the inclusion of structural or methodological features in a particular figure does not imply that such features are required in all embodiments, and in some embodiments, these features may be omitted or may be combined with other features.

[0196] It should be noted that in the examples and description of this patent, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one" does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0197] Although this application has been illustrated and described with reference to certain embodiments thereof, those skilled in the art will understand that various changes in form and detail may be made thereto without departing from the scope of this application.

Claims

1. An image processing method, characterized in that, Applied to electronic devices, including: Acquire the first depth image and color image of the object captured by the electronic device; Based on the image features of the color image, a reference depth image is obtained, wherein the reference depth image includes the depth relationship between each pixel in the color image; For a first pixel in the first depth image that has missing depth information, the depth information of the first pixel is determined based on the reference depth image to obtain a second depth image; The depth information of each pixel in the second depth image is offset based on the depth offset to obtain the third depth image. The depth offset is determined based on the shooting distance between the electronic device and the object and a first relationship, wherein the first relationship is a one-to-one correspondence between the shooting distance and the depth offset. Based on the scaling factor and coordinate transformation relationship, the third depth image is transformed to obtain the fourth depth image. Wherein, the position of the third pixel representing the first position of the object in the fourth depth image is the same as the position of the fourth pixel representing the first position in the color image, and the coordinate transformation relationship is the transformation relationship between the coordinate system of the depth camera that captures the first depth image and the coordinate system of the color camera that captures the color image. The target image is obtained by fusing the fourth depth image and the color image; The scaling factor is determined in the following way: Determine at least one key point from the second depth image; The overall depth of the second depth image is determined based on the depth information of the at least one key point; The ratio of the total depth to the overall depth is used as the scaling factor, wherein the total depth is the sum of the overall depth of the second depth image and the depth offset.

2. The method according to claim 1, characterized in that, Determining the depth information of the first pixel based on the reference depth image includes: Multiple associated pixels in the reference depth image are identified that are associated with the first pixel. The multiple associated pixels include a first associated pixel corresponding to the first pixel and at least one second associated pixel. At least some of the at least one second associated pixels have corresponding pixels in the first depth image that do not lack depth information. The depth information of the first pixel is determined based on the depth information of the pixels corresponding to the multiple associated pixels in the first depth image.

3. The method according to claim 2, characterized in that, The determination of multiple associated pixels in the reference depth image that are associated with the first pixel includes: At least one second pixel is determined from the first depth image, wherein the second pixel is a pixel that does not lack depth information; The pixel corresponding to the second pixel in the reference depth image is taken as the second associated pixel.

4. The method according to claim 3, characterized in that, The step of determining the depth information of the first pixel based on the depth information of the multiple associated pixels corresponding to the multiple pixels in the first depth image includes: The depth information of the first pixel is obtained based on the depth relationship between the first associated pixel and the second associated pixel in the reference depth image, and the depth information of the second pixel in the first depth image.

5. The method according to claim 2, characterized in that, The determination of multiple associated pixels in the reference depth image that are associated with the first pixel includes: At least one pixel in the reference depth image whose similarity to the depth information of the first associated pixel is higher than a similarity threshold and whose corresponding pixel in the first depth image does not lack depth information is identified as the second associated pixel.

6. The method according to claim 5, characterized in that, The step of determining the depth information of the first pixel based on the depth information of the multiple associated pixels corresponding to the multiple pixels in the first depth image includes: Multiple pixels corresponding to the multiple associated pixels are determined from the first depth image; The depth information of the first pixel is determined based on the depth information of the plurality of pixels.

7. The method according to claim 6, characterized in that, Corresponding to the object being a face, the key points include pixels corresponding to at least one of the following features in the face: Left corner of the eye, right corner of the eye, left corner of the mouth, right corner of the mouth, tip of the nose.

8. An electronic device, characterized in that, include: A memory for storing instructions executed by one or more processors of the electronic device, and a processor, being one of one or more processors of the electronic device, for performing the image processing method according to any one of claims 1-7.

9. A readable storage medium, characterized in that, The readable storage medium stores instructions that, when executed on an electronic device, cause the electronic device to perform the image processing method according to any one of claims 1-7.

10. A computer program product, characterized in that, The computer program product includes computer instructions, which, when executed by an electronic device, execute the computer program code of the image processing method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Image depth completion method and device, computer equipment and storage medium

    CN117788546A

  • Parallax truth value calculation method and device and storage medium

    CN118297887A