Image Processing Method and Its Apparatus

By collecting posture information at the part of the subject to perform three-dimensional reconstruction, the problem of low accuracy and reliability in extreme scenes in the prior art is solved, and stable three-dimensional reconstruction in various scenes is achieved.

CN115272574BActive Publication Date: 2025-07-22VIVO MOBILE COMM CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210896910.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-28
Publication Date
2025-07-22
Estimated Expiration
2042-07-28

AI Technical Summary

Technical Problem

When existing electronic devices are reconstructed in three-dimensionally, they are particularly in extreme scenarios (such as dark light, occlusion, and large-angle motion). The reconstruction accuracy is reduced and the reliability is low.

Method used

The pose information is collected by the device worn on the part of the subject, and the information is used for three-dimensional reconstruction to generate the three-dimensional reconstruction image.

Benefits of technology

It improves the reliability and accuracy of three-dimensional reconstruction, and can carry out 3-dimensional reconstruction stably in various scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115272574B_ABST
    Figure CN115272574B_ABST
Patent Text Reader

Abstract

The present application discloses an image processing method and apparatus thereof, belonging to the field of communication technologies. The method includes: obtaining a first image, where the first image includes a first region of a first object, the first object being the imaging of a first photographing object in the first image, and the first region being the imaging region of a first part of the first photographing object in the first image; obtaining first attitude information of the first part at a first moment collected by a first device, where the first device is worn on a second part of the first photographing object, and the first moment is the acquisition moment of the first image; performing three-dimensional reconstruction on the first part according to the first attitude information to generate a second image including the three-dimensionally reconstructed first part.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of image processing, and particularly relates to an image processing method and apparatus thereof. Background Art

[0002] Three-dimensional (3D) face reconstruction technology can restore the three-dimensional shape of a face in a two-dimensional (2D) face image, and has high application value and broad application prospects in fields such as animation production and online games.

[0003] Currently, the 3D face reconstructed by an electronic device in a normal scenario with normal illumination, no occlusion, and no large-angle rapid movement can achieve relatively high reconstruction accuracy. However, for the 3D face reconstructed in extreme scenarios such as low light, occlusion, and large-angle movement, the reconstruction accuracy is extremely likely to decrease, the point positions are inaccurate, and there are floating situations. It can be seen that the reliability of the three-dimensional reconstruction of existing electronic devices is relatively low. Summary of the Invention

[0004] The purpose of the embodiments of this application is to provide an image processing method and apparatus thereof, which can solve the problem of relatively low reliability of electronic devices in three-dimensional reconstruction in the prior art.

[0005] In a first aspect, the embodiments of this application provide an image processing method, which is applied to an electronic device. The method includes:

[0006] Obtain a first image, where the first image includes a first region of a first object, the first object is the imaging of a first shooting object in the first image, and the first region is the imaging region of a first part of the first shooting object in the first image;

[0007] Obtain first pose information of the first part at a first moment collected by a first device, where the first device is worn on a second part of the first shooting object, and the first moment is the collection moment of the first image;

[0008] Perform three-dimensional reconstruction on the first part according to the first pose information to generate a second image including the three-dimensionally reconstructed first part.

[0009] In a second aspect, the embodiments of this application provide an image processing apparatus, which is applied to an electronic device and includes:

[0010] A first acquisition module, configured to obtain a first image, where the first image includes a first region of a first object, the first object is the imaging of a first shooting object in the first image, and the first region is the imaging region of a first part of the first shooting object in the first image;

[0011] A second acquisition module, configured to acquire first pose information of the first part at a first moment collected by the first device, where the first device is worn on a second part of the first photographed object, and the first moment is the acquisition moment of the first image;

[0012] A first generation module, configured to perform three-dimensional reconstruction on the first part according to the first pose information, and generate a second image including the three-dimensionally reconstructed first part.

[0013] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor and a memory. The memory stores a program or instruction that can run on the processor. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.

[0014] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.

[0015] In a fifth aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is configured to run a program or instruction to implement the method described in the first aspect.

[0016] In a sixth aspect, an embodiment of the present application provides a computer program product, which is stored in a storage medium and is executed by at least one processor to implement the method described in the first aspect.

[0017] In an embodiment of the present application, after acquiring a first image of a first region including a first object, it is possible to acquire the pose information of a first part of the first photographed object at a first moment collected by a first device worn on a second part of the first photographed object, and perform three-dimensional reconstruction on the first part based on the pose information to obtain a second image of the three-dimensionally reconstructed first part, where the first object is the imaging of the first photographed object in the first image, the first region is the imaging region of the first part in the first image, and the first moment is the acquisition moment of the first image. It can be seen that the embodiment of the present application can perform three-dimensional reconstruction on the first part of the photographed object by virtue of the first pose information of the first part collected by the first device worn on the second part of the photographed object. Since the pose information of the first part is collected by a device worn on the photographed object, it can be not limited by the image acquisition scene, ensuring the accuracy of obtaining the pose information of the first part, and thus improving the reliability of three-dimensional reconstruction. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 is one of the flowcharts of the image processing method provided by the embodiment of the present application;

[0019] Figure 2 It is a schematic diagram of the attitude information provided by the embodiments of the present application;

[0020] Figure 3 It is a schematic diagram of the image acquisition scenario provided by the embodiments of the present application;

[0021] Figure 4a It is one of the schematic diagrams of the image processing algorithm provided by the embodiments of the present application;

[0022] Figure 4b It is another schematic diagram of the image processing algorithm provided by the embodiments of the present application;

[0023] Figure 5 It is the second flowchart of the image processing method provided by the embodiments of the present application;

[0024] Figure 6a It is a schematic diagram of the first image provided by the embodiments of the present application;

[0025] Figure 6b It is a schematic diagram of the first area provided by the embodiments of the present application;

[0026] Figure 7 It is one of the schematic diagrams of the image processing effect provided by the embodiments of the present application;

[0027] Figure 8a It is another schematic diagram of the image processing effect provided by the embodiments of the present application;

[0028] Figure 8b It is the third schematic diagram of the image processing effect provided by the embodiments of the present application;

[0029] Figure 9 It is the structural diagram of the image processing device provided by the embodiments of the present application;

[0030] Figure 10 It is one of the structural diagrams of the electronic device provided by the embodiments of the present application;

[0031] Figure 11 It is another structural diagram of the electronic device provided by the embodiments of the present application. Detailed implementation manners

[0032] Next, the technical solutions in the embodiments of the present application will be clearly described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application belong to the scope of protection of the present application.

[0033] The terms "first", "second", etc. in the description and claims of this application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same category, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, "and / or" in the description and claims means at least one of the connected objects, and the character " / ", generally indicates an "or" relationship between the associated objects before and after.

[0034] With the development of Augmented Reality (AR) and Virtual Reality (VR) technologies, current image processing is beginning to develop from 2D to 3D. In electronic devices, especially mobile phones, three-dimensional face reconstruction is an important technology in this development trend.

[0035] The following will, with reference to the accompanying drawings, explain in detail the image processing method provided by the embodiments of this application through specific embodiments and their application scenarios.

[0036] Figure 1 It is one of the schematic flowcharts of the image processing method provided by the embodiments of this application. The image processing method of the embodiments of this application can be applied to an electronic device or be executed by an electronic device. In practical applications, the electronic device can be a mobile phone, a tablet personal computer, a laptop computer or a notebook computer, a personal digital assistant (PDA), a handheld computer, a netbook, an ultra-mobile personal computer (UMPC), a mobile Internet device (MID) held by a user, an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, a vehicle-mounted device (VUE), a pedestrian terminal (PUE), etc.

[0037] As Figure 1 shown, the image processing method may include the following steps:

[0038] Step 101: Obtain a first image, where the first image includes a first region of a first object, the first object is the imaging of a first photographing object in the first image, and the first region is the imaging region of a first part of the first photographing object in the first image.

[0039] In an embodiment of the present application, the first image may be any image including a first part of a first photographing object collected by an image acquisition device of an electronic device. In one implementation, the first image may be an image obtained by the image acquisition device operating in a photographing mode; in another implementation, the first image may be an image obtained by the image acquisition device operating in a video recording mode. At this time, the first image may be any image frame in the recorded video.

[0040] Optionally, the first image may include at least one object, and the first object may be any object in the first image. That is, for any object in the first image, the image processing method of the embodiment of the present application can be used to perform three-dimensional reconstruction on its corresponding first part. In practical applications, the first object may be randomly determined by the electronic device or determined by the user independently.

[0041] The first object may include at least one region, such as a face region, a body region (which can be further divided into a hand region, a leg region, etc.). The first region may be any region in the first object. That is, for any region in the first object, the image processing method of the embodiment of the present application can be used to perform three-dimensional reconstruction on its corresponding part. In practical applications, the first region may be randomly determined by the electronic device or determined by the user independently. It can be understood that when the first region is a face region, the first part is a face part; when the first region is a hand region, the first part is a hand.

[0042] Step 102: Obtain first pose information of the first part at a first moment collected by a first device, where the first device is worn on a second part of the first photographing object, and the first moment is the acquisition moment of the first image.

[0043] In an embodiment of the present application, the first device is worn on a second part of the first photographing object and can collect the pose information (or called spatial information, spatial angle information) of the first part of the first photographing object. That is, the first device is worn on a second part of the first photographing object, but the collected pose information is applicable to the first part. Specifically, the first device can collect the pose information of the first part through a gyroscope or other elements that can collect pose information carried by itself.

[0044] To make the attitude information collected by the first device applicable to the first part, in one implementation, the second part and the first part can be the same part. For example, when the first part is the hand (or leg), the first device can be a wearable device worn on the hand (or leg) and equipped with a gyroscope. In another implementation, the motion attitude of the second part is the same as or similar to that of the first part. For example, when the first part is the human face, the first device can be an earphone or glasses worn on the ear and equipped with a gyroscope, etc.

[0045] In the embodiments of the present application, the first device and the electronic device can communicate, and the first device is a peripheral device wired or wirelessly connected to the electronic device. After the first device collects the attitude information of the first part, it can transmit it to the electronic device so that the electronic device can perform three-dimensional reconstruction on the first part according to the received attitude information of the first part.

[0046] Further, when the first device transmits the attitude information, it can simultaneously transmit the time stamps corresponding to each attitude information. In this way, the electronic device can use the time stamps corresponding to the attitude information to determine and use the true attitude information of the first part at the acquisition moment of the image to perform three-dimensional reconstruction on the first part of the image, thereby improving the accuracy of three-dimensional reconstruction.

[0047] The embodiments of the present application do not limit the specific manifestation form of the attitude information, which can be specifically set according to actual needs. In one implementation, as Figure 2 shown, the attitude information can be represented as a spatial rotation angle (or called Euler rotation angle), which can specifically include: the pitch angle (pitch) of rotation around the X-axis, the yaw angle (yaw) of rotation around the Y-axis, and the roll angle (roll) of rotation around the Z-axis.

[0048] In some embodiments, the spatial angle at a certain moment collected by the first device can be directly determined as the spatial rotation angle of the first part at that moment. In other embodiments, the difference between the spatial angle at a certain moment collected by the gyroscope and the reference spatial angle can be used as the spatial rotation angle of the first part at that moment. For example, when the first attitude information R' can be represented as the Euler rotation angle R t of the first part at the first moment, it can be calculated by formula (1):

[0049] R t = R' t - R' t-1 (1)

[0050] where, R t ' represents the spatial angle of the first part at the first moment, which is collected by the first device; R t-1' represents the reference spatial angle of the first part. When the first image is an image frame in a video, the reference spatial angle can be the spatial angle of the first part at the acquisition moment of the previous image frame of the first image, which is acquired by the first device; when the first image is not an image frame in a video, the reference spatial angle can be set to zero, that is, the spatial angle of the first part at the first moment can be determined as the spatial rotation angle of the first part at the first moment.

[0051] Step 103: Perform three-dimensional reconstruction on the first part according to the first pose information to generate a second image including the three-dimensionally reconstructed first part.

[0052] Specifically, the electronic device can first reconstruct the 3D shape S of the first part based on the target information whole , where the target information at least includes the shape information of the first part. Further, when the first part is a human face, the target information can also include the expression information of the face. In one implementation, the target information can be obtained by inputting the first region into the first model and regressing; in another implementation, the target information can be pre-acquired reference information, and specific details can be referred to the following related descriptions and will not be described here.

[0053] After that, the first pose information can be used to rotate the position of S whole to obtain the three-dimensionally reconstructed first part S final .

[0054] The main difference between the above three-dimensional reconstruction and the three-dimensional reconstruction in the related art is that the pose information used for three-dimensional reconstruction is different. In the embodiments of the present application, the pose information of the first part collected by the first device worn on the second part of the photographed object is used. Other implementations can refer to the related art and will not be described here.

[0055] In the image processing method of the embodiments of the present application, after obtaining the first image including the first region of the first object, the pose information of the first part of the first photographed object at the first moment collected by the first device worn on the second part of the first photographed object can be obtained, and based on this pose information, three-dimensional reconstruction of the first part is performed to obtain a second image of the three-dimensionally reconstructed first part, where the first object is the imaging of the first photographed object in the first image, the first region is the imaging region of the first part in the first image, and the first moment is the acquisition moment of the first image. It can be seen that the embodiments of the present application can perform three-dimensional reconstruction on the first part of the photographed object by virtue of the first pose information of the first part collected by the first device worn on the second part of the photographed object. Since the pose information of the first part is collected by the device worn on the photographed object, it can be not limited by the image acquisition scene, ensuring the acquisition accuracy of the pose information of the first part, thereby improving the reliability of the three-dimensional reconstruction.

[0056] In practical applications, when the electronic device captures the first image, in one scenario, the electronic device may be placed in a fixed position. Exemplarily, as Figure 3 described, the electronic device is mounted on a support. In Figure 3 this case, the electronic device is a mobile phone, the first part is a human face, and the first device is a headset. However, this does not limit the forms of the electronic device, the first part, and the first device. In this scenario, the electronic device can directly perform three-dimensional reconstruction on the first part according to the first pose information.

[0057] In another scenario, the electronic device may be held by a user, such as: held by a certain subject to be photographed, or held by the person being photographed. In this scenario, performing three-dimensional reconstruction on the first part according to the first pose information to generate a second image including the three-dimensionally reconstructed first part may include:

[0058] When the electronic device is held by a user, obtaining second pose information of the electronic device captured at the first moment;

[0059] Determining first difference information between the first pose information and the second pose information;

[0060] Performing three-dimensional reconstruction on the first part according to the first difference information to generate a second image including the three-dimensionally reconstructed first part.

[0061] In this scenario, since the electronic device is held by a user, the pose information of the electronic device may change during the image capture process. Therefore, by calculating the difference information between the second pose information of the electronic device at the first moment and the first pose information, the true pose information of the first part at the first moment can be obtained, that is, the above first difference information is the true pose information of the first part at the first moment. After that, the first part in 3D can be reconstructed using the first difference information. In this way, the reliability of the 3D reconstruction of the first part can be improved.

[0062] It should be noted that the image processing method of the embodiments of the present application can be applicable to normal shooting scenarios with normal lighting conditions, no occlusion of the first part, and no large-angle rapid movement of the first part, and can also be applicable to extreme shooting scenarios with abnormal lighting conditions (such as the presence of low light or bright light), occlusion of the first part, and large-angle movement of the first part (hereinafter referred to as target shooting scenarios).

[0063] In some embodiments, the electronic device may not pay attention to the current shooting scenario and directly perform three-dimensional reconstruction through the image processing method of the embodiments of the present application.

[0064] In some other embodiments, the three-dimensional reconstruction of the first part according to the first pose information to generate a second image including the three-dimensionally reconstructed first part may include:

[0065] When a first condition is satisfied, performing three-dimensional reconstruction on the first part according to the first pose information to generate a second image including the three-dimensionally reconstructed first part;

[0066] Wherein, satisfying the first condition indicates that the electronic device is in a target shooting scene.

[0067] In this embodiment, the electronic device can pay attention to the current shooting scene and perform three-dimensional reconstruction in different ways according to different shooting scenes.

[0068] When the current shooting scene is a normal shooting scene, three-dimensional reconstruction can be performed through a three-dimensional reconstruction algorithm in related technologies. Taking human face three-dimensional reconstruction as an example, in one way, the algorithm predicts the 3D spatial coordinates of each point corresponding to the 3D human face. For ease of understanding, please refer to Figure 4a , such as Figure 4a shown, an rgb (red - green - blue) image of the human face area can be input into a neural network regressor to obtain the 3D human face coordinates (the data volume is several thousand to tens of thousands), that is, the output of this neural network regressor is the 3D human face coordinates. After that, a 3D human face is reconstructed based on the 3D human face coordinates. In another way, the algorithm regresses the basis model coefficients of the human face, such as shape coefficients (also called shape parameters), expression coefficients (also called expression parameters), pose coefficients (also called pose parameters), etc., and then calculates the final 3D spatial coordinates through a predefined model. For ease of understanding, refer to Figure 4b , such as Figure 4b shown, an rgb image of the human face area can be input into a neural network regressor to obtain the basis model coefficients (the data volume is dozens to hundreds), that is, the output of this neural network regressor is the basis model coefficients. After that, a 3D human face is reconstructed by combining the basis model coefficients and the predefined model. Among them, the predefined model can be understood as a basis model database, and the basis model can include a shape basis and an expression basis. Specifically in implementation, the basis model coefficients can be multiplied by the corresponding basis models to reconstruct a 3D human face.

[0069] When the current shooting scene is a target shooting scene, the image processing method of this application embodiment can be used for three-dimensional reconstruction.

[0070] In this way, in a normal shooting scene, the electronic device does not need to obtain the pose information collected by the first device and can disconnect from the first device, thereby saving the power consumption of the electronic device.

[0071] In the embodiments of the present application, the electronic device can determine whether it is in a target shooting scenario in various ways.

[0072] In some embodiments, the electronic device can determine whether it is in a target shooting scenario by comparing the information of two consecutive frames. For example, if the shape change value of the first part in two consecutive frames exceeds a threshold, it indicates that there is a large-angle movement of the first part, and it can be determined that the electronic device is in a target shooting scenario. Or, the electronic device can identify whether there is occlusion of the first part through an occlusion detection algorithm to determine whether the electronic device is in a target shooting scenario. Or, the electronic device can identify whether the brightness value in the image acquisition environment is within a preset range to determine whether the lighting condition is normal, and further determine whether the electronic device is in a target shooting scenario. Or, the electronic device can determine whether the electronic device is in a target shooting scenario based on a user-defined shooting scenario.

[0073] In some other embodiments, before performing three-dimensional reconstruction on the first part according to the first pose information to generate a second image including the three-dimensionally reconstructed first part, the method may further include:

[0074] a1. Input the first region into a first model to regress and obtain the third pose information of the first part at the first moment;

[0075] a2. Determine the second difference information between the third pose information and the target pose information of the first part at the first moment, where the target pose information is related to the first pose information;

[0076] Performing three-dimensional reconstruction on the first part according to the first pose information to generate a second image including the three-dimensionally reconstructed first part includes:

[0077] In the case where the second difference information is greater than the threshold information, perform three-dimensional reconstruction on the first part according to the first pose information to generate a second image including the three-dimensionally reconstructed first part.

[0078] In the embodiments of the present application, after the electronic device acquires the first image, it can detect the first regions of various objects in the first image through a region detection algorithm and crop out the first regions.

[0079] The electronic device can input the first region into a first model and use the first model to regress and obtain the pose information of the first part at the first moment, that is, the third pose information. The first model can be Figure 4b a neural network regressor in, but not limited thereto.

[0080] The target pose information represents the true pose information of the first part at the first moment. The target pose information is related to the first pose information, specifically, the target pose information is the first pose information, or the target pose information is determined based on the first pose information.

[0081] In an alternative implementation, before determining the second difference information between the third pose information and the target pose information of the first part at the first moment, the method may further include at least one of the following:

[0082] b1. When the electronic device is placed at a fixed position, determining the first pose information as the target pose information;

[0083] b2. When the electronic device is held by a user, determining the first difference information as the target pose information, where the first difference information is the difference information between the second pose information collected by the electronic device at the first moment and the first pose information.

[0084] That is, when the electronic device is placed at a fixed position, directly using the first pose information collected by the first device at the first moment as the true pose information of the first part at the first moment. When the electronic device is held by a user, using the difference information between the first pose information collected by the first device at the first moment and the second pose information of the electronic device at the first moment as the true pose information of the first part at the first moment. In this way, the reliability of determining the true pose information of the first part can be ensured.

[0085] After the electronic device obtains the third pose information and the target pose information, it can compare the difference information between the two, that is, the second difference information, with the threshold information to determine the robustness of the first model, and then determine the shooting scene where the electronic device is located. It can be understood that the robustness of the first model is related to the shooting scene where the electronic device is located. Specifically, when the electronic device is in the target shooting scene, the robustness of the first model is poor; when the electronic device is in the normal shooting scene, the robustness of the first model is good.

[0086] When the second difference information is greater than the threshold information, it can be determined that the robustness of the first model is poor and the electronic device is in the target shooting scene. In this case, the electronic device can perform 3D reconstruction based on the target pose information.

[0087] When the second difference information is less than the threshold information, it can be determined that the robustness of the first model is good and the electronic device is in the normal shooting scene. In this case, the electronic device can perform 3D reconstruction based on the third pose information.

[0088] It can be seen that in the embodiment, by calculating the difference information between the third pose information of the first part obtained by the first model based on regression in the first area at the first moment and the target pose information that can reflect the true pose information of the first part at the first moment, and comparing the difference information with the threshold information, the shooting scene where the electronic device is located can be determined, so that the determined shooting scene can conform to the actual shooting scene. Furthermore, a suitable method can be selected for 3D reconstruction, improving the reliability of 3D reconstruction.

[0089] In some embodiments, after determining the second difference information between the third pose information and the target pose information of the first part at the first moment, the method may further include:

[0090] c1. When the second difference information is less than or equal to the threshold information, obtain the confidence weight of the first device;

[0091] c2. Adjust the third pose information according to the confidence weight and the first pose information to obtain the fourth pose information;

[0092] c3. Perform 3D reconstruction on the first area according to the fourth pose information to generate a third image including the 3D reconstructed first area.

[0093] In this embodiment, in a normal shooting scene, the true pose information of the first part determined based on the first device at the first moment can be used to fine-tune the third pose information obtained by regression of the first model, further improving the reliability of 3D reconstruction.

[0094] In a specific implementation, in an alternative implementation, the fourth pose information can be calculated by formula (2):

[0095] R = (1 - w) * R + w * R′ (2)

[0096] Where R on the left side of the equation represents the fourth pose information, R on the right side of the equation represents the third pose information, w represents the confidence weight of the first electronic device, and * represents the multiplication sign. That is, the third pose information and the first pose information can be weighted and averaged to obtain the fourth pose information.

[0097] In another alternative implementation, the fourth pose information can be obtained by adding or subtracting the confidence weight and the first pose information from the third pose information.

[0098] In this embodiment, the true pose information of the first part determined based on the first device at the first moment can be used to assist 3D reconstruction in a normal shooting scene, thereby further improving the reliability of 3D reconstruction.

[0099] In some embodiments, performing three-dimensional reconstruction on the first part according to the first pose information to generate a second image including the three-dimensionally reconstructed first part includes:

[0100] Obtaining reference shape information of the first part;

[0101] Performing three-dimensional reconstruction on the first part according to the first pose information and the reference shape information to generate a second image including the three-dimensionally reconstructed first part.

[0102] In this embodiment, to improve the restoration degree of the three-dimensional shape of the first part, reference shape information of the first part can be obtained based on pre-acquired images with relatively high quality. The images with relatively high quality can meet the following conditions: the lighting condition of the first part is normal, the first part is not blocked, and there is no large-angle movement. Specifically, in implementation, the images with relatively high quality can be: high-quality image frames selected during the three-dimensional reconstruction of a video, or high-quality images prepared in advance by the user.

[0103] In this way, performing three-dimensional reconstruction using the reference shape information can make the shape of the three-dimensionally reconstructed first part match the actual shape, and further improve the image processing effect.

[0104] In other embodiments, the shape information used for reconstructing the first part can be determined based on the first region. For example, the first region is input into a first model to regress the shape information of the first part, and this shape information is used to perform three-dimensional reconstruction on the first part.

[0105] It should be noted that the various optional embodiments described in the embodiments of the present application can be combined and implemented with each other or implemented separately without conflict, and the embodiments of the present application do not make any limitations in this regard.

[0106] The following describes the information storage method provided by the embodiments of the present application in combination with a specific application example.

[0107] To facilitate understanding of the image processing method provided in the above embodiments, the above image processing method is described below with a specific scenario embodiment. Figure 5 It is a schematic flowchart of a scenario embodiment of the image processing method provided by the embodiments of the present application.

[0108] The application scenario of this scenario embodiment can be: leveraging the hardware advantages of mobile phone manufacturers, combining the spatial angle information collected by the gyroscopes in the user's worn earphones, designing an extreme shooting scenario for occlusion, low light, large-angle movement, or a combination of the above multiple situations, performing real-time, stable, and well-fitting 3D face video reconstruction, and further, this information can be used to assist 3D face video reconstruction in normal shooting scenarios.

[0109] The physical usage scenario of the embodiment of this scenario is as follows Figure 3 shown. The user wears a headset with a gyroscope and faces the mobile phone. The mobile phone is placed on a support. If it is held by hand, the gyroscope difference between the headset and the mobile phone needs to be calculated to obtain the spatial information of the user's head.

[0110] As Figure 5 shown, the embodiment of this scenario may specifically include the following steps

[0111] Step 1: Obtain the reference shape basis s'.

[0112] Step 2: Obtain the image frame.

[0113] In the embodiment of this scenario, the input data is a single-frame video image, that is, a single 2D image, as Figure 6a shown.

[0114] Step 3: Detect the face region in the image frame.

[0115] The face detection algorithm detects the face position and crops out the face region, as Figure 6b shown.

[0116] Step 4: Input the face region into the first model.

[0117] Specifically, as Figure 4b shown, the rgb image of the face region is input into the Figure 4b neural network regressor in.

[0118] Step 5: The first model outputs the 3D shape s, expression e, and spatial rotation angle R of the face.

[0119] The neural network regressor can calculate the coefficient α of the shape basis s of the face in the current picture frame, the coefficient β of the expression basis e, and the spatial rotation angle R. Among them, the shape basis s represents the shape of the face. For the video stream of the same person, this shape is almost invariant physically. Therefore, a picture with higher quality in the reconstruction process or a high-quality picture prepared by the user can be selected to obtain s, that is Figure 5 the s' obtained in step 1 in, as the backup shape basis. Using s and e, the 3D shape of the face can be reconstructed. The reconstruction formula (3) is as follows

[0120] S whole = ∑α i * s i + ∑β i * e i (3)

[0121] Among them, α i is the coefficient of the shape basis of the i-th face in the database, si is the shape basis of the i-th face in the database, β i is the coefficient of the expression basis of the i-th face in the database, e i is the expression basis of the i-th face in the database, S whole are the 3D space coordinates to be obtained.

[0122] Then, by combining the spatial rotation parameter R and performing position rotation, the final spatial pose S can be obtained final , and the calculation formula (4) is as follows:

[0123] S final = R * S whole (4)

[0124] Step 6: The headphone gyroscope outputs the face spatial rotation angle R'.

[0125] The headphone gyroscope transmits the spatial rotation angle R' of the user in each frame, which can be calculated according to the following formula (5):

[0126] R t = R′ t - R′ t-1 (5)

[0127] where R′ t is the current frame angle; R′ t-1 is the previous frame angle; R t is the rotation Euler angle of the current frame, corresponding to Figure 5 R' in

[0128] The rotation information of the current frame is calculated based on the difference between the previous and current frames. Here, if it is a handheld shooting, the gyroscope parameters in the handheld camera need to be further obtained, and then the relative value between the two is calculated to obtain the real information of the shooter. If it is a tripod shooting, there is no impact

[0129] Step 7: Robustness judgment: Whether R - R' is less than f

[0130] The difference e between R output by the neural network regressor and R' provided by the gyroscope is calculated through formula (6).

[0131] e = R - R′ (6)

[0132] Judge the magnitude of the difference e and the threshold f. When e is less than or equal to f, execute Step 8; when e is greater than f, execute Step 9

[0133] Step 8: Reconstruct the face with s, e, and R

[0134] If the difference between R and R’ is less than the threshold f, it is considered that the neural network regressor has high robustness, and the reconstruction is performed using s, e, and R given by the neural network regressor. Here, R can be slightly adjusted according to R’ given by the gyroscope to assist in improving the quality. The fine-tuning can be done through formula (7):

[0135] R = (1 - w) * R + w * R′ (7)

[0136] where w is the current confidence weight of the gyroscope.

[0137] The overall reconstruction formulas (8) and (9) are as follows:

[0138] S whole = ∑α i * s i + ∑β i * e i (8)

[0139] S final = R * S whole (9)

[0140] where α i is the coefficient of the shape basis, s i is the shape basis, β i is the coefficient of the expression basis, e i is the expression basis, and S whole is the 3D space coordinates to be obtained, and R is the pose parameter.

[0141] Step 9: Reconstruct the human face using s’, e, and R’.

[0142] If the difference between R and R’ is greater than the threshold f, it is considered that the neural network regressor has poor robustness. This situation is very likely to occur when there are occlusions, large-angle movements, or poor low-light image quality in the picture. At this time, it is considered that both the predicted R and the shape parameter s of the neural network regressor have deviated. Therefore, the reconstruction is performed using R’ provided by the gyroscope and the spare s’. At this time, the e of the previous frame is reused. The reconstruction formulas (10) and (11) are as follows:

[0143] S whole = ∑α′ i s′ i + ∑β i * e i (10)

[0144] S final = R′ * S whole (11)

[0145] where α′ i is the spare coefficient of the shape basis, s′ i is the spare shape basis, and ∑α′ is i ′ is the aforementioned reference shape information; β i is the coefficient of the expression basis, e i is the expression basis, S whole is the 3D spatial coordinates to be obtained, and R’ is the pose parameter.

[0146] As Figure 7 shown, in the case of face occlusion, an external device, i.e., the gyroscope parameter of the first device, is used, and the 3D reconstruction effect of the left figure assisted by it is better than that of the right figure without assistance. As Figure 8a shown, in the case of extremely dark conditions, i.e., the brightness value of the shooting environment is lower than the brightness threshold, where the brightness threshold is the lowest value of the brightness in a normal shooting scene, the 3D reconstruction assisted by the gyroscope parameter of the first device can also ensure the accuracy of 3D reconstruction. As Figure 8b shown, during large-angle movement, the right Figure Three 3D reconstruction effect assisted by the gyroscope of the first device is better than that of the left figure without assistance.

[0147] In the embodiment of this scenario, by combining the spatial angle information of the gyroscopes in the user's worn earphones and the mobile phone gyroscope, a 3D reconstruction effect with better robustness than the algorithm scenario with only single-frame image input can be obtained, which can improve the reliability of 3D reconstruction.

[0148] For the image processing method provided in the embodiment of this application, the execution subject can be an image processing device. In the embodiment of this application, taking the image processing device executing the image processing method as an example, the image processing device provided in the embodiment of this application is described.

[0149] See Figure 9 , the image processing device 900 may include:

[0150] A first acquisition module 901, configured to acquire a first image, where the first image includes a first region of a first object, the first object is the imaging of a first shooting object in the first image, and the first region is the imaging region of a first part of the first shooting object in the first image;

[0151] A second acquisition module 902, configured to acquire first pose information of the first part collected by a first device at a first moment, where the first device is worn on a second part of the first shooting object, and the first moment is the acquisition moment of the first image;

[0152] A first generation module 903, configured to perform 3D reconstruction on the first part according to the first pose information, and generate a second image including the 3D reconstructed first part.

[0153] In some embodiments, the first generation module includes:

[0154] A first acquisition unit, configured to acquire second pose information collected by the electronic device at the first moment when the electronic device is held by a user;

[0155] A determination unit, configured to determine first difference information between the first pose information and the second pose information;

[0156] A first generation unit, configured to perform three-dimensional reconstruction on the first part according to the first difference information, and generate a second image including the three-dimensionally reconstructed first part.

[0157] In some embodiments, the apparatus further includes:

[0158] An input module, configured to input the first region into a first model, and regress to obtain third pose information of the first part at the first moment;

[0159] A first determination module, configured to determine second difference information between the third pose information and target pose information of the first part at the first moment, where the target pose information is related to the first pose information;

[0160] The first generation module is specifically configured to, when the second difference information is greater than a threshold information, perform three-dimensional reconstruction on the first part according to the first pose information, and generate a second image including the three-dimensionally reconstructed first part.

[0161] In some embodiments, the apparatus further includes a second determination module, configured to perform at least one of the following:

[0162] When the electronic device is placed at a fixed position, determine the first pose information as the target pose information;

[0163] When the electronic device is held by a user, determine the first difference information as the target pose information, where the first difference information is a difference information between second pose information collected by the electronic device at the first moment and the first pose information.

[0164] In some embodiments, the apparatus further includes:

[0165] A third acquisition module, configured to acquire a confidence weight of the first device when the second difference information is less than or equal to the threshold information;

[0166] An adjustment module, configured to adjust the third pose information according to the confidence weight and the first pose information to obtain fourth pose information;

[0167] A second generation module, configured to perform three-dimensional reconstruction on the first area according to the fourth attitude information, and generate a third image including the three-dimensionally reconstructed first area.

[0168] The image processing device in the embodiments of the present application may be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device may be a terminal or other devices other than terminals. Exemplarily, the electronic device may be a mobile phone, a tablet computer, a laptop computer, a handheld computer, a vehicle-mounted electronic device, a Mobile Internet Device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., which are not specifically limited in the embodiments of the present application.

[0169] The image processing device in the embodiments of the present application may be a device with an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiments of the present application.

[0170] The image processing device provided in the embodiments of the present application can implement Figure 1 each process implemented by the method embodiments. To avoid repetition, it will not be elaborated here.

[0171] Optionally, as Figure 10 shown, the embodiments of the present application further provide an electronic device 1000, including a processor 1001 and a memory 1002. A program or instruction that can run on the processor 1001 is stored on the memory 1002. When the program or instruction is executed by the processor 1001, it implements each step of the above image processing method embodiments and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0172] It should be noted that the electronic devices in the embodiments of the present application include the above-mentioned mobile electronic devices and non-mobile electronic devices.

[0173] Figure 11 A schematic diagram of the hardware structure of an electronic device for implementing the embodiments of the present application.

[0174] The electronic device 1100 includes, but is not limited to, components such as a radio frequency unit 1101, a network module 1102, an audio output unit 1103, an input unit 1104, a sensor 1105, a display unit 1106, a user input unit 1107, an interface unit 1108, a memory 1109, and a processor 1110.

[0175] Those skilled in the art can understand that the electronic device 1100 may further include a power source (such as a battery) for supplying power to each component. The power source can be logically connected to the processor 1110 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. Figure 11 The structure of the electronic device shown does not limit the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0176] Among them, the processor 1110 is used for:

[0177] Obtain a first image, where the first image includes a first region of a first object, the first object is the imaging of a first shooting object in the first image, and the first region is the imaging region of the first part of the first shooting object in the first image;

[0178] Obtain first attitude information of the first part at a first moment collected by a first device, where the first device is worn on a second part of the first shooting object, and the first moment is the acquisition moment of the first image;

[0179] Perform three-dimensional reconstruction on the first part according to the first attitude information to generate a second image including the three-dimensionally reconstructed first part.

[0180] The electronic device provided by the embodiments of the present application can implement Figure 1 Each process implemented by the method embodiments will not be elaborated here to avoid repetition.

[0181] It should be understood that in the embodiments of the present application, the input unit 1104 may include a Graphics Processing Unit (GPU) 11041 and a microphone 11042. The graphics processor 11041 processes the image data of static pictures or videos obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The display unit 1106 may include a display panel 11061, and the display panel 11061 may be configured in the form of a liquid crystal display, an organic light emitting diode, etc. The user input unit 1107 includes at least one of a touch panel 11071 and other input devices 11072. The touch panel 11071 is also referred to as a touch screen. The touch panel 11071 may include two parts: a touch detection device and a touch controller. The other input devices 11072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and a joystick, which will not be elaborated here.

[0182] The memory 1109 can be used to store software programs and various data. The memory 1109 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data. Among them, the first storage area can store an operating system, application programs or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 1109 can include volatile memory or non-volatile memory, or the memory 1109 can include both volatile and non-volatile memory. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDR SDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synch link dynamic random access memory (SLDRAM), and a direct rambus random access memory (DRRAM). The memory 1109 in the embodiments of the present application includes, but is not limited to, these and any other suitable types of memory.

[0183] The processor 1110 may include one or more processing units; optionally, the processor 1110 integrates an application processor and a modem processor. Among them, the application processor mainly processes operations related to the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above modem processor may not be integrated into the processor 1110 either.

[0184] The embodiments of the present application also provide a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, it implements each process of the above embodiment of the image processing method and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0185] Among them, the processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory ROM, random access memory RAM, magnetic disks, or optical discs, etc.

[0186] Another embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement each process of the above-mentioned embodiment of the image processing method and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0187] It should be understood that the chip mentioned in the embodiments of the present application may also be referred to as a system-on-chip, system chip, chip system, or system-on-chip, etc.

[0188] The embodiments of the present application provide a computer program product. The program product is stored in a storage medium and is executed by at least one processor to implement each process of the above-mentioned embodiment of the image processing method and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0189] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without more limitations, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0190] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a computer software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present application.

[0191] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them fall within the protection scope of the present application.

Claims

1. An image processing method, characterized in that, including: obtaining a first image, where the first image includes a first region of a first object, the first object being an image of a first photographing object in the first image, and the first region being an imaging region of a first part of the first photographing object in the first image; obtaining first pose information of the first part at a first moment collected by a first device, where the first device is worn on a second part of the first photographing object, and the first moment is the collection moment of the first image; performing three-dimensional reconstruction on the first part according to the first pose information to generate a second image including the three-dimensionally reconstructed first part; The performing three-dimensional reconstruction on the first part according to the first pose information to generate a second image including the three-dimensionally reconstructed first part includes: when the electronic device is held by a user, obtaining second pose information collected by the electronic device at the first moment; determining first difference information between the first pose information and the second pose information; performing three-dimensional reconstruction on the first part according to the first difference information to generate a second image including the three-dimensionally reconstructed first part.

2. The method according to claim 1, wherein Before the performing three-dimensional reconstruction on the first part according to the first pose information to generate a second image including the three-dimensionally reconstructed first part, it further includes: inputting the first region into a first model to regress and obtain third pose information of the first part at the first moment; determining second difference information between the third pose information and target pose information of the first part at the first moment; The performing three-dimensional reconstruction on the first part according to the first pose information to generate a second image including the three-dimensionally reconstructed first part includes: when the second difference information is greater than threshold information, performing three-dimensional reconstruction on the first part according to the first pose information to generate a second image including the three-dimensionally reconstructed first part.

3. The method according to claim 2, wherein Before the determining second difference information between the third pose information and target pose information of the first part at the first moment, it further includes at least one of the following: when the electronic device is placed at a fixed position, determining the first pose information as the target pose information; when the electronic device is held by a user, determining the first difference information as the target pose information, where the first difference information is the difference information between the second pose information collected by the electronic device at the first moment and the first pose information.

4. The method according to claim 2, wherein After the determining second difference information between the third pose information and target pose information of the first part at the first moment, it further includes: when the second difference information is less than or equal to the threshold information, obtaining a confidence weight of the first device; adjusting the third pose information according to the confidence weight and the first pose information to obtain fourth pose information; performing three-dimensional reconstruction on the first part according to the fourth pose information to generate a third image including the three-dimensionally reconstructed first part.

5. An image processing apparatus, applied to an electronic device, characterized in that, including: A first acquisition module, configured to acquire a first image, where the first image includes a first region of a first object, the first object is the imaging of a first shooting object in the first image, and the first region is the imaging region of a first part of the first shooting object in the first image; A second acquisition module, configured to acquire first pose information of the first part at a first moment collected by a first device, where the first device is worn on a second part of the first shooting object, and the first moment is the acquisition moment of the first image; A first generation module, configured to perform three-dimensional reconstruction on the first part according to the first pose information, and generate a second image including the three-dimensionally reconstructed first part; The first generation module includes: A first acquisition unit, configured to acquire second pose information collected by the electronic device at the first moment when the electronic device is held by a user; A determination unit, configured to determine first difference information between the first pose information and the second pose information; A first generation unit, configured to perform three-dimensional reconstruction on the first part according to the first difference information, and generate a second image including the three-dimensionally reconstructed first part.

6. The device according to claim 5, characterized in that, It further includes: An input module, configured to input the first region into a first model, and regress to obtain third pose information and target information of the first part at the first moment, where the target information includes shape information; A first determination module, configured to determine second difference information between the third pose information and target pose information of the first part at the first moment, where the target pose information is related to the first pose information; The first generation module is specifically configured to perform three-dimensional reconstruction on the first part according to the first pose information and generate a second image including the three-dimensionally reconstructed first part when the second difference information is greater than a threshold information.

7. The device according to claim 6, characterized in that, It further includes a second determination module, configured to perform at least one of the following: When the electronic device is placed at a fixed position, determine the first pose information as the target pose information; When the electronic device is held by a user, determine the first difference information as the target pose information, where the first difference information is the difference information between the second pose information collected by the electronic device at the first moment and the first pose information.

8. The device according to claim 6, characterized in that, It further includes: A third acquisition module, configured to acquire a confidence weight of the first device when the second difference information is less than or equal to the threshold information; An adjustment module, configured to adjust the third pose information according to the confidence weight and the first pose information to obtain fourth pose information; A second generation module, configured to perform three-dimensional reconstruction on the first region according to the fourth pose information, and generate a third image including the three-dimensionally reconstructed first region.

Citation Information

Patent Citations

  • Gesture recognition system and method based on VR virtual office

    CN112083801A

  • Sign language data acquisition method, equipment and storage medium

    CN113496168A