Image processing method and model training method

By acquiring and rendering facial feature template information and facial contour template information, a digital human face image is generated, and a facial feature parameter determination model is trained. This solves the problems of low generation efficiency and severe homogenization in existing technologies, and realizes personalized and efficient digital human generation.

CN120808425BActive Publication Date: 2026-01-06ALIBABA (CHINA) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511320864.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-16
Publication Date
2026-01-06
Estimated Expiration
2045-09-16

AI Technical Summary

Technical Problem

In existing technologies, methods for generating digital humans are difficult to generate personalized digital human images efficiently, and the low generation efficiency leads to serious homogenization of digital human images, making it difficult to match the user's real image.

Method used

By acquiring the face image of the digital human to be generated and the template face image, using facial feature template information and facial contour template information, a facial feature mask is obtained and rendered onto the facial contour to generate the face image of the digital human. The facial feature parameters are then used to determine the model training and adjust the facial feature parameters to improve generation efficiency and personalization.

Benefits of technology

It enables more accurate capture of the personality information of the digital human to be generated, improves the personalization and efficiency of digital human generation, simplifies the operation process, and increases user satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808425B_ABST
    Figure CN120808425B_ABST
Patent Text Reader

Abstract

Embodiments of the present specification provide image processing methods and model training methods. The image processing method comprises: obtaining a face image of a to-be-generated digital human, and a template face image, wherein the template face image comprises five facial feature template information and face contour template information; based on the five facial feature template information, obtaining a face feature mask in the face image of the to-be-generated digital human; based on the face contour template information and the face image of the to-be-generated digital human, determining a face contour of the to-be-generated digital human; and rendering the face feature mask to the face contour of the to-be-generated digital human to generate a face image of a digital human. The image processing method provided by the present specification improves the efficiency of the digital human generation process and the personalization of the digital human generation result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments in this specification relate to the field of computer technology, and in particular to image processing methods and model training methods. Background Technology

[0002] A digital human is a virtual avatar generated by computer technology that possesses human-like characteristics in terms of vision, speech, and behavior. With the continuous development of computer technology, people are increasingly using digital humans in their daily lives.

[0003] In related technologies, the process of generating digital humans typically relies on user-uploaded videos to create corresponding digital humans, or on adjusting preset digital human templates to generate digital humans with a uniform appearance. However, these methods are inefficient, and the resulting digital humans are highly homogenized, making it difficult to closely resemble the user's real appearance. Summary of the Invention

[0004] In view of the above, embodiments of this specification provide an image processing method and a model training method. One or more embodiments of this specification also relate to an image processing apparatus, a model training apparatus, a computing device, a computer-readable storage medium, and a computer program product, to address the technical deficiencies existing in the prior art.

[0005] According to a first aspect of the embodiments of this specification, an image processing method is provided, comprising:

[0006] Obtain the face image of the digital human to be generated, and the template face image, wherein the template face image includes facial feature template information and facial contour template information;

[0007] Based on the facial feature template information, a facial feature mask is obtained from the face image of the digital human to be generated;

[0008] Based on the facial contour template information and the face image of the digital human to be generated, the facial contour of the digital human to be generated is determined;

[0009] The facial feature mask is rendered onto the facial contour of the digital human to be generated, thus generating the digital human's facial image.

[0010] According to a second aspect of the embodiments of this specification, a model training method is provided, comprising:

[0011] Acquire sample data, wherein the sample data includes sample facial feature information and sample facial feature parameters;

[0012] The sample facial feature information is input into the facial feature parameter determination model to obtain the predicted facial feature parameters output by the facial feature parameter determination model;

[0013] The model loss value is calculated based on the predicted facial feature parameters and the sample facial feature parameters;

[0014] The model parameters of the facial feature determination model are adjusted according to the model loss value, and the model training of the facial feature determination model continues until the model training stop condition is reached. The facial feature determination model is used to determine the facial feature parameters of the digital human to be generated, and the facial feature parameters are used to render the facial feature mask onto the facial contour of the digital human to be generated.

[0015] According to a third aspect of the embodiments of this specification, an image processing method is provided, applied to a cloud-side device, comprising:

[0016] The receiving end device sends a face image of the digital human to be generated, and obtains a template face image, wherein the template face image includes facial feature template information and facial contour template information;

[0017] Based on the facial feature template information, a facial feature mask is obtained from the face image of the digital human to be generated;

[0018] Based on the facial contour template information and the face image of the digital human to be generated, the facial contour of the digital human to be generated is determined;

[0019] The facial feature mask is rendered onto the facial contour of the digital human to be generated, thus generating the digital human's facial image;

[0020] The face image is sent to the edge device.

[0021] According to a fourth aspect of the embodiments of this specification, a task platform is provided, including a request interface and a response unit;

[0022] The request interface is used to receive the face image of the digital human to be generated sent by the end device;

[0023] The response unit is configured to acquire a template face image, wherein the template face image includes facial feature template information and facial contour template information; based on the facial feature template information, acquire a facial feature mask in the face image of the digital human to be generated; based on the facial contour template information and the face image of the digital human to be generated, determine the facial contour of the digital human to be generated; render the facial feature mask onto the facial contour of the digital human to be generated, and generate the face image of the digital human.

[0024] According to a fifth aspect of the embodiments of this specification, an image processing apparatus is provided, comprising:

[0025] The acquisition unit is configured to acquire a face image of the digital human to be generated and a template face image, wherein the template face image includes facial feature template information and facial contour template information.

[0026] The processing unit is configured to obtain a facial feature mask from the face image of the digital human to be generated based on the facial feature template information; and to determine the facial contour of the digital human to be generated based on the facial contour template information and the face image of the digital human to be generated.

[0027] The generation unit is configured to render the facial feature mask onto the facial contour of the digital human to be generated, thereby generating a facial image of the digital human.

[0028] According to a sixth aspect of the embodiments of this specification, a model training apparatus is provided, comprising:

[0029] The acquisition unit is configured to acquire sample data, wherein the sample data includes sample facial feature information and sample facial feature parameters;

[0030] The prediction unit is configured to input the sample facial information into the facial parameter determination model to obtain the predicted facial parameters output by the facial parameter determination model.

[0031] The calculation unit is configured to calculate the model loss value based on the predicted facial features parameters and the sample facial features parameters;

[0032] The adjustment unit is configured to adjust the model parameters of the facial feature parameter determination model according to the model loss value, and continue to train the facial feature parameter determination model until the model training stop condition is reached. The facial feature parameter determination model is used to determine the facial feature parameters of the digital human to be generated, and the facial feature parameters are used to render the facial feature mask onto the facial contour of the digital human to be generated.

[0033] According to a seventh aspect of the embodiments of this specification, a computing device is provided, comprising:

[0034] Memory and processor;

[0035] The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the above method.

[0036] According to an eighth aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores a computer program / instructions that, when executed by a processor, implement the steps of the above-described method.

[0037] According to a ninth aspect of the embodiments of this specification, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described method.

[0038] According to the image processing method provided in this manual, the facial image of the digital human to be generated, along with facial feature template information and facial contour information from a template facial image, is used to more accurately capture the individual characteristics of the digital human to be generated. Based on this, in the process of determining the facial contour using the facial contour information from the template facial image and the digital human's facial image, the generated digital image is structurally similar to the digital human to be generated, thereby enhancing the personalization of the digital human generation method. Furthermore, the method of generating a corresponding digital human's facial image from the digital human's facial image is simple to operate, highly efficient in digital human generation, and improves user satisfaction and experience. Attached Figure Description

[0039] Figure 1 A flowchart of an image processing method according to an embodiment of this specification is shown;

[0040] Figure 2 A schematic diagram of a model training method for a facial sensory parameter determination model provided in one embodiment of this specification is shown.

[0041] Figure 3 A schematic diagram of a sample data acquisition method according to an embodiment of this specification is shown;

[0042] Figure 4 A flowchart of a model training method provided according to an embodiment of this specification is shown;

[0043] Figure 5 A schematic diagram illustrating the processing procedure of an image processing method provided in one embodiment of this specification is shown;

[0044] Figure 6 This specification shows a schematic diagram of the structure of an image processing apparatus according to one embodiment;

[0045] Figure 7 This specification shows a schematic diagram of the structure of a model training device according to one embodiment.

[0046] Figure 8 This specification shows an architecture diagram of an image processing system provided in one embodiment;

[0047] Figure 9 A structural block diagram of a computing device according to an embodiment of this application is shown. Detailed Implementation

[0048] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.

[0049] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0050] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."

[0051] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this manual are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant regions, and corresponding operation portals are provided for users to choose to authorize or refuse.

[0052] First, the terms and concepts used in one or more embodiments of this specification will be explained.

[0053] Mask: The area outside the selection box. The masks referred to in this manual are layer masks. For example, a facial feature mask can be understood as a separate layer mask set for different facial features within a face.

[0054] Digital human: can be understood as a virtual character.

[0055] The image processing methods provided in this specification are applied to scenarios involving the generation of virtual avatars. For example, they are applied to scenarios where interaction occurs based on virtual avatars.

[0056] With the continuous development of computer technology, the generation of virtual avatars often relies on manual settings. For example, multiple layered materials are hand-drawn by technicians, and then disassembled, bound, and debugged using Live2D rendering technology to generate a digital human avatar. This process is time-consuming, has a high barrier to entry, and is highly repetitive, making it difficult to meet the efficient and personalized customization needs in the application of digital humans.

[0057] To address the above issues, related technologies can also acquire a user's video and generate a virtual avatar based on the video content. However, this method requires acquiring a considerable amount of video footage beforehand, resulting in long processing times and low efficiency. Alternatively, existing digital avatar templates can be selected. However, this method leads to highly homogenized digital avatars, and the generated avatars' movements are limited by the templates, making personalized adjustments difficult and resulting in a lack of diversity in the digital avatar designs.

[0058] Therefore, this specification provides an image processing method, and also relates to an image processing apparatus, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.

[0059] In one specific embodiment provided in this specification, an image processing method includes:

[0060] Obtain the face image of the digital human to be generated, and the template face image, wherein the template face image includes facial feature template information and facial contour template information;

[0061] Based on the facial feature template information, a facial feature mask is obtained from the face image of the digital human to be generated;

[0062] Based on the facial contour template information and the face image of the digital human to be generated, the facial contour of the digital human to be generated is determined;

[0063] The facial feature mask is rendered onto the facial contour of the digital human to be generated, thus generating the digital human's facial image.

[0064] In one specific embodiment provided in this specification, a face image containing a clear facial contour and facial feature information is acquired, and a template face image containing facial feature template information and facial contour information is acquired. Based on the facial feature template information containing different facial features in the template face image, facial feature masks corresponding to different facial features are acquired from the face image of the digital human to be generated. Furthermore, based on the facial contour template information in the template face image and the face image of the digital human to be generated, the facial contour of the digital human to be generated is determined. Based on the facial feature mask and the facial contour of the digital human to be generated, the face image of the digital human is generated.

[0065] According to a specific embodiment provided in this specification, using the facial image of the digital human to be generated, as well as the facial feature template information and facial contour information from the template facial image, allows for more accurate capture of the individual information of the digital human to be generated, enhancing the personalization of the digital human generation method. Furthermore, the method of generating a corresponding digital human's facial image based on the facial image of the digital human to be generated is simple to operate and highly efficient in digital human generation.

[0066] For ease of understanding, this instruction manual uses the following... Figure 1 The image methods mentioned above are explained in the manner shown.

[0067] See Figure 1 , Figure 1 A flowchart of an image processing method according to an embodiment of this specification is shown, including steps 102-108.

[0068] Step 102: Obtain the face image of the digital human to be generated, and the template face image.

[0069] The template face image includes facial feature template information and facial contour template information. Facial feature template information can be understood as including at least the positional information of the facial features in the template face image, while facial contour template information can be understood as including at least the size information of the facial contour in the template face image.

[0070] A facial image can be understood as an image containing a clear facial outline and information about facial features.

[0071] The facial image of the digital human to be generated can be understood as the image corresponding to the real facial image determined by the user who is using the digital human image for adaptive photography.

[0072] A template face image can be understood as a pre-defined standard face image in which facial features are evenly and symmetrically distributed.

[0073] The digital human to be generated can be understood as the facial image uploaded by a user who has the need to generate a virtual character.

[0074] The five senses can be understood as facial features including the eyes, nose, mouth, eyebrows, and ears.

[0075] In one specific embodiment of this specification, a real facial image uploaded by the user is used as the facial image of the digital human to be generated. A pre-generated facial image containing standard and symmetrical facial features is used as a template facial image.

[0076] It should be understood that the method of obtaining the facial image of the digital human to be generated by uploading a real face image taken by the user, as described in this specification, is a specific example. The method of obtaining the facial image of the digital human to be generated in this specification is not limited to this. Similarly, this specification does not limit the method of obtaining template facial images.

[0077] In one specific embodiment provided in this specification, obtaining the face image of the digital human to be generated includes:

[0078] Obtain the initial face image and determine the position information of the left eye and the right eye among the facial features in the initial face image;

[0079] The initial face image is adjusted so that the left eye position information and the right eye position information are at the same horizontal position, and the adjusted initial face image is used as the face image of the digital human to be generated.

[0080] The initial face image can be understood as the unprocessed face image uploaded by the user. The left eye position information can be understood as the information representing the left eye in the initial face image. Specifically, it can be understood as information stored in pixels representing the pixel position and coordinates of the left eye in the face. The right eye position information can be understood as the information representing the right eye in the initial face image. Specifically, it can be understood as information stored in pixels representing the pixel position and coordinates of the right eye.

[0081] In one specific embodiment of this specification, the initial face image uploaded by the user may be skewed due to differences in shooting angle or shooting method. However, the template face image used in this specification is a pre-defined standard face image with evenly distributed and symmetrical facial features. Therefore, to ensure that the digital human image corresponding to the user's face image generated subsequently is more accurate, the position information of each part of the facial features is used to determine whether the initial face image is skewed. For ease of judgment, the position information of the left and right eyes is used as the judgment criteria. By determining whether the horizontal positions of the pixels representing the horizontal position in the left and right eye position information are at the same horizontal coordinate, it is determined whether the initial face image is skewed. If the initial face image is skewed, the position information of the left and right eyes is adjusted to obtain a corrected initial face image, which is then used as the face image of the digital human to be generated.

[0082] For example, in the initial facial image, the horizontal positions of the pixels corresponding to the left and right pupils are determined. If the horizontal positions of the pixels corresponding to the left and right pupils are not at the same level, the initial facial image is considered skewed. Based on this, the horizontal positions of the pixels corresponding to the left and right eyes are rotated to ensure that, in the adjusted initial facial image, the horizontal positions of the left and right eyes are at the same level.

[0083] According to the specific implementation provided in this specification, by determining the horizontal position of the left eye and the horizontal position of the right eye in the initial face image, the initial face image is adjusted to eliminate the misalignment of the subsequently generated digital human caused by the distortion of the initial face image, thereby improving the naturalness and accuracy of the subsequently generated digital human image.

[0084] Step 104: Based on the facial feature template information, obtain the facial feature mask from the face image of the digital human to be generated.

[0085] In this context, a facial feature mask can be understood as a separate mask image generated for each part of the facial features (e.g., a mask image for the eyes, a mask image for the mouth, etc.). Within these separate mask images, adjustments to one feature do not affect the masks corresponding to other features. For example, adjusting the mask image for the eyes does not affect the mask image for the nose. Alternatively, a facial feature mask can also be understood as a set of mask images generated for each part of the facial features. This means that adjustments to different parts of the face are made within the same mask image.

[0086] In one specific embodiment of this specification, key points corresponding to each part of the facial features in the template face image are determined, so that the contours of each part of the facial features in the template face image can be obtained based on the key points of the facial features. Based on this, to facilitate the accuracy of subsequent segmentation of each part of the facial features in the face image of the digital human to be generated, according to the facial feature template information corresponding to each part of the facial features in the template face image, the corresponding facial feature parts are cropped from the face image of the digital human to be generated, resulting in a facial feature mask.

[0087] In one specific embodiment provided in this specification, based on the facial feature template information, a facial feature mask is obtained from the facial image of the digital human to be generated, including S1042-S1044:

[0088] S1042. Align the facial features of the digital human face image to be generated with the facial features of the template face image.

[0089] In one specific embodiment of this specification, there are discrepancies between the proportions of the facial features and the arrangement of the various parts of the facial features in the face image of the digital person to be generated and the template face image. Therefore, to facilitate the subsequent accurate generation of a digital person whose facial features match the proportions and arrangement of the facial features of the digital person to be generated based on the template face image, the facial features of the digital person to be generated are aligned with the facial features of the template face image by aligning the key points of the facial features. For example, the eyes in the face image of the digital person to be generated are aligned with the eyes in the template face image.

[0090] In one specific embodiment provided in this specification, aligning the facial features of the face image of the digital human to be generated with the facial features of the template face image includes:

[0091] Identify the facial features to be treated.

[0092] The facial features to be processed can be any one of the facial features.

[0093] Obtain the first key point information of the facial features to be processed in the face image of the digital human to be generated, and obtain the second key point information of the facial features to be processed in the template face image.

[0094] The first keypoint information can be understood as the positional information of each keypoint in the set of keypoints constituting the facial features to be processed in the face image of the digital human to be generated. The second keypoint information can be understood as the positional information of each keypoint in the set of keypoints constituting the same facial features as those to be processed in the template face image.

[0095] Determine the transformation relationship between the first key point information and the second key point information.

[0096] The transformation relationship can be understood as the translation relationship between the first key point information and the second key point information, and / or the scaling relationship.

[0097] The face image of the digital human to be generated is adjusted according to the transformation relationship so that the facial features to be processed in the face image of the digital human to be generated are aligned with the facial features to be processed in the template face image.

[0098] Taking the eyes as an example, the above content will be explained and illustrated by way of example.

[0099] For example, in the face image of the digital human to be generated, a first set of keypoints constituting the eyes is determined, and first keypoint information corresponding to this first set of keypoints is determined. The first keypoint information can represent the position and size of the eyes in the face image of the digital human to be generated. Similarly, in the template face image, a second set of keypoints constituting the eyes is determined, and second keypoint information corresponding to this second set of keypoints is determined. The second keypoint information can represent the position and size of the eyes in the template face image. The offset between the position of the eyes in the face image of the digital human to be generated represented by the first keypoint information and the position of the eyes in the template face image represented by the second keypoint information is calculated, as is the scaling between the size of the eyes in the face image of the digital human to be generated represented by the first keypoint information and the size of the eyes in the template face image represented by the second keypoint information. Based on the offset and scaling, the first keypoint information is adjusted so that the adjusted first keypoint information corresponds to the second keypoint information. Based on this, the eye region in the face image of the digital human to be generated is aligned with the eye region in the template face image.

[0100] The treatment methods for other parts of the face are consistent with those for the eyes described above, and will not be repeated in this instruction manual.

[0101] According to the specific implementation method provided in this specification, the change relationship of key points in any part of the facial features to be processed is calculated, so that the facial features of the digital human to be generated can be accurately marked on the corresponding position of the template facial image, thereby improving the accuracy of subsequent digital human generation.

[0102] S1044. Based on the facial feature template information, extract the facial features from the face image of the digital human to be generated after the facial features are aligned, and obtain a facial feature mask.

[0103] In one specific embodiment of this specification, since the facial features of the digital human to be generated are already aligned with those of the template facial features in the face image of the digital human to be generated (e.g., the position and proportion of the eyes in the face image of the digital human to be generated are consistent with the position and proportion of the eyes in the template facial image), to allow each part of the facial features to be adjusted independently, the cropping boundaries of each part of the facial features are determined by using the facial feature template information in the template facial image as a reference and combining the features of each part of the facial features in the face image of the digital human to be generated. Based on these cropping boundaries, a facial feature mask corresponding to the digital human to be generated is cropped. The resulting facial feature mask includes facial feature masks obtained by cropping each part of the facial features individually. (For example, the facial feature mask includes at least a mask corresponding to the nose, a mask corresponding to the eyes, and a mask corresponding to the mouth, etc.)

[0104] According to this implementation method, in the face image of the digital human to be generated after alignment, the facial feature template information is used to crop the image, thereby avoiding the offset between the facial features in the template face image and the facial features in the actual face image due to errors in the alignment process.

[0105] In one specific embodiment provided in this specification, in the face image of the digital human to be generated after the facial features are aligned, the facial features are cropped according to the facial feature template information, and facial feature masks of different parts of the corresponding facial features are obtained respectively.

[0106] To facilitate understanding, the following example is used in this manual to explain how to obtain a facial feature mask.

[0107] Taking the acquisition of a facial feature mask corresponding to the eye area as an example, template information related to the eye area in the template face is used, such as the position, proportion, and angle of the eyes in the template face image. Based on the above information, a facial feature mask is extracted from the eye area of ​​the digital human's face image to be generated, aligned with the eye area.

[0108] For facial feature masks of other parts of the face, the same method as for the facial feature mask of the eye part is used to crop them, thus obtaining multiple facial feature masks corresponding to different facial feature parts.

[0109] According to the specific implementation method provided in this specification, by aligning the facial features of the digital human to be generated with the template image, the personalized facial features of the digital human to be generated are mapped onto the standard coordinate system corresponding to the template face image, ensuring that the position, proportion, and angle of the facial features conform to the specifications of digital human facial structure. Furthermore, for each facial feature, a facial feature mask generated based on the facial feature template information is used to clearly define the corresponding area of ​​each facial feature in the face image of the digital human to be generated, thereby achieving separation of each facial feature from other facial features. This facilitates subsequent adjustments to any facial feature without misoperating other features, improving the stability of facial feature adjustments.

[0110] Step 106: Determine the facial contour of the digital human to be generated based on the facial contour template information and the face image of the digital human to be generated.

[0111] In one specific embodiment of this specification, the facial contour corresponding to the face image of the digital human to be generated is determined based on the size of the facial contour contained in the facial contour template information of the template face image and the position of the key points constituting the facial contour.

[0112] For example, the facial contour dimensions corresponding to the template face image are determined, and the facial contour dimensions corresponding to the face image of the digital human to be generated are also determined. The facial contour dimensions corresponding to the face image of the digital human to be generated are adjusted so that the facial contour dimensions corresponding to the template face image are greater than or equal to the facial contour dimensions corresponding to the face image of the digital human to be generated. Based on this, using the facial contour dimensions corresponding to the template face image as the cropping basis, the facial contour of the digital human to be generated is cropped from the face image of the digital human to be generated, so that the cropped facial contour of the digital human to be generated includes the complete facial contour of the digital human to be generated.

[0113] Step 108: Render the facial feature mask onto the facial contour of the digital human to be generated, and generate the digital human's facial image.

[0114] In one specific embodiment of this specification, as described above, the facial feature mask is obtained based on the facial image of the digital human to be generated and the template facial image. To obtain a complete facial image of the digital human, the facial features in the facial feature mask are mapped onto the facial contour of the digital human to be generated, resulting in a facial image that conforms to the facial features of the digital human to be generated.

[0115] In one specific embodiment provided in this specification, the facial feature mask is rendered onto the facial contour of the digital human to be generated, including steps S1082-S1086:

[0116] S1082. Determine the facial features parameters of the digital human to be generated.

[0117] In one specific embodiment of this specification, to facilitate rendering the facial feature mask onto the facial contour of the digital human to be generated, the facial feature parameters of the digital human to be generated are determined, so as to determine the offset and scaling of the facial feature mask relative to the facial features of the template facial image. This facilitates the subsequent rendering of the facial feature mask onto the facial contour of the digital human to be generated based on the offset and scaling.

[0118] The facial feature parameters of the digital human to be generated are used to determine the position of the facial feature mask in the facial contour of the digital human to be generated.

[0119] In one specific embodiment provided in this specification, determining the facial features parameters of the digital human to be generated includes:

[0120] In the facial contour of the digital human to be generated, the facial features information of the digital human to be generated is determined;

[0121] The facial features information of the digital human to be generated is input into the facial feature parameter determination model to obtain the facial feature parameters of the digital human to be generated output by the facial feature parameter determination model.

[0122] For example, in the facial contour of the digital human to be generated, the position and size information of each facial feature of the digital human to be generated are determined. The position and size information of each facial feature is input into the facial feature parameter determination model. Based on the facial feature parameter determination model, the offset and scaling of each facial feature of the digital human to be generated relative to the corresponding facial feature in the template face image are obtained.

[0123] According to the specific implementation method provided in this specification, the facial features information of the digital human to be generated is input into the facial feature parameter determination model to obtain the facial feature parameters of the digital human to be generated. Compared with the method of manually annotating facial feature parameters, the facial feature parameter determination model makes the obtained facial feature parameters more accurate.

[0124] In one specific embodiment provided in this specification, see [link to specific embodiment]. Figure 2 . Figure 2 This diagram illustrates a model training method for a facial sensory parameter determination model provided in one embodiment of this specification, as shown below. Figure 2 As shown, the method includes steps 202-208:

[0125] Step 202: Obtain sample data.

[0126] The sample data includes information about the facial features of the sample and parameters of the facial features.

[0127] In one specific embodiment of this specification, sample facial information and sample facial parameters are obtained, and sample pairs are obtained based on their correspondence with the sample facial information and sample facial parameters, and the sample pairs are used as sample data.

[0128] In one specific embodiment provided in this specification, the sample data is obtained in the following manner:

[0129] Obtain the template face image;

[0130] The template face image includes initial facial feature information and initial facial feature parameters corresponding to the initial facial feature information.

[0131] The initial facial features parameters are randomly adjusted to obtain sample facial features parameters, and the sample facial features information corresponding to the sample facial features parameters is obtained.

[0132] In one specific embodiment of this specification, any one of the facial feature parameters in the sample is randomly adjusted (e.g., the left eye is shifted 30mm horizontally), and the initial facial feature parameters are re-determined based on the adjusted parameters to obtain the sample facial feature information. The facial feature parameters and the sample facial feature information are used as a sample pair.

[0133] The sample facial features parameters and data are saved to obtain the sample data.

[0134] For ease of understanding, this instruction manual combines... Figure 3 The content shown explains the methods for obtaining sample data mentioned above.

[0135] Figure 3 A schematic diagram of a sample data acquisition method provided according to an embodiment of this specification is shown.

[0136] like Figure 3 As shown, Figure 3 The leftmost image shown is used as a template face image. The initial facial feature parameters corresponding to the face image in the leftmost image can be understood as 0.

[0137] In one example, based on the image on the left, the left eye parameters corresponding to the left eye of the template face image are adjusted, for example, the horizontal parameter of the left eye (left_eye) is adjusted. x Shifting left by 30 (-30) yields the second image from the left. In the second image from the left, due to the different facial features (face_param) parameters (excluding the left eye parameter) for the template face image... others The parameters for the other facial features of the template face image shown on the left (second from the left) have not been adjusted, therefore the parameters for the other facial features are (face_param). others=0). Based on this, the key points of the facial features are obtained from the adjusted template face image (second image from the left), and the data corresponding to the key points of the facial features are used as facial feature data, and the facial feature data and facial feature parameters are used as sample data.

[0138] In another example, based on the image on the left, the mouth parameters corresponding to the mouth in the template face image are adjusted. For example, the mouth size is scaled. As shown in the third image on the left, the mouth size is scaled by 30 compared to the original mouth size. z The image is enlarged to a scale of 30. In the three images on the left, all facial features (face_param) except for the mouth parameter are preserved in the template face image. others (face_param) remains unchanged others =0). Based on this, the key points of the facial features are obtained from the adjusted template face image (third image on the left), and the data corresponding to the key points of the facial features are used as facial feature data, and the facial feature data and facial feature parameters are used as sample data.

[0139] In another example, based on the three images on the left, the position of the mouth corresponding to the mouth in the template face image is offset. As shown in the image on the right, the vertical position corresponding to the mouth is shifted down by 30 units compared to the original mouth position. y =-30). And in the rightmost image, maintain all facial features parameters (face_param) except for the size and position of the mouth in the template face image. others (face_param) remains unchanged others =0). Based on this, the key points of the facial features are obtained from the adjusted template face image (right image), and the data corresponding to the key points of the facial features are used as facial feature data, and the facial feature data and facial feature parameters are used as sample data.

[0140] Step 204: Input the sample facial feature information into the facial feature parameter determination model to obtain the predicted facial feature parameters output by the facial feature parameter determination model.

[0141] In one specific embodiment of this specification, taking a multilayer perceptron as an example, the facial sensory parameter determination model takes sample facial sensory information as input and outputs predicted facial sensory parameters.

[0142] Step 206: Calculate the model loss value based on the predicted facial features parameters and the sample facial features parameters.

[0143] In one specific embodiment of this specification, to facilitate determining the accuracy of the model's predictions based on the facial sensory parameters, the difference between the predicted facial sensory parameters and the sample facial sensory parameters is compared, and the loss value is determined by the magnitude of the difference. The greater the difference, the greater the loss value.

[0144] Step 208: Adjust the facial sensory parameters to determine the model parameters based on the model loss value, and continue training the model determined by the facial sensory parameters until the model training stops.

[0145] Model parameters can be understood as the model's weights, biases, etc. Adjusting these parameters improves the model's predictive ability.

[0146] In one specific embodiment of this specification, the influence of the loss on each parameter of the model is calculated based on the loss value using the backpropagation algorithm, and then the model parameters are adjusted so that the predicted facial features parameters output by the adjusted model parameters are consistent with the sample facial features parameters.

[0147] To facilitate understanding, the training process of the model for determining facial sensory parameters mentioned above is illustrated in the following examples.

[0148] For example, the facial sensory parameters and information from the obtained sample data are normalized and cleaned, and the sample data is randomly divided into training and validation sets. The facial sensory parameters are input into a multilayer perceptron (MLP), which is then used to train the MLP to predict the facial sensory parameters. The predicted parameters are compared with the actual sample parameters to calculate the loss value of the MLP. The parameters of the MLP are adjusted based on the loss value so that the error between the predicted facial sensory parameters and the actual sample parameters is less than a set error, thus meeting the training stopping condition of the MLP. The adjusted MLP is then used as the model for determining the facial sensory parameters.

[0149] S1084. Based on the position of the facial feature mask in the facial contour of the digital human to be generated, the facial features in the facial contour of the digital human to be generated are masked to obtain the image of the face to be synthesized.

[0150] In one specific embodiment of this specification, in order to avoid interference from redundant information in the facial contour of the digital human to be generated, and to facilitate accurate positioning of the facial features in the various parts of the face according to the digital human to be synthesized, the facial features in the facial contour of the digital human to be generated are cropped according to the facial feature parameters, or the facial features are covered in the form of a facial texture map, thereby obtaining the face image to be synthesized.

[0151] For example, the facial contours of the digital human to be generated contain the base of facial features. If facial feature masking is not performed on this basis, there is a risk of generating dual facial features during the subsequent process of generating the digital human from the image of the face to be synthesized.

[0152] In one specific embodiment of this specification, based on the position of the facial feature mask within the facial contour of the digital human to be generated, the facial features in the facial contour of the digital human to be generated are masked, including:

[0153] The facial features to be processed are identified, and the positions of the facial features to be processed are determined based on the positions of the facial feature mask in the facial contour of the digital human to be generated.

[0154] Wherein, the facial features to be processed are any one of the facial features;

[0155] Determine the reference occlusion information corresponding to the facial features to be processed;

[0156] Based on the location of the facial features to be processed, the facial features to be processed are occluded according to the reference occlusion information.

[0157] The reference occlusion information can be understood as reference information used to occlude different parts of the face. For example, when the face to be processed is the mouth, the reference occlusion information could be the skin texture corresponding to the mouth template in the face feature template information. In this case, the mouth is occluded according to the skin texture corresponding to the mouth template. As another example, when the face to be processed is the nose, the reference occlusion information could be the skin texture corresponding to the nose template in the face feature template information. In this case, the nose is occluded according to the skin texture corresponding to the nose template.

[0158] In one specific embodiment of this specification, the facial features to be processed are determined, and the position of the outline of the facial features to be processed within the facial contour of the digital human to be generated is determined based on the position of the facial feature mask within the facial contour of the digital human to be generated. For example, based on the position of the mouth in the facial feature mask within the facial contour of the digital human to be generated, the position of the key point corresponding to the mouth within the facial contour of the digital human to be generated is determined. The mouth is then masked based on its position within the facial contour (e.g., by cropping the mouth to make its corresponding position a transparent layer, or by filling the mouth with a skin texture, etc.). The facial features of the digital human to be generated are processed sequentially in the above manner until all facial features of the digital human to be generated are completely masked using the method shown for the mouth. The parts of the facial contour that do not include the facial features are masked with black pixels, thereby obtaining the image of the synthesized face.

[0159] To facilitate understanding, the following examples are used in this instruction manual to explain the aforementioned facial feature obscuring.

[0160] In one example, the eyes are taken as the facial feature to be processed. Based on the occlusion information corresponding to the eyes in the reference occlusion information, the eyes are occluded so that the displayed content for the eye area is skin texture, while other non-eye facial features are erased, thus obtaining the composite face image corresponding to the eye area. For the other facial features in the composite face image besides the eyes, they are determined in the same way as the eyes, thus obtaining the composite face images for each of the corresponding facial features.

[0161] According to the specific implementation method provided in this specification, by masking the facial features to be processed, the part to be processed in the facial features is accurately located, and the part to be processed in the facial features is separated from other parts of the facial features. This avoids the influence of other non-processed parts during the generation of the facial features of the digital human, thereby improving the accuracy of digital human generation and making it more in line with the facial features of the digital human to be generated.

[0162] S1086. Render the facial feature mask onto the face image to be synthesized.

[0163] In one specific embodiment of this specification, a facial feature mask corresponding to each part of the facial features is mapped onto the image of the face to be synthesized. For example, based on the location of each part of the facial features in the image of the face to be synthesized, the facial feature mask is mapped according to the corresponding position.

[0164] In one specific embodiment of this specification, rendering the facial feature mask onto the image of the face to be synthesized includes:

[0165] In the image of the face to be synthesized, determine the facial features to be rendered;

[0166] Wherein, the facial features to be rendered are any part of the facial features;

[0167] Based on the facial features to be rendered, a facial feature mask corresponding to the facial features to be rendered is determined in the facial feature mask.

[0168] The mask of the facial features to be processed is rendered onto the facial features to be rendered in the image of the face to be synthesized.

[0169] In one specific embodiment of this specification, in the image of a face to be synthesized, any part of the facial features is determined as the facial feature to be rendered. A mask for the facial feature to be processed is determined in the facial feature mask, and this mask is overlapped with the corresponding facial feature in the image of the face to be synthesized. Based on this, the mask for the facial feature to be processed is rendered onto the image of the face to be synthesized.

[0170] The above rendering operation is performed on any facial feature to render the facial feature mask onto the image of the face to be synthesized.

[0171] In one specific embodiment of this specification, after rendering the facial feature mask onto the facial contour of the digital human to be generated, the method further includes:

[0172] Repair or fill in the facial contours of the digital human to be generated.

[0173] In one specific embodiment of this specification, after rendering the facial feature mask onto the facial contour of the digital human to be generated, to avoid errors in the above process that could cause the facial contour of the digital human to be generated to be difficult to fit the facial contour of the template facial image, thus leading to defects in the dynamic driving process of the digital human to be generated, the facial contour of the digital human to be generated is repaired or filled in so that the facial contour of the digital human to be generated fits the facial contour of the template facial image more naturally.

[0174] The repair or filling of the facial contours of the digital human to be generated can be understood as, for example, applying textures to the facial contours of the digital human to be generated based on its facial texture. This specification does not limit the implementation method for repairing or filling the facial contours of the digital human to be generated.

[0175] According to the specific implementation method provided in this specification, the obtained facial feature mask is repaired or filled, so that the resulting digital human is more closely aligned with the face of the digital human to be generated, thereby avoiding the exposure of flaws during the dynamic driving process of the generated digital human.

[0176] In one specific embodiment of this specification, the face image of the digital human to be generated further includes the hair of the digital human to be generated;

[0177] Before generating the digital human's facial image, the following steps are included:

[0178] Hair segmentation is performed on the face image of the digital human to be generated to obtain a hair mask;

[0179] The hair mask is rendered onto the face image to generate a face image of a digital human containing hair.

[0180] In one specific embodiment of this specification, when the facial contour of the digital human to be generated is determined, and the face image of the digital human to be generated includes hair, the hair in the face image of the digital human to be generated is segmented to obtain a hair mask. Based on this, the hair mask is rendered onto the facial contour of the digital human to be generated, resulting in the facial contour of the digital human to be generated that includes a hair layer.

[0181] According to the specific implementation method provided in this specification, by obtaining the hair mask of the digital human to be generated, and rendering the hair mask and the facial feature mask together onto the facial contour of the digital human to be generated, the final generated image of the digital human is consistent with the image of the digital human to be generated, thereby realizing the personalized generation of the digital human to be generated.

[0182] In the image processing method described above, the initial face image of the digital human to be generated is corrected based on the template face image, making the subsequently generated digital human more closely match the template face image. Furthermore, using the facial feature template information and facial contour information of the template face image, a facial feature mask and facial contour are extracted from the face image of the digital human to be generated. This ensures that the same template face image is used as the basis for subsequent digital human generation, eliminating the need to set a separate template face image and improving the generation efficiency of the template face image. Moreover, since the digital human is generated based on the face image of the digital human to be generated, the generated digital human closely matches the facial features of the digital human to be generated, thereby achieving personalized customization of the generated digital human and obtaining a digital human that accurately matches the face image of the digital human to be generated.

[0183] In one specific embodiment of this specification, in conjunction with the appendix Figure 4 This paper takes the application of the model training method provided in this manual in determining the facial features parameters of the digital human to be generated as an example to illustrate the model training method. Figure 4 A flowchart of a model training method according to an embodiment of this specification is shown, including steps 402-408.

[0184] Step 402: Obtain sample data, wherein the sample data includes sample facial feature information and sample facial feature parameters.

[0185] In one specific embodiment of this specification, the sample data is obtained in the following manner:

[0186] Obtain a template face image, wherein the template face image includes initial facial feature information and initial facial feature parameters corresponding to the initial facial feature information;

[0187] The initial facial features parameters are randomly adjusted to obtain sample facial features parameters, and the sample facial features information corresponding to the sample facial features parameters is obtained.

[0188] The sample facial features parameters and data are saved to obtain the sample data.

[0189] Step 404: Input the sample facial feature information into the facial feature parameter determination model to obtain the predicted facial feature parameters output by the facial feature parameter determination model.

[0190] Step 406: Calculate the model loss value based on the predicted facial features parameters and the sample facial features parameters.

[0191] Step 408: Adjust the facial sensory parameters to determine the model parameters based on the model loss value, and continue training the model determined by the facial sensory parameters until the model training stops.

[0192] The facial feature parameter determination model is used to determine the facial feature parameters of the digital human to be generated, and the facial feature parameters are used to render the facial feature mask onto the facial contour of the digital human to be generated.

[0193] In one specific embodiment provided in this specification, sample data containing sample facial feature information and sample facial feature parameters is obtained. The sample facial feature information is used as input to the facial feature parameter determination model to obtain the predicted facial feature parameters output by the model. The loss value of the model is calculated based on the predicted facial feature parameters and the sample facial feature parameters. The facial feature parameter determination model is adjusted based on the loss value until the model training stopping condition is reached. At this point, the facial feature parameters of the digital human to be generated are obtained based on the facial feature parameter determination model that has reached the model training stopping condition. The facial feature mask of the digital human to be generated is then rendered onto the facial contour of the digital human to be generated using the facial feature parameters of the digital human to be generated.

[0194] To facilitate understanding of the image processing methods provided in this manual, the following description is in conjunction with the appendix. Figure 5 Taking the application of the image processing method provided in this specification on a server as an example, the image processing method will be further explained. Figure 5 This specification illustrates a schematic diagram of the processing procedure of an image processing method according to an embodiment of the present specification, which specifically includes the following contents.

[0195] exist Figure 5 In the process, the system acquires the user's facial image and determines the key point information corresponding to the facial features. The key point information corresponding to the eyes is rotated to achieve facial correction. The user's facial image is adjusted so that the key points corresponding to the left and right eyes are horizontal, and the adjusted facial image is used as the facial image for generating the digital human.

[0196] In the facial image of the digital human to be generated, key points corresponding to each facial feature are detected separately. From the key points corresponding to each facial feature, a subset of key points are selected as markers for that feature (i.e., first key points). In the template facial image, markers corresponding to the same facial features as those in the digital human's facial image are determined (i.e., second key points). The first key point information corresponding to the first key point is compared with the second key point information corresponding to the second key point. For any key point corresponding to a facial feature, the scaling and translation of the first key point information in the digital human's facial image relative to the second key point information are calculated using least squares. This ensures that the key points corresponding to the same facial features are aligned between the digital human's facial image and the template facial image. This facilitates subsequent cropping of the digital human's facial image using the second key point information from the template facial image.

[0197] In the template face image, for each part of the facial features, the facial feature template information is used to extract the facial feature part from the aligned face image of the digital human to be generated, and the facial features are transferred to the texture to obtain the facial feature mask corresponding to that facial feature part.

[0198] Furthermore, by using the key points of the facial contour corresponding to the facial contour of the digital human to be generated, and the facial contour template information in the template face image, the facial contour of the digital human to be generated is transferred to the texture. That is, the facial contour of the digital human to be generated is transferred to the template face image and placed at the bottom layer of the template face image, so that the current template face image contains the facial contour of the digital human to be generated without facial features, i.e., the texture does not contain facial feature rendering.

[0199] When the template face image contains a facial contour of the digital human to be generated that does not contain facial features, a model (e.g., a multilayer perceptron) is used to process the facial feature parameters based on the facial feature parameters. This process is then used to mask the facial features in the facial contour of the digital human to be generated (e.g., by smoothing out the facial features on the facial contour), resulting in the image of the digital human to be synthesized. The facial feature mask corresponding to the facial feature area is then rendered onto the image of the digital human to be synthesized (e.g., a combination of facial feature and facial contour textures). The facial features in the rendered image of the digital human to be synthesized are then smoothed out to prevent imperfections from being exposed during dynamic driving. Finally, the hair mask corresponding to the face image of the digital human to be generated is obtained and combined with the facial features in the rendered image of the digital human to be synthesized for final rendering, generating the face image of the digital human.

[0200] The generated digital human's facial image is adjusted, such as by applying color transfer to the facial features and exposed skin of the generated digital human based on the skin color, to improve the skin color consistency between the final digital human's facial image and the original digital human's facial image, thus obtaining a digital human that is consistent with the original digital human's facial image.

[0201] Regarding the above, the process of obtaining facial feature parameters by determining the model using facial feature parameters is described in the above-mentioned section of this manual. Figure 2 The training method shown will not be described in detail in this manual.

[0202] According to the image processing method provided in this specification, generating a corresponding digital human from an acquired facial image improves the efficiency of digital human generation compared to generating digital human content from video. Furthermore, generating a digital human based on the facial image of the person to be generated ensures that the generated digital human's image closely matches the actual appearance of the person to be generated, thereby enabling personalized expression of the generated digital human.

[0203] Corresponding to the above method embodiments, this specification also provides embodiments of an image processing apparatus. Figure 6 A schematic diagram of the structure of an image processing apparatus provided in one embodiment of this specification is shown. Figure 6 As shown, the device includes:

[0204] The acquisition unit 602 is configured to acquire a face image of the digital human to be generated and a template face image, wherein the template face image includes facial feature template information and facial contour template information.

[0205] The processing unit 604 is configured to obtain a facial feature mask from the face image of the digital human to be generated based on the facial feature template information; and to determine the facial contour of the digital human to be generated based on the facial contour template information and the face image of the digital human to be generated.

[0206] The generation unit 606 is configured to render the facial feature mask onto the facial contour of the digital human to be generated, thereby generating a facial image of the digital human.

[0207] Optionally, the generation unit 606 is further configured to:

[0208] The facial features parameters of the digital human to be generated are determined, wherein the facial features parameters of the digital human to be generated are used to determine the position of the facial feature mask in the facial contour of the digital human to be generated;

[0209] Based on the position of the facial feature mask in the facial contour of the digital human to be generated, the facial features in the facial contour of the digital human to be generated are masked to obtain the image of the face to be synthesized.

[0210] The facial feature mask is rendered onto the image of the face to be synthesized.

[0211] Optionally, the generation unit 606 is further configured to:

[0212] In the facial contour of the digital human to be generated, the facial features information of the digital human to be generated is determined;

[0213] The facial features information of the digital human to be generated is input into the facial feature parameter determination model to obtain the facial feature parameters of the digital human to be generated output by the facial feature parameter determination model.

[0214] Optionally, the device further includes a training module configured to:

[0215] Acquire sample data, wherein the sample data includes sample facial feature information and sample facial feature parameters;

[0216] The sample facial feature information is input into the facial feature parameter determination model to obtain the predicted facial feature parameters output by the facial feature parameter determination model;

[0217] The model loss value is calculated based on the predicted facial feature parameters and the sample facial feature parameters;

[0218] The model parameters are determined by adjusting the facial sensory parameters based on the model loss value, and the model is trained by continuing to train the model based on the facial sensory parameters until the model training stops.

[0219] Optionally, the training module is further configured to:

[0220] Obtain a template face image, wherein the template face image includes initial facial feature information and initial facial feature parameters corresponding to the initial facial feature information;

[0221] The initial facial features parameters are randomly adjusted to obtain sample facial features parameters, and the sample facial features information corresponding to the sample facial features parameters is obtained.

[0222] The sample facial features parameters and data are saved to obtain the sample data.

[0223] Optionally, the generation unit 606 is further configured to:

[0224] The facial features to be processed are determined, and the position of the facial features to be processed is determined based on the position of the facial feature mask in the facial contour of the digital human to be generated, wherein the facial features to be processed are any part of the facial features;

[0225] Determine the reference occlusion information corresponding to the facial features to be processed;

[0226] Based on the location of the facial features to be processed, the facial features to be processed are masked according to the reference value masking information.

[0227] Optionally, the generation unit 606 is further configured to:

[0228] In the face image to be synthesized, the facial features to be rendered are determined, wherein the facial features to be rendered are any part of the facial features;

[0229] Based on the facial features to be rendered, a facial feature mask corresponding to the facial features to be rendered is determined in the facial feature mask.

[0230] The mask of the facial features to be processed is rendered onto the facial features to be rendered in the image of the face to be synthesized.

[0231] Optionally, the processing unit 604 is further configured to:

[0232] Align the facial features of the digital human to be generated with the facial features of the template facial image;

[0233] Based on the facial feature template information, facial features are extracted from the face image of the digital human to be generated after the facial features are aligned, and a facial feature mask is obtained.

[0234] Optionally, the processing unit 604 is further configured to:

[0235] The facial features to be processed are determined, wherein the facial features to be processed are any part of the facial features;

[0236] Obtain the first key point information of the facial features to be processed in the face image of the digital human to be generated, and obtain the second key point information of the facial features to be processed in the template face image;

[0237] Determine the transformation relationship between the first key point information and the second key point information;

[0238] The face image of the digital human to be generated is adjusted according to the transformation relationship so that the facial features to be processed in the face image of the digital human to be generated are aligned with the facial features to be processed in the template face image.

[0239] Optionally, the face image of the digital human to be generated may also include the hair of the digital human to be generated;

[0240] The generation unit 606 is further configured to:

[0241] Hair segmentation is performed on the face image of the digital human to be generated to obtain a hair mask;

[0242] The hair mask is rendered onto the face image to generate a face image of a digital human containing hair.

[0243] Optionally, the acquisition unit 602 is further configured to:

[0244] Obtain the initial face image and determine the position information of the left eye and the right eye among the facial features in the initial face image;

[0245] The initial face image is adjusted so that the left eye position information and the right eye position information are at the same horizontal position, and the adjusted initial face image is used as the face image of the digital human to be generated.

[0246] The above is an illustrative scheme of an image processing apparatus according to this embodiment. It should be noted that the technical solution of this image processing apparatus and the technical solution of the image processing method described above belong to the same concept. For details not described in detail in the technical solution of the image processing apparatus, please refer to the description of the technical solution of the image processing method described above.

[0247] Corresponding to the above method embodiments, this specification also provides embodiments of a model training device. Figure 7 A schematic diagram of a model training apparatus according to one embodiment of this specification is shown. Figure 7 As shown, the device includes:

[0248] The acquisition unit 702 is configured to acquire sample data, wherein the sample data includes sample facial feature information and sample facial feature parameters;

[0249] The prediction unit 704 is configured to input the sample facial information into the facial parameter determination model to obtain the predicted facial parameters output by the facial parameter determination model.

[0250] The calculation unit 706 is configured to calculate the model loss value based on the predicted facial features parameters and the sample facial features parameters;

[0251] The adjustment unit 708 is configured to adjust the model parameters of the facial feature parameter determination model according to the model loss value, and continue to train the facial feature parameter determination model until the model training stop condition is reached. The facial feature parameter determination model is used to determine the facial feature parameters of the digital human to be generated, and the facial feature parameters are used to render the facial feature mask onto the facial contour of the digital human to be generated.

[0252] Acquisition unit 702 is further configured as follows:

[0253] Obtain a template face image, wherein the template face image includes initial facial feature information and initial facial feature parameters corresponding to the initial facial feature information;

[0254] The initial facial features parameters are randomly adjusted to obtain sample facial features parameters, and the sample facial features information corresponding to the sample facial features parameters is obtained.

[0255] The sample facial features parameters and data are saved to obtain the sample data.

[0256] See Figure 8 , Figure 8 This specification illustrates an architecture diagram of an image processing system according to one embodiment of the present specification. The image processing system may include a client 100 and a server 200.

[0257] Client 100 is used to send the face image of the digital human to be generated to server 200;

[0258] Server 200 is used to obtain template face images, wherein the template face images include facial feature template information and facial contour template information;

[0259] Based on the facial feature template information, a facial feature mask is obtained from the face image of the digital human to be generated;

[0260] Based on the facial contour template information and the face image of the digital human to be generated, the facial contour of the digital human to be generated is determined;

[0261] The facial feature mask is rendered onto the facial contour of the digital human to be generated, generating the digital human's facial image; the digital human's facial image is then sent to the client 100.

[0262] Client 100 is also used to receive digital human face images sent by server 200.

[0263] An image processing system may include multiple clients 100 and a server 200. Clients 100 can be referred to as edge devices, and the server 200 can be referred to as cloud devices. Multiple clients 100 can establish communication connections through the server 200. In an image processing scenario, the server 200 is used to provide image processing services between the multiple clients 100. Each client 100 can act as a sender or receiver, communicating through the server 200.

[0264] Users can interact with server 200 through client 100 to receive data sent by other clients 100, or send data to other clients 100, etc. In image processing scenarios, users can publish data streams to server 200 through client 100, server 200 can generate a digital human's face image based on the data stream, and push the generated digital human's face image to other clients that have established communication.

[0265] In this system, client 100 and server 200 establish a connection via a network. The network provides the medium for communication between client 100 and server 200. The network can include various connection types, such as wired or wireless communication links or fiber optic cables. Data transmitted by client 100 may need to undergo encoding, transcoding, compression, or other processing before being published to server 200.

[0266] Client 100 can be a browser, an app (application), a web application such as an H5 (HyperText Markup Language 5) application, a lightweight application (also known as a mini-program), or a cloud application. Client 100 can be developed based on the software development kit (SDK) of the corresponding service provided by server 200, such as a real-time communication (RTC) SDK. Client 100 can be deployed on a computing device and depends on the device or certain apps on the device to run. The computing device may have a display screen and support information browsing, such as a personal mobile terminal like a mobile phone, tablet, or personal computer. Various other types of applications can also be configured on the computing device, such as human-computer interaction applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, and social media platform software.

[0267] Server 200 may include servers providing various services, such as servers providing communication services to multiple clients, servers supporting backend training of models used on clients, and servers processing data sent by clients. It should be noted that server 200 can be implemented as a distributed server cluster composed of multiple servers, or as a single server. The server can also be a server in a distributed system, or a server integrated with blockchain. The server can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.

[0268] It is worth noting that the image processing methods provided in the embodiments of this specification are generally executed by the server. However, in other embodiments of this specification, the client may also have similar functions to the server, thereby executing the image processing methods provided in the embodiments of this specification. In other embodiments, the image processing methods provided in the embodiments of this specification may also be executed jointly by the client and the server.

[0269] Figure 9 A structural block diagram of a computing device 900 according to an embodiment of this application is shown. The components of the computing device 900 include, but are not limited to, a memory 910 and a processor 920. The processor 920 is connected to the memory 910 via a bus 930, and a database 950 is used to store data.

[0270] The computing device 900 also includes an access device 940, which enables the computing device 900 to communicate via one or more networks 960. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 940 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Wi-MAX (Worldwide Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.

[0271] In one embodiment of this application, the aforementioned components of the computing device 900 and Figure 9 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 9 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this application. Those skilled in the art can add or replace other components as needed.

[0272] The computing device 900 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 900 can also be a mobile or stationary server.

[0273] The processor 920 is used to execute the following computer program / instructions, which, when executed by the processor, implement the steps of the above-described image processing method.

[0274] The above is an illustrative scheme of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the image processing method described above belong to the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the image processing method described above.

[0275] An embodiment of this specification also provides a computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the image processing method described above.

[0276] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on its differences from other embodiments. In particular, the computer-readable storage medium embodiments are basically similar to the image processing method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the image processing method embodiments.

[0277] An embodiment of this specification also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described image processing method.

[0278] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the image processing method described above belong to the same concept. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solution of the image processing method described above.

[0279] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0280] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or certain intermediate forms. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.

[0281] It should be noted that the above description describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in a different order than that shown in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous. Secondly, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of this specification.

[0282] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0283] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the embodiments described herein. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the embodiments, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.

Claims

1. An image processing method, comprising: obtaining a face image of a real person to be generated into a digital person, and a template face image, wherein the template face image comprises template information of facial features and template information of a face contour; obtaining a face feature mask in the face image of the real person to be generated into a digital person based on the template information of facial features; determining a face contour of the real person to be generated into a digital person based on the template information of the face contour and the face image of the real person to be generated into a digital person, wherein the determining comprises: determining a face feature part to be processed, wherein the face feature part to be processed is any part of facial features; obtaining first key point information of the face feature part to be processed in the face image of the real person to be generated into a digital person, and obtaining second key point information of the face feature part to be processed in the template face image; determining a transformation relationship between the first key point information and the second key point information; adjusting the face image of the real person to be generated into a digital person according to the transformation relationship, so that the face feature part to be processed in the face image of the real person to be generated into a digital person is aligned with the face feature part to be processed in the template face image; and obtaining a face feature mask based on the template information of facial features in the face image of the real person to be generated into a digital person after alignment of facial features; rendering the face feature mask to the face contour of the real person to be generated into a digital person, and generating a face image of a digital person, wherein the digital person is a virtual character generated based on a real face image of the real person to be generated into a digital person.

2. The method of claim 1, wherein the rendering the face feature mask to the face contour of the real person to be generated into a digital person comprises: determining facial feature parameters of the real person to be generated into a digital person, wherein the facial feature parameters of the real person to be generated into a digital person are used to determine a position of the face feature mask in the face contour of the real person to be generated into a digital person; masking facial features in the face contour of the real person to be generated into a digital person based on the position of the face feature mask in the face contour of the real person to be generated into a digital person, and obtaining a face image to be combined; rendering the face feature mask to the face image to be combined.

3. The method of claim 2, wherein the determining the facial feature parameters of the real person to be generated into a digital person comprises: determining facial feature information of the real person to be generated into a digital person in the face contour of the real person to be generated into a digital person; and inputting the facial feature information of the real person to be generated into a digital person into a facial feature parameter determination model, and obtaining the facial feature parameters of the real person to be generated into a digital person output by the facial feature parameter determination model.

4. The method of claim 3, wherein the facial feature parameter determination model is obtained by training through the following steps: acquiring sample data, wherein, the sample data comprises sample facial feature information and sample facial feature parameters; inputting the sample facial feature information into a facial feature parameter determination model, and obtaining predicted facial feature parameters output by the facial feature parameter determination model; calculating a model loss value based on the predicted facial feature parameters and the sample facial feature parameters; and Adjusting model parameters of the five-wink parameter determination model according to the model loss value, and continuing to train the five-wink parameter determination model until a model training stop condition is reached.

5. The method of claim 4, wherein the sample data is obtained in the following manner: acquiring a template face image, wherein, The template face image includes initial five-wink information and initial five-wink parameters corresponding to the initial five-wink information; The initial five-wink parameters are randomly adjusted to obtain sample five-wink parameters, and sample five-wink information corresponding to the sample five-wink parameters is obtained; The sample five-wink parameters and the sample five-wink information are saved to obtain the sample data.

6. The method of claim 2, wherein the five-wink in the face contour of the to-be-generated digital person is shielded based on the position of the face five-wink mask in the face contour of the to-be-generated digital person, comprising: determining a to-be-processed five-wink part, and determining the position of the to-be-processed five-wink part based on the position of the face five-wink mask in the face contour of the to-be-generated digital person, wherein the to-be-processed five-wink part is any part of the five-wink; determining reference shielding information corresponding to the to-be-processed five-wink part; shielding the to-be-processed five-wink part according to the reference shielding information based on the position of the to-be-processed five-wink part.

7. The method of claim 2, wherein the face five-wink mask is rendered to the to-be-combined face image, comprising: determining a to-be-rendered five-wink part in the to-be-combined face image, wherein the to-be-rendered five-wink part is any part of the five-wink; determining a to-be-processed five-wink part mask corresponding to the to-be-rendered five-wink part in the face five-wink mask based on the to-be-rendered five-wink part; rendering the to-be-processed five-wink part mask to the to-be-rendered five-wink part in the to-be-combined face image.

8. The method of any one of claims 1 to 7, wherein the face image of the to-be-generated digital person further comprises hair of the to-be-generated digital person; Before generating the face image of the digital person, comprising: performing hair segmentation on the face image of the to-be-generated digital person to obtain a hair mask; rendering the hair mask to the face image to generate the face image of the digital person containing hair.

9. The method of any one of claims 1 to 7, wherein obtaining the face image of the to-be-generated digital person comprises: obtaining an initial face image and determining left eye position information and right eye position information in the five-wink of the initial face image; adjusting the initial face image so that the left eye position information and the right eye position information are at the same horizontal position, and taking the adjusted initial face image as the face image of the to-be-generated digital person.

10. A model training method, comprising: obtaining sample data, wherein the sample data comprises sample five-wink information and sample five-wink parameters; inputting the sample five-wink information into a five-wink parameter determination model to obtain predicted five-wink parameters output by the five-wink parameter determination model; calculating a model loss value according to the predicted five-wink parameters and the sample five-wink parameters; According to the model loss value, the model parameters of the five-feature parameter determination model are adjusted, and the five-feature parameter determination model is continuously trained until a model training stop condition is reached, wherein the five-feature parameter determination model is used to determine the five-feature parameters of a digital person to be generated, the five-feature parameters are used to render a face feature mask to the face contour of the digital person to be generated, and a face image of the digital person is generated based on the face feature mask rendered to the face contour of the digital person to be generated, wherein the digital person is a virtual character generated based on a real face image of the digital person to be generated; Wherein, the face feature mask is obtained in the following way: Determine the face feature part to be processed, wherein the face feature part to be processed is any part of the face feature; obtain the first key point information of the face feature part to be processed in the face image of the digital person to be generated, obtain the second key point information of the face feature part to be processed in the template face image; determine the transformation relationship between the first key point information and the second key point information; adjust the face image of the digital person to be generated according to the transformation relationship, so that the face feature part to be processed in the face image of the digital person to be generated is aligned with the face feature part to be processed in the template face image; based on the five-feature template information, the face feature is intercepted in the face feature aligned face image of the digital person to be generated to obtain the face feature mask.

11. The method of claim 10, wherein the sample data is obtained in the following way: acquiring a template face image, wherein The template face image includes initial five-feature information and initial five-feature parameters corresponding to the initial five-feature information; Randomly adjust the initial five-feature parameters to obtain sample five-feature parameters, and obtain sample five-feature information corresponding to the sample five-feature parameters; Save the sample five-feature parameters and the sample five-feature information to obtain the sample data.

12. An image processing method applied to a cloud-side device, comprising: receiving a face image of a digital person to be generated sent by an end-side device, and obtaining a template face image, wherein the template face image includes five-feature template information and face contour template information; based on the five-feature template information, obtaining a face feature mask in the face image of the digital person to be generated; based on the face contour template information and the face image of the digital person to be generated, determining the face contour of the digital person to be generated; The face contour of the digital human to be generated is determined based on the face contour template information and the face image of the digital human to be generated, including: determining a face feature part to be processed, wherein the face feature part to be processed is any part of the face feature; obtaining first key point information of the face feature part to be processed in the face image of the digital human to be generated, and obtaining second key point information of the face feature part to be processed in the template face image; determining a transformation relationship between the first key point information and the second key point information; adjusting the face image of the digital human to be generated according to the transformation relationship, so that the face feature part to be processed in the face image of the digital human to be generated is aligned with the face feature part to be processed in the template face image; based on the five-wink template information, face features are intercepted in the face image of the digital human after the face features are aligned, to obtain a face feature mask; and the face five-wink mask is rendered to the face contour of the digital human to be generated, to generate a face image of a digital human, wherein the digital human is a virtual character image generated based on a real face image of the digital human to be generated. The face image of the digital human is sent to the terminal device.

13. A task platform, comprising a request interface and a response unit; The request interface is configured to receive a face image of a digital human to be generated sent by a terminal device. The response unit is configured to obtain a template face image, wherein the template face image comprises five-wink template information and face contour template information; based on the five-wink template information, a face five-wink mask is obtained in the face image of the digital human to be generated; based on the face contour template information and the face image of the digital human to be generated, a face contour of the digital human to be generated is determined; and the face five-wink mask is rendered to the face contour of the digital human to be generated, to generate a face image of a digital human, wherein the digital human is a virtual character image generated based on a real face image of the digital human to be generated; wherein the face contour of the digital human to be generated is determined based on the face contour template information and the face image of the digital human to be generated, including: determining a face feature part to be processed, wherein the face feature part to be processed is any part of the face feature; obtaining first key point information of the face feature part to be processed in the face image of the digital human to be generated, and obtaining second key point information of the face feature part to be processed in the template face image; determining a transformation relationship between the first key point information and the second key point information; adjusting the face image of the digital human to be generated according to the transformation relationship, so that the face feature part to be processed in the face image of the digital human to be generated is aligned with the face feature part to be processed in the template face image; based on the five-wink template information, face features are intercepted in the face image of the digital human after the face features are aligned, to obtain a face feature mask.

14. An image processing apparatus, comprising: The acquisition unit is configured to acquire a face image of a person to be generated and a template face image, wherein the template face image comprises five facial feature template information and face contour template information; The processing unit is configured to acquire a face five feature mask in the face image of the person to be generated based on the five facial feature template information, determine a face contour of the person to be generated based on the face contour template information and the face image of the person to be generated, wherein the determination of the face contour of the person to be generated based on the face contour template information and the face image of the person to be generated comprises: determining a face feature part to be processed, wherein the face feature part to be processed is any part of a face feature; acquiring first key point information of the face feature part to be processed in the face image of the person to be generated, acquiring second key point information of the face feature part to be processed in the template face image, determining a transformation relationship between the first key point information and the second key point information, and adjusting the face image of the person to be generated according to the transformation relationship to align the face feature part to be processed in the face image of the person to be generated with the face feature part to be processed in the template face image; and based on the five facial feature template information, the face feature mask is obtained by intercepting a face feature in the face image of the person to be generated after the face feature alignment; and the generation unit is configured to render the face five feature mask to the face contour of the person to be generated to generate a face image of a digital person, wherein the digital person is a virtual character image generated based on a real face image of the person to be generated.

15. A computing device comprising: a memory and a processor; the memory is configured to store computer programs / instructions, and the processor is configured to execute the computer programs / instructions, and the computer programs / instructions, when executed by the processor, implement the steps of the method of any one of claims 1 to 11.

16. A computer readable storage medium storing computer programs / instructions, and the computer programs / instructions, when executed by a processor, implement the steps of the method of any one of claims 1 to 11.

17. A computer program product comprising computer programs / instructions, and the computer programs / instructions, when executed by a processor, implement the steps of the method of any one of claims 1 to 11.

Citation Information

Patent Citations

  • Image processing method and apparatus, electronic device and medium

    CN108564526A

  • Makeup trying processing method and device for face image, computer equipment and storage medium

    CN111369644A