Training method of image generation model, generation method and device of digital portrait, electronic device, and program

By training an image generation model with multiple face images and background fusion, the method addresses the inconsistency issue in digital portrait generation, achieving improved facial consistency and style stability in generated images.

JP2025137686AActive Publication Date: 2025-09-19BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025122064
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-18
Filing Date
2025-07-22
Publication Date
2025-09-19
Estimated Expiration
2045-07-22

Smart Images

  • Figure 2025137686000001_ABST
    Figure 2025137686000001_ABST
Patent Text Reader

Abstract

To provide a training method of an image generation model, a generation method and device of a digital portrait, an electronic device, and a program especially related to technical fields, such as artificial intelligence, a large-scale model, and big data.SOLUTION: A concrete solution is to acquire N target face images for a target face, where N is an integer larger than 1, a target digital portrait in which a target face and each target background image are merged is obtained by inputting N target face images and at least one target background image to a preset image generation model, and a target image generation mode is obtained by performing training to a preset image generation model on the basis of a degree of difference between a first face characteristic in a target digital portrait and a second face characteristic of a target face in a target face image.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to the field of image processing technology, and more particularly to the field of artificial intelligence, large-scale models, big data, and the like. [Background technology]

[0002] With the rapid advancement of Artificial Intelligence Generated Content (AIGC) technology and the growing demand for digital portrait generation, open source platforms play a key role in driving innovation in this field. However, in the field of digital portrait generation, current technology is unable to accurately capture facial details, which affects the quality and authenticity of the generated images. Summary of the Invention [Problem to be solved by the invention]

[0003] The present disclosure provides a method for training an image generation model, a method for generating a digital human image, and an apparatus, electronic device, and program therefor. [Means for solving the problem]

[0004] In one aspect of the present disclosure, there is provided a method for training an image generation model, the method comprising: acquiring N target face images of a target face, where N is an integer greater than 1; Inputting N target face images and at least one target background image into a preset image generation model to obtain a target digital person image in which the target face and each target background image are fused; and training a preset image generation model based on the degree of difference between the first facial feature in the target digital person image and the second facial feature of the target face in the target face image to obtain a target image generation model.

[0005] In another aspect of the present disclosure, there is provided a method for generating a digital person image, the method comprising: Obtaining a plurality of target face images for processing for a preset face; inputting a plurality of processing target face images and at least one preset background image into a target image generation model, and obtaining a digital person image in which the preset face and each preset background image are fused; Here, the target image generation model is obtained after training a preset image generation model based on the degree of difference between a first facial feature in the target digital person image and a second facial feature of the target face in the target face image, the target digital person image is obtained after inputting at least N target face images into the preset image generation model, and the N target face images are obtained by extending M initial face images, where N is an integer greater than 1 and M is a natural number less than or equal to N.

[0006] In another aspect of the present disclosure, there is provided an apparatus for training an image generation model, the apparatus comprising: an image enhancement unit for obtaining N target face images of a target face, where N is an integer greater than 1; a first generation unit for inputting N target face images and at least one target background image into a preset image generation model to obtain a target digital person image in which the target face and each target background image are fused; and a training unit for training the preset image generation model based on the degree of difference between the first facial feature in the target digital person image and the second facial feature of the target face in the target face image to obtain a target image generation model.

[0007] In another aspect of the present disclosure, there is provided an apparatus for generating a digital person image, the apparatus comprising: an image acquisition unit for acquiring a plurality of target face images for the preset face; a second generation unit for inputting a plurality of target face images and at least one preset background image into a target image generation model to obtain a digital person image in which the preset face and each preset background image are fused; Here, the target image generation model is obtained after training a preset image generation model based on the degree of difference between a first facial feature in the target digital person image and a second facial feature of the target face in the target face image, the target digital person image is obtained after inputting at least N target face images into the preset image generation model, and the N target face images are obtained by extending M initial face images, where N is an integer greater than 1 and M is a natural number less than or equal to N.

[0008] In another aspect of the present disclosure, there is provided an electronic device, the device comprising: at least one processor; a memory communicatively coupled to the at least one processor; The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, cause the implementation of any one of the methods in the embodiments of the present disclosure.

[0009] Another aspect of the present disclosure provides a non-transitory computer-readable storage medium having stored thereon computer instructions for causing a computer to perform any one of the methods in the embodiments of the present disclosure.

[0010] Another aspect of the present disclosure provides a program that, when executed by a processor, performs any of the methods of the embodiments of the present disclosure.

[0011] The present disclosure uses multiple images with the same face to provide facial details under different conditions to a preset image generation model, so as to obtain a target digital person image output by the preset image generation model; and further utilizes the difference between the target digital person image output by the preset image generation model and the output image to perform model training for the preset image generation model, thereby enhancing the generalization ability in the training process of the preset image generation model and effectively improving the facial consistency and style stability of the generated target digital person image, laying the foundation for further improving the user experience.

[0012] It should be understood that the contents described herein are not intended to describe key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will be better understood through the following specification. [Brief explanation of the drawings]

[0013] The accompanying drawings are for better understanding of the solutions of the present disclosure and are not to be construed as limiting the present disclosure.

[0014] [Figure 1] 1 is a schematic flowchart of a method for training an image generation model according to an embodiment of the present disclosure; [Figure 2] 1 is a second schematic flowchart of a method for training an image generation model according to an embodiment of the present disclosure. [Figure 3A] 1 is a schematic diagram illustrating target digital person image generation according to an embodiment of the present disclosure; [Figure 3B] FIG. 2 is a second schematic diagram illustrating target digital person image generation according to an embodiment of the present disclosure. [Figure 4] 10 is a third schematic flowchart of a method for training an image generation model according to an embodiment of the present disclosure. [Figure 5] FIG. 1 is a schematic diagram illustrating image expansion using an initial face image according to an embodiment of the present disclosure. [Figure 6] 1 is a schematic flowchart of a method for generating a digital person image according to an embodiment of the present disclosure. [Figure 7] FIG. 1 is a schematic diagram illustrating a configuration of a training device 700 for an image generation model according to an embodiment of the present disclosure. [Figure 8] FIG. 7 is a second schematic diagram illustrating the configuration of a training device 700 for an image generation model according to an embodiment of the present disclosure. [Figure 9] 9 is a schematic diagram illustrating a configuration of a digital person image generating device 900 according to an embodiment of the present disclosure. [Figure 10] FIG. 10 is a block diagram of an electronic device 1000 for implementing an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0015]

[0023] Exemplary embodiments of the present disclosure will now be described with reference to the accompanying drawings. These drawings include various details of the embodiments of the present disclosure to facilitate understanding, and should be considered as illustrative only. Therefore, it should be understood that those skilled in the art can make various changes and modifications to the embodiments described herein without departing from the scope of the present disclosure. Similarly, descriptions of well-known features and structures are omitted in the following description for clarity and conciseness.

[0016] The term "and / or" in the present disclosure merely describes the relationship between related objects and indicates that three types of relationships may exist. For example, A and / or B refer to three situations: A exists alone, A and B exist simultaneously, and B exists alone. The term "at least one" in the present disclosure indicates any combination of at least two of any one or more of a plurality of elements. For example, at least one of A, B, and C indicates that any one element or multiple elements can be selected from the set consisting of A, B, and C. The terms "first" and "second" in the present disclosure are used to refer to and distinguish multiple similar terms and do not limit the order or limit the number to only two. For example, a first feature and a second feature mean the existence of two types / two features, and the first feature may be one or more, and the second feature may be one or more.

[0017] Furthermore, in order to better explain the present disclosure, numerous specific details are set forth in the following specific embodiments. It should be understood by those skilled in the art that the present disclosure can be similarly practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art are not described in detail in order to emphasize the gist of the present disclosure.

[0018] The following describes related technologies of the embodiments of the present disclosure. The following related technologies can be arbitrarily combined with the technical solutions of the embodiments of the present disclosure as optional solutions, and all of them fall within the scope of protection of the embodiments of the present disclosure.

[0019] With the rapid development of AIGC technology, open source platforms are finding wider application in the field of digital portrait generation, becoming an important driving force for innovation in this field. Using open source platforms, people can generate realistic and creative digital portraits and perform complex face swapping operations, meeting the increasing demand for personalization.

[0020] Despite significant progress in digital person image generation, AIGC technology still faces several technical bottlenecks. For example, current face swapping and consistency generation technologies primarily rely on a single input image, i.e., model training and image generation are performed using a single input image. This approach suffers from insufficient sampling and is unable to fully capture all facial details and features. Furthermore, due to insufficient sampling, when attempting digital person image generation or face swapping, the system often has difficulty accurately simulating the facial details of a person in different environments. This results in inconsistencies in the generated images in terms of facial features, facial expressions, and other aspects, which impacts the quality and authenticity of the final output image and significantly degrades the user experience.

[0021] Based on this, the present disclosure provides a method for training an image generation model and a method for generating a digital person image using the trained image generation model, wherein the training method of the present disclosure is based on a consistent fusion technique of multiple target facial images and combines a feature change and dataset amplification strategy to effectively improve the quality and quantity of the target facial images, and further improve the facial consistency and style stability of the target digital person images. Specifically, the training method of the present disclosure can combine a feature change and dataset amplification strategy to improve the quality and quantity of the target facial images, and can also perform a consistent fusion process on the amplified multiple target facial images, further effectively improving the facial consistency and style stability of the generated target digital person images.

[0022] 1 is a schematic flowchart of a method for training an image generation model according to an embodiment of the present disclosure. The method can be selectively applied to electronic devices such as personal computers, servers, and server clusters.

[0023] Furthermore, the method includes at least part of the following content: As shown in FIG. In step S101, N target face images of a target face are obtained.

[0024] Here, N is an integer greater than 1. For example, in one example, the N target face images are N person face images including the same person's face.

[0025] In step S102, N target face images and at least one target background image are input to a preset image generation model to obtain a target digital person image in which the target face and each target background image are fused.

[0026] That is, in one example, multiple target face images of the same target face can be input into a preset image generation model together with multiple target background images, and target digital person images can be obtained by fusing the target face with each target background image. In this example, it can be seen that the number of generated target digital person images is the same as the number of input target background images, thus providing strong support for realizing a unified face modification.

[0027] Here, in one example, the target background image includes, but is not limited to, an indoor environment, a natural landscape, a promotional poster, etc. In practical applications, the target background image can be determined based on the specific needs of digital portrait generation, and the present disclosure does not impose any specific limitations thereon.

[0028] In step S103, a preset image generation model is trained based on the degree of difference between the first facial feature in the target digital portrait image and the second facial feature of the target face in the target face image to obtain a target image generation model.

[0029] In this way, the present disclosure uses multiple images of the same face (e.g., images of the same face with different details) to provide facial details under different conditions to a preset image generation model, obtain a target digital person image output by the preset image generation model, and further utilize the difference between the target digital person image output by the preset image generation model and the output image (i.e., the target facial image) to perform model training on the preset image generation model, thereby enhancing the generalization ability in the training process of the preset image generation model and effectively improving the facial consistency and style stability of the generated target digital person image, and laying the foundation for further improving the user experience.

[0030] Furthermore, in one embodiment, the image features of different target face images are different. Further, in one example, the image features may include, but are not limited to, at least one of angle, lighting, facial details, and the like.

[0031] For example, in one example, different target face images are positioned at different angles (eg, front, side, or semi-side, etc.). Alternatively, in another example, different target face images may have different lighting environments, or different target face images may have different facial details (such as facial texture or facial expressions).

[0032] In this way, since the image features of different target face images are different, a richer and more diverse training sample can be constructed, which further effectively improves the generalization ability of the target image generation model obtained after training, and provides data support for improving the facial consistency and style stability of the generated target digital person images.

[0033] 2 is a second schematic flowchart of a method for training an image generation model according to an embodiment of the present disclosure. This method can be selectively applied to electronic devices such as personal computers, servers, and server clusters. It can be understood that the relevant content of the method shown in FIG. 1 above can also be applied to this example. The relevant content will not be repeated in this example.

[0034] Furthermore, the method includes at least part of the following content: As shown in FIG. In step S201, N target face images for a target face are obtained.

[0035] Here, N is an integer greater than 1. For related examples of target face images, please refer to the above description, and the description will not be repeated here.

[0036] In step S202, N target face images and at least one target background image are input into the image generation network of the preset image generation model, facial feature extraction is performed on the N target face images to obtain a facial feature set for the N target face images, and background feature extraction is performed on each target background image to obtain background features for each target background image.

[0037] That is, in this example, the preset image generation model includes an image generation network, which in one example may specifically be a stable diffusion network.

[0038] Furthermore, an image generation network can be used to extract facial features for each input target face image, where in one example, facial features can include details such as the five senses, facial contours, and facial texture.

[0039] Furthermore, after performing facial feature extraction on each target face image using the image generation network, a set including the facial features of each target face image, i.e., a facial feature set, can be obtained.

[0040] Additionally, an image generation network can be used to extract background features for each input target background image, where, in one example, background features can include information useful for representing the overall appearance and details of the background image, such as color, pattern, and shape.

[0041] In step S203, the facial feature set and the background features of each target background image are input into the consistency fusion network of the preset image generation model, and a facial consistency constraint is performed on the facial feature sets of the N target face images. After the facial consistency constraint, feature fusion is performed with the background features of each target background image, and a target digital person image is obtained in which the target face and each target background image are fused.

[0042] That is, in this example, the preset image generation model may also include a consistency fusion network. In this case, the consistency fusion network can be used to enforce a facial consistency constraint on a facial feature set including facial features of N target facial images, with the aim of ensuring that the finally generated target digital person image accurately reflects these facial features. In other words, the consistency fusion network can be used to ensure that the finally generated target digital person image and the facial features of the target facial image are consistent (e.g., highly consistent in structure, appearance, and details), thereby improving facial consistency and appearance stability.

[0043] Here, in one example, a consistency fusion network can be used to perform consistency constraint on a facial feature set including facial features of N target face images. Furthermore, the consistency fusion network can be used to fuse the consistency constraint result of the facial feature set (i.e., the facial features after consistency constraint) with the background features of each target background image, to obtain a target digital person image that fuses the target face and each target background image.

[0044] Alternatively, in another example, the preset image generation model may include an image fusion module. In this example, a consistency fusion network is used to output a consistency constraint result for the facial feature set. Furthermore, an image fusion module is used to fuse the consistency constraint result for the facial feature set with the background features of each target background image, and further obtain a target digital person image that is a fusion of the target face and each target background image.

[0045] It should be noted that, depending on actual inference needs, the preset image generation model may include other necessary modules, such as a decoder, etc. The image fusion module included in the above preset image generation model is merely an example, and the present disclosure does not limit whether other modules are further included in the preset image generation model.

[0046] It should be noted that the target background image may specifically include a facial region, for example, in one example, the target background image may be a related image including a human face and a background in which the human face exists, such as a poster image, etc. In this case, in this scenario, in the process of feature fusing the facial feature (e.g., the facial feature after consistency fusion) and the background feature, the facial feature after consistency fusion can be merged with the facial region of the background feature, thereby realizing the exchange or adjustment of the facial feature.

[0047] 3A is a schematic diagram illustrating target digital person image generation according to an embodiment of the present disclosure. In one example, as shown in FIG. 3A, a target face image set (including N target face images) and one target background image are input to a preset image generation model. Here, the preset image generation model may include an image generation network and a consistency fusion network. Furthermore, feature extraction is performed on the target face image set using the image generation network to obtain a face feature set corresponding to the target face image set, and background feature extraction is performed on the target background image using the image generation network to obtain background features.

[0048] Additionally, in one example, the background features may include face location information, facilitating accurate identification of swap locations corresponding to face swap operations.

[0049] Furthermore, the facial feature set is input to a consistency fusion network to perform a facial consistency constraint on the facial feature set. Finally, the consistency constraint result of the facial feature set and the background features of the target background image are fused to obtain the target digital person image.

[0050] 3B is a second schematic diagram illustrating target digital person image generation according to an embodiment of the present disclosure. In one example, as shown in FIG. 3B , a target face image set (including N target face images) and a target background image set (e.g., including P target background images, which can be respectively referred to as the first target background image, the second target background image, ..., the Pth target background image, where P is an integer greater than or equal to 2) are input into a preset image generation model. Here, the preset image generation model may include an image generation network and a consistency fusion network. Furthermore, feature extraction is performed on the target face image set using the image generation network to obtain a face feature set corresponding to the target face image set, and background feature extraction is performed on each target background image in the target background image set using the image generation network to obtain background features for each target background image.

[0051] Additionally, in one example, the background features of each target background image can include face position information, which can facilitate precise identification of the swap position corresponding to the face swap operation.

[0052] Furthermore, to perform a facial consistency constraint on the facial feature set, the facial feature set is input to a consistency fusion network. Finally, after feature fusion of the consistency constraint result of the facial feature set and the background features of each target background image, a target digital person image set is obtained. This target digital person image set includes target digital person images in which the target face and each target background image are fused, and there are a total of P target digital person images, which can be denoted as, for example, the first target digital person image, the second target digital person image, ..., the Pth target digital person image.

[0053] It should be noted that for a related example of a method for generating a target digital person based on the consistency constraint result of the facial feature set and the background features, reference can be made to the above description, which will not be repeated here.

[0054] In some examples, preset prompt information can also be input to the preset image generation model along with a target face image set and a target background image so that the preset image generation model uses the input prompt information to generate a target digital person image that meets the requirements, such as to reveal details of the digital person's position, posture, and interaction with the surrounding environment in the image.

[0055] In step S204, a preset image generation model is trained based on the degree of difference between the first facial feature in the target digital person image and the second facial feature of the target face in the target face image to obtain a target image generation model.

[0056] That is, after generating a target digital person image using a preset image generation model, the degree of difference between the first facial feature in this target digital person image and the second facial feature of the target face in each target face image is used as a reference element for model training, and the target image generation model can be obtained when the degree of difference between the facial feature in the target digital person image and the facial feature of the target face in each target face image meets the preset requirement or when the training reaches the training end condition.

[0057] It should be noted that the method of the present disclosure can also employ a multi-stage training strategy, for example, to improve the model's ability to recover face consistency under different environments.

[0058] In this way, the present disclosure uses an image generation network to extract facial features of multiple target face images, and uses a consistency fusion network to constrain the facial features corresponding to the multiple target face images with consistency so that the generated target digital person image matches the input target face image in terms of facial features, thereby improving the quality of the target digital person image generation.

[0059] In addition, the present disclosure introduces multiple target face images for the same target face, which is convenient for expressing the features of the target face from multiple dimensions through the multiple target face images. For example, the features of the target face can be expressed in terms of face angle, lighting conditions, facial details, etc. This solves the problem of large discrepancies between the generated digital portrait image and the input facial image due to a lack of facial features, and significantly improves the detailed expression and overall visual effect of the generated digital portrait image.

[0060] At the same time, compared with conventional AIGC generation technology, the present disclosure enables the preset image generation model to learn richer and more diverse information based on the features of the target face, further enhancing the stability of the generation results.

[0061] In other words, the target image generation model obtained after training in the present disclosure can not only improve the accuracy of digital person image generation, but also has high applicability, for example, being able to maintain high-quality visual generation effects under different platforms, equipment, and application scenarios.

[0062] Furthermore, in one embodiment, the preset image generation model can be trained as follows: Specifically, training the preset image generation model based on the degree of difference between the first facial feature in the target digital person image and the second facial feature of the target face in the target face image (for example, step S204) can specifically include:

[0063] In step S204-1, the similarity between at least the first facial feature in the target digital person image and the second facial feature of the target face in each target face image is calculated to obtain similarity information.

[0064] Here, for example, a facial contrast algorithm such as Learned Perceptual Image Patch Similarity (LIPS) or Frechet Inception Distance (FID) is used to calculate the similarity between the first facial feature in the target digital person image and the second facial feature in each target face image, and based on the similarity calculation results, similarity information between the first facial feature in the target digital person image and the second facial feature in each target face image can be obtained.

[0065] In step S204-2, fine adjustments are made to the image generation parameters in the preset image generation model based on the similarity information.

[0066] For example, a total of N similarity values ​​are calculated between the first facial feature in the target digital person image and the second facial feature in each target face image. Furthermore, when P target digital person images are generated, N×P similarity values ​​are obtained, and at this time, these N×P similarity values ​​can be directly used to fine-tune the image generation parameters in the preset image generation model.

[0067] The image generation parameters may be adjustable parameters that control the output results of a preset image generation model, and may include, for example, neural network weights and bias terms in the model, as well as parameters that affect the consistency of the generated target digital person image.

[0068] Furthermore, based on the similarity information obtained above, fine adjustments are made to the image generation parameters based on the preset image generation model in order to improve the output quality of the preset image generation model.

[0069] In this way, by calculating the similarity between the facial features of the target digital person image and each target face image, the difference between the output image and the facial features in the output image can be accurately quantified, thus providing an accurate direction for fine-tuning the subsequent preset image generation model. Then, by fine-tuning the preset image generation model based on the similarity information, the output quality of the accurate model can be improved, and therefore the generation quality of the target digital person image can be improved, and the generated target digital person image can be kept consistent in facial features with the input target face image.

[0070] Specifically, Figure 4 is a third schematic flowchart of a method for training an image generation model according to an embodiment of the present disclosure. This method can be selectively applied to electronic devices such as personal computers, servers, and server clusters. It can be understood that the relevant content of the methods shown in Figures 1 to 3 above can also be applied to this example. The relevant content will not be repeated in this example.

[0071] Furthermore, the method includes at least part of the following content: As shown in FIG. In step S401, M initial face images for a target face are obtained.

[0072] Here, M is a natural number equal to or less than N, and N is an integer equal to or greater than 1.

[0073] In one example, the M initial face images may be face images of the same target person at different angles, lighting conditions, or facial details, in which case N target face images can be obtained by extending them based on the face images at different angles, lighting conditions, or facial details, and thus the model can conveniently represent the features of the target face from multiple dimensions, such as from the aspects of face angle, lighting conditions, facial details, etc., thus providing effective support for solving the problem of large discrepancies between the generated digital person image and the input face image due to the lack of facial features.

[0074] Alternatively, in another example, the M initial facial images may be M identical images, in which case, in this example, the feature random variation or amplification strategy provided by the method of the present disclosure can be used to perform feature expansion, which can provide effective support for solving the problem of large discrepancies between the generated digital person image and the input facial image due to a lack of facial features.

[0075] In step S402, at least one face augmented image for a target face is obtained to obtain N target face images based on the face augmented image according to at least one of the following methods (ie, at least one of the three methods):

[0076] For related examples of target face images, please refer to the above description, and the description will not be repeated here.

[0077] Furthermore, in Method 1, local perturbations are performed on the target face based on the facial key features of the initial face image. In other words, in Method 1, a new face image can be obtained by performing feature expansion through local perturbations, such as random perturbations, on the target face.

[0078] In this example, the new face image obtained by the feature expansion can be generally referred to as an "augmented face image." Furthermore, this augmented face image can be used as a target face image, thereby realizing the amplification of the training data set.

[0079] Furthermore, in one embodiment, the following method can be used to perform local perturbation on the target face, specifically, performing local perturbation on the target face based on the facial key features of the above-mentioned initial facial image can specifically include performing fine adjustment and local perturbation on non-key points (such as non-five senses) on the target face based on the facial key features of the initial facial image. In other words, this local perturbation (for example, performing random or specific small changes) can keep the key facial features unchanged and fine-tune other details, thus increasing the diversity of data.

[0080] Here, in one example, the non-key points of the initial face image may include detailed information such as facial contours, skin texture, and facial movements, but the present disclosure does not specifically limit this.

[0081] In this way, the present disclosure can enrich the details of the initial facial image by fine-tuning the non-key points of the initial facial image without changing the overall facial structure, and can effectively improve the data diversity, while at the same time providing strong support for improving the generalization ability of model learning and further improving the model output quality.

[0082] In Method 2, the angle of view of the target face is fine-tuned based on the key facial features of the initial face image. In other words, Method 2 can fine-tune the angle of view of the target face, for example, by random adjustment, to obtain a new face image, thereby simulating the face changes at different angles of view and effectively improving the diversity of data.

[0083] Furthermore, in one specific example, the angle of view of the target face can be fine-tuned as follows: Specifically, making fine adjustments to the angle of view of the target face based on the face key features of the initial face image may include making fine adjustments to the face key features of the initial face image based on the target face under a preset angle of view, and fine-tuning the angle of view of the target face.

[0084] Here, in this example, the preset angle of view can be understood as the angle of view that the target face is desired to reach, and may be a specific angle such as front, side, elevation, depression, etc., and the present disclosure does not impose any limitations thereon and can be set based on actual production needs.

[0085] Here, in one example, the method of making fine adjustments to the facial key features so that the target face appears at the preset view angle can include moving, rotating, scaling, etc. the facial key features.

[0086] In this way, the present disclosure can generate facial images with different angles of view, increasing the diversity of facial images and thus facilitating preset image generation models to acquire facial features from target facial images with different angles of view, while at the same time providing strong support for enhancing the generalization ability of model learning and thus improving the output quality of the model.

[0087] In Method 3, the lighting environment of the target face is transformed based on the key facial features of the initial face image. In other words, Method 3 can fine-tune the lighting environment of the target face, for example, by random adjustment, to obtain a new face image, thereby simulating the changes of the face under different lighting conditions and effectively improving the diversity of data.

[0088] Furthermore, in one embodiment, the lighting environment of the target face can be transformed as follows: Specifically, performing transformation on the lighting environment of the target face based on the face key features of the above-mentioned initial face image specifically includes fine-tuning the face key features of the initial face image based on preset light conditions, and transforming the lighting environment of the target face.

[0089] Here, in one example, the preset lighting conditions may be one or more specific lighting conditions that are set, such as parameters such as the position, intensity, and color of the light source, thereby simulating different natural light or artificial light source environments, which is convenient for the target face to adapt to different light environments.

[0090] For example, in one example, ways in which key facial features can be fine-tuned can include changing the brightness, contrast, etc. of facial regions to simulate the effect of preset lighting conditions on the facial features.

[0091] In this way, the present disclosure can simulate facial images with different light and shadow effects, increasing the diversity of facial images and further facilitating preset image generation models to acquire facial features from target facial images with different angles of view, while at the same time providing strong support for enhancing the generalization ability of model learning and thus improving the output quality of the model.

[0092] 5 is a schematic diagram illustrating image expansion using initial face images according to an embodiment of the present disclosure. As shown in FIG. 5, an initial face image set includes M initial face images, and a target face image set (including N target face images) is obtained by expanding the M initial face images based on an expansion strategy (e.g., including at least one of local perturbation, fine adjustment of angle of view, and lighting condition conversion).

[0093] In the process of performing expansion on M initial face images, a specified expansion strategy (for example, one of the above three methods) may be selected, or one, two, etc. may be randomly selected from the above three expansion strategies using a random method, or the above three expansion strategies may be used as is, and the present disclosure does not limit the specific method of selecting an expansion strategy.

[0094] It should be noted that the number M of initial face images and the number N of target face images in FIG. 5 are merely examples, and in actual applications, M may be equal to or less than N, and the present disclosure is not limited thereto.

[0095] In step S403, N target face images and at least one target background image are input to a preset image generation model to obtain a target digital person image in which the target face and each target background image are fused.

[0096] For a related example of generating a target digital person image, please refer to the above description, and the description will not be repeated here.

[0097] In step S404, a preset image generation model is trained based on the degree of difference between the first facial feature in the target digital person image and the second facial feature of the target face in the target face image to obtain a target image generation model.

[0098] For related examples of preset image generation model training, please refer to the above description, and will not be repeated here.

[0099] In this way, the present disclosure can use initial facial images to generate the same number or more target facial images (N, where N is greater than or equal to M), thus increasing the diversity and richness of the facial image dataset, helping the preset image generation model to comprehensively understand facial features, further improving the generalization ability of model training, and making the generated digital person images more realistic.

[0100] Furthermore, compared with digital human image generation methods based on diffusion models or adversarial networks, the present disclosure can acquire target face images with multiple viewing angles, different facial details, or different lighting environments, and optimize a preset image generation model based on the target face images, enhancing the robustness and adaptability of the model in actual applications.

[0101] Furthermore, in a specific example, the N target face images are determined based on the face expansion images, and the face expansion images are determined based on the face key features of the initial face image, so that in order to accurately obtain the face key features of the initial face image, the method includes: performing preprocessing on the initial face image; performing feature coding on the preprocessed initial face image; The method further includes performing feature extraction on the key points in the feature-encoded initial face image to obtain facial key features of the initial face image.

[0102] Here, the initial face image can be subjected to a series of preset operations and transformations, i.e., pre-processing, before being used for further processing. These pre-processing steps can improve image quality, enhance image features, reduce noise and interference, etc.

[0103] Specifically, in one example, the step of performing pre-processing on the initial face image may include:

[0104] (1) Noise reduction: Noise in the initial face image is reduced by means of a filter or the like, improving the quality of the image. (2) Image standardization: Adjusting the pixel value range of an image to fit a specific distribution or range. (3) Image alignment: The facial region is identified in the initial facial image, and operations such as rotation and scaling are performed so that the facial features are in the same position and scale.

[0105] It should be noted that the above pre-processing steps are merely exemplary, and the present disclosure does not impose any specific limitations on the flow and method of image pre-processing.

[0106] Furthermore, based on the preprocessed initial face image, a feature extraction model (e.g., a Contrastive Language-Image Pre-training (CLIP) model, a deep learning network for human face recognition (e.g., a face network FaceNet), etc.) can be used to convert the facial features in the initial face image into a digitized representation format to complete feature encoding. Furthermore, feature extraction can be performed on the key points in the feature-encoded initial face image to obtain the facial key features of the initial face image.

[0107] In this way, by performing preprocessing on the initial face image, the present disclosure can effectively improve the quality of the initial face image, making the subsequent feature encoding and feature extraction more accurate and reliable.

[0108] Furthermore, feature coding can convert a high-dimensional initial face image into low-dimensional feature data, thereby improving the expressive power of the features of the initial face image.

[0109] In addition, after feature coding is performed on the initial face image, feature extraction is further performed on the key points, so that the facial key features of the initial face image can be accurately obtained, providing a data basis for obtaining the subsequent N target face images.

[0110] The present disclosure further provides a method for generating a digital person image, and specifically, Figure 6 is a schematic flowchart of a method for generating a digital person image according to an embodiment of the present disclosure. The method can be selectively applied to electronic devices such as personal computers, servers, server clusters, etc.

[0111] Furthermore, the method includes at least part of the following content: As shown in FIG. In step S601, a plurality of face images to be processed for a preset face are obtained.

[0112] Note that the related examples regarding the processing target face image are similar to the related examples regarding the target face image, and therefore, please refer to the above description regarding the target face image and will not be described again here.

[0113] In step S602, a plurality of target face images and at least one preset background image are input to a target image generation model to obtain a digital person image in which the preset face and each preset background image are fused.

[0114] Here, the target image generation model is obtained by training a preset image generation model based on the degree of difference between a first facial feature in the target digital person image and a second facial feature of the target face in the target face image, where the target digital person image is obtained by inputting at least N target face images into the preset image generation model.

[0115] Furthermore, N target face images are obtained by performing extension on M initial face images, where N is an integer greater than 1, and M is a natural number equal to or less than N. Furthermore, in one example, the target image generation model is obtained using the training method described above.

[0116] In one example, multiple target face images for a preset face can be input into the target image generation model simultaneously with multiple preset background images, and a digital person image can be obtained by fusing the preset face with each preset background image. In this example, it can be understood that the number of generated digital person images is the same as the number of input preset background images, thereby realizing a collective face modification and improving the efficiency of the collective modification of the digital person face.

[0117] For examples of the relationship between the target digital person image, the preset image generation model, the target face image, and the initial face image, please refer to the above description, and the description will not be repeated here.

[0118] In this way, the present disclosure can utilize a target image generation model to generate a digital person image that has facial consistency and style stability with the input target face image based on multiple target face images for a preset face, thereby efficiently improving the user experience.

[0119] In addition, since multiple target face images are used in the process of training the target image generation model, the target image generation model can accurately extract facial features from these images, effectively alleviating the problem of reduced image generation quality due to insufficient sampling.In addition, the target image generation model can maintain the consistency of the generated digital person images, thereby effectively preventing the occurrence of style drift.

[0120] Furthermore, the target image generation model proposed in this disclosure is highly scalable and can be incorporated into multiple AIGC generation frameworks, such as stable diffusion.

[0121] Based on the above advantages, the target image generation model of the present disclosure can be applied to various scenarios. For example, in the field of digital character creation, the target image generation model of the present disclosure can improve the consistency of character images; in the field of virtual anchors, the target image generation model of the present disclosure can ensure the consistency of face swapping effects during live streaming; in the field of film and television production, the target image generation model of the present disclosure can efficiently generate high-quality character images and accelerate the subsequent production process; and in the field of social media avatar creation, the target image generation model of the present disclosure can maintain the consistency of user avatar styles across different platforms.

[0122] The present disclosure further provides an apparatus 700 for training an image generation model, as shown in FIG. an image expansion unit 701 for obtaining N target face images for a target face, where N is an integer greater than 1; a first generation unit 702 for inputting N target face images and at least one target background image into a preset image generation model to obtain a target digital person image in which the target face and each target background image are fused; and a training unit 703 for training the preset image generation model based on the degree of difference between the first facial feature in the target digital person image and the second facial feature of the target face in the target face image to obtain a target image generation model.

[0123] In one embodiment of the present disclosure, the image expansion unit 701 specifically includes: Obtaining M initial face images of a target face, where M is a natural number less than or equal to N; To obtain N target face images based on the face augmented images, performing local perturbations on the target face based on facial key features of the initial face image; fine-tuning the angle of view of the target face based on the face key features of the initial face image; Transforming the target face based on the facial key features of the initial face image to the lighting environment of the target face; The method is used to obtain at least one face augmentation image of the target face based on at least one of the methods.

[0124] In one embodiment of the present disclosure, the image expansion unit 701 specifically includes: Based on the facial key features of the initial face image, it is used to make fine adjustments and local perturbations to the non-key points in the target face.

[0125] In one embodiment of the present disclosure, the image expansion unit 701 specifically includes: Based on the target face at the preset angle of view, fine adjustment is made to the face key features of the initial face image, and the angle of view of the target face is fine-tuned.

[0126] In one embodiment of the present disclosure, the image expansion unit 701 specifically includes: Based on the preset lighting conditions, fine-tuning is performed on the face key characteristics of the initial face image, and the lighting environment of the target face is transformed.

[0127] In one embodiment of the present disclosure, the image features of different target face images are different.

[0128] In one embodiment of the present disclosure, as shown in FIG. 8, the image generation model training apparatus 700 further includes a feature extraction unit 704, where the feature extraction unit 704 is performing preprocessing on the initial face image; performing feature coding on the preprocessed initial face image; Feature extraction is performed on the key points in the feature-encoded initial face image to obtain the facial key features of the initial face image.

[0129] In one embodiment of the present disclosure, the first generating unit 702 specifically includes: Inputting N target face images and at least one target background image into an image generation network of a preset image generation model, performing facial feature extraction on the N target face images to obtain a facial feature set for the N target face images, and performing background feature extraction on each target background image to obtain background features for each target background image; The facial feature set and the background features of each target background image are input into the consistency fusion network of the preset image generation model, and a facial consistency constraint is applied to the facial feature sets of N target face images. After the facial consistency constraint, feature fusion is performed with the background features of each target background image, so as to obtain a target digital person image in which the target face and each target background image are fused.

[0130] In one embodiment of the present disclosure, the training unit 703 specifically includes: calculating a similarity between at least a first facial feature in the target digital person image and a second facial feature of the target face in each target face image to obtain similarity information; The similarity information is used to fine-tune the image generation parameters in the preset image generation model.

[0131] For a description of the specific functions and examples of each unit of the apparatus in the embodiments of the present disclosure, please refer to the relevant descriptions of the corresponding steps in the above method embodiments, and they will not be repeated here.

[0132] The present disclosure further provides a digital person image generating apparatus 900, as shown in FIG. an image acquisition unit 901 for acquiring a plurality of target face images for the preset face; a second generation unit 902 for inputting a plurality of target face images and at least one preset background image into a target image generation model to obtain a digital person image in which the preset face and each preset background image are fused; Here, the target image generation model is obtained after training a preset image generation model based on the degree of difference between a first facial feature in the target digital person image and a second facial feature of the target face in the target face image, the target digital person image is obtained after inputting at least N target face images into the preset image generation model, and the N target face images are obtained by extending M initial face images, where N is an integer greater than 1 and M is a natural number less than or equal to N.

[0133] For a description of the specific functions and examples of each unit of the apparatus in the embodiments of the present disclosure, please refer to the relevant descriptions of the corresponding steps in the above method embodiments, and they will not be repeated here. In the technical solution disclosed herein, the acquisition, storage, and application of users' personal information all comply with the provisions of relevant laws and regulations and do not violate public order and morals.

[0134] According to embodiments of the present disclosure, the present disclosure further provides an electronic device, a non-transitory computer-readable storage medium, and a program product.

[0135] 10 is a block diagram of an electronic device 1000 for implementing an embodiment of the present disclosure. The electronic device refers to various types of digital computers, including laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The electronic device also refers to various types of mobile devices, including personal digital assistants, cellular phones, intelligent phones, wearable devices, and other similar computing devices. The components, their connections, and functions described in this disclosure are merely exemplary and do not limit the implementation of what is described and specified in this disclosure.

[0136] 10, device 1000 includes a computing unit 1001 that can perform various appropriate operations and processes based on computer program instructions stored in a read-only memory (ROM) 1002 or loaded from a storage unit 1008 into a random access memory (RAM) 1003. The RAM 1003 can further store various programs and data required for the operation of device 1000. The computing unit 1001, ROM 1002, and RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.

[0137] Multiple components in device 1000 are connected to an I / O interface 1005, including an input unit 1006 such as a keyboard or mouse, an output unit 1007 such as various displays and speakers, a storage unit 1008 such as a magnetic disk or optical disk, and a communication unit 1009 such as a network card, modem, wireless communication transceiver, etc. The communication unit 1009 allows device 1000 to exchange information / data with other devices via a computer network such as the Internet and / or various carrier networks.

[0138] The computing unit 1001 may be various general-purpose and / or special-purpose processing components having processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, computing units that execute various machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 performs each of the methods and processes described above, such as the method for training an image generation model or generating a digital person image. For example, in some embodiments, the method for training an image generation model or the method for generating a digital person image may be implemented as a computer software program tangibly embodied in a machine-readable medium such as the storage unit 1008. In some embodiments, some or all of the computer program may be loaded and / or installed into the device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into the RAM 1003 and executed by the computing unit 1001, it may perform one or more steps of the method for training an image generation model or generating a digital person image described above. Additionally, in other embodiments, the computing unit 1001 may be configured to perform the training of the image generation model or the method for generating a digital person image in any other suitable manner (e.g., firmware).

[0139] Various embodiments of the systems or techniques described in this disclosure may be implemented using digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. Each of these embodiments may involve execution by one or more computer programs executed and / or interpreted by a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor, capable of receiving data and instructions from, and transferring data and instructions to, a storage system, at least one input device, and at least one output device.

[0140] Program code for carrying out the methods of the present disclosure can be written in any combination of one or more programming languages. Such program code can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programming data processing apparatus, such that when the program code is executed by the processor or controller, it can perform the functions / acts specified in the flowcharts and / or block diagrams. The program code can be executed entirely on-site, partially on-site, partially on-site and partially on a remote site as a separate soft encapsulation, or entirely on a remote site or server.

[0141] In the context of this disclosure, a machine-readable medium may be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. Further examples of machine-readable storage media include one or more hard-wired electrical connections, a portable computer disk cartridge, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any combination of the foregoing.

[0142] To provide for user interaction, the systems and techniques described herein can be implemented on a computer that includes a display device (e.g., a cathode ray tube (CRT) or liquid crystal display (LCD) monitor, etc.) for displaying information to the user, and a keyboard and pointing device (e.g., a mouse or trackball, etc.) for the user to provide input to the computer. Other types of devices can also be used to provide for user interaction; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, haptic feedback, etc.), and input from the user can be received in any form (e.g., acoustic input, voice input, tactile input, etc.).

[0143] The systems and techniques described herein can be implemented in a computing system that includes background components (e.g., as a data server), middleware components (e.g., an application server), front-end components (e.g., a user computer having a graphical user interface or network browser through which a user can interact with embodiments of the systems and techniques described herein), or any combination of such background, middleware, or front-end components. Components of the system can be connected to each other via any form or medium of digital data communication (e.g., a communications network). Examples of communications networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0144] The computer system may include a client and a server. Typically, the client and server are remote from each other and generally interact via a communication network. The client-server relationship is created by a computer program running on a corresponding computer. The server may be a cloud server, a server in a distributed system, or a server incorporating a blockchain.

[0145] It should be understood that steps can be newly ranked, added, or deleted using the various aspects of the flow shown above. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order. This disclosure is not limited thereto, as long as the technical solutions disclosed in this disclosure can achieve the desired results.

[0146] The above specific examples do not constitute limitations on the scope of protection of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions are possible depending on design considerations and other factors. Any modifications, equivalent replacements, improvements, etc. within the spirit and principles of the present disclosure should be included in the scope of protection of the present disclosure.

Claims

1. A method for training an image generation model, comprising: acquiring N target face images of a target face, where N is an integer greater than 1; inputting N target face images and at least one target background image into a preset image generation model to obtain a target digital person image in which the target face and each target background image are fused; training a preset image generation model based on a degree of difference between the first facial feature in the target digital person image and the second facial feature of the target face in the target face image to obtain a target image generation model; How to train image generation models.

2. The acquisition of N target face images of the target face includes: Obtaining M initial face images of a target face, where M is a natural number less than or equal to N; To obtain N target face images based on the face augmentation images, performing local perturbations on the target face based on facial key features of the initial face image; fine-tuning the angle of view of the target face based on the face key features of the initial face image; Transforming the target face based on the facial key features of the initial face image to the lighting environment of the target face; and obtaining at least one augmented face image for the target face based on at least one of the following methods: The method for training an image generation model according to claim 1 .

3. performing a local perturbation on a target face based on facial key features of the initial face image, performing fine-tuning and local perturbations on non-key points in the target face based on facial key features of the initial face image; The method for training an image generation model according to claim 2 .

4. The method of fine-tuning the angle of view of the target face based on the face key features of the initial face image includes: fine-tuning the face key features of the initial face image based on the target face at the preset angle of view, and fine-tuning the angle of view of the target face; The method for training an image generation model according to claim 2 .

5. Transforming the lighting environment of the target face based on the facial key features of the initial face image includes: and fine-tuning the face key characteristics of the initial face image based on the preset light conditions to transform the lighting environment of the target face. The method for training an image generation model according to claim 2 .

6. The method for training the image generation model includes: The image features of different target face images are different. The method for training an image generation model according to claim 2 .

7. The method for training the image generation model includes: performing preprocessing on the initial face image; performing feature coding on the preprocessed initial face image; performing feature extraction on key points in the feature-encoded initial face image to obtain facial key features of the initial face image; The method for training an image generation model according to claim 2 .

8. inputting the N target face images and at least one target background image into a preset image generation model, and obtaining a target digital person image in which the target face and each target background image are fused together; Inputting N target face images and at least one target background image into an image generation network of a preset image generation model, performing facial feature extraction on the N target face images to obtain a facial feature set for the N target face images, and performing background feature extraction on each target background image to obtain background features for each target background image; Inputting the facial feature set and the background features of each target background image into a consistency fusion network of a preset image generation model, performing a facial consistency constraint on the facial feature set of the N target face images, and performing feature fusion with the background features of each target background image after the facial consistency constraint, to obtain a target digital person image in which the target face and each target background image are fused; The method for training an image generation model according to claim 2 .

9. Training a preset image generation model based on a degree of difference between a first facial feature in the target digital person image and a second facial feature of a target face in the target face image includes: calculating a similarity between at least a first facial feature in the target digital person image and a second facial feature of the target face in each target face image to obtain similarity information; and fine-tuning image generation parameters in the preset image generation model based on the similarity information. The method for training an image generation model according to claim 8.

10. 1. A method for generating a digital person image, comprising: Obtaining a plurality of target face images for processing for a preset face; inputting a plurality of processing target face images and at least one preset background image into a target image generation model, and obtaining a digital person image in which the preset face and each preset background image are fused; wherein the target image generation model is obtained after training a preset image generation model based on the degree of difference between a first facial feature in the target digital person image and a second facial feature of the target face in the target face image, the target digital person image is obtained after inputting at least N target face images into the preset image generation model, and the N target face images are obtained by extending M initial face images, where N is an integer greater than 1 and M is a natural number equal to or less than N; A method for generating digital portraits.

11. An apparatus for training an image generation model, comprising: an image enhancement unit for obtaining N target face images of a target face, where N is an integer greater than 1; a first generation unit for inputting N target face images and at least one target background image into a preset image generation model to obtain a target digital person image in which the target face and each target background image are fused; a training unit for training the preset image generation model based on a degree of difference between the first facial feature in the target digital person image and the second facial feature of the target face in the target face image to obtain a target image generation model; A training device for image generation models.

12. A device for generating a digital portrait image, comprising: an image acquisition unit for acquiring a plurality of target face images for the preset face; a second generation unit for inputting a plurality of target face images and at least one preset background image into a target image generation model to obtain a digital person image in which the preset face and each preset background image are fused; wherein the target image generation model is obtained after training a preset image generation model based on the degree of difference between a first facial feature in the target digital person image and a second facial feature of the target face in the target face image, the target digital person image is obtained after inputting at least N target face images into the preset image generation model, and the N target face images are obtained by extending M initial face images, where N is an integer greater than 1 and M is a natural number equal to or less than N; A device for generating digital portrait images.

13. at least one processor; a memory communicatively coupled to the at least one processor; The memory stores instructions executable by the at least one processor, the instructions, when executed by the at least one processor, causing the at least one processor to perform the method of any one of claims 1 to 10. Electronic devices.

14. A non-transitory computer readable storage medium having stored thereon computer instructions that cause a computer to perform the method of any one of claims 1 to 10.

15. A program for implementing the method of any one of claims 1 to 10 when executed by a processor in a computer.

Citation Information

Patent Citations

  • Imaging device, image composition method, and program

    JP2012244226A

  • Image processing method and device, computer device, and computer program

    JP2024515907A