Training method of image generation model and generation method and device of digital human image
By acquiring multiple target facial images and fusion and training with preset image generation models, the problem of insufficient facial details capture in digital human image generation is solved, and a higher quality and consistent digital human image generation is achieved.
Patent Information
- Application Number
- CN202510323752.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-17
AI Technical Summary
Current digital human image generation technology is difficult to accurately capture facial details, affecting the quality and authenticity of image generation.
By acquiring multiple target facial images, using a preset image generation model to fuse it with the target background image, generate the target digital human image, and train the model based on the difference in facial features to improve the generalization ability of the model and the consistency of the generated image.
It enhances the generalization ability of the image generation model during the training process, improves the facial consistency and style stability of the generated digital human images, thereby improving the user experience.
Smart Images

Figure CN120164071A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing technologies, and particularly to technical fields such as artificial intelligence, large models, and big data. Background Art
[0002] With the rapid progress of Artificial Intelligence Generated Content (AIGC) technology and the increasing demand for digital human image generation, open-source platforms have played a core role in promoting innovation in this field. However, in the field of digital human image generation, current technologies cannot accurately capture facial details, affecting the quality and authenticity of image generation. Summary of the Invention
[0003] The present disclosure provides a training method for an image generation model, a method and apparatus for generating digital human images.
[0004] According to one aspect of the present disclosure, there is provided a training method for an image generation model, including:
[0005] Obtaining N target facial images for a target face; where N is an integer greater than 1;
[0006] Inputting the N target facial images and at least one target background image into a preset image generation model to obtain a target digital human image after fusing the target face with each target background image;
[0007] Training the preset image generation model based on the degree of difference between the first facial features in the target digital human image and the second facial features of the target face in the target facial image to obtain a target image generation model.
[0008] According to another aspect of the present disclosure, there is provided a method for generating digital human images, including:
[0009] Obtaining a plurality of to-be-processed facial images for a preset face;
[0010] Inputting the plurality of to-be-processed facial images and at least one preset background image into the target image generation model to obtain a digital human image after fusing the preset face with each preset background image;
[0011] Wherein, the target image generation model is obtained by training the preset image generation model based on the degree of difference between the first facial features in the target digital human image and the second facial features of the target face in the target facial image; the target digital human image is obtained by inputting at least N target facial images into the preset image generation model; the N target facial images are obtained by expanding M initial facial images; N is an integer greater than 1; M is a natural number less than N.
[0012] According to another aspect of the present disclosure, there is provided a training device for an image generation model, including:
[0013] An image expansion unit, configured to obtain N target facial images for a target face; wherein, N is an integer greater than 1;
[0014] A first generation unit, configured to input the N target facial images and at least one target background image into a preset image generation model to obtain a target digital human image after fusing the target face with each target background image;
[0015] A training unit, configured to train the preset image generation model based on the degree of difference between the first facial feature in the target digital human image and the second facial feature of the target face in the target facial image, so as to obtain a target image generation model.
[0016] According to another aspect of the present disclosure, there is provided a generation device for a digital human image, including:
[0017] An image acquisition unit, configured to obtain a plurality of to-be-processed facial images for a preset face;
[0018] A second generation unit, configured to input the plurality of to-be-processed facial images and at least one preset background image into the target image generation model to obtain a digital human image after fusing the preset face with each preset background image;
[0019] Wherein, the target image generation model is obtained by training the preset image generation model based on the degree of difference between the first facial feature in the target digital human image and the second facial feature of the target face in the target facial image; the target digital human image is obtained by inputting at least the N target facial images into the preset image generation model; the N target facial images are obtained by expanding M initial facial images; N is an integer greater than 1; M is a natural number less than N.
[0020] According to another aspect of the present disclosure, there is provided an electronic device, including:
[0021] At least one processor; and
[0022] A memory communicatively connected to the at least one processor; wherein,
[0023] The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute any method in the embodiments of the present disclosure.
[0024] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method according to any one of the embodiments of the present disclosure.
[0025] According to another aspect of the present disclosure, there is provided a computer program product including a computer program, which implements the method according to any one of the embodiments of the present disclosure when executed by a processor.
[0026] The solution of the present disclosure can use multiple images with the same face to provide facial details under different conditions for a preset image generation model, and obtain the target digital human image output by the preset image generation model. Furthermore, by using the difference between the target digital human image output by the preset image generation model and the output image, the preset image generation model is trained. In this way, the generalization ability of the preset image generation model during the training process is enhanced. At the same time, the facial consistency and style stability of the generated target digital human image are effectively improved, thereby laying a foundation for improving the user experience.
[0027] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings
[0028] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0029] Figure 1 is a schematic flowchart of a method for training an image generation model according to an embodiment of the present application Figure 1 ;
[0030] Figure 2 is a schematic flowchart of a method for training an image generation model according to an embodiment of the present application Figure 2 ;
[0031] FIG. 3(a) is a schematic diagram of generating a target digital human image according to an embodiment of the present application Figure 1 ;
[0032] FIG. 3(b) is a schematic diagram of generating a target digital human image according to an embodiment of the present application Figure 2 ;
[0033] Figure 4 is the third schematic flowchart of a method for training an image generation model according to an embodiment of the present application;
[0034] Figure 5 is a schematic diagram of image expansion using an initial facial image according to an embodiment of the present application;
[0035] Figure 6 is a schematic flowchart of a method for generating a digital human image according to an embodiment of the present application;
[0036] Figure 7 is a schematic structure of a training device 700 for an image generation model according to an embodiment of the present disclosure Figure 1 ;
[0037] Figure 8 is a schematic structure of a training device 700 for an image generation model according to an embodiment of the present disclosure Figure 2 ;
[0038] Figure 9 is a schematic diagram of the structure of a generating device 900 for a digital human image according to an embodiment of the present disclosure;
[0039] Figure 10 shows a schematic block diagram of an example electronic device 1000 that can be used to implement the embodiments of the present disclosure. Detailed Embodiments
[0040] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to assist understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0041] The term "and / or" in this document is merely a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The term "at least one" in this document means any one of multiple or any combination of at least two of multiple. For example, including at least one of A, B, and C can represent including any one or more elements selected from the set composed of A, B, and C. The terms "first" and "second" in this document represent referring to multiple similar technical terms and distinguishing them, and do not mean limiting the order, or limiting to only two. For example, the first feature and the second feature refer to two categories / two features. The first feature can be one or more, and the second feature can also be one or more.
[0042] In addition, for better illustration of the present disclosure, numerous specific details are given in the following detailed embodiments. Those skilled in the art should understand that the present disclosure can still be implemented without some specific details. In some instances, methods, means, elements, and circuits well-known to those skilled in the art are not described in detail to highlight the gist of the present disclosure.
[0043] The related technologies of the embodiments of the present disclosure are described below. The following related technologies can be arbitrarily combined with the technical solutions of the embodiments of the present disclosure as optional solutions, and all of them fall within the protection scope of the embodiments of the present disclosure.
[0044] With the booming development of AIGC technology, the application of open-source platforms in the field of digital human image generation has become increasingly widespread, becoming an important force driving innovation in this field. Using open-source platforms, realistic and creative digital human images can be generated, and even complex face-swapping operations can be performed to meet the growing personalized needs.
[0045] Although AIGC technology has made significant progress in digital human image generation, it still faces some technical bottlenecks. For example, current face-swapping or consistency generation technologies mainly rely on a single input image. For example, a single input image is used for model training or image generation. This approach has the problem of insufficient sampling and cannot fully capture all the details and features of the face. Moreover, due to insufficient sampling, when attempting to generate digital human images or perform face-swapping, the system often has difficulty accurately simulating the facial details of a person in different environments. This results in inconsistencies in facial features, expressions, etc. in the generated images, thereby affecting the quality and authenticity of the final output images and severely reducing the user experience.
[0046] Based on this, the present disclosure provides a method for training an image generation model and a method for generating digital human images using the trained image generation model. The training method of the present disclosure can effectively improve the quality and quantity of target facial images based on the consistency fusion technology of multiple target facial images, combined with feature variation and dataset augmentation strategies, and further improve the facial consistency and style stability of target digital human images. Specifically, the training method of the present disclosure can combine feature variation and dataset augmentation strategies to improve the quality and quantity of target facial images, and moreover, this training method can also perform consistency fusion processing on multiple amplified target facial images, thereby effectively improving the facial consistency and style stability of the generated target digital human images.
[0047] Specifically, Figure 1 is a schematic flowchart of a method for training an image generation model according to an embodiment of the present application Figure 1 . This method is optionally applied to electronic devices, such as personal computers, servers, server clusters, and other electronic devices.
[0048] Furthermore, this method at least includes at least some of the following content. As Figure 1 shown, it includes:
[0049] Step S101: Obtain N target facial images of a target face.
[0050] Here, N is an integer greater than 1. For example, in one example, the N target facial images are N facial images of the same human face.
[0051] Step S102: Input the N target facial images and at least one target background image into a preset image generation model to obtain a target digital human image after fusing the target face with each target background image.
[0052] That is to say, in one example, multiple target facial images of the same target face can be input into the preset image generation model simultaneously with multiple target background images, and then a target digital human image after fusing the target face with each target background image can be obtained. It can be understood that in this example, the number of generated target digital human images is the same as the number of input target background images. In this way, it provides strong support for batch implementation of facial changes.
[0053] Here, in one example, the target background image may include but is not limited to indoor environments, natural scenery, promotional posters, etc. In practical applications, the target background image can be determined according to the specific requirements of digital human image generation, and the present disclosure solution does not make specific limitations on this.
[0054] Step S103: Train the preset image generation model based on the difference degree between the first facial feature in the target digital human image and the second facial feature of the target face in the target facial image to obtain a target image generation model.
[0055] In this way, the solution of the present disclosure can use multiple images with the same face (for example, images with different details of the same face) to provide facial details under different conditions for the preset image generation model, and obtain the target digital human image output by the preset image generation model. Then, the difference between the target digital human image output by the preset image generation model and the output image (that is, the target facial image) is used to train the preset image generation model. In this way, the generalization ability of the preset image generation model during the training process is enhanced, and at the same time, the facial consistency and style stability of the generated target digital human image are effectively improved, thus laying a foundation for improving the user experience.
[0056] Furthermore, in a specific example, the image features of different target facial images are different.
[0057] Furthermore, in one example, the image features may include but are not limited to at least one of the following: angle, light, facial details, etc.
[0058] For example, in one example, the angles at which different target facial images are located (such as front, side, or semi-side, etc.) are different.
[0059] Alternatively, in another example, the lighting environments of different target facial images are different. Or, in yet another example, the facial details (such as facial texture or expression, etc.) of different target facial images are different.
[0060] In this way, due to the different image features of different target facial images, a more abundant and diverse set of training samples can be constructed, thereby effectively improving the generalization ability of the target image generation model obtained after training. At the same time, it also provides data support for improving the facial consistency and style stability of the generated target digital human images.
[0061] Figure 2 is a schematic flowchart of a method for training an image generation model according to an embodiment of the present application Figure 2 . This method is optionally applied to an electronic device, such as a personal computer, a server, a server cluster, or other electronic devices. It can be understood that the relevant content of the above Figure 1 shown method can also be applied to this example, and the relevant associated content will not be elaborated in this example.
[0062] Furthermore, this method at least includes at least some of the following content. As Figure 2 shown, it includes:
[0063] Step S201: Obtain N target facial images for a target face.
[0064] Here, N is an integer greater than 1.
[0065] It should be noted that for relevant examples of target facial images, reference can be made to the above description, and details will not be elaborated here.
[0066] Step S202: Input the N target facial images and at least one target background image into the image generation network of a preset image generation model to extract facial features for the N target facial images, obtaining a facial feature set for the N target facial images, and extract background features for each target background image, obtaining background features for each target background image.
[0067] That is to say, in this example, the preset image generation model includes an image generation network. For example, in one example, the image generation network can specifically be a Stable Diffusion network.
[0068] Furthermore, the image generation network can be used to extract facial features from each input target facial image. Here, in one example, the facial features may include details such as facial features, facial contours, and facial textures.
[0069] Further, after using the image generation network to extract facial features from each target facial image, a set of facial features containing the facial features of each target facial image can be obtained, that is, the facial feature set.
[0070] Further, the image generation network can also extract background features from each input target background image. Here, in one example, the background features may include information such as color, texture, and shape that help describe the overall style and details of the background image.
[0071] Step S203: Input the facial feature set and the background features of each target background image into the consistency fusion network of the preset image generation model to perform facial consistency constraints on the facial feature set of N target facial images, and perform feature fusion with the background features of each target background image after the facial consistency constraints to obtain the target digital human image after fusing the target face with each target background image.
[0072] That is to say, in this example, the preset image generation model may also include a consistency fusion network. At this time, the consistency fusion network can be used to perform facial consistency constraints on the facial feature set containing the facial features of N target facial images, aiming to ensure that the finally generated target digital human image can accurately reflect these facial features. In other words, using the consistency fusion network can make the finally generated target digital human image tend to be consistent with the facial features of the target facial image (for example, highly consistent in structure, style, and details), thereby improving facial consistency and style stability.
[0073] Here, in one example, the consistency fusion network can be used to perform consistency constraints on the facial feature set containing the facial features of N target facial images. Further, the consistency fusion network can also be used to fuse the consistency constraint result of the facial feature set (that is, the facial features after consistency constraints) with the background features of each target background image, thereby obtaining the target digital human image after fusing the target face with each target background image.
[0074] Alternatively, in another example, the preset image generation model may also include an image fusion module. In this example, the consistency fusion network is used to output the consistency constraint result of the facial feature set. Further, using the image fusion module, the consistency constraint result of the facial feature set and the background features of each target background image are fused, thereby obtaining the target digital human image after fusing the target face with each target background image.
[0075] Here, it should be noted that according to actual inference requirements, the preset image generation model may also include other necessary modules, such as a decoder, etc. The image fusion module included in the above-described preset image generation model is only an exemplary display, and the present disclosure solution does not limit whether other modules are additionally included in the preset image generation model.
[0076] Here, it should be noted that the target background image may specifically include a facial region. For example, in one example, the target background image may specifically be an image related to a human face and the background where the face is located, such as a poster image, etc. At this time, in this scenario, during the process of feature fusion of the facial feature (such as the facial feature after consistency fusion) and the background feature, the facial feature after consistency fusion can be specifically fused into the facial region of the background feature, thus realizing the replacement or adjustment of the facial feature.
[0077] FIG. 3(a) is a schematic diagram of generating a target digital human image according to an embodiment of the present application Figure 1 In one example, as shown in FIG. 3(a), a set of target facial images (including N target facial images) and a target background image are input into a preset image generation model. Here, the preset image generation model may include an image generation network and a consistency fusion network. Further, the image generation network is used to extract features from the set of target facial images to obtain a set of facial features corresponding to the set of target facial images, and the image generation network is used to extract background features from the target background image to obtain background features.
[0078] Further, in one example, the background features may include facial position information to facilitate accurately locating the replacement position corresponding to the face replacement operation.
[0079] Further, the set of facial features is input into the consistency fusion network to perform facial consistency constraint on the set of facial features. Finally, after feature fusion of the consistency constraint result of the set of facial features and the background features of the target background image, a target digital human image is obtained.
[0080] FIG. 3(b) is a schematic diagram of generating a target digital human image according to an embodiment of the present application Figure 2, in one example, as shown in FIG. 3(b), the target facial image set (including N target facial images) and the target background image set (for example, including P target background images, which can be respectively denoted as: the first target background image, the second target background image,..., the Pth target background image; P is an integer greater than or equal to 2) are input into the preset image generation model. Here, the preset image generation model may include an image generation network and a consistency fusion network. Further, the image generation network is used to extract features from the target facial image set to obtain the facial feature set corresponding to the target facial image set, and the image generation network is used to extract background features from each target background image in the target background image set to obtain the background features of each target background image.
[0081] Further, in one example, the background features of each target background image may include facial position information to facilitate accurately locating the replacement position corresponding to the face replacement operation.
[0082] Further, the facial feature set is input into the consistency fusion network to perform facial consistency constraints on the facial feature set. Finally, after fusing the consistency constraint results of the facial feature set and the background features of each target background image, a target digital human image set is obtained. The target digital human image set includes P target digital human images after fusing the target face with each target background image, which can be respectively denoted as the first target digital human image, the second target digital human image,..., the Pth target digital human image.
[0083] It should be noted that for relevant examples of the method of generating a target digital human based on the consistency constraint results of the facial feature set and the background features, reference can be made to the above description, and details will not be elaborated here.
[0084] In addition, in some examples, in order to clarify the position, pose of the digital human in the image and the interaction details with the surrounding environment, etc., the preset prompt information may also be input into the preset image generation model together with the target facial image set and the target background image, so as to use the input prompt information to guide the preset image generation model to generate the target digital human image that meets the requirements.
[0085] Step S204: Train the preset image generation model based on the difference degree between the first facial features in the target digital human image and the second facial features of the target face in the target facial image to obtain the target image generation model.
[0086] That is to say, after generating the target digital human image by using the preset image generation model, the degree of difference between the first facial feature in the target digital human image and the second facial feature of the target face in each target facial image can be used as a reference factor for model training, so as to obtain the target image generation model when the degree of difference between the facial features in the target digital human image and the facial features of the target face in each target facial image meets the preset requirements or the training reaches the training termination condition.
[0087] It should be noted that the solution of the present disclosure can also adopt a multi-stage training strategy, so as to enhance the model's facial consistency restoration ability in different environments.
[0088] In this way, the solution of the present disclosure can extract the facial features of multiple target facial images by using the image generation network, and use the consistency fusion network to perform consistency constraints on the facial features corresponding to the multiple target facial images, so that the generated target digital human image is consistent with the input target facial image in terms of facial features, thus improving the quality of the generated target digital human image.
[0089] In addition, since the solution of the present disclosure introduces multiple target facial images for the same target face, it is convenient to describe the features of the target face from multiple dimensions through the multiple target facial images. For example, the features of the target face can be described from aspects such as facial angle, lighting conditions, and facial details. Therefore, the problem that the generated digital human image has a large deviation from the input facial image due to insufficient facial features can be solved, and the generated digital human image is significantly improved in terms of detail performance and overall visual effect.
[0090] At the same time, compared with the traditional AIGC generation technology, the solution of the present disclosure can enable the preset image generation model to learn richer and more diverse information based on the features of the target face, thereby enhancing the stability of the generation result.
[0091] In other words, the target image generation model obtained after training the solution of the present disclosure can not only improve the accuracy of digital human image generation, but also has strong applicability. For example, it can maintain a high-quality visual generation effect on different platforms, devices, and application scenarios.
[0092] Furthermore, in a specific example, the following method can be used to train the preset image generation model; specifically, training the preset image generation model based on the degree of difference between the first facial feature in the target digital human image and the second facial feature of the target face in the target facial image (such as step S204) can specifically include:
[0093] Step S204-1: Calculate the similarity between at least the first facial feature in the target digital human image and the second facial feature of the target face in each target facial image to obtain similarity information.
[0094] Here, facial comparison algorithms such as Learned Perceptual Image Patch Similarity (LPIPS), Fréchet Inception Distance (FID), etc. can be used to calculate the similarity between the first facial feature in the target digital human image and the second facial feature in each target facial image, so as to obtain the similarity information between the first facial feature in the target digital human image and the second facial features in each target facial image based on the similarity calculation results.
[0095] Step S204-2: Fine-tune the image generation parameters in the preset image generation model based on the similarity information.
[0096] For example, calculate the similarity values between the first facial feature in the target digital human image and the second facial feature in each target facial image, a total of N. Further, if a total of P target digital human images are generated, then a total of N×P similarity values are obtained. At this time, the N×P similarity values can be directly used to fine-tune the image generation parameters in the preset image generation model.
[0097] Here, it should be noted that the image generation parameters can be adjustable parameters that control the output results of the preset image generation model. For example, they can include the weights and bias terms of the neural network in the model, and can also include parameters that affect the consistency of the generated target digital human image, etc.
[0098] Further, fine-tune the image generation parameters based on the obtained similarity information above on the basis of the preset image generation model to improve the output quality of the preset image generation model.
[0099] In this way, by calculating the similarity of facial features between the target digital human image and each target facial image, the difference between the output image and the facial features in the output image can be accurately quantified. Thus, it provides an accurate direction for the subsequent fine-tuning of the preset image generation model. And fine-tuning the preset image generation model based on the similarity information can specifically improve the output quality of the model, and further improve the generation quality of the target digital human image, so that the generated target digital human image is consistent with the input target facial image in terms of facial features.
[0100] Specifically, Figure 4FIG. 3 is a schematic flowchart of a method for training an image generation model according to an embodiment of the present application. This method is optionally applied to an electronic device, such as a personal computer, a server, a server cluster, or other electronic devices. It can be understood that the relevant content of the method shown in FIG. -3 can also be applied to this example, and the relevant associated content will not be elaborated here. Figure 1 The relevant content of the method shown in FIG. -3 can also be applied to this example, and the relevant associated content will not be elaborated here.
[0101] Furthermore, this method at least includes at least part of the following content. As shown in FIG. Figure 4 It includes:
[0102] Step S401: Obtain M initial facial images of a target face.
[0103] Here, M is a natural number less than or equal to N, and N is an integer greater than 1.
[0104] In one example, the M initial facial images can be facial images of the same target person at different angles, lighting conditions, or facial details. At this time, in this example, N target facial images can be obtained by expanding based on the facial images at different angles, lighting conditions, or facial details. In this way, it is convenient for the model to describe the features of the target face from multiple dimensions through these multiple target facial images. For example, the features of the target face can be described from aspects such as facial angle, lighting condition, and facial details. In this way, it provides strong support for solving the problem that the generated digital human image has a large deviation from the input facial image due to insufficient facial features.
[0105] Or, in another example, the M initial facial images can also be M identical images. At this time, in this example, the feature random variation or amplification strategy provided by the present disclosure can be used for feature expansion. In this way, it also provides strong support for solving the problem that the generated digital human image has a large deviation from the input facial image due to insufficient facial features.
[0106] Step S402: Obtain at least one facial extended image of the target face based on at least one of the following methods (that is, at least one of three methods), so as to obtain N target facial images based on the facial extended image.
[0107] It should be noted that for relevant examples of target facial images, please refer to the above description, and will not be elaborated here.
[0108] Furthermore, Method 1: Perform local perturbation on the target face based on the facial key features of the initial facial image. In other words, in this method, feature expansion can be performed by local perturbation of the target face, such as random perturbation, to obtain a new facial image.
[0109] It should be noted that in this example, the new facial images obtained after expanding the features can be collectively referred to as facial expansion images. Further, the facial expansion images can be used as target facial images, so as to achieve the amplification of the training dataset.
[0110] Further, in a specific example, the following method can be used to perform local perturbation on the target face; specifically, the local perturbation of the target face based on the facial key features of the initial facial image can specifically include: based on the facial key features of the initial facial image, performing fine-tuning on the non-key points (such as non-facial features like eyes, nose, mouth, etc.) of the target face to perform local perturbation. In other words, this local perturbation (such as making random or specific small changes) can keep the key facial features unchanged and fine-tune other details, so as to improve data diversity.
[0111] Here, in an example, the non-key points of the initial facial image can include details such as facial contours, skin textures, and facial expressions. The present disclosure does not specifically limit this.
[0112] In this way, since the present disclosure fine-tunes the non-key points of the initial facial image, it is possible to enrich the details of the initial facial image without changing the overall facial structure, thereby effectively improving data diversity. At the same time, it also provides strong support for improving the generalization ability of model learning and thus enhancing the output quality of the model.
[0113] Method 2: Fine-tuning the perspective of the target face based on the facial key features of the initial facial image. In other words, in this method, the perspective of the target face can be fine-tuned, such as making random adjustments, etc., to obtain a new facial image, so as to simulate facial changes under different perspectives and effectively improve data diversity.
[0114] Further, in a specific instance, the following method can be used to fine-tune the perspective of the target face; specifically, the fine-tuning of the perspective of the target face based on the facial key features of the initial facial image can specifically include: based on the target face under a preset perspective, fine-tuning the facial key features of the initial facial image to fine-tune the perspective of the target face.
[0115] Here, in this example, the preset perspective can be understood as the perspective that the target face is expected to reach. For example, it can be a specific angle, such as the front view, side view, elevation angle, or depression angle, etc. The present disclosure does not limit this and can be set based on actual generation needs.
[0116] Here, in an example, in order to make the target face presented under the preset perspective, the methods of fine-tuning the facial key features can include moving, rotating, scaling, etc. of the facial key features.
[0117] In this way, the disclosed solution can generate facial images from different perspectives, increasing the diversity of facial images. Furthermore, it facilitates the preset image generation model to obtain facial features from target facial images from different perspectives. At the same time, it also provides strong support for improving the generalization ability of model learning and thus enhancing the output quality of the model.
[0118] Method 3: Transform the lighting environment of the target face based on the facial key features of the initial facial image. In other words, in this method 3, the lighting environment of the target face can be fine-tuned, for example, by making random adjustments, etc., to obtain new facial images. In this way, the facial changes under different lighting conditions are simulated, effectively enhancing the data diversity.
[0119] Furthermore, in a specific example, the following method can be used to transform the lighting environment of the target face; specifically, the transformation of the lighting environment of the target face based on the facial key features of the initial facial image as described above can specifically include: fine-tuning the facial key features of the initial facial image based on the preset light conditions to transform the lighting environment of the target face.
[0120] Here, in an example, the preset lighting conditions can be one or more specific lighting conditions that are set, such as parameters like the position, intensity, and color of the light source. In this way, it is convenient to simulate different natural light or artificial light source environments so that the target face can adapt to different light environments.
[0121] For example, in an example, the ways to fine-tune the facial key features can include changing the brightness, contrast, etc. of the facial area to simulate the influence of the preset lighting conditions on the facial features.
[0122] In this way, the disclosed solution can simulate facial images with different light and shadow effects, increasing the diversity of facial images. Furthermore, it facilitates the preset image generation model to obtain facial features from target facial images from different perspectives. At the same time, it also provides strong support for improving the generalization ability of model learning and thus enhancing the output quality of the model.
[0123] Figure 5 It is a schematic diagram of image expansion using an initial facial image according to an embodiment of the present application. As Figure 5 shown, the initial facial image set contains M initial facial images. Based on the expansion strategy (for example, including at least one of local perturbation, perspective fine-tuning, and lighting condition transformation), the M initial facial images are expanded to obtain a target facial image set (containing N target facial images).
[0124] It should be noted that during the process of expanding the M initial facial images, a specified expansion strategy (such as the above three methods) can be selected, or one, two, etc. can be randomly selected from the above three expansion strategies in a random manner, or directly use the above three expansion strategies, etc. The specific selection method of the expansion strategy in the present disclosure is not limited.
[0125] It should be noted that Figure 5 The number M of the initial facial images and the number N of the target facial images are only exemplary displays. In practical applications, M can be less than or equal to N, and the present disclosure does not limit this.
[0126] Step S403: Input the N target facial images and at least one target background image into a preset image generation model to obtain a target digital human image after fusing the target face with each target background image.
[0127] It should be noted that for relevant examples of generating the target digital human image, refer to the above description and will not be elaborated here.
[0128] Step S404: Train the preset image generation model based on the difference degree between the first facial features in the target digital human image and the second facial features of the target face in the target facial image to obtain a target image generation model.
[0129] It should be noted that for relevant examples of training the preset image generation model, refer to the above description and will not be elaborated here.
[0130] In this way, the present disclosure can use the initial facial images to generate the same number or more target facial images (N, where N is greater than or equal to M). Thus, the diversity and richness of the facial image dataset are increased, which helps the preset image generation model to comprehensively understand facial features, and further improves the generalization ability of model training, making the generated digital human image more realistic.
[0131] In addition, compared with the digital human image generation method based on the diffusion model or the generative adversarial network, the present disclosure can obtain target facial images from multiple perspectives, with different facial details or different lighting environments, and optimize the preset image generation model based on the target facial images, enhancing the robustness and adaptability of the model in practical applications.
[0132] Furthermore, in a specific example, since the N target facial images are determined based on the facial expansion images, and the facial expansion images are determined based on the facial key features of the initial facial images, in order to accurately obtain the facial key features of the initial facial images, the method further includes:
[0133] Preprocess the initial facial images;
[0134] Perform feature encoding on the preprocessed initial facial image;
[0135] Extract features of key points in the initial facial image after feature encoding to obtain the facial key features of the initial facial image.
[0136] Here, before using the initial facial image for subsequent processing, a series of preset operations or transformations, i.e., preprocessing, can be performed on it. These preprocessing steps can improve the image quality, enhance image features, reduce noise or interference, etc.
[0137] Specifically, in one example, the steps for preprocessing the initial facial image may include:
[0138] (1) Noise removal: Reduce the noise in the initial facial image through means such as filters to improve the image quality;
[0139] (2) Image normalization: Adjust the pixel value range of the image to make it conform to a specific distribution or range;
[0140] (3) Image alignment: Locate the facial region in the initial facial image and perform operations such as rotation or scaling to make the facial features in the same position and scale.
[0141] It should be noted that the above preprocessing steps are only exemplary demonstrations, and the present disclosure scheme does not specifically limit the process and method of image preprocessing.
[0142] Furthermore, based on the preprocessed initial facial image, a feature extraction model (such as a Contrastive Language-Image Pre-training (CLIP) model, a deep learning network for face recognition (such as a FaceNet), etc.) can be used to convert the facial features in the initial facial image into a numerical representation form to complete feature encoding. Further, extract features of the key points in the initial facial image after feature encoding. In this way, the facial key features of the initial facial image can be obtained.
[0143] In this way, through preprocessing the initial facial image, the present disclosure scheme can effectively improve the quality of the initial facial image, making subsequent feature encoding and feature extraction more accurate and reliable.
[0144] Moreover, through feature encoding, the high-dimensional initial facial image can be converted into low-dimensional feature data, thereby enhancing the expression ability of the features of the initial facial image.
[0145] In addition, based on the feature encoding of the initial facial image, further feature extraction is performed on the key points, enabling the accurate acquisition of the facial key features of the initial facial image, providing a data basis for obtaining N target facial images subsequently.
[0146] The present disclosure also provides a method for generating a digital human image. Specifically, Figure 6 FIG. is a schematic flowchart of a method for generating a digital human image according to an embodiment of the present application. This method is optionally applied to an electronic device, such as a personal computer, a server, a server cluster, or other electronic devices.
[0147] Furthermore, this method at least includes at least part of the following content. As Figure 6 shown, it includes:
[0148] Step S601: Obtain a plurality of facial images to be processed for a preset face.
[0149] It should be noted that the relevant examples of the facial images to be processed are similar to those of the target facial images. For details, refer to the above description of the target facial images, which will not be elaborated here.
[0150] Step S602: Input the plurality of facial images to be processed and at least one preset background image into a target image generation model to obtain a digital human image after fusing the preset face and each preset background image.
[0151] Here, the target image generation model is obtained by training a preset image generation model based on the degree of difference between the first facial features in the target digital human image and the second facial features of the target face in the target facial image. Among them, the target digital human image is obtained by inputting at least N target facial images into the preset image generation model.
[0152] Furthermore, the N target facial images are obtained by expanding M initial facial images. Wherein, N is an integer greater than 1; M is a natural number less than or equal to N.
[0153] Furthermore, in one example, the target image generation model is obtained by using the above-described training method.
[0154] In one example, a plurality of facial images to be processed for a preset face can be input into the target image generation model simultaneously with a plurality of preset background images, and then a digital human image after fusing the preset face and each preset background image can be obtained. It can be understood that in this example, the number of generated digital human images is the same as the number of input preset background images. In this way, batch face change can be achieved, improving the efficiency of batch face change of the digital human.
[0155] It should be noted that for relevant examples of the target digital human image, the preset image generation model, the target facial image, and the initial facial image, please refer to the above description, which will not be elaborated here.
[0156] In this way, the present disclosure solution can utilize the target image generation model and generate a digital human image based on multiple to-be-processed facial images for a preset face. The digital human image has facial consistency and style stability with the input to-be-processed facial images. Thus, the user experience is effectively improved.
[0157] In addition, since multiple target facial images are utilized during the training of the target image generation model, the target image generation model can accurately extract the facial features in these images, effectively alleviating the problem of low image generation quality caused by insufficient sampling. And it enables the target image generation model to maintain the style consistency of the generated digital human images, thereby effectively preventing the occurrence of style drift.
[0158] Furthermore, the target image generation model proposed in the present disclosure solution has high scalability and can be integrated into various AIGC generation frameworks such as StableDiffusion.
[0159] Based on the above advantages, the target image generation model of the present disclosure solution can be applied to various scenarios. For example, in the field of digital human production, the target image generation model of the present disclosure solution can improve the consistency of character images; in the field of virtual anchors, the target image generation model of the present disclosure solution can ensure the coherence of face replacement effects during live broadcasts; in the field of film and television production, the target image generation model of the present disclosure solution can efficiently generate high-quality character images, accelerating the post-production process; in the field of social media avatar generation, the target image generation model of the present disclosure solution can enable users to maintain the consistency of avatar styles on different platforms.
[0160] The present disclosure solution also provides a training device 700 for an image generation model, as Figure 7 shown, including:
[0161] An image expansion unit 701, configured to obtain N target facial images for a target face; where N is an integer greater than 1;
[0162] A first generation unit 702, configured to input the N target facial images and at least one target background image into a preset image generation model to obtain a target digital human image after fusing the target face with each target background image;
[0163] A training unit 703, configured to train the preset image generation model based on the difference degree between the first facial features in the target digital human image and the second facial features of the target face in the target facial image to obtain a target image generation model.
[0164] In a specific example of the present disclosure solution, the image expansion unit 701 is specifically configured to:
[0165] Obtain M initial facial images of the target face, where M is a natural number less than or equal to N;
[0166] Based on at least one of the following methods, obtain at least one facial expansion image for the target face, so as to obtain N target facial images based on the facial expansion image:
[0167] Perform local perturbation on the target face based on the facial key features of the initial facial image;
[0168] Fine-tune the perspective of the target face based on the facial key features of the initial facial image;
[0169] Transform the lighting environment of the target face based on the facial key features of the initial facial image.
[0170] In a specific example of the present disclosure solution, the image expansion unit 701 is specifically configured to:
[0171] Fine-tune the non-key points on the target face based on the facial key features of the initial facial image to perform local perturbation.
[0172] In a specific example of the present disclosure solution, the image expansion unit 701 is specifically configured to:
[0173] Fine-tune the facial key features of the initial facial image based on the target face under a preset perspective to fine-tune the perspective of the target face.
[0174] In a specific example of the present disclosure solution, the image expansion unit 701 is specifically configured to:
[0175] Fine-tune the facial key features of the initial facial image based on a preset light condition to transform the lighting environment of the target face.
[0176] In a specific example of the present disclosure solution, the image features of different target facial images are different.
[0177] In a specific example of the present disclosure solution, as Figure 8 shown, the training device 700 of the image generation model may further include a feature extraction unit 704, where the feature extraction unit 704 is used to:
[0178] Preprocess the initial facial image;
[0179] Perform feature encoding on the preprocessed initial facial image;
[0180] Extract features of key points in the initial facial image after feature encoding to obtain the facial key features of the initial facial image.
[0181] In a specific example of the present disclosure solution, the first generation unit 702 is specifically configured to:
[0182] Input N target facial images and at least one target background image into the image generation network of a preset image generation model to perform facial feature extraction on the N target facial images to obtain a facial feature set for the N target facial images, and perform background feature extraction on each target background image to obtain background features for each target background image;
[0183] Input the facial feature set and the background features of each target background image into the consistency fusion network of the preset image generation model to perform facial consistency constraint on the facial feature set of the N target facial images, and perform feature fusion with the background features of each target background image after the facial consistency constraint to obtain a target digital human image after fusing the target face with each target background image.
[0184] In a specific example of the present disclosure solution, the training unit 703 is specifically configured to:
[0185] At least calculate the similarity between the first facial feature in the target digital human image and the second facial feature of the target face in each target facial image to obtain similarity information;
[0186] Based on the similarity information, fine-tune the image generation parameters in the preset image generation model.
[0187] For the specific functions and example descriptions of the units of the device in the embodiments of the present disclosure, reference may be made to the relevant descriptions of the corresponding steps in the above method embodiments, which will not be elaborated here.
[0188] The present disclosure solution also provides a digital human image generation device 900, as Figure 9 shown, including:
[0189] An image acquisition unit 901 for acquiring a plurality of to-be-processed facial images for a preset face;
[0190] A second generation unit 902 for inputting the plurality of to-be-processed facial images and at least one preset background image into a target image generation model to obtain a digital human image after fusing the preset face with each preset background image;
[0191] Among them, the target image generation model is obtained by training a preset image generation model based on the degree of difference between the first facial feature in the target digital human image and the second facial feature of the target face in the target facial image; the target digital human image is obtained by inputting at least N target facial images into the preset image generation model; the N target facial images are obtained by expanding M initial facial images; N is an integer greater than 1; M is a natural number less than or equal to N.
[0192] For the specific functions and example descriptions of the units of the device according to the embodiments of the present disclosure, reference may be made to the relevant descriptions of the corresponding steps in the above method embodiments, which will not be elaborated here.
[0193] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0194] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0195] Figure 10 FIG. shows a schematic block diagram of an exemplary electronic device 1000 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital assistant, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0196] As Figure 10 shown, the device 1000 includes a computing unit 1001, which can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 1002 or the computer program loaded from the storage unit 1008 into the random access memory (RAM) 1003. In the RAM 1003, various programs and data required for the operation of the device 1000 can also be stored. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other through a bus 1004. The input / output (I / O) interface 1005 is also connected to the bus 1004.
[0197] Multiple components in device 1000 are connected to I / O interface 1005, including: an input unit 1006, such as a keyboard, a mouse, etc.; an output unit 1007, such as various types of displays, speakers, etc.; a storage unit 1008, such as a disk, an optical disc, etc.; and a communication unit 1009, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1009 allows device 1000 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0198] The computing unit 1001 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 executes the various methods and processes described above, such as the training method of an image generation model or the generation method of a digital human image. For example, in some embodiments, the training method of an image generation model or the generation method of a digital human image can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into the RAM 1003 and executed by the computing unit 1001, one or more steps of the training method of the image generation model or the generation method of the digital human image described above can be executed. Alternatively, in other embodiments, the computing unit 1001 can be configured to execute the training method of the image generation model or the generation method of the digital human image by any other suitable means (e.g., by means of firmware).
[0199] The various embodiments of the systems and techniques described above in this article can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs, which can be executed and / or interpreted on a programmable system including at least one programmable processor, where the programmable processor can be a special or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0200] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.
[0201] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0202] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0203] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0204] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, or a server of a distributed system, or a server incorporating blockchain.
[0205] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitations are imposed herein.
[0206] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A training method for an image generation model, comprising: Obtain N target facial images of the target face; wherein N is an integer greater than 1; Inputting N target facial images and at least one target background image into a preset image generation model to obtain a target digital human image after the target facial image is fused with each target background image; Based on the difference between the first facial feature in the target digital human image and the second facial feature of the target face in the target facial image, the preset image generation model is trained to obtain the target image generation model.
2. The method according to claim 1, wherein: The step of obtaining N target facial images of a target face includes: Obtain M initial facial images of the target face, where M is a natural number less than or equal to N; At least one face extension image for the target face is obtained based on at least one of the following methods, so as to obtain N target face images based on the face extension image: Based on the key facial features of the initial facial image, the target face is locally perturbed; Based on the key facial features of the initial facial image, the perspective of the target face is fine-tuned; Based on the key facial features of the initial facial image, the lighting environment of the target face is transformed.
3. The method according to claim 2, wherein: The method of locally perturbing the target face based on the key facial features of the initial facial image includes: Based on the facial key features of the initial face image, the non-key points in the target face are fine-tuned to perform local perturbations.
4. The method according to claim 2, wherein: The step of fine-tuning the viewing angle of the target face based on the key facial features of the initial facial image comprises: Based on the target face at a preset viewing angle, key facial features of the initial facial image are fine-tuned to fine-tune the viewing angle of the target face.
5. The method according to claim 2, wherein: The step of transforming the illumination environment of the target face based on the key facial features of the initial facial image includes: Based on preset lighting conditions, the key facial features of the initial facial image are fine-tuned to transform the lighting environment of the target face.
6. The method according to any one of claims 2 to 5, wherein: Different target facial images have different image features.
7. The method according to any one of claims 2 to 5, further comprising: Preprocessing the initial facial image; Perform feature encoding on the preprocessed initial facial image; Feature extraction is performed on key points in the feature-encoded initial facial image to obtain facial key features of the initial facial image.
8. The method according to any one of claims 2 to 5, wherein: The step of inputting N target facial images and at least one target background image into a preset image generation model to obtain a target digital human image after the target facial image is fused with each target background image comprises: Inputting N target facial images and at least one target background image into an image generation network of a preset image generation model to extract facial features from the N target facial images to obtain a facial feature set for the N target facial images, and extracting background features from each target background image to obtain background features for each target background image; The facial feature set and the background features of each target background image are input into the consistency fusion network of the preset image generation model to perform facial consistency constraints on the facial feature set of N target facial images, and then perform feature fusion with the background features of each target background image after the facial consistency constraints, so as to obtain the target digital human image after the target face is fused with each target background image.
9. The method according to claim 8, wherein: The step of training a preset image generation model based on the difference between the first facial feature in the target digital human image and the second facial feature of the target face in the target facial image includes: At least calculating the similarity between the first facial feature in the target digital human image and the second facial feature of the target face in each target facial image to obtain similarity information; Based on the similarity information, the image generation parameters in the preset image generation model are fine-tuned.
10. A method for generating a digital human image, comprising: Acquire a plurality of facial images to be processed for a preset face; Inputting a plurality of facial images to be processed and at least one preset background image into a target image generation model to obtain a digital human image after the preset facial image and each preset background image are fused; Among them, the target image generation model is obtained by training the preset image generation model based on the degree of difference between the first facial feature in the target digital human image and the second facial feature of the target face in the target facial image; the target digital human image is obtained by inputting at least N target facial images into the preset image generation model; the N target facial images are obtained by expanding the M initial facial images; N is an integer greater than 1; M is a natural number less than or equal to N.
11. A training device for an image generation model, comprising: An image expansion unit, used to obtain N target facial images for a target face; wherein N is an integer greater than 1; A first generating unit, used for inputting N target facial images and at least one target background image into a preset image generating model, to obtain a target digital human image after the target facial image is fused with each target background image; The training unit is used to train a preset image generation model based on the difference between a first facial feature in a target digital human image and a second facial feature of a target face in a target facial image to obtain a target image generation model.
12. A device for generating a digital human image, comprising: An image acquisition unit, used for acquiring a plurality of facial images to be processed for a preset face; A second generation unit is used to input a plurality of facial images to be processed and at least one preset background image into a target image generation model to obtain a digital human image after the preset facial image is fused with each preset background image; Among them, the target image generation model is obtained by training the preset image generation model based on the degree of difference between the first facial feature in the target digital human image and the second facial feature of the target face in the target facial image; the target digital human image is obtained by inputting at least N target facial images into the preset image generation model; the N target facial images are obtained by expanding the M initial facial images; N is an integer greater than 1; M is a natural number less than N.
13. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 10.
14. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-10.
15. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 10.