Model training method and device, electronic equipment and computer readable storage medium
By obtaining a sample set containing face images, edited face images and description text, the initial image editing model is trained, and the noise difference optimization model is used to solve the problem that the image editing model and the expected modified text in the prior art are not matched, achieving higher matching and user experience.
Patent Information
- Application Number
- CN202510192605.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-13
AI Technical Summary
When generating face images, the existing image editing model does not match the expected modified text, resulting in poor user experience.
By obtaining a sample set containing face images, edited face images and description text, the initial image editing model is trained, and the noise difference is used to optimize the model to obtain an image editing model that can better match the text description.
It improves the matching degree between the face images generated by the image editing model and the expected modified text, and improves the editing effect and user experience of the model.
Smart Images

Figure CN120147109A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly relates to a model training method, apparatus, electronic device, and computer-readable storage medium. Background Art
[0002] In a game scenario, in order to improve the user experience, it is possible to enable users to create game characters with expected facial features according to their own expectations.
[0003] Currently, generally, the expected modification text is input into an image editing model to modify the face image.
[0004] However, the face images generated by the current image editing model do not match well with the expected modification text, resulting in a poor user experience. Summary of the Invention
[0005] This application provides a model training method, apparatus, electronic device, and computer-readable storage medium, which can improve the matching degree between the face images generated by the image editing model and the expected modification text. The specific solutions are as follows:
[0006] In a first aspect, an embodiment of this application provides a model training method, and the method includes:
[0007] Obtain a first sample set; the first sample set includes a plurality of first samples, and each first sample includes a first face image, a second face image obtained by editing the first face image, and a first description text, where the first description text is used to describe the changes that occur from the first face image to the second face image;
[0008] Input the first face image and the first description text into an initial image editing model, and train the initial image editing model according to the difference between the noise added to the second face image and the noise predicted based on the first face image and the first description text, to obtain a first image editing model for modifying an image according to text;
[0009] Determine a plurality of positive samples and a plurality of negative samples from the plurality of first samples;
[0010] Input the first face image and the first description text into the first image editing model, and train the first image editing model according to the difference between the noise added to the second face image in the positive samples and the noise predicted based on the first face image and the first description text, and the difference between the noise added to the second face image in the negative samples and the noise predicted based on the first face image and the first description text, to obtain a second image editing model.
[0011] In a second aspect, an embodiment of this application provides a model training apparatus, and the apparatus includes: an obtaining unit, a training unit, and a determining unit.
[0012] An acquisition unit, configured to acquire a first sample set; the first sample set includes a plurality of first samples, and each first sample includes a first face image, a second face image obtained by editing the first face image, and a first description text, where the first description text is used to describe the changes that occur from the first face image to the second face image;
[0013] A training unit, configured to input the first face image and the first description text into an initial image editing model, and train the initial image editing model according to the difference between the noise added to the second face image and the noise predicted based on the first face image and the first description text, so as to obtain a first image editing model for modifying an image according to text;
[0014] A determination unit, configured to determine a plurality of positive samples and a plurality of negative samples from the plurality of first samples;
[0015] The training unit is further configured to input the first face image and the first description text into the first image editing model, and train the first image editing model according to the difference between the noise added to the second face image in the positive samples and the noise predicted based on the first face image and the first description text, and the difference between the noise added to the second face image in the negative samples and the noise predicted based on the first face image and the first description text, so as to obtain a second image editing model.
[0016] In a third aspect, the present application further provides an electronic device, including:
[0017] A processor; and
[0018] A memory, configured to store a data processing program. After the electronic device is powered on and runs the program through the processor, the method described in the first aspect is executed.
[0019] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, storing a data processing program, and when the program is run by a processor, the method described in the first aspect is executed.
[0020] The model training method provided by the embodiment of the present application includes the following steps:
[0021] Obtain a first sample set; the first sample set includes multiple first samples, and each first sample includes a first face image, a second face image obtained by editing the first face image, and a first description text, where the first description text is used to describe the changes from the first face image to the second face image; input the first face image and the first description text into an initial image editing model, and train the initial image editing model according to the difference between the noise added to the second face image and the noise predicted based on the first face image and the first description text, to obtain a first image editing model for modifying an image according to text; determine multiple positive samples and multiple negative samples from the multiple first samples; input the first face image and the first description text into the first image editing model, and train the first image editing model according to the difference between the noise added to the second face image in the positive samples and the noise predicted based on the first face image and the first description text, and the difference between the noise added to the second face image in the negative samples and the noise predicted based on the first face image and the first description text, to obtain a second image editing model. It can be seen that in this application, first, an initial image editing model constructed by a first sample set including a first face image, a second face image obtained by editing the first face image, and a first description text is trained to obtain a first image editing model, and multiple positive samples and multiple negative samples are determined from the multiple first samples according to preset conditions; then, the first image editing model is trained according to the difference between the noise added to the second face image in the positive and negative samples and the predicted noise output by the first image editing model based on the first face image and the first description text, to obtain a second image editing model. The trained second image editing model can edit a face image according to the input expected description text and face image, and output a new face image that conforms to the expected description text. At the same time, since the second image editing model is obtained by training the first image editing model with positive and negative samples, the result output by the second image editing model is closer to the positive samples, that is, the matching degree between the face image output by the second image editing model and the expected description text is higher, thereby improving the editing effect of the model and the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 It is a flowchart of the model training method provided by an embodiment of this application;
[0023] FIG. 2 is a schematic diagram of a first sample provided by an embodiment of this application;
[0024] Figure 3 It is a schematic diagram of the training process of the first image editing model provided by an embodiment of this application;
[0025] Figure 4 It is a schematic diagram of the inference process of the first image editing model provided by an embodiment of this application;
[0026] Figure 5 The structural block diagram of an example of the model training device provided by an embodiment of the present application;
[0027] Figure 6 The structural block diagram of an example of the electronic device for model training provided by an embodiment of the present application. Detailed implementation manners
[0028] In the following description, many specific details are set forth in order to provide a thorough understanding of the present application. However, the present application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the connotation of the present application. Therefore, the present application is not limited by the specific implementations disclosed below.
[0029] It should be noted that the terms "first", "second", "third", etc. in the claims, the description and the drawings of the present application are used to distinguish similar objects and are not used to describe a specific order or sequence. The data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than that shown or described herein. In addition, the terms "include", "have" and their variants are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these process, method, product or device.
[0030] It should be understood that in the embodiments of the present application, "at least one" means one or more, and "a plurality" means two or more. "And / or" is merely a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after. "Including A, B and / or C" means including any one or any two or all three of A, B, and C.
[0031] It should be understood that in the embodiments of the present application, "B corresponding to A", "B corresponding to A relatively", "A corresponding to B relatively" or "B corresponding to A relatively" means that B is associated with A, and B can be determined according to A. Determining B according to A does not mean determining B only according to A, but B can also be determined according to A and / or other information.
[0032] In a game scenario, to enhance the user experience, users should be able to create game characters with expected facial features according to their own expectations. Currently, generally, the modification of a face image is achieved by inputting the expected modification text into an image editing model. However, the matching degree between the face images generated by the current image editing model and the expected modification text is not high, resulting in a poor user experience.
[0033] It can be understood that creating customized game characters that can meet specific needs for players in a game scenario has become an important research focus. However, accurately editing the facial features of game characters created by users still faces many challenges. When players perform character editing, there are mainly two aspects of requirements: on the one hand, they hope to be able to precisely adjust specific features (such as the shape of facial features, makeup, etc.); on the other hand, they hope that abstract style descriptions (such as handsome, cute, cold, etc.) can conform to general aesthetic standards. Therefore, the importance of realizing game character face editing based on user input is self-evident.
[0034] The image editing methods in the prior art often focus on the editing of generalized scene images, and modify specific content images through text input. However, this method often has deficiencies in meeting the editing accuracy requirements of game character faces, specifically manifested in the lack of consistency between the edited content and the text description, and the inability to ensure the consistency of the content in the unedited area. In addition, such methods usually have difficulty ensuring the aesthetics of the results, thus affecting the adoption rate of users.
[0035] Exemplarily, the editing of generalized scene images refers to the global transformation or adjustment of the entire image without specifying a specific object or area. These operations are usually based on some predefined rules, templates, or simple algorithms, aiming to change the overall style, tone, brightness, and other attributes of the image. The image editing methods in the prior art focus on the editing of generalized scene images and rely on predefined rules or templates. These methods may not be able to accurately convert text descriptions into corresponding image modifications. For example, when the user inputs "increase the curvature of the eyebrows", the system may not be able to accurately understand and execute this instruction. Another example, assuming that the user hopes to change the eye color of the game character from blue to green, the image editing methods in the prior art may change the eye color, but may also affect the surrounding skin tone or other areas at the same time, resulting in an unnatural overall effect. When performing local editing, the image editing methods in the prior art are difficult to ensure that the content of the unedited area remains unchanged, which means that when modifying a specific part, unnecessary changes may occur in other parts. For example, when the user requests to adjust the hairstyle of the game character, the image editing methods in the prior art may accidentally change the facial contour or skin texture of the character, destroying the overall consistency of the character.
[0036] Due to the reasons of the background technology, in order to improve the matching degree between the face images generated by the image editing model and the expected modified text, the first embodiment of the present application provides a model training method, which is applied to an electronic device. The electronic device can be a desktop computer, a laptop computer, a mobile phone, a tablet computer, a smart watch, etc., or other electronic devices capable of model training. The embodiments of the present application do not specifically limit this.
[0037] The technical solution of the present application will be described in detail below through specific embodiments. It should be noted that these specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0038] Figure 1 It is a flowchart of the model training method provided by the embodiments of the present application. As Figure 1 shown, the method may include S101-S104.
[0039] S101. Obtain a first sample set; the first sample set includes a plurality of first samples, and each first sample includes a first face image, a second face image obtained by editing the first face image, and a first description text, where the first description text is used to describe the changes that occur from the first face image to the second face image.
[0040] Exemplarily, a large amount of face parameter data can be collected in advance, and each piece of face parameter data is rendered through a game engine to obtain a corresponding face image. The face parameter data is classified to form a facial attribute set (such as including eyes, mouth, eyebrows, etc.), and each facial attribute contains multiple dimensions (such as the upper eyelid position, lower eyelid position, eye size, etc. in the eye attribute, the mouth width, lip color, etc. in the mouth attribute, and the eyebrow thickness, eyebrow color, eyebrow angle, eyebrow length, etc. in the eyebrow attribute). After obtaining a large amount of face parameter data, such as 200,000 sets of face parameter data, two pieces of face parameter data can be selected from the face parameter data, such as face parameter data A and face parameter data B, and several target attributes are randomly selected from the facial attribute set, such as eyes and eyebrows; then the attribute values of the target attributes (i.e., eyes and eyebrows) in the previous piece of face parameter data (i.e., face parameter data A) are replaced with the attribute values of the target attributes (i.e., eyes and eyebrows) in the latter piece of face parameter data (i.e., face parameter data B) to obtain a new piece of face parameter data (which can be denoted as face parameter data C); the new piece of face parameter data (i.e., face parameter data C) is rendered through the game engine to obtain a corresponding face image; the face image corresponding to the previous piece of face parameter data (i.e., face parameter data A) is used as the first face image, the face image corresponding to the new piece of face parameter data (i.e., face parameter data C) is used as the second face image, and the text describing the changes from the first face image (i.e., face parameter data A) to the second face image (i.e., face parameter data C) is used as the first description text (which can be an artificial description); the above process can be repeated multiple times to obtain the first sample set.
[0041] Exemplarily, the first description text can be a description of the specific facial feature changes from the first face image to the second face image, or a description of the style changes from the first face image to the second face image, and there is no limitation thereto. For example, the first description text can include "the eyes become bigger", "the lipstick turns green", "the face becomes rounder and looks cuter". FIG. 2 is a schematic diagram of the first sample provided by an embodiment of the present application. As shown in FIG. 2, FIG. 2(A) represents the first face image, the first description text is "the eyes become bigger", and FIG. 2(B) represents the second face image.
[0042] S102. Input the first face image and the first description text into the initial image editing model, and train the initial image editing model according to the difference between the noise added to the second face image and the noise predicted based on the first face image and the first description text to obtain a first image editing model for modifying an image according to text.
[0043] Exemplarily, the output of the first image editing model may include predicted noise and a second face image. The initial image editing model may include a generative diffusion sub-model. The training process of the initial image editing model may include forward diffusion process training, reverse generation process training, and instruction embedding and feature fusion training. Forward diffusion process training can be understood as gradually adding noise to the first face image until it finally becomes a pure noise image. This process is usually designed as a series of discrete steps, where in each step, a part of the noise is added to the first face image, and a series of samples can be generated that gradually transition from the clear first face image to the pure noise image. This sample can be used for subsequent reverse generation process training to learn how to recover the first face image from the pure noise image. After the forward diffusion process training, the reverse generation process training begins. The goal of the reverse generation process is to learn a mapping that gradually denoises the pure noise image and restores it to the clear first face image. Contrary to the forward process, the reverse process starts from the pure noise image and gradually removes the noise to restore the first face image or generate a second face image. The goal of the training is to let the model learn how to reverse each step in the forward process, so as to be able to generate meaningful data from random noise. Instruction embedding and feature fusion training can, either during or after the reverse generation process training, convert the first description text and the first face image into feature vectors, and based on the feature vectors of the first face image and the first description text, output predicted noise and a second face image. The initial image editing model can be trained according to the difference between the noise added to the second face image and the predicted noise output by the initial image editing model based on the first face image and the first description text, to obtain the first image editing model for modifying the image according to the text.
[0044] Exemplarily, Figure 3 is a schematic diagram of the training process of the first image editing model provided by an embodiment of this application. As Figure 3As shown, the first face image, the first description text, and the second face image are input into the initial image editing model. The VAE encoder in the initial image editing model is used to perform transformation processing on the first face image and the second face image respectively, obtaining the feature vector of the first face image (i.e., the first feature image) and the feature vector of the second face image (i.e., the second feature image). The CLIP text encoder in the initial image editing model is used to transform the first description text, obtaining the text feature vector of the first description text. Gaussian noise is added to the second feature image through the image noise addition layer in the initial image editing model, obtaining the second feature image after noise addition. The second feature image after noise addition, the first feature image, and the text feature vector of the first description text are input into the noise estimation layer in the initial image editing model, and the noise estimation layer applies a U-net network. The second feature image after noise addition and the first feature image are subjected to splicing processing, obtaining the spliced feature image, and this spliced feature image is used as the input of the U-net network. The text feature vector of the first description text is used as the guiding condition to perform noise prediction on the second feature image after noise addition, outputting the predicted noise image. The first feature image can provide high-resolution detail information through skip connections, helping to restore the important features masked by noise in the second feature image, guiding the model to focus on specific regions or features, helping to retain details such as edges and textures, and improving the denoising effect. The text feature vector of the first description text can be embedded into a certain layer of the U-net network according to the requirements of the actual application scenario, thereby changing the feature values of this layer, affecting the feature extraction and reconstruction process at a deeper level, enabling the model to dynamically adjust the degree of attention to different regions according to the text feature vector, providing high-level semantic information, helping the model better understand the image content, guiding the denoising process, and enhancing the classification and segmentation capabilities.
[0045] Figure 4 This is a schematic diagram of the inference process of the first image editing model provided by the embodiment of the present application. As Figure 4As shown in the figure, the fourth face image, the first edited text, and the pure Gaussian noise image can be input into the first image editing model. The VAE encoder in the first image editing model is used to perform transformation processing on the fourth face image and the pure noise image respectively, obtaining the feature vector of the fourth face image (i.e., the fourth feature image) and the feature vector of the pure Gaussian noise image (i.e., the pure Gaussian noise feature image). The CLIP text encoder in the first image editing model is used to transform the first edited text, obtaining the text feature vector of the first edited text. The pure Gaussian noise feature image, the fourth feature image, and the text feature vector of the first edited text are input into the noise estimation layer in the first image editing model. The noise estimation layer is used to predict the noise of the pure Gaussian noise feature image, obtaining the predicted noise image. According to the predicted noise image, the noise removal layer in the first image editing model is used to perform denoising processing on the pure Gaussian noise feature image, obtaining the feature vector of the fifth face image. Then, the VAE encoder in the first image editing model is used to process the feature vector of the fifth face image, obtaining the fifth face image, where the fifth face image represents the image obtained by editing the fourth face image according to the first edited text.
[0046] S103. Determine a plurality of positive samples and a plurality of negative samples from a plurality of first samples.
[0047] Exemplarily, obtain the label of each first sample in the first sample set; determine a plurality of positive samples and a plurality of negative samples from the plurality of first samples according to the label; the positive sample is a sample whose value represented by the label of the second face image is greater than the first preset threshold, and the negative sample is a sample whose value corresponding to the label of the second face image is less than the second preset threshold. The label includes an aesthetic label and / or a semantic consistency label. The aesthetic label can represent the visual aesthetics of the edited image evaluated by manual scoring; the semantic consistency label can be used to evaluate whether the edited image meets the editing requirements described in the text.
[0048] Exemplarily, each first sample in the first sample set can be scored based on aesthetics and semantic consistency, and the score is set to 1-5 points. Based on the scores in these two aspects, a comprehensive evaluation can be performed on each sample in the first sample set. Through a preset threshold, the first sample set is divided into a positive sample set and a negative sample set. When the first preset threshold is 3 and the second preset threshold is 2, the first samples with scores greater than 3 points can be recorded as positive samples, and the samples with scores less than 2 points can be recorded as negative samples. The positive sample set contains samples that meet the user's preferences, while the negative sample set contains samples that do not meet the user's preferences.
[0049] S104. Input the first face image and the first description text into the first image editing model, and train the first image editing model according to the difference between the noise added to the second face image in the positive sample and the noise predicted based on the first face image and the first description text, and the difference between the noise added to the second face image in the negative sample and the noise predicted based on the first face image and the first description text, to obtain the second image editing model.
[0050] Exemplarily, one sample can be randomly selected from the positive sample and the negative sample respectively, to determine the difference between the noise added to the second face image in the positive sample by the first image editing model and the noise predicted by the first image editing model based on the first face image and the first description text, and the difference between the noise added to the second face image in the negative sample by the first image editing model and the noise predicted by the first image editing model based on the first face image and the first description text; then, according to the difference between the noise added to the second face image in the positive sample by the first image editing model and the noise predicted by the first image editing model based on the first face image and the first description text, and the difference between the noise added to the second face image in the negative sample by the first image editing model and the noise predicted by the first image editing model based on the first face image and the first description text, use the Direct Preference Optimization (DPO) algorithm to train the first image editing model to obtain the second image editing model. The second image editing model can output a second face image according to the input first face image and the expected modification text for the first face image, and the second face image is the face image obtained by modifying the first face image according to the expected modification text. For example, input the face image X and the expected modification text "enlarge the eyes and change the lipstick color to red" into the second image editing model, and the second image editing model can output the face image Y, which is the face image obtained by enlarging the eyes in the face image X and changing the lipstick color to red. Among them, DPO is a method of directly optimizing user preferences, which can directly optimize the parameters of the model by collecting the preference data of users for the model output, so that the model output is more in line with the user's preferences. By introducing the positive and negative sample comparison into the loss function, by comparing the difference between the predicted noise and the actually added noise in the positive and negative samples, the second image editing model can produce results that are closer to the positive sample and farther from the negative sample.
[0051] In an optional embodiment, before inputting the first face image and the first description text into the initial image editing model, the above method further includes: obtaining a second sample set, where the second sample set includes a plurality of second samples, and each second sample includes a third face image and a second description text, and the second description text is used to describe the features of the third face image; training a pre-trained text-to-image model according to the second sample set; and constructing an initial image editing model according to the pre-trained text-to-image model.
[0052] Exemplarily, a preset image editing model can be obtained; the preset image editing model includes a generative diffusion sub-model; replacing the generative diffusion sub-model part in the preset image editing model with the text-to-image model to obtain the initial image editing model. The structure of the preset image editing model is based on the open-source InstructPix2Pix model structure and can include a generative diffusion sub-model, a variational autoencoder (VAE), and a multimodal model (Contrastive Language–Image Pretraining, CLIP). The VAE can be used to encode and decode images, and the CLIP model can be used to convert text into text feature vectors. The type of the text-to-image model can be a generative diffusion model. Optionally, the above text-to-image model can be trained according to the second sample set.
[0053] Exemplarily, obtain a second sample set, where the second sample set includes a plurality of second samples, and each second sample includes a third face image and a second description text, and the second description text is used to describe the features of the third face image; train a text-to-image model according to the second sample set. Specifically, a large amount of face parameter data can be collected in advance, and each face parameter data is rendered through a game engine to obtain the corresponding face image. After obtaining a large amount of face parameter data, such as 200,000 sets of face parameter data, 50,000 sets of face parameter data can be randomly selected from the 200,000 sets of face parameter data, and the face images corresponding to the 50,000 sets of face parameter data are used as the third face images, that is, there are 50,000 third face images. Feature description (which can be manual description) is performed on these 50,000 third face images, and the obtained description text is used as the second description text, and thus the second sample set can be obtained.
[0054] Exemplarily, the second description text can only describe the characteristic parts and does not describe all the facial features in the face image, and there is no limitation in this regard. For example, the second description text can include "female, blue eyes, purple eyeshadow, gray-black lipstick, fair skin", "female, moist lips, slanting eyebrows, looks like a cute little girl", "male, yellow vertical pupils, a golden pattern between the eyebrows", "lips painted with dark red lipstick", etc.
[0055] In an alternative embodiment, the above method further includes: deleting target samples from a plurality of first samples according to the labels; the target samples include samples whose values corresponding to the labels of the second face images are less than or equal to a first preset threshold and greater than or equal to a second preset threshold.
[0056] Exemplarily, the above positive samples, that is, the samples whose values represented by the labels of the second face images are greater than the first preset threshold, can be understood as samples that meet the user's preferences, and the above negative samples, that is, the samples whose values represented by the labels of the second face images are less than the first preset threshold, can be understood as samples that do not meet the user's preferences. The target samples, that is, the samples whose values corresponding to the labels of the second face images are less than or equal to the first preset threshold and greater than or equal to the second preset threshold, can be understood as fuzzy samples, that is, samples that are not very distinguishable whether they meet the user's preferences. The target samples can be directly deleted to improve the distinguishability between positive and negative samples.
[0057] Exemplarily, when the value represented by the label of the second face image is 1-5, this value can represent the score of the second face image. When the first preset threshold is 3 and the second preset threshold is 2, the first samples whose values represented by the labels of the second face images are greater than 3 can be recorded as positive samples, and those whose values represented by the labels of the second face images are less than 2 can be recorded as negative samples. Then, the samples whose values represented by the labels of the second face images are greater than or equal to 2 and less than or equal to 3 are target samples, and the samples whose values represented by the labels of the second face images are greater than or equal to 2 and less than or equal to 3 can be directly deleted.
[0058] In an alternative embodiment, the above method further includes: determining the difference between the noise added to the second face image and the predicted noise output by the initial image editing model based on the first face image and the first description text through a first loss function.
[0059] Exemplarily, the first loss function can be determined by formula (1).
[0060]
[0061] Wherein, represents the expectation operator, ω(t) represents the weight function related to the sampling time t, x t represents the image obtained after adding noise to the second face image at the sampling time t, c T represents the feature vector corresponding to the first description text, c I represents the feature vector corresponding to the first face image, ∈ represents the noise added to the second face image at the sampling time t, ∈ ref (x t , t, c T , c I ) represents the noise predicted by the first image editing model at the sampling time t, ‖‖2 Denotes the Euclidean norm, also known as the L2 norm, which is used to measure the difference between two vectors.
[0062] Exemplarily, the process of adding noise to the second face image can be understood as gradually adding noise to the second face image within the time period from 0 to T, adding noise to the second face image once every preset time interval within the time period from 0 to T, and the moment of adding noise is the sampling moment t.
[0063] In an optional embodiment, the above method further includes: determining, through a second loss function, the difference between the noise added to the second face image in the positive samples and the predicted noise output by the initial image editing model based on the first face image and the first description text, and the difference between the noise added to the second face image in the negative samples and the predicted noise output by the initial image editing model based on the first face image and the first description text.
[0064] Exemplarily, the second loss function can be determined by formula (2).
[0065]
[0066] Among them, Denotes the expectation operator, ∈ dpo Denotes the noise predicted by the second image editing model at the sampling moment t, ω(t) denotes the weight function related to the sampling moment t, Denotes the image obtained after adding noise to the second face image at the sampling moment t in the positive samples, Denotes the image obtained after adding noise to the second face image at the sampling moment t in the negative samples, ∈ w Denotes the noise added to the second face image at the sampling moment t in the positive samples, ∈ l Denotes the noise added to the second face image at the sampling moment t in the negative samples, ∈ ref Denotes the noise predicted by the first image editing model at the sampling moment t, ‖‖ 2 Denotes the Euclidean norm, also known as the L2 norm, which is used to measure the difference between two vectors.
[0067] The above is a specific implementation manner of model training. In specific implementation, the model can also be trained by other feasible methods, and the embodiments of the present application are not limited thereto.
[0068] Corresponding to the model training method provided in the first embodiment of the present application, the second embodiment of the present application also provides a model training device, as Figure 5 shown, the model training device 500 includes:
[0069] An acquisition unit 501, configured to acquire a first sample set; the first sample set includes a plurality of first samples, and each first sample includes a first face image, a second face image obtained by editing the first face image, and a first description text, where the first description text is used to describe the changes from the first face image to the second face image;
[0070] A training unit 502, configured to input the first face image and the first description text into an initial image editing model, and train the initial image editing model according to the difference between the noise added by the initial image editing model to the second face image and the noise predicted by the initial image editing model based on the first face image and the first description text, so as to obtain a first image editing model for modifying an image according to text;
[0071] A determination unit 503, configured to determine a plurality of positive samples and a plurality of negative samples from the plurality of first samples;
[0072] The training unit 502 is further configured to input the first face image and the first description text into the first image editing model, and train the first image editing model according to the difference between the noise added by the first image editing model to the second face image in the positive samples and the noise predicted by the first image editing model based on the first face image and the first description text, and the difference between the noise added by the first image editing model to the second face image in the negative samples and the noise predicted by the first image editing model based on the first face image and the first description text, so as to obtain a second image editing model.
[0073] Optionally, as Figure 5 shown, the model training device 500 further includes: a processing unit 504.
[0074] The acquisition unit 501 is further configured to acquire a preset image editing model; the preset image editing model includes a generative diffusion sub-model; the processing unit 504 is configured to replace the generative diffusion sub-model in the preset image editing model with a pre-trained text-to-image model to obtain an initial image editing model.
[0075] Optionally, the training unit 502 is specifically configured to determine the difference between the noise added by the first image editing model to the second face image in the positive sample and the noise predicted by the first image editing model based on the first face image and the first description text, and the difference between the noise added by the first image editing model to the second face image in the negative sample and the noise predicted by the first image editing model based on the first face image and the first description text; and train the first image editing model according to the difference between the noise added by the first image editing model to the second face image in the positive sample and the noise predicted by the first image editing model based on the first face image and the first description text, and the difference between the noise added by the first image editing model to the second face image in the negative sample and the noise predicted by the first image editing model based on the first face image and the first description text, to obtain the second image editing model.
[0076] Optionally, the determining unit 503 is specifically configured to obtain the label of each first sample in the first sample set; determine a plurality of positive samples and a plurality of negative samples from the plurality of first samples according to the labels; the positive sample is a sample in which the value represented by the label of the second face image is greater than the first preset threshold, and the negative sample is a sample in which the value corresponding to the label of the second face image is less than the second preset threshold.
[0077] As Figure 3 shown, the model training device 500 further includes: a deleting unit 505.
[0078] The deleting unit 505 is configured to delete target samples from the plurality of first samples according to the labels; the target samples include samples in which the value corresponding to the label of the second face image is less than or equal to the first preset threshold and greater than or equal to the second preset threshold.
[0079] Optionally, the label includes an aesthetic label and / or a semantic consistency label.
[0080] Optionally, the determining unit 503 is further configured to determine the difference between the noise added by the initial image editing model to the second face image and the noise predicted by the initial image editing model based on the first face image and the first description text through the first loss function.
[0081] Optionally, the first loss function includes: Wherein, represents the expectation operator, ω(t) represents the weight function related to the sampling time t, t ∈ (0, T), x t represents the image obtained after adding noise to the second face image at the sampling time t, c T represents the feature vector corresponding to the first description text, c I represents the feature vector corresponding to the first face image, ∈ represents the noise added to the second face image at the sampling time t, ∈ ref(x t , t, c T , c I ) represents the noise predicted by the first image editing model for the sampling time t.
[0082] Optionally, the determination unit 503 is further configured to determine, by using a second loss function, the difference between the noise added by the first image editing model to the second face image in the positive sample and the noise predicted by the first image editing model based on the first face image and the first description text, and the difference between the noise added by the first image editing model to the second face image in the negative sample and the noise predicted by the first image editing model based on the first face image and the first description text.
[0083] Optionally, the second loss function includes: Wherein, represents the expectation operator, ∈ dpo represents the noise predicted by the second image editing model for the sampling time t, ω(t) represents the weight function related to the sampling time t, represents the image obtained after adding noise to the second face image at the sampling time t in the positive sample, represents the image obtained after adding noise to the second face image at the sampling time t in the negative sample, ∈ w represents the noise added to the second face image at the sampling time t in the positive sample, ∈ l represents the noise added to the second face image at the sampling time t in the negative sample, ∈ ref represents the noise predicted by the first image editing model for the sampling time t.
[0084] Optionally, the first description text is used to describe the facial image feature changes occurring from the first face image to the second face image; the facial image features include at least one of facial features, face shape, and makeup.
[0085] Optionally, as Figure 5 shown, the model training device 500 further includes: a construction unit 506.
[0086] The acquisition unit 501 is further configured to acquire a second sample set, the second sample set includes a plurality of second samples, each second sample includes a third face image and a second description text, the second description text is used to describe the features of the third face image; and train a text-to-image model according to the second sample set; the construction unit 506 is configured to construct an initial image editing model according to the pre-trained text-to-image model.
[0087] Corresponding to the model training method provided in the first embodiment of the present application, the third embodiment of the present application further provides an electronic device for model training.
[0088] As Figure 6 shown, the following is a block diagram of an example of an electronic device for model training provided by an embodiment of the present application.
[0089] In this embodiment, an optional hardware structure of the electronic device 600 may be as Figure 4 shown, including: at least one processor 601, at least one memory 602, and at least one communication bus 605; the memory 602 contains a program 603 and data 604.
[0090] The bus 605 may be a communication device for transmitting data between components inside the electronic device 600, such as an internal bus (e.g., a CPU-memory bus, where the processor is the central processing unit, abbreviated as CPU), an external bus (e.g., a universal serial bus port, a peripheral component interconnect express port), etc.
[0091] In addition, the electronic device further includes: at least one network interface 606 and at least one peripheral interface 607. The network interface 606 provides wired or wireless communication related to an external network 608 (e.g., the Internet, an intranet, a local area network, a mobile communication network, etc.); in some embodiments, the network interface 606 may include any combination of any number of network interface controllers (abbreviated as NIC in English), radio frequency (abbreviated as RF in English) modules, repeaters, transceivers, modems, routers, gateways, wired network adapters, wireless network adapters, Bluetooth adapters, infrared adapters, near field communication (abbreviated as NFC in English) adapters, cellular network chips, etc.
[0092] The peripheral interface 607 is used to connect to peripherals, and the peripherals may be, for example, peripheral 1 ( Figure 6 609 in Figure 6 ), peripheral 2 ( Figure 6 610 in
[0093] The processor 601 may be a CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present application.
[0094] The memory 602 may include a high-speed RAM (Random Access Memory) memory, and may also include a non-volatile memory, such as at least one disk memory.
[0095] Wherein, the processor 601 invokes the programs and data stored in the memory 602 and executes the following steps:
[0096] Obtain a first sample set; the first sample set includes a plurality of first samples, and each first sample includes a first face image, a second face image obtained by editing the first face image, and a first description text, where the first description text is used to describe the changes from the first face image to the second face image;
[0097] Input the first face image and the first description text into an initial image editing model, and train the initial image editing model according to the difference between the noise added to the second face image and the noise predicted based on the first face image and the first description text, to obtain a first image editing model for modifying an image according to text;
[0098] Determine a plurality of positive samples and a plurality of negative samples from the plurality of first samples;
[0099] Input the first face image and the first description text into the first image editing model, and train the first image editing model according to the difference between the noise added to the second face image in the positive samples and the noise predicted based on the first face image and the first description text, and the difference between the noise added to the second face image in the negative samples and the noise predicted based on the first face image and the first description text, to obtain a second image editing model.
[0100] Corresponding to the data processing method in the human-computer interaction provided in the first embodiment of the present application, the fourth embodiment of the present application provides a computer-readable storage medium storing a program for the data processing method in the human-computer interaction, and the program is run by a processor to execute the following steps:
[0101] Obtain a first sample set; the first sample set includes a plurality of first samples, and each first sample includes a first face image, a second face image obtained by editing the first face image, and a first description text, where the first description text is used to describe the changes from the first face image to the second face image;
[0102] Input the first face image and the first description text into the initial image editing model, and train the initial image editing model according to the difference between the noise added to the second face image and the noise predicted based on the first face image and the first description text, so as to obtain the first image editing model for modifying the image according to the text;
[0103] Determine a plurality of positive samples and a plurality of negative samples from a plurality of first samples;
[0104] Input the first face image and the first description text into the first image editing model, and train the first image editing model according to the difference between the noise added to the second face image in the positive sample and the noise predicted based on the first face image and the first description text, and the difference between the noise added to the second face image in the negative sample and the noise predicted based on the first face image and the first description text, so as to obtain the second image editing model.
[0105] It should be noted that for the detailed descriptions of the devices, electronic devices and computer-readable storage media provided in the second, third and fourth embodiments of the present application, reference may be made to the relevant descriptions of the first embodiment of the present application, which will not be repeated here.
[0106] Although the present application is disclosed above with preferred embodiments, it is not intended to limit the present application. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of the present application. Therefore, the protection scope of the present application should be defined by the scope defined in the claims of the present application.
[0107] In a typical configuration, the node devices in the blockchain include one or more processors (CPUs), input / output interfaces, network interfaces and memories.
[0108] The memory may include non-permanent memory in the computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of, for example, read-only memory (ROM) or flash memory (flash RAM). The memory is an example of the computer-readable medium.
[0109] 1. A computer-readable medium includes both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), random access memory (RAM) of other properties, read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage media, or any other non-transitory medium that can be used to store information accessible by a computing device. As defined herein, a computer-readable medium does not include transitory media such as modulated data signals and carrier waves.
[0110] 2. Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0111] Although the present application is disclosed above in preferred embodiments, it is not intended to limit the present application. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of the present application. Therefore, the protection scope of the present application should be determined by the scope defined by the claims of the present application.
Claims
1. A model training method, characterized in that: The method comprises: Acquire a first sample set; the first sample set includes a plurality of first samples, the first sample includes a first face image, a second face image edited from the first face image, and a first description text, the first description text is used to describe the changes from the first face image to the second face image; Inputting the first facial image and the first description text into an initial image editing model, and training the initial image editing model according to a difference between the noise added to the second facial image and the noise predicted based on the first facial image and the first description text, to obtain a first image editing model for modifying an image according to text; Determine a plurality of positive samples and a plurality of negative samples from the plurality of first samples; The first facial image and the first descriptive text are input into the first image editing model, and the first image editing model is trained based on the difference between the noise added to the second facial image in the positive sample and the noise predicted based on the first facial image and the first descriptive text, and the difference between the noise added to the second facial image in the negative sample and the noise predicted based on the first facial image and the first descriptive text to obtain a second image editing model.
2. The method according to claim 1, characterized in that Before inputting the first face image and the first description text into the initial image editing model, the method includes: Acquire a preset image editing model; the preset image editing model includes generating a diffusion sub-model; The generation diffusion sub-model in the preset image editing model is replaced with a pre-trained text generation image model to obtain the initial image editing model.
3. The method according to claim 1, characterized in that Before training the first image editing model according to the difference between the noise added by the first image editing model to the second face image in the positive sample and the noise predicted by the first image editing model based on the first face image and the first description text, and the difference between the noise added by the first image editing model to the second face image in the negative sample and the noise predicted by the first image editing model based on the first face image and the first description text to obtain the second image editing model, the method further includes: Determine the difference between the noise added by the first image editing model to the second face image in the positive sample and the noise predicted by the first image editing model based on the first face image and the first descriptive text, and the difference between the noise added by the first image editing model to the second face image in the negative sample and the noise predicted by the first image editing model based on the first face image and the first descriptive text.
4. The method according to claim 1, characterized in that: The determining a plurality of positive samples and a plurality of negative samples from the plurality of first samples comprises: Obtaining a label of each first sample in the first sample set; According to the label, multiple positive samples and multiple negative samples are determined from the multiple first samples; the positive samples are samples whose numerical values represented by the labels of the second facial images are greater than a first preset threshold, and the negative samples are samples whose numerical values corresponding to the labels of the second facial images are less than a second preset threshold.
5. The method according to claim 4, characterized in that The method further comprises: Delete target samples from the multiple first samples according to the label; the target samples include samples whose values corresponding to the labels of the second face images are less than or equal to the first preset threshold and greater than or equal to the second preset threshold.
6. The method according to claim 4, characterized in that The tags include aesthetic tags and / or semantic consistency tags.
7. The method according to claim 1, characterized in that The method further comprises: The difference between the noise added to the second facial image and the noise predicted based on the first facial image and the first description text is determined by a first loss function.
8. The method according to claim 7, characterized in that The first loss function includes: in, represents the expectation operator, ω(t) represents the weight function related to the sampling time t, t∈(0,T), x t represents the image obtained after adding noise to the second face image at sampling time t, c T represents the feature vector corresponding to the first description text, c I represents the feature vector corresponding to the first face image, ∈ represents the noise added to the second face image at sampling time t, ∈ ref (x t ,t,c T ,c I ) represents the noise predicted by the first image editing model for sampling time t.
9. The method according to claim 1, characterized in that: The method further comprises: The difference between the noise added to the second face image in the positive sample and the noise predicted based on the first face image and the first description text, and the difference between the noise added to the second face image in the negative sample and the noise predicted based on the first face image and the first description text are determined by a second loss function.
10. The method according to claim 9, characterized in that The second loss function includes: in, represents the expectation operator, ∈ dpo represents the noise predicted by the second image editing model for sampling time t, ω(t) represents the weight function related to sampling time t, represents the image obtained after adding noise to the second face image in the positive sample at sampling time t, represents the image obtained by adding noise to the second face image in the negative sample at sampling time t, ∈ w represents the noise added to the second face image in the positive sample at sampling time t, ∈ l represents the noise added to the second face image in the negative sample at sampling time t, ∈ ref Represents the noise predicted by the first image editing model for sampling time t.
11. The method according to claim 1, characterized in that: The first description text is used to describe the changes in facial image features from the first facial image to the second facial image; the facial image features include at least one of facial features, face shape and makeup.
12. The method according to claim 1, characterized in that Before inputting the first face image and the first description text into the initial image editing model, the method further includes: Acquire a second sample set, where the second sample set includes a plurality of second samples, where the second samples include a third face image and a second description text, where the second description text is used to describe features of the third face image; Obtaining a pre-trained text-to-image model based on the second sample set; An initial image editing model is constructed based on the pre-trained text-generated image model.
13. A model training device, characterized in that: The device comprises: an acquisition unit, configured to acquire a first sample set; the first sample set includes a plurality of first samples, the first sample includes a first face image, a second face image edited from the first face image, and a first description text, the first description text being used to describe changes from the first face image to the second face image; a training unit, configured to input the first facial image and the first description text into the initial image editing model, and train the initial image editing model according to a difference between the noise added to the second facial image and the noise predicted based on the first facial image and the first description text, so as to obtain a first image editing model for modifying an image according to text; A determining unit, configured to determine a plurality of positive samples and a plurality of negative samples from the plurality of first samples; The training unit is also used to input the first facial image and the first description text into the first image editing model, and train the first image editing model according to the difference between the noise added to the second facial image in the positive sample and the noise predicted based on the first facial image and the first description text, and the difference between the noise added to the second facial image in the negative sample and the noise predicted based on the first facial image and the first description text, so as to obtain a second image editing model.
14. An electronic device, characterized in that: include: processor; as well as The memory is used to store a data processing program. After the electronic device is powered on and the program is run by the processor, the method according to any one of claims 1 to 12 is executed.
15. A computer-readable storage medium, characterized in that: A data processing program is stored, and the program is run by a processor to execute the method according to any one of claims 1 to 12.