Picture generation method and device, model training method and device, equipment and storage medium

By receiving image generation instructions and determining the target makeup parameters, the original portrait image is processed using an image generation model, which solves the problem of mismatch between makeup and generated style and achieves matching between makeup and style.

CN121600092APending Publication Date: 2026-03-03GUANGZHOU SHIYINLIAN SOFTWARE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511770799.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-28
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

The existing technology addresses the issue of mismatched makeup in generated portrait images with the selected generation style.

Method used

By receiving image generation instructions, the target makeup parameters are determined, and the original portrait image is processed using an image generation model based on the target makeup parameters and the target generation style to generate the target portrait image.

Benefits of technology

Ensure that the makeup in the generated portrait image matches the target style to avoid mismatches.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121600092A_ABST
    Figure CN121600092A_ABST
Patent Text Reader

Abstract

The invention discloses a picture generation method and device, a model training method and device, equipment and a storage medium, and belongs to the field of picture generation. The method comprises the steps that a picture generation instruction is received, the picture generation instruction comprises an original figure picture and a target generation style, and the target generation style is used for indicating the style of a picture generated according to the original figure picture; a target makeup parameter matched with the target generation style is determined in multiple preset makeup parameters, and the preset makeup parameters are used for indicating a processing mode for adding makeup to the portrait picture; and processing the original figure picture through an image generation model according to the target makeup parameters and the target generation style, generating a target portrait picture corresponding to the original figure picture, the image generation model being obtained by training a plurality of sample portrait pictures under a plurality of preset makeup parameters. The problem that the makeup of the generated portrait picture is not matched with the selected generation style can be avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image generation, and in particular to an image generation method, model training method, apparatus, device, and storage medium. Background Technology

[0002] Image generation technology can be used to generate images that meet the requirements of a user account. For example, images provided by a user account can be processed according to the user account's requirements to generate new images that match the user account's specifications.

[0003] In related technologies, image generation models can be used to generate portrait images according to the style selected by the user account, such as generating retro-style or sweet-style portrait images. During the portrait image generation process, the user account needs to select a generation style and provide an original portrait image. The image generation model then processes the provided portrait image according to the selected style to generate a portrait image that matches the chosen style.

[0004] When generating portrait images using the above method, there may be a mismatch between the makeup of the generated portrait image and the selected generation style. Summary of the Invention

[0005] This application provides an image generation method, a model training method, an apparatus, a device, and a storage medium, which can avoid the problem of mismatch between the makeup in the generated portrait image and the selected generation style. The technical solution is as follows: According to one aspect of this application, an image generation method is provided, the method comprising: Receive an image generation instruction, the image generation instruction including an original portrait image and a target generation style, the target generation style being used to indicate the style of the image generated based on the original portrait image; Among a variety of preset makeup parameters, a target makeup parameter that matches the target generation style is determined. The preset makeup parameters are used to indicate the processing method for adding makeup to portrait images. The image generation model processes the original portrait image according to the target makeup parameters and the target generation style to generate the target portrait image corresponding to the original portrait image. The image generation model is trained on multiple sample portrait images under the various preset makeup parameters.

[0006] According to another aspect of this application, a model training method is provided, the method comprising: Acquire portrait images; The captured portrait image is processed according to each of the various preset makeup parameters to obtain multiple sample portrait images; The image generation model is trained using the aforementioned multiple sample portrait images; The image generation model is used to process the original portrait image according to the target makeup parameters and the target generation style to generate the target portrait image corresponding to the original portrait image. The target makeup parameters are determined from a variety of preset makeup parameters according to the target generation style. The original portrait image and the target generation style belong to the image generation instructions.

[0007] According to another aspect of this application, an image generation apparatus is provided, the apparatus comprising: A receiving module is used to receive an image generation instruction, the image generation instruction including an original portrait image and a target generation style, the target generation style being used to indicate the style of the image generated based on the original portrait image; The determination module is used to determine the target makeup parameters that match the target generation style from a variety of preset makeup parameters. The preset makeup parameters are used to indicate the processing method for adding makeup to portrait images. The generation module is used to process the original portrait image according to the target makeup parameters and the target generation style through an image generation model to generate a target portrait image corresponding to the original portrait image. The image generation model is trained on multiple sample portrait images under the various preset makeup parameters.

[0008] In an optional design, the determining module is configured to: determine the target makeup parameters based on the similarity between the target generated style and different preset makeup parameters among the multiple preset makeup parameters.

[0009] In an optional design, the determining module is configured to: convert the target generated style into a first feature vector; convert the preset makeup parameters into a second feature vector; determine the similarity between the first feature vector and different second feature vectors; and determine the preset makeup parameters corresponding to the second feature vector with the highest similarity as the target makeup parameters.

[0010] In an optional design, the determining module is used to: determine the preset makeup parameters corresponding to the target generation style in the preset mapping relationship as the target makeup parameters; wherein, the preset mapping relationship includes the correspondence between different generation styles and different preset makeup parameters.

[0011] In an optional design, the image generation model is trained using multiple sample portrait images under the various preset makeup parameters and parameter identifiers corresponding to the various preset makeup parameters respectively; the determining module is used to: determine the target parameter identifier corresponding to the target generation style from the parameter identifiers corresponding to the various preset makeup parameters respectively; The generation module is used to: process the original portrait image according to the target parameter identifier and the target generation style through the image generation model, and generate a target portrait image corresponding to the original portrait image.

[0012] In an optional design, the generation instruction further includes generation prompts corresponding to the target generation style; the generation module is used to: process the original portrait image through the image generation model according to the target makeup parameters, the target generation style and the generation prompts, and generate a target portrait image corresponding to the original portrait image.

[0013] In an optional design, the device further includes: a retrieval module, used to retrieve portrait images that match the target generation style as keywords; The generation module is used to: process the original portrait image according to the target makeup parameters, the target generation style, and the retrieved portrait image using the image generation model, and generate a target portrait image corresponding to the original portrait image.

[0014] In an optional design, the determining module is configured to: determine the similarity between the face regions in the original portrait image and the face regions in different retrieved portrait images; and determine the retrieved portrait image with the highest similarity as the reference portrait image; The generation module is used to: process the original portrait image using the image generation model according to the target makeup parameters, the target generation style, and the reference portrait image, and generate a target portrait image corresponding to the original portrait image.

[0015] In an optional design, the device further includes: an acquisition module, configured to acquire historical interaction images corresponding to the first user account, the historical interaction images including images of past interactions with the first user account; The determining module is used to: determine the matching degree between the historical interaction image and the target generation style; and determine the historical interaction image with the highest matching degree as the reference image; The generation module is used to: process the original portrait image according to the target makeup parameters, the target generation style, and the reference image using the image generation model, and generate a target portrait image corresponding to the original portrait image.

[0016] According to another aspect of this application, a model training apparatus is provided, the apparatus comprising: The acquisition module is used to acquire captured portrait images; The processing module is used to process the captured portrait image according to each of the multiple preset makeup parameters to obtain multiple sample portrait images; The training module is used to train an image generation model using the multiple sample portrait images; The image generation model is used to process the original portrait image according to the target makeup parameters and the target generation style to generate the target portrait image corresponding to the original portrait image. The target makeup parameters are determined from a variety of preset makeup parameters according to the target generation style. The original portrait image and the target generation style belong to the image generation instructions.

[0017] According to another aspect of this application, a computer device is provided, the computer device including a processor and a memory, the memory storing at least one program, the at least one program being loaded and executed by the processor to implement the image generation method or model training method as described above.

[0018] According to another aspect of this application, a computer-readable storage medium is provided, wherein at least one program is stored therein, the at least one program being loaded and executed by a processor to implement the image generation method or model training method as described above.

[0019] According to another aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the image generation method or model training method provided in various alternative implementations of the above aspects.

[0020] The beneficial effects of the technical solution provided in this application include at least the following: An image generation model is used to process the original portrait image based on the target makeup parameters and the target generation style to generate the target portrait image. Since the target makeup parameters match the target generation style, and the image generation model is trained on portrait images with various preset makeup parameters, it has the ability to generate portrait images with multiple preset makeup parameters. Therefore, it can be guaranteed that the makeup of the generated target portrait image matches the target generation style, avoiding the problem of mismatch between the generated portrait image's makeup and the selected generation style. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 This is a structural block diagram of a computer system provided in an exemplary embodiment of this application; Figure 2 This is a schematic diagram illustrating the process of generating a human portrait image provided in an exemplary embodiment of this application; Figure 3 This is a flowchart illustrating an exemplary embodiment of the image generation method provided in this application; Figure 4 This is a flowchart illustrating an exemplary embodiment of the image generation method provided in this application; Figure 5 This is a schematic diagram of the structure of a machine learning model provided in an exemplary embodiment of this application; Figure 6 This is a flowchart illustrating a model training method provided in an exemplary embodiment of this application; Figure 7 This is a schematic diagram of the model training and inference process provided in an exemplary embodiment of this application; Figure 8 This is a schematic diagram of a model training process provided in an exemplary embodiment of this application; Figure 9 This is a schematic diagram of a model reasoning process provided in an exemplary embodiment of this application; Figure 10 This is a schematic diagram of the structure of an image generation apparatus provided in an exemplary embodiment of this application; Figure 11 This is a schematic diagram of the structure of an image generation apparatus provided in an exemplary embodiment of this application; Figure 12 This is a schematic diagram of the structure of an image generation apparatus provided in an exemplary embodiment of this application; Figure 13 This is a schematic diagram of the structure of a model training apparatus provided in an exemplary embodiment of this application; Figure 14 This is a schematic diagram of the structure of a computer device provided in an exemplary embodiment of this application.

[0023] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. Detailed Implementation

[0024] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0025] Figure 1 This is a structural block diagram of a computer system provided in an exemplary embodiment of this application. The computer system 100 includes: a terminal 110 and a server 120.

[0026] Terminal 110 has an application 111 installed and running that supports image generation. This application 111 includes any one of the following: Artificial Intelligence (AI) assistant applications, image beautification applications, photography applications, music listening applications, song applications, singing applications, live streaming applications, social applications, office applications, game applications, food delivery applications, online shopping applications, video-on-demand applications, short video applications, financial applications, lifestyle service applications, navigation applications, medical applications, learning applications, and mini-programs. In some embodiments, application 111 is a client. Terminal 110 is the terminal used by user 112, and user account 112 can be logged into application 111. Terminal 110 can refer to one of multiple terminals. Optionally, the device type of terminal 110 includes at least one of the following: smartphone, tablet, smartwatch, e-book reader, MP3 player, MP4 player, laptop, and desktop computer.

[0027] Terminal 110 is connected to server 120 via wireless or wired network.

[0028] Server 120 includes at least one of a single server, multiple servers, a cloud computing platform, and a virtualization center. Server 120 provides background services for application 111 that supports image generation. Optionally, server 120 undertakes the primary computing task, and terminal 110 undertakes the secondary computing task; or, server 120 undertakes the secondary computing task, and terminal 110 undertakes the primary computing task; or, server 120 and terminal 110 collaborate on computing using a distributed computing architecture.

[0029] In an illustrative example, server 120 includes processor 121, user account database 122, image generation module 123, and user-facing input / output interface (I / O interface) 124. Processor 121 loads instructions stored in server 120 and processes data in user account database 122 and image generation module 123. User account database 122 stores user account data used by terminal 110 and other terminals, such as user account avatars, nicknames, and group memberships. Image generation module 123 provides services for image generation in application 111. User-facing I / O interface 124 establishes communication and exchanges data with terminal 110 via wireless or wired network.

[0030] Based on the above description of the implementation environment involved in this application, the method provided in the embodiments of this application will be described below.

[0031] Figure 2 This is a schematic diagram illustrating the process of generating a human portrait image provided in an exemplary embodiment of this application. For example... Figure 2 As shown, the computer device receives an image generation instruction 201, which includes an original portrait image 202 and a target generation style 203. In some embodiments, the image generation instruction 201 is triggered by a first user account, and the original portrait image 202 is a portrait image provided by the first user account, such as a portrait image corresponding to a user using the first user account. The target generation style 203 is used to indicate the style of the image generated based on the original portrait image 202. When the image generation instruction 201 is triggered by the first user account, the target generation style 203 can be selected by the first user account from multiple preset generation styles. The computer device then determines a target makeup parameter 205 that matches the target generation style 203 from multiple preset makeup parameters 204. The preset makeup parameters 204 are used to indicate the processing method for adding makeup to the portrait image, including parameters corresponding to makeup under different processing dimensions. In some embodiments, the computer device determines the target makeup parameter 205 based on the similarity between the target generation style 203 and different preset makeup parameters 204 among the multiple preset makeup parameters 204. In some embodiments, the computer device determines the preset makeup parameters 204 corresponding to the target generation style 203 in the preset mapping relationship as the target makeup parameters 205. The preset mapping relationship includes the correspondence between different generation styles and different preset makeup parameters 204.

[0032] Given the original portrait image 202, the target generation style 203, and the target makeup parameters 205, the computer device uses an image generation model 206 to process the original portrait image 202 according to the target makeup parameters 205 and the target generation style 203, generating a target portrait image 207 corresponding to the original portrait image 202. The image generation model 206 is trained using multiple sample portrait images under various preset makeup parameters 204. These multiple sample portrait images are obtained by processing photographed portrait images according to each of the various preset makeup parameters 204. For example, the image generation model 206 is trained using multiple sample portrait images corresponding to a first user account under various preset makeup parameters 204. These multiple sample portrait images are obtained by processing photographed portrait images of the first user account according to each of the various preset makeup parameters 204. In some embodiments, the image generation model 206 is trained using sample portrait images under various preset makeup parameters 204 and parameter identifiers corresponding to each of the various preset makeup parameters 204. In this case, the computer device determines the target parameter identifier corresponding to the target generation style 203 from the parameter identifiers corresponding to various preset makeup parameters 204. The correspondence between different generation styles and different parameter identifiers can be preset. Then, the computer device processes the original portrait image 202 according to the target parameter identifier and the target generation style 203 through the image generation model 206 to generate the target portrait image 207 corresponding to the original portrait image 202.

[0033] An image generation model is used to process the original portrait image based on the target makeup parameters and the target generation style to generate the target portrait image. Since the target makeup parameters match the target generation style, and the image generation model is trained on portrait images with various preset makeup parameters, it has the ability to generate portrait images with multiple preset makeup parameters. Therefore, it can be guaranteed that the makeup of the generated target portrait image matches the target generation style, avoiding the problem of mismatch between the generated portrait image's makeup and the selected generation style.

[0034] Figure 3 This is a schematic flowchart illustrating an exemplary embodiment of an image generation method provided in this application. The method can be used in a computer device, such as one used for... Figure 1 The server shown. (As shown) Figure 3 As shown, the method includes: Step 302: Receive image generation instructions.

[0035] The image generation instruction includes an original portrait image and a target generation style, or it can be understood as carrying an original portrait image and a target generation style. The original portrait image is a portrait image to be processed, and it includes at least one of the following: a facial image, an upper body image, and a full-body image. It should be noted that the person in the original portrait image can be a real person or a virtual person, and this application embodiment does not impose any restrictions on this.

[0036] The target generation style is used to indicate the style of the image generated from the original portrait image. Optionally, the target generation style is used to indicate at least one of the background style, clothing style, and styling style corresponding to the image generated from the original portrait image. For example, the target generation style includes at least one of retro style, professional style, and sweet style. It should be noted that the styles mentioned in the above examples are only used as examples and are not intended to limit the target generation style in the embodiments of this application. It is understood that, in addition to the styles mentioned in the above examples, the target generation style may include more styles, and the embodiments of this application do not limit this.

[0037] In some embodiments, the image generation instruction is triggered by a first user account, for example, by the first user account within a client, which includes clients that support image generation functionality. The original portrait image is a portrait image provided by the first user account, which may be uploaded by the first user account within the client or taken through the client, for example, a portrait image corresponding to the user using the first user account. The target generation style is the style selected by the first user account from a variety of preset generation styles provided by the client.

[0038] Step 304: Among various preset makeup parameters, determine the target makeup parameters that match the target generated style.

[0039] Preset makeup parameters are pre-defined makeup settings, such as manually set makeup parameters. Each preset makeup parameter is different. Preset makeup parameters are used to indicate how makeup is applied to portrait images. For example, they indicate how makeup is applied to the facial area of ​​a portrait image. Different preset makeup parameters indicate different ways of applying makeup to a portrait image.

[0040] In some embodiments, the preset makeup parameters include parameters corresponding to the makeup under different processing dimensions, such as parameters corresponding to the degree of skin smoothing, parameters corresponding to the eyeshadow color, parameters corresponding to the filter intensity, parameters corresponding to the lipstick color, parameters corresponding to the blush color, and parameters corresponding to the skin tone brightness. It should be noted that the parameters mentioned in the above examples are for illustrative purposes only and are not intended to limit the preset makeup parameters in the embodiments of this application. It is understood that, in addition to the parameters mentioned in the above examples, the preset makeup parameters may include more parameters, and the embodiments of this application do not impose any limitations on this.

[0041] Optionally, the computer device determines the target makeup parameters based on the similarity between the target generated style and different preset makeup parameters among a variety of preset makeup parameters. Alternatively, the computer device determines the target makeup parameters based on the correspondence between different preset makeup parameters and different generated styles, combined with the target generated style.

[0042] Step 306: Process the original portrait image using an image generation model based on the target makeup parameters and the target generation style to generate the target portrait image corresponding to the original portrait image.

[0043] Image generation models include AI models that support the generation of images that match input information. Exemplarily, the image generation models in this application embodiment include at least one of the following: a Stable Diffusion model (a text-to-image generation model), a Stable Diffusion model fine-tuned based on Low-Rank Adaptation (LoRA), and a Vision Language Model (VLM). It should be noted that the models mentioned in the above examples are for illustrative purposes only and are not intended to limit the image generation models in this application embodiment. It is understood that, in addition to the models mentioned in the above examples, image generation models may also include other types of models, and this application embodiment does not impose any limitations on them.

[0044] The image generation model is trained using multiple sample portrait images under various preset makeup parameters. For example, it can be trained on an open-source or pre-trained model using these sample portrait images. The multiple sample portrait images include portrait images corresponding to each of the preset makeup parameters. Each sample portrait image under a preset makeup parameter is a portrait image with the makeup corresponding to that parameter. These multiple sample portrait images are obtained by processing photographed portrait images according to each of the preset makeup parameters. The photographed portrait images can be taken directly, such as portrait images taken for a first user account. By training the model using multiple sample portrait images under various preset makeup parameters, the image generation model gains the ability to process portrait images according to different preset makeup parameters to generate portrait images with the makeup corresponding to those preset parameters.

[0045] In some embodiments, the image generation model is trained using multiple sample portrait images corresponding to a first user account under various preset makeup parameters. The multiple sample portrait images are obtained by processing the portrait images taken by the first user account according to each of the various preset makeup parameters. The portrait images taken by the first user account can be portrait images taken by the first user account through the client.

[0046] It should be noted that the portrait image captured for the first user account can be the same as or different from the original portrait image for the same user account. If they are the same, the computer device can first generate a sample portrait image based on the captured image to train the image generation model, and then use the image generation model to generate the target portrait image based on the captured image. This is the case when the first user account uses the image generation function for the first time. If they are different, the computer device can first generate a sample portrait image based on the captured image to train the image generation model, and then use the image generation model to generate the target portrait image based on the original portrait image. This is the case when the first user account uses the image generation function again after its initial use.

[0047] Optionally, by inputting target makeup parameters, target generation style, and the original portrait image into an image generation model, the computer device can add makeup to the original portrait image based on the target makeup parameters, and process the original portrait image to conform to the target generation style, thereby generating a target portrait image corresponding to the original portrait image. The target portrait image is a portrait image that possesses the makeup corresponding to the target makeup parameters and conforms to the target generation style.

[0048] It should be noted that the method provided in this application embodiment can be applied to a server or a client in a terminal. When the method is applied to a server, the image generation instruction is triggered by the client and sent to the server. When a target portrait image is generated, the server sends the generated target portrait image to the client so that the client can display the generated target portrait image. When the method is applied to a client, the image generation instruction is triggered by the client itself. When a target portrait image is generated, the client can display the generated target portrait image.

[0049] In summary, the method provided in this embodiment generates a target portrait image by using an image generation model to process the original portrait image according to the target makeup parameters and the target generation style. Since the target makeup parameters match the target generation style, and the image generation model is trained on portrait images with various preset makeup parameters, it has the ability to generate portrait images with multiple preset makeup parameters. Therefore, it can ensure that the makeup corresponding to the generated target portrait image matches the target generation style, avoiding the problem of mismatch between the makeup of the generated portrait image and the selected generation style.

[0050] Figure 4 This is a schematic flowchart illustrating an exemplary embodiment of an image generation method provided in this application. The method can be used in a computer device, such as one used for... Figure 1 The server shown. (As shown) Figure 4 As shown, the method includes: Step 402: Receive image generation instructions.

[0051] The image generation instruction includes an original portrait image and a target generation style. The original portrait image is the image of the person to be processed, and it includes at least one of the following: a facial image, an upper body image, and a full-body image. The target generation style indicates the style of the image to be generated from the original portrait image. Optionally, the target generation style indicates at least one of the following: a background style, a clothing style, and a styling style corresponding to the image generated from the original portrait image.

[0052] In some embodiments, the image generation instruction is triggered by a first user account, for example, by the first user account within the client. The original portrait image is a portrait image provided by the first user account, which can be uploaded by the first user account within the client or taken through the client, for example, a portrait image corresponding to the user using the first user account. The target generation style is the style selected by the first user account from a variety of preset generation styles provided by the client.

[0053] Step 404: Determine the target makeup parameters based on the similarity between the target generated style and different preset makeup parameters among various preset makeup parameters.

[0054] Preset makeup parameters are pre-defined makeup parameters, such as manually set makeup parameters. Each preset makeup parameter is different. Preset makeup parameters are used to indicate the processing method for adding makeup to a portrait image, for example, to indicate the processing method for adding makeup to the facial area of ​​a portrait image. Different preset makeup parameters indicate different processing methods for adding makeup to a portrait image. In some embodiments, preset makeup parameters include parameters corresponding to makeup under different processing dimensions.

[0055] The similarity between the target generated style and different preset makeup parameters among various preset makeup parameters is used to reflect the degree of similarity between the target generated style and different preset makeup parameters. The computer device determines the preset makeup parameter with the highest similarity as the target makeup parameter.

[0056] In some embodiments, the computer device converts the target generated style into a first feature vector and the preset makeup parameters into a second feature vector. Optionally, the computer device maps the target generated style to a vector space to obtain the first feature vector and maps the preset makeup parameters to a vector space to obtain the second feature vector. For example, the computer device uses an embedding model to convert the target generated style into the first feature vector and the preset makeup parameters into the second feature vector. The embedding model is used to convert high-dimensional, unstructured data into a low-dimensional vector representation. The computer device then determines the similarity between the first feature vector and different second feature vectors. For example, the computer device determines the similarity between the first feature vector and different second feature vectors based on the cosine distance between the first feature vector and different second feature vectors. The computer device then determines the preset makeup parameters corresponding to the second feature vector with the highest similarity as the target makeup parameters.

[0057] Step 406: Determine the preset makeup parameters corresponding to the target generation style in the preset mapping relationship as the target makeup parameters.

[0058] The preset mapping relationship is a pre-set mapping relationship, such as a manually set mapping relationship. The preset mapping relationship includes the correspondence between different generation styles and different preset makeup parameters, such as the correspondence between different preset generation styles and different preset makeup parameters among multiple preset generation styles on the client side. When the target generation style is obtained, the computer device uses the target generation style to query the preset mapping relationship to obtain the corresponding preset makeup parameters under the preset mapping relationship, which are then used as the target makeup parameters.

[0059] It should be noted that steps 404 and 406 are parallel steps. When implementing the method provided in the embodiments of this application, one of steps 404 and 406 can be selected for execution.

[0060] Step 408: Process the original portrait image using an image generation model based on the target makeup parameters and the target generation style to generate the target portrait image corresponding to the original portrait image.

[0061] Image generation models include AI models that support the generation of images that match input information. For example, the image generation models in this application include at least one of the following: a Stable Diffusion model (a text-to-image generation model), a LoRA-based fine-tuned Stable Diffusion model, and a VLM.

[0062] For example, Figure 5 This is a schematic diagram of the structure of a machine learning model provided in an exemplary embodiment of this application. For example... Figure 5 As shown, the machine learning model 501 includes an encoding network 502 and a decoding network 503. By inputting input information into the encoding network 502, the feature extraction result of the encoding network 502 on the input information can be obtained, i.e., the encoded information. By inputting the encoded information into the decoding network 503, the output information of the decoding network 503 can be obtained, i.e., the predicted information that matches the input information. The encoding network 502 and decoding network 503 are N-layer structures. The encoding network 502 is a cascaded structure of N encoders, and the decoding network 503 is a cascaded structure of N decoders. The structure of each layer in the encoding network 502 is consistent, and the structure of each layer in the decoding network 503 is similar to that in the decoding network 503.

[0063] Continue to refer to Figure 5 Each layer (encoder) of the 502 encoding network typically includes a multi-head self-attention module, i.e. Figure 5 The left-hand encoder structure contains "self-attention" and a fully connected feedforward network (also known as a feedforward network, FFN)... Figure 5 The "feedforward full connection" in the encoder structure on the left side of the middle.

[0064] Continue to refer to Figure 5 Each layer (decoder) of the 503 decoding network typically includes a Mask Multi-Head Self-Attention Module (which can be considered a type of multi-head self-attention module). Figure 5The "self-attention" module at the bottom of the decoder structure on the right side. This self-attention module of a cross encoder and decoder (also known as a cross self-attention module, which can be considered a multi-head self-attention module) is... Figure 5 The self-attention mechanism in the middle of the right-hand decoder structure, along with a feedforward fully connected module. Figure 5 The "feedforward fully connected" at the top of the decoder structure on the right side.

[0065] In this network, the multi-head self-attention module of the encoding network 502 is used to obtain the weight relationship of each word in the input text relative to other words in the input text. The feedforward fully connected module of the encoding network 502 is used to perform non-linear transformation on the input features. The masked multi-head self-attention module of the decoding network 503 has a similar function to the multi-head self-attention module of the encoding network 502, except that it is also used to prevent the decoding network 503 from obtaining the prediction results corresponding to the words following that word in the input text when generating a prediction result that matches a certain word in the input text (during training, this is done). Figure 5 The lower position in the decoder structure on the right side represents the prediction result corresponding to the input text. The cross-self-attention module of the decoding network 503 functions similarly to the multi-head self-attention module of the encoding network 502, the difference being that its input consists of the output information of the previous module in the decoding network 503 and the output information of the last layer in the encoding network 502. The feedforward fully connected module of the decoding network 503 functions similarly to the feedforward fully connected module of the encoding network 502.

[0066] In addition, continue to refer to Figure 5 Each of the modules (multi-head self-attention module, feedforward fully connected module) in the encoding network 502 and decoding network 503 of the machine learning model 501 is equipped with a residual connection and a layer normalization (LayerNorm) layer (i.e. Figure 5 The model employs residual and normalization (Add & Norm) layers. Residual connections can be viewed as structures that allow the output of one module to serve as the input to a subsequent, non-adjacent module, reducing model complexity and preventing gradient vanishing. Normalization layers are used to normalize the input information, such as through standardization. Both residual connections and normalization layers stabilize the model's training. During the training of Machine Learning Model 501, the computer acquires a large amount of textual data for pre-training, enabling Model 501 to generalize well to texts from various domains.

[0067] In some embodiments, the image generation model in this application can be implemented based on the structure of the machine learning model 501 described above.

[0068] The image generation model is trained using multiple sample portrait images under various preset makeup parameters. For example, it can be trained on an open-source or pre-trained model using these sample portrait images. The multiple sample portrait images include portrait images corresponding to each of the preset makeup parameters. Each sample portrait image under a preset makeup parameter is a portrait image with the makeup corresponding to that parameter. These multiple sample portrait images are obtained by processing photographed portrait images according to each of the preset makeup parameters. The photographed portrait images can be taken directly, such as portrait images taken for a first user account. By training the model using multiple sample portrait images under various preset makeup parameters, the image generation model gains the ability to process portrait images according to different preset makeup parameters to generate portrait images with the makeup corresponding to those preset parameters.

[0069] In some embodiments, the image generation model is trained using multiple sample portrait images corresponding to a first user account under various preset makeup parameters. The multiple sample portrait images are obtained by processing the portrait images taken by the first user account according to each of the various preset makeup parameters. The portrait images taken by the first user account can be portrait images taken by the first user account through the client.

[0070] Optionally, by inputting target makeup parameters, target generation style, and the original portrait image into an image generation model, the computer device can add makeup to the original portrait image based on the target makeup parameters, and process the original portrait image to conform to the target generation style, thereby generating a target portrait image corresponding to the original portrait image. The target portrait image is a portrait image that possesses the makeup corresponding to the target makeup parameters and conforms to the target generation style.

[0071] Optionally, the generation instructions also include generation prompts corresponding to the target generation style. These prompts can be text prompts indicating the requirements of the image generation model for generating the target portrait image. These text prompts may or may not be associated with the target generation style. In this case, the computer device processes the original portrait image using the image generation model based on the target makeup parameters, the target generation style, and the generation prompts, generating the target portrait image corresponding to the original portrait image.

[0072] In some embodiments, the computer device uses a target generation style as a keyword to retrieve portrait images that match that keyword. For example, the computer device uses the target generation style as a keyword to retrieve images through a search website, and identifies the top N images (where N is a positive integer) as the target portrait images. During the generation of the target portrait image, the computer device uses an image generation model to process the original portrait image based on the target makeup parameters, the target generation style, and the retrieved portrait image, generating the target portrait image corresponding to the original portrait image. For example, the computer device inputs the target makeup parameters, the target generation style, the retrieved portrait image, and the original portrait image into the image generation model to generate the target portrait image corresponding to the original portrait image. During the generation of the target portrait image, the image generation model references the retrieved portrait image when processing the original portrait image.

[0073] Optionally, the retrieved portrait image includes multiple images. During the process of generating a target portrait image corresponding to the original portrait image by processing the original portrait image using an image generation model based on target makeup parameters, target generation style, and the retrieved portrait image, the computer device determines the similarity between the face regions in the original portrait image and the face regions in different retrieved portrait images, and determines the retrieved portrait image with the highest similarity as the reference portrait image. Optionally, the computer device extracts the face regions from the original portrait image and the face regions in the retrieved portrait images using a face recognition model. The face recognition model is built based on a Convolutional Neural Network (CNN) and is trained using sample face images labeled with face regions. After obtaining the face regions in the original portrait image and the face regions in the retrieved portrait images, the computer device extracts a first feature from the face region in the original portrait image and a second feature from the face region in the retrieved portrait image. Then, it calculates the similarity between the first and second features to obtain the similarity between the face regions in the original portrait image and the face regions in different retrieved portrait images. After obtaining a reference portrait image, the computer device uses an image generation model to process the original portrait image based on the target makeup parameters, the target generation style, and the reference portrait image, generating a target portrait image corresponding to the original portrait image. For example, the computer device inputs the target makeup parameters, the target generation style, the reference portrait image, and the original portrait image into the image generation model to generate a target portrait image corresponding to the original portrait image. During the generation of the target portrait image, the image generation model processes the original portrait image with reference to the reference portrait image.

[0074] In some embodiments, the computer device acquires historical interaction images corresponding to a first user account. These historical interaction images include images from past interactions with the first user account, such as at least one of images liked, saved, downloaded, and commented on by the first user account. The computer device then determines the matching degree between the historical interaction images and the target generation style, and identifies the historical interaction image with the highest matching degree as the reference image. Optionally, the computer device extracts features from the historical interaction images, maps the target generation style into a vector, and calculates the similarity between the features of the historical interaction images and the vector of the target generation style to obtain the matching degree between the historical interaction images and the target generation style. After obtaining the reference image, the computer device processes the original portrait image using an image generation model based on the target makeup parameters, the target generation style, and the reference image to generate a target portrait image corresponding to the original portrait image. For example, the computer device inputs the target makeup parameters, the target generation style, the reference image, and the original portrait image into the image generation model to generate the target portrait image corresponding to the original portrait image. During the generation of the target portrait image, the image generation model processes the original portrait image with reference to the reference image.

[0075] In some embodiments, the image generation model is trained using multiple sample portrait images with various preset makeup parameters and parameter identifiers corresponding to each preset makeup parameter. Specifically, during the training process using sample portrait images corresponding to preset makeup parameters, the computer device inputs the parameter identifiers corresponding to the preset makeup parameters of the sample portrait images into the image generation model. Different preset makeup parameters correspond to different parameter identifiers, which can be the identity document (ID) of the preset makeup parameters. The computer device determines the target parameter identifier corresponding to the target generation style from among the parameter identifiers corresponding to the various preset makeup parameters; the correspondence between different generation styles and different parameter identifiers can be pre-set. Then, the computer device processes the original portrait image using the image generation model based on the target parameter identifier and the target generation style to generate the target portrait image corresponding to the original portrait image. During the generation of the target portrait image, the image generation model processes the original portrait image according to the processing method corresponding to the learned target parameter identifier.

[0076] In summary, the method provided in this embodiment generates a target portrait image by using an image generation model to process the original portrait image according to the target makeup parameters and the target generation style. Since the target makeup parameters match the target generation style, and the image generation model is trained on portrait images with various preset makeup parameters, it has the ability to generate portrait images with multiple preset makeup parameters. Therefore, it can ensure that the makeup corresponding to the generated target portrait image matches the target generation style, avoiding the problem of mismatch between the makeup of the generated portrait image and the selected generation style.

[0077] The method provided in this embodiment further determines the target makeup parameters based on the similarity between the target generation style and different preset makeup parameters in the preset makeup parameters. This enables the generation of target portrait images using makeup parameters that match the target generation style, avoiding the problem of mismatch between the makeup of the generated portrait image and the selected generation style. By determining the target makeup parameters based on the target generation style and preset mapping relationships, it is possible to generate target portrait images according to preset rules based on the target generation style, avoiding the problem of mismatch between the makeup of the generated portrait image and the selected generation style. By retrieving portrait images using the target generation style and generating target portrait images based on the retrieved portrait images, it is possible to ensure that the makeup corresponding to the generated target portrait image matches the target generation style, and it also enriches the diversity of the generated target portrait images.

[0078] Figure 6 This is a schematic flowchart illustrating a model training method provided in an exemplary embodiment of this application. The method can be used with a computer device, such as one used for… Figure 1 The server shown. (As shown) Figure 3 As shown, the method includes: Step 602: Obtain the portrait image taken.

[0079] The portrait image can be a photographed image, such as a portrait image taken for the first user account. The portrait image includes at least one of the following: a view of the person's face, a view of the person's upper body, and a view of the person's full body. It should be noted that the person in the portrait image can be a real person or a virtual person; this application embodiment does not impose any limitations on this.

[0080] Step 604: Process the portrait images according to each preset makeup parameter among a variety of preset makeup parameters to obtain multiple sample portrait images.

[0081] Preset makeup parameters are pre-defined makeup settings, such as manually set makeup parameters. Each preset makeup parameter is different. Preset makeup parameters are used to indicate how makeup is applied to portrait images. For example, they indicate how makeup is applied to the facial area of ​​a portrait image. Different preset makeup parameters indicate different ways of applying makeup to a portrait image.

[0082] For example, the preset makeup parameters include five preset makeup parameters, and the portrait images include eight portrait images taken by the first user account through the client. The computer device processes each of the eight portrait images using one of the five preset makeup parameters, resulting in 5*8=40 sample portrait images. Optionally, during the process of the first user account taking portrait images through the client, a shooting preview will be displayed in the client. The portrait in the shooting preview can be a beautified portrait, while the portrait in the acquired portrait images is the original image.

[0083] Step 606: Train the image generation model using multiple sample portrait images.

[0084] By training the model with multiple sample portrait images under various preset makeup parameters, the image generation model can be equipped with the ability to process portrait images according to different preset makeup parameters in order to generate portrait images with makeup corresponding to the preset makeup parameters.

[0085] In some embodiments, the image generation model is trained using multiple sample portrait images corresponding to a first user account under various preset makeup parameters. The multiple sample portrait images are obtained by processing the portrait images taken by the first user account according to each of the various preset makeup parameters. The portrait images taken by the first user account can be portrait images taken by the first user account through the client.

[0086] The image generation model processes the original portrait image based on the target makeup parameters and the target generation style to generate a target portrait image corresponding to the original portrait image. The target makeup parameters are determined from a variety of preset makeup parameters based on the target generation style. The original portrait image and the target generation style are image generation instructions. For a description of the image generation model and the process of generating the target portrait image, please refer to the relevant content above; this embodiment will not elaborate further.

[0087] In summary, the method provided in this embodiment generates a target portrait image by using an image generation model to process the original portrait image according to the target makeup parameters and the target generation style. Since the target makeup parameters match the target generation style, and the image generation model is trained on portrait images with various preset makeup parameters, it has the ability to generate portrait images with multiple preset makeup parameters. Therefore, it can ensure that the makeup corresponding to the generated target portrait image matches the target generation style, avoiding the problem of mismatch between the makeup of the generated portrait image and the selected generation style.

[0088] Taking the method provided in this application embodiment for generating portraits as an example, this application embodiment provides an AI portrait generation method and system based on makeup enhancement. The main contents include: first, uploading the original high-fidelity image; then, performing controllable and diverse makeup preprocessing in the cloud; and finally, using the enhanced dataset for model training and inference. For example, Figure 7 This is a schematic diagram illustrating the model training and inference process provided in an exemplary embodiment of this application. For example... Figure 7 As shown: In steps A1-A3, data acquisition and uploading are performed. The client guides the user to take multiple photos, for example, eight raw facial images from different angles and with different expressions. To enhance user experience, the client provides a real-time basic beautification effect during the preview, but this effect is only for previewing; the final image uploaded to the cloud server is the raw image data without any beautification or compression. This separation of preview and upload is crucial to ensuring the quality of the data source.

[0089] In steps A4-A5, cloud-based structured makeup preprocessing and model training are performed. The cloud server receives the original image set uploaded by the user. The server has multiple pre-set sets of structured makeup parameters. Each parameter set represents a complete makeup style, such as "retro" or "professional," and includes several quantifiable and adjustable basic parameters, such as filter intensity, RGB values ​​of lipstick shades, eyeshadow color and intensity, and blush color. It's important to note that this cloud-based beauty preprocessing only involves simple beauty parameter configuration, such as adding different filters, blush, lipstick, and eyeshadow, as well as slight skin smoothing and brightening. Its purpose is to provide diverse and high-quality input data for the AI ​​model, rather than excessively beautifying the user's image. Its core principle is style exploration while preserving feature fidelity. That is, exploring the diversity of stylistic elements such as makeup and color while preserving the user's real facial features to the greatest extent possible. This ensures that the AI ​​model can learn a stable mapping relationship between the user's accurate facial features and various stylized elements. This allows for the generation of diverse and stylized portraits while maintaining a high degree of similarity to the user in the final result, resolving the conflict between beautification and resemblance to the real person. The server uses each set of makeup parameters in parallel to perform batch beautification processing on the same set of original images of the user. For example, using 5 sets of parameters to process 8 original images will generate 5*8=40 intermediate images. These intermediate images have the same content but different makeup styles, forming a high-quality, highly diverse training dataset. During the training phase, these generated sets of intermediate images are used as training data and input into the image generation model for training. This process explicitly teaches the AI ​​model to learn the mapping relationship between the user's facial features and various makeup styles. The training process is as follows: Figure 8 As shown.

[0090] In steps A6-A10, model inference is performed to generate a photograph and feedback is provided. The inference (generation) phase includes the following: (1) The system triggers a style-makeup matching mechanism to determine the makeup parameters that best match the user's selected target style, such as "beach wedding dress". This mechanism can be implemented in one or more of the following ways: (a) Rule mapping: The system backend maintains a mapping table that directly defines the correspondence between "target style" and "makeup parameter set ID" (e.g., target style = "beach wedding dress" corresponds to parameter set ID = "makeup_set_005"). This is the most direct and stable method.

[0091] (b) Intelligent Matching: Each set of makeup parameters and each target style tag are converted into a high-dimensional feature vector through an embedding model. By calculating the similarity (e.g., cosine similarity) between the style vector and all makeup vectors, the makeup parameter set with the highest similarity is selected as the best match. This method is more suitable for scenarios with a large number of style tags.

[0092] (c) End-to-end control: During model training, the makeup parameter set IDs used are trained together as conditional labels. During inference, after the user selects a style, the system finds the corresponding parameter set ID through mapping rules and uses this ID as a conditional input to the model to directly control the generation process.

[0093] (2) After determining the target parameter set, the system automatically calls the makeup parameters and the trained model, combines them with the text prompts of the target style, and generates the final highly stylized and harmonious portrait photo. The generation process is as follows: Figure 9 As shown.

[0094] The results are then fed back, and a set (multiple) of photos matching the user's selected style are returned to the user's terminal for display.

[0095] The system includes: a user terminal module for image acquisition, local preview beautification, uploading of original images, and display of results; a cloud processing module including a communication unit, an original image storage unit, a makeup parameter set management unit, and a parallel image processing unit; and an AI engine module including a model training unit and an image inference and generation unit.

[0096] The method provided in this application embodiment has at least the following beneficial effects: (1) By uploading the original image and performing unified and controllable high-quality beautification processing in the cloud, the highest fidelity is guaranteed from the source of the data, avoiding the cumulative damage to image quality caused by multiple processing.

[0097] (2) By applying multiple sets of structured, high-level makeup parameter sets, semantic-level data augmentation was achieved, directly providing the AI ​​model with diverse learning samples that are strongly correlated with the target style. Controllable and structured preprocessing parameters were used to precisely guide and enhance the diversity of AI model generation.

[0098] (3) By preprocessing images, abstract text descriptions are transformed into concrete and precisely controllable visual parameters (such as RGB color values ​​of lipstick shades), thereby stably and accurately guiding the model to generate the expected makeup, achieving the effect of "visual cues" being superior to "text cues".

[0099] It should be noted that this application may display prompt interfaces, pop-ups, or output voice prompts before and during the collection of user data. These prompt interfaces, pop-ups, or voice prompts are used to inform the user that their data is being collected. This ensures that the application only begins the steps for collecting user data after receiving confirmation from the user regarding the prompt interface or pop-up; otherwise (i.e., without user confirmation), the steps for collecting user data end, meaning no user data is collected. In other words, all user data collected in this application is collected with the user's consent and authorization, and the collection, use, and processing of related user data must comply with the relevant laws, regulations, and standards of the relevant countries and regions.

[0100] It should be noted that the order of the method steps provided in the embodiments of this application can be appropriately adjusted, and the steps can also be added or removed as appropriate. Any method variations that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the protection scope of this application, and therefore will not be elaborated further.

[0101] Figure 10 This is a schematic diagram of the structure of an image generation apparatus provided in an exemplary embodiment of this application. For example... Figure 10 As shown, the device includes: The receiving module 1001 is used to receive an image generation instruction, the image generation instruction including an original portrait image and a target generation style, the target generation style being used to indicate the style of the image generated based on the original portrait image; The determination module 1002 is used to determine, among a variety of preset makeup parameters, a target makeup parameter that matches the target generation style, wherein the preset makeup parameter is used to indicate the processing method for adding makeup to a portrait image; The generation module 1003 is used to process the original portrait image according to the target makeup parameters and the target generation style through an image generation model to generate a target portrait image corresponding to the original portrait image. The image generation model is trained by multiple sample portrait images under the various preset makeup parameters.

[0102] In an optional design, the determining module 1002 is used to: determine the target makeup parameters based on the similarity between the target generated style and different preset makeup parameters among the multiple preset makeup parameters.

[0103] In an optional design, the determining module 1002 is configured to: convert the target generation style into a first feature vector; convert the preset makeup parameters into a second feature vector; determine the similarity between the first feature vector and different second feature vectors; and determine the preset makeup parameters corresponding to the second feature vector with the highest similarity as the target makeup parameters.

[0104] In an optional design, the determining module 1002 is used to: determine the preset makeup parameters corresponding to the target generation style in the preset mapping relationship as the target makeup parameters; wherein, the preset mapping relationship includes the correspondence between different generation styles and different preset makeup parameters.

[0105] In an optional design, the image generation model is trained using multiple sample portrait images under the various preset makeup parameters and parameter identifiers corresponding to the various preset makeup parameters respectively; the determining module 1002 is used to: determine the target parameter identifier corresponding to the target generation style from the parameter identifiers corresponding to the various preset makeup parameters respectively. The generation module 1003 is used to: process the original portrait image according to the target parameter identifier and the target generation style through the image generation model, and generate a target portrait image corresponding to the original portrait image.

[0106] In an optional design, the generation instruction further includes generation prompts corresponding to the target generation style; the generation module 1003 is used to: process the original portrait image through the image generation model according to the target makeup parameters, the target generation style and the generation prompts, and generate a target portrait image corresponding to the original portrait image.

[0107] In an optional design, such as Figure 11 As shown, the device further includes: a retrieval module 1004, used to retrieve portrait images that match the target generation style as keywords; The generation module 1003 is used to: process the original portrait image according to the target makeup parameters, the target generation style and the retrieved portrait image through the image generation model, and generate a target portrait image corresponding to the original portrait image.

[0108] In an optional design, the determining module 1002 is configured to: determine the similarity between the face regions in the original portrait image and the face regions in different retrieved portrait images; and determine the retrieved portrait image with the highest similarity as the reference portrait image; The generation module 1003 is used to: process the original portrait image according to the target makeup parameters, the target generation style and the reference portrait image through the image generation model, and generate a target portrait image corresponding to the original portrait image.

[0109] In an optional design, such as Figure 12 As shown, the device further includes: an acquisition module 1005, used to acquire historical interaction images corresponding to the first user account, the historical interaction images including images of past interactions with the first user account; The determining module 1002 is used to: determine the matching degree between the historical interaction image and the target generation style; and determine the historical interaction image with the highest matching degree as the reference image. The generation module 1003 is used to: process the original portrait image according to the target makeup parameters, the target generation style and the reference image through the image generation model, and generate a target portrait image corresponding to the original portrait image.

[0110] Figure 13 This is a schematic diagram of the structure of a model training apparatus provided in an exemplary embodiment of this application. Figure 13 As shown, the device includes: The acquisition module 1301 is used to acquire captured portrait images; Processing module 1302 is used to process the captured portrait image according to each preset makeup parameter among a variety of preset makeup parameters to obtain multiple sample portrait images; Training module 1303 is used to train an image generation model using the multiple sample portrait images; The image generation model is used to process the original portrait image according to the target makeup parameters and the target generation style to generate the target portrait image corresponding to the original portrait image. The target makeup parameters are determined from a variety of preset makeup parameters according to the target generation style. The original portrait image and the target generation style belong to the image generation instructions.

[0111] It should be noted that the image generation device provided in the above embodiments is only an example of the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the image generation device and the image generation method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0112] Similarly, the model training device provided in the above embodiments is only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the model training device and the model training method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0113] Embodiments of this application also provide a computer device, which includes a processor and a memory. The memory stores at least one instruction, at least one program, code set, or instruction set. The processor loads and executes the at least one instruction, at least one program, code set, or instruction set to implement the image generation method or model training method provided in the above-described method embodiments.

[0114] For example, Figure 14 This is a schematic diagram of the structure of a computer device provided in an exemplary embodiment of this application.

[0115] The computer device 1400 includes a central processing unit (CPU) 1401, a system memory 1404 including random access memory (RAM) 1402 and read-only memory (ROM) 1403, and a system bus 1405 connecting the system memory 1404 and the CPU 1401. The computer device 1400 also includes a basic input / output system (I / O system) 1406 to facilitate information transfer between various components within the computer device, and a mass storage device 1407 for storing the operating system 1413, application programs 1414, and other program modules 1415.

[0116] The basic input / output system 1406 includes a display 1408 for displaying information and an input device 1409 for user input, such as a mouse or keyboard. Both the display 1408 and the input device 1409 are connected to the central processing unit 1401 via an input / output controller 1410 connected to the system bus 1405. The basic input / output system 1406 may also include the input / output controller 1410 for receiving and processing input from multiple other devices such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 1410 also provides output to a display screen, printer, or other types of output devices.

[0117] The mass storage device 1407 is connected to the central processing unit 1401 via a mass storage controller (not shown) connected to the system bus 1405. The mass storage device 1407 and its associated computer-readable storage media provide non-volatile storage for the computer device 1400. That is, the mass storage device 1407 may include computer-readable storage media (not shown) such as a hard disk or a compact disc read-only memory (CD-ROM) drive.

[0118] Without loss of generality, the computer-readable storage medium may include computer storage media and communication media. Computer storage media include volatile and non-volatile, removable and non-removable media implemented using any method or technology for storing information such as computer-readable storage instructions, data structures, program modules, or other data. Computer storage media include RAM, ROM, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other solid-state storage devices, CD-ROM, digital versatile disc (DVD) or other optical storage, magnetic tape cassettes, magnetic tape, disk storage, or other magnetic storage devices. Of course, those skilled in the art will recognize that the computer storage medium is not limited to the above-mentioned types. The system memory 1404 and mass storage device 1407 described above can be collectively referred to as memory.

[0119] The memory stores one or more programs, which are configured to be executed by one or more central processing units 1401. The one or more programs contain instructions for implementing the above method embodiments, and the central processing unit 1401 executes the one or more programs to implement the methods provided by the various method embodiments described above.

[0120] According to various embodiments of this application, the computer device 1400 can also be connected to a remote computer device on a network, such as the Internet. That is, the computer device 1400 can be connected to a network 1412 via a network interface unit 1411 connected to the system bus 1405, or the network interface unit 1411 can be used to connect to other types of networks or remote computer device systems (not shown).

[0121] The memory further includes one or more programs stored in the memory, and the one or more programs include steps performed by a computer device in the methods provided in the embodiments of this application.

[0122] This application also provides a computer-readable storage medium storing at least one instruction, at least one program, code set, or instruction set. When the at least one instruction, at least one program, code set, or instruction set is loaded and executed by the processor of a computer device, the image generation method or model training method provided in the above-described method embodiments is implemented.

[0123] This application also provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the image generation method or model training method provided in the above-described method embodiments.

[0124] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0125] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent switching, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A method for generating an image, characterized in that, The method includes: Receive an image generation instruction, the image generation instruction including an original portrait image and a target generation style, the target generation style being used to indicate the style of the image generated based on the original portrait image; Among a variety of preset makeup parameters, a target makeup parameter that matches the target generation style is determined. The preset makeup parameters are used to indicate the processing method for adding makeup to portrait images. The image generation model processes the original portrait image according to the target makeup parameters and the target generation style to generate the target portrait image corresponding to the original portrait image. The image generation model is trained on multiple sample portrait images under the various preset makeup parameters.

2. The method according to claim 1, characterized in that, The step of determining the target makeup parameters that match the target generated style from a variety of preset makeup parameters includes: The target makeup parameters are determined based on the similarity between the target generated style and different preset makeup parameters among the multiple preset makeup parameters.

3. The method according to claim 2, characterized in that, The step of determining the target makeup parameters based on the similarity between the target generated style and different preset makeup parameters among the multiple preset makeup parameters includes: The target style is converted into a first feature vector; the preset makeup parameters are converted into a second feature vector. Determine the similarity between the first feature vector and different second feature vectors; The preset makeup parameters corresponding to the second feature vector with the highest similarity are determined as the target makeup parameters.

4. The method according to claim 1, characterized in that, The step of determining the target makeup parameters that match the target generated style from a variety of preset makeup parameters includes: The preset makeup parameters corresponding to the target generation style in the preset mapping relationship are determined as the target makeup parameters; The preset mapping relationship includes the correspondence between different generation styles and different preset makeup parameters.

5. The method according to any one of claims 1 to 4, characterized in that, The image generation model is trained using multiple sample portrait images under various preset makeup parameters, and parameter labels corresponding to each preset makeup parameter; the method further includes: Among the parameter identifiers corresponding to the various preset makeup parameters, the target parameter identifier corresponding to the target generation style is determined; The image generation model processes the original portrait image according to the target parameter identifier and the target generation style to generate a target portrait image corresponding to the original portrait image.

6. The method according to any one of claims 1 to 4, characterized in that, The generation instruction also includes generation prompts corresponding to the target generation style; the step of processing the original portrait image using an image generation model based on the target makeup parameters and the target generation style to generate a target portrait image corresponding to the original portrait image includes: The image generation model processes the original portrait image according to the target makeup parameters, the target generation style, and the generation prompts to generate a target portrait image corresponding to the original portrait image.

7. The method according to any one of claims 1 to 4, characterized in that, The method further includes: Using the target generation style as keywords, retrieve portrait images that match the keywords; The step of processing the original portrait image using an image generation model based on the target makeup parameters and the target generation style to generate a target portrait image corresponding to the original portrait image includes: The image generation model processes the original portrait image according to the target makeup parameters, the target generation style, and the retrieved portrait image to generate a target portrait image corresponding to the original portrait image.

8. A model training method, characterized in that, The method includes: Acquire portrait images; The captured portrait image is processed according to each of the various preset makeup parameters to obtain multiple sample portrait images; The image generation model was trained using the aforementioned multiple sample portrait images; The image generation model is used to process the original portrait image according to the target makeup parameters and the target generation style to generate the target portrait image corresponding to the original portrait image. The target makeup parameters are determined from a variety of preset makeup parameters according to the target generation style. The original portrait image and the target generation style belong to the image generation instructions.

9. An image generation device, characterized in that, The device includes: A receiving module is used to receive an image generation instruction, the image generation instruction including an original portrait image and a target generation style, the target generation style being used to indicate the style of the image generated based on the original portrait image; The determination module is used to determine the target makeup parameters that match the target generation style from a variety of preset makeup parameters. The preset makeup parameters are used to indicate the processing method for adding makeup to portrait images. The generation module is used to process the original portrait image according to the target makeup parameters and the target generation style through an image generation model to generate a target portrait image corresponding to the original portrait image. The image generation model is trained on multiple sample portrait images under the various preset makeup parameters.

10. A model training device, characterized in that, The device includes: The acquisition module is used to acquire captured portrait images; The processing module is used to process the captured portrait image according to each of the multiple preset makeup parameters to obtain multiple sample portrait images; The training module is used to train an image generation model using the multiple sample portrait images; The image generation model is used to process the original portrait image according to the target makeup parameters and the target generation style to generate the target portrait image corresponding to the original portrait image. The target makeup parameters are determined from a variety of preset makeup parameters according to the target generation style. The original portrait image and the target generation style belong to the image generation instructions.

11. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing at least one program, the at least one program being loaded and executed by the processor to implement the image generation method as described in any one of claims 1 to 7, or the model training method as described in claim 8.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one program, which is loaded and executed by a processor to implement the image generation method as described in any one of claims 1 to 7, or the model training method as described in claim 8.

13. A computer program product, characterized in that, The computer program product includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the image generation method as described in any one of claims 1 to 7, or the model training method as described in claim 8.