Method and system for generating cartoon portrait picture based on LDM sketch, and storage medium
Through the LDM-based method, the modular sketch encoder and hidden diffusion model are trained, which solves the problem of poor quality when generating anime portrait pictures in the existing technology, and achieves high-quality and accurate anime portrait pictures.
Patent Information
- Application Number
- CN202510460210.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing methods and systems for generating pictures of sketches have problems with poor image generation quality when generating cartoon portrait pictures, including shape deviation, color mismatch, texture distortion and loss of details, and lack of sufficient contextual understanding and error correction capabilities, resulting in the generated pictures being unable to accurately reflect the intention and characteristics of the sketch.
Using an LDM-based method, the sketch encoder and hidden diffusion model are trained by obtaining the data set of sketches and animation portrait images. The sketch encoder adopts a modular design to handle facial features such as face profile, mouth, nose, eyes, etc. The hidden diffusion model generates high-quality anime portrait images by gradually adding and removing noise.
High-quality animation portrait image generation is realized, the accuracy and credibility of image generation is improved, the usability and credibility of the system are enhanced, and the problem of poor image generation quality in the prior art is solved.
Smart Images

Figure CN119991911A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence, and specifically relates to a method, system and storage medium for generating cartoon portrait images based on LDM sketches. Background Art
[0002] Sketching is one of the most natural and flexible ways for humans to express and convey visual information. In the field of modern human-computer interaction, especially with the popularity of touch-screen devices such as smartphones and tablets, hand-drawn sketches of simple lines to express users' visual needs have become an important means of interaction. The high degree of abstraction of sketches makes them a key tool for many art designs, user interface designs, and creative expressions. However, the abstract and diverse characteristics of sketches also bring great challenges to the development of related technologies.
[0003] In recent years, with the rapid development of deep learning technology, research and applications related to sketches have made great progress. Sketch-to-image translation, as an important image generation task, has received widespread attention. Through sketch-to-image translation technology, users can draw or upload sketches and generate high-quality images that are consistent with the shape and semantics of the sketches. This technology has wide application value in character design, virtual creation, and entertainment.
[0004] However, sketch-to-image translation still faces many challenges, especially the task of generating sketches into cartoon portraits. When training models, a large number of annotated datasets are required. Currently, there are large-scale sketch-to-image pairing datasets, but the datasets are mainly about animals, plants, and daily necessities. There is currently no large-scale sketch-to-cartoon portrait pairing dataset.
[0005] In addition, the current methods and systems for generating pictures from sketches face many challenges in terms of picture generation quality. These problems not only affect the user experience, but also limit the further development and application of related technologies. In addition to the possible deviation in shape between the input sketch and the generated picture, the existing sketch-generated picture system often encounters problems such as color mismatch, texture distortion, and loss of details. These problems often stem from the algorithm's insufficient understanding or improper processing of the sketch information, resulting in the generated picture being unable to accurately reflect the intention and characteristics of the sketch; more seriously, sometimes the system may even generate pictures that are completely inconsistent with the sketch category. This situation usually occurs when the sketch information is too abstract or vague, and the algorithm lacks sufficient contextual understanding and error correction capabilities, which not only violates the original intention of generating pictures from sketches, but also greatly reduces the availability and credibility of the system. Therefore, the present invention provides a method, system, and storage medium for generating anime portrait pictures from sketches based on LDM. Summary of the invention
[0006] The present invention aims to solve at least one of the technical problems existing in the prior art; to this end, the present invention proposes a method, system and storage medium for generating cartoon portrait images from sketches based on LDM, which are used to solve the technical problem of poor image generation quality in the existing methods and systems for generating images from sketches.
[0007] To achieve the above object, a first aspect of the present invention provides a method for generating an animated portrait image based on a sketch of an LDM, comprising the following steps: Obtain several sketches and the corresponding cartoon portrait images of the sketches to construct a data set; wherein, data enhancement is performed on the sketches; the data set includes a training set and a test set; Input the sketches in the training set into the sketch encoder to generate a reconstructed sketch, calculate a loss function value 1 between the sketches in the training set and the corresponding reconstructed sketches, compare the loss function value 1 with a preset loss threshold 1, and obtain a trained sketch encoder; The sketches in the training set are data enhanced, and the data-enhanced sketches are input into the trained sketch encoder to obtain sketch features; based on the sketch features and the latent diffusion model, an anime portrait image is generated; the loss function value 2 between the generated anime portrait image and the real anime portrait image is calculated; the loss function value 2 is compared with the preset loss threshold 2 to obtain the image generation model.
[0008] Among them, the sketch features include face contour features, mouth features, nose features, and eye features.
[0009] Preferably, the preprocessing of the sketch includes: Select clear sketches and cartoon portraits to build training and test sets; and perform data augmentation on each sketch.
[0010] Preferably, the sketch encoder includes a head and face contour encoder, a mouth encoder, a nose encoder, a left eye encoder and a right eye encoder.
[0011] The sketch encoder of the present invention adopts a modular design, in which different facial features such as head / face contour, mouth, nose, and eyes are processed by independent encoders respectively. This design makes the system highly flexible and scalable, and each encoder focuses on capturing the fine details of its corresponding facial features, thereby improving the recognition precision and accuracy of the entire system. When the encoding algorithm of a certain facial feature needs to be updated or optimized, only the corresponding encoder needs to be modified without reconstructing the entire system, which greatly reduces the maintenance cost and update difficulty of the system.
[0012] Preferably, calculating the loss function value between the sketch in the training set and the corresponding reconstructed sketch comprises: ; in, is the loss function value 1, is the set of sketch encoders, i is the sketch encoder number, i=0,1,…,5, is the sketch encoder i, and x is the sketch corresponding to the sketch encoder i.
[0013] Preferably, the loss function value between the calculated generated cartoon portrait picture and the real cartoon portrait picture comprises: Extract sketch features, and calculate the loss function value 2 based on the implicit diffusion model, as follows: ; in, is the loss function value 2, E is the expected value, t is the time step of the current iteration, t [1,T], T is the total time steps of the entire diffusion process; is a pre-trained encoder; is the real noise; is the noise predicted by the model; is the coded sketch; For the decoder; A collection of sketch features.
[0014] Preferably, comparing the loss function value 1 with a preset loss threshold 1 includes: Determine whether the loss function value 1 is less than a preset loss threshold 1; if yes, mark the training result of the sketch encoding as completed; if no, mark the training result of the sketch encoding as not completed; The comparing the loss function value 2 with the preset loss threshold 2 includes: It is determined whether the loss function value 2 is less than the preset loss threshold 2; if so, the invisible diffusion model is marked as an image generation model; if not, the invisible diffusion model is not marked.
[0015] Preferably, the loss function of the implicit diffusion model includes: .
[0016] Preferably, the implicit diffusion model includes a forward diffusion model and a backward diffusion model; The forward diffusion model is as follows: ; in, is a multivariate normal distribution; The noise scheduling parameters , is the noise variance at step t, is the initial image data, is the image data after the t-th noise, , is the random noise of standard normal distribution; is the identity matrix; The back diffusion model is as follows: ; in, is the image data after the t-1th noise, is the mean of the Gaussian distribution, is the variance of the Gaussian distribution.
[0017] Preferably, the second aspect of the present invention provides a system for generating cartoon portrait images from sketches based on LDM, comprising an acquisition module and a model training module; Acquisition module: used to obtain several sketches and the corresponding anime portrait images to build a data set; Model training module: used to input the sketches in the training set into the sketch encoder, generate a reconstructed sketch, calculate the loss function value 1 between the sketches in the training set and the corresponding reconstructed sketch, compare the loss function value 1 with a preset loss threshold 1, and obtain a trained sketch encoder; The sketches in the training set are data enhanced, and the data-enhanced sketches are input into the trained sketch encoder to obtain sketch features; based on the sketch features and the latent diffusion model, an anime portrait image is generated; the loss function value 2 between the generated anime portrait image and the real anime portrait image is calculated; the loss function value 2 is compared with the preset loss threshold 2 to obtain the image generation model.
[0018] Preferably, the third aspect of the present invention provides a storage medium, wherein the storage medium stores a computer program, and when the computer program is executed, the method described in the first aspect of the present invention runs.
[0019] Compared with the prior art, the present invention has the following beneficial effects: In the first stage, the present invention focuses on the training of the sketch encoder, which contains a series of sketches and their corresponding cartoon portrait pictures. The sketch encoder is trained to understand and capture the key features in the sketch. In this process, the sketch is input into the encoder to generate a reconstructed sketch, and the difference between the reconstructed sketch and the original sketch is evaluated by calculating the loss function value one, thereby continuously optimizing the performance of the encoder. After this stage of training, the sketch encoder can accurately and efficiently extract and reconstruct sketch features; in the second stage, the hidden diffusion model is introduced to drive the generation of cartoon portrait pictures based on the sketch features generated by the sketch encoder obtained by the first stage of training. In this stage, in order to further improve the generalization ability and robustness of the model, the sketches in the training set are firstly subjected to data enhancement processing, and the enhanced sketches are input into the trained sketch encoder to generate sketch features. These codes contain the key information required to generate cartoon portrait pictures. Subsequently, the hidden diffusion model uses these sketch features to gradually generate high-quality cartoon portrait pictures by learning and simulating the process of adding and removing noise. At the same time, the difference between the generated image and the target anime portrait image is evaluated by calculating the loss function value 2, and the performance of the latent diffusion model is further optimized; efficient and accurate sketch feature extraction, enhanced model generalization ability and robustness, and the generation of high-quality anime portrait images are achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0021] Figure 1 This is a schematic diagram of a method for generating anime portrait images based on sketches according to an embodiment of the present application; Figure 2 This is a flowchart of a method for generating anime portrait images based on sketches according to an embodiment of the present application; Figure 3 A network structure diagram of a sketch-based cartoon portrait image generation method according to an embodiment of the present application; Figure 4 This is a structural diagram of the sketch encoding stage of the sketch-based cartoon portrait image generation method according to an embodiment of the present application. DETAILED DESCRIPTION
[0022] The technical solution of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0023] Sketch-based anime portrait generation technology based on LDM (Latent Diffusion Models) is a method that uses deep learning and diffusion models to transform simple sketches into high-quality anime-style portraits.
[0024] Diffusion Models: Diffusion models are a type of generative model that works by gradually adding noise to the data to destroy the data structure, and then training a neural network to reverse this process to recover the original data from the noise.
[0025] See also Figure 1-Figure 2 The first embodiment of the present invention provides a method for generating an animated portrait picture based on a sketch of an LDM, comprising the following steps: Obtain several sketches and the corresponding cartoon portrait images of the sketches to construct a data set; wherein, data enhancement is performed on the sketches; the data set includes a training set and a test set; The sketch includes a basic line outline specifying the general face shape, hair length and posture of the anime character, and has sufficient information to indicate that the sketch is a portrait of an anime character; And, preprocess the data set as follows: Since the collected data sets contain samples that do not meet the requirements, such as pictures containing multiple cartoon characters, the sample data is first preprocessed, and clear sketches and color pictures are obtained by screening to construct training and test sets; Therefore, we perform data augmentation on the dataset, processing each sketch into several sketches with different resolutions, which can enhance the robustness of arbitrary abstract sketch inputs; because the sketches in the dataset are only a few basic lines that outline the general outline of the object, while the anime portraits are high-definition color pictures. Through experiments, we can observe that the abstract level of the sketch edge depends on the image resolution, so in order to enhance the robustness of arbitrary abstract sketch inputs, we use data augmentation methods; For example, if the resolution of the cartoon portrait images in the sample is 512×512, the sketches in the dataset are enhanced. For example, a 512×512 original sketch is processed to generate two other sketches with resolutions of 128×128 and 256×256, which are also used as the data of the dataset to expand the sketch dataset.
[0026] Since directly jointly training the sketch encoder to map the sketch domain to the image domain may result in a non-smooth distribution, such as for a pair of mirrored sketches, the expectation is to generate a pair of mirrored anime portraits with the same identity, but the result is to generate two images with large identity differences. Therefore, the present invention adopts a two-stage training strategy, in the first stage, the sketch encoder is trained separately, and in the second stage, the entire network is trained. Please refer to Figure 3 ; Through two-stage training, sketches can be generated into anime portrait images more accurately, and two-stage training can reduce the computational cost required for training, thereby promoting model optimization.
[0027] In order to further improve the robustness of the model to different sketches, a random mask conditional training strategy is adopted, which randomly masks part of the sketch to train the generative network part.
[0028] The first stage of training process is as follows: the sketches in the training set are input into the sketch encoder to generate the reconstructed sketches, and the loss function value between the sketches in the training set and the corresponding reconstructed sketches is calculated; Determine whether the loss function value 1 is less than the preset loss threshold 1; if yes, stop training and mark the training result of the sketch encoding as completed; if no, mark the training result of the sketch encoding as not completed; Among them, the sketch encoder is composed of a head and face contour encoder, a mouth encoder, a nose encoder, a left eye encoder and a right eye encoder; this can better extract the features of different structures of the sketch to guide the diffusion generation process of anime portrait images.
[0029] In the second stage, the dataset will be augmented and a portion of the input sketch will be randomly masked to train the generation process.
[0030] The specific training process of the second stage is as follows: data augmentation is performed on the sketches in the training set, and the data augmented sketches are input into the trained sketch encoder to obtain sketch features; based on the sketch features and the latent diffusion model, anime portrait images are generated; the loss function value 2 between the generated anime portrait image and the real anime portrait image is calculated; it is determined whether the loss function value 2 is less than the preset loss threshold 2; if yes, the training is stopped and the latent diffusion model is marked as the image generation model; otherwise, the latent diffusion model is not marked.
[0031] The key idea of the implicit diffusion model is to randomly add noise to the initial image data and distribute the data according to the Markov chain to simulate the diffusion in non-equilibrium thermodynamics; The forward diffusion process gradually adds noise to the initial image data to obtain the noisy image data, usually by adding Gaussian noise. Assuming the time step is t, the forward process is defined as a Markov chain: ; By recursive application, the above formula simplifies to: ; in, is a multivariate normal distribution; is the noise scheduling parameter, , is the noise variance at step t, t is the number of the time step, t [1,T], T is the total time steps of the entire diffusion process, that is, the number of complete iterations required for the model to gradually reconstruct the real data from complete noise, is the initial image data, is the image data after the t-1th noise, is the image data after the t-th noise, , is the random noise of standard normal distribution; is the identity matrix; The back diffusion process is from Gradually recover to , achieved by learning the reverse distribution, the reverse process is also a Markov chain, as follows: ; in, is the mean of the Gaussian distribution, is the variance of the Gaussian distribution.
[0032] The loss function of the implicit diffusion model is established based on the Gaussian noise term: ; After simplification, the loss function can be written as: ; Where E is the expected value, is the real noise; is the noise predicted by the model; In order to reduce the computational cost, a latent diffusion model is used, which helps the features of perceptual details and conceptual semantic relevance to remain in the latent code after being encoded. Therefore, a pre-trained encoder is set up to convert the h, w, The encoding is a low-dimensional latent encoding; where h is the height of the image and w is the width of the image. is the number of channels, which in the present invention are the three color channels (red, green, and blue) in the RGB image.
[0033] The pre-trained decoder decodes the image from the latent encoder, and the loss function for: ; Generate anime portrait images according to the given sketch input, regard the sketch as the condition to guide the model during denoising, and the loss function of the latent diffusion model Build as: ; in, is the loss function value 2, is a pre-trained encoder, is the coded sketch; is a decoder, used to reverse the diffusion process; A collection of sketch features.
[0034] The present invention does not simply train a sketch encoder to provide conditional feature maps to process Denoising, but by training a multi-encoder network architecture to build a conditional module; the entire encoder It consists of 5 encoder groups, ; in, is the left eye encoder, is the right eye encoder, For the nose encoder, is the mouth encoder, For the head and face contour encoder, please refer to Figure 3 , Figure 4 ;in, Figure 4 The detailed structure of the sketch decoder in corresponds to that of the sketch encoder.
[0035] In the first sketch encoding phase of the two-stage training, the loss function value between the sketches in the training set and the corresponding reconstructed sketches is calculated by optimizing the sum of the MSE loss function thresholds of each partial encoder. : ; in, is the loss function value 1, is the set of sketch encoders, i is the sketch encoder number, i=0,1,…,5, is the sketch encoder i, and x is the sketch corresponding to the sketch encoder i.
[0036] It should be noted that the sketch encoder corresponds to the corresponding face area, and the set of sketch encoders includes an encoder and a corresponding decoder. The decoder is specifically used in the training phase of the multiple encoders and is not used in the training phase or the overall reasoning phase of the conditional latent diffusion model network.
[0037] A second aspect of the present invention provides a system for generating cartoon portrait images from sketches based on LDM, comprising an acquisition module and a model training module; Acquisition module: used to obtain several sketches and the corresponding anime portrait images to build a data set; Model training module: used for inputting the sketches in the training set into the sketch encoder, generating a reconstructed sketch, calculating the loss function value 1 between the sketches in the training set and the corresponding reconstructed sketches, comparing the loss function value 1 with a preset loss threshold 1, and obtaining a trained sketch encoder; The sketches in the training set are data enhanced, and the data-enhanced sketches are input into the trained sketch encoder to obtain sketch features; based on the sketch features and the latent diffusion model, an anime portrait image is generated; the loss function value 2 between the generated anime portrait image and the real anime portrait image is calculated; the loss function value 2 is compared with the preset loss threshold 2 to obtain the image generation model.
[0038] The third aspect of the present invention provides a storage medium storing a computer program. When the computer program is executed, it runs the method described in the first aspect of the embodiment above.
[0039] This embodiment also provides an electronic device, including: one or more processors, one or more memories, and one or more computer programs; wherein the processor is connected to the memory, and the one or more computer programs are stored in the memory. When the electronic device is running, the processor executes the one or more computer programs stored in the memory so that the electronic device executes the method described in the above embodiment one.
[0040] Part of the data in the above formula is calculated by removing the dimension and taking its numerical value. The formula is a formula closest to the actual situation obtained by software simulation of a large amount of collected data; the preset parameters and preset thresholds in the formula are set by technical personnel in this field according to actual conditions or obtained through simulation of a large amount of data.
[0041] The working principle of the present invention is as follows: a plurality of sketches and cartoon portrait images corresponding to the sketches are obtained to construct a data set; the sketches in the training set are input into a sketch encoder to generate a reconstructed sketch, and a loss function value 1 between the sketches in the training set and the corresponding reconstructed sketch is calculated, and the loss function value 1 is compared with a preset loss threshold 1 to obtain a trained sketch encoder; Data enhancement is performed on the sketches in the training set, and the data-enhanced sketches are input into the trained sketch encoder to obtain sketch features; based on the sketch features and the latent diffusion model, an anime portrait image is generated; a loss function value 2 between the generated anime portrait image and the real anime portrait image is calculated; the loss function value 2 is compared with a preset loss threshold 2 to obtain an image generation model; Input the sketch into the image generation model and output the anime portrait image corresponding to the sketch.
[0042] The above embodiments are only used to illustrate the technical method of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical method of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical method of the present invention.
Claims
1. A method for generating anime portrait images based on LDM sketches, characterized in that: The following steps are involved: Obtain a number of sketches and the corresponding cartoon portrait images of the sketches to construct a data set; wherein the sketches are preprocessed; the data set includes a training set and a test set; Input the sketches in the training set into the sketch encoder to generate a reconstructed sketch, calculate a loss function value 1 between the sketches in the training set and the corresponding reconstructed sketches, compare the loss function value 1 with a preset loss threshold 1, and obtain a trained sketch encoder; The sketches in the training set are data enhanced, and the data-enhanced sketches are input into the trained sketch encoder to obtain sketch features; based on the sketch features and the latent diffusion model, an anime portrait image is generated; the loss function value 2 between the generated anime portrait image and the real anime portrait image is calculated; the loss function value 2 is compared with the preset loss threshold 2 to obtain the image generation model.
2. The method for generating anime portrait images from sketches based on LDM according to claim 1, characterized in that: The preprocessing of the sketch is to perform data enhancement on each sketch.
3. The method for generating anime portrait images from sketches based on LDM according to claim 1, characterized in that: The sketch encoder includes a head and face contour encoder, a mouth encoder, a nose encoder, a left eye encoder and a right eye encoder.
4. The method for generating anime portrait images from sketches based on LDM according to claim 1 or 3, characterized in that: The calculating a loss function value between a sketch in a training set and a corresponding reconstructed sketch includes: ; in, is the loss function value between the sketch and the corresponding reconstructed sketch, is the set of sketch encoders, i is the sketch encoder number, i=0,1,…,5, is the sketch encoder i, and x is the sketch corresponding to the sketch encoder i.
5. The method for generating anime portrait images from sketches based on LDM according to claim 4, characterized in that: The second loss function value between the calculated generated cartoon portrait picture and the real cartoon portrait picture includes: Extract sketch features, and calculate the loss function value between the generated anime portrait image and the real anime portrait image based on the loss function of the latent diffusion model, as follows: ; in, is the loss function value between the generated anime portrait image and the real anime portrait image, E is the expected value, t is the time step of the current iteration, and t [1,T], T is the total time steps of the entire diffusion process; is a pre-trained encoder; is the real noise; is the noise predicted by the model; is the coded sketch; For the decoder; A collection of sketch features.
6. The method for generating anime portrait images from sketches based on LDM according to claim 5, characterized in that: The comparing the loss function value 1 with a preset loss threshold 1 includes: Determine whether the loss function value 1 is less than a preset loss threshold 1; if yes, mark the training result of the sketch encoding as completed; if no, mark the training result of the sketch encoding as not completed; The comparing the loss function value 2 with the preset loss threshold 2 includes: It is determined whether the loss function value 2 is less than the preset loss threshold 2; if so, the invisible diffusion model is marked as an image generation model; if not, the invisible diffusion model is not marked.
7. The method for generating anime portrait images from sketches based on LDM according to claim 1 or 5, characterized in that: The loss function of the implicit diffusion model includes: 。 8. The method for generating anime portrait images from sketches based on LDM according to claim 1 or 5, characterized in that: The implicit diffusion model includes a forward diffusion model and a backward diffusion model, including: The forward diffusion model is as follows: ; in, is a multivariate normal distribution; The noise scheduling parameters , is the noise variance at step t, is the initial image data, is the image data after the t-th noise, , is the random noise of standard normal distribution; is the identity matrix; The back diffusion model is as follows: ; in, is the image data after the t-1th noise, is the mean of the Gaussian distribution, is the variance of the Gaussian distribution.
9. A system for generating cartoon portrait images from sketches based on LDM, operating based on the method for generating cartoon portrait images from sketches based on LDM according to any one of claims 1 to 8, characterized in that: Includes acquisition module and model training module; Acquisition module: used to obtain several sketches and the corresponding anime portrait images to build a data set; Model training module: used to input the sketches in the training set into the sketch encoder, generate a reconstructed sketch, calculate the loss function value 1 between the sketches in the training set and the corresponding reconstructed sketch, compare the loss function value 1 with a preset loss threshold 1, and obtain a trained sketch encoder; The sketches in the training set are data-augmented, and the data-augmented sketches are input into the trained sketch encoder to obtain sketch features; Generate anime portrait images based on sketch features and latent diffusion model; calculate loss function value 2 between the generated anime portrait images and the real anime portrait images; compare loss function value 2 with preset loss threshold 2 to obtain image generation model.
10. A storage medium storing a computer program, wherein the computer program, when executed, runs the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Face sketch coloring method based on reference picture
CN116168107A
Generating novel images using sketch image representations
US20230419551A1