A multi-style picture book generation method based on diffusion models
By introducing preprocessing modules, style consistency modules, role consistency modules and attention-based Unet modules in the diffusion model, the problem of insufficient character and background consistency when generating multi-frame story picture books in the existing technology is solved, and efficient and flexible generation of multi-style picture books is achieved, improving image quality and style diversity.
Patent Information
- Application Number
- CN202411734471.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-29
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-11-29
AI Technical Summary
When existing image generation technology based on diffusion model generates complex and coherent multi-frame story picture books, it is difficult to ensure the consistency of characters and backgrounds, affect the coherence of stories and the fluency of narratives, and lacks the flexibility and consistency of generating multiple styles of images in continuous picture books.
A multi-style picture book generation method based on diffusion model is proposed, including preprocessing module, style consistency module, role consistency module and attention mechanism-based Unet module. Through these modules, text, characters, styles and images are encoded and embedded to ensure the consistency and diversity of the generated picture book images in style and roles.
It realizes efficient and flexible generation of continuous picture book images of multiple artistic styles, ensuring the consistency and consistency of generated images, and improving the quality and style diversity of picture book images.
Smart Images

Figure CN119228633B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and particularly relates to a multi-style picture book generation method based on a diffusion model. Background Art
[0002] Currently, manually creating story picture books with different artistic styles is time-consuming and laborious, and requires artists to have high painting skills and the ability to control styles. With the rapid development of deep learning technology, the capabilities of computer vision and generative models have been greatly improved, especially in the fields of image generation and style transfer, where breakthrough progress has been made.
[0003] Existing image generation technologies based on the Diffusion Model can already generate high-quality single images according to text prompts. However, when applied to generating complex and coherent multi-frame story picture books, these technologies face some challenges. The main problem lies in the lack of consistency of characters and backgrounds between different frames, resulting in inconsistent changes in character images and scenes, which affects the coherence of the story and the smoothness of the narrative.
[0004] In addition, although existing style transfer technologies can convert images into specific artistic styles, they usually can only handle single-style conversions and lack the flexibility and consistency required to generate images of multiple styles in continuous picture books.
[0005] In the task of story picture book generation, it is necessary to ensure that the generated characters maintain the consistency of appearance and features in the front and back frames, and avoid image changes or distortions. At the same time, it is necessary to ensure that the overall style of the generated results is unified and the background transition is smooth and natural to maintain the sense of space and the smoothness of the narrative. Therefore, the existing technologies still need to be improved in the following aspects: they cannot efficiently and flexibly generate continuous picture book images of multiple artistic styles; it is difficult to ensure the coherence and consistency of the generated images in the picture book narrative; most existing image style transfer algorithms are limited to single scenes and do not have the ability of dynamic multi-style transfer. Summary of the Invention
[0006] To solve these problems, the present invention proposes a multi-style picture book generation method based on a diffusion model, including the following steps:
[0007] Step S1: Construct a picture book data set, which includes several picture book images and corresponding story texts; construct a style reference data set, which includes several style reference images;
[0008] Step S2: Construct a picture book generation model based on a diffusion model. The model includes a preprocessing module, a style consistency module, a character consistency module, and an Unet module based on an attention mechanism. The preprocessing module encodes the picture book images and corresponding story texts in Step S1 to obtain text embeddings, character masks, and character images.
[0009] Step S3: Import the style reference image in Step S1 into the style consistency module to obtain style feature embeddings.
[0010] Step S4: Import the text embeddings, character masks, and character images in Step S2 into the character consistency module to obtain character embeddings and layout embeddings.
[0011] Step S5: Import the picture book images in Step S1, the style feature embeddings in Step S3, the character embeddings and layout embeddings in Step S4 into the attention blocks in the Unet module based on the attention mechanism to predict the noise of the picture book images and obtain the predicted picture book images.
[0012] Step S6: Construct a loss function and minimize the loss function to optimize the parameters of the picture book generation model.
[0013] Furthermore, the preprocessing module in Step S1 includes a CLIP encoder and an image segmentation model GSA.
[0014] Step S1 is specifically as follows:
[0015] Step S11: Encode the story text through a CLIP encoder to obtain text embeddings with a preset data dimension, expressed as: , expressed as:
[0016] ;
[0017] Among them, represents regularization, represents a multi-layer perceptron, n represents the number of internal operations of the multi-layer perceptron, represents a self-attention operation, represents tokenizing the text, represents position embedding encoding, represents the input story text;
[0018] Step S12: Use the image segmentation model GSA to segment the picture book images to obtain character masks and character images , expressed as:
[0019] ;
[0020] Among them, represents the picture book image; represents a selection function for obtaining mask information of specific characters in a picture book image, represents an image segmentation module, represents dot multiplication.
[0021] Furthermore, step S3 is specifically as follows:
[0022] Step S31: Process the input style reference image specifically as follows:
[0023] Call the text - image large - language model to generate the text semantic content of the style reference image, and then use the image encoder and text encoder of the CLIP encoder to encode the style reference image and the corresponding text semantic content respectively, obtain the image encoding and text encoding of the style reference image map the two to the same semantic space, subtract the text encoding from the image encoding, so as to obtain the style feature embedding without text semantic content which is expressed as:
[0024] ;
[0025] wherein, represents the text encoder of the CLIP encoder, represents the multi - modal large - language model, represents the image encoder of the CLIP encoder, represents the multi - layer perceptron;
[0026] Step S32: Further process the style feature embedding to obtain the style feature, which is expressed as:
[0027] ;
[0028] wherein, represents the learnable embedding, represents the self - attention operation, represents the cross - attention operation, represents the fully - connected layer, represents the style feature.
[0029] Furthermore, the character consistency module includes a resampling module and a layout embedding module,
[0030] Step S4 is specifically as follows:
[0031] Step S41: The character image and the character mask Input character consistency module, obtain the resampled embeddings corresponding to each character, and then perform cross-attention calculation with the intermediate noise of the diffusion model through MLP mapping to obtain the character embeddings , expressed as:
[0032] ;
[0033] Among them, represents the resampled embedding, represents the intermediate noise of the diffusion model, represents the character embedding, represents the cross-attention mechanism, represents the resampling operation, represents the multi-layer perceptron;
[0034] Step S42: Input text embedding and resampled embedding to the layout control module for processing to obtain the layout embedding , expressed as:
[0035] ;
[0036] Among them, represents the fully connected layer, and noise represents the input noise.
[0037] Furthermore, the Unet module based on the attention mechanism in step S5 includes several attention blocks for guiding image generation;
[0038] Step S5 is specifically:
[0039] ;
[0040] Among them, represents the predicted picture book image, represents the Unet module, represents the noisy picture book image, represents the number of noise addition steps.
[0041] Furthermore, the loss function in step S6 is expressed as:
[0042] ;
[0043] Among them, , , represents the weight coefficients of different losses, , , are respectively the loss of the diffusion model, the loss of character consistency, and the loss of style consistency.
[0044] The positive and progressive effects of the present invention are as follows:
[0045] Based on the diffusion model, the present invention proposes an efficient multi-style picture book generation method, which can generate story picture books in any artistic style according to style reference images, greatly simplifies the creation process, and at the same time improves the quality and style diversity of the generated picture book images. Brief Description of the Drawings
[0046] Figure 1 It is a flowchart of the steps of a multi-style picture book generation method based on the diffusion model of the present invention. Detailed Embodiments
[0047] The following uses specific specific examples to illustrate the embodiments of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention.
[0048] Refer to Figure 1 , a multi-style picture book generation method based on the diffusion model. In one example, step S1: Construct a picture book data set, which includes several picture book images and corresponding story texts; construct a style reference data set, which includes several style reference images;
[0049] Step S2: Construct a picture book generation model based on the diffusion model. The model includes a preprocessing module, a style consistency module, a character consistency module, and a Unet module based on the attention mechanism; the preprocessing module encodes the picture book images and corresponding story texts in step S1 to obtain text embeddings, character masks, and character images;
[0050] Step S3: Import the style reference image in step S1 into the style consistency module to obtain style feature embeddings;
[0051] Step S4: Import the text embeddings, character masks, and character images in step S2 into the character consistency module to obtain character embeddings and layout embeddings;
[0052] Step S5: Import the picture book images in step S1, the style feature embeddings in step S3, the character embeddings and layout embeddings in step S4 into the attention block in the Unet module based on the attention mechanism to perform picture book image noise prediction, and obtain the predicted picture book images;
[0053] Step S6: Construct a loss function and minimize the loss function to optimize the parameters of the picture book generation model.
[0054] Further, in one example, the preprocessing module in step S1 includes a CLIP encoder and an image segmentation model GSA;
[0055] Step S1 is specifically as follows:
[0056] Step S11: Encode the story text through the CLIP encoder to obtain a text embedding with a preset data dimension, , expressed as:
[0057] ;
[0058] Among them, represents regularization, represents a multi-layer perceptron, n represents the number of internal operations of the multi-layer perceptron, represents a self-attention operation, represents a tokenization operation on the text, represents a positional embedding encoding, represents the input story text;
[0059] Step S12: Use the image segmentation model GSA to segment the picture book image to obtain a character mask and a character image , expressed as:
[0060] ;
[0061] Among them, represents the picture book image; represents a selection function for obtaining mask information of a specific character in the picture book image, represents an image segmentation module, represents a dot product.
[0062] Further, in one example, step S3 is specifically as follows:
[0063] Step S31: Process the input style reference image , specifically as follows:
[0064] Call the text-image large language model to generate the text semantic content of the style reference image , and then use the image encoder and text encoder of the CLIP encoder to encode the style reference image and the corresponding text semantic content respectively to obtain the image encoding and text encoding of the style reference image , map the two to the same semantic space, and subtract the text encoding from the image encoding to obtain a style feature embedding without text semantic content , expressed as:
[0065] ;
[0066] Among them, represents the text encoder of the CLIP encoder, represents the multimodal large language model, represents the image encoder of the CLIP encoder, represents the multi-layer perceptron;
[0067] Step S32: Further process the style feature embedding to obtain the style feature, expressed as:
[0068] ;
[0069] Among them, represents the learnable embedding, represents the self-attention operation, represents the cross-attention operation, represents the fully connected layer, represents the style feature.
[0070] Furthermore, in one example, the character consistency module includes a resampling module and a layout embedding module.
[0071] Step S4 is specifically as follows:
[0072] Step S41: Input the character image and the character mask into the character consistency module to obtain the resampling embedding corresponding to each character, and then perform cross-attention calculation with the intermediate noise of the diffusion model through MLP mapping to obtain the character embedding , expressed as:
[0073] ;
[0074] Among them, represents the resampling embedding, represents the intermediate noise of the diffusion model, represents the character embedding, represents the cross-attention mechanism, represents the resampling operation, represents the multi-layer perceptron;
[0075] Step S42: Input the text embedding and the resampling embedding into the layout control module for processing to obtain the layout embedding , expressed as:
[0076] ;
[0077] Among them, represents the fully connected layer, and noise represents the input noise.
[0078] Furthermore, in one example, the Unet module based on the attention mechanism in step S5 includes several attention blocks for guiding image generation;
[0079] Step S5 is specifically as follows:
[0080] ;
[0081] Among them, represents the predicted picture book image, represents the Unet module, represents the picture book image after adding noise, represents the number of noise addition steps.
[0082] Furthermore, in one example, the loss function in step S6 is expressed as:
[0083] ;
[0084] Among them, , , represents the weight coefficients of different losses, , , are respectively the loss of the diffusion model, the loss of character consistency, and the loss of style consistency.
[0085] The present invention has been described in detail above in conjunction with the embodiments with reference to the drawings. Those of ordinary skill in the art can make various variations of the present invention according to the above description. Therefore, some details in the embodiments should not constitute a limitation to the present invention, and the present invention will be protected by the scope defined by the appended claims.
Claims
1. A method for generating multi-style picture books based on a diffusion model, characterized in that: The following steps are involved: Step S1: construct a picture book dataset, which includes a number of picture book images and corresponding story texts; construct a style reference dataset, which includes a number of style reference images; Step S2: construct a picture book generation model based on a diffusion model, which includes a preprocessing module, a style consistency module, a role consistency module, and a Unet module based on an attention mechanism; the preprocessing module encodes the picture book image and the corresponding story text in step S1 to obtain text embedding, role mask, and role image; Step S3: import the style reference image of step S1 into the style consistency module to obtain style feature embedding; Step S4: import the text embedding, role mask and role image of step S2 into the role consistency module to obtain the role embedding and layout embedding; Step S5: import the picture book image in step S1, the style feature embedding in step S3, the character embedding and the layout embedding in step S4 into the attention block in the Unet module based on the attention mechanism to perform picture book image noise prediction and obtain the predicted picture book image; Step S6: construct a loss function, and minimize the loss function to optimize the parameters of the picture book generation model; Step S4 is specifically as follows: Step S41: Transform the character image and role mask Input the role consistency module to obtain the resampled embedding corresponding to each role, and then perform cross-attention calculation through MLP mapping and diffusion model intermediate noise to obtain the role embedding , expressed as: ; in, represents the resampled embedding, represents the intermediate noise of the diffusion model, represents role embedding, represents the cross attention mechanism, represents the resampling operation, represents a multi-layer perceptron; Step S42: Input text embedding and role embedding Go to the layout control module for processing and obtain the layout embedding , expressed as: ; in, represents the fully connected layer, and noise represents the input noise.
2. The method for generating a multi-style picture book based on a diffusion model as claimed in claim 1, characterized in that: The preprocessing module in step S1 includes the CLIP encoder and the image segmentation model GSA, specifically: Step S11: Encode the story text using the CLIP encoder to obtain text embedding with preset data dimensions , expressed as: ; in, represents regularization, represents a multilayer perceptron, n represents the number of internal operations performed by the multilayer perceptron, represents the self-attention operation, Indicates the word segmentation operation on the text. represents the position embedding code, Represents the input story text; Step S12: Use the image segmentation model GSA to segment the picture book image and obtain the character mask and character images , expressed as: ; in, Indicates picture book images; represents the selection function, which is used to obtain the mask information of a specific character in the picture book image. represents the image segmentation module, Represents dot product.
3. The method for generating a multi-style picture book based on a diffusion model as claimed in claim 2, characterized in that: Step S3 is specifically as follows: Step S31: For the input style reference image To process, specifically: Call the large language model of images and text to generate the reference image of the style The text semantic content of the style reference image is then encoded using the image encoder and text encoder of the CLIP encoder. Encode the corresponding text semantic content to obtain the style reference image The image encoding and text encoding are mapped to the same semantic space, and the text encoding is subtracted from the image encoding to obtain the style feature embedding without the text semantic content. , expressed as: ; in, represents the text encoder of the CLIP encoder, Representing a large language model that is multimodal, represents the image encoder of the CLIP encoder, represents a multi-layer perceptron; Step S32: embedding style features Further processing is performed to obtain the style features, which are expressed as: ; in, represents a learnable embedding, represents the self-attention operation, represents the cross attention operation, represents the fully connected layer, Indicates style characteristics.
4. The method for generating a multi-style picture book based on a diffusion model as claimed in claim 3, characterized in that: The attention-based Unet module of step S5 includes several attention blocks for guiding image generation; Step S5 is specifically as follows: ; in, represents the predicted picture book image, Represents the Unet module, represents the picture book image after adding noise, Indicates the number of noise adding steps.
5. The method for generating a multi-style picture book based on a diffusion model as claimed in claim 1, characterized in that: Loss function in step S6 It is expressed as: ; in, , , Represents the weight coefficient of different losses, , , They are the loss of the diffusion model, the loss of role consistency, and the loss of style consistency.
Citation Information
Patent Citations
Text illustrating method and device
CN116385597A
Image redrawing model training method, image redrawing method and device
CN116664719A