Frame generation method and system based on stable diffusion model

By constructing a target border generation model, combining deep learning and image generation algorithms, the problems of lack of data resources and poor direction in traditional cultural image generation are solved, and efficient and customized traditional cultural border generation are achieved.

CN119941890APending Publication Date: 2025-05-06SHANDONG NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411991920.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art has problems such as lack of data resources and poor directionality of AI-generated images when generating traditional cultural images, which is difficult to meet the needs of precise design.

Method used

The border generation method based on the stable diffusion model is adopted to build a target border generation model, combine deep learning and image generation algorithms to receive user text prompts or images to generate borders with traditional culture.

Benefits of technology

It achieves more accurate capture and reproducing traditional cultural and artistic styles, providing highly customized border design options, reducing time and cost during the design process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941890A_ABST
    Figure CN119941890A_ABST
Patent Text Reader

Abstract

The invention provides a frame generation method and system based on a stable diffusion model, relates to the technical field of image frame design, and aims to solve the problem that an existing model is lack of Chinese wind element data resources and is difficult to meet the requirement of generating a Chinese wind frame. The method comprises the steps of obtaining a picture containing a target feature, and receiving a text prompt of a user; processing the picture to obtain a processed picture; generating image features related to a preset style by using the fine-tuned stable diffusion model; the text prompt of the user is converted into text output which can be processed by the image generator through a text encoder; and inputting the image features related to the preset style to an image generator to generate a target frame image. According to the method, the frame with the Chinese style is generated by using the deep learning algorithm, highly customized frame design options are provided for the user, the art style of the Chinese style is more accurately captured and reproduced, the frame design options of the Chinese style are provided, and the time and the cost of the design process are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of image frame generation, and in particular relates to a frame generation method and system based on a stable diffusion model. Background Art

[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.

[0003] With the continuous advancement of technology, the application of artificial intelligence in the field of design is becoming more and more extensive and in-depth. In the design process, it can help designers create and edit visual elements. However, in the field of design, when AI technology is used to explore the traditional cultural image content in AI-generated content, we can see that there are problems in the combination of AI technology and traditional culture on multiple levels.

[0004] There is a lack of data resources on traditional cultural elements. Text and image data provide rich training materials for generative models. By learning large-scale text descriptions and corresponding images, the model can generate high-quality images that match the text descriptions. However, due to different data sources, resource allocation and cultural differences, there is a lack of data resources on traditional cultural elements in AI databases.

[0005] AI-generated images have poor directivity. When AI generates images, especially low-detail images, the composition is basically determined randomly. This means that it is difficult for users to specify the specific location of something in the image. In the UI field, we often need to specify elements accurately to the pixel. The randomness and poor directivity of AI make it difficult to meet precise design requirements.

[0006] To sum up, the problems with current technology are:

[0007] (1) The current models lack data resources for traditional cultural elements, and the generated effects of models for specific styles may not meet the needs.

[0008] (2) AI-generated images have poor directionality and are unable to express precise design elements. Summary of the invention

[0009] In order to overcome the shortcomings of the above-mentioned prior art, the present invention provides a border generation method and system based on a stable diffusion model. By constructing a target border generation model, it is possible to provide users with text prompts or images and generate borders with traditional culture. This technology combines deep learning and image generation algorithms, and by simulating traditional cultural art elements, creates visual art works that combine traditional cultural elements with design elements, solving the problem of insufficient regional cultural adaptation and lack of design elements when AI is used as an assistant.

[0010] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions: A first aspect of the present invention provides a border generation method based on a stable diffusion model, comprising:

[0011] Get pictures containing target features and receive text prompts from users;

[0012] Processing the image to obtain a processed image;

[0013] The processed image is input into the fine-tuned stable diffusion model to generate image features related to the preset style;

[0014] Through the text encoder, the user's text prompt is converted into text output that can be processed by the image generator;

[0015] The text output and image features related to the preset style are input into the image generator to generate the target border image.

[0016] As an implementation mode, the image is processed to obtain a processed image, specifically:

[0017] Marking the image content to obtain a marked image;

[0018] The marked image information is translated into Chinese and English to obtain the processed image.

[0019] As an implementation mode, after the image is processed, the stable diffusion model is trained, and the parameters of the stable diffusion model are set and adjusted by a trainer.

[0020] As an implementation mode, receiving a text prompt from a user is specifically as follows: in a stable diffusion model, a large model is selected as a base film, and a text prompt is input in a user interface of the stable diffusion model.

[0021] As an implementation method, a text prompt is input in the user interface of the stable diffusion model, specifically: a trigger word is input in the user interface, and a positive prompt word and a reverse prompt word are input in the text input box;

[0022] Set the usage weight, adoption method, guide word coefficient, magnification factor, iteration step number and magnification algorithm of the stable diffusion model.

[0023] As an implementation method, the text output and the image features related to the preset style are input into an image generator to generate a target frame image, wherein the image generator includes an image information creator and an image decoder, specifically:

[0024] Inputting text input and image features related to a preset style into an image information creator to obtain border image information of the preset style;

[0025] The preset style border image information is input into the image decoder to generate the target border.

[0026] As an implementation method, the stable diffusion model is fine-tuned. Specifically, the stable diffusion model is fine-tuned through a super model fusion tool.

[0027] A second aspect of the present invention provides a border generation system based on a stable diffusion model, comprising:

[0028] The data acquisition module is used to obtain pictures containing target features and receive text prompts from users;

[0029] An image processing module, used to process the image to obtain a processed image;

[0030] The target bounding box generation module is used to input the processed image into the fine-tuned stable diffusion model to generate image features related to the preset style;

[0031] Through the text encoder, the user's text prompt is converted into text output that can be processed by the image generator;

[0032] The text output and image features related to the preset style are input into the image generator to generate the target border image.

[0033] A third aspect of the present invention provides a computer device, comprising a memory, a processor, and a program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps in the method described in the first aspect of the present invention are implemented.

[0034] The fourth aspect of the present invention aims to provide a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the steps in the method described in the first aspect of the present invention.

[0035] One or more of the above technical solutions have the following beneficial effects:

[0036] In this embodiment, a Chinese-style border generation model based on Stable Diffusion technology is constructed. The model can receive text prompts or images as input, use deep learning algorithms to generate Chinese-style borders, and provide users with highly customized border design options. Compared with traditional image generation models, the model in this embodiment can more accurately capture and reproduce the Chinese art style, provide Chinese-style border design options, and reduce time and cost in the design process.

[0037] Advantages of additional aspects of the present invention will be given in part in the following description, and in part will become obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0039] Figure 1 A schematic diagram of a border generation method based on a stable diffusion model according to the first embodiment of the present invention;

[0040] Figure 2 This is an accurate workflow diagram before the target border model is generated in the first embodiment;

[0041] Figure 3 This is a user interface flow chart of the first embodiment;

[0042] Figure 4 This is an example diagram of different iteration steps for generating a target border in the first embodiment of the present invention;

[0043] Figure 5 This is an example diagram of different sampling methods for generating a target border in the first embodiment of the present invention;

[0044] Figure 6 This is an example image of a Chinese red, paper-cut style frame generated in the first embodiment;

[0045] Figure 7 This is an example diagram of training settings for a stable diffusion model of the first embodiment. DETAILED DESCRIPTION

[0046] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.

[0047] It should be noted that the terms used herein are for describing specific embodiments only and are not intended to be limiting of exemplary embodiments according to the present invention.

[0048] In the absence of conflict, the embodiments of the present invention and the features of the embodiments may be combined with each other.

[0049] Terminology explanation:

[0050] Stablediffusion: The Stable Diffusion Model is a text-to-image generation model based on deep learning, which can generate high-quality and high-resolution images based on text descriptions.

[0051] Embodiment 1

[0052] This embodiment discloses a border generation method based on a stable diffusion model.

[0053] In order to more clearly illustrate this embodiment, the implementation process of a border generation method based on a stable diffusion model can be specifically described as follows:

[0054] A border generation method based on a stable diffusion model, comprising:

[0055] S1, obtain the picture containing the target feature and receive the user's text prompt;

[0056] S2, processing the image to obtain a processed image;

[0057] S3, inputting the processed image into the fine-tuned stable diffusion model to generate image features related to the preset style;

[0058] S4, converting the user's text prompt into a text output that can be processed by the image generator through a text encoder;

[0059] S5. Input the text output and the image features related to the preset style into the image generator to generate a target border image.

[0060] like Figure 1 As shown, in step S1, a picture containing target features is obtained, and a text prompt from a user is received.

[0061] S1-1. Obtain a picture containing target features.

[0062] In this embodiment, a Chinese style frame image is generated, wherein the Chinese style includes Chinese style elements and Chinese style colors. The Chinese style elements include at least: paper-cutting, auspicious clouds, Peking opera, Chinese knots, lanterns and oil-paper umbrellas. The Chinese style colors include at least: Chinese red, rouge, concubine color and water green.

[0063] Pictures containing Chinese elements and colors can be obtained through manual drawing, Internet collection and other means.

[0064] Through the above steps, the set Chinese style pictures are collected and summarized to provide data support for the subsequent extraction of elements of the set style.

[0065] S1-2. Receive text prompts from the user.

[0066] The user's text prompt is received, specifically: in the stable diffusion model, the large model is selected as the base film, and the text prompt is input in the user interface of the stable diffusion model.

[0067] Specifically, first enter the trigger word in the user interface, and enter the positive and reverse prompt words in the text input box, and then set the model usage weight, adoption method, guide word coefficient, magnification factor, iteration steps and magnification algorithm.

[0068] In this embodiment, the user selects the ReVAnimated_v122_V122 large model as the base model in the Stable Diffusion operation interface, and then enters the prompt word. The trigger word of the present invention is: "biankuang", the model usage weight is "1", and the general reverse prompt word. The sampling method "Euler a" is the most preferred after testing. The number of iterations "30" is the best after testing. The guide word coefficient is "7" and the resolution defaults to "512×512". Select high-resolution restoration, the magnification is "2", the magnification algorithm is "R-ESRGAN 4x+" and the redrawing amplitude is "0.2".

[0069] In this embodiment, the Stable Diffusion operation interface uses Web UI or Comfy UI to design the user interface.

[0070] like Figure 1 As shown, in step S2, the picture is processed to obtain a processed picture.

[0071] In this embodiment, the image content is marked to obtain a marked image;

[0072] The marked image information is translated into Chinese and English to obtain the processed image.

[0073] After the above steps, the accuracy of the picture is ensured to meet the needs of subsequent operations.

[0074] like Figure 1 As shown, in step S3, the processed image is input into the fine-tuned stable diffusion model to generate image features related to the preset style.

[0075] In this embodiment, S3-1, construct a target bounding box generation model.

[0076] The target bounding box generation model includes a stable diffusion model, a text encoder, and an image generator.

[0077] In order to solve the problem of insufficient data resources for Chinese style elements, poor directionality of AI raw images, and inability to generate Chinese style borders, a stable diffusion model is used to extract Chinese style elements and color features.

[0078] In order to improve the extraction effect, the stable diffusion model is trained and fine-tuned; the text encoder and image generator can be trained.

[0079] S3-2, input the processed image into the stable diffusion model for training, obtain the trained stable diffusion model, and fine-tune the stable diffusion model.

[0080] S3-2-1. Train the stable diffusion model.

[0081] In this embodiment, a pre-trained Stable Diffusion model is used, which can understand and generate image elements and color features related to Chinese style art.

[0082] Based on the particularity of the Stable Diffusion model, the SD-Trainer is imported to adjust the parameters reasonably, such as Figure 7 As shown, the training of the stable diffusion model is completed.

[0083] S3-2-2. Fine-tune the stable diffusion model.

[0084] In this embodiment, the super model fusion tool is used to set parameters in the stable diffusion model to achieve fine-tuning of the stable bolt model. Set the parameters and set the hierarchical control weights to: down lr_weight-"1,0.2,1,1,0.2,1,1,0.2,1,1,1,1", mid lr_weight."1", up lr_weight-"1,1,1,1,1,1,1,1,1,1,1,1".

[0085] After the above steps, the hierarchical control of the model can keep the background and style consistent as much as possible in various environments, standardize the border shape, and better integrate with the border when generating new elements.

[0086] S3-3. Test the target bounding box generation model.

[0087] like Figure 6 As shown, in this embodiment, taking the generation of a Chinese red, paper-cut style border as an example, the trigger word "biankuang" is input, and then the instruction "Red theme, paper cutouts, flowers" is input.

[0088] The input instructions are fed into the text encoder, where they are converted into text output that can be processed by the image generator.

[0089] The text output and image features related to the preset style are passed through the image generator to generate a Chinese red, paper-cut style border.

[0090] After testing, a tested target border generation model was obtained. The availability rate of the target border generated by this model is about 80%. The target border generation model can be used to generate various Chinese-style borders containing Chinese elements and Chinese colors.

[0091] like Figure 1 As shown, in step S3, the processed image is input into a stable diffusion model to generate image features related to traditional culture.

[0092] In this embodiment, generation of a "rouge-colored, auspicious cloud-style border" is taken as an example.

[0093] In the Stable Diffusion model operation interface, enter the trigger word "biankuang", and the received user text prompt is "Carmine theme, auspicious cloud". Enter the text prompt into the prompt word filling box of the operation interface.

[0094] Collect and summarize pictures containing Chinese style elements and colors, extract elements and color features from the pictures, so that they meet the principles and methods of border design and conform to the specifications of the fine-tuning model.

[0095] like Figure 1 As shown, in step S4, the user's text prompt is converted into a text output that can be processed by the target bounding box generation model through a text encoder.

[0096] In this embodiment, the user's text prompt received in the stable diffusion model is input into the text encoder CLIPTextModel, and the user's text prompt is converted into a format that can be understood by the target bounding box generation model. The formula is:

[0097] Text_embedding=CLIPTextModel(H i );

[0098] Among them, Text_embedding represents text output, H i Represents the i-th user text prompt.

[0099] The text prompt is input into the text encoder for encoding, where it is converted into text output that can be processed by the image generator.

[0100] like Figure 1 As shown, in step S5, the text output and the image features related to the preset style are input into the image generator to generate the target frame image.

[0101] In this embodiment, an image generator is used to generate a target frame image according to the output of the text encoder, wherein the image generator includes an image information creator and an image decoder.

[0102] Generates a rouge-colored, auspicious cloud-style Chinese border.

[0103] S5-1. Input text input and image features related to a preset style into an image information creator to obtain frame image information of the preset style.

[0104] In this embodiment, in the image information creator, the text input of "rouge color, auspicious cloud style border" and the image elements and colors of the preset style are created as the border image information features of the preset style. The formula is:

[0105] Image_embedding=CLIPImageEncoder(Text_embedding,I i );

[0106] Among them, Text_embedding represents text output, I i Represents the image features related to the preset style. A Gaussian noise matrix is ​​generated using the random function, i.e., the potential features of the border of the preset style, and image optimization is performed iteratively to obtain the optimized potential features of the border of the preset style, i.e., the border image information of the preset style. Here, the border image information of the rouge color and auspicious cloud style is obtained.

[0107] S5-2, inputting the frame image information of the preset style into the image decoder for decoding to generate a target frame.

[0108] In this embodiment, the rouge color and auspicious cloud style border image information is input into the image decoder for decoding, and the rouge color and auspicious cloud style border image information is reconstructed into a pixel level image, that is, a rouge color and auspicious cloud style border image is obtained.

[0109] This embodiment combines deep learning and image generation algorithms to create visual art works that combine Chinese style elements with design elements by simulating Chinese style artistic elements and colors. It emphasizes the integration of the culture, environment, tradition and lifestyle of a specific region in the design process, and proposes a method to solve the problem of AI as an assistant, which is insufficiently adapted to Chinese regional culture and lacks design elements in the design process.

[0110] Embodiment 2

[0111] The purpose of this embodiment is to provide a border generation system based on a stable diffusion model, including:

[0112] The data acquisition module is used to obtain pictures containing target features and receive text prompts from users;

[0113] An image processing module, used to process the image to obtain a processed image;

[0114] The target bounding box generation module is used to input the processed image into the fine-tuned stable diffusion model to generate image features related to the preset style;

[0115] Through the text encoder, the user's text prompt is converted into text output that can be processed by the image generator;

[0116] The text output and image features related to the preset style are input into the image generator to generate the target border image.

[0117] Based on providing a border generation system based on a stable diffusion model, the method steps in embodiment 1 are implemented.

[0118] Embodiment 3

[0119] The purpose of this embodiment is to provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the program.

[0120] Embodiment 4

[0121] The purpose of this embodiment is to provide a computer-readable storage medium.

[0122] A computer-readable storage medium stores a computer program, which executes the steps of the above method when executed by a processor.

[0123] The steps involved in the apparatus of the above embodiment correspond to the method embodiment 1, and the specific implementation method can refer to the relevant description part of embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood to include any medium that can store, encode or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.

[0124] Those skilled in the art should understand that the modules or steps of the present invention described above can be implemented by a general-purpose computer device, or alternatively, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.

[0125] Although the above describes the specific implementation mode of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without creative work are still within the scope of protection of the present invention.

Claims

1. A border generation method based on a stable diffusion model, characterized in that: include: Get pictures containing target features and receive text prompts from users; Processing the image to obtain a processed image; The processed image is input into the fine-tuned stable diffusion model to generate image features related to the preset style; Through the text encoder, the user's text prompt is converted into text output that can be processed by the image generator; The text output and image features related to the preset style are input into the image generator to generate the target border image.

2. A border generation method based on a stable diffusion model as claimed in claim 1, characterized in that: The image is processed to obtain a processed image, specifically: Marking the image content to obtain a marked image; The marked image information is translated into Chinese and English to obtain the processed image.

3. The border generation method based on the stable diffusion model according to claim 1, characterized in that: After the image is processed, the stable diffusion model is trained, and the parameters of the stable diffusion model are set and adjusted through the trainer.

4. The border generation method based on the stable diffusion model according to claim 1, characterized in that: The user's text prompt is received, specifically: in the stable diffusion model, the large model is selected as the base film, and the text prompt is input in the user interface of the stable diffusion model.

5. A border generation method based on a stable diffusion model as claimed in claim 4, characterized in that: Enter text prompts in the user interface of the stable diffusion model, specifically: enter a trigger word in the user interface, and enter positive and reverse prompt words in the text input box; Set the usage weight, adoption method, guide word coefficient, magnification factor, iteration step number and magnification algorithm of the stable diffusion model.

6. A border generation method based on a stable diffusion model as claimed in claim 1, characterized in that: The text output and the image features related to the preset style are input into the image generator to generate a border image, wherein the image generator includes an image information creator and an image decoder, specifically: Inputting text input and image features related to a preset style into an image information creator to obtain border image information of the preset style; The preset style border image information is input into the image decoder to generate the target border.

7. A border generation method based on a stable diffusion model as claimed in claim 1, characterized in that: The stable diffusion model is fine-tuned. Specifically, the stable diffusion model is fine-tuned through a super model fusion tool.

8. A border generation system based on a stable diffusion model, characterized in that: include: The data acquisition module is used to obtain pictures containing target features and receive text prompts from users; An image processing module, used to process the image to obtain a processed image; The target bounding box generation module is used to input the processed image into the fine-tuned stable diffusion model to generate image features related to the preset style; Through the text encoder, the user's text prompt is converted into text output that can be processed by the image generator; The text output and image features related to the preset style are input into the image generator to generate the target border image.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the method described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method described in any one of claims 1 to 7 are performed.