AI-based text cartoon device and method

Through the AI-based Wensheng comic device and method, the script analysis module and LoRA model training are used, combined with the Stable Diffusion framework and plug-in, the multi-dimensional control and detail repair of characters and scenes are realized, which solves the problem of inconsistent images generated by existing AI tools, improves the controllability and efficiency of generated content, and is suitable for the automated production of comic creation and plot commentary videos.

CN120339457APending Publication Date: 2025-07-18SHANGHAI 2345 NETWORK TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510202316.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing AI Wensheng comic tools have flaws in the controllability of generating pictures, inconsistent character images and scenes, difficulty in controlling actions and expressions, insufficient detailed repair capabilities, resulting in poor coherence of comics or videos, and the existing solutions have problems such as high training costs or low efficiency.

Method used

Using AI-based Wensheng comics device and method, image parameters are extracted through the script analysis module, combined with LoRA model training and Stable Diffusion framework, Regional Prompter, ControlNet and Adetailer plug-ins are used to realize multi-dimensional control of character position, action, and scene style, generate highly consistent keyframe pictures, and repair details through the local redraw module.

Benefits of technology

It significantly improves the controllability and efficiency of generated content, reduces the cost of manual adjustment, and is suitable for the automated generation of comic creation and plot commentary videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339457A_ABST
    Figure CN120339457A_ABST
Patent Text Reader

Abstract

The invention discloses an AI-based text cartoon device and method, and the method comprises the steps: extracting image parameters through a script analysis module, combining with the training of a LoRA model and a Stable Diffusion frame, achieving the multi-dimensional control of the standing position, action and scene style of a character through a Regional Prompter plug-in, a ControlNet plug-in and an Adtailer plug-in, and generating a key frame picture with high consistency; and the local redrawing module repairs details through segmentation and Inpaining technologies, and finally synthesizes a cartoon or a video. According to the method, the problems of inconsistent roles and incoherent scenes of pictures generated by an existing AI tool are solved, the controllability and efficiency of the generated content are remarkably improved, the manual adjustment cost is reduced, and the method is suitable for automatic generation of cartoon creation and story explanation videos.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of artificial intelligence and image generation, and particularly relates to an AI-based text-to-comic device and method for achieving consistency in characters, actions, and scenes in comic or video production through controllable image generation technology. Background Art

[0002] With the development of artificial intelligence technology, the application of AI in the field of image generation is becoming increasingly widespread. Although existing AI text-to-comic tools can generate comic pictures with certain artistic effects, they have defects in the controllability of the generated pictures. Specifically, it is manifested as follows:

[0003] Inconsistency between character images and scenes: Due to the randomness of the generation model, it is difficult to maintain consistency in character images and scenes in different pictures. For example, the facial features, clothing details, or body types of the same character vary in different pictures, and background elements (such as architectural styles, lighting effects) cannot be unified in consecutive frames. This results in poor coherence of comics or videos.

[0004] Difficulty in controlling actions and expressions: Existing AI tools have difficulty precisely controlling the actions and expressions of each character when generating multi-character scenes. For example, when generating a multi-person interaction scene, the character positions, limb movements deviate greatly from the description of the prompt words. This causes the generated pictures to not meet expectations.

[0005] Insufficient detail repair ability: The generated pictures are prone to breakdown in details, such as distortion in parts of the character's face, limbs, etc., affecting the viewing experience.

[0006] In the prior art, some solutions improve the generation effect by fine-tuning the pre-trained model, but there are problems of high training costs and poor flexibility; other solutions rely on manual post-correction, with low efficiency. Therefore, there is an urgent need for an AI text-to-comic technology that can achieve high controllability and ensure the coherence and consistency of comics or videos. Summary of the Invention

[0007] The present invention aims to provide an AI-based text-to-comic device and method, which solve the problems of poor controllability and insufficient front-back consistency of the generated pictures in the prior art by training a dedicated LoRA model and combining Stable Diffusion and its functional plugins, and significantly improve the efficiency and quality of comic or video production.

[0008] To achieve the above object, the present invention adopts the following technical solutions:

[0009] In a first aspect, the present application provides an AI-based text-to-comic device, including:

[0010] A script analysis module for analyzing the input comic script and extracting the character images, actions, and scene parameters required for each picture;

[0011] A model training module for training LoRA models of characters and scenes based on the output of the script analysis module and storing the trained models in a model library;

[0012] An image generation module integrating the Stable Diffusion framework, configured with a Regional Prompter plugin, an Adetailer plugin, and at least one ControlNet plugin, for controlling the character positions, actions, expressions, and scenes where they are located to generate key-frame pictures that meet the script requirements;

[0013] A local redrawing module that segments image regions through the Segment Anything plugin and calls the StableDiffusion Inpainting function to repair specified regions;

[0014] An output module for synthesizing the processed picture sequence into a comic or video and supporting the addition of text, subtitles, and multi-format export.

[0015] In a preferred embodiment, the model training module specifically includes:

[0016] A character model training unit for training LoRA models of each character to control the character images;

[0017] A scene model training unit for training LoRA models of each scene to control the scene style.

[0018] In a preferred embodiment, the image generation module supports directly calling pre-trained LoRA models.

[0019] In a preferred embodiment, the image generation module specifically includes:

[0020] A prompt control unit for introducing LoRA models in the prompts and setting weight coefficients to control the character images and scene styles;

[0021] A region control unit that divides the picture into multiple regions through the Regional Prompter plugin, and binds the standing position coordinates and action prompts of specific characters to each region;

[0022] A detail repair unit that automatically detects characters through the Adetailer plugin and repairs the facial and limb details of the characters;

[0023] Style control unit: Add reference images through the ControlNet Reference Only plugin and adjust the StyleFidelity parameter to control the consistency of the scene style;

[0024] Action control unit: Capture the human skeleton actions in the reference image through the ControlNet OpenPose plugin to control the poses and standing positions of the characters in the generated images.

[0025] In a preferred embodiment, the local redrawing module specifically includes:

[0026] Segmentation unit: Use the SegmentAnything plugin to segment the image area to be adjusted to generate a mask;

[0027] Redrawing unit: Call the Stable Diffusion Inpainting function to perform local redrawing on the masked area.

[0028] In a preferred embodiment, the output module supports the OCR text recognition function, automatically matches the script dialogue to the picture timeline and generates a subtitle file.

[0029] In a preferred embodiment, the output module supports export in MP4, GIF or PDF formats, and the adjustable resolution range is from 720p to 4K.

[0030] In a second aspect, the present application also provides an AI-based text-to-comic method, including the following steps:

[0031] Step S1: Analyze the comic script and plan the character images, actions and scene parameters of each picture;

[0032] Step S2: Train the LoRA model of the character and scene based on the parameters, or call the pre-trained LoRA model library;

[0033] Step S3: Use Stable Diffusion to generate key frame images, divide the picture area through the Regional Prompter plugin and bind the character standing positions, and control the character actions and scene styles through the ControlNet plugin;

[0034] Step S4: Segment and redraw the inconsistent local areas in the generated images to ensure the coherence of the character and scene details in the front and back pictures;

[0035] Step S5: Synthesize the corrected pictures into a comic or video in the script order and perform post-production.

[0036] In a preferred embodiment, the specific steps of the step S3 include the following steps:

[0037] When generating a multi-person scene, the Regional Prompter plug-in is used to divide the picture into multiple regions for partition control, controlling the standing positions and interaction actions of the characters in each region;

[0038] When generating full-body photos and long-distance photos, the Adetailer plug-in is used to automatically detect the characters and repair their expressions and limbs;

[0039] The ControlNet Reference Only plug-in is used to control the consistency of the picture scene style. Add a reference picture and adjust the Style Fidelity parameter to control the similarity between the generated picture style and the reference picture;

[0040] The ControlNet OpenPose plug-in is used to control the actions and standing positions of the characters in the picture.

[0041] In a preferred embodiment, in step S3, after generating the key-frame picture, the method further includes: automatically screening pictures through a consistency scoring rule, and the consistency scoring rule includes the matching degree of character costumes, the integrity of background elements, and the accuracy of actions.

[0042] In a preferred embodiment, step S4 specifically includes the following steps:

[0043] Use the SegmentAnything plug-in to segment the part of the picture that needs to be fine-tuned to generate a mask;

[0044] Use the Stable Diffusion Inpainting function to perform local redrawing on the part inside or outside the mask until the generated picture effect meets the expectations.

[0045] Compared with the prior art, the technical solution of the present invention has the following beneficial effects:

[0046] The present invention provides an AI-based text-to-comic device and method. The image parameters are extracted through a script parsing module, combined with the LoRA model training and the Stable Diffusion framework, and the Regional Prompter, ControlNet, and Adetailer plug-ins are used to achieve multi-dimensional control of the standing positions, actions, and scene styles of the characters, generating key-frame pictures with high consistency; the local redrawing module repairs the details through segmentation and Inpainting technologies, and finally synthesizes comics or videos. The present invention solves the problems of inconsistent characters and discontinuous scenes in the pictures generated by existing AI tools, significantly improves the controllability and efficiency of the generated content, reduces the manual adjustment cost, and is applicable to comic creation and the automated generation of plot explanation videos. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0048] Figure 1 is the structural block diagram of the AI-based text-to-comic device in the embodiment of the present invention;

[0049] Figure 2 is the flowchart of the AI-based text-to-comic method in the embodiment of the present invention. Specific Embodiments

[0050] In order to make the above and other features and advantages of the present invention clearer, the present invention will be further described below with reference to the drawings. It should be understood that the specific embodiments given herein are for the purpose of explaining to those skilled in the art and are merely exemplary, not restrictive.

[0051] It should be noted that the terms "first", "second", etc. in the specification, claims and drawings of the present invention are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0052] Embodiment 1:

[0053] As Figure 1 shown, this embodiment provides an AI-based text-to-comic device, including: a script parsing module, a model training module, an image generation module, a local redrawing module, and an output module. Each module will be introduced in detail below.

[0054] 1.1. Script Parsing Module

[0055] The script parsing module is used to parse the input comic script and extract the character images, actions, and scene parameters required for each picture.

[0056] Character Setting: According to the analysis result of the comic script, set the image characteristics of each character, such as hairstyle, clothing, facial features, etc.

[0057] Scene setting: Set the background features of each scene, such as location, time, weather, etc.

[0058] Action setting: Set the actions and expressions of each character in each scene to ensure that the actions and expressions meet the requirements of the script.

[0059] 1.2. Model training module

[0060] The model training module is used to train the LoRA models of characters and scenes according to the output of the script parsing module and store the trained models in the model library.

[0061] The purpose of training the LoRA model is to be able to more precisely control image generation. Only using prompts to shape the character and scene in the generated images has a large randomness and cannot guarantee the consistency of the character image and scene before and after. Of course, some ready-made LoRA models on the Internet can also be directly used.

[0062] Specifically, the model training module includes:

[0063] The character model training unit is used to train the LoRA models of each character to control the character image;

[0064] The scene model training unit is used to train the LoRA models of each scene to control the scene style.

[0065] 1.3. Image generation module

[0066] The image generation module integrates the Stable Diffusion framework and is configured with Regional Prompter plugin, Adetailer plugin and at least one ControlNet plugin. The following units cooperate to control the generation of key frame images:

[0067] (1) Prompt control unit, which is used to introduce the LoRA model in the prompt and set the weight coefficient to control the character image and scene style. For example, the prompt can include "Male character A stands in the park wearing red clothes".

[0068] (2) Region control unit, which divides the screen into multiple regions through plugins such as Regional Prompter, and binds the standing positions and action prompts of specific characters to each region.

[0069] The image generation accuracy is relatively low when drawing multi-character scenes. For example, when the prompt generates a picture of a man and a woman talking face to face, the number, gender, and standing positions of the characters in the generated picture probably do not meet expectations. By using plugins such as Regional Prompter, the picture is divided into multiple regions for partition control, and the standing positions and interaction actions of the characters in each region are controlled to improve the success rate of image generation. For example, "Left area: Character A raises his hand; Right area: Character B runs."

[0070] (3) Detail repair unit, which automatically detects and repairs the facial and limb details of the characters through the Adetailer plugin.

[0071] For example, when generating full-body photos or long-distance photos, the faces and limbs of the characters are prone to breakdowns. Use Adetailer to automatically detect the characters and repair their expressions and limbs.

[0072] (4) Style control unit, which adds reference pictures and adjusts the Style Fidelity parameter through the ControlNet Reference Only plugin to control the consistency of the scene style. The higher the coefficient, the closer the style of the generated picture is to the reference picture.

[0073] (5) Action control unit, which captures the skeleton actions of the characters in the reference picture through the ControlNet OpenPose plugin to control the poses and standing positions of the characters in the generated picture.

[0074] For example, during the image generation process, if a picture is obtained where the character actions meet the requirements but other aspects do not, this picture can be used as a reference picture. Through OpenPose, the positions of the characters in the picture are captured, and the poses and standing positions of the characters are restored in the generated picture.

[0075] 1.4. Local redrawing module

[0076] The local redrawing module further adjusts pictures with some defects. Specifically, it uses the SegmentAnything plugin to segment the image area and calls the Stable Diffusion Inpainting function to repair the specified area.

[0077] The local redrawing module specifically includes:

[0078] Segmentation unit, which uses the SegmentAnything plugin to segment the image area to be fine-tuned to generate a mask. For example, the tie part of the character can be separated.

[0079] The redrawing unit calls the Stable Diffusion Inpainting function to perform local redrawing on the part inside or outside the mask until the generated image effect meets the expectations. For example, only the tie part can be redrawn until the generated image effect meets the expectations.

[0080] 1.5. Output module

[0081] The output module is used to process, edit the processed image sequence, synthesize it into a comic or video, and add text or subtitles to produce a comic or a plot explanation video.

[0082] Specifically, it includes:

[0083] Image processing: Process the generated key-frame images, including operations such as color adjustment, sharpening, and blurring, to improve the image quality.

[0084] Editing: Edit the processed key-frame images in the order of the script to form continuous pictures.

[0085] Adding text or subtitles: Add text or subtitles at appropriate positions according to the script requirements to enhance the expressiveness of the pictures.

[0086] Video synthesis: Synthesize the edited image sequence into a video, and add elements such as background music and sound effects to produce a complete comic or plot explanation video.

[0087] Embodiment 2:

[0088] As Figure 2 shown, this embodiment provides an AI-based method for generating comics from text, including the following steps:

[0089] Step S1: Analyze the comic script and plan the character images, actions, and scene parameters for each picture.

[0090] Step S2: Train the LoRA model for the character and scene based on the parameters, or call the pre-trained LoRA model library.

[0091] Step S3: Use Stable Diffusion to control the character's standing position, actions, expressions, and the scene where the character is located to generate pictures, and screen out the key-frame pictures that are more in line with the expectations.

[0092] The specific implementation of Step S3 includes the following steps:

[0093] Step S3.1, introduce the LoRA model in the prompt and set the weight coefficient to control the character image and scene style.

[0094] Step S3.2: Use some functional plugins of Stable Diffusion to further enhance the controllability of the generated images, including:

[0095] When drawing multi-character scenes, the accuracy of the generated images is relatively low. For example, when the prompt generates a picture of a man and a woman talking face to face, the number, gender, and standing positions of the people in the generated picture mostly do not meet the expectations. Use plugins such as Regional Prompter to divide the picture into multiple regions for partition control, control the standing positions and interaction actions of the people in each region, and improve the success rate of image generation.

[0096] When generating full-body photos or long-distance photos, the faces and limbs of the people are prone to distortion. Use Adetailer to automatically detect the people and repair their expressions and limbs.

[0097] Use the ControlNet Reference Only plugin to control the consistency of the picture scene style. Add a reference picture to set the stylization parameter (Style Fidelity). The higher the coefficient, the closer the style of the generated picture is to the reference picture.

[0098] Use the ControlNet OpenPose plugin to control the actions and standing positions of the people in the picture. For example, if a picture is obtained during the image generation process where the actions of the people meet the requirements but other aspects do not, this picture can be used as a reference picture. Use OpenPose to capture the positions of the people in the picture and restore the postures and standing positions of the people in the generated picture.

[0099] Step S3.3: After steps S3.1 and S3.2, the controllability and efficiency of the generated images are greatly improved. Generate images through StableDiffusion and select key frame images that are more in line with expectations.

[0100] In a preferred embodiment, after generating the key frame images, the method may further include: automatically screening the images through a consistency scoring rule, and the consistency scoring rule includes the matching degree of character costumes, the integrity of background elements, and the accuracy of actions.

[0101] Step S4: Segment and redraw the inconsistent local areas in the generated images to ensure the coherence of the details of the people and the scene in the front and back images.

[0102] Make further adjustments to some pictures with partial defects, including:

[0103] Step S4.1: Use plugins such as SegmentAnything to segment the parts of the picture that need to be fine-tuned to generate a mask.

[0104] Step S4.2: Use the Stable Diffusion Inpainting function to perform local redrawing on the part inside or outside the mask until the generated image effect meets the expectations. For example, if the color of the character's tie is inconsistent before and after, for the image area of the tie in the separation area, only redraw the tie, and keep the other areas of the picture unchanged; in this way, the effect of replacing the picture background or replacing the character can also be achieved.

[0105] Step S5: Combine the corrected pictures into a comic or video in the order of the script and perform post-production.

[0106] After generating the key-frame pictures that meet the script, process and edit the pictures, and add text or subtitles to make a comic or a plot explanation video for subsequent promotion and operation.

[0107] To better understand the technical solution of the present invention, the following will be described in detail through specific cases.

[0108] Case example:

[0109] Suppose there is a comic script describing the story of two friends taking a walk in the park. The script is divided into three scenes:

[0110] Scene 1: Two friends meet at the park entrance.

[0111] Scene 2: Two friends take a walk on the path in the park.

[0112] Scene 3: Two friends rest on a bench in the park.

[0113] The following details the implementation process based on the technical solution of the present invention.

[0114] Step 1: Script planning and parameter extraction.

[0115] Input script: Parse the script text through the NLP module and extract the key elements of each scene:

[0116] 1) Scene 1:

[0117] Characters: Character A (male, blue sweatshirt, short hair), Character B (female, red dress, ponytail);

[0118] Actions: The two stand at the park entrance and smile at each other;

[0119] Background: There is an arched iron gate and a green plant flower bed at the park entrance.

[0120] 2) Scene 2:

[0121] Actions: Character A is in front and Character B is behind, walking along the path and talking;

[0122] Background: A stone path with cherry trees on both sides.

[0123] 3) Scene Three:

[0124] Action: Character A sits on the left side of the bench, and Character B sits on the right side of the bench, looking out at the lake together;

[0125] Background: A wooden bench, a lake and a setting sun in the distance.

[0126] Output parameters:

[0127] Character costumes, hairstyles, standing position coordinates (e.g., Character A is at 30% on the left side of the screen);

[0128] Keywords for background style (e.g., "cherry tree - pink petals", "setting sun - warm tone").

[0129] Step 2: Model training.

[0130] Collect a large amount of image data related to Character A and Character B, and preprocess the data. Then use a deep learning framework to train the LoRA model. During the training process, by adjusting the weight coefficients, the model can more accurately control the image features of Character A and Character B.

[0131] Step 3: Image generation.

[0132] Use the Stable Diffusion algorithm and various plugins to generate images for each scene:

[0133] 1) Generation of Scene One (Park entrance):

[0134] Prompt: "There are two friends at the park entrance. Character A is wearing a blue sweatshirt and smiling, and Character B is wearing a red dress and smiling. There is an iron arch, flower beds, and bright sunlight."

[0135] Use the Regional Prompter to divide the image into two regions, respectively controlling the positions and actions of Character A and Character B.

[0136] For example: Divide into left (30%), right (30%), and background (40%) regions; the left region is bound to "Character A stands on the left, wearing a blue sweatshirt"; the right region is bound to "Character B stands on the right, wearing a red dress".

[0137] Use Adetailer to repair the faces and limbs of Character A and Character B.

[0138] Use the ControlNet Reference Only plugin to control the background style to be the park entrance.

[0139] Use the ControlNet OpenPose plugin to control the actions and positions of Character A and Character B. For example, import the "standing and looking at each other" skeleton diagram.

[0140] 2) Generation of Scenario Two (Cherry Blossom Path):

[0141] Prompt: "Two friends are walking on a path. Character A is in front and Character B is behind. There are cherry trees on both sides, and pink petals are floating in the wind."

[0142] Use the Regional Prompter to divide the picture into two regions to control the positions and actions of Character A and Character B respectively.

[0143] For example: Divide into the front (40%), back (30%), and background (30%) regions; bind "Character A is in front" to the front region; bind "Character B is behind" to the back region.

[0144] Use Adetailer to repair the faces and limbs of Character A and Character B.

[0145] Use the ControlNet Reference Only plugin to use the background of Scenario One as a reference to maintain the consistency of the cherry blossom color tone.

[0146] Use the ControlNet OpenPose plugin to control the actions and positions of Character A and Character B. For example, bind the "walking pose" skeleton diagram to Character A and Character B respectively.

[0147] 3) Generation of Scenario Three (Resting on a Bench):

[0148] Prompt: "Two friends are sitting on a wooden bench. Character A is sitting on the left and Character B is sitting on the right. The two are looking at the lake as the sun is setting."

[0149] Use the Regional Prompter to divide the picture into two regions to control the positions and actions of Character A and Character B respectively.

[0150] For example: Divide into the left (40%), right (40%), and background (20%) regions; bind "Character A is sitting on the left, looking at the lake" to the left region; bind "Character B is sitting on the right, looking at the lake" to the right region.

[0151] Use Adetailer to repair the hand details of Character A and Character B (to avoid finger distortion).

[0152] Use the ControlNet OpenPose plugin to control the actions and positions of Character A and Character B. For example, import the "sitting and looking out" skeleton diagram and adjust the head angles of the characters.

[0153] Step 4: Local redrawing.

[0154] Perform local redrawing on the generated images:

[0155] 1) Correction for Scenario 1:

[0156] Problem: The sweatshirt color of Character A is too gray (needs to be adjusted to bright blue).

[0157] Operation: Use SegmentAnything to select the sweatshirt area and generate a mask; Inpainting prompt: "A vivid blue sweatshirt with a detailed zipper"; Denoising strength 0.3, enable "ColorTransfer" to maintain the original image's shadow.

[0158] 2) Correction for Scenario 2:

[0159] Problem: The position of Character B's arm is unnatural.

[0160] Operation: Use SegmentAnything to select the arm area and generate a mask; Inpainting prompt: "A natural arm pose, slightly bent"; After redrawing, use OpenPose to verify the action coherence.

[0161] 3) Correction for Scenario 3:

[0162] Problem: The reflection on the lake surface is too strong, obscuring the distant view.

[0163] Operation: Use SegmentAnything to select the lake area and generate a mask; Inpainting prompt: "A calm lake with a faint reflection of the setting sun"; Adjust the denoising strength to 0.5 and retain the outline of the distant mountain range.

[0164] Step 5: Output and composition.

[0165] Processing of the image sequence: Arrange the corrected key-frame images in the script order, and use image processing software to crop, resize, etc. the images to meet the layout requirements of the comic. Add transition effects between the images to enhance the coherence of the comic. For example, use "fade in and fade out" (duration 0.5 seconds) between scenes and use "pan shot" within the scene to simulate a dynamic perspective.

[0166] Subtitle addition: Recognize the script dialogue text through OCR and automatically match the time axis to generate a subtitle file.

[0167] In summary, the present invention provides an AI-based text-to-comic device and method. The image parameters are extracted by the script parsing module, combined with LoRA model training and the Stable Diffusion framework, and the Regional Prompter, ControlNet, and Adetailer plugins are used to achieve multi-dimensional control of character standing positions, actions, and scene styles, generating key frame images with high consistency. The local redrawing module repairs details through segmentation and Inpainting techniques, and finally synthesizes comics or videos. The present invention solves the problems of inconsistent character generation and discontinuous scenes in existing AI tools, significantly improves the controllability and efficiency of the generated content, reduces the manual adjustment cost, and is applicable to comic creation and the automated production of plot explanation videos.

[0168] The specific embodiments of the present invention have been described in detail above, but they are only examples, and the present invention is not limited to the specific embodiments described above. For those skilled in the art, any equivalent modifications and substitutions to the present invention are also within the scope of the present invention. Therefore, all equivalent transformations and modifications made without departing from the spirit and scope of the present invention should be covered within the scope of the present invention.

Claims

1. An AI-based text-to-comic device, characterized in that, Including: A script analysis module for analyzing the input comic script and extracting the character images, actions, and scene parameters required for each picture; A model training module for training LoRA models of characters and scenes based on the output of the script analysis module and storing the trained models in a model library; An image generation module integrating the Stable Diffusion framework, configured with a Regional Prompter plugin, an Adetailer plugin, and at least one ControlNet plugin, for controlling the character's standing position, actions, expressions, and the scene where they are located to generate key-frame pictures that meet the script requirements; A local redrawing module that segments image regions through the Segment Anything plugin and calls the Stable Diffusion Inpainting function to repair the specified regions; An output module for synthesizing the processed picture sequence into a comic or video and supporting the addition of text, subtitles, and multi-format export.

2. The AI-based text-to-comic device according to claim 1, characterized in that, The model training module specifically includes: A character model training unit for training LoRA models of each character to control the character image; A scene model training unit for training LoRA models of each scene to control the scene style.

3. The AI-based text-to-comic device according to claim 1, wherein The image generation module specifically includes: A prompt control unit for introducing LoRA models into the prompts and setting weight coefficients to control the character image and scene style; A region control unit that divides the picture into multiple regions through the Regional Prompter plugin, and binds the standing position coordinates and action prompts of specific characters to each region; A detail repair unit that automatically detects characters through the Adetailer plugin and repairs the facial and limb details of the characters; A style control unit that adds reference pictures and adjusts the StyleFidelity parameter through the ControlNet Reference Only plugin to control the consistency of the scene style; An action control unit that captures the skeleton actions of the characters in the reference picture through the ControlNet OpenPose plugin to control the poses and standing positions of the characters in the generated pictures.

4. The AI-based text-to-comic device according to claim 1, characterized in that, The local redrawing module specifically includes: A segmentation unit that uses the Segment Anything plugin to segment the image regions to be adjusted to generate masks; A redrawing unit that calls the Stable Diffusion Inpainting function to perform local redrawing on the masked regions.

5. The AI-based text-to-comic device according to claim 1, wherein The output module supports OCR text recognition function, automatically matches the script dialogues to the picture timeline, and generates subtitle files.

6. The AI-based text-to-comic device according to claim 1, wherein, The output module supports export in MP4, GIF, or PDF formats, and the adjustable resolution range is from 720p to 4K.

7. An AI-based method for generating comics from text, characterized in that, Including the following steps: Step S1: Analyze the comic script and plan the character images, actions, and scene parameters for each picture; Step S2: Train LoRA models of characters and scenes based on the parameters, or call a pre-trained LoRA model library; Step S3: Use Stable Diffusion to generate key-frame images, divide the picture area through the Regional Prompter plug-in and bind the character positions, and control the character actions and scene styles through the ControlNet plug-in; Step S4: Segment and redraw the inconsistent local areas in the generated images to ensure the coherence of the character and scene details in the front and back images; Step S5: Combine the corrected images into a comic or video in the script order and perform post-production.

8. The method for generating comics from text based on AI according to claim 7, characterized in that, The specific steps of the said Step S3 are as follows: When generating a multi-person scene, divide the picture into multiple areas through the Regional Prompter plug-in for partition control, and control the character positions and interaction actions in each area; When generating full-body pictures and long-shot pictures, automatically detect the characters through the Adetailer plug-in and repair their expressions and limbs; Control the consistency of the picture scene style through the ControlNet Reference Only plug-in, add a reference picture, and adjust the Style Fidelity parameter to control the similarity between the generated picture style and the reference picture; Control the actions and positions of the characters in the picture through the ControlNet OpenPose plug-in.

9. The method for generating comics from text based on AI according to claim 7, characterized in that, In the said Step S3, after generating the key-frame images, the method further includes: Automatically screen the images through the consistency scoring rules, and the consistency scoring rules include the matching degree of character costumes, the integrity of background elements, and the accuracy of actions.

10. A method for generating comics from text based on AI according to claim 7, characterized in that, The specific steps of the said Step S4 are as follows: Use the SegmentAnything plug-in to segment the parts that need to be fine-tuned in the picture to generate a mask; Perform local redrawing on the part inside or outside the mask through the Stable Diffusion Inpainting function until the generated picture effect meets the expectations.

Citation Information

Cited By

  • An AI cartoon production pipeline construction method based on multi-agent cooperation

    CN122597595A