AIGC-based idiom story picture book generation method and system
By pre-processing and post-processing of idiom story text and picture data, establishing and fine-tuning the literary and biographical diagram model, combining the prompt word engineering and the Lora model, the problems of logical errors, lack of details and limited cultural element generation capabilities in picture book creation in the existing technology are solved, and high-quality picture book generation that conforms to Chinese cultural characteristics is achieved.
Patent Information
- Application Number
- CN202510283969.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-27
AI Technical Summary
The existing generative artificial intelligence technology has logical errors, lack of details, and limited understanding and generation ability of Chinese traditional cultural elements in picture book creation, making it difficult to accurately generate elemental content with Chinese cultural characteristics.
By collecting and preprocessing the idiom story text, filtering the target elements, and post-processing their related image data, we obtain the training tag set. Create a literary graph model and use the acquired training labels to fine-tune the model. Introduce prompt word engineering, use large language models to create scripts, and select appropriate Lora models to load according to the script content to generate images that match the script content.
It significantly improves the logic, consistency and richness of the generated picture book story text and images, and can more accurately generate elemental content with Chinese cultural characteristics.
Smart Images

Figure CN120218239A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and specifically to a method and system for generating idiom story picture books based on AIGC. Background Art
[0002] Idiom stories have rich historical backgrounds and educational significance. Traditional creation of idiom story picture books requires the author to have high literary qualities and painting skills, and manually design the plot, draw the pictures and write the text. The creation process is cumbersome and the creation cycle is long, making it difficult to meet the need for quickly generating personalized picture books.
[0003] With the rapid development of generative artificial intelligence AIGC (Artificial Intelligence Generated Content), it has become possible to use AI technology to assist or even automatically generate picture book content. Although the existing generative artificial intelligence technology shows great potential in the field of picture book creation, there are still some deficiencies. For example, the generated text and images may have logical errors and lack of details, making it difficult to fully meet the creator's intentions, and the understanding and generation ability of traditional Chinese cultural elements is limited, making it difficult to accurately generate element content with Chinese cultural characteristics, etc.
[0004] Therefore, there is an urgent need for a method and system for generating idiom story picture books based on AIGC to solve the above problems. Summary of the Invention
[0005] The purpose of the present invention is to provide a method and system for generating idiom story picture books based on AIGC, which provides an efficient and convenient solution for the creation of idiom story picture books and improves the quality and controllability of the generated content.
[0006] To achieve the above object, the present invention is realized through the following technical solutions:
[0007] On the one hand, the present invention provides a method for generating idiom story picture books based on AIGC, including the following steps:
[0008] S1: Collect and preprocess the idiom story text, and screen out the target elements;
[0009] S2: Post-process the target elements and their related picture data screened out in step S1 to obtain a training label set;
[0010] S3: Establish a text-to-image model, and use the training labels obtained in step S2 to fine-tune the text-to-image model;
[0011] S4: Perform inference on the text-to-image model established in step S3.
[0012] Preferably, in step S1, preprocessing the idiom story text includes the following steps:
[0013] Clean the data of the collected idiom story text, removing irrelevant characters, punctuation marks, and spaces;
[0014] Use a word segmentation tool to perform word segmentation on the cleaned text.
[0015] Preferably, screening the target elements includes: performing a named ontology recognition operation on the text after word segmentation, specifically:
[0016] Select the spacy library to perform a named entity recognition library on all idiom story documents D = {d1, d2, …, d n}, to obtain an entity set E = {e1, e2, …, e n}.
[0017] Preferably, count the entity word frequencies in the entity set E, specifically:
[0018] Use the Count() method to count the word frequencies of all recognized named entities, and count the number of times each entity appears in the dataset:
[0019] f(e i ) = ∑ d∈D ∑ t∈d I(t = e i );
[0020] Among them, I(t = e i ) is an exponential function, when the word t is equal to the entity e i , the value is 1, otherwise it is 0;
[0021] According to the word frequency statistics results, filter out entity names with low frequencies and irrelevant to the camera switching, and mark the remaining entity names as elements.
[0022] Preferably, step S2 includes the following steps:
[0023] S21: Screen the elements obtained in step S1, specifically:
[0024] Use a text-to-image model to perform a generation test on the elements, and according to the generation results, identify the problem elements that occur, and store the problem elements in the target element list of the dataset U;
[0025] S22: Collect data on the elements that cannot be correctly generated by the text-to-image model, specifically:
[0026] Use web crawler technology to crawl images of the target elements from various angles on the platform, and the platform includes but is not limited to: image websites and museum websites;
[0027] S23: Post-process the collected image data, specifically:
[0028] Filter the collected images, screening out images with a file size less than 100k and low resolution;
[0029] Use the Codeformer algorithm to perform high-definition reconstruction of the resolution of the remaining images;
[0030] Perform scaling processing on the images after high-definition reconstruction, and the resolution of the scaled images is less than 1024*1024;
[0031] S24: Annotate the image data, specifically:
[0032] Use labels to annotate the images in the dataset in sequence, generate a batch of label files, use the BooruDatasetTagManager tool to manage the annotated labels, and delete incorrect labels.
[0033] Preferably, the step S3 includes the following steps:
[0034] S31: Adjust the format of the label set obtained in step S2, specifically:
[0035] Store the folders of multiple labels in the same-level folder directory, save the images and texts of each label in the folder corresponding to the label name, and train the images in each label using a set of trigger words of class identifier;
[0036] S32: Classify the label set, and the classification methods include but are not limited to: classification by dynasty, classification by character, classification by painting style;
[0037] S33: Fine-tune the model using the classified label set data.
[0038] Preferably, the step S4 includes the following steps:
[0039] S41: Introduce prompt engineering and use a large language model to create a script;
[0040] S42: Use a large language model to generate a description of the characters in the story;
[0041] S43: Select the required Lora model to load according to the script content;
[0042] S44: Use the image generation model with text input, taking the description of the picture as the input to generate pictures.
[0043] On the other hand, a system based on the above-mentioned idiom story picture book generation method based on AIGC is also provided, including:
[0044] A data collection and preprocessing module, which is used for: collecting and preprocessing idiom story texts, and screening target elements;
[0045] A data post-processing module, which is used for: post-processing the screened target elements and their related picture data to obtain a training label set;
[0046] A model establishment and training module, which is used for: establishing a text-to-image model, and using the obtained training labels to finely tune and train the text-to-image model;
[0047] A model inference module. It is used for: inferring the established text-to-image model.
[0048] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0049] By introducing prompt engineering and fine-tuning training, the generated picture book story texts and images are significantly improved in terms of logic, consistency, and detail richness:
[0050] 1. Collect idiom stories, use tools to segment the text and count entities, and according to the word frequency statistics results, filter out entity concepts with low occurrence frequencies and irrelevant to the camera shots, and retain the entity names that finally need to be shown in the pictures;
[0051] 2. In the present invention, through text-to-image model generation tests, a data set is made for elements with poor generation effects. Use web crawler technology to collect pictures, use super-resolution algorithms to improve clarity, and use tools to manage and correct label content to provide a high-quality label set for fine-tuning training;
[0052] 3. Process the label set into a format suitable for fine-tuning, and classify the data set to meet the picture performance requirements of different idiom stories. Use the classified training data to finely tune the large text-to-image model to provide high-quality pictures for the scenes of idiom stories;
[0053] 4. Introduce prompt engineering, input the designed prompt template into the large language model to generate an idiom story script and character images that meet the requirements, and select a suitable Lora model to load. Use the generated picture description as the input, and dynamically adjust the details in the prompt until the generated image conforms to the script content. Finally, the generated idiom story pictures will be accompanied by Chinese and English captions to form a complete story picture book. Brief Description of the Drawings
[0054] Figure 1 is the method flow chart of the present invention;
[0055] Figure 2It is the flow chart for processing the idiom story text data of the present invention;
[0056] Figure 3 It is the flow chart for entity screening of the present invention;
[0057] Figure 4 It is the schematic diagram of the dataset construction process of the present invention;
[0058] Figure 5 It is the schematic diagram of the matrix for training the Lora model of the present invention;
[0059] Figure 6 It is the flow chart for text-to-image inference of the present invention;
[0060] Figure 7 It is the schematic diagram of the system structure of the present invention. Detailed implementation manners
[0061] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by this application.
[0062] In the present invention, terms such as "upper", "lower", "left", "right", "front", "rear", "vertical", "horizontal", "side", "bottom", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings. They are only relationship terms determined for the convenience of describing the structural relationship of each component or element of the present invention, and do not specifically refer to any component or element in the present invention, and should not be construed as a limitation to the present invention.
[0063] In the present invention, terms such as "fixed connection", "connected", "connected" should be understood in a broad sense, which may mean a fixed connection, an integral connection or a detachable connection; it may be directly connected or indirectly connected through an intermediate medium. For those relevant scientific research or technical personnel in the field, the specific meanings of the above terms in the present invention can be determined according to specific circumstances, and should not be construed as a limitation to the present invention.
[0064] Embodiment:
[0065] As Figure 1 shown, this embodiment provides an AIGC-based idiom story picture book generation method, including the following steps:
[0066] S1: Collect and preprocess the idiom story text, and screen the target elements;
[0067] S2: Post-process the target elements and their related picture data screened in step S1 to obtain a training label set;
[0068] S3: Establish a text-to-image model, and use the training labels obtained in step S2 to fine-tune the text-to-image model;
[0069] S4: Perform inference on the text-to-image model established in step S3.
[0070] According to the collection and preprocessing of the idiom story text described in step S1, filter the target elements. The specific details are as Figure 2 shown, mainly including the following data processing steps:
[0071] 1. Collection of idiom stories, specifically:
[0072] Preliminarily collect the original data of 1,000 idiom stories, and continuously expand the data scale according to actual needs. Use web crawler technology to crawl the original data of idiom stories from platforms such as idiom story websites and online ancient book libraries, collect and organize paper idiom story books, and use technologies such as OCR (Optical Character Recognition) to convert scanned paper documents into electronic texts for digital processing;
[0073] 2. Preprocessing of text data, specifically:
[0074] First, perform text cleaning. Clean the collected idiom story data to remove irrelevant characters, punctuation marks, spaces, etc. Then use word segmentation tools such as Jieba to perform word segmentation on the cleaned text to prepare for subsequent named entity recognition;
[0075] 3. Named entity recognition, specifically:
[0076] Select the spacy library to perform named entity recognition on all idiom story documents D = {d1, d2,..., d n}. The spacy library provides efficient and accurate Chinese named entity recognition functions. In this embodiment, we focus on named entities of people and items, such as historical figures like Confucius and Mencius, and ancient artifacts like bamboo slips and jade seals. Finally, obtain the entity set E = {e1, e2,..., e n};
[0077] 4. Entity word frequency statistics, specifically:
[0078] Use the Count() method to perform word frequency statistics on all recognized named entities, and count the number of times each entity appears in the dataset:
[0079] f(e i ) = ∑ d∈D ∑ t∈d I(t = e i );
[0080] Where, I(t=e i ) is the indicator function, when word t is equal to entity e i The value is 1, otherwise it is 0; according to the word frequency statistics, filter out entity concepts with low frequency and irrelevant to the camera image, such as abstract concepts, emotional words, etc., and retain the entity names that ultimately need to be expressed on the screen. The obtained entity names are elements.
[0081] Post-process the image data related to the target element as described in step S2 to obtain a training label set. The specific details are as follows: Figure 3 , Figure 4 As shown, the data set preparation steps mainly include the following:
[0082] 1. Filter the retained entity names (elements), specifically:
[0083] Use the text graph model to generate test on the retained entity names (elements), observe the generation results, identify elements with poor generation effect or that cannot be generated normally, such as bamboo slips and carriages, and include the identified problematic elements in the target element list of the special data set to be produced;
[0084] 2. Collect data for elements that cannot be correctly generated by the Wensheng graph model, specifically:
[0085] Use crawler technology to crawl pictures of target elements from various angles from picture websites, museum websites and other platforms. For the rare and difficult-to-obtain element pictures on the Internet, take them yourself or entrust professional organizations to take them. Finally, for each target element, collect 50-100 pictures to ensure the diversity and representativeness of the data set;
[0086] 3. Post-process the collected image data, specifically:
[0087] First, the collected images are filtered to remove images with a file size less than 100k and low resolution to ensure image quality. Then, the Codeformer algorithm is used to reconstruct the images in high resolution to improve the clarity and detail of the images. Finally, the images are scaled to control the image resolution within 1024*1024 to meet the training specifications required by the model.
[0088] 4. Data annotation, specifically:
[0089] With the help of the image tagging model developed by the developer "SmilingWolf", the images in the dataset are tagged one by one to generate a batch of tag files. Then, the BooruDatasetTagManager tool is used to manage the tagged tags, and incorrect tags are deleted. The tag content should include information such as the name, attributes, and features of the target element to ensure the accuracy of the tags.
[0090] Establish a text-to-image model according to the description in step S3, and use the training tags obtained in step S2 to fine-tune the text-to-image model. The specific process is as Figure 5 shown, which mainly includes the following data processing steps:
[0091] 1. Process the obtained tag set into a suitable fine-tuning format, specifically:
[0092] For the tag set obtained in the previous step, the folders of multiple tags are in the same-level folder directory. The images and texts of each tag are saved in the folder corresponding to the element name. Each image is trained using a set of trigger words of class identifier. Class is the general category of the training target, and identifier is the word used to identify the training target and learn (such as the bamboo slips of ancient Chinese items). The model can be trained to identify and learn the required "bamboo slips" target from "ancient Chinese items". When generating images, using "ancient Chinese item bamboo slips" will generate images containing bamboo slip elements learned.
[0093] Since the number of image data for each element ranges from 50 to 100, it is necessary to set the repetition times of each training element. Generally, each element needs to be repeated 5 to 8 times, and finally the ratio of each element during training reaches a balance. The total training data volume is "the repetition times of training images × the number of training images". The sample balance formula is as follows:
[0094]
[0095] Among them, C is the total number of element categories ("bamboo slips", "jade seal", etc.), r c ∈[5, 8] is the repetition times of category c, N c is the number of original images of category c (50 - 100). By adjusting r c to make achieve category balance;
[0096] 2. Classify the tag set, specifically:
[0097] Since the content of each idiom story varies greatly, in order to meet the needs of the visual representation of each idiom story, we classify the ancient Chinese elements collected. The classification methods are:
[0098] (1) Classify by dynasty. For example, the items related to clothing, food, housing, and transportation in the Tang Dynasty, Song Dynasty, Spring and Autumn Period, Warring States Period, etc. are grouped into a training set, which can ensure the accuracy of the item representation in the dynasties where the idiom stories are set.
[0099] (2) Classify by figure. For example, famous historical figures such as Confucius, Mencius, and Dayu. Since such historical celebrities have relatively definite IP images, for the related figures with higher frequencies of appearance, we train a unique character Lora model independently to ensure the consistency of the figures in the series of stories. For the figures with lower frequencies of appearance, we use the model to generate the character images automatically.
[0100] (3) Classify by painting style. For example, ink painting style, meticulous painting style, Q-version painting style, etc. Different types of painting styles are needed to produce picture books for different target audiences such as children's books and teenagers' books.
[0101] 3. Fine-tune the model using the classified training data of ancient Chinese elements. Specifically:
[0102] Use the classified training data of ancient Chinese elements obtained from the previous steps to fine-tune the text-to-image large model so that it can better generate shot images that match the picture descriptions. Fine-tuning can be achieved by using the data to perform additional training on the existing open-source text-to-image large model and adjusting the model parameters to improve its performance and adaptability. The full name of fine-tuning training is Parameter-Efficient Fine-tuning (PEFT), which means fine-tuning the base large model with a small number of parameters, which can save costs.
[0103] Commonly used parameter-efficient fine-tuning methods include:
[0104] (1) FREE, that is, parameter freezing. Freeze some parameters of the original base large model and only train some parameters to enable training of the base large model on a single card.
[0105] (2) LORA, freeze the pre-trained model weights, and inject the trainable rank decomposition matrix into each weight of the Transformer layer, greatly reducing the number of trainable parameters for downstream tasks.
[0106] (3) P-Tuning, add prompt parameters to each layer of the transformer for fine-tuning.
[0107] Through the above steps, the data of ancient Chinese elements can be processed into a training dataset that meets the format requirements for model fine-tuning. Then, the text-to-image large model is fine-tuned using the classified label set to obtain a LoRA model matrix of various dynasties, painting styles, and characters. Such a process can endow the text-to-image large model with the ability to accurately generate images of ancient Chinese elements and provide more accurate and higher-quality shot images for idiom story scenes.
[0108] Perform inference on the text-to-image model established in step S3 according to the description in step S4. The specific details are as Figure 6 shown, mainly including the following steps:
[0109] 1. Introduce prompt engineering and use a large language model to create a script. Specifically:
[0110] Input the designed prompt template into the large language model to generate an idiom story script that meets the requirements. An example of the prompt template for the idiom story script is as follows:
[0111] " 'Idiom story: During the Spring and Autumn Period, Duke Huan of Qi, one of the Five Hegemons, loved to wear purple clothes, and later the ministers followed suit...';
[0112] The above is the story of the Chinese idiom 'If the upper classes behave a certain way, the lower classes will follow suit'. Adapt it into a picture book suitable for preschool children with the following requirements:
[0113] (1) A description of 15 pictures, and be sure to use the rules of AIGC text-to-image prompts to describe the picture content in detail;
[0114] (2) The number of characters in the picture book cannot exceed 2;
[0115] (3) If the protagonist appears in the picture, start with ["Protagonist's name"];
[0116] (4) Camera perspectives, such as: close-up, medium shot, long shot, extreme long shot, wide angle, aerial shot, etc.;
[0117] (5) The narration corresponding to 15 pictures;
[0118] (6) Please output in the following format:
[0119] 'Match Figure 1 ':
[0120] 'Picture description':
[0121] 'Camera perspective':
[0122] 'Chinese narration':
[0123] 'English narration': ";
[0124] The example script generated by the large language model is as follows:
[0125] “‘Match Figure 1 ’:
[0126] ‘Scene description’ ["Duke Huan of Qi"] is sitting inside the palace, surrounded by magnificent decorations. The sunlight shines through the window onto the carpet. Duke Huan of Qi is wearing a purple robe and has a satisfied smile on his face.
[0127] ‘Shot perspective’ Medium shot,
[0128] ‘Chinese narration’ Once upon a time, there was a king named Duke Huan of Qi who especially liked purple clothes.
[0129] ‘English narration’ Once upon a time, there was a king named Duke Huan of Qi, who loved wearing purple clothes.
[0130] ‘Match Figure 2 ’:
[0131] ‘Scene description’ The ministers are standing outside the palace, each wearing purple clothes. They are about to enter the palace for the morning court.
[0132] ‘Shot perspective’ Panoramic view,
[0133] ‘Chinese narration’ The ministers saw that the king liked purple, so they also started wearing purple clothes.
[0134] ‘English narration’ Seeing the king's preference, his ministers also started to wear purple clothes.”;
[0135] Conduct a manual review and optimization of the generated script to ensure its logic, interestingness, and educational significance;
[0136] 2. Use the large language model to generate character descriptions in the story. Specifically:
[0137] Input the designed character description prompt template into the large language model to generate characters that meet the story requirements. An example of the character description prompt template is as follows:
[0138] “Output the character settings based on the above content: Briefly describe the character's gender, facial features, clothing color, hairstyle and other features. The format is as follows: ["Protagonist's name"]: ”;
[0139] An example of the character generated by the large language model:
[0140] "Qi Huan Gong", male, with distinct facial features, short beard, deep-set eyes, wearing traditional purple clothing and a ceremonial cap symbolizing his status, and having a neat hairstyle.
[0141] Through these steps, it can be ensured that the generated character descriptions not only conform to the background and plot of the story but also provide sufficient details for the image generation model to generate images that meet expectations.
[0142] 3. Select and load the required Lora model according to the script content. Specifically:
[0143] Analyze the characters, plot, scene descriptions, etc. in the script, extract the required visual styles and detailed requirements, and select the appropriate Lora for loading according to the requirements. For example, in the content of the idiom story "Those above set the example and those below follow suit", the main character "Qi Huan Gong" is not a frequently appearing character in the story. Therefore, we use a language model to shape the character image and do not need to use a character-specific Lora model. Since Qi Huan Gong lived in the State of Qi during the Spring and Autumn period in the story, we load the Lora of the Spring and Autumn period. All Lora models and the text-to-image large model are loaded into the video memory.
[0144] 4. Use the text-to-image model to generate images with the scene description as the input. Specifically:
[0145] According to the analysis results, select the appropriate Lora model, load the selected Lora model into the generation process, and adjust the prompt words for the visual performance of the scene according to the script, such as details of the character's facial expression, posture, clothing color, actions, etc. In this process, the prompt words and the weight parameters of the Lora model can be continuously and dynamically adjusted according to different scenes and plot changes until the generated image fits the script content. Finally, the generated idiom story images are paired with corresponding Chinese and English captions and output as a story picture book.
[0146] As Figure 7 shown, this embodiment also provides a system based on the above-mentioned AIGC-based idiom story picture book generation method, including:
[0147] A data collection and preprocessing module for collecting and preprocessing idiom story texts and screening target elements;
[0148] A data post-processing module for post-processing the selected target elements and their related picture data to obtain a training label set;
[0149] A model establishment and training module for establishing a text-to-image model and fine-tuning and training the text-to-image model using the obtained training labels;
[0150] Model thrust module. Function: perform inference on the established text-to-image model.
[0151] The above is a specific description of the preferred embodiment of the present invention. However, the present invention is not limited to the described embodiment. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present invention. These equivalent deformations or substitutions are all included within the scope defined by the claims of this application.
Claims
1. A method for generating idiom story picture books based on AIGC, characterized in that: The following steps are involved: S1: Collect and preprocess idiom story texts and filter target elements; S2: Post-process the target elements and their related image data screened out in step S1 to obtain a training label set; S3: Establish a document graph model, and use the training labels obtained in step S2 to fine-tune the document graph model; S4: Perform reasoning on the text graph model established in step S3.
2. The method for generating an idiom story picture book based on AIGC according to claim 1, characterized in that: In step S1, the idiom story text is preprocessed, including the following steps: The collected idiom story texts were cleaned to remove irrelevant characters, punctuation marks, and spaces; Use the word segmentation tool to segment the cleaned text.
3. The method for generating an idiom story picture book based on AIGC according to claim 2, characterized in that: The screening of target elements includes: performing a naming ontology recognition operation on the text after word segmentation processing, specifically: Select spacy library for all idiom story documents D = {d1, d2, ..., d n } to perform named entity recognition and obtain the entity set E = {e1, e2, ..., e n }.
4. The method for generating an idiom story picture book based on AIGC according to claim 3, characterized in that: The entity word frequency in the entity set E is counted, specifically: Use the Count() method to perform word frequency statistics on all identified named entities and count the number of times each entity appears in the dataset: f(e i )=∑ d∈D ∑ t∈d I(t=e i ); Where I(t=e i ) is an exponential function, when word t is equal to entity e i When , the value is 1, otherwise it is 0; According to the word frequency statistics, entity names with low frequency of occurrence and irrelevant to the shot change are filtered out, and the remaining entity names are marked as elements.
5. The method for generating idiom story picture books based on AIGC according to claim 1, characterized in that: The step S2 comprises the following steps: S21: Screen the elements obtained in step S1, specifically: Use the document graph model to generate test elements, identify the elements with generation problems based on the generation results, and store the problem elements in the target element list of the dataset U; S22: Collect data for elements that cannot be correctly generated by the text graph model, specifically: Use crawler technology to crawl pictures of target elements from various angles from platforms, including but not limited to: picture websites and museum websites; S23: Post-processing the collected image data, specifically: Filter the collected images and filter out those with a file size less than 100k and low resolution; Use the Codeformer algorithm to reconstruct the remaining images at high resolution; The high-definition reconstructed image is scaled, and the resolution of the scaled image is less than 1024*1024; S24: Label the image data, specifically: Use tags to label the images in the dataset one by one, generate batch tag files, use the BooruDatasetTagManager tool to manage the labeled tags and delete incorrect tags.
6. The method for generating idiom story picture books based on AIGC according to claim 1, characterized in that: The step S3 comprises the following steps: S31: Adjust the format of the tag set obtained in step S2, specifically: Store folders of multiple labels in the same folder directory, save the images and texts of each label in the folder corresponding to the label name, and use a set of trigger words of the class identifier to train the images in each label; S32: classifying the data of the tag set, wherein the classification method includes but is not limited to: classification by dynasty, classification by character, and classification by painting style; S33: Fine-tune the model using the classified label set data.
7. The method for generating idiom story picture books based on AIGC according to claim 1, characterized in that: The step S4 comprises the following steps: S41: Introducing the prompt word project and using a large language model to create scripts; S42: Generate character descriptions in stories using large language models; S43: Select the required Lora model to load according to the script content; S44: Take the picture description as input and generate the picture using the text-to-picture model.
8. A system based on the AIGC-based idiom story picture book generation method as claimed in claim 1, characterized in that: include: The data collection and preprocessing module is used to collect and preprocess the idiom story texts and filter the target elements; The data post-processing module is used to post-process the screened target elements and their related image data to obtain a training label set; The model building and training module is used to: build a document graph model and use the acquired training labels to fine-tune the document graph model; The model reasoning module is used to: reason about the established text graph model.
Citation Information
Patent Citations
Keyword-based child picture book story generation method, system and device
CN110941960A
Image automatic generation method and device based on AIGC, equipment and medium
CN117496302A
Illustration generation method and device, equipment and storage medium
CN117689751A
Electronic picture book generation method and device, electronic equipment and storage medium
CN119399320A
Evaluating System for Capability of Biometric Terminal to Detect Fake Biometric and Evaluating Method Thereof
KR102834663B1
Cited By
AIGC-based sports explanation generation method
CN122153804A