Story image generation method and system with inference chain and electronic equipment
By constructing a method that combines a storyteller model and an image rendering model, multi-semantic dimension story images are generated and evaluated. This solves the problem of lack of creativity and logic in the generation of single story images in existing technologies, and realizes the automated production of high-quality story images.
Patent Information
- Application Number
- CN202511777730.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-28
- Publication Date
- 2026-02-27
AI Technical Summary
Existing technologies cannot generate single semantically rich story images, lack narrative logic and creativity, and lack effective evaluation criteria.
A storyteller model is constructed using a large language model to generate image story corpora with inference chains, and a multi-semantic dimension story image is generated through an image rendering model. The story image is then evaluated using a story image evaluator.
It enables automated batch production of single semantically rich story images, solving the problem of lack of creativity and logic in existing technologies and providing an effective evaluation standard.
Smart Images

Figure CN121582377A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multimodal artificial intelligence, and more particularly to a method, system, and electronic device for generating story images with reasoning chains. Background Technology
[0002] An image can convey a compelling story by presenting rich and logically connected visual cues. The relationships between these cues within the story form a Chain-of-Reasoning (CoR) within the image, enabling the viewer to infer events, causal relationships, and other information, thereby understanding the narrative behind the image. Such semantically rich images are defined as storytelling images. Figure 1 As shown, this is a classic example of an image with a rich semantic reasoning chain, telling the story of two children secretly taking cookies from the cupboard while their mother is washing dishes. Due to the rich semantics of such images, they are widely used in picture description tasks to assess human cognitive and linguistic abilities. These types of images can be used not only for illustration and cognitive assessment but also, thanks to their ability to visually convey multi-layered information and stimulate active understanding, for a wider range of applications. Similar existing techniques for generating story images include: 1. Story Visualization: Story visualization is a technique that aims to generate a coherent sequence of images based on a series of chronologically arranged textual descriptions of a story. Each image corresponds to a step or scene in the story.
[0003] 2. T2I (Text-to-Image, general text-to-image generation): Examples include DALL-E and StableDiffusion models. Their core function is to receive a single text prompt from the user and generate a matching image.
[0004] In the process of realizing this invention, the inventors discovered at least the following problems in the related technology: 1. Story Visualization: Inability to Generate a Single Semantically Rich Image: Its design goal is multi-image storytelling, resulting in the semantics of the story being scattered across the image sequence. The focus is on the temporal coherence between images, rather than the semantic complexity and logical depth within a single image. The generated single images are typically semantically simple, failing to "tell a story with a single image." Its technical architecture is designed to solve the problem of "how to tell a story with multiple images," with its optimization goal being the coherence of the sequence, not the semantic density of a single frame. This represents a different task generation direction compared to generating a story image with a single chain of reasoning.
[0005] 2. General T2I generation: Lack of "creativity" and "narrative logic": Such models are passive "renderers" that do not have the ability to actively conceive complex stories. Users must provide extremely detailed, narrative-logic-filled prompt words for the model to generate story pictures. Therefore, the creative bottleneck is the user. They cannot solve the core creative problem of "how to conceive a good story", but instead completely pass this most difficult burden to the user, "prompt engineering".
[0006] 3. Lack of evaluation criteria for story image generation: Since there is no clear task paradigm of "single story image generation" in existing research, the existing evaluation system mainly faces "image quality" or "text-image consistency" and does not pay attention to narrative semantic complexity and other elements. Therefore, the current technology lacks an evaluation index system for story image generation, which leads to a lack of quantifiable control benchmarks for research. SUMMARY
[0007] In order to at least solve the problems of being unable to generate a single semantic-rich image, lacking narrative logic, and poor ability to present complex logic in the prior art, in a first aspect, embodiments of the present application provide a story image generation method with a reasoning chain, comprising: generating an image story corpus with reasoning chains of roles' execution actions anchored at the same time point and with different semantic dimensions by a storyteller model constructed by a large language model, wherein the reasoning chains with different semantic dimensions are used to constitute multi-dimensional semantics of the image story corpus; inputting the image story corpus with the reasoning chains as prompt words to an image rendering model to generate a story image presenting a multi-semantic dimension narrative scene; evaluating the image story corpus and the story image based on a story image evaluator, and obtaining a semantic-rich story image after evaluation.
[0008] In a second aspect, embodiments of the present application provide a story image generation system with a reasoning chain, comprising: a storyteller module configured to generate an image story corpus with reasoning chains of roles' execution actions anchored at the same time point and with different semantic dimensions by a storyteller model constructed by a large language model, wherein the reasoning chains with different semantic dimensions are used to constitute multi-dimensional semantics of the image story corpus; a painter module configured to input the image story corpus with the reasoning chains as prompt words to an image rendering model to generate a story image presenting a multi-semantic dimension narrative scene; An evaluation module is configured to evaluate the image story corpus and the story image based on the story image evaluator, and obtain a semantic-rich story image after the evaluation.
[0009] In a third aspect, an electronic device is provided, which includes at least one processor and a memory connected to the at least one processor in communication, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the steps of the story image generation method with reasoning chains according to any one of the embodiments of the present application.
[0010] In a fourth aspect, an embodiment of the present application provides a storage medium having a computer program stored thereon, wherein the program is executed by a processor to implement the steps of the story image generation method with reasoning chains according to any one of the embodiments of the present application.
[0011] In a fifth aspect, an embodiment of the present application provides a computer program product, which includes a computer program / instruction, and the computer program / instruction is executed by a processor to implement the steps of the story image generation method with reasoning chains according to any one of the embodiments of the present application.
[0012] The method can automatically produce a standard rare "story image" in batches. The image contains rich semantics, multiple logical reasoning chains and complex narration in a single picture. The user does not need to think of complex prompt words, and the defect that the general image rendering model depends on the prompt word is solved. BRIEF DESCRIPTION OF DRAWINGS
[0013] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings described below are some embodiments of the present application, and those skilled in the art can obtain other drawings according to these drawings without creative labor.
[0014] Figure 1 is a schematic diagram of a story image "stealing cookies" created by an existing technology; Figure 2 is a flowchart of a story image generation method with reasoning chains according to an embodiment of the present application; Figure 3 is a generation architecture diagram of a story image generation method with reasoning chains according to an embodiment of the present application; Figure 4 is a summarizer prompt word schematic diagram in a KNN-based diversity evaluator of a story image generation method with reasoning chains according to an embodiment of the present application; Figure 5 is a prompt word diagram of the alignment score stage in the story image alignment evaluator of a story image generation method with a reasoning chain provided by an embodiment of the present application; Figure 6 is a model performance diagram of a story image generation method with a reasoning chain provided by an embodiment of the present application; Figure 7 is a distribution diagram of visual cues in seven dimensions of a story image generation method with a reasoning chain provided by an embodiment of the present application; Figure 8 is a story image alignment evaluation diagram of a painter model of a story image generation method with a reasoning chain provided by an embodiment of the present application; Figure 9 is a generated result comparison diagram of the same story under different image rendering models of a story image generation method with a reasoning chain provided by an embodiment of the present application; Figure 10 is a generated story and corresponding image diagram of a story image generation method with a reasoning chain provided by an embodiment of the present application; Figure 11 is an effectiveness diagram of a verification diversity evaluator of a story image generation method with a reasoning chain provided by an embodiment of the present application; Figure 12 is a Mini-Storyteller model performance diagram of a story image generation method with a reasoning chain provided by an embodiment of the present application; Figure 13 is a corresponding image diagram of a story generated by using a Mini-Storyteller model of a story image generation method with a reasoning chain provided by an embodiment of the present application; Figure 14 is a structure diagram of a story image generation system with a reasoning chain provided by an embodiment of the present application; Figure 15 is a structure diagram of an electronic device for a story image generation with a reasoning chain provided by an embodiment of the present application. DETAILED DESCRIPTION
[0015] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.
[0016] As Figure 2 Fig. 1 shows a flowchart of a story image generation method with reasoning chain according to an embodiment of the present application, including the following steps: S11: generating an image story corpus with reasoning chain of the execution action of the role anchored on the same time point and with different semantic dimensions by using a storyteller model constructed by a large language model, wherein the reasoning chain with different semantic dimensions is used to constitute the multi-dimensional semantics of the image story corpus; S12: inputting the image story corpus with reasoning chain as a prompt word to an image rendering model to generate a story image presenting a multi-semantic dimension narrative scene; S13: evaluating the image story corpus and the story image based on a story image evaluator, and obtaining a semantic-rich story image after evaluation.
[0017] The present method finds that high-quality story images are relatively scarce and difficult to create, and the field of artificial intelligence currently lacks related work focusing on the automatic generation of such images. Traditional story image creation mainly has two modes: either completely relying on professional painters to create, or simply giving the artificially conceived story outline to the T2I model for rendering. The core innovation of the present method is to realize the automatic "decoupling" generation of story images: the present method completely separates the two core links of "story creativity and logical reasoning" and "visual rendering". It introduces a new, independent and personified stage called "storyteller".
[0018] The method explicitly points out that the existing image rendering model is not good at "creativity" or "logical thinking" by nature, so the method finds that it should not be forced to play the role of such human intelligence. On the contrary, the method can use a powerful large language model to first perform creative reasoning and generate a detailed story text rich in CoR (Chain-of-Reasoning, reasoning chain) and precisely anchored at a single moment. At this time, the image rendering model only needs to play a passive "rendering brush" role and strictly execute this clear and detailed story text. The two-stage combination of "LLM (as a creative brain)" and "image rendering model (as a rendering brush)" aims to further improve the generation of story images with rich reasoning chains. To further improve the generation of story images with rich reasoning chains, the two-stage story is evaluated, and the "reasoning chain guided mode" (CoR-Guided Mode) implemented by guiding the LLM to achieve a rigorous logic provides a non-obvious system-level solution to the problem. Simply put, using an easily understandable anthropomorphic description, the method includes three parts: the first stage: Storyteller - to solve the "creativity" defect; the second stage: Painter - to solve the "rendering" defect; and the Evaluator - to solve the "quality evaluation" defect.
[0019] For step S11, to solve the "creativity" defect, the method uses a Storyteller model constructed by a large language model to automatically generate a narrative story. This does not rely on user input of complex prompt words, but uses a large language model to automatically generate a narrative story. As shown in Figure 3 In the first stage, the method uses the imagination and reasoning ability of the LLM to generate a story, and this part is called the Storyteller model. The key difference between this type of story and existing conventional narratives is the temporal structure. The story generated here by the method is anchored at the same time point, which means that each character in the story can only perform one action. In contrast, existing conventional stories usually unfold along a sequence of events, enabling characters to perform multiple actions in succession.
[0020] In a specific implementation, a semantically rich story image can be defined as an image that embeds a complex CoR. These reasoning chains logically connect various visual cues or premises in the image to construct a narrative. For a reasoning chain, let be a set of different visual cues identified in the image. Let K be a conclusion, and the CoR is defined by a logical structure, where the combination of these cues implies the conclusion K: where ∧ denotes a logical conjunction, meaning that all cues are considered simultaneously; represents a logical implication. This structure indicates that the combined presence of all visual cues can logically support the conclusion K.
[0021] In general, in the story image generation task: given a set of N generation instructions where a single instruction is not necessarily different from the others (i.e., where it is possible that N stories are first created, and then in the subsequent image generation step, N corresponding story images are generated. Each story should describe a dramatic, instantaneous scene that can be depicted in an image, rather than a long story with a time sequence or character dialogue. It should also be semantically as different as possible from all other stories in the set. Each generated image is expected to independently and accurately convey the story .
[0022] As an implementation, the semantic dimensions include: time, location, character identity, character relationship, event, event causality, and character mental state.
[0023] In this implementation, the storyteller model of the present method contains two different prompting modes, as shown in Figure 3 : 1. Naive mode: In the naive mode, the present method uses relatively simple prompts for the storyteller model. The prompts specify the requirement to imagine a story that can be told from an image and provide the “cookie thief” story as an example and some formatting instructions. The formatting instructions specify the requirements of the present method on the structure of the generated story. For example, the preference for stories to be concise rather than lengthy, they should be imaginative explanations of a single image rather than traditional narratives with a sequence of events or dialogue.
[0024] 2. Reasoning chain guidance mode: The reasoning chains that constitute a story image often involve multiple semantic dimensions. By analyzing the “cookie thief” and other story images, the present method identifies seven core dimensions that capture the essence of story images. With these semantic dimensions of reasoning chains as prior knowledge, the storyteller model generates more coherent and semantically richer stories. The seven dimensions are as follows: [Time]: The specific time when the story takes place, such as Christmas, Thanksgiving, or 3 am.
[0025] [Location]: The specific setting where the story takes place, such as a school, grocery store, or baseball field.
[0026] [Character Role]: Refers to the identity or profession of a character, such as a police officer, teacher, or firefighter.
[0027] [Character Relationship]: Refers to the relationship between characters, such as mother and son, lovers, etc.
[0028] [Event]: Refers to a key event that occurs in the story. For example, two boys steal cookies from a cabinet.
[0029] [Event Causal Relationship]: Refers to the causal connection between events. For example, a boy is scolded by his mother because he scribbled on the wall.
[0030] [Mental State]: Refers to the emotional or psychological state of a character, such as happy, angry, or sad.
[0031] Then, using the descriptions in the story images in the CogBench dataset, and further using the dynamic ICL (In-Context Learning) strategy to enhance diversity, three descriptions were randomly extracted from CogBench as prompt examples for the story home model, and through the above method, an image story corpus with reasoning chains was obtained, in which the actions of the characters were anchored at the same time point.
[0032] As an implementation, the method comprises optimizing training of the large language model in natural language narration, including: Based on the prepared positive story samples, the large language model is supervised fine-tuning training for creative reasoning guidance; Based on the prepared positive story samples and the rejected story samples, the large language model is preference optimization training for preference guidance.
[0033] In this implementation, since story narration mainly concerns the basic elements of natural language, such as grammar, vocabulary, facts, and reasoning, small language models (SLMs) have the potential to generate fluent and coherent stories. In view of this, the present method considers a fine-tuning strategy for the model, using the story corpus generated by the specialized model (GPT-4.1) in the reasoning chain guidance mode in the way of knowledge distillation to improve the quality of story generation of open source models. The open source models can be selected: Llama-3.2-1B-Instruct, Llama-3.2-3B-Instruct, Llama-3.1-8B-Instruct, Qwen2.5-1.5B-Instruct and Qwen2.5-7B-Instruct. After fine-tuning, a series of lightweight Mini-Storyteller models are obtained.
[0034] Specifically, the present method attempts to use the "supervised fine-tuning" (SFT) scheme. That is, 2000 high-quality stories are generated by GPT-4.1 as a storyteller model, and then these stories are used to fine-tune open source models such as Llama-3.1-8B-Instruct, expecting to replace expensive specialized models such as GPT-4.1 with small models.
[0035] The present method also trains lightweight storyteller models based on DPO (Direct Preference Optimization), which can be called Mini-Storyteller models. In order to construct the preference pair for training, the present method uses the stories generated by the open source model itself as rejected samples, and uses the output of the stronger model (GPT-4.1) as preferred samples, in order to encourage the open source model to align with the higher quality generated results.
[0036] Further, the present method further trains the DPO-Mix Mini-Storyteller model. For the two models of Llama-3.1-8B-Instruct and Qwen2.5-7B-Instruct, the present method uses the output of their corresponding small-scale models as rejected samples. Specifically, the rejected samples of Qwen2.5-7B-Instruct come from its own output and the output of Qwen2.5-1.5B-Instruct. And the rejected samples of Llama-3.1-8B-Instruct come from its own output and the output of Llama-3.2-1B-Instruct.
[0037] By the above-mentioned manner, the Mini-Storyteller model constructed by using the large language model and the fine-tuning training optimization is used to construct the storyteller model, and a richer image story corpus with a reasoning chain at the same time point is realized.
[0038] For step S12, the image story corpus with a reasoning chain determined in step S11 is input as a prompt word to an image rendering model, wherein the image rendering model can select a T2I model. It should be noted that the T2I model of the present method is only used as an image rendering model and does not participate in the generation of the story corpus in step S11. At this time, the image generation function of the T2I model is used to create an image describing the generated story. In order to facilitate understanding, the image rendering model can be personified as a painter model.
[0039] For step S13, in order to evaluate the story image presenting the semantic dimension narrative scene, a story image evaluation framework is specially designed.
[0040] Specifically, the story image evaluator includes a semantic complexity evaluator, a content diversity evaluator, and an alignment evaluator. The evaluation of the image story corpus and the story image based on the story image evaluator includes: In the semantic complexity evaluator, the story image is converted into a text feature description, and a semantic complexity score is determined based on the text feature description. In the content diversity evaluator, the image story corpus is structurally processed to obtain a structured text of an incomplete story, and an embedding vector of the structured text is determined. The average KNN cosine distance using the embedding vector is determined as a diversity score. In the alignment evaluator, an alignment score of the image story corpus and the story image is determined. In the alignment evaluator, the alignment score of the image story corpus and the story image includes: The structured information extractor is used to extract multiple semantic dimensions including time, location, character identity, character relationship, event, event causality, and character psychological state from the image story corpus. The multi-modal scorer performs multi-dimensional fine-grained alignment scoring on the image story corpus and the story image based on the multiple semantic dimensions, and determines the alignment score through the multi-dimensional fine-grained alignment scoring.
[0041] In the present embodiment, the story image evaluator of the present method includes three evaluators: a semantic complexity evaluator for evaluating the semantic complexity of the generated story and image; a KNN-based diversity evaluator for evaluating the diversity of the story; and an alignment evaluator for the story-image for measuring the semantic gap between the story and the image.
[0042] 1. Semantic complexity evaluator: This method automatically evaluates the storytelling ability of images by defining a semantic complexity score. This method employs the ISA (Image Semantic Assessment) model as an evaluator. Specifically, it uses CoT VLISA (BERT) in the ISA model as an evaluator to automatically calculate the semantic complexity score. CoT VLISA (BERT) will first extract features from the image in text form using GPT-4o, and then use a fine-tuned BERT to calculate the semantic score based on the text features. The extracted text features include the seven dimensions listed above. Based on the calculated semantic score, by fixing the other two variables, the effectiveness of the storyteller model and the painter model under different prompt patterns can be observed.
[0043] 2. KNN-based diversity evaluator: This method uses distance-based semantic diversity to measure the diversity of the stories generated by this method.
[0044] To calculate the diversity score, the potential bias introduced by different story lengths generated by different models is first addressed. To ensure a fair comparison of core narrative components, this method uses a custom-designed summarizer (gpt-4.1-mini) that structurally processes each generated story to extract a relatively standardized structured representation, including four CoR elements: [time], [place], [character], and [event]. The prompts used by this method are shown in Figure 4 .
[0045] Then, the extracted information (rather than the complete story text) is input into the text embedding model Qwen3-Embedding-0.6B to obtain its embedding e. The diversity score is then defined as the average KNN (K-Nearest Neighbors) cosine distance of all embeddings. The formula is as follows: where N is the number of generated stories, K is the number of nearest neighbors, are the embeddings of the extracted key information of stories i and j, respectively. denotes their cosine similarity. For example, K can be set to 5, or other values.
[0046] 3. Story-image alignment evaluator: This method uses a two-stage approach to evaluate the alignment between the generated image and its corresponding story.
[0047] In the first stage, key point extraction is performed. The method uses LLM (GPT-4.1) as a structured extractor. The model parses the generated story to extract information in seven predefined dimensions: [time], [location], [character identity], [character relationship], [event], [event causality], and [character psychological state].
[0048] In the second stage, the alignment score is determined. The method uses LVLM (GPT-4o) as a score evaluator. The score evaluator evaluates the generated image based on the structured key points extracted in the previous stage. The evaluation is multi-faceted: a fine-grained consistency score is provided for each of the seven dimensions, and an overall score summarizes the overall consistency. The prompt is as shown in Figure 5 .
[0049] Finally, if the evaluation is passed (for example, the score exceeds a preset threshold, or other settings can be made), a story-image semantically enhanced story image is obtained. If the evaluation is not passed, it means that the generated story image does not meet the standard of a story image.
[0050] As can be seen from the implementation, the method can automatically and in batches produce rare "story images" that meet the standard. These images contain rich semantics, multiple logical reasoning chains, and complex narratives in a single picture. Users do not need to think of complex prompt words, solving the defect of general image rendering models relying on prompt words.
[0051] Furthermore, the method can be used to realize an efficient illustration tool, automatically generating high-quality, "soulful", and "story-like" illustrations for blogs, novels, children's books, advertising ideas, and other fields, greatly improving the appeal of text-image content, and realizing the empowerment of content creation. The generated "story images" are used to evaluate human cognitive ability and advanced reasoning ability of large visual language models, promoting the development of AI cognitive assessment and human cognitive assessment.
[0052] Experiments are conducted to illustrate the method. Regarding the experimental setup, for the storyteller model Storyteller, the method uses five different parameter sizes of open-source LLM: Llama-3.2-1B-Instruct, Llama-3.2-3B-Instruct, Llama-3.1-8B-Instruct, Qwen2.5-1.5B-Instruct, and Qwen2.5-7B-Instruct, as well as two most advanced proprietary LLMs GPT-4o and GPT-4.1. For the image rendering model (artist model), the method uses two representative T2I models: DALL·E 3 and GPT-image 1.
[0053] As for the evaluation metrics, in addition to the three main evaluation metrics mentioned above, the method further evaluates according to the following standards: Human semantic complexity score, although the ISA model can automatically calculate the semantic complexity score and maintain relatively high consistency with human judgment, it still inevitably contains bias. Therefore, in order to more accurately evaluate the effectiveness of different methods, in addition to the ISA model, the method also combines manual evaluation to evaluate the semantic complexity of the generated images. Specifically, the annotation standard established in ISA is used to train human evaluators, and the original 1 to 5 scoring scale is refined to include 0.5 increments to obtain finer scoring granularity. Each image is scored independently by three evaluators, and the final manual semantic score is obtained by averaging their scores. To measure the reliability between raters, the ICC (Intra-class Correlation) of the annotation is calculated. The model selected by the method is a two-way mixed effects model. The ICC of human evaluation is 0.889, which means that there is high consistency between evaluators.
[0054] Number of visual cues, in the feature extraction stage of CoT VLISA (BERT), seven dimensions of visual cues are extracted. Therefore, the number of visual cues can be calculated according to the features as a reference.
[0055] Word count, since the generated story aims to be short and suitable for image generation, rather than a long narrative with a time progression, the method calculates the number of words in each story. In other words, in many cases, an excessively long story length can be a signal that the model has failed to follow the given instructions.
[0056] For each story painter configuration, the method generates 30 stories and their corresponding images, and calculates the average of the evaluation metrics as the final result. CogBench contains many story images created by human illustrators, which are used as a reference for human ability. Considering that the training data of the ISA model contains some CogBench images, the method determines 48 overlapping images from the ISA test set as a reference set. For the reasoning chain guide mode, CogBench is split in a ratio of 3:2, where the two parts are used to build training data and test data, respectively.
[0057] For experimental results, the method compares the effectiveness of different prompt modes: such as Figure 6The performance of various configurations of our method story-image generation is shown (wherein the semantic scores are expressed in %, with a maximum score of 100%. The quality configuration is "medium" for "*"; the quality setting is "automatic" for "†". "↑30" indicates that 30 samples with the highest evaluation scores are selected. The highest is shown in bold line, and the second highest is shown in underlined). By comparing the semantic scores of the images generated under the two different prompt modes, it can be observed that the inference chain guidance mode significantly improves the semantic complexity of the generated stories and corresponding images. Specifically, according to the number of visual cues in Figure 6 and Figure 7 , it can be seen that, compared with the naive mode, the inference chain guidance mode of our method does indeed increase the number of visual cues in multiple dimensions. In terms of diversity, the diversity score of the inference chain guidance mode of our method is also significantly higher than that of the naive mode, which can be attributed to the dynamic ICL strategy.
[0058] GPT-4o and GPT-4.1 combined with the image rendering model GPT-Image-1 as storyteller models can achieve the best performance. By comparing their semantic scores with CogBench, it can be seen that they can generate some qualified story images. However, the gap in semantic scores between the images generated by GPT-4 and the top 30 artificially created images in CogBench indicates that the ability of this model is still far below the top human performance. The performance of smaller open-source models is significantly inferior to GPT-4o and GPT-4.1. Qwen2.5-1.5B-Instruct, Llama3.2-1B-Instruct, and Llama-3.2-3B-Instruct generate significantly more story words than other models. This is because they cannot understand the instructions and generate long stories with a time sequence. Other larger open-source models, although they can understand the task instructions of our method and occasionally generate detailed stories, are still not very good at building logically coherent and semantically rich narratives. This indicates the high challenge of the story-image generation task with inference chains.
[0059] Regarding the impact of the painter model (image rendering model) on the whole: As for the painter model, different images generated by different painter models for the same story will produce different semantic scores. GPT-Image-1 is a newer T2I model, and its performance is superior to DALL·E 3. When the "quality" hyperparameter of GPT-Image-1 changes from "medium" to "automatic", a significant improvement in performance can also be observed.
[0060] This is attributed to the different text-to-image alignment capabilities of the painter models. Higher quality images generated by better painter models can more accurately depict events in the image, character emotions, other visual cues, and their interconnections, making it easier for an ISA model or a human to understand the story happening in the image. This conclusion can be illustrated by Figure 8 the story-image alignment scores of GPT-image-1, which are significantly higher than those of DALL-E 3. The performance gap is particularly significant in the four dimensions of [characters], [character relationships], [events], and [event causality]. More specifically, as shown in Figure 9 the top half, GPT-image-1 successfully presents the text “Strawberries: Buy One Get One Free” in the image, while DALL-E 3 does not. In the bottom half, GPT-image-1 also more accurately depicts the emotions of the characters, capturing the excitement and nervousness of the children as they sneakily open the presents, and the satisfaction of the father. Thus, it can be seen that the story-image generation task also places higher demands on the capabilities of T2I models to meet the requirement of telling a semantically rich story using only a single image.
[0061] Figure 10 shows the story images generated by the present method. It can be seen that GPT-4o is able to generate logically coherent, semantically rich stories and subsequently generate corresponding images. For example, the story generated by GPT-4o takes place in a grocery store on a weekend discount day, where a police officer arrests a thief who is trying to steal an item. The generated image contains relatively rich visual cues and reasoning chains CoR. For the open-source Storyteller models Qwen2.5-7B-Instruct, Llama-3.2-3B-Instruct, and Qwen2.5-1.5B-Instruct, the stories they generate are clearly less engaging and coherent in terms of interest and logical consistency. While Qwen2.5-7B-Instruct can generate logically coherent stories, the semantic complexity is clearly lower. For the image generated by Qwen2.5-7B-Instruct, it only describes a scene where a family is celebrating Christmas, without providing rich visual cues or causal relationships between events. Llama-3.2-3B-Instruct and Qwen2.5-1.5B-Instruct perform worse. The stories they produce are often primarily a superposition of multiple events. In both cases, the models fail to correctly follow the instructions and generate a long story with a temporal sequence or even a dialogue, instead of a story that can be described in a single image. Overall, these cases highlight the key role of story models in generating story images, while also revealing their limitations.
[0062] To verify the effectiveness of the diversity evaluator of this method, the method first constructed a test set, which consisted of 8 groups of data with different degrees of repetition. Each group contained 30 samples, and human ratings of diversity for each group were obtained. Subsequently, the proposed diversity metrics were analyzed for correlation with these human evaluations. The results are as Figure 11 shown. Different measurement methods using KNN or simple average distance (Avg.) are compared in the figure, and the evaluation is carried out with or without a summarizer respectively. According to the results, the correlation between the KNN distance of this method combined with the summarizer and human judgment is the highest, achieving an excellent Spearman coefficient of 0.958 and a strongly correlated Pearson coefficient of 0.912, significantly outperforming each ablation experiment. It is worth noting that the same KNN metric only reaches 0.802 and 0.842 without a summarizer, highlighting the key role of the summarizer in extracting relevant elements. In addition, the superior performance of KNN compared to Avg. distance verifies its effectiveness. These strongly correlated results indicate that the measurement metrics of this method can robustly capture the diversity performance of generated stories.
[0063] To verify the effectiveness of the alignment evaluator, the method constructed a test set by extracting 209 key points and manually annotating their true alignment scores. After benchmark testing with this manually annotated dataset, the evaluator of this method achieved an overall accuracy of 0.737. This result shows that the evaluator has strong accuracy in judging whether a certain key point is visually presented in the image.
[0064] Regarding the Mini-Storyteller model of this method, in terms of training, the method uses two fine-tuning strategies: supervised fine-tuning (SFT) and direct preference optimization (DPO).
[0065] 1. SFT: For SFT, the student model is trained using the standard causal language modeling objective, which aims to maximize the probability of the target answer. The SFT loss function is the standard cross-entropy loss, calculated as the negative log-likelihood of the target token: where D is the SFT dataset of this method, (x, y) is a prompt-response pair, T is the number of tokens in the response y, is the probability of the T-th token given the prompt x and the previous token y<T as predicted by the model with parameters θ. The training data directly utilizes the high-quality instruction-response pair dataset generated by GPT-4.1 for SFT.
[0066] 2. DPO: The loss function used in DPO training is defined as: where, and denote the probabilities assigned by the model (parameterized by θ) to the two candidate responses for a given input x, where, is the preferred response, is the rejected response. Similarly, and denote the probabilities assigned by the reference policy. To construct the training preference pairs, the self-generated stories of the student model are used as rejected samples, and the output of the stronger model is used as the preferred sample, encouraging the student to align with the higher-quality generation. Additionally, for Llama-3.1-8B-Instruct and Qwen2.5-7B-Instruct, the method also attempts to leverage the output of their respective small-scale models as rejected samples. That is, the rejected samples for Qwen2.5-7B-Instruct come from its own output and the output of Qwen2.5-1.5B-Instruct, while for Llama-3.1-8B-Instruct, they come from its own output and Llama-3.2-1B-Instruct. This strategy is named DPO-Mix. For each preference pair, the selected and rejected samples are generated using the same prompt to ensure direct comparability.
[0067] During the dataset construction process, 2000 training samples generated by GPT-4.1 are created for SFT and DPO. Additionally, 4000-sample training sets are constructed for Llama3.1-8B-Instruct and Qwen2.5-7B-Instruct. For each of these models, its data includes 2000 samples from its respective smaller-scale model Llama-3.2-1B-Instruct and Qwen2.5-1.5B-Instruct, and another 2000 samples generated by the model itself. The method trains the models using LLaMA Factory on Tesla V100 GPUs. Each model is trained for one epoch using LoRA with a learning rate chosen from {1e-4, 1e-5} and a total training batch size of 16. A cosine learning rate scheduler with a warmup ratio of 0.1 is also applied. During training, 10% of the training data is also set aside for validation.
[0068] As Figure 12Results of the Mini-Storyteller model of the present method are shown. It can be observed that the semantic complexity has been significantly improved after training. The increase in semantic score indicates that both SFT and DPO can enhance the Storyteller model’s ability to generate semantically rich stories. In addition, Llama-3.2-1B-Instruct, Llama-3.2.3B-Instruct, and Qwen2.5-1.5B-Instruct have significantly reduced the number of words. This indicates that these models successfully generated the required concise stories following the instructions, rather than long narratives driven by time series. Specifically, SFT is more effective than DPO in terms of semantic complexity. For example, with SFT, Mini-Storyteller-LLaMA-3.1-8B-SFT achieved a performance level comparable to GPT-4 and improved by 24.4 and 29 points in automatic and human semantic scoring, respectively, over Llama-3.1-8B-Instruct, reaching the best performance.
[0069] DPO-Mix shows better performance for Llama-3.1-8B-Instruct and Qwen2.5-7B-Instruct. While DPO successfully trained small models (Llama-3.2-1B-Instruct, Llama-3.2-3B-Instruct, and Qwen2.5-1.5B-Instruct) to generate the required short stories, it was less effective for larger models. Vanilla DPO failed to reduce the number of words for Llama-3.1-8B-Instruct and Qwen2.5-7B-Instruct, and even increased the number of words for the latter. This is because it disrupted their original understanding of the task instructions, causing it to generate lengthy, time-series-driven stories. DPO-Mix can effectively address this issue by introducing rejected samples from small-scale models and further improve the semantic score. As shown in Table 2, DPO-Mix significantly reduces the number of words for Llama-3.1-8B-Instruct and Qwen2.5-7B-Instruct, and improves the semantic score for both. Figure 13 Four cases generated by four different Mini-Storyteller models are shown.
[0070] It can be seen that after training, Mini-Storyteller-LLaMA-3.2-1B-SFT and Mini-Storyteller-LLaMA-3.2-3B-DPO can successfully generate concise and engaging stories without time sequences and dialogues. The images generated by the Mini-Storyteller model obviously contain more visual cues that form the CoR. For example, "Today: Free cupcakes at the school science fair" in the image generated by Mini-Storyteller-LLaMA-3.2-1B-SFT sets the background for the main activity and echoes the actions of the students outside the window and the staff inside. In the image generated by Mini-Storyteller-LLaMA-3.2-3B-DPO, the baseball in the girl's hand and the children playing baseball in the background both suggest the reason for her fall. For the image generated by Mini-Storyteller-Qwen2.5-7B-DPO-Mix, the sleep mask on the man's head indicates that it is bedtime, and he is either preparing to sleep or has just been awakened by the boy. The story in the image also contains various semantic dimensions, including [time], [location], [character identity], [character relationship], [event], [event causality], and [character mental state].
[0071] Overall, the present method proposes a new story image generation task aimed at generating semantically rich images that can tell coherent stories. To this end, the Story Painter task is also proposed, which is a two-stage process that combines the creativity of LLMs with the visual synthesis capabilities of T2I models. The present method also proposes an evaluation framework consisting of three evaluators. To improve the storytelling ability of small open-source LLMs and narrow the performance gap with proprietary LLMs, the models are fine-tuned using SFT and DPO, resulting in Mini-Storyteller models. The story image generation task has broad application potential, including cognitive assessment of humans and models, book illustrations, and enhancing the ability of multi-modal models to understand and generate semantically complex images.
[0072] As Figure 14 The structure of the story image generation system with reasoning chain provided by an embodiment of the present application is shown in the structural schematic diagram of the story image generation system with reasoning chain. The system can execute the story image generation method with reasoning chain described in any of the embodiments above and is configured in a terminal.
[0073] The story image generation system with reasoning chain 10 provided by the present embodiment includes a Storyteller module 11, a Painter module 12, and an Evaluation module 13.
[0074] The story maker module 11 is configured to generate an image story corpus of reasoning chains of execution actions of a role anchored at the same time point and with different semantic dimensions, by using a story maker model constructed by a large language model, wherein the reasoning chains of different semantic dimensions are used to constitute multi-dimensional semantics of the image story corpus.
[0075] The embodiment of the present application also provides a non-volatile computer storage medium, which stores computer executable instructions, and the computer executable instructions are used for executing the story image generation method with reasoning chains in any method embodiment. As an implementation manner, the non-volatile computer storage medium of the present application stores computer executable instructions, and the computer executable instructions are configured to: generate an image story corpus of reasoning chains of execution actions of a role anchored at the same time point and with different semantic dimensions, by using a story maker model constructed by a large language model, wherein the reasoning chains of different semantic dimensions are used to constitute multi-dimensional semantics of the image story corpus; input the image story corpus with reasoning chains into an image rendering model as prompt words, to generate a story image presenting a multi-semantic dimension narrative scene; evaluate the image story corpus and the story image based on a story image evaluator, and obtain a story image with rich semantics after evaluation.
[0076] As a non-volatile computer readable storage medium, the non-volatile computer readable storage medium can be used for storing non-volatile software programs, non-volatile computer executable programs and modules, such as program instructions / modules corresponding to the method in the embodiment of the present application. One or more program instructions are stored in the non-volatile computer readable storage medium, and when executed by a processor, the story image generation method with reasoning chains in any method embodiment is executed.
[0077] Figure 15 is a hardware structure schematic diagram of an electronic device with the story image generation method with reasoning chains provided by another embodiment of the present application, as shown in the figure, the device comprises: Figure 15 as shown in the figure, the device comprises: one or more processors 1510 and a memory 1520, Figure 15 In the figure, the processor 1510 is taken as an example. The device with the story image generation method with reasoning chains can further comprise an input device 1530 and an output device 1540.
[0078] The processor 1510, the memory 1520, the input device 1530 and the output device 1540 can be connected by a bus or other means, Figure 15 The connection by the bus is taken as an example.
[0079] The memory 1520, as a non-volatile computer readable storage medium, can be used to store non-volatile software programs, non-volatile computer executable programs and modules, such as program instructions / modules corresponding to the story image generation method with reasoning chain in the embodiments of the present application. The processor 1510 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions and modules stored in the memory 1520, that is, implements the story image generation method with reasoning chain in the above method embodiments.
[0080] The memory 1520 can include a program storage area and a data storage area, wherein the program storage area can store an operating system and at least one application required by a function; the data storage area can store data and the like. In addition, the memory 1520 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state memory device. In some embodiments, the memory 1520 can optionally include a memory disposed remotely with respect to the processor 1510, and these remote memories can be connected to the mobile device through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0081] The input device 1530 can receive input digital or character information. The output device 1540 can include a display device such as a display screen.
[0082] The one or more modules are stored in the memory 1520, and when executed by the one or more processors 1510, the story image generation method with reasoning chain in any of the above method embodiments is executed.
[0083] The above product can execute the method provided by the embodiments of the present application, and has the corresponding function modules and beneficial effects of executing the method. Technical details not described in detail in the embodiments can be referred to the method provided by the embodiments of the present application.
[0084] Non-volatile computer-readable storage media may include a stored program area and a stored data area, wherein the stored program area may store an operating system and an application program required for at least one function; the stored data area may store data created based on the use of the device, etc. Furthermore, the non-volatile computer-readable storage medium may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, the non-volatile computer-readable storage medium may optionally include memory remotely located relative to the processor, and these remote memories may be connected to the device via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0085] This invention also provides an electronic device comprising: at least one processor and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the steps of the story image generation method with inference chain according to any embodiment of this invention.
[0086] The electronic devices described in this application exist in various forms, including but not limited to: (1) Mobile communication devices: These devices are characterized by their mobile communication capabilities and primarily aim to provide voice and data communication. These terminals include smartphones, multimedia phones, feature phones, and low-end phones.
[0087] (2) Ultra-mobile personal computer devices: These devices fall under the category of personal computers, possessing computing and processing capabilities, and generally also have mobile internet access features. These terminals include PDAs, MIDs, and UMPCs, such as tablet computers.
[0088] (3) Portable entertainment devices: These devices can display and play multimedia content. This category includes audio and video players, handheld game consoles, e-book readers, as well as smart toys and portable car navigation devices.
[0089] (4) Other electronic devices with data processing functions.
[0090] In this document, relational terms such as first and second and the like can be used solely to distinguish one entity or action from another entity or action without necessarily requiring or implying any actual such relationship or order between such entities or actions. Moreover, the terms "comprises", "comprising", "includes", "including" and the like can be used herein to indicate either an inclusion of a few items in a fully detailed list of items, or to indicate that items are included at least to some extent. Lack of inclusion of such terms should not be interpreted as excluding the possibility of inclusion.
[0091] The above-described apparatus embodiments are merely illustrative, and the units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., can be located in one place or distributed on multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purposes of the embodiments. Those skilled in the art can understand and implement without creative labor.
[0092] From the above description of the embodiments, those skilled in the art can clearly understand that the embodiments can be implemented by means of software and necessary universal hardware platforms, and of course can also be implemented by hardware. Based on such understanding, the above technical solutions, essentially or in terms of contribution to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, and the like, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0093] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements for some technical features; and these modifications or replacements do not make the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for generating story images with reasoning chains, comprising: The image story corpus generated by the storyteller model constructed from a large language model has characters whose actions are anchored at the same point in time and have inference chains with different semantic dimensions. The inference chains with different semantic dimensions are used to constitute the multi-dimensional semantics of the image story corpus. The image story corpus with the inference chain is input as prompt words into the image rendering model to generate story images that present multi-semantic dimension narrative scenes; The image story corpus and the story images are evaluated using a story image evaluator, and semantically rich story images are obtained after the evaluation.
2. The story image generation method with reasoning chain according to claim 1, characterized in that, Before the execution actions of characters generated by the storyteller model built from the large language model are anchored to reasoning chains with different semantic dimensions at the same point in time, the method includes optimizing the large language model for natural language narrative training, including: The large language model is subjected to supervised fine-tuning training guided by creative reasoning based on pre-prepared positive story samples. The large language model is trained using preference-guided preference optimization based on pre-prepared positive story samples and rejection story samples.
3. The story image generation method with reasoning chain according to claim 1, characterized in that, The story image evaluator includes: a semantic complexity evaluator, a content diversity evaluator, and an alignment evaluator; The evaluation of the image story corpus and the story images by the story image evaluator includes: In the semantic complexity estimator, the story image is converted into a text feature description, and a semantic complexity score is determined based on the text feature description; In the content diversity evaluator, the image story corpus is structured to obtain structured text of incomplete stories, the embedding vector of the structured text is determined, and the average KNN cosine distance of the embedding vector is used to determine the diversity score. In the alignment evaluator, the alignment score between the image story corpus and the story image is determined.
4. The story image generation method with reasoning chain according to claim 3, characterized in that, The method further includes: The evaluation scores of the image story corpus and the story image are determined based on the semantic complexity score, diversity score, and alignment score. If the evaluation score reaches the preset evaluation score, the evaluation is passed.
5. The story image generation method with reasoning chain according to claim 1, characterized in that, The semantic dimensions include: time, location, role identity, role relationship, event, causal relationship of the event, and role psychological state.
6. The story image generation method with reasoning chain according to claim 3, characterized in that, In the alignment evaluator, determining the alignment score between the image story corpus and the story image includes: The structured information extractor is used to extract multiple semantic dimensions from image story corpora, including time, location, character identity, character relationship, event, event causality, and character psychological state. The image story corpus and story images are aligned using a multi-dimensional fine-grained alignment score based on multiple semantic dimensions by a multi-modal scorer, and the alignment score is determined by the multi-dimensional fine-grained alignment score.
7. A story image generation system with a reasoning chain, comprising: The storyteller module is used to generate image story corpora with inference chains that anchor the character's actions at the same point in time and have different semantic dimensions, using a storyteller model built from a large language model. The inference chains with different semantic dimensions are used to constitute the multi-dimensional semantics of the image story corpora. The painter module is used to input the image story corpus with the reasoning chain as prompt words into the image rendering model to generate story images that present multi-semantic dimension narrative scenes. The evaluation module is used to evaluate the image story corpus and the story images based on the story image evaluator, and obtain semantically rich story images after evaluation.
8. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method described in any one of claims 1-6.
9. An electronic device comprising: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor to enable the at least one processor to perform the steps of the method according to any one of claims 1-6.
10. A storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the steps of the method described in any one of claims 1-6.