A picture book generation method, device, vehicle, storage medium and program product

By generating a blueprint for picture book creation and constructing visual feature data, the problem of visual consistency in picture book generation using the diffusion model was solved, achieving visual consistency and narrative coherence in picture books and improving the quality of picture book generation.

CN122289442APending Publication Date: 2026-06-26CHONGQING CHANGAN AUTOMOBILE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHONGQING CHANGAN AUTOMOBILE CO LTD
Filing Date
2026-03-31
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

Existing diffusion models, when generating continuous picture books, suffer from random sampling features that prevent the same character from being accurately reproduced in multiple images, thus disrupting the visual unity and coherence of the picture book.

Method used

By generating a blueprint for picture book creation, visual objects that appear more frequently than a preset frequency requirement are extracted, corresponding visual feature data is constructed, and an image generation model is used to strictly follow the pre-constructed visual features to generate a picture book image set, ensuring visual consistency and narrative coherence.

Benefits of technology

It achieves visual consistency of the same core visual object throughout the entire picture book, strengthens the narrative coherence of the picture book and the precise matching of visual style with content plot, and improves the quality of picture book generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122289442A_ABST
    Figure CN122289442A_ABST
Patent Text Reader

Abstract

This invention relates to the field of artificial intelligence technology and discloses a picture book generation method, apparatus, vehicle, storage medium, and program product. Responding to a user-input picture book theme, the invention conceptualizes a story based on the theme, generates a picture book generation blueprint containing pagination scripts and visual object classification information, extracts visual objects whose frequency of occurrence exceeds a preset frequency requirement from the blueprint, and constructs corresponding visual feature data for these visual objects. Then, based on the visual feature data and the script information corresponding to each page, a picture book image set is generated. By strictly adhering to the pre-constructed visual features to generate picture book images, visual consistency of the same core visual object is achieved across all pages of the picture book, establishing visual connections between pages, strengthening the narrative coherence of the picture book, and achieving precise adaptation of visual style to content plot, thereby improving the quality of picture book generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, specifically to a picture book generation method, apparatus, vehicle, storage medium, and program product. Background Technology

[0002] As a visual storytelling medium, picture books place extremely high demands on the consistency of the core characters' identities, appearances, and overall artistic style. Although advanced models, such as the Stable Diffusion XL (SDXL) model, can generate single images, when directly applied to content creation tasks with continuous plots, the generation process of the diffusion model is essentially a process of gradually denoising and restoring the image by eliminating noise. Each step involves random sampling, and even with identical text prompts, slight differences in the initial noise will be amplified during the denoising process, resulting in differences in details in the generated images. For example, for character generation, the randomness directly leads to the inability to accurately reproduce the same character in multiple consecutive images, thus disrupting the overall visual unity and coherence of the picture book. Summary of the Invention

[0003] This invention provides a picture book generation method, apparatus, vehicle, storage medium, and program product to solve the problem of the overall visual unity of a picture book being destroyed due to random sampling features.

[0004] In a first aspect, the present invention provides a picture book generation method, the method comprising: responding to a picture book theme input by a user, conceiving a story for the picture book theme, generating a picture book generation blueprint including a pagination script and visual object classification information, wherein the visual object classification information includes at least one visual object and a category corresponding to each visual object, the category being determined based on the frequency of occurrence of the visual object in the story; extracting a list of visual objects of a target category from the picture book generation blueprint, and constructing corresponding visual feature data for each visual object in the list of visual objects, wherein the visual objects of the target category are visual objects that appear more frequently than a preset frequency requirement in multiple pages of the pagination script; and generating a picture book image set based on the visual feature data and the script information corresponding to each page, wherein the images in the picture book image set maintain visual consistency when presenting the visual objects of the target category.

[0005] The picture book generation method provided by this invention responds to a picture book theme input by the user, conceives a story for the picture book theme, generates a picture book generation blueprint containing pagination scripts and visual object classification information, extracts visual objects that appear more frequently than a preset frequency requirement from the picture book generation blueprint, and constructs corresponding visual feature data for these visual objects. Then, based on the visual feature data and the script information corresponding to each page, a picture book image set is generated. By strictly following the pre-constructed visual features to generate picture book images, the method achieves visual consistency of the same core visual object across all pages of the picture book, establishes visual connections between pages, strengthens the narrative coherence of the picture book, and achieves precise adaptation of visual style and content plot, thereby improving the quality of picture book generation.

[0006] In one optional implementation, the step of conceiving a story for the picture book theme and generating a picture book generation blueprint containing pagination scripts and visual object classification information includes: based on the picture book theme, calling a preset knowledge base to perform story retrieval and conception, and generating a story outline; expanding the story outline into a pagination script, and generating a corresponding script description for each page of the pagination script; identifying visual objects in the pagination script and marking all appearing visual objects; classifying each visual object according to its frequency of occurrence across multiple pages to obtain visual object classification information; converting the script description of each page into prompts that conform to the input specifications of the image generation model, wherein the input specifications include at least a token length limit; and integrating the pagination script, the visual object classification information, and the prompts corresponding to each page to generate the picture book generation blueprint.

[0007] This invention achieves precise matching of picture book themes and story materials through a built-in knowledge base, improving the logic and efficiency of story conception, ensuring the quality of picture book plots, and clearly converting prompt words according to the model's input specifications. This ensures that the prompt words fully retain the core visual and plot information, avoids token overload, avoids attention distraction problems, and improves image generation accuracy.

[0008] In one optional implementation, when the current page is a complex page that does not conform to the preset generation constraints, the step of converting the script description of each page into prompt words that conform to the input specifications of the image generation model includes: decomposing the script description corresponding to the current page into multiple sub-descriptions; and converting each sub-description into prompt words that conform to the input specifications of the image generation model.

[0009] When encountering scenarios that are too complex for a single generation, this invention can automatically break down the script description into multiple sub-descriptions, ensuring that the instruction prompts for each page of the picture book are concise and clear, thereby improving the quality of the picture book images.

[0010] In one optional implementation, constructing corresponding visual feature data for each visual object in the list of visual objects includes: generating a corresponding standard image for the current visual object; extracting a first type of feature information and a second type of feature information from the standard image, wherein the first type of feature information is used to characterize the essential identity information of the visual object, and the second type of feature information is used to characterize the appearance information of the visual object; and performing vector concatenation and fusion processing on the first type of feature information and the second type of feature information to obtain the visual feature data of the current visual object.

[0011] This invention performs deep mapping and fusion of essential identity features and outward appearance features, preserving the core essence of identity and rich details of appearance, thereby improving the accuracy of picture book generation.

[0012] In an optional implementation, when it is detected that the multiple consecutive pages currently being processed in the picture book generation blueprint belong to a continuous action scene, the generation of the picture book image set further includes: when generating the image of the current page, obtaining the page image of the previous page as a reference frame; during the denoising process of the image generation model, extracting the mask region of the target visual object in the reference frame, wherein the target visual object is the object performing the continuous action; calculating the pixel-level dense spatial mapping relationship between the current page image and the reference frame regarding the target visual object for the mask region; according to the pixel-level dense spatial mapping relationship, fusing the preset visual detail features belonging to the target visual object in the reference frame into the features of the current page image according to the preset weights; and based on the pixel-level dense spatial mapping relationship, sharing the appearance style features corresponding to the target visual object in the reference frame to the current page image through a shared attention mechanism, wherein the appearance style features include at least illumination distribution and color.

[0013] This invention uses the previous page image as a reference frame and integrates the visual detail features of the target visual object in the reference frame into the features of the current page image according to a preset weight, ensuring the natural continuation of details. Furthermore, through a shared attention mechanism, the appearance style features of the target visual object in the reference frame are shared to the current page, ensuring that the appearance style attributes remain consistent in the same continuous action sequence. This guarantees the accurate reproduction of the core characters and the overall visual unity, while also enhancing narrative coherence and the visual smoothness of continuous actions, thereby improving the accuracy, adaptability, and efficiency of picture book generation.

[0014] In one optional implementation, calculating the pixel-level dense spatial mapping relationship between the current page and the reference frame regarding the target visual object includes: extracting the first deep semantic feature map and the second deep semantic feature map corresponding to the current page image and the reference frame respectively during the denoising process; calculating the feature similarity between each feature position in the first deep semantic feature map and each feature position in the second deep semantic feature map; and matching the corresponding feature position with the highest similarity in the reference frame for each feature position in the current page image based on the feature similarity, so as to construct a pixel-level dense spatial mapping relationship.

[0015] This invention extracts semantic feature maps and performs subsequent dense spatial mapping, which can accurately capture the semantic part information of the target visual object, unaffected by changes in lighting. Then, based on the dense spatial mapping relationship, it calculates the similarity of each feature position in the two feature maps, achieving pixel-level comprehensive correspondence. This ensures that every detail of the core visual object can be accurately reproduced across pages, further enhancing the visual coherence of the picture book across pages.

[0016] In an optional implementation, the method further includes: performing feature consistency verification and semantic verification on the current page images in the picture book image set, wherein the feature consistency verification is used to evaluate the feature similarity between the visual object to be tested in the current page image and the visual object to be tested in the standard image; the semantic verification is used to evaluate the semantic conformity between the current page image and the script description of the corresponding page; if all page images in the picture book image set pass the feature consistency verification and semantic verification, the final picture book image set is output.

[0017] This invention maps the generated image and the standard image to the same feature space for inner product operation, quantitatively evaluating the similarity between the newly generated visual object and the standard image. This transforms subjective similarity judgment into a rigorous and fast mathematical comparison that the system can perform, forming the core of automated quality inspection.

[0018] In an optional implementation, if the current page image fails the feature consistency check and / or semantic check, the method further includes: generating an image editing instruction based on the reason for the check failure; calling an image editing engine to make local or global adjustments to the current page image based on the image editing instruction; and repeatedly performing the feature consistency check and semantic check steps on the adjusted page image until the current page image passes the feature consistency check and semantic check.

[0019] This invention automatically triggers a correction process after a verification failure, eliminating the need for manual intervention and improving quality control efficiency. Specifically, it accurately locates the cause of the failure, generates targeted editing instructions, and performs subsequent corrections. This achieves dynamic generation of precise editing instructions that fit the failure scenario, avoiding blind corrections and improving the overall efficiency and quality stability of picture book generation.

[0020] Secondly, the present invention provides a picture book generation device, the device comprising: a story conception module, configured to, in response to a picture book theme input by a user, conceive a story for the picture book theme and generate a picture book generation blueprint containing a pagination script and visual object classification information, wherein the visual object classification information includes at least one visual object and a category corresponding to each visual object, the category being determined based on the frequency of occurrence of the visual object in the story; a visual feature construction module, configured to extract a list of visual objects of a target category from the picture book generation blueprint and construct corresponding visual feature data for each visual object in the list of visual objects, wherein the visual objects of the target category are visual objects that appear more frequently than a preset frequency requirement on multiple pages of the pagination script; and a picture book generation module, configured to, based on the visual feature data and the script information corresponding to each page, generate a picture book image set, wherein the images in the picture book image set maintain visual consistency when presenting the visual objects of the target category.

[0021] Thirdly, the present invention provides a vehicle, the vehicle including an interactive interface and a controller, the interactive interface being used to display a set of picture book images, the controller including a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the computer instructions to perform the picture book generation method of the first aspect or any corresponding embodiment described above.

[0022] Fourthly, the present invention provides a computer-readable storage medium storing computer instructions for causing a computer to execute the picture book generation method of the first aspect or any corresponding embodiment thereof.

[0023] Fifthly, the present invention provides a computer program product, including computer instructions for causing a computer to execute the picture book generation method of the first aspect or any corresponding embodiment described above.

[0024] The present invention has the following technical effects: The picture book generation method provided by this invention responds to a picture book theme input by the user, conceives a story for the picture book theme, generates a picture book generation blueprint containing pagination scripts and visual object classification information, extracts visual objects that appear more frequently than a preset frequency requirement from the picture book generation blueprint, and constructs corresponding visual feature data for these visual objects. Then, based on the visual feature data and the script information corresponding to each page, a picture book image set is generated. By strictly following the pre-constructed visual features to generate picture book images, the method achieves visual consistency of the same core visual object across all pages of the picture book, establishes visual connections between pages, strengthens the narrative coherence of the picture book, and achieves precise adaptation of visual style and content plot, thereby improving the quality of picture book generation. Attached Figure Description

[0025] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0026] Figure 1 This is a structural block diagram of a picture book generation system according to an embodiment of the present invention; Figure 2 This is an example diagram illustrating the interaction logic of each component in a picture book generation system according to an embodiment of this aspect; Figure 3 This is a schematic diagram of the first type of picture book generation method according to an embodiment of the present invention; Figure 4 This is a flowchart illustrating the process of generating picture book images based on visual feature data and prompts according to an embodiment of the present invention. Figure 5 This is a schematic diagram of a second process for a picture book generation method according to an embodiment of the present invention; Figure 6 This is a flowchart illustrating the process of generating a picture book blueprint according to an embodiment of the present invention; Figure 7 This is a flowchart illustrating the construction of visual feature data according to an embodiment of the present invention; Figure 8 This is a flowchart illustrating image verification according to an embodiment of the present invention; Figure 9 This is a flowchart illustrating a picture book generation method according to an embodiment of the present invention; Figure 10 This is a structural block diagram of a vehicle according to an embodiment of the present invention; Figure 11 This is a structural block diagram of a picture book generation device according to an embodiment of the present invention; Figure 12 This is a schematic diagram of the hardware structure of the controller according to an embodiment of the present invention. Detailed Implementation

[0027] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0028] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0029] According to an embodiment of the present invention, a method for generating picture books is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0030] This embodiment provides a picture book generation method, which can be used in a picture book generation system, such as... Figure 1As shown, the picture book generation system includes a story architecture agent 1, a cross-image identity consistency generation engine 2, and an automated verification and correction module 3 (Quality Assurance, QA). The story architecture agent 1 is a pure text processing system whose goal is to optimize and standardize all text instructions that may lead to generation failure before generating images. This includes conceptualizing the story based on the user-input picture book theme, generating pagination scripts and image generation instructions for each page (including at least corresponding prompts), and identifying visual objects of the target category requiring cross-page consistency through character classification. This provides the cross-image identity consistency generation engine with a list of visual objects of the target category, forming a picture book generation blueprint. The cross-image identity consistency generation engine 2 is a visual execution unit that receives the picture book generation blueprint from the story architecture agent 1 and constructs corresponding visual feature data (Graphical Image) for each visual object of the target category. The system uses an entity (GIE) and stores the constructed visual feature data in the GIE feature library to ensure visual consistency across pages for each visual object. Then, based on the visual feature data and image generation instructions corresponding to each visual object, identity-injected image generation is performed to obtain the initial draft image. The automated verification and correction module 3 performs multi-dimensional verification on the initial draft image. If the verification passes, it outputs a highly consistent picture book image set. If the verification fails, the image can be corrected and verified again until all images pass the verification. The collaborative working logic between the modules can be found in [link to module description]. Figure 2 As shown, please refer to the following embodiments for a detailed description, which will not be repeated here.

[0031] Figure 3 This is a flowchart of a picture book generation method according to an embodiment of the present invention, such as... Figure 3 As shown, the process includes the following steps: Step S301: In response to the picture book theme input by the user, the story of the picture book theme is conceived, and a picture book generation blueprint containing pagination scripts and visual object classification information is generated.

[0032] The visual object classification information includes at least one visual object and the category corresponding to each visual object. The categories are determined based on the frequency of the visual object's appearance in the story.

[0033] In this embodiment of the invention, the story architecture agent 1 receives a picture book theme input by the user. The method of receiving the user-inputted picture book theme is not limited, but is set according to the application field and scenario. For example, in a car infotainment system, the user's voice input of the picture book theme can be monitored via a sound sensor. On a mobile device, the user can input the picture book theme via text or voice. For example, if the received picture book theme is "The Little Rabbit's Picnic," this is just an example. A short story outline can be created around the picture book theme, and the story outline can be expanded into a detailed paginated script (e.g., a 6-page script, with each page corresponding to 1-2 core plot points). Then, visual objects appearing in each page of the script (such as the little rabbit, the...) can be extracted. (Images of carrots, grass, spring rain, etc.) The frequency of each visual object appearing across all pages is statistically analyzed. Visual objects are then categorized based on frequency, such as core objects (appearing in over 80% of pages, like rabbits and grass), supporting objects (appearing in 30-80% of pages, like carrots), and scene objects (appearing in less than 30% of pages, like spring rain). This categorization information forms the visual object classification information, ultimately generating a picture book blueprint containing pagination scripts and visual objects (including categories). Alternatively, other classification methods can be used, such as categorizing visual objects appearing in the picture book theme as core protagonists and those not appearing in the theme as supporting characters; this is just one example.

[0034] Step S302: Extract the list of visual objects of the target category from the picture book generation blueprint, and construct corresponding visual feature data for each visual object in the list of visual objects.

[0035] The target category of visual objects refers to visual objects that appear more frequently than a preset frequency requirement across multiple pages in the pagination script.

[0036] The cross-image identity consistency generation engine 2 of this invention can extract a list of visual objects of the target category from the picture book generation blueprint. Among them, the visual objects of the target category are visual objects that appear more frequently than a preset frequency requirement in multiple pages of the pagination script (such as visual objects of the core category, which may include not only characters, but also continuously appearing buildings, scenes, etc.). Then, unified visual feature data can be formulated for the visual objects of the target category to ensure visual coherence in subsequent images. It can rely on the general public's common perception of visual objects (conventional appearance) and decompose the conventional appearance into feasible visual features to obtain visual feature data. Alternatively, a reference image that meets the expectations can be generated for the visual object first, and then the visual feature data can be extracted by image analysis tools. For example, a rabbit: white fur, pink inner ears, short round tail, etc., are just examples.

[0037] Step S303: Generate a picture book image set based on visual feature data and script information corresponding to each page.

[0038] Among them, the images in the picture book image collection maintain visual consistency when presenting visual objects of the target category.

[0039] This invention can use an image generation model (taking the SDXL dual-encoder model as an example) to generate picture book image sets. This includes utilizing the visual encoder (CLIP-ViT-G) and text encoder (CLIP-Text-G) included in the SDXL model to achieve initial separation of content and visual style at the instruction level. Using the diffusion model denoising network UNet as the central hub, visual feature data of visual objects is locked through GIE identity feature embedding. An attention mechanism is then used to combine cue words and visual feature data of visual objects to generate single-page images, ensuring that the core visual objects maintain high consistency in shape and color throughout the entire picture book. Specifically, the visual encoder reads the pre-constructed visual feature data of visual objects and sets the core visual objects... Visual features such as appearance, texture, and color are converted into fixed VisualEmbeddings to lock in the appearance, color, and other identity features of core visual objects. This provides a standard visual template for the drawing model, forcing the image generation model to converge to this unique standard visual model in each random denoising process, thus effectively solving the problem of inconsistent identities. At the same time, the text encoder reads the pagination scripts and prompts in the picture book blueprint and converts them into TextEmbeddings to guide the AI ​​in generating scene compositions that conform to the plot. This ensures that the generated feature data strictly follows the structured pattern of "core text content + visual style," allowing the two encoders to better perform their respective functions.

[0040] After transformation, both visual and textual feature data flow into the diffusion model denoising network UNet. UNet uses TextEmbeddings to guide the composition of the image (e.g., "Where is the bunny?", "Where is the carrot?") and VisualEmbeddings to guide entity rendering (e.g., "This bunny has white fur", "Short round tail"), ensuring visual consistency of the visual objects. It denoises the initial Gaussian noise, restores the image, and outputs the generated single-page image. The self-attention module is responsible for the logical coordination within the image, ensuring the natural integration of the background, lighting, and subject, maintaining the overall style of the picture book. The logic of picture book image generation can be found in [link to relevant documentation]. Figure 4 The example shown is for illustrative purposes only.

[0041] The picture book generation method provided by this invention responds to a picture book theme input by the user, conceives a story for the picture book theme, generates a picture book generation blueprint containing pagination scripts and visual object classification information, extracts visual objects that appear more frequently than a preset frequency requirement from the picture book generation blueprint, and constructs corresponding visual feature data for these visual objects. Then, based on the visual feature data and the script information corresponding to each page, a picture book image set is generated. By strictly following the pre-constructed visual features to generate picture book images, the method achieves visual consistency of the same core visual object across all pages of the picture book, establishes visual connections between pages, strengthens the narrative coherence of the picture book, and achieves precise adaptation of visual style and content plot, thereby improving the quality of picture book generation.

[0042] This embodiment provides a picture book generation method, which can be used in a picture book generation system. Figure 5 This is a flowchart of a picture book generation method according to an embodiment of the present invention, such as... Figure 5 As shown, the process includes the following steps: Step S501: In response to the user-inputted picture book theme, a story is conceived based on the picture book theme, generating a picture book generation blueprint containing pagination scripts and visual object classification information. The visual object classification information includes at least one visual object and the corresponding category for each visual object, with the categories derived based on the frequency of the visual object's appearance in the story.

[0043] Specifically, the following detailed approach can be used to generate a picture book generation blueprint containing pagination scripts and visual object classification information: Based on the picture book theme, a story retrieval and conceptualization are performed using a pre-defined knowledge base to generate a story outline; the story outline is expanded into a pagination script, and a corresponding script description is generated for each page of the pagination script; visual objects in the pagination script are identified, and all appearing visual objects are marked; based on the frequency of each visual object's appearance across multiple pages, each visual object is categorized to obtain visual object classification information; the script description for each page is converted into prompts that conform to the input specifications of the image generation model, with the input specifications including at least a token length limit; the pagination script, visual object classification information, and prompts corresponding to each page are integrated to generate the picture book generation blueprint.

[0044] The story architecture agent 1 of this embodiment has a built-in picture book story resource library. This resource library contains structured information such as story frameworks, character settings, plot logic, and scene elements corresponding to various picture book themes. It can employ a retrieval-augmented generation knowledge base architecture. The Base (RAG) enables precise retrieval and information reuse. Based on a picture book theme, it can call upon story materials, character settings, and plot templates related to the theme from the knowledge base, filter core materials that are suitable for the target type of picture book style (such as children's picture books, sketching style, etc.) and logically coherent, and then combine the retrieved materials to develop a story, determine the core storyline, core characters, and key plots, and generate a concise and clear story outline. The story outline is then broken down into paginated scripts according to preset picture book page number and other decomposition rules and retrieved plot details, ensuring that the paginated scripts not only conform to the logic of the story outline but also have visual appeal. For each page of the paginated script, combined with visual element materials from the knowledge base, a detailed script description is generated, clearly defining the character actions, scene elements, and picture composition of the page, ensuring that the script description can accurately guide the subsequent image generation.

[0045] This invention can identify visual objects (including but not limited to characters, scenes, buildings, and scene elements) in pagination scripts, mark all appearing visual objects, and classify each visual object based on its frequency of appearance across multiple pages. For example, core objects must appear at a frequency of ≥80%, while one-time supporting objects must appear at a frequency of <30%. This is just an example.

[0046] This invention clarifies the input specifications for the SDXL model, including compliance with SDXL principles and adherence to token length limits. For example, by constraining all output text instructions (prompts) to a safe range of 77 tokens, the core information is extracted, key descriptions of the main character and style are preserved, secondary elements are simplified, and attention-grabbing issues are fundamentally avoided, ensuring the complete presentation of the visual content. During application, the script descriptions for each page are converted into prompts, extracting core information (main objects, actions, style, etc.) and eliminating redundant expressions. This ensures that the prompts retain core visual and plot information without exceeding the token length limit. Simultaneously, the prompt conversion must integrate the overall style requirements of the picture book (taking a cartoon picture book style as an example) to ensure a consistent style across all pages. Finally, the pagination script, visual object classification information, and prompts from each page are integrated to generate a picture book generation blueprint, typically in JSON or XML format.

[0047] This invention achieves precise matching of picture book themes and story materials through a built-in knowledge base, improving the logic and efficiency of story conception, ensuring the quality of picture book plots, and clearly converting prompt words according to the model's input specifications. This ensures that the prompt words fully retain the core visual and plot information, avoids token overload, avoids attention distraction problems, and improves image generation accuracy.

[0048] In one optional implementation, if the current page is a simple page that conforms to preset generation constraints, the script description can be directly converted into prompt words. If the current page is a complex page that does not conform to preset generation constraints, the script description corresponding to the current page can be decomposed into multiple sub-descriptions; each sub-description can be converted into prompt words that conform to the input specifications of the image generation model.

[0049] In this embodiment of the invention, if a page's script description is found to contain multi-character interactions, complex actions, or multiple scene elements, thus failing to meet preset generation constraints, the current page is determined to be a complex page that does not meet these constraints. The script description corresponding to the current page can then be decomposed into multiple sub-descriptions. The decomposition principle can be that each sub-description corresponds to a core scene / action / element combination, ensuring that each sub-description is concise and its core information is clear. Each sub-description can then be converted into sub-hint words that conform to the input specifications of the image generation model, ensuring that each sub-hint word retains the core information of its corresponding sub-description. The number of tokens for each sub-hint word is verified one by one to ensure that none exceed the 77-token limit. Simultaneously, all sub-hint words are associated and integrated, and the corresponding complex pages are labeled to ensure that the combined sub-hint words can completely reconstruct the script description of the complex page. When inputting into the SDXL model, all sub-hint words are called synchronously to achieve accurate generation of complex pages. Alternatively, multiple sub-hint words can be used to generate corresponding picture book images separately. This is just an example; for the specific blueprint generation process logic, please refer to [link to relevant documentation]. Figure 6 As shown.

[0050] When encountering scenarios that are too complex for a single generation, this invention can automatically break down the script description into multiple sub-descriptions, ensuring that the instruction prompts for each page of the picture book are concise and clear, thereby improving the quality of the picture book images.

[0051] Step S502: Extract a list of visual objects for the target category from the picture book generation blueprint, and construct corresponding visual feature data for each visual object in the list. Visual objects for the target category are those that appear more frequently than a preset frequency requirement across multiple pages in the pagination script.

[0052] Specifically, step S502 includes: Step S5021: Generate a corresponding standard image for the current visual object.

[0053] Step S5022: Extract the first type of feature information and the second type of feature information from the standard image. The first type of feature information is used to represent the essential identity information of the visual object, and the second type of feature information is used to represent the appearance information of the visual object.

[0054] Step S5023: Perform vector concatenation and fusion processing on the first type of feature information and the second type of feature information to obtain the visual feature data of the current visual object.

[0055] In this embodiment of the invention, the visual object of the target category in the picture book generation blueprint is read. First, it can be determined whether the visual feature data of the current visual object already exists in the GIE feature library. If not, a "style-neutral reference image" strategy can be automatically constructed, calling a preset text-based image model to generate a standard image (i.e., a high-quality reference image, or "final portrait") for the current visual object. This ensures that the GIE itself does not carry excessive scene and style information, thereby achieving separation of identity and scene / style at the feature level. Then, the first type of feature information can be extracted through an identity recognition model (such as InsightFace). The first type of feature information is used to characterize the essential identity information of a visual object, mainly capturing the character's pure facial bone structure, facial proportions, and other features. This part of the feature information is highly robust to changes in lighting and pose. At the same time, the second type of feature information can be extracted from standard image images using a general visual encoder (such as CLIP Image Encoder). The second type of feature information is used to characterize the appearance of visual objects. This includes not only the face, but more importantly, it captures "macroscopic appearance" features such as hairstyle, clothing color, and overall demeanor. This is just one example; for the logical architecture of constructing visual feature data, please refer to [link to relevant documentation]. Figure 7 As shown.

[0056] In this embodiment of the invention, the first type of feature information and the second type of feature information can be vector-concatenated and fused to obtain the visual feature data of the current visual object. In the subsequent generation of picture book spread pages, this data represents the unique and immutable "digital identity fingerprint" of the character, as shown in the following formula:

[0057] in, This represents the visual feature data of the currently viewed object; This represents a vector concatenation operation; () indicates a fusion projection model (usually a multilayer perceptron).

[0058] This invention performs deep mapping and fusion of essential identity features and outward appearance features, preserving the core essence of identity and rich details of appearance, thereby improving the accuracy of picture book generation.

[0059] Step S503: Based on the visual feature data and the script information corresponding to each page, a picture book image set is generated. The images in the picture book image set maintain visual consistency when presenting visual objects of the target category.

[0060] In one optional implementation, when it is detected that multiple consecutive pages in the picture book generation blueprint indicate that they belong to a continuous action scene, the page image of the previous page can be obtained as a reference frame when generating the image of the current page. During the denoising process of the image generation model, the mask region of the target visual object in the reference frame is extracted, and the target visual object is the object performing the continuous action. For the mask region, the pixel-level dense spatial mapping relationship between the current page image and the reference frame with respect to the target visual object is calculated. According to the pixel-level dense spatial mapping relationship, the preset visual detail features belonging to the target visual object in the reference frame are fused into the features of the current page image according to the preset weight. Based on the pixel-level dense spatial mapping relationship, the appearance style features corresponding to the target visual object in the reference frame are shared to the current page image through a shared attention mechanism. The appearance style features include at least illumination distribution and color.

[0061] In this embodiment of the invention, when the blueprint for generating a picture book indicates that multiple consecutive pages being processed belong to a continuous action scene, the cross-image identity consistency generation engine 2 can activate the built-in ConsiStory relationship synchronization module. This module establishes a temporary "cross-page memory" by performing cross-image attention sharing and feature injection during batch generation, ensuring that dynamic details such as light and shadow, clothing folds, etc., of visual objects in continuous actions maintain logical continuity. Specifically, when generating the image of the current page, the image of the previous page can be obtained as a reference frame. During the denoising process of the image generation model, the mask region of the target visual object in the current page image and the reference frame can be extracted respectively. For the mask region, the pixel-level dense spatial mapping relationship of the target visual object between the current page image and the reference frame is calculated.

[0062] Specifically, the pixel-level dense spatial mapping relationship of the target visual object can be calculated through the following steps: extract the first deep semantic feature map and the second deep semantic feature map corresponding to the current page image and the reference frame respectively during the denoising process; calculate the feature similarity between each feature position in the first deep semantic feature map and each feature position in the second deep semantic feature map; based on the feature similarity, match the corresponding feature position with the highest similarity in the reference frame for each feature position in the current page image to construct the pixel-level dense spatial mapping relationship.

[0063] In this embodiment of the invention, the first deep semantic feature maps corresponding to the current page image and the reference frame during the denoising process can be extracted respectively. Second deep semantic feature map Then, the similarity between each feature position in the first deep semantic feature map and each feature position in the second deep semantic feature map is calculated. Based on the feature similarity, the corresponding feature position with the highest similarity in the reference frame is matched for each feature position in the current page image to construct a pixel-level dense spatial mapping relationship, as shown in the following formula:

[0064] in, This represents the calculated pixel-level dense spatial mapping relationship; This represents the optimal pixel matching criterion; This represents the pixel-level spatial position (any single pixel) in the current page image i. Indicates the pixel-level spatial position in reference frame j; Indicates the similarity between feature locations; This represents the masked region of the target visual object in reference frame j, used to exclude background interference and ensure that the corresponding relationship is found only for the character itself.

[0065] In continuous actions in picture books (such as Xiaoming going from standing to running), the model needs to accurately determine which position in the reference frame corresponds to Xiaoming's left sleeve in the current page image. By calculating the distance between features, the model can accurately find the densely corresponding coordinates of the same semantic part of the character in the two images, providing coordinate anchors for subsequent detail transfer.

[0066] This invention extracts semantic feature maps and performs subsequent dense spatial mapping, which can accurately capture the semantic part information of the target visual object, unaffected by changes in lighting. Then, based on the dense spatial mapping relationship, it calculates the similarity of each feature position in the two feature maps, achieving pixel-level comprehensive correspondence. This ensures that every detail of the core visual object can be accurately reproduced across pages, further enhancing the visual coherence of the picture book across pages.

[0067] After obtaining the pixel-level dense spatial mapping relationship, this invention can directly inject specific visual details (such as specific clothing textures or local clothing wrinkles) on the character in the reference frame into the feature positions of the current page image according to a weight ratio, truly realizing dynamic detail continuity between multiple frames. Specifically, based on the pixel-level dense spatial mapping relationship, corresponding part features corresponding to each position in the current page image are extracted from the deep semantic feature map of the reference frame. The original denoising features of the current page image and the extracted corresponding part features are fused according to a preset injection intensity coefficient to generate fused denoising features. The fused denoising features are used as input for subsequent denoising processes so that the current page image inherits the microscopic detail features in the reference frame at the corresponding positions. The preset visual detail features belonging to the target visual object in the reference frame can be fused into the features of the current page image according to a preset weight using the following formula:

[0068] in, This indicates that the fused features will continue to participate in subsequent noise reduction. Indicates the feature injection weights; This represents the original features during the denoising process of the current page image; This represents the corresponding feature that was precisely extracted from the reference frame.

[0069] This invention can obtain the query vector of the current generated frame in the attention layer of the diffusion model based on pixel-level dense spatial mapping relationships; the key vector and value vector of the current generated frame are concatenated with the key vector and value vector of the reference frame to generate expanded key vector and expanded value vector; based on the attention calculation of the query vector and the expanded key vector, features are aggregated from the expanded value vector, so that the current generated frame can simultaneously pay attention to its own features and the features of the reference frame during the generation process, thereby realizing the sharing of illumination features and global texture features. Specifically, the sharing of appearance style features can be achieved through the following formula:

[0070] in, The query vector representing the image on the current page; and These represent the expanded key and value vectors, respectively, which are the concatenation of key-value pairs of features from the current page image i and the reference frame j (i.e., the attention mechanism can "see" both images simultaneously).

[0071] This invention forces the image generation model to borrow attention from the features of the same character on the previous page when generating the character on the current page by splicing and expanding the key vector and value vector. This ensures a high degree of consistency in the overall lighting distribution, color tendency, and other macroscopic styles of the character in continuous shots. Furthermore, it uses mathematical formulas to force and constrain the random sampling of the model, ensuring a high degree of consistency in the physical details of the core protagonist in continuous actions.

[0072] This invention uses the previous page image as a reference frame and integrates the visual detail features of the target visual object in the reference frame into the features of the current page image according to a preset weight, ensuring the natural continuation of details. Furthermore, through a shared attention mechanism, the appearance style features of the target visual object in the reference frame are shared to the current page, ensuring that the appearance style attributes remain consistent in the same continuous action sequence. This guarantees the accurate reproduction of the core characters and the overall visual unity, while also enhancing narrative coherence and the visual smoothness of continuous actions, thereby improving the accuracy, adaptability, and efficiency of picture book generation.

[0073] In one optional implementation, feature consistency verification and semantic verification can also be performed on the current page images in the picture book image set. Feature consistency verification is used to evaluate the feature similarity between the visual object to be tested in the current page image and the visual object to be tested in the standard image image; semantic verification is used to evaluate the semantic conformity between the current page image and the script description of the corresponding page; if all page images in the picture book image set pass the feature consistency verification and semantic verification, the final picture book image set is output.

[0074] After receiving the picture book image set, the automated verification and correction module 3 of this embodiment of the invention can perform dual verification on the current page image, including but not limited to feature consistency verification and semantic verification. Feature consistency verification is used to evaluate the feature similarity between the visual object to be tested in the current page image and the visual object to be tested in the standard image. Semantic verification is used to evaluate the semantic conformity between the current page image and the script description of the corresponding page. Specifically, a specialized feature extraction model (such as InsightFace) can be used to extract the facial feature vector of the visual object to be tested in the generated image and the facial feature vector of the visual object to be tested in its standard image. Then, the feature similarity is calculated using the following formula:

[0075] in, This represents the facial feature vector of the visual object to be tested, extracted from the generated image. This represents the baseline vector of facial features of the visual object under test extracted from the standard image. This represents the cosine similarity between the two in a high-dimensional feature space.

[0076] This invention maps the generated image and the standard image to the same feature space for inner product operation, quantitatively evaluating the similarity between the newly generated visual object and the standard image. This transforms subjective similarity judgment into a rigorous and fast mathematical comparison that the system can perform, forming the core of automated quality inspection.

[0077] This invention can utilize advanced multimodal large models (such as qwen-vl) to compare the generated page image with the original text prompts to determine whether the screen content accurately responds to the instructions logically and semantically, as shown in the following formula:

[0078] in, A consistency score out of 100, indicating whether semantic validation is successful.

[0079] This invention maps scores to a standard range of 0%-100%. In practical applications, a hard threshold can be set (taking 85% as an example). If the value is below the hard threshold, it indicates that the semantic verification has failed and will automatically trigger the subsequent correction process, ensuring that the final output picture book maintains consistency and accuracy in the appearance of the characters.

[0080] In this embodiment of the invention, if all picture book images in the picture book image set pass feature consistency verification and semantic verification, the final version of the picture book image set, which has passed verification and quality certification, and is highly consistent and accurate, can be output.

[0081] This invention performs feature consistency verification and semantic verification on the generated page images, which can automatically complete quality verification before outputting the picture book, improve the overall reliability of the picture book images, and feature consistency verification can quantitatively assess whether the character identity, appearance and structure are consistent, ensuring that visual objects are stable and consistent across pages, and semantic verification can ensure that the content of the picture strictly fits the script description, ultimately ensuring the high quality and consistency of the output picture book images.

[0082] In one optional implementation, if the current page image fails the feature consistency check and / or semantic check, an image editing instruction can be generated based on the reason for the check failure; the image editing engine can be invoked to make local or global adjustments to the current page image based on the image editing instruction; the adjusted page image can be repeatedly subjected to the steps of feature consistency check and semantic check until the current page image passes the feature consistency check and semantic check.

[0083] In this embodiment of the invention, if the current page image fails the feature consistency check and / or semantic check, the Large Language Model (LLM) module can receive the failure details (including check type, failure index, and specific deviation content) output by the check module. Combined with the current page script description and standard image features, the failure reason can be accurately analyzed. Based on the analyzed failure reason, the LLM module generates targeted and executable image editing instructions. These instructions must clearly define the object to be corrected, the direction of correction, and the correction standard to ensure the correction engine can accurately identify and execute them. For example, an image editing instruction could be "adjust the fur color of the white rabbit on the current page to white to match the standard image, without changing the character's posture, actions, or other elements of the image." This is just an example; the image editing instructions generated by the LLM module... The editing command is input into the preset image editing engine (COSXL_edit) to start the image correction process. Based on the editing command, the correction engine makes targeted adjustments to the current page image, prioritizing local fine-tuning to avoid global adjustments that could damage the original style and narrative expression of the image. For example, only the fur color might be adjusted without affecting other elements of the image. After the correction engine completes the adjustments, it outputs the adjusted current page image and can also record the corrections simultaneously. The adjusted current page image can then be automatically sent back to the verification module for feature consistency verification and semantic verification, achieving iterative closed-loop verification until both verifications are passed. The closed-loop workflow of verification-correction can be found in [link to documentation]. Figure 8 As shown.

[0084] This invention automatically triggers a correction process after a verification failure, eliminating the need for manual intervention and improving quality control efficiency. Specifically, it accurately locates the cause of the failure, generates targeted editing instructions, and performs subsequent corrections. This achieves dynamic generation of precise editing instructions that fit the failure scenario, avoiding blind corrections and improving the overall efficiency and quality stability of picture book generation.

[0085] In a specific embodiment, the user provides a simple picture book theme. The story architecture agent 1 queries its built-in RAG knowledge base, combines the retrieved knowledge to conceive a complete story outline, expands the story outline into a detailed pagination script, and automatically analyzes and marks all visual objects, classifying them into core classes that need to maintain consistency and supporting character classes that appear less frequently. The script descriptions of the pagination script images are compiled into structured prompts that conform to the SDXL principle and strictly adhere to token restrictions, and integrated to obtain the picture book generation blueprint. The cross-image identity consistency generation engine 2 reads the core class visual objects in the picture book generation blueprint and automatically uses the "style neutral" prompt. A standard image is generated, and visual feature data is extracted from it and stored in the GIE feature library. Based on script descriptions and prompts, and combined with the visual feature data, images are generated to obtain a preliminary picture book image set. The automated verification and correction module 3 performs automated double verification on the preliminary picture book image set. If all images pass verification, the final picture book image set is output. If any image fails verification, a correction program is automatically initiated. The LLM module generates precise editing instructions based on the reason for the failure and calls the editing engine to fine-tune the image. The corrected image is then verified again, forming a closed loop until all images meet the standards. All images that pass verification are compiled and output. The picture book generation logic architecture can be found in [reference needed]. Figure 9 As shown, please refer to the above embodiments for detailed explanation, which will not be repeated here.

[0086] This embodiment also provides a vehicle, such as Figure 10 As shown, the vehicle includes an interactive interface 1001 and a controller 1002. The interactive interface 1001 is used to display a set of picture book images. The controller 1002 includes a memory and a processor, which are interconnected. The memory stores computer instructions, and the processor executes the picture book generation method described above.

[0087] This embodiment also provides a picture book generation device for implementing the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0088] This embodiment provides a picture book generation device, such as... Figure 11As shown, it includes: a story conception module 1101, used to respond to the picture book theme input by the user, to conceive a story for the picture book theme, and generate a picture book generation blueprint containing pagination scripts and visual object classification information. The visual object classification information includes at least one visual object and the category corresponding to each visual object. The category is obtained based on the frequency of the visual object in the story; a visual feature construction module 1102, used to extract a list of visual objects of the target category from the picture book generation blueprint, and to construct corresponding visual feature data for each visual object in the list of visual objects. The visual objects of the target category are visual objects that appear more frequently than a preset frequency requirement in multiple pages of the pagination script; and a picture book generation module 1103, used to generate a picture book image set based on the visual feature data and the script information corresponding to each page. The images in the picture book image set maintain visual consistency when presenting visual objects of the target category.

[0089] In some optional implementations, the picture book generation module 1103 includes: an outline generation unit, used to retrieve and conceive a story based on the picture book theme by calling a preset knowledge base; an object tagging unit, used to expand the story outline into a pagination script and generate a corresponding script description for each page of the pagination script; to identify visual objects in the pagination script and tag all appearing visual objects; an object classification unit, used to classify each visual object according to the frequency of its appearance on multiple pages to obtain visual object classification information; a prompt word conversion unit, used to convert the script description of each page into prompt words that conform to the input specifications of the image generation model, the input specifications including at least a token length limit; and to integrate the pagination script, visual object classification information, and prompt words corresponding to each page to generate a picture book generation blueprint.

[0090] In some optional implementations, when the current page is a complex page that does not conform to the preset generation constraints, the prompt word conversion unit includes: a description decomposition subunit, used to decompose the script description corresponding to the current page into multiple sub-descriptions; and a prompt word conversion subunit, used to convert each sub-description into a prompt word that conforms to the input specification of the image generation model.

[0091] In some optional implementations, the visual feature construction module 1102 includes: an image generation unit for generating a corresponding standard image image of the current visual object; a feature extraction unit for extracting a first type of feature information and a second type of feature information from the standard image image, wherein the first type of feature information is used to characterize the essential identity information of the visual object, and the second type of feature information is used to characterize the appearance information of the visual object; and a feature fusion unit for performing vector concatenation and fusion processing on the first type of feature information and the second type of feature information to obtain the visual feature data of the current visual object.

[0092] In some optional implementations, when the picture book generation blueprint indicates that multiple consecutive pages being processed belong to a continuous action scene, the picture book generation module 1103 further includes: a reference frame acquisition unit, used to acquire the page image of the previous page as a reference frame when generating the image of the current page; a mask region extraction unit, used to extract the mask region of the target visual object in the reference frame during the denoising process of the image generation model, wherein the target visual object is the object performing the continuous action; a mapping relationship calculation unit, used to calculate the pixel-level dense spatial mapping relationship between the current page image and the reference frame regarding the target visual object for the mask region; a feature fusion unit, used to fuse the preset visual detail features belonging to the target visual object in the reference frame into the features of the current page image according to the preset weights based on the pixel-level dense spatial mapping relationship; and a feature sharing unit, used to share the appearance style features corresponding to the target visual object in the reference frame to the current page image through a shared attention mechanism based on the pixel-level dense spatial mapping relationship, wherein the appearance style features include at least illumination distribution and color.

[0093] In some optional implementations, the mapping relationship calculation unit includes: a feature map extraction subunit, used to extract the first deep semantic feature map and the second deep semantic feature map corresponding to the current page image and the reference frame respectively during the denoising process; a feature similarity calculation subunit, used to calculate the feature similarity between each feature position in the first deep semantic feature map and each feature position in the second deep semantic feature map; and a feature matching subunit, used to match each feature position in the current page image with the corresponding feature position in the reference frame with the highest similarity based on the feature similarity, so as to construct a pixel-level dense spatial mapping relationship.

[0094] In some optional implementations, the picture book generation device further includes: a verification module, used to perform feature consistency verification and semantic verification on the current page image in the picture book image set; the feature consistency verification is used to evaluate the feature similarity between the visual object to be tested in the current page image and the visual object to be tested in the standard image; the semantic verification is used to evaluate the semantic conformity between the current page image and the script description of the corresponding page; and an image output module, used to output the final picture book image set if all page images in the picture book image set pass the feature consistency verification and semantic verification.

[0095] In some optional implementations, if the current page image fails the feature consistency check and / or semantic check, the picture book generation device further includes: an image editing instruction generation module, used to generate image editing instructions based on the reason for the check failure; an image adjustment module, used to call the image editing engine and make local or global adjustments to the current page image based on the image editing instructions; and an image verification module, used to repeatedly perform the feature consistency check and semantic check steps on the adjusted page image until the current page image passes the feature consistency check and semantic check.

[0096] The picture book generation apparatus provided in this embodiment of the invention can execute the picture book generation method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method. Further functional descriptions of the various modules and units described above are the same as in the corresponding embodiments described above, and will not be repeated here.

[0097] Figure 12 This is a schematic diagram of the structure of a controller 1002 provided in an embodiment of the present invention.

[0098] The following is a detailed reference. Figure 12 The diagram illustrates a structural schematic suitable for implementing a controller in an embodiment of the present invention. The controller may include a processor (e.g., a central processing unit, a graphics processing unit, etc.) 1201, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 1202 or a program loaded from memory 1208 into random access memory (RAM) 1203. The RAM 1203 also stores various programs and data required for controller operation. The processor 1201, ROM 1202, and RAM 1203 are interconnected via a bus 1204. An input / output (I / O) interface 1205 is also connected to the bus 1204.

[0099] Typically, the following devices can be connected to I / O interface 1205: input devices 1206 including, for example, a touchscreen, touchpad, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc.; output devices 1207 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; memory 1208 including, for example, magnetic tape, hard disk, etc.; and communication devices 1209. Communication device 1209 allows the controller to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 12 A controller with various devices is shown, but it should be understood that it is not required to implement or have all of the devices shown, and may alternatively implement or have more or fewer devices.

[0100] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 1209, or installed from a memory 1208, or installed from a ROM 1202. When the computer program is executed by the processor 1201, it performs the functions defined in the picture book generation method of the embodiments of the present invention.

[0101] Figure 12 The controller shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.

[0102] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, the picture book generation method shown in the above embodiments is implemented.

[0103] A portion of this invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the invention through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.

[0104] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and all such modifications and variations fall within the defined scope.

Claims

1. A method for generating picture books, characterized in that, The method includes: In response to the picture book theme input by the user, a story is conceived for the picture book theme, and a picture book generation blueprint containing pagination scripts and visual object classification information is generated. The visual object classification information includes at least one visual object and the category corresponding to each visual object. The category is obtained based on the frequency of the visual object in the story. Extract a list of visual objects of the target category from the picture book generation blueprint, and construct corresponding visual feature data for each visual object in the list of visual objects. The visual objects of the target category are visual objects that appear more frequently than a preset frequency requirement in multiple pages of the pagination script. Based on the visual feature data and the script information corresponding to each page, a picture book image set is generated. The images in the picture book image set maintain visual consistency when presenting visual objects of the target category.

2. The method according to claim 1, characterized in that, The process of conceiving a story for the picture book theme and generating a picture book generation blueprint that includes pagination scripts and visual object classification information includes: Based on the picture book theme, a story outline is generated by calling a preset knowledge base to retrieve and brainstorm stories. Expand the story outline into a paginated script, and generate a corresponding script description for each page of the paginated script; Identify and mark all visual objects that appear in the pagination script; Based on the frequency of each visual object appearing on multiple pages, each visual object is categorized to obtain visual object classification information; The script description for each page is converted into prompts that conform to the input specifications of the image generation model, which include at least a token length limit. By integrating the pagination script, the visual object classification information, and the prompts corresponding to each page, a picture book generation blueprint is generated.

3. The method according to claim 2, characterized in that, When the current page is a complex page that does not conform to the preset generation constraints, the step of converting the script description of each page into prompt words that conform to the input specifications of the image generation model includes: The script description corresponding to the current page is broken down into multiple sub-descriptions; Each sub-description is converted into a prompt word that conforms to the input specification of the image generation model.

4. The method according to claim 2, characterized in that, The step of constructing corresponding visual feature data for each visual object in the list of visual objects includes: Generate a standard image corresponding to the current visual object; Extract a first type of feature information and a second type of feature information from the standard image, wherein the first type of feature information is used to characterize the essential identity information of the visual object, and the second type of feature information is used to characterize the appearance information of the visual object; The first type of feature information and the second type of feature information are vector-concatenated and fused to obtain the visual feature data of the current visual object.

5. The method according to claim 1, characterized in that, When it is detected that multiple consecutive pages currently being processed in the picture book generation blueprint belong to a continuous action scene, the generation of the picture book image set further includes: When generating the image for the current page, the image of the previous page is used as a reference frame. In the denoising process of the image generation model, the mask region of the target visual object in the reference frame is extracted, and the target visual object is an object that performs continuous actions; For the masked region, calculate the pixel-level dense spatial mapping relationship between the current page image and the reference frame regarding the target visual object; Based on the pixel-level dense spatial mapping relationship, the preset visual detail features belonging to the target visual object in the reference frame are fused into the features of the current page image according to the preset weights; Based on the pixel-level dense spatial mapping relationship, the appearance style features corresponding to the target visual object in the reference frame are shared to the current page image through a shared attention mechanism. The appearance style features include at least illumination distribution and color.

6. The method according to claim 5, characterized in that, The calculation of the pixel-level dense spatial mapping relationship between the current page and the reference frame regarding the target visual object includes: Extract the first and second deep semantic feature maps corresponding to the current page image and the reference frame during the denoising process, respectively; Calculate the feature similarity between each feature position in the first deep semantic feature map and each feature position in the second deep semantic feature map; Based on the feature similarity, each feature position in the current page image is matched with the corresponding feature position with the highest similarity in the reference frame to construct a pixel-level dense spatial mapping relationship.

7. The method according to claim 4, characterized in that, The method further includes: The current page image in the picture book image set is subjected to feature consistency verification and semantic verification. The feature consistency verification is used to evaluate the feature similarity between the visual object to be tested in the current page image and the visual object to be tested in the standard image image. The semantic verification is used to evaluate the semantic conformity between the current page image and the script description of the corresponding page. If all page images in the picture book image set pass the feature consistency check and semantic check, the final picture book image set is output.

8. The method according to claim 7, characterized in that, If the current page image fails the feature consistency check and / or semantic check, the method further includes: Based on the reason for the verification failure, generate image editing instructions; The image editing engine is invoked to make local or global adjustments to the current page image based on the image editing instructions. Repeat the steps of performing feature consistency verification and semantic verification on the adjusted page image until the current page image passes the feature consistency verification and semantic verification.

9. A picture book generation device, characterized in that, The device includes: The story conception module is used to respond to the picture book theme input by the user, to conceive a story for the picture book theme, and to generate a picture book generation blueprint containing pagination scripts and visual object classification information. The visual object classification information includes at least one visual object and the category corresponding to each visual object. The category is obtained based on the frequency of the visual object in the story. The visual feature construction module is used to extract a list of visual objects of the target category from the picture book generation blueprint, and construct corresponding visual feature data for each visual object in the list of visual objects. The visual objects of the target category are visual objects that appear more frequently than a preset frequency requirement in multiple pages of the pagination script. The picture book generation module is used to generate a picture book image set based on the visual feature data and the script information corresponding to each page. The images in the picture book image set maintain visual consistency when presenting visual objects of the target category.

10. A vehicle, characterized in that, The vehicle includes an interactive interface and a controller. The interactive interface is used to display a set of picture book images. The controller includes a memory and a processor. The memory and the processor are communicatively connected to each other. The memory stores computer instructions. The processor executes the computer instructions to perform the picture book generation method according to any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a computer to perform the picture book generation method according to any one of claims 1 to 8.

12. A computer program product, characterized in that, Includes computer instructions for causing a computer to perform the picture book generation method according to any one of claims 1 to 8.