Method and device for generating a panning shot, storage medium and electronic device

By leveraging the collaborative work of large language models and multimodal models, the system automatically generates storyboard panels for animated series, solving the problem of fragmented AI animated series production tools and achieving efficient and coherent storyboard panel generation, thereby improving creative efficiency and quality.

CN122134865APending Publication Date: 2026-06-02SHANGHAI IQIYI NEW MEDIA TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI IQIYI NEW MEDIA TECH CO LTD
Filing Date
2026-02-10
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing AI animation production tools are highly fragmented, leading to frequent switching between creation tools, making it difficult to maintain consistency in visual elements and creative coherence, thus affecting efficiency and quality.

Method used

The script text is structured and decomposed using a large language model. Combined with a multimodal large model and an image generation model, scene design images and character design images are generated. The automatic generation of storyboard images is achieved through text-to-image model and image editing model.

Benefits of technology

It achieves unified scheduling from script text to storyboard images, reduces the cost of tool switching and manual debugging during the creation process, improves the consistency and generation efficiency of storyboard images, and enhances the overall efficiency and quality of AI comic book production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122134865A_ABST
    Figure CN122134865A_ABST
Patent Text Reader

Abstract

This application relates to a method, apparatus, storage medium, and electronic device for generating storyboard images for comics. The method includes: performing structured parsing and storyboard decomposition on the script text to obtain scene design description information, character design information, and shot storyboard description information for each shot; combining this with reference art style images to generate scene design images and character design images for each character; analyzing and predicting information about physical objects and camera positions in the scene to generate shot-level scene background images for each shot; performing semantic parsing on the shot-level scene background image, shot storyboard description information, and corresponding character design images for the target shot to generate image description information for the target shot; and generating the storyboard image for the target shot based on the image description information, shot-level scene background image, and character design images. This application solves the technical problem of severe fragmentation in comic production tools.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of multimedia intelligent processing technology, and in particular to a method, apparatus, storage medium, and electronic device for generating comic book storyboard images. Background Technology

[0002] With the rapid development of generative artificial intelligence technology, AI-generated comic book production is experiencing explosive growth. Currently, a creative team of only 6-8 people can complete the entire adaptation and production of a top-tier IP comic book, significantly lowering the barrier to content production. However, while efficiency has improved, the existing AI comic book creation ecosystem has exposed significant technical bottlenecks, particularly in the highly fragmented nature of creative tools. Current production processes typically employ a "toolchain" workflow, requiring creators to frequently switch between multiple independent AI tools to complete tasks such as script analysis, character design, scene generation, and storyboard construction. This model not only fragments the workflow but also places high technical demands on creators, requiring dedicated personnel to repeatedly adjust model parameters and prompts to ensure consistent output. Due to the lack of a unified scheduling and control mechanism, core visual elements such as character design, art style, color scheme, and lighting effects are difficult to maintain consistency when migrating creative assets between different tools, often requiring multiple rounds of rework and correction. Simultaneously, creators must constantly adapt to the technical limitations of different tools when conceiving plots and storyboards, frequently interrupting their creative thinking and affecting the overall creative continuity. Especially in AI animation production based on the first and last frame method, it is necessary to pre-build highly consistent storyboards, which further amplifies the efficiency and quality problems brought about by the collaboration of multiple tools. Summary of the Invention

[0003] This application provides a method, apparatus, storage medium, and electronic device for generating comic book storyboards, in order to solve the technical problem of severe fragmentation in comic book production tools.

[0004] Firstly, this application provides a method for generating storyboard images for a comic book, comprising: calling a target large language model to perform structured parsing and storyboard decomposition on the script text input by the user, obtaining scene design description information, character image design information of each character, and shot storyboard description information of each shot; calling a target image generation model to generate scene design images and character design images of each character based on the aforementioned scene design description information and character image design information, combined with the aforementioned reference art style image input by the user; and calling a target multimodal large model to generate images of physical objects in the scene based on the aforementioned scene design images and shot storyboard description information of each shot. The system analyzes and predicts information and camera positions, and calls the target image model to generate shot-level scene background images for each shot based on the prediction results and the aforementioned scene design images. It then calls the aforementioned target multimodal model to perform semantic analysis on the shot-level scene background images, shot storyboard descriptions, and corresponding character design images of the target shots, generating image description information for the target shots, where the target shot is any one of the shots. Finally, it calls the target image editing model to generate the storyboard images for the target shots based on the image description information, shot-level scene background images, and character design images.

[0005] Secondly, this application provides a device for generating storyboard images for a comic book, comprising: a parsing module, used to call a target large language model to perform structured parsing and storyboard decomposition on the script text input by the user, to obtain scene design description information, character image design information of each character, and shot storyboard description information of each shot, and to call a target image generation model to generate scene design images and character design images of each character based on the aforementioned scene design description information and character image design information, combined with the aforementioned reference art style image input by the user; and an analysis module, used to call a target multimodal large model to analyze the information of physical objects in the scene based on the aforementioned scene design images and shot storyboard description information of each shot. The system analyzes and predicts camera positions, and calls the target image model to generate shot-level scene background images for each shot based on the prediction results and the aforementioned scene design images. The first generation module calls the aforementioned target multimodal model to perform semantic analysis on the shot-level scene background images, shot storyboard description information, and corresponding character design images of the target shots to generate the image description information of the target shots, where the target shots are any one of the shots. The second generation module calls the target image editing model to generate the storyboard images of the target shots based on the image description information, shot-level scene background images, and character design images of the target shots.

[0006] As an optional example, the above-mentioned parsing module includes: a first extraction unit, used to perform structured parsing of the script text according to preset scene decomposition rules, so as to decompose the script text into at least one scene, and extract the corresponding scene semantic information for each scene to generate scene design description information corresponding to each scene; a second extraction unit, used to extract the appearance features, clothing features and personality features of each character based on the character description information in the script text, and generate character image design information for each character; and a decomposition unit, used to perform shot-level decomposition on each scene to generate at least one corresponding shot storyboard description information, wherein the shot storyboard description information includes the main subject of the shot, the characters appearing, the characters' behavior, the shot size information and the shot's expressive intent.

[0007] As an optional example, the above-mentioned parsing module includes: a first generation unit, used to generate scene design images corresponding to each scene based on the scene design description information and the reference style images, and to associate and store them with the corresponding scene identifiers, wherein the scene design images are used to represent the spatial structure, environmental elements and overall visual style of the scene; and a second generation unit, used to generate character design images corresponding to each character based on the character image design information and the reference style images, and to associate and store them with the corresponding character identifiers, wherein the character design images are simplified background or white background character images.

[0008] As an optional example, the above analysis module includes: a recognition unit, used to recognize visible physical objects contained in the scene based on the above scene design image, and obtain corresponding physical object category information and their spatial distribution information in the scene; a first prediction unit, used to jointly analyze the shot description information of each shot and the recognition result of the above scene design image, predict the set of physical objects to be presented in each shot and the relative positional relationship of the physical objects in the shot image, and obtain the physical object prediction result; and a second prediction unit, used to predict the camera position parameters corresponding to each shot based on the shot size information and shot expression intention in the above shot description information, wherein the above camera position parameters are used to indicate the shooting direction, shooting angle and shooting distance of the virtual camera.

[0009] As an optional example, the analysis module includes: a processing unit, configured to identify each unprocessed shot as the current shot, and call the target text image big model to perform the following processing on the current shot: based on the entity object prediction results and shot position parameters of the current shot, construct shot-level scene generation conditions that match the shot segmentation description information of the current shot; according to the shot-level scene generation conditions of the current shot, reconstruct or crop the scene design image to generate a shot-level scene background image that conforms to the shot position parameters of the current shot; and associate and store the shot-level scene background image of the current shot with its shot identifier.

[0010] As an optional example, the first generation module includes a third generation unit, which is used to analyze the spatial structure relationship, the relative positional relationship between characters and the scene, the standing position relationship, action state, and interaction relationship between characters of the target shot based on the shot-level scene background image, shot storyboard description information, and the character design images of each character therein. Based on the shot size information and shot expression intention in the shot storyboard description information of the target shot, it generates picture description information to describe the composition method, perspective characteristics, and emotional expression of the target shot.

[0011] As an optional example, the above-mentioned device further includes: a detection module, used to perform human target detection on the storyboard of the target shot after generating the storyboard of the target shot, and compare the number of detected human targets with the number of human targets in the storyboard description information of the target shot; a third generation module, used to trigger the regeneration of the storyboard of the target shot if the number of human targets is inconsistent; and a fourth generation module, used to call the target multimodal large model if the number of human targets is consistent, to perform consistency analysis on the storyboard of the target shot, the scene description information and the shot storyboard description information, and to correct the scene description information of the target storyboard and regenerate the storyboard based on the analysis results, or to confirm the current storyboard as the final storyboard.

[0012] Thirdly, this application provides a storage medium storing a computer program, wherein the computer program is executed by a processor to perform the above-described method for generating comic book storyboards.

[0013] Fourthly, this application also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the above-described method for generating comic book storyboards through the computer program.

[0014] The technical solutions provided in this application have the following advantages compared with the prior art: This application employs a target large-scale language model to perform structured parsing and storyboard decomposition of the user-input script text, obtaining scene design descriptions, character design information, and shot storyboard descriptions for each scene. It then calls a target image generation model to generate scene design images and character design images based on the aforementioned scene design descriptions and character design information, combined with the user-input reference art style images. Finally, it calls a target multimodal large-scale model to analyze and predict the information of physical objects and camera positions in the scene based on the scene design images and shot storyboard descriptions. Finally, it calls a target text-to-image large-scale model to generate images based on the prediction results and the aforementioned scene design images. The process involves generating shot-level scene background images for each shot; calling the aforementioned target multimodal large model to perform semantic analysis on the shot-level scene background images, shot storyboard descriptions, and corresponding character design images of the target shots, generating the scene description information for the target shots, where the target shots can be any one of the shots; and calling the target image editing model to generate the storyboard images for the target shots based on the scene description information, shot-level scene background images, and character design images. This method achieves automated generation from script text to storyboard images by uniformly scheduling the large language model, multimodal large model, text-to-image model, and image editing model. First, the large language model performs structured analysis and storyboard decomposition of the script text, generating scene design descriptions, character designs, and shot storyboard descriptions, and combining these with reference art styles to generate scene and character design images. Then, the multimodal large model analyzes the scene content and camera positions, and the text-to-image model generates shot-level scene backgrounds. Next, based on multimodal semantic analysis, scene description information is generated, and finally, the image editing model generates the storyboard images. This achieves unified collaboration of multiple model capabilities, reduces the cost of tool switching and manual debugging during the creation process, and improves the consistency and generation efficiency of storyboard images in terms of character, scene and shot expression, thereby solving the technical problem of severe fragmentation of comic book production tools. Attached Figure Description

[0015] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0016] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] One or more embodiments are illustrated by way of example with reference numerals in the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.

[0018] Figure 1 This is a flowchart of an optional method for generating comic book storyboard frames according to an embodiment of this application; Figure 2 This is a schematic diagram of a reference art style for an optional method of generating comic book storyboards according to an embodiment of this application; Figure 3 This is a scene design diagram of an optional method for generating comic book storyboards according to an embodiment of this application; Figure 4 This is a schematic diagram of character design images for an optional method of generating comic book storyboards according to an embodiment of this application; Figure 5 This is a schematic diagram of a scene background image at the lens level, representing an optional method for generating comic book storyboards according to an embodiment of this application. Figure 6 This is a schematic diagram of a storyboard for an optional method of generating storyboards for a comic strip according to an embodiment of this application. Figure 7 This is a schematic diagram of the structure of an optional comic strip storyboard generation device according to an embodiment of this application; Figure 8 This is a schematic diagram of an optional electronic device according to an embodiment of this application. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0020] The following disclosure provides numerous different embodiments or examples for implementing various structures of this application. To simplify the disclosure, specific examples of components and arrangements are described below. These are merely examples and are not intended to limit the scope of this application. Furthermore, reference numerals and / or letters may be repeated in different examples. Such repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed.

[0021] According to a first aspect of the embodiments of this application, a method for generating storyboard frames for an animated series is provided, optionally, as follows: Figure 1 As shown, the above method includes: S102, call the target large language model to perform structured parsing and storyboard decomposition on the script text input by the user, obtain scene design description information, character image design information of each character and shot storyboard description information of each shot, and call the target image generation model to generate scene design images and character design images of each character based on the scene design description information and character image design information, combined with the reference art style image input by the user. S104, call the target multimodal large model, analyze and predict the information of physical objects and camera positions in the scene based on the scene design image and the shot description information of each shot, and call the target text image large model to generate the shot-level scene background image of each shot based on the prediction results and the scene design image. S106, call the target multimodal large model, perform image semantic analysis on the shot-level scene background image, shot storyboard description information and the corresponding character design images of each character in the target shot, and generate the image description information of the target shot, where the target shot is any one of the shots; S108 calls the target image editing model to generate the storyboard of the target shot based on the image description information of the target shot, the shot-level scene background image, and the character design image.

[0022] Optionally, in this embodiment, a method for generating storyboard frames for an animated series is provided, aiming to solve the problems of severe tool fragmentation, low efficiency in storyboard construction, and difficulty in ensuring image consistency in the existing AI animated series production process. By introducing multiple artificial intelligence models and performing unified scheduling, efficient and continuous generation from script text to storyboard frames can be achieved. The method includes the following steps.

[0023] First, the target large language model is invoked to perform structured parsing and scene breakdown of the user-input script text. The large language model performs semantic understanding of the script text, breaking it down into several scenes and shots according to pre-defined script parsing rules. It extracts and generates scene design descriptions for each scene, character design information for each character in the script, and scene storyboard descriptions for each shot. The character design information describes the character's appearance, clothing, and overall style, while the scene storyboard descriptions describe the subject of the shot, the characters appearing, their actions, the shot size, and the shot's intended meaning. Then, based on the scene design descriptions and character design information, and combined with a reference art style image input by the user (see the example image below), the process is repeated. Figure 2As shown, the target image generation model is invoked to generate scene design images corresponding to each scene and character design images corresponding to each character, thereby constructing basic visual assets. A schematic diagram of the scene design images is shown below. Figure 3 As shown, the character design illustration is as follows: Figure 4 As shown.

[0024] Secondly, the target multimodal large model is invoked to analyze and predict the information of physical objects and camera positions in the scene based on the scene design image and the shot description information of each shot. The multimodal large model, through visual understanding of the scene design image, identifies the categories and spatial distribution information of physical objects in the scene, and, combined with the shot description information, predicts the set of physical objects to be presented in each shot and the corresponding camera position parameters. Based on this, the target text-based large model is invoked to reconstruct or adjust the perspective of the scene design image based on the predicted physical object information, camera position parameters, and scene design image, generating a shot-level scene background image corresponding to each shot. A schematic diagram of the shot-level scene background image is shown below. Figure 5 As shown, this enables the generation of background images from different camera angles within the same scene.

[0025] Then, for each target shot, the target multimodal big model is invoked to perform semantic analysis on the corresponding shot-level scene background image, shot storyboard description information, and character design images of each character. The multimodal big model analyzes the spatial relationship between characters and scene by jointly analyzing the character design images and scene background images, and combines the shot storyboard description information to determine the standing relationships, action states, character interaction relationships, and composition of the characters in the target shot, thereby generating image description information to describe the overall composition, perspective characteristics, and emotional expression of the target shot.

[0026] Finally, the target image editing model is invoked to fuse the character image with the scene background based on the image description information of the target shot, the shot-level scene background image, and the character design image, generating the corresponding storyboard for the target shot. A schematic diagram of the storyboard is shown below. Figure 6 As shown. By following the steps above, the storyboard frames for each shot can be generated sequentially, ultimately resulting in a complete sequence of animated film storyboard frames.

[0027] Optionally, in this embodiment, by introducing the collaborative work of a large language model, a multimodal large model, a text-to-image model, and an image editing model, unified scheduling and automated generation of script text to storyboard images are achieved, effectively reducing the efficiency loss caused by frequent switching of multiple tools during the creation process; by introducing a unified semantic parsing and image description mechanism in the character design, scene design, and shot-level generation processes, the consistency of storyboard images in terms of character image, art style, and shot expression is significantly improved; at the same time, it reduces the creator's dependence on operating complex AI toolchains, allowing creators to focus more on content conception, thereby improving the overall efficiency and quality of AI comic storyboard production.

[0028] As an optional example, the target large language model is invoked to perform structured parsing and storyboard decomposition of the script text input by the user, obtaining scene design description information, character design information for each character, and storyboard description information for each shot, including: According to the preset scene breakdown rules, the script text is structured and parsed to break it down into at least one scene. For each scene, the corresponding scene semantic information is extracted and scene design description information for each scene is generated. Based on the character descriptions in the script, we extract the physical features, clothing features, and personality traits of each character to generate character design information. For each scene, it is broken down at the shot level to generate at least one corresponding shot description information, which includes the main subject of the shot, the characters appearing, the characters' behavior, the shot size information, and the shot's expressive intent.

[0029] Optionally, in this embodiment, the target large language model is invoked to perform structured parsing and scene breakdown of the script text input by the user, thereby automatically converting the script text into structured semantic data that can be used for subsequent scene generation. Specifically, the script text is first parsed according to preset scene breakdown rules, which are used to limit continuous plot content occurring at the same time and place, thus breaking down the script text into at least one scene. For each scene, the target large language model performs semantic understanding on the scene-related descriptions in the script, extracting scene semantic information such as spatial structure, environmental elements, and atmosphere features contained in the scene, and generating scene design description information corresponding to that scene, providing a clear semantic basis for subsequent scene image generation.

[0030] After scene-by-scene decomposition, the target large language model further analyzes the character information in the script text. By analyzing the characters' descriptions and behaviors in the script, it extracts the characters' physical features, clothing characteristics, and personality traits, generating standardized character design information. This character design information is used to uniformly describe the characters' visual image and temperament, thereby avoiding inconsistencies in character appearance during subsequent multi-shot generation.

[0031] Furthermore, for each scene that has been decomposed, the target large language model continues to perform shot-level decomposition processing, further breaking down a scene into at least one continuous shot unit. For each shot, corresponding shot description information is generated. This shot description information includes at least the shot subject, characters, character actions, shot size, and the shot's expressive intent, used to clarify the narrative focus and expressive method of the shot. Through the above structured parsing and shot decomposition process, the original script text is transformed into scene and shot description data with clear hierarchy and complete semantics, thus providing a reliable data foundation for subsequent generation of storyboard images based on a multimodal model.

[0032] As an optional example, the target image generation model is invoked to generate scene design images and character design images for each character, based on scene design description information and character image design information, combined with reference style images input by the user. Based on the scene design description information and reference art style images, generate scene design images corresponding to each scene, and associate and store them with the corresponding scene identifier. The scene design images are used to represent the spatial structure, environmental elements and overall visual style of the scene. Based on the character design information and reference art style images of each character, generate corresponding character design images for each character, and associate and store them with the corresponding character identifiers. The character design images are simplified or white background images of the characters.

[0033] Optionally, in this embodiment, by invoking the target image generation model, scene design images and character design images for subsequent storyboard construction are automatically generated based on the scene design description information and character image design information obtained in the preceding steps, combined with the reference art style image input by the user. This process aims to transform abstract semantic descriptions into basic visual materials with a unified style, thereby providing a stable and consistent visual benchmark for shot-level image generation.

[0034] Specifically, for each scene, the corresponding scene design description information and the user-inputted reference style image are used as input to call the target image generation model to perform image generation processing. The scene design description information is used to constrain the spatial structure, environmental layout, and key environmental elements of the scene, while the reference style image is used to limit the overall visual style, color tone, and expression of the image. The scene design image generated in this way can intuitively represent the overall environment and visual atmosphere of the corresponding scene and serve as the basic background for generating each shot in that scene. After generation, the scene design image is associated with the corresponding scene identifier and stored so that it can be called as needed in subsequent shot-level background generation or storyboard construction.

[0035] Secondly, for each character parsed from the script text, character design information and reference art style images are used as input to generate corresponding character design images using a target image generation model. The character design information defines the character's appearance, clothing style, and overall temperament, while the reference art style images ensure stylistic consistency between the character and scene designs. During generation, simplified or white backgrounds are preferred for the character design images to highlight the main character, reduce background interference, and facilitate reuse and combination in different scenes. The generated character design images are associated with corresponding character identifiers and stored for subsequent fusion generation of characters and scenes in storyboards. Through this method, standardized generation and management of scene and character visual assets are achieved, providing a reliable foundation for the efficient construction of comic book storyboards.

[0036] As an optional example, the target multimodal large model is invoked to analyze and predict the information of physical objects and camera positions in the scene based on the scene design images and the shot description information of each shot, including: Based on the scene design image, the visible physical objects contained in the scene are identified to obtain the corresponding physical object category information and its spatial distribution information in the scene; By jointly analyzing the shot description information of each shot and the recognition results of the scene design image, the set of physical objects to be presented in each shot and the relative positional relationship of the physical objects in the shot are predicted, and the prediction results of physical objects are obtained. Based on the shot composition information and shot expression intent in the shot description information, the camera position parameters corresponding to each shot are predicted. The camera position parameters are used to indicate the shooting direction, shooting angle and shooting distance of the virtual camera.

[0037] Optionally, in this embodiment, by invoking the target multimodal large model, the generated scene design images and the shot description information corresponding to each shot are jointly analyzed to achieve intelligent understanding and prediction of shot-level image components, providing precise control basis for subsequent shot-level scene background generation. This process fully integrates image information and text semantic information, avoiding comprehension bias caused by relying on only a single modality.

[0038] Specifically, the scene design image is first input into the target multimodal large model, and the visible objects in the scene image are identified and analyzed. These objects include, but are not limited to, buildings, props, natural elements, and other environmental components. Through visual understanding of the scene design image, the category information corresponding to each object and its spatial distribution information in the scene are obtained, thereby forming a basic understanding of the overall composition of the scene.

[0039] Building upon this, the shot description information of each scene is further input into the target multimodal large model for joint analysis along with the scene image recognition results. The shot description information includes semantic content such as the main subject of the shot, characters appearing, character actions, and focal points of the scene. Based on these semantic constraints, the multimodal large model predicts the set of physical objects that need to be presented in the corresponding shot from the identified physical objects, and determines the relative position of each physical object in the shot, such as foreground, midground, or background, thereby obtaining shot-level physical object prediction results.

[0040] Furthermore, the target multimodal large model predicts the camera positions corresponding to each shot based on the shot composition information and the expressive intent of the shot description. Camera position parameters indicate the shooting direction, angle, and distance of the virtual camera when generating the image, reflecting the narrative effect and emotional atmosphere expressed by different camera techniques. Through the aforementioned object prediction and camera position prediction processes, the subsequently generated shot-level scene background images are highly consistent with the script's semantics in terms of content composition and shot expression, thereby significantly improving the accuracy and expressiveness of the storyboard images.

[0041] As an optional example, the target text image large model is invoked to generate shot-level scene background images for each shot based on the prediction results and scene design images, including: Each unprocessed shot is identified as the current shot, and the target raw image large model is invoked to perform the following processing on the current shot: Based on the predicted results of physical objects in the current shot and the camera position parameters, construct shot-level scene generation conditions that match the shot description information of the current shot. Based on the current shot's shot-level scene generation conditions, the scene design image is reconstructed or cropped to generate a shot-level scene background image that conforms to the current shot's camera position parameters. Associate and store the current shot's scene background image with its shot identifier.

[0042] Optionally, in this embodiment, by calling the target text-based image big model and combining the previously obtained object prediction results, camera position parameters, and scene design images, a shot-level scene background image consistent with the semantics of each shot's storyboard is automatically generated, thereby achieving refined generation and management from overall scene design to shot-level visuals. Specifically, for each shot, the generation process of the text-based image big model is triggered. During the generation process, based on the object prediction results and camera position parameters corresponding to the current shot, shot-level scene generation conditions matching the shot's storyboard description information are constructed. The shot-level scene generation conditions not only include the categories, quantities, and relative spatial relationships of the objects to be presented in the scene, but also comprehensively consider the shot type information, composition method, and shot expression intention in the storyboard description, thus forming constraining generation conditions for the content and visual structure of the scene. Through these generation conditions, the framing range, viewing direction, and focus of the current shot in the overall scene can be clearly defined.

[0043] Subsequently, the target image model is invoked, and based on the constructed shot-level scene generation conditions, the original scene design image is reconstructed or cropped accordingly. On one hand, when camera position parameters require local focus on the scene, the image model can perform cropping and perspective transformation operations based on the original scene design image to generate a background image that conforms to the specified shooting distance and angle. On the other hand, when the shot expression requires supplementation or adjustment of image elements, the image model can reconstruct the scene design image while maintaining the consistency of the overall scene style, thereby enhancing the integrity of the image and the expressiveness of the shot. The resulting shot-level scene background image is highly consistent with the corresponding shot description in terms of spatial structure, visual content, and semantics.

[0044] Finally, the generated lens-level scene background image of the current shot is associated with its corresponding lens identifier and stored to form an indexable and reusable lens-level scene resource.

[0045] As an optional example, the target multimodal large model is invoked to perform semantic parsing on the shot-level scene background image, shot storyboard description information, and the corresponding character design images of each character in the target shot, generating the shot description information of the target shot, including: Based on the scene background image, shot description information, and character design images of each character in the target shot, the spatial structure relationship, the relative positional relationship between characters and scene, the standing position relationship, action state, and interaction relationship between characters in the target shot are analyzed. Based on the shot size information and shot expression intention in the shot description information of the target shot, the image description information used to describe the composition, perspective characteristics, and emotional expression of the target shot is generated.

[0046] Optionally, in this embodiment, by invoking the target multimodal large model, the shot-level scene background image of the target shot, the corresponding shot storyboard description information, and the character design images of each character involved in the storyboard description are jointly input and semantically understood, thereby realizing automated semantic parsing of the shot and generating structured target shot description information. The target multimodal large model has the ability to perform unified modeling and joint reasoning on multimodal information such as images and text, and can establish a correspondence at the level of visual features and semantic features, providing an accurate semantic foundation for subsequent image generation and verification.

[0047] Specifically, firstly, based on the shot-level scene background image of the target shot, the overall spatial structure of the scene is analyzed to identify the main areas, foreground and background layers, and key visual anchor points. Then, combined with the shot storyboard description information, the framing range and center of gravity of the shot are determined. On this basis, character design images corresponding to each character from the shot storyboard description information are introduced, and semantic modeling is performed on the characters' appearance features, body proportions, and identity attributes. This enables the model to distinguish different characters and understand their functional roles in the shot.

[0048] Subsequently, the target multimodal big data model analyzes the relative positional relationships between characters and the scene, clarifying the characters' standing positions, hierarchical relationships, and spatial correspondences with scene elements. Simultaneously, it understands and infers the characters' action states, orientation information, and interactions between them, such as dialogue, confrontation, cooperation, or emotional resonance, thereby reconstructing the narrative relationships and plot points expressed in the shots. By analyzing the relationships between characters, the consistency of narrative logic and emotional expression within the shots can be effectively supported.

[0049] Furthermore, by combining the shot size information and the expressive intent of the shot in the shot description information, the target multimodal big data model comprehensively judges the composition and perspective characteristics of the shot, generating compositional semantics to describe the target shot, including the proportion of people in the frame, the layout of the main subject and supporting elements, and the height and direction of the viewpoint. At the same time, based on the expressive intent of the shot, the emotional atmosphere to be conveyed by the image is semantically summarized, such as tension, warmth, oppression, or openness, and incorporated into the image description information.

[0050] The final generated image description information is output in a structured or semi-structured form, which can comprehensively characterize the semantics of the target shot in terms of spatial structure, character relationships, compositional features, and emotional expression. Through the above method, the automatic transformation of the shot image from "visible image" to "understandable semantics" is realized, providing key semantic support and technical guarantee for subsequent character compositing, image consistency verification, automatic shot generation, and intelligent production of film and television content.

[0051] As an optional example, after generating the storyboard for the target shot, the above method further includes: Perform character target detection on the storyboard of the target shot, and compare the number of detected characters with the number of characters in the storyboard description information of the target shot for consistency. If the number of characters is inconsistent, the storyboard of the target shot will be regenerated. When the number of characters is consistent, the target multimodal large model is invoked to perform consistency analysis on the target shot's storyboard, image description information, and shot storyboard description information. Based on the analysis results, the image description information of the target storyboard is corrected and the storyboard is regenerated, or the current storyboard is confirmed as the final storyboard.

[0052] Optionally, in this embodiment, after generating the storyboard for the target shot, the above method further introduces a consistency verification and self-feedback optimization mechanism based on multimodal understanding to improve the accuracy, narrative consistency, and stability of the generated results. This step forms a closed-loop optimization process by automatically detecting, semantically analyzing, and correcting the generated images, thereby reducing manual intervention costs and improving the quality of storyboard generation.

[0053] Specifically, the process begins by performing human target detection on the storyboard generated from the target shot. This process can be based on a pre-trained human detection model or a multimodal visual perception model to identify and count the human targets appearing in the storyboard, thus obtaining information on the actual number of people presented in the shot. Subsequently, the detected number of people is compared with the number of people appearing as specified in the storyboard description information of the target shot to determine whether the generated shot meets the storyboard design requirements in terms of the completeness of the characters.

[0054] If an inconsistency in the number of characters is detected, such as missing characters, extra characters, or duplicate characters, it is determined that the current storyboard does not meet the basic constraints of the shot description information. At this time, the process of regenerating the storyboard of the target shot is triggered. By reconstructing the generation conditions or adjusting the generation parameters, the target shot is regenerated to correct the deviation in the number of characters and ensure that the storyboard is consistent with the storyboard design at the basic composition level.

[0055] When the number of characters is consistent, the target multimodal large model is further invoked to perform joint input and consistency analysis on the storyboard of the target shot, the previously generated image description information, and the corresponding shot storyboard description information. This consistency analysis not only focuses on the number of characters, but also evaluates the degree of matching between the image content and the semantics of the storyboard at a higher semantic level, including whether the characters' positional relationships, action states, interaction relationships, composition methods, and emotional expressions conform to the expressive intent defined in the shot storyboard description information.

[0056] Based on the consistency analysis results, when a deviation between the scene and the storyboard semantics is detected but can be corrected by adjusting the description information, the scene description information of the target storyboard is adaptively corrected, and the storyboard scene is regenerated accordingly to further align the scene semantics with the storyboard design; when the consistency analysis results show that the current storyboard scene meets the constraints of the shot storyboard description information and the scene description information, the storyboard scene is confirmed as the final storyboard scene output.

[0057] The above technical solution enables automatic consistency verification and iterative optimization during the storyboard generation process, effectively ensuring a high degree of consistency between character presentation, semantic expression and storyboard design, and improving the reliability and controllability of the storyboard generation results. It is applicable to application scenarios such as automatic generation of film and television storyboards, animation pre-visualization and intelligent production of digital content.

[0058] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0059] According to another aspect of the embodiments of this application, an apparatus for generating comic book storyboard frames is also provided, such as... Figure 7 As shown, it includes: The parsing module 702 is used to call the target large language model to perform structured parsing and storyboard decomposition on the script text input by the user, to obtain scene design description information, character image design information of each character, and shot storyboard description information of each shot. It also calls the target image generation model to generate scene design images and character design images of each character based on the scene design description information and character image design information, combined with the reference art style image input by the user. Analysis module 704 is used to call the target multimodal large model, analyze and predict the information of physical objects and camera positions in the scene based on the scene design image and the shot description information of each shot, and call the target text image large model to generate the shot-level scene background image of each shot based on the prediction results and the scene design image. The first generation module 706 is used to call the target multimodal large model to perform image semantic analysis on the shot-level scene background image, shot storyboard description information and the corresponding character design images of each character in the target shot, and generate the image description information of the target shot, wherein the target shot is any one of the shots. The second generation module 708 is used to call the target image editing model to generate the storyboard of the target shot based on the image description information of the target shot, the shot-level scene background image and the character design image.

[0060] It should be noted that the parsing module 702 in this embodiment can be used to execute step S102 in this application embodiment, the analysis module 704 in this embodiment can be used to execute step S104 in this application embodiment, the first generation module 706 in this embodiment can be used to execute step S106 in this application embodiment, and the second generation module 708 in this embodiment can be used to execute step S108 in this application embodiment.

[0061] As an optional example, the parsing module includes: The first extraction unit is used to perform structured parsing of the script text according to the preset scene decomposition rules, so as to decompose the script text into at least one scene, and extract the corresponding scene semantic information for each scene, and generate scene design description information corresponding to each scene. The second extraction unit is used to extract the appearance features, clothing features and personality features of each character based on the description information of the characters in the script text, and generate the character image design information of each character. The decomposition unit is used to decompose each scene at the shot level and generate at least one corresponding shot description information, wherein the shot description information includes the main subject of the shot, the characters appearing, the characters' behavior, the shot size information, and the shot's expressive intent.

[0062] As an optional example, the parsing module includes: The first generation unit is used to generate scene design images corresponding to each scene based on scene design description information and reference style images, and to associate and store them with the corresponding scene identifier. The scene design images are used to represent the spatial structure, environmental elements and overall visual style of the scene. The second generation unit is used to generate character design images corresponding to each character based on the character image design information and reference art style images, and to associate and store them with the corresponding character identifiers. The character design images are simplified or white background character images.

[0063] As an optional example, the analysis module includes: The recognition unit is used to identify visible physical objects in the scene based on the scene design image, and obtain the corresponding physical object category information and its spatial distribution information in the scene; The first prediction unit is used to jointly analyze the shot description information of each shot and the recognition results of the scene design image to predict the set of physical objects that need to be presented in each shot and the relative positional relationship of the physical objects in the shot image, and obtain the prediction result of physical objects. The second prediction unit is used to predict the camera position parameters corresponding to each shot based on the shot size information and shot expression intention in the shot description information. The camera position parameters are used to indicate the shooting direction, shooting angle and shooting distance of the virtual camera.

[0064] As an optional example, the analysis module includes: The processing unit is used to identify each unprocessed shot as the current shot and call the target text image large model to perform the following processing on the current shot: Based on the predicted results of physical objects in the current shot and the camera position parameters, construct shot-level scene generation conditions that match the shot description information of the current shot. Based on the current shot's shot-level scene generation conditions, the scene design image is reconstructed or cropped to generate a shot-level scene background image that conforms to the current shot's camera position parameters. Associate and store the current shot's scene background image with its shot identifier.

[0065] As an optional example, the first generation module includes: The third generation unit is used to analyze the spatial structure relationship of the target shot, the relative positional relationship between characters and the scene, the standing relationship, action state, and interaction relationship between characters, based on the shot-level scene background image, shot storyboard description information, and the corresponding character design images of each character in the target shot. Based on the shot size information and shot expression intention in the shot storyboard description information of the target shot, it generates picture description information to describe the composition method, perspective characteristics, and emotional expression of the target shot.

[0066] As an optional example, the above-described apparatus further includes: The detection module is used to detect human targets in the storyboard of the target shot after the storyboard of the target shot is generated, and to compare the number of detected human targets with the number of human targets in the storyboard description information of the target shot. The third generation module is used to trigger the regeneration of the storyboard of the target shot when the number of characters is inconsistent. The fourth generation module is used to call the target multimodal large model when the number of characters is consistent, to perform consistency analysis on the storyboard, scene description information and shot storyboard description information of the target shot, and to correct the scene description information of the target storyboard and regenerate the storyboard based on the analysis results, or to confirm the current storyboard as the final storyboard.

[0067] For other examples of this embodiment, please refer to the examples above, which will not be repeated here.

[0068] Figure 8 This is a schematic diagram of an optional electronic device according to an embodiment of this application, such as... Figure 8 As shown, it includes a processor 802, a communication interface 804, a memory 806, and a communication bus 808. The processor 802, communication interface 804, and memory 806 communicate with each other via the communication bus 808. Memory 806 is used to store computer programs; When processor 802 executes a computer program stored in memory 806, it performs the following steps: The target large language model is invoked to perform structured parsing and storyboard decomposition on the script text input by the user, so as to obtain scene design description information, character image design information of each character, and shot storyboard description information of each shot. The target image generation model is invoked to generate scene design images and character design images of each character based on the scene design description information and character image design information, combined with the reference art style image input by the user. The target multimodal large model is invoked to analyze and predict the information of physical objects and camera positions in the scene based on the scene design image and the shot description information of each shot. The target text image large model is also invoked to generate the shot-level scene background image of each shot based on the prediction results and the scene design image. The target multimodal large model is invoked to perform image semantic analysis on the shot-level scene background image, shot storyboard description information and the corresponding character design images of each character in the target shot, and generate the image description information of the target shot, where the target shot is any one of the shots; The target image editing model is invoked to generate the storyboard of the target shot based on the image description information of the target shot, the shot-level scene background image, and the character design image.

[0069] Optionally, in this embodiment, the communication bus can be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 8 The symbol is represented by a single thick line, but this does not indicate that there is only one bus or one type of bus. The communication interface is used for communication between the aforementioned electronic devices and other devices.

[0070] The memory may include RAM, or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0071] As an example, the memory 806 described above may include, but is not limited to, the parsing module 702, analysis module 704, first generation module 706, and second generation module 708 in the aforementioned anime storyboard generation device. Furthermore, it may include, but is not limited to, other module units in the aforementioned anime storyboard generation device, which will not be elaborated upon in this example.

[0072] The processor mentioned above can be a general-purpose processor, including but not limited to: CPU (Central Processing Unit), NP (Network Processor), etc.; it can also be DSP (Digital Signal Processor), ASIC (Application Specific Integrated Circuit), FPGA (Field-Programmable Gate Array) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0073] Optionally, specific examples in this embodiment can refer to the examples described in the above embodiments, and will not be repeated here.

[0074] Those skilled in the art will understand that Figure 8The structure shown is for illustrative purposes only. The device that implements the above method for generating comic book storyboards can be a terminal device, such as a smartphone (e.g., Android phone, iOS phone), tablet computer, PDA, mobile Internet Devices (MID), PAD, etc. Figure 8 This does not limit the structure of the aforementioned electronic devices. For example, the electronic device may also include components that are more... Figure 8 The more or fewer components shown (such as network interfaces, display devices, etc.), or having the same Figure 8 The different configurations shown.

[0075] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, ROM, RAM, disk or optical disk, etc.

[0076] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, wherein a computer program is stored in the computer program, which, when executed by a processor, performs the steps in the above-described method for generating comic book storyboard frames.

[0077] Optionally, in this embodiment, those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0078] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0079] If the integrated units in the above embodiments are implemented as software functional units and sold or used as independent products, they can be stored in the aforementioned computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause one or more computer devices (which may be personal computers, servers, or network devices, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application.

[0080] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0081] In the several embodiments provided in this application, it should be understood that the disclosed client can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between units or modules, and may be electrical or other forms.

[0082] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0083] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0084] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A method for generating storyboard frames for a comic strip, characterized in that, include: The target large language model is invoked to perform structured parsing and storyboard decomposition on the script text input by the user, so as to obtain scene design description information, character image design information of each character, and shot storyboard description information of each shot. The target image generation model is invoked to generate scene design images and character design images of each character based on the scene design description information and the character image design information, combined with the reference art style image input by the user. The target multimodal large model is invoked, and based on the scene design image and the shot description information of each shot, the information of physical objects in the scene and the camera position are analyzed and predicted. The target text image large model is invoked, and based on the prediction results and the scene design image, the shot-level scene background image of each shot is generated. The target multimodal large model is invoked to perform image semantic analysis on the shot-level scene background image, shot storyboard description information and the corresponding character design images of each character in the target shot, and generate the image description information of the target shot, wherein the target shot is any one of the shots; The target image editing model is invoked to generate the storyboard of the target shot based on the image description information of the target shot, the shot-level scene background image, and the character design image.

2. The method according to claim 1, characterized in that, The target large language model is invoked to perform structured parsing and storyboard decomposition of the script text input by the user, obtaining scene design description information, character image design information for each character, and storyboard description information for each shot, including: According to the preset scene breakdown rules, the script text is structured and parsed to break it down into at least one scene. For each scene, the corresponding scene semantic information is extracted to generate scene design description information for each scene. Based on the character descriptions in the script text, the appearance, clothing and personality traits of each character are extracted to generate character design information for each character. For each scene, it is broken down at the shot level to generate at least one corresponding shot description information, wherein the shot description information includes the main subject of the shot, the characters appearing, the characters' behavior, the shot size information, and the shot's expressive intent.

3. The method according to claim 2, characterized in that, The target image generation model is invoked, and based on the scene design description information and the character image design information, combined with the reference style image input by the user, scene design images and character design images of each character are generated, including: Based on the scene design description information and the reference art style image, a scene design image corresponding to each scene is generated and associated with the corresponding scene identifier for storage. The scene design image is used to represent the spatial structure, environmental elements and overall visual style of the scene. Based on the character design information of each character and the reference art style image, generate the corresponding character design image for each character, and associate and store it with the corresponding character identifier. The character design image is a character image with a simplified background or a white background.

4. The method according to claim 1, characterized in that, The target multimodal large model is invoked, and based on the scene design image and the shot description information of each shot, the information of physical objects and camera positions in the scene are analyzed and predicted, including: Based on the scene design image, the visible physical objects contained in the scene are identified to obtain the corresponding physical object category information and their spatial distribution information in the scene; The shot description information of each shot is jointly analyzed with the recognition result of the scene design image to predict the set of physical objects to be presented in each shot and the relative positional relationship of the physical objects in the shot image, so as to obtain the prediction result of physical objects. Based on the shot composition information and shot expression intent in the shot description information, the camera position parameters corresponding to each shot are predicted, wherein the camera position parameters are used to indicate the shooting direction, shooting angle and shooting distance of the virtual camera.

5. The method according to claim 4, characterized in that, The target text image model is invoked, and based on the prediction results and the scene design image, lens-level scene background images for each shot are generated, including: The unprocessed footage is identified as the current footage, and the target raw image large model is invoked to perform the following processing on the current footage: Based on the predicted object results and camera position parameters of the current shot, construct shot-level scene generation conditions that match the shot description information of the current shot. Based on the current shot's shot-level scene generation conditions, the scene design image is reconstructed or cropped to generate a shot-level scene background image that conforms to the current shot's camera position parameters; The current shot's scene background image is associated with and stored with its shot identifier.

6. The method according to claim 1, characterized in that, The target multimodal large model is invoked to perform image semantic analysis on the shot-level scene background image, shot storyboard description information, and corresponding character design images of the target shot, generating the image description information of the target shot, including: Based on the scene background image, shot description information, and character design images of each character in the target shot, the spatial structure relationship, the relative positional relationship between characters and scene, the standing position relationship, action state, and interaction relationship between characters in the target shot are analyzed. Based on the shot size information and shot expression intention in the shot description information of the target shot, a picture description information is generated to describe the composition, perspective characteristics, and emotional expression of the target shot.

7. The method according to any one of claims 1 to 6, characterized in that, After generating the storyboard of the target shot, the method further includes: Human targets are detected in the storyboard of the target shot, and the number of detected human targets is compared with the number of human targets in the storyboard description information of the target shot. If the number of characters is inconsistent, the storyboard of the target shot will be regenerated. When the number of characters is consistent, the target multimodal large model is invoked to perform consistency analysis on the storyboard, scene description information and shot storyboard description information of the target shot. Based on the analysis results, the scene description information of the target storyboard is corrected and the storyboard is regenerated, or the current storyboard is confirmed as the final storyboard.

8. A device for generating storyboard frames for a comic strip, characterized in that, include: The parsing module is used to call the target large language model to perform structured parsing and storyboard decomposition on the script text input by the user, to obtain scene design description information, character image design information of each character, and shot storyboard description information of each shot. It also calls the target image generation model to generate scene design images and character design images of each character based on the scene design description information and the character image design information, combined with the reference art style image input by the user. The analysis module is used to call the target multimodal large model, analyze and predict the information of physical objects and camera positions in the scene based on the scene design image and the shot description information of each shot, and call the target text image large model to generate the shot-level scene background image of each shot based on the prediction results and the scene design image. The first generation module is used to call the target multimodal large model to perform image semantic analysis on the shot-level scene background image, shot storyboard description information and the corresponding character design images of each character in the target shot, and generate the image description information of the target shot, wherein the target shot is any one of the shots; The second generation module is used to call the target image editing model to generate the storyboard of the target shot based on the image description information of the target shot, the shot-level scene background image and the character design image.

9. A computer-readable storage medium storing a computer program, characterized in that, The computer program is executed by the processor to perform the method described in any one of claims 1 to 7.

10. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the method described in any one of claims 1 to 7 through the computer program.