Story board generation method and device, electronic equipment and computer storage medium

By automatically generating story scripts and storyboard image sequences, the problem of low efficiency in user storyboard production is solved, and efficient and personalized storyboard generation is achieved.

CN120876642APending Publication Date: 2025-10-31ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510957912.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-10
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

In existing technologies, the user storyboard production process relies on manual user profile abstraction and plot construction, which is time-consuming and labor-intensive, has efficiency bottlenecks, and the multi-role collaboration and visual generation stages are inefficient.

Method used

By acquiring user-uploaded story requirements, the system automatically generates story scripts and storyboard image sequences using preset models, including semantic parsing, narrative structure generation, and shot grammar structuring, generating storyboard information and reducing issues related to switching between multiple tools and style inconsistencies.

Benefits of technology

It improves the personalization and semantic consistency of story script information, enhances the generation efficiency and consistency of storyboard information, and realizes the integrated generation from story requirements to storyboard information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120876642A_ABST
    Figure CN120876642A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a story board generation method and device, electronic equipment and a computer storage medium. The method comprises the following steps: acquiring story demand information uploaded by a user, and performing story script generation processing on the story demand information by adopting a preset model to obtain story script information corresponding to the story demand information; and further performing image generation processing on the story script information by using the preset model to obtain a split image sequence corresponding to the story script information, further displaying the split image sequence, and when a determination operation for the split image sequence is detected, generating story board information according to the split image sequence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of information technology, and in particular to a storyboard generation method, apparatus, electronic device, and computer storage medium. Background Technology

[0002] User storyboards, as an important tool connecting user research and product design implementation, are widely used to depict the behavioral paths and interaction processes of typical users in specific scenarios.

[0003] In related technologies, the user storyboard production process usually relies on manual user profile abstraction and plot construction. This process is not only time-consuming and labor-intensive, but also has obvious efficiency bottlenecks in multi-role collaboration and visual generation, resulting in low storyboard generation efficiency. Summary of the Invention

[0004] This specification provides a storyboard generation method, apparatus, electronic device, and computer storage medium, which can achieve precise face-to-face matching of users, encourage users to bring people around them to participate in activities, issue benefits after successful matching, and promote offline payment methods.

[0005] Firstly, embodiments of this specification provide a storyboard generation method, including: Obtain user-uploaded story request information; Based on the above story requirements information, a preset model is used to generate the story script, resulting in the corresponding story script information. Based on the above story script information, the above preset model is used to perform image generation processing to obtain the story script information corresponding to the story image sequence. The above-mentioned storyboard image sequence is displayed, and in response to the determination operation of the above-mentioned storyboard image sequence, storyboard information is generated based on the above-mentioned storyboard image sequence.

[0006] In one possible implementation, the story script generation process based on the aforementioned story requirement information uses a preset model to obtain the corresponding story script information, including: Based on the above story requirement information, a preset model is used to perform semantic parsing to obtain script element information, which includes at least one of the following: behavioral appeal information, story scene information, and behavioral path information. Determine user type information based on the above script element information; Based on the above user type information and the above script element information, the above preset model is used to generate the narrative structure, and the initial narrative structure corresponding to the above user type information is obtained. The initial narrative structure described above is processed using shot grammar structuring to obtain the story script information corresponding to the user type information.

[0007] In one possible implementation, the narrative structure generation process based on the aforementioned user type information and script element information, using the aforementioned preset model, yields an initial narrative structure corresponding to the aforementioned user type information, including: Based on the above user type information and the above script element information, the above preset model is used to perform preference optimization processing to obtain the target plot information and expression style information corresponding to the above user type information; Based on the aforementioned target plot information and expressive style information, the aforementioned preset model is used to generate the narrative structure, resulting in an initial narrative structure.

[0008] In one possible implementation, the initial narrative structure is subjected to shot grammar structuring to obtain the story script information corresponding to the user type information, including: The initial narrative structure is broken down based on preset shot grammar rules to obtain multiple initial script storyboard units; each of the multiple initial script units includes at least one of the following information: character information, behavior description information, and scene information. Based on the above-mentioned multiple initial script storyboard units, the above-mentioned preset model is used to perform storyboard generation processing to obtain multiple script storyboard information. The above-mentioned script storyboard information is used to describe the visual expression content of the storyboard image frame. Based on the above multiple script storyboard information, determine the story script information corresponding to the above user type information.

[0009] In one possible implementation, the image generation process based on the aforementioned story script information and using the aforementioned preset model to obtain the story script information corresponding to the story script information includes: Generate prompts based on the story script information above; Based on the storyboard information of each script scene in the above storyboard information, the above preset model is used to perform feature parsing processing to obtain the image content information corresponding to the script scene information. Based on the above-mentioned script storyboard information and the above-mentioned prompt information, the above-mentioned preset model is used to perform image frame generation processing to obtain the storyboard image frames corresponding to the above-mentioned script storyboard information. Arrange the multiple storyboard image frames corresponding to the multiple storyboard images in the above storyboard information to generate the storyboard image sequence corresponding to the above storyboard information.

[0010] In one possible implementation, the above prompt information includes user description information and / or scene description information to guide image generation; The above-mentioned prompt information generated based on the story script information includes: Extract character information from the aforementioned story script information, and generate corresponding user description information based on the aforementioned character information using a preset user description template; and / or, Extract the background information from the story script information above, and generate the corresponding scene description information based on the background information using a preset scene description template.

[0011] In one possible implementation, the above method also includes: Display style switching controls for the above storyboard image sequence; Upon detecting the triggering operation of the aforementioned style switching control, the aforementioned preset model is used to perform style adjustment processing based on the aforementioned storyboard image sequence and the target image style corresponding to the aforementioned style switching control, thereby generating an adjusted storyboard image sequence. Based on the adjusted storyboard image sequence described above, the steps of displaying the storyboard image sequence and generating storyboard information based on the storyboard image sequence are performed in response to the determination operation for the storyboard image sequence.

[0012] In one possible implementation, after displaying the above-described storyboard image sequence, the method further includes: In response to a preset operation targeting a local image region in any storyboard image frame, image processing instructions are obtained; Based on the above image processing instructions and the above local image regions, the above preset model is used to modify the image regions to obtain the target storyboard image frame. In response to the determination instruction for the target storyboard image frame, the storyboard image sequence is updated based on the target storyboard image frame to obtain an updated storyboard image sequence. Based on the updated storyboard image sequence, the steps of displaying the storyboard image sequence and generating storyboard information based on the storyboard image sequence are performed in response to the determination operation for the storyboard image sequence.

[0013] In one possible implementation, the above method also includes: In response to the determination command for the aforementioned target storyboard image frame, modification record information is generated; Based on the above modification record information, reinforcement learning feedback data is generated. Based on the reinforcement learning feedback data, the model parameters of the preset model are optimized and trained, and the preset model is updated based on the optimized model parameters.

[0014] Secondly, embodiments of this specification provide a storyboard generation apparatus, comprising: The first acquisition module is used to acquire story request information uploaded by users; The first processing module is used to generate a story script based on the above story requirements information using a preset model, and obtain the corresponding story script information. The second processing module is used to perform image generation processing based on the above story script information using the above preset model to obtain the story script information corresponding to the story image sequence. The first generation module is used to display the above-mentioned storyboard image sequence and, in response to the determination operation of the above-mentioned storyboard image sequence, generate storyboard information based on the above-mentioned storyboard image sequence.

[0015] Thirdly, embodiments of this specification provide an electronic device, including: a processor and a memory; The processor is connected to the memory. The aforementioned memory is used to store executable program code; The processor reads the executable program code stored in the memory to run the program corresponding to the executable program code, so as to execute the method provided by the first aspect of the embodiments of this specification or any possible implementation of the first aspect.

[0016] Fourthly, embodiments of this specification provide a computer storage medium storing a plurality of instructions adapted for loading by a processor and executing the method provided by the first aspect of the embodiments of this specification or any possible implementation thereof.

[0017] Fifthly, embodiments of this specification provide a computer program product containing instructions that, when run on a computer or processor, cause the computer or processor to perform the method provided by the first aspect of the embodiments of this specification or any possible implementation thereof.

[0018] This specification's embodiments obtain user-uploaded story requirement information and automatically generate a story script based on this information using a preset model. This script better aligns with the user's intent, improving the personalization and semantic consistency of the story script information. Furthermore, the corresponding story script and storyboard image sequence are automatically generated based on the preset model, ultimately producing storyboard information. The storyboard image sequence generated from the script ensures a close integration between the story script information and the storyboard images. This integrated generation method from story requirement information to storyboard information reduces issues related to switching between multiple tools and style inconsistencies, effectively improving the generation efficiency of storyboard information and the consistency between storyboard information and story requirement information. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of this specification, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 A schematic diagram of the architecture of a storyboard generation system provided as an exemplary embodiment of this specification; Figure 2 A flowchart illustrating a storyboard generation method provided as an exemplary embodiment of this specification; Figure 3 A flowchart illustrating a method for determining story script information provided as an exemplary embodiment of this specification; Figure 4 A flowchart illustrating a method for determining a storyboard image sequence as provided in an exemplary embodiment of this specification; Figure 5 A flowchart illustrating a storyboard generation method provided as an exemplary embodiment of this specification; Figure 6 A flowchart illustrating another storyboard generation method provided as an exemplary embodiment of this specification; Figure 7 A detailed flowchart illustrating a storyboard generation method provided as an exemplary embodiment of this specification; Figure 8 A schematic diagram of a story requirement information acquisition interface provided for an exemplary embodiment of this specification; Figure 9 A schematic diagram of a first storyboard image sequence display page provided for an exemplary embodiment of this specification; Figure 10 A schematic diagram of a second storyboard image sequence display page provided for an exemplary embodiment of this specification; Figure 11 A schematic diagram of the structure of a storyboard generation apparatus provided in an exemplary embodiment of this specification; Figure 12 This is a schematic diagram of the structure of an electronic device provided as an exemplary embodiment of this specification. Detailed Implementation

[0021] The technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings.

[0022] The terms "first," "second," "third," etc., used in this specification, claims, and the foregoing drawings are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to such processes, methods, products, or apparatus.

[0023] Please refer to Figure 1 , Figure 1 This is a schematic diagram of the architecture of a storyboard generation system provided for an exemplary embodiment of this specification. Figure 1 As shown, the storyboard generation system 10 may include a terminal 101, a network 102, and a server 103. The network 102 serves as the medium for providing a communication link between the terminal 101 and the server 103. The network 102 may include various types of wired or wireless communication links, such as wired communication links including fiber optic cables, twisted-pair cables, or coaxial cables, and wireless communication links including Bluetooth communication links, Wireless-Fidelity (Wi-Fi) communication links, or microwave communication links, etc.

[0024] Terminal 101 can interact with server 103 via network 102 to receive messages from or send messages to server 103. Alternatively, terminal 101 can interact with server 103 via network 102 to receive messages or data sent to server 103 by other users. Terminal 101 can be hardware or software. When terminal 101 is hardware, it can be various electronic devices, including but not limited to tablet computers, laptops, and desktop computers. When terminal 101 is software, it can be installed in the aforementioned electronic devices and can be implemented as multiple software programs or software modules (e.g., to provide distributed services) or as a single software program or software module; no specific limitation is made here.

[0025] In this embodiment of the specification, terminal 101 first obtains the story requirement information uploaded by the user, and then sends the story requirement information to server 103 through network 102. Server 103 then performs story script generation processing based on the story requirement information using a preset model to obtain the corresponding story script information, and performs image generation processing based on the story script information using the preset model to obtain the storyboard image sequence corresponding to the story script information. Then, server 103 sends the storyboard image sequence to terminal 102 through network 102. Terminal 102 displays the storyboard image sequence and, in response to the determination operation of the storyboard image sequence, generates storyboard information based on the storyboard image sequence.

[0026] Server 103 can be a server that provides various services. It should be noted that server 103 can be hardware or software. When server 103 is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. When server 103 is software, it can be implemented as multiple software programs or software modules (e.g., used to provide distributed services), or as a single software program or software module; no specific limitations are made here.

[0027] Alternatively, the system architecture may not include server 103. In other words, server 103 may be an optional device in the embodiments of this specification. That is, the method provided in the embodiments of this specification can be applied to a system structure that only includes terminal 101. The embodiments of this specification do not limit this.

[0028] It should be understood that Figure 1 The number of terminals, networks, and servers shown is only illustrative; the number can be any number of terminals, networks, and servers depending on the implementation requirements.

[0029] Please see Figure 2 , Figure 2 This is a flowchart illustrating a storyboard generation method provided in an embodiment of this specification. The execution entity in this embodiment can be a terminal executing the storyboard generation method, a processor within the terminal executing the storyboard generation method, or a storyboard generation service within the terminal executing the storyboard generation method. For ease of description, the following example uses a processor within a terminal as the execution entity to illustrate the specific execution process of the storyboard generation method.

[0030] Please see Figure 2 , Figure 2 This is a flowchart illustrating a storyboard generation method provided in an embodiment of this specification. Figure 2 As shown, storyboard generation methods can include at least: S202: Obtain story request information uploaded by users.

[0031] Users can be natural persons or institutional users. The identity of a user may include, but is not limited to, designers, screenwriters, animators, marketing planners, content platform operators, ordinary consumer users, or third-party platform systems that call this system through an Application Programming Interface (API).

[0032] Optionally, story requirement information may include multimodal information (such as image information, text information, audio information, video information, etc.) to describe the user's creative intent and content settings from multiple dimensions, thereby improving the accuracy and diversity of generated scripts and images. Story requirement information can be uploaded by the user in at least one of the following ways: file upload, text input, structured form input, image pasting or dragging and dropping, link reference, and voice input, wherein the file can be in Portable Document Format (PDF).

[0033] In some embodiments, the content of the story requirement information may specifically include at least one of the following: character profile information, natural language description information, preset industry template information, and reference material information. Character profile information may include, but is not limited to, the identity category, interests, content preferences, and professional background of the story characters; natural language description information may be free text information determined by the user, used to describe the story background, characters, plot, style, theme, etc.; preset industry template information may be structured template information provided by the platform, which are standardized scene elements designed for specific industries (such as film, animation, games, advertising, etc.); reference material information may include existing materials such as images, audio, video, or text uploaded or referenced by the user, used to assist in constructing the story background or character settings.

[0034] It is worth noting that the specific content of the story requirements information can be obtained through online uploading or offline collection. Offline collection methods include, but are not limited to, collecting user requirements through offline interviews, face-to-face communication or telephone surveys, and then forming structured or unstructured input content.

[0035] S204: Based on the above story requirements information, a preset model is used to generate the story script, and the corresponding story script information is obtained.

[0036] The preset model can be a pre-trained or pre-configured artificial intelligence model (such as a large language model), which has the ability to reason and generate text based on input requirements.

[0037] Optionally, the aforementioned story script information includes multiple storyboard information. The storyboard information can be segmented plot and visual guidance information units extracted and structured from the complete story script information. Each storyboard information can correspond to a plot node in the story, and may include, but is not limited to: shot sequence information, scene description information, character behavior information, dialogue text information, and image generation prompts, to guide the subsequent generation and processing of storyboard image sequences.

[0038] In one embodiment, the aforementioned preset model may integrate a large language model, such as a natural language generation model based on a Transformer architecture, for understanding and generating user-input story requirement information. Specifically, the story requirement information can be input into the large language model within the preset model. The large language model organizes and formats the story requirement information based on preset prompt word templates to construct highly guiding input content, and then outputs story script information corresponding to the story requirement information. The aforementioned large language model may include, but is not limited to, Generative Pre-trained Transformer (GPT) series models, Natural Language Processing (NLP) models, or other language models capable of generating long texts.

[0039] S206: Based on the above story script information, the above preset model is used to perform image generation processing to obtain the storyboard image sequence corresponding to the above story script information.

[0040] The storyboard image sequence can be a set of image frames arranged chronologically to express the story's progression. Each frame corresponds to a plot segment or shot. Specifically, the storyboard image sequence includes multiple storyboard image frames corresponding to the aforementioned multiple script storyboard information. These multiple storyboard image frames are arranged in the logical order of the story script information, sequentially presenting the visual representation of each plot segment, including but not limited to: character actions, background scenes, shot composition, emotional atmosphere, etc., thus forming a complete visual flow of the story.

[0041] Specifically, the preset model can integrate a diffusion model, which can be used to automatically generate high-quality storyboard image frames based on text prompts or structured script storyboard information, so as to accurately reproduce visual features such as character movements, scene details, composition layout, and emotional atmosphere. The aforementioned diffusion model can specifically be an image generation diffusion model based on text prompts (such as StableDiffusion, Flux model, etc.).

[0042] S208: Display the above storyboard image sequence, and in response to the determination operation for the above storyboard image sequence, generate storyboard information based on the above storyboard image sequence.

[0043] Optionally, multiple storyboard image frames can be presented on the user interface for user preview or interactive confirmation. The user interface may include a confirmation control. The confirmation operation for the aforementioned storyboard image sequence can be triggered by the confirmation control, and the system can then respond to the confirmation operation by taking the finally determined storyboard image sequence as input to generate corresponding storyboard information.

[0044] Specifically, generating storyboard information based on the aforementioned storyboard image sequence may include: arranging multiple generated storyboard image frames in a structured manner according to a logical order, and binding and displaying them with the script storyboard information corresponding to each image frame, thereby forming standardized storyboard information for previewing, editing, or exporting. In one embodiment, the storyboard information may include, but is not limited to: image frame order and number, script description or text description corresponding to each storyboard image frame, timeline, emotion markers, shot descriptions, and exportable editing formats (such as PDF, proprietary templates, etc.).

[0045] This specification's embodiments obtain user-uploaded story requirement information and automatically generate a story script based on this information using a preset model. This script better aligns with the user's intent, improving the personalization and semantic consistency of the story script information. Furthermore, the corresponding story script and storyboard image sequence are automatically generated based on the preset model, ultimately producing storyboard information. The storyboard image sequence generated from the script ensures a close integration between the story script information and the storyboard images. This integrated generation method from story requirement information to storyboard information reduces issues related to switching between multiple tools and style inconsistencies, effectively improving the generation efficiency of storyboard information and the consistency between storyboard information and story requirement information.

[0046] In one embodiment, the process of determining the story script information in step S204 above can be found in [reference needed]. Figure 3 This is a flowchart illustrating a method for determining story script information provided in an embodiment of this specification. Figure 3 As shown, the method for determining the story script information includes the following steps: S302: Based on the above story requirement information, a preset model is used to perform semantic parsing to obtain script element information. The above script element information includes at least one of the following: behavioral appeal information, story scene information, and behavioral path information.

[0047] Understandably, semantic parsing can be seen as the process of using models to structurally extract implicit meanings from natural language.

[0048] Optionally, behavioral appeal information can represent the key actions or conflicts that users want to be presented in the story, such as "the character must escape" or "there must be a turning point"; story scene information can include the time, place, and environmental background of the story (such as an ancient palace or a future city); behavioral path information can represent the interaction logic between characters and the plot development route.

[0049] S304: Determine user type information based on the above script element information.

[0050] User type information can be used to characterize the system's identification and classification results of user attributes, so as to guide subsequent content preference optimization and narrative style selection.

[0051] Specifically, user type information can be preset user classification tags, which can be automatically identified based on script element information or selected by the user. The aforementioned user type information includes, but is not limited to, the following types: international exchange students (with multicultural exchange needs, preferring cross-cultural themes and coming-of-age plots), culture enthusiasts (preferring narrative styles with elements of traditional culture, customs, and historical background), and families with children (content can lean towards heartwarming, educational, and safe storylines).

[0052] S306: Based on the above user type information and the above script element information, the above preset model is used to perform narrative structure generation processing to obtain the initial narrative structure corresponding to the above user type information.

[0053] The initial narrative structure can refer to the draft of the textual plot structure that has not yet been refined into detailed shots and visual mappings, or the overall structured plot framework of script elements and information.

[0054] In some embodiments, in S306, the narrative structure generation process is performed based on the user type information and the script element information using the preset model to obtain the initial narrative structure corresponding to the user type information, including: performing preference optimization processing based on the user type information and the script element information using the preset model to obtain the target plot information and expression style information corresponding to the user type information; and performing narrative structure generation processing based on the target plot information and the expression style information using the preset model to obtain the initial narrative structure.

[0055] Among them, target plot information can refer to plot content that better meets user needs by combining the original script elements with user preferences; expression style information can refer to the personalized settings of the story in terms of language style, narrative tone, and text tone, such as fairy tale style, realistic style, light and humorous style, etc.

[0056] Specifically, user type information and script element information can be input into the preference sub-model in the preset model to obtain the target plot information and expression style information output by the preference sub-model. For example, if the user type information is a parent-child family, the output will be a heartwarming and educational target plot information, as well as a warm and positive expression style information; if the user type information is a culture enthusiast, the output will be a target plot information with historical background and traditional elements, as well as an expression style information with elegant language and rigorous narrative style.

[0057] Furthermore, the target plot information and the aforementioned expression style information can be input into the narrative sub-model in the preset model. The narrative sub-model then processes the input information to generate a narrative structure, thereby generating the initial narrative structure corresponding to the aforementioned user type information.

[0058] In the embodiments described in this specification, by introducing user type information and script element information during the narrative structure generation process, and by performing phased processing based on the preference sub-model and narrative sub-model in the preset model, it is possible to generate more suitable target plot information and expression style information for the preference characteristics of different types of users, thereby constructing an initial narrative structure with a personalized narrative style. This effectively improves the customization of the story content and enhances the diversity and expressiveness of the story's narrative style.

[0059] S308: Perform shot grammar structuring on the above initial narrative structure to obtain the story script information corresponding to the above user type information.

[0060] In some embodiments, in S308, the above-mentioned shot grammar structuring processing of the initial narrative structure to obtain the story script information corresponding to the user type information includes: splitting the initial narrative structure based on preset shot grammar rules to obtain multiple initial script storyboard units; wherein, each of the multiple initial script units includes at least one of the following information: character information, behavior description information, and scene information; performing storyboard generation processing using the preset model based on the multiple initial script storyboard units to obtain multiple script storyboard information, wherein the script storyboard information is used to describe the visual expression content of the storyboard image frames; and determining the story script information corresponding to the user type information based on the multiple script storyboard information.

[0061] The shot grammar rules can be a built-in set of rules used to guide how to segment the initial narrative structure into shot units, i.e., initial script storyboard units. Specifically, the shot grammar rules can be split based on at least one of action changes, character switching, scene changes, and emotional moments to obtain multiple initial script storyboard units.

[0062] Optionally, in each initial script unit, the character information refers to the main characters or roles that participate in the appearance or interaction in that initial script unit; the behavior description information is used to describe the specific actions or states performed by the characters in the initial script unit; and the scene information is used to provide the background environment in which the initial script unit takes place, including time, location, atmosphere, etc.

[0063] In one embodiment, the preset model may include a script segmentation model. Specifically, multiple initial script segmentation units can be input into the script segmentation model included in the initial model to obtain the script segmentation information corresponding to each initial script segmentation unit output by the script segmentation model, thereby obtaining multiple script segmentation information. The script segmentation information corresponds one-to-one with the initial script segmentation units.

[0064] Furthermore, multiple structured script storyboard information can be integrated and organized in any way according to preset narrative logic relationships, shot order, content association, etc., to generate a structured script data set for fully expressing the story content, which can be used as story script information for subsequent image generation, script export, or editing presentation.

[0065] This specification's embodiments introduce a joint modeling mechanism for user type information and script element information into the story script generation process. Based on the preference sub-model and narrative sub-model set in the preset model, it can differentiate content preferences and language style characteristics for different user types, thereby generating a highly personalized initial narrative structure. Building upon this, the initial narrative structure is further decomposed using shot grammar rules, and structured script shot information is generated by combining it with a script shot model, thus constructing logically complete, clearly expressed, and semantically refined story script information. This significantly improves the quality of the generated story content in terms of structural rigor, visual guidance, and style matching, achieving systematic and automated processing of story narration from semantic understanding and structural construction to shot expression. It also enhances the system's adaptability to story generation for different user types and the efficiency of visual conversion of script content.

[0066] In one embodiment, the process of determining the storyboard image sequence in step S206 above can be found in [reference needed]. Figure 4 This is a flowchart illustrating a method for determining a storyboard image sequence, provided in an embodiment of this specification. Figure 4 As shown, the method for determining the storyboard image sequence includes the following steps: S402: Generate prompt information based on the above story script information.

[0067] In some embodiments, the above-mentioned prompt information includes user description information and / or scene description information for guiding image generation; wherein, the user description information can be used to describe the character information in the image, such as appearance, posture, clothing, emotion, etc.; the scene description information can be used to describe the background setting, such as location, time, atmosphere, environment, etc.

[0068] Optionally, in S402, the above-mentioned generation of prompt information based on the story script information includes: extracting character information from the story script information and generating corresponding user description information based on the character information using a preset user description template; and / or, extracting background information from the story script information and generating corresponding scene description information based on the background information using a preset scene description template.

[0069] Specifically, character names, actions, clothing, identities, and emotions are extracted from the storyboard information of each scene. A standardized description is then constructed using a preset user cue template to generate user description information. Additionally, elements such as location, time, atmosphere, and environment are extracted from the storyboard information of each scene. A standardized description is then constructed using a preset scene cue template to generate scene description information. For example, the preset user cue template is "A (character name), Clothing (clothing), Expression (emotion), Is (action)". From the storyboard information, we can extract that the character name is "student," wearing a school uniform, the emotion is "excited," and the action is "waving." Therefore, the user description information is: "A student, wearing a school uniform, with an excited expression, is waving."

[0070] S404: Based on the storyboard information of each script in the above storyboard information, the above preset model is used to perform feature parsing processing to obtain the image content information corresponding to the storyboard information.

[0071] The image content information may include, but is not limited to: human characteristics (appearance, expression, posture), scene characteristics (time, location, environmental composition), composition information (viewpoint, shot size, shot type), style information (cartoon / realistic, bright / dark, etc.), semantic tags (action, emotion, relationship), etc.

[0072] Optionally, the preset model may include a parsing sub-model, which can input the storyboard information of each script in the storyboard information into the parsing sub-model included in the preset model to obtain the image content information corresponding to the storyboard information output by the parsing sub-model.

[0073] S406: Based on the above-mentioned script storyboard information and the above-mentioned prompt information, the above-mentioned preset model is used to perform image frame generation processing to obtain the storyboard image frames corresponding to the above-mentioned script storyboard information.

[0074] The preset model may include an image generation sub-model, which can input the storyboard information of each script and the above-mentioned prompt information into the image generation sub-model included in the preset model to obtain the storyboard image frames corresponding to the storyboard information of each script output by the image generation sub-model.

[0075] It is understandable that for multiple different storyboard scenes in the above storyboard information, the user description information and scene description information in the prompt information can be reused or kept stable during the generation of multiple storyboard image frames, so that the same character has consistent appearance characteristics in multiple shots, and the same scene maintains consistent visual style in different shots.

[0076] S408: Arrange the multiple storyboard image frames corresponding to the multiple storyboard information in the above storyboard information to generate the storyboard image sequence corresponding to the above storyboard information.

[0077] Optionally, arranging multiple storyboard image frames corresponding to multiple storyboard information in the script refers to sequentially organizing and combining the generated image frames according to the chronological order or logical structure of each storyboard information in the story script to generate a sequence of storyboard images for display or export. This image sequence can visually express the complete story process and can be used to assist in script previewing, storyboard construction, or subsequent content editing.

[0078] This specification's embodiments convert structured content (such as storyboard information) in the story script into prompts that the image generation model can recognize. Combined with feature analysis and image generation processes, this achieves automatic generation from text scripts to image frames. By extracting character and background information and using preset templates to standardize user description and scene description information, consistency and completeness of the prompt input are ensured, guaranteeing the accuracy and stylistic uniformity of image generation. Simultaneously, a sub-model is used to perform fine-grained semantic analysis of the storyboard information, further enhancing the ability to reproduce image content. Each storyboard and prompt jointly drives the image generation sub-model to output high-quality storyboard image frames. Multiple image frames are arranged in a logical order to form a storyboard image sequence, constituting a visual story flow. This effectively improves the efficiency of automatic text-to-image conversion, enhances the visual expressiveness and editability of the story generation system, and is applicable to various scenarios such as intelligent script-assisted creation, animation preview, and storyboard generation, demonstrating good practicality and scalability.

[0079] In one embodiment, such as Figure 5 As shown, another method for generating storyboards is provided, which includes at least the following steps: S502: Obtain story request information uploaded by users.

[0080] Specifically, S502 is the same as S202, and will not be described again here.

[0081] S504: Based on the above story requirements information, a preset model is used to generate the story script, and the corresponding story script information is obtained.

[0082] Specifically, S504 is the same as S204, and will not be described again here.

[0083] S506: Based on the above story script information, the above preset model is used to perform image generation processing to obtain the storyboard image sequence corresponding to the above story script information.

[0084] Specifically, S506 is the same as S206, and will not be repeated here.

[0085] S508: Displays style switching controls for the above storyboard image sequence.

[0086] Optionally, the style switching control can be at least one of the following: a drop-down menu, a slider button, a preview thumbnail style selection area, etc. The drop-down menu can include selections of various styles such as graffiti style, line drawing, and painted style; the slider button can be used by the user to adjust the level of detail, color saturation, etc. The style switching control provides a visual interactive interface for the user, allowing them to apply different visual styles to existing storyboard image sequences.

[0087] It is understandable that the style switching control can correspond to different image post-processing models, used to apply specified visual style transfer processing to the current storyboard image frame. Image post-processing models include, but are not limited to, image style transfer models, generative adversarial network models, diffusion model parameter fine-tuning combinations, and image filter enhancement modules.

[0088] S510: Upon detecting the triggering operation of the style switching control, the step of displaying the storyboard image sequence and generating storyboard information based on the storyboard image sequence is performed based on the adjusted storyboard image sequence. In response to the determination operation of the storyboard image sequence, the step of generating storyboard information based on the storyboard image sequence is performed.

[0089] Understandably, after previewing the adjusted style version, users can click the "OK" or "Apply this style" button to trigger an operation on the style switching control. When the system detects this trigger operation, it confirms that the current style version is the style for the storyboard to be exported later.

[0090] This specification's embodiments introduce a style switching control, providing users with more flexible and controllable visual style customization capabilities. Users can easily select or adjust the image's expressive style, such as cartoon, sketch, watercolor, or realistic, through interactive methods like drop-down menus, sliders, or preview images, achieving personalized customization of the storyboard image sequence's appearance. The image post-processing model associated with the style switching control can perform visual style conversion on existing storyboard image frames, thus achieving diverse image style expressions without altering the content. After the user confirms the selected style interactively, the final storyboard information can be directly generated based on the current style version, significantly improving the system's usability and the artistic expressiveness of the generated results, enhancing the system's technical advantages in human-computer interaction, generation controllability, and output diversity.

[0091] In one embodiment, such as Figure 6 As shown, another method for generating storyboards is provided, which includes at least the following steps: S602: Obtain story request information uploaded by users.

[0092] Specifically, S602 is the same as S202, and will not be described again here.

[0093] S604: Based on the above story requirements information, a preset model is used to generate the story script, and the corresponding story script information is obtained.

[0094] Specifically, S604 is the same as S204, and will not be described again here.

[0095] S606: Based on the above story script information, the above preset model is used to perform image generation processing to obtain the storyboard image sequence corresponding to the above story script information.

[0096] Specifically, S606 is the same as S206, and will not be described in detail here.

[0097] S608: Display the above storyboard image sequence, and in response to a preset operation for a local image region in any storyboard image frame, obtain image processing instructions.

[0098] The local image region can refer to a portion of the image selected by the user (such as a person's face, props, or a corner of the background). Preset operations can be specific actions that the user can perform on the image region, such as circling, clicking, dragging a box, or marking.

[0099] In some embodiments, the above image processing instructions can be used to express the user's modification requirements for an image region, such as changing the color, replacing objects, adjusting textures, etc.

[0100] S610: Based on the above image processing instructions and the above local image region, the above preset model is used to modify the image region to obtain the target storyboard image frame.

[0101] Optionally, the target storyboard image frame can be a new image frame obtained by modifying a local region based on the original image frame.

[0102] Specifically, the preset model may include a modification sub-model. The aforementioned image processing instructions and the aforementioned local image regions can be input into the modification sub-model included in the preset model. The modification sub-model then performs image region modification processing, thereby obtaining the target storyboard image frame output by the modification sub-model. The modification sub-model can be an image restoration model, a local redrawing model, or an image completion model based on the diffusion principle, supporting high-fidelity modification based on the original image content context.

[0103] In one embodiment, the modified sub-model described above can be fine-tuned using the Low-Rank Adaptation of Large Language Models (LoRA) method to support efficient adaptive optimization in image region modification tasks. This method combines user-input image processing commands with the local modification region, introducing a low-rank trainable matrix to instrument and update key layers in the model. While maintaining the stability of the original model, targeted adjustments are made, thereby improving the quality and consistency of the generated local images and meeting users' personalized editing needs for visual content.

[0104] S612: In response to the determination instruction for the target storyboard image frame, the storyboard image sequence is updated based on the target storyboard image frame to obtain an updated storyboard image sequence.

[0105] Optionally, after the user confirms the target storyboard image frame, a confirmation command can be generated. In response to the confirmation command, the original storyboard image frame can be replaced with the target image frame, and the content of the current storyboard image sequence can be updated based on the replacement operation, thereby obtaining an updated storyboard image sequence including the modified image frame.

[0106] S614: Based on the updated storyboard image sequence, perform the step of displaying the storyboard image sequence and generating storyboard information based on the storyboard image sequence in response to the determination operation for the storyboard image sequence.

[0107] In some embodiments, the updated storyboard image sequence can be re-displayed for user review and confirmation. If the user performs a confirmation operation on the current sequence, the system can generate updated storyboard information based on all image frames in the sequence and their corresponding script descriptions for subsequent export, rendering, or editing.

[0108] S616: In response to the determination instruction for the above-mentioned target storyboard image frame, generate modification record information.

[0109] The modification log information can be structured data used to record image editing operations, which facilitates version management and learning feedback.

[0110] In one embodiment, after the user confirms the target storyboard image frame, operation information related to the modification of that image is automatically recorded, generating structured modification record information. This recorded information may include the original image frame identifier, the location of the modified area, the image processing instruction content, processing model parameters, user confirmation time, etc., for version management, traceability control, or feedback mechanism support.

[0111] S618: Based on the above modification record information, generate reinforcement learning feedback data.

[0112] The aforementioned reinforcement learning feedback data is used to guide the aforementioned preset model to adaptively optimize image generation parameters.

[0113] Understandably, reinforcement learning feedback data consists of training feedback samples extracted from user modification behavior, used to optimize the image generation model to make it more closely resemble user preferences.

[0114] S620: Optimize and train the model parameters of the preset model based on the reinforcement learning feedback data, and update the preset model based on the optimized model parameters.

[0115] In some embodiments, reinforcement learning feedback data can be constructed based on the aforementioned modification record information to guide the preset model in adaptive optimization for future image generation tasks. This feedback data can be used to train the reward model, update the policy network, or optimize generation parameters, thereby enabling the model to better match the user's actual preferences and editing behavior, achieving personalized adaptation and continuous learning.

[0116] This specification's embodiments significantly improve the system's image generation accuracy, user engagement, and model adaptability by introducing an interactive image modification mechanism and a user behavior-based reinforcement learning feedback mechanism. On one hand, it supports users in performing preset operations such as selecting and clicking on local areas (e.g., facial features, prop details) in any storyboard image frame. Based on image processing instructions, it calls the image modification sub-model in the preset model to achieve high-fidelity modification and redrawing of local image areas, improving the visual detail quality and expressive accuracy of the storyboard image frame. On the other hand, after the user confirms the modification result, structured modification record information is generated, and reinforcement learning feedback data is constructed based on this information to guide model parameter optimization training. Through this mechanism, the preset model can continuously absorb user preferences, adjust strategies, and evolve its generation capabilities, thereby achieving personalized and dynamically optimized content production capabilities. Overall, it not only enhances the editability and interactivity of the storyboard generation process but also promotes continuous model optimization through a closed-loop learning mechanism, effectively improving user satisfaction with the system's output results, model robustness, and the professionalism of content generation.

[0117] This specification also provides a method for training a preset model, which includes: creating an initial model based on a basic large model; obtaining story requirement sample information and a target storyboard image sequence corresponding to the story requirement sample information; inputting the story requirement sample information into the initial model for model training; performing story script generation processing on the story requirement sample information through the initial model to obtain corresponding story script sample information; and performing image generation processing on the story script sample information through the initial model to obtain a reference storyboard image sequence corresponding to the story script sample information; during the model training process, adjusting the model parameters of the initial model based on the target storyboard image sequence and the reference storyboard image sequence to obtain the preset model.

[0118] It is understandable that in the above model training process, a model loss function is constructed based on the parameters corresponding to the target storyboard image sequence and the above reference storyboard image sequence, thereby using the target storyboard image sequence and the above reference storyboard image sequence to determine the model loss value of the initial model based on the model loss function. Then, the model parameters of the initial model are adjusted using the model loss value, and then the model is trained again until the initial model converges after adjusting the model parameters, thus obtaining the preset model.

[0119] The base model can be a multimodal model, which is a large-scale deep learning model capable of processing multiple modalities of data (such as text, images, and audio). It can consist of multiple sub-models, each specializing in processing one modality of data. Through cross-modal learning, it integrates information from different modalities to achieve more accurate and comprehensive information extraction and understanding. When creating an initial model using the base model, the base model can be initially trained to obtain the initial model; or the model parameters of the base model can be adaptively adjusted to obtain the initial model.

[0120] In the embodiments of this specification, an initial model is constructed based on a multimodal basic large model, and supervised training is performed by combining story requirement sample information labeled with target storyboard image sequences. The model parameters are then continuously adjusted to minimize the difference between the target storyboard image sequence and the aforementioned reference storyboard image sequence, and finally a converged preset model is obtained, thus achieving high-precision and strong generalization ability model construction.

[0121] Next, combine Figure 7 This document describes the specific storyboard generation method provided in the embodiments of this specification. Please refer to [link / reference needed] for details. Figure 7 This is a schematic diagram illustrating a specific process for a storyboard generation method provided in the embodiments of this specification. For example... Figure 7 As shown, the story request information uploaded by the user is obtained through the front-end user interface. The front-end user interface may include a story request information retrieval interface; please refer to [link to relevant documentation]. Figure 8 The story requirement information acquisition interface 81 includes a first input box 811, which displays the prompt "Enter the desired user storyboard". The first input box 811 can be used for users to upload story requirement information, specifically including a first file addition control 812 and a first file upload control 813. Additionally, the story requirement information acquisition interface 81 also includes an initial dialogue prompt such as "What kind of storyboard do you want? Tell me quickly~". Through this story requirement information acquisition interface 81, user-uploaded story requirement information can be obtained, and then this information can be input into a preset model in the backend. A story script information is generated through a large language model (which may include the aforementioned preference sub-model, narrative sub-model, script storyboard model, etc.), and a storyboard image sequence is generated through a diffusion model (which may include the aforementioned parsing sub-model, image generation sub-model, etc.), and then transmitted to the front-end user interface to display the storyboard image sequence.

[0122] Please refer to Figure 9 , Figure 9The first storyboard image sequence display page 91 includes a second input box 911, which can be used for users to upload supplementary story requirements. Specifically, it can include a second file addition control 912 and a second file upload control 913. In addition, the first storyboard image sequence display page 91 can also include dialogue information, such as the user inputting natural language description information "Help me generate a museum user storyboard", and the feedback information of the preset model "The museum mainly has the following three typical users: international exchange students, culture enthusiasts, and families with children. How do they visit the museum? Take a look at the pictures~". The feedback information can also include the storyboard image sequence corresponding to each user type information (such as international exchange students, culture enthusiasts, and families with children) (for example, the sequence corresponding to international exchange students includes storyboard image frame 1, storyboard image frame 2, storyboard image frame 3, storyboard image frame 4, storyboard image frame 5, and storyboard image frame 6). The first storyboard image sequence display page 91 may also include a first download control 914. When the first download control 914 is triggered, the corresponding storyboard information is generated based on the currently displayed storyboard image sequence, and the relevant file is generated and downloaded to the local terminal. The file can be a PDF file or a slide presentation file.

[0123] It's worth noting that when a user's command to view any storyboard image frame is detected (such as the user controlling the mouse cursor to hover over the display area of ​​a specific storyboard image frame), the storyboard image frame details page can be displayed on the current page. Please refer to [link to details] for further information. Figure 10 , Figure 10 The second storyboard image sequence display page 101 includes a third input box 1011, a third file addition control 1012, a third file upload control 1013, a third download control 1014, and dialogue information. The dialogue information includes a storyboard image frame details page 1015 corresponding to storyboard image frame 1. The storyboard image frame details page 1015 displays user description information and scene description information corresponding to storyboard image frame 1, and can also display the image generation process of storyboard image frame 1, such as intermediate image frames involved in the image generation process.

[0124] Furthermore, the storyboard image frame details page 1015 also allows users to perform secondary editing operations on specified storyboard image frames, including but not limited to: modifying the original prompt information, applying style changes, and adjusting local areas (such as partial redrawing, color changes, object replacement, etc.). After the editing operation is completed, the target storyboard image frame can be regenerated based on the updated prompt information or image processing instructions, and the content at the corresponding position can be updated. Finally, users can confirm the entire storyboard image sequence, and the system supports packaging the current version of the complete storyboard image sequence and its associated text information into a standard format user storyboard file for export, download, and subsequent use. This file can be used in scenarios such as creative proposals, content production, and visual scripts, and has high visual appeal and interactivity.

[0125] This specification covers the entire process from structured user profile analysis, semantic-driven story script generation, shot-level image content construction, to the visualization and interactive editing and export of storyboard image sequences. By integrating a large language model, diffusion model, image editing model, and reinforcement learning feedback mechanism, it achieves automatic generation and controllable optimization from user need intent recognition to visual expression content, significantly reducing the time cost of traditional manual conception and storyboard drawing. It is particularly suitable for user experience design, marketing creativity, cultural tourism guidance, education and training, and other scenarios, and can assist designers in quickly generating story content and visual scripts that conform to the preferences and expression styles of the target audience based on specific user profiles. In addition, the front-end interface provides a visual storyboard preview, style switching controls, single-frame image details, and editing operation entry points, so that users are not only passive recipients in the end-to-end process, but also active participants and real-time feedback providers, enhancing the interactivity and flexibility of the entire content production process. Therefore, this end-to-end integrated generation and editing process significantly simplifies the complex operations of multiple stages in traditional storyboard design, helping content creators or user experience designers to quickly achieve a closed loop from story conception to final visual content output, improving creation efficiency and content consistency, and demonstrating good practicality.

[0126] Please refer to Figure 11 The invention provides a storyboard generation apparatus. The storyboard generation apparatus 1100 includes: The first acquisition module 1110 is used to acquire story requirement information uploaded by users; The first processing module 1120 is used to generate a story script based on the above-mentioned story requirement information using a preset model to obtain the corresponding story script information. The second processing module 1130 is used to perform image generation processing based on the above story script information using the above preset model to obtain the story script information corresponding to the story script image sequence. The first generation module 1140 is used to display the above-mentioned storyboard image sequence and, in response to the determination operation of the above-mentioned storyboard image sequence, generate storyboard information based on the above-mentioned storyboard image sequence.

[0127] In one possible implementation, the first processing module 1120 includes: The first parsing unit is used to perform semantic parsing based on the above story requirement information using a preset model to obtain script element information. The script element information includes at least one of the following: behavioral appeal information, story scene information, and behavioral path information. The first determining unit is used to determine user type information based on the above-mentioned script element information; The second determining unit is used to perform narrative structure generation processing based on the above-mentioned user type information and script element information using the above-mentioned preset model to obtain the initial narrative structure corresponding to the above-mentioned user type information. The third determining unit is used to perform shot grammar structuring on the above initial narrative structure to obtain the story script information corresponding to the above user type information.

[0128] In one possible implementation, the second determining unit includes: The first processing subunit is used to perform preference optimization processing based on the above-mentioned user type information and script element information using the above-mentioned preset model to obtain the target plot information and expression style information corresponding to the above-mentioned user type information. The second processing subunit is used to generate the narrative structure based on the aforementioned target plot information and the aforementioned expression style information using the aforementioned preset model, thereby obtaining the initial narrative structure.

[0129] In one possible implementation, the third determining unit includes: The sub-units are used to split the initial narrative structure based on preset shot grammar rules to obtain multiple initial script storyboard units; wherein each of the multiple initial script units includes at least one of the following information: character information, behavior description information, and scene information; The third processing subunit is used to perform storyboard generation processing based on the above-mentioned multiple initial script storyboard units using the above-mentioned preset model to obtain multiple script storyboard information. The above-mentioned script storyboard information is used to describe the visual expression content of the storyboard image frame. The first determining subunit is used to determine the story script information corresponding to the user type information based on the above multiple script storyboard information.

[0130] In one possible implementation, the second processing module 1130 includes: The first generation unit is used to generate prompt information based on the aforementioned story script information; The second parsing unit is used to perform feature parsing processing based on the storyboard information of each script scene in the storyboard information, and to obtain the image content information corresponding to the script scene information using the preset model. The processing unit is used to perform image frame generation processing based on the above-mentioned script storyboard information and the above-mentioned prompt information, using the above-mentioned preset model to obtain the storyboard image frames corresponding to the above-mentioned script storyboard information. The second generation unit is used to arrange multiple storyboard image frames corresponding to multiple storyboard information in the story script information to generate a storyboard image sequence corresponding to the story script information.

[0131] In one possible implementation, the above prompt information includes user description information and / or scene description information to guide image generation; The aforementioned first generation unit includes: The first extraction subunit is used to extract character information from the aforementioned story script information, and generate corresponding user description information based on the aforementioned character information using a preset user description template; and / or, The second extraction subunit is used to extract background information from the story script information and generate corresponding scene description information based on the background information using a preset scene description template.

[0132] In one possible implementation, the device 1100 further includes: The display module is used to display the style switching controls for the above storyboard image sequence; The third processing module is used to perform style adjustment processing based on the preset model according to the above-mentioned storyboard image sequence and the target image style corresponding to the above-mentioned style switching control when the trigger operation of the above-mentioned style switching control is detected, and to generate the adjusted storyboard image sequence. The first execution module is configured to perform the steps of displaying the storyboard image sequence based on the adjusted storyboard image sequence, and generating storyboard information based on the storyboard image sequence in response to the determination operation for the storyboard image sequence.

[0133] In one possible implementation, the device 1100 further includes: The second acquisition module is used to acquire image processing instructions in response to a preset operation on a local image region in any storyboard image frame after displaying the above storyboard image sequence. The fourth processing module is used to modify the image region based on the above image processing instructions and the above local image region using the above preset model to obtain the target storyboard image frame. The update module is used to update the storyboard image sequence based on the target storyboard image frame in response to the determination instruction for the target storyboard image frame, so as to obtain the updated storyboard image sequence. The second execution module is used to perform the steps of displaying the storyboard image sequence based on the updated storyboard image sequence, and generating storyboard information based on the storyboard image sequence in response to the determination operation for the storyboard image sequence.

[0134] In one possible implementation, the device 1100 further includes: The second generation module is used to generate modification record information in response to the determination instruction for the above-mentioned target storyboard image frame; The third generation module is used to generate reinforcement learning feedback data based on the above-mentioned modification record information; The training module is used to optimize and train the model parameters of the preset model based on the reinforcement learning feedback data, and update the preset model based on the optimized model parameters.

[0135] The division of modules in the storyboard generation device described above is for illustrative purposes only. In other embodiments, the storyboard generation device can be divided into different modules as needed to complete all or part of the functions of the storyboard generation device. The implementation of each module in the storyboard generation device provided in the embodiments of this specification can be in the form of a computer program. This computer program can run on a terminal or server. The program modules constituted by this computer program can be stored in the memory of the terminal or server. When the computer program is executed by a processor, it implements all or part of the steps of the storyboard generation method described in the embodiments of this specification.

[0136] This specification also provides an electronic device, which may be a server, and its internal structure diagram may be as follows: Figure 12 As shown, this electronic device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. The processor executes computer programs to implement a storyboard generation method.

[0137] Those skilled in the art will understand that Figure 12 The structures shown are merely block diagrams of some structures related to the embodiments of this specification, and do not constitute a limitation on the electronic devices to which the embodiments of this specification are applied. Specific electronic devices may include more or fewer components than those shown in the figures, or combine certain components, or have different component arrangements.

[0138] In one possible implementation, a computer storage medium is provided that stores instructions, which, when executed on a computer or processor, cause the computer or processor to perform one or more steps in the above embodiments. If the constituent modules of the above-described electronic device are implemented as software functional units and sold or used as independent products, they can be stored in the aforementioned computer-readable storage medium.

[0139] In one possible implementation, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0140] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When these computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this specification are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer storage medium or transmitted through the computer storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, Digital Subscriber Line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The aforementioned available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., Digital Versatile Discs (DVDs)), or semiconductor media (e.g., Solid State Disks (SSDs)).

[0141] It should be noted that the information (including but not limited to story requirement information), data (including but not limited to data used for analysis, stored data, and displayed data), and signals involved in the embodiments of this specification are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the modification record information involved in this specification was obtained under full authorization.

[0142] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. The aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks. Unless otherwise specified, the technical features of this embodiment and its implementation can be combined arbitrarily.

[0143] The above-described embodiments are merely preferred embodiments of the embodiments in this specification and are not intended to limit the scope of the embodiments in this specification. Various modifications and improvements made by those skilled in the art to the technical solutions of the embodiments in this specification without departing from the design spirit of the embodiments in this specification should fall within the protection scope defined by the claims.

[0144] The foregoing has described specific embodiments of the embodiments described in this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired results. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

Claims

1. A storyboard generation method, comprising: Obtain user-uploaded story request information; Based on the story requirement information, a preset model is used to generate the story script, and the corresponding story script information is obtained. Based on the story script information, the preset model is used to perform image generation processing to obtain the story script information corresponding to the story image sequence. The storyboard image sequence is displayed, and in response to a determination operation for the storyboard image sequence, storyboard information is generated based on the storyboard image sequence.

2. The method as described in claim 1, wherein the step of generating a story script based on the story requirement information using a preset model to obtain the corresponding story script information includes: Based on the story requirement information, a preset model is used to perform semantic parsing to obtain script element information, which includes at least one of the following: behavioral appeal information, story scene information, and behavioral path information; Determine user type information based on the script element information; Based on the user type information and the script element information, the preset model is used to generate the narrative structure, thereby obtaining the initial narrative structure corresponding to the user type information. The initial narrative structure is processed using shot grammar structuring to obtain the story script information corresponding to the user type information.

3. The method as described in claim 2, wherein the step of performing narrative structure generation processing based on the user type information and the script element information using the preset model to obtain the initial narrative structure corresponding to the user type information includes: Based on the user type information and the script element information, the preset model is used to perform preference optimization processing to obtain the target plot information and expression style information corresponding to the user type information; Based on the target plot information and the expression style information, the narrative structure is generated using the preset model to obtain an initial narrative structure.

4. The method as described in claim 2, wherein performing shot grammar structuring on the initial narrative structure to obtain the story script information corresponding to the user type information includes: The initial narrative structure is broken down based on preset shot grammar rules to obtain multiple initial script storyboard units; wherein, each of the multiple initial script units includes at least one of the following information: character information, behavior description information, and scene information; The multiple initial script storyboard units are used to perform storyboard generation processing using the preset model to obtain multiple script storyboard information, which is used to describe the visual expression content of the storyboard image frames. The story script information corresponding to the user type information is determined based on the multiple script storyboard information.

5. The method as described in claim 1, wherein the step of performing image generation processing based on the story script information using the preset model to obtain the story script information storyboard image sequence includes: Generate prompt information based on the story script information; Based on the story script information, the preset model is used to perform feature parsing to obtain the image content information corresponding to the story script information. Based on the storyboard information of each script and the prompt information, the preset model is used to perform image frame generation processing to obtain the storyboard image frame corresponding to each storyboard information. Arrange the multiple storyboard image frames corresponding to the multiple storyboard image information in the story script information to generate the storyboard image sequence corresponding to the story script information.

6. The method of claim 5, wherein the prompting information includes user description information and / or scene description information for guiding image generation; The step of generating prompt information based on the story script information includes: Extract character information from the story script information, and generate corresponding user description information based on the character information using a preset user description template; And / or, Extract the background information from the story script information, and generate corresponding scene description information based on the background information using a preset scene description template.

7. The method of claim 1, further comprising: Display style switching controls for the storyboard image sequence; Upon detecting the triggering operation of the style switching control, the preset model is used to perform style adjustment processing based on the storyboard image sequence and the target image style corresponding to the style switching control, thereby generating an adjusted storyboard image sequence. The steps of displaying the storyboard image sequence based on the adjusted storyboard image sequence, and generating storyboard information based on the storyboard image sequence in response to a determination operation for the storyboard image sequence.

8. The method of claim 1, further comprising, after displaying the storyboard image sequence: In response to a preset operation targeting a local image region in any storyboard image frame, image processing instructions are obtained; Based on the image processing instructions and the local image region, the preset model is used to modify the image region to obtain the target storyboard image frame; In response to the determination command for the target storyboard image frame, the storyboard image sequence is updated based on the target storyboard image frame to obtain an updated storyboard image sequence; The steps of displaying the storyboard image sequence based on the updated storyboard image sequence, and generating storyboard information based on the storyboard image sequence in response to a determination operation for the storyboard image sequence.

9. The method of claim 8, further comprising: In response to a determination command for the target storyboard image frame, modification record information is generated; Based on the modified record information, reinforcement learning feedback data is generated; The model parameters of the preset model are optimized and trained based on the reinforcement learning feedback data, and the preset model is updated based on the optimized model parameters.

10. A storyboard generation device, comprising: The first acquisition module is used to acquire story request information uploaded by users; The first processing module is used to generate a story script based on the story requirement information using a preset model to obtain the corresponding story script information. The second processing module is used to perform image generation processing based on the story script information using the preset model to obtain the story script information corresponding to the story script information; The first generation module is used to display the storyboard image sequence and, in response to a determination operation on the storyboard image sequence, generate storyboard information based on the storyboard image sequence.

11. An electronic device, comprising: Processor and memory; The processor is connected to the memory; The memory is used to store executable program code; The processor runs a program corresponding to the executable program code stored in the memory to perform the method as described in any one of claims 1-9.

12. A computer storage medium storing a plurality of instructions adapted for loading by a processor and executing the method as claimed in any one of claims 1-9.

13. A computer program product comprising instructions that, when run on a computer or processor, cause the computer or processor to perform the method as described in any one of claims 1-9.