Stroma video generation method and related device
By visualizing the generation process of storyboard images and videos, the problem of insufficient transparency of intermediate results in traditional narrative video generation is solved, enabling users to intuitively verify and adjust, thus improving generation quality and efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-20
- Publication Date
- 2026-03-27
AI Technical Summary
Traditional methods of generating narrative videos lack transparency in intermediate results, making it difficult for users to verify and adjust the reasonableness of intermediate results, thus affecting the quality of the generated content.
This paper provides a method for generating storyboard videos. By visualizing the generation process of storyboard images and videos, the intermediate results are made transparent, and reference information is provided to users during the generation process to facilitate verification and adjustment.
It improves the quality of generated story videos, ensures the reasonableness of intermediate results, avoids common problems such as character misplacement, and enhances generation efficiency and quality.
Smart Images

Figure CN121750951A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computers, in particular to a plot video generation method and related device. BACKGROUND
[0002] Plot video generation is a core creation link in the fields of films, games, short videos, etc., and its goal is to deliver a specific theme or emotion through building coherent narrative logic and visual presentation.
[0003] The traditional creation process highly depends on manual work and needs to be completed by multiple roles such as screenwriters, directors, actors, and editors, which has problems such as long cycle, high cost, and limited creativity. With the rapid development of Artificial Intelligence Generated Content (AIGC) technology, automatic plot video generation has become an industry demand.
[0004] In order to realize automatic plot video generation, related technologies generate plot videos in an end-to-end video manner, but this manner often lacks transparency of intermediate results, which makes it difficult for users to check and adjust the rationality of the intermediate results, thereby affecting the quality of the generated plot videos. SUMMARY
[0005] To solve the above technical problems, the present application provides a plot video generation method and related device, the generation process of the storyboard picture and the storyboard video is visualized, the intermediate result, i.e., the target storyboard picture and the target storyboard video, is transparent, and reference information is provided to the user during the generation process of the storyboard picture and the storyboard video, so that the user can intuitively and accurately check and adjust the intermediate result according to the reference information, to ensure the rationality of the intermediate result, and thus improve the quality of the generated plot video.
[0006] The present application embodiment discloses the following technical solutions: In one aspect, the present application embodiment provides a plot video generation method, which comprises: displaying a storyboard picture generation interface, the storyboard picture generation interface comprising a first plot description of a current storyboard; in response to a picture generation operation performed on the storyboard picture generation interface, displaying a target storyboard picture of the current storyboard, and the first plot description is used to check whether the target storyboard picture conforms to the storyboard intention; in response to a first confirmation operation on the target storyboard picture, entering a storyboard video generation interface, the storyboard video generation interface comprising a second plot description of the current storyboard and attribute information of a dubbing role; In response to a video generation operation performed on the shot video generation interface, a target shot video of the current shot generated based on the target shot picture is displayed, the attribute information is used to verify a matching degree of a current dubbing role in the target shot video and the second plot description, and the target shot video is used to generate the plot video.
[0007] In an aspect, an embodiment of the present application provides a plot video generation device, the device comprising a display unit and an entering unit: The display unit is configured to display a shot picture generation interface, the shot picture generation interface comprising a first plot description of a current shot. The display unit is further configured to, in response to a picture generation operation performed on the shot picture generation interface, display a target shot picture of the current shot generated, and the first plot description is used to verify whether the target shot picture meets a shot intention. The entering unit is configured to, in response to a first confirmation operation on the target shot picture, enter a shot video generation interface, the shot video generation interface comprising a second plot description of the current shot and attribute information of a dubbing role. The display unit is further configured to, in response to a video generation operation performed on the shot video generation interface, display a target shot video of the current shot generated based on the target shot picture, the attribute information is used to verify a matching degree of a current dubbing role in the target shot video and the second plot description, and the target shot video is used to generate the plot video.
[0008] In an aspect, an embodiment of the present application provides a computer device, the computer device comprising a processor and a memory: The memory is configured to store a computer program and transmit the computer program to the processor. The processor is configured to execute the method according to the instructions in the computer program.
[0009] In an aspect, an embodiment of the present application provides a computer readable storage medium, the computer readable storage medium being configured to store a computer program, the computer program causing a processor to execute the method according to any one of the preceding aspects when executed by the processor.
[0010] In an aspect, an embodiment of the present application provides a computer program product comprising a computer program, the computer program being configured to implement the method according to any one of the preceding aspects when executed by a processor.
[0011] It can be seen from the technical solution that the generation process of the plot video provided by the application is visualized, and the intermediate result is transparent and can be checked and adjusted by the user. Specifically, when the plot video needs to be generated, a shot picture generation interface can be displayed, and the shot picture generation interface includes a first plot description of the current shot. The first plot description can be used as reference information for shot picture generation. After the target shot picture of the current shot is displayed in response to the picture generation operation performed on the shot picture generation interface, the user can check whether the target shot picture meets the shot intention according to the first plot description, thereby facilitating the user to check and adjust the rationality of the generated shot picture, so as to obtain a target shot picture that meets the shot intention. If the first confirmation operation is performed on the target shot picture, the shot video generation interface is entered in response to the first confirmation operation, and the shot video generation interface includes a second plot description of the current shot and attribute information of the voice-over role. The second plot description and the attribute information of the voice-over role can be used as reference information for shot video generation. After the target shot video of the current shot generated based on the target shot picture is displayed in response to the video generation operation performed on the shot video generation interface, the user can verify the matching degree between the current voice-over role in the generated shot video and the second plot description according to the attribute information, thereby facilitating the user to check and adjust the rationality of the generated shot video, so as to avoid common problems such as role misplacement, so as to obtain a reasonable target shot video, and then generate a reasonable plot video according to the target shot video. It can be seen that the generation process of the shot picture and the shot video is visualized, the intermediate result, i.e., the target shot picture and the target shot video, is transparent, and reference information is provided to the user during the generation process of the shot picture and the shot video, thereby facilitating the user to intuitively and accurately check and adjust the intermediate result according to the reference information, so as to ensure the rationality of the intermediate result, and thereby improve the quality of the generated plot video. BRIEF DESCRIPTION OF DRAWINGS
[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and those skilled in the art can also obtain other drawings according to these drawings without creative labor.
[0013] Figure 1 An architecture diagram of a computer system provided by an embodiment of the present application; Figure 2 An application scenario architecture diagram of a plot video generation method provided by an embodiment of the present application; Figure 3 A flowchart of a plot video generation method provided by an embodiment of the present application; Figure 4 An example diagram of a split picture generation interface provided for an embodiment of the present application; Figure 5 An example diagram of another split picture generation interface provided for an embodiment of the present application; Figure 6 An example diagram of a split video generation interface provided for an embodiment of the present application; Figure 7 An example diagram of yet another split picture generation interface provided for an embodiment of the present application; Figure 8 An example diagram of an interface showing each hierarchical node provided for an embodiment of the present application; Figure 9 An example diagram of a video prompt word provided for an embodiment of the present application; Figure 10 An example diagram of another split video generation interface provided for an embodiment of the present application; Figure 11 An example diagram of still another split picture generation interface provided for an embodiment of the present application; Figure 12 An example diagram of a split director planning interface provided for an embodiment of the present application; Figure 13 An example diagram of a subject design interface provided for an embodiment of the present application; Figure 14 An example diagram of inserting a subject reference picture provided for an embodiment of the present application; Figure 15 An example diagram of a process of generating a target split picture provided for an embodiment of the present application; Figure 16 An example diagram of another process of generating a target split picture provided for an embodiment of the present application; Figure 17 A detailed design diagram of a context management project provided for an embodiment of the present application; Figure 18 An example diagram of a complete process of generating a plot video provided for an embodiment of the present application; Figure 19 An example diagram of a script creation interface provided for an embodiment of the present application; Figure 20 A flowchart of a hierarchical iterative optimization mechanism provided for an embodiment of the present application; Figure 21 A structural diagram of a plot video generation apparatus provided for an embodiment of the present application; Figure 22 A structural diagram of a terminal provided for an embodiment of the present application; Figure 23A structural diagram of a server provided in an embodiment of the present application. DETAILED DESCRIPTION
[0014] Embodiments of the present application are described below with reference to the accompanying drawings.
[0015] The related art generates a plot video in an end-to-end video manner, however, this manner often lacks transparency of intermediate results, which makes it difficult for a user to check and adjust the rationality of the intermediate results, thereby affecting the quality of the generated plot video.
[0016] To solve the above technical problems, the embodiments of the present application provide a plot video generation method, the generation process of the plot video is visualized, and the obtained intermediate results, i.e., target shot pictures and target shot videos, are transparent, which can be checked and adjusted by the user. In addition, reference information is provided to the user during the generation process of the shot pictures and the shot videos, so that the user can intuitively and accurately check and adjust the intermediate results according to the reference information, so as to ensure the rationality of the intermediate results, and further improve the quality of the generated plot video.
[0017] It should be noted that the plot video generation method provided in the embodiments of the present application can be applied to various scenes that need to generate a plot video, such as film and television production, advertisement shooting, game development, short video creation, virtual reality, etc. The plot video can be a video that tells a complete plot by combining elements such as pictures, sounds, and performances. In the above scenes, a video reflecting a complete plot may be needed, so the plot video generation method provided in the embodiments of the present application can be applied to the above scenes, and the following scenes are briefly introduced as follows: (1) Film and television production scene: In film and television production, there are more and more complex scenes, especially in science fiction, fantasy, and animation films and television dramas, the construction of characters and scenes requires a lot of manpower and resources. The AIGC technology can automatically generate a plot video, thereby reducing the generation cost of the plot video and improving the generation efficiency. In order to further ensure the quality of film and television production, the plot video can be generated by the method provided in the embodiments of the present application, and then the film and television production can be completed. The AIGC can refer to the use of artificial intelligence technology to automatically generate text, images, videos, and other multimedia content. The plot video generated in the film and television production scene can be a film and television drama.
[0018] (2) Short video creation scene: In the short video creation scenario, many short video creators can be individual creators, and can even not be professional video producers. If traditional means are used to create videos, the video quality of the short video platform can be affected. If the video quality is limited, the threshold for short video creation can be very high, which is contrary to the original intention of short videos. Therefore, AIGC technology can be used to automatically generate a plot video. To further ensure the quality of short video creation, a plot video can be generated by the method provided in the embodiments of the present application, and then short video creation is completed. The plot video generated in the short video creation scenario can be a short video.
[0019] (Three) advertisement shooting scenario: In the advertisement shooting scenario, the advertisement shooting usually needs to be preceded by the conception of an advertisement idea. In order to complete the advertisement shooting, scriptwriting, directing, camera shooting, post-production and other posts need to be coordinated, which involves a large amount of manpower, material resources and time investment. From the conception of the idea to the output of the finished product, it often takes several weeks or even months, which is difficult to adapt to the rapidly changing market demand. At the same time, limited by the physical rules and shooting conditions of the real world, the space for creative expression is limited. Therefore, AIGC technology can be used to automatically generate a plot video (i.e., an advertisement). To further ensure the quality of the advertisement, a plot video can be generated by the method provided in the embodiments of the present application, and then the advertisement shooting is completed. The plot video generated in the advertisement shooting scenario can be an advertisement.
[0020] It should be noted that the method for generating a plot video provided in the embodiments of the present application can be executed by a computer device, which can be a terminal. The terminal can independently execute the method for generating a plot video provided in the embodiments of the present application. In some possible implementation manners, the method for generating a plot video provided in the embodiments of the present application also needs the support of related processing of the background, therefore, the computer device can also be a server, which can be a server providing a plot video generation service, and the server can execute the related processing of the background, such as the generation of a target split shot picture, the generation of a target split shot video, the generation of a target planning result, and the like. In this case, the terminal and the server, as two computer devices, can cooperate to execute the method for generating a plot video provided in the embodiments of the present application.
[0021] In order to facilitate understanding of the method for generating a plot video provided in the embodiments of the present application, the computer system used to implement the method for generating a plot video is first exemplarily introduced as follows.
[0022] Referring to Figure 1 , Figure 1 the architecture diagram of the computer system provided in the embodiments of the present application. Figure 1 The computer system 100 in the computer system 100 is a system architecture for implementing the method for generating a plot video in the embodiments of the present application. The computer system 100 includes a terminal 120 and a server 140.
[0023] The terminal 120 can install and run a plot video generation client, through which the plot video can be generated. The plot video generation client can include, but is not limited to, an application (App) in the terminal 120, a mini program, and the like, and can also be in the form of a webpage. The embodiments of the present application do not limit the form of the plot video generation client.
[0024] The terminal 120 can be a terminal used by a user. The plot video generation client running in the terminal 120 displays a split picture generation interface, which includes a first plot description of a current split. In response to a picture generation operation performed on the split picture generation interface, a target split picture of the generated current split is displayed. The first plot description is used to verify whether the target split picture meets the split intention. In response to a first confirmation operation on the target split picture, a split video generation interface is entered, which includes a second plot description of the current split and attribute information of a dubbing role. In response to a video generation operation performed on the split video generation interface, a target split video of the current split generated based on the target split picture is displayed. The attribute information is used to verify the matching degree of the current dubbing role in the target split video and the second plot description. The target split video is used to generate a plot video.
[0025] The terminal 120 can refer to the terminal described above. The terminal 120 can generally refer to one of a plurality of terminals. The embodiments of the present application only take the terminal 120 as an example. The device type of the terminal 120 includes, but is not limited to, at least one of a vehicle-mounted terminal, a smart phone, a tablet computer, a computer, a smart voice interaction device, a smart home appliance, a flying device, a wearable device, a personal computer (PC), a laptop computer, and a desktop computer. Those skilled in the art can know that the number of the terminal 120 described above can be more or less. For example, the terminal 120 described above can be only one, or the terminal 120 described above can be multiple or more. The number and device type of the terminal 120 are not limited in the embodiments of the present application.
[0026] The terminal 120 can be connected to the server 140 through a network to perform network communication. The network can be a wireless network or a wired network. The network uses standard communication technologies and / or protocols, and is usually the Internet, but can be any network, including but not limited to any combination of Bluetooth, a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), mobile, a private network, or a virtual private network. In some embodiments, custom or dedicated data communication technologies can be used instead of or in addition to the above data communication technologies.
[0027] The server 140 can be a stand-alone physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services, a cloud database, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and basic cloud computing services such as big data and artificial intelligence platforms.
[0028] For example, the server 140 is used to provide background services for the scenario video generation client running on the terminal 120. The server 140 includes a processor 144 and a memory 142, the memory 142 is used to store related computer programs, for example, computer programs related to the background services to be provided, and the processor 144 is used to execute the computer programs stored in the memory 142 to provide corresponding background services for the client, such as the generation of target split shot pictures, the generation of target split shot videos, the generation of target planning results, and the like, to provide for display on the terminal 120.
[0029] Optionally, the server 140 undertakes the main computing work, and the terminal 120 undertakes the secondary computing work; or the server 140 undertakes the secondary computing work, and the terminal 120 undertakes the main computing work; or the server 140 and the terminal 120 adopt a distributed computing architecture for collaborative computing.
[0030] The application scenario of the scenario video generation method provided by the embodiments of the present application will be introduced below in conjunction with the drawings. As shown in Figure 2 Figure 2 An application scenario architecture diagram of a scenario video generation method is shown. In this application scenario, it can include a terminal 21 and a server 22.
[0031] The terminal 21 can be a terminal used by a user, and a scenario video generation client can run on the terminal 21.
[0032] When a user opens the story video generation client on terminal 21 and wants to generate a story video through the story video generation client, they can enter the storyboard image generation interface, which will then be displayed on terminal 21.
[0033] The storyboard image generation interface can include a first plot description of the current storyboard. This first plot description serves as reference information for generating the storyboard image. After terminal 21 responds to the image generation operation performed on the storyboard image generation interface and displays the generated target storyboard image for the current storyboard, the user can verify whether the target storyboard image conforms to the storyboard intent based on the first plot description. This facilitates the user's verification and adjustment of the generated storyboard image's rationality, thereby obtaining a target storyboard image that meets the storyboard intent. The generation of the target storyboard image based on the image prompts of the current storyboard can be performed by server 22, which then sends the generated target storyboard image to terminal 21.
[0034] In one possible implementation, the storyboard image generation interface can also include a main reference image. This main reference image could be, for example, a character image. For instance, if the first scene description is "Close-up of boy A holding a fire-tipped spear, a provocative smile on his face, his sharp eyes staring ahead, with rolling waves behind him," then the character image could be an image of boy A. In this case, the storyboard image generation interface can refer to... Figure 2 As shown in 2111. In the storyboard image generation interface, the first plot description can be seen in 2111, the main reference image can be seen in 2112, and the generated target storyboard image can be seen in 2113.
[0035] If, after user verification or adjustment, the target storyboard image conforms to the storyboard intent, the user can perform a first confirmation operation on the target storyboard image. Terminal 21 can then respond to this first confirmation operation by entering the storyboard video generation interface. This interface includes a second plot description of the current storyboard and the attribute information of the voice-over character. The second plot description and the voice-over character's attribute information can serve as reference information for storyboard video generation. Thus, after terminal 21 responds to the video generation operation performed on the storyboard video generation interface and displays the target storyboard video generated based on the target storyboard image, the user can verify the matching degree between the current voice-over character and the second plot description in the generated storyboard video based on the attribute information. This facilitates the user's verification and adjustment of the generated storyboard video's rationality, avoiding common problems such as character misalignment, and obtaining a reasonable target storyboard video. Furthermore, a reasonable plot video can be generated based on the target storyboard video. The generation of the target storyboard video based on the video prompts and the target storyboard image can be performed by server 22, which then sends the generated target storyboard video to terminal 21.
[0036] Among them, the attribute information of the voice actor can be used to reflect the characteristics of the voice actor, such as the voice actor's identity information, emotional state, and dialogue description.
[0037] In one possible implementation, the target storyboard image can be displayed in the storyboard video generation interface, thereby further facilitating the verification of the generated storyboard video. In another possible implementation, if the storyboard video is relatively short, the second plot description can be similar to the first plot description; for example, the second plot description could be "Close-up of boy A holding a fire-tipped spear, a provocative smile on his face, his sharp eyes staring ahead, with rolling giant waves behind him." The storyboard video generation interface can be found in [reference needed]. Figure 2 As shown in 212, in the storyboard video generation interface, the second plot description can be seen in 2121, the attribute information of the voice actor can be seen in 2122, the identity information of the voice actor can be "Boy A", the dialogue description of the voice actor can be "Heh, is that all you've got?", the target storyboard image can be seen in 2123, the target storyboard video can be seen in 2124, the duration of the target storyboard video can be 4 seconds, and the generated target storyboard video is played in the storyboard video generation interface. The target storyboard video is currently playing at the 1-second mark.
[0038] In this embodiment, the generation process of storyboard images and storyboard videos is visualized, and the intermediate results, namely the target storyboard images and target storyboard videos, are transparent. Furthermore, reference information is provided to the user during the generation process of storyboard images and storyboard videos, which facilitates the user to intuitively and accurately verify and adjust the intermediate results based on the reference information, thereby ensuring the rationality of the intermediate results and improving the quality of the generated narrative video.
[0039] It should be noted that in the specific implementation of this application, the entire process may involve user information and other related data. When the above embodiments of this application are applied to specific products or technologies, separate consent or permission from the user is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0040] Next, with reference to the accompanying drawings and using a computer device as the terminal as an example, the method for generating narrative videos provided in this application will be described in detail. See also Figure 3 , Figure 3 A flowchart of a method for generating a narrative video is shown, the method including steps S301-S304, as detailed below: S301. Display the storyboard image generation interface, which includes the first plot description of the current storyboard.
[0041] The generation of narrative videos mainly includes the following six stages: scriptwriting, main design, storyboard directing and planning, storyboard image generation, storyboard video generation, and video editing. These six stages work together to automatically generate a complete narrative video from the user's creative inspiration. This application embodiment mainly describes the main implementation process of entering the storyboard image generation stage and the storyboard video generation stage.
[0042] Storyboarding is a visual script used in film and television production, typically presented in the form of a comic strip or cartoon. It transforms the written descriptions in the script into concrete visual language through a series of panels (each panel representing a shot) combined with textual descriptions (such as camera movement, duration, dialogue, sound effects, etc.). However, the script only provides a general description of the story and doesn't specify how to present it visually. For example, the phrase "two people arguing fiercely" might correspond to different approaches such as close-ups, over-the-shoulder shots, and push-in / pull-out camera movements. Therefore, to better translate textual descriptions into concrete visual language, storyboard videos can be created first, and then the actual film can be generated based on the storyboard videos, thus avoiding rework due to misunderstandings during filming.
[0043] In this embodiment, the number of scenes can be determined based on the desired length of the generated narrative video. The length of each scene can be within a certain range, such as approximately 4 seconds. If the length of the narrative video is approximately 30 seconds, it can be divided into 7 scenes, each with a length of approximately 4 seconds.
[0044] To better generate storyboard videos, the narrative descriptions of the storyboards can be used. The narrative descriptions transform the written narration in the script into concrete visual descriptions. They revolve around each shot, detailing the content, character actions, scene details, camera language, and technical parameters, providing precise guidance for storyboard creation and actual filming. Shot numbers can be sequentially labeled (e.g., "Shot 1," "Shot 2") for easy team positioning and collaboration; scene names clearly indicate the location of the shot (e.g., "living room," "street," "seaside"); time can describe the time of the shot, such as day / night / dusk, or a specific time (e.g., "6 AM"); character actions can describe the character's appearance, position, actions, and expressions, for example, character A (30 years old, wearing a black trench coat) stands by the window, looking down at his watch, his brow furrowed; camera language can reflect the different shot sizes, such as close-up, medium shot, long shot, overhead shot, and low-angle shot.
[0045] The current storyboard can be the storyboard for which a storyboard video needs to be generated. To generate the storyboard video for the current storyboard, you can first enter the storyboard image generation interface to generate the storyboard image for the current storyboard. The storyboard image generation interface can include the first plot description of the current storyboard. The first plot description can be a textual description of the script's narrative converted into a storyboard image. The first plot description can serve as reference information for generating the storyboard image, providing a basis for correcting or adjusting the generated storyboard image, thereby verifying whether the target storyboard image conforms to the storyboard's intent. The storyboard image generation interface can be found here. Figure 4 As shown, in Figure 4 In the middle, the first plot description can be seen in 401, which is described as "a close-up of boy A holding a fire-tipped spear, with a provocative smile on his face, staring sharply ahead, with rolling giant waves behind him."
[0046] Storyboard intent refers to the overall information the creator wants to convey to the audience through the storyboard, which may include aspects such as visual presentation, character development, and thematic expression. Verifying whether a target storyboard image conforms to the storyboard intent involves checking whether the target storyboard image includes the overall information the creator wants to convey to the audience. For example, the first plot description can reveal the visual effects the creator wants to convey to the audience. Users can check whether the target storyboard image has these visual effects. If the first plot description is "Close-up of boy A holding a fire-tipped spear, with a provocative smile on his face, staring sharply ahead, with rolling waves behind him," then the target storyboard image can be checked to see if it includes boy A holding a fire-tipped spear, having a provocative smile on his face, staring ahead, and having rolling waves behind him. For example, the first plot description can reveal the character traits that the creator wants to convey to the audience. Users can check whether the character in the target storyboard has these traits. If the first plot description describes boy A with a provocative smile and sharp eyes staring ahead, users can verify whether boy A in the target storyboard has a provocative smile and sharp eyes, and so on.
[0047] S302. In response to the image generation operation performed on the storyboard image generation interface, the target storyboard image of the generated current storyboard is displayed.
[0048] When storyboard images need to be generated, users can perform image generation operations on the storyboard image generation interface. The terminal responds to the image generation operation and displays the target storyboard image of the current storyboard.
[0049] In one possible implementation, the image generation operation can be any operation that triggers the generation of storyboard images, such as a click, double-click, or long-press. In some cases, the image generation operation may be, for example, clicking the "Image Generation" control on the storyboard image generation interface. See alsoFigure 4 As shown, the "Image Generation" control can be found in 402, and the generated target storyboard image can be found in 403.
[0050] In one possible implementation, in response to the image generation operation performed on the storyboard image generation interface, the way to display the target storyboard image of the current storyboard can be in response to the image generation operation, displaying the target storyboard image generated according to the image prompts of the current storyboard, that is, the target storyboard image is generated based on the image prompts.
[0051] A prompt is a set of symbols or text that the user needs to input. These prompts serve as instructions or descriptive text to guide artificial intelligence (AI) in generating content that meets expectations. In AIGC scenarios, a prompt typically refers to a series of statements or characters that the user inputs into the model to guide it in generating relevant content. Its core function is to translate human needs into language that AI can understand, controlling the style, theme, and details of the generated results through precise descriptions.
[0052] Image prompts can be instructions or descriptive text that guide AI to generate storyboard images. In this embodiment, if the first plot description is "Close-up of boy A holding a fire-tipped spear, with a provocative smile on his face, staring sharply ahead, with rolling giant waves behind him," in order to generate storyboard images that conform to the first plot description, the image prompts can be "General, anime style, close-up, boy A standing on the beach, holding a fire-tipped spear, with a provocative smile on his face, staring sharply ahead, with rolling giant waves behind him, during the day, the waves are brightly lit."
[0053] In one possible implementation, image cues can be displayed in the storyboard image generation interface, allowing for intuitive correction of whether the generated storyboard images follow the guidance of the image cues. For example, if the image cues are "General, Anime Style, Close-up, Boy A stands on the beach, holding a musket, a defiant smile on his face, his eyes sharply fixed ahead, with rolling waves behind him; it's daytime, and the waves are brightly lit," then the image cues could be found in [reference needed]. Figure 4 As shown in Figure 404.
[0054] In one possible implementation, the generation of target storyboard images can be achieved through an intelligent agent. An intelligent agent is a closed-loop AI system integrating perception, memory, reasoning, and action capabilities. It can perceive environmental changes through multimodal interfaces, autonomously plan task paths based on a built-in knowledge base and decision-making algorithms, invoke tools to execute operations, and continuously optimize strategies based on feedback to ultimately achieve the preset goal. The intelligent agent that generates storyboard images can be called a storyboard image generation intelligent agent.
[0055] It should be noted that there may be multiple storyboards obtained from the splitting process. In this embodiment, to improve the efficiency of storyboard image generation, storyboard images can be generated in batches. In this case, the storyboard image generation interface can be found here. Figure 5 As shown, this includes the generation of storyboard images for multiple storyboards, such as Storyboard 1, Storyboard 2, ... 501 shows the generation of storyboard images for Storyboard 1, including the first plot description of Storyboard 1 ("Close-up of boy A holding a fire-tipped spear, with a provocative smile on his face, staring sharply ahead, with rolling waves behind him"), the image prompt for Storyboard 1 ("General, anime style, close-up shot, boy A stands on the beach, holding a fire-tipped spear, with a provocative smile on his face, staring sharply ahead, with rolling waves behind him, daytime, bright light on the waves"), and the generated target storyboard image for Storyboard 1. 502 shows the generation of storyboard images for Storyboard 2, including the first plot description of Storyboard 2 ("Long shot of boy B wearing dragon scale armor, holding a Qiankun Ring, floating on the sea"), the image prompt for Storyboard 2 ("General, anime style, long shot, boy B floats on the sea, holding a Qiankun Ring, with a serious expression and wary eyes, daytime, soft light on the sea"), and the generated target storyboard image for Storyboard 1.
[0056] S303. In response to the first confirmation operation of the target storyboard image, enter the storyboard video generation interface, which includes the second plot description of the current storyboard and the attribute information of the voice actor.
[0057] Understandably, the first plot description serves as reference information for generating storyboard images, used to verify whether the target storyboard image conforms to the storyboard intent. Thus, after receiving the storyboard image on the terminal, the user can use the first plot description to verify whether the target storyboard image conforms to the storyboard intent, thereby facilitating the user's verification and adjustment of the generated storyboard image's rationality, ultimately obtaining a target storyboard image that meets the storyboard intent.
[0058] For example Figure 4 The target storyboard image shown in S303 indicates that the main character is a boy holding a fire-tipped spear, sporting a defiant smile, and staring sharply ahead. Behind him are rolling waves, which largely aligns with the storyboard intent described in the first plot. If any parts of the target storyboard image do not conform to the intended storyboard intent, they can be adjusted, for example, by re-executing S303 to regenerate the target storyboard image.
[0059] After user verification or adjustment, if the target storyboard image matches the storyboard intent, the user can perform a first confirmation operation on the target storyboard image. When the user performs the first confirmation operation, it indicates that the user approves the generated target storyboard image, and the storyboard video generation can continue. The terminal can then respond to the first confirmation operation and enter the storyboard video generation interface, which includes a second plot description of the current storyboard and the attribute information of the voice-over character.
[0060] The second plot description can be a textual description of the script's narrative converted into a storyboard video. The voice actor's attribute information can be used to represent the character's characteristics, such as their identity, emotional state, and dialogue. Both the second plot description and attribute information serve as reference information for generating the storyboard video. The second plot description primarily verifies whether the storyboard video conforms to the storyboard's intent, while the attribute information primarily verifies the match between the current voice actor in the generated storyboard video and the second plot description.
[0061] In one possible implementation, if the storyboard video is relatively short, the second plot description can be similar to the first plot description. If the storyboard image generation interface is as follows... Figure 4 As shown, the second plot description could be: "Close-up of boy A holding a fire-tipped spear, a defiant smile on his face, his sharp eyes fixed ahead, with rolling giant waves behind him," see [reference]. Figure 6 As shown in Figure 601. The attribute information of a voice-over character can include the character's identity information, dialogue descriptions, etc. For example, if the voice-over character's identity information is "Boy A," the dialogue description could be, "Heh, is that all you've got?" See [reference needed]. Figure 6 As shown in Figure 602.
[0062] The first confirmation operation can be any operation that confirms the target storyboard image conforms to the storyboard's intent, such as clicking, double-clicking, or long-pressing. For the storyboard image generation interface, please refer to [link / reference needed]. Figure 7 As shown in Figure 701, the storyboard image generation interface includes a "Confirm Import" control. The first confirmation operation can be clicking the "Confirm Import" control. The storyboard image generation interface described above is merely an example, and this application does not limit the scope of the storyboard image generation interface.
[0063] In one possible implementation, the generation of the target storyboard image can be achieved by an intelligent agent, and the intelligent agent that generates the storyboard video can be called the storyboard video generation intelligent agent.
[0064] S304. In response to the video generation operation performed on the storyboard video generation interface, display the target storyboard video of the current storyboard generated based on the target storyboard image.
[0065] When storyboard videos need to be generated, users can perform video generation operations on the storyboard video generation interface. The terminal responds to the video generation operation by displaying the target storyboard image for the current storyboard. It's understandable that the attribute information serves as reference information for storyboard video generation, used to verify the matching degree between the current voice-over character in the target storyboard video and the second plot description. In this way, after the terminal receives the storyboard video, users can verify the matching degree between the current voice-over character in the target storyboard video and the second plot description based on the attribute information, thereby avoiding common problems such as character mismatch.
[0066] The matching degree between the current voice-over character and the second plot description can refer to the degree of consistency between the current voice-over character and the description of the character in the second plot description. For example, whether the identity of the character is consistent with the setting of the current voice-over character and the setting of the second plot description, and whether the words spoken by the current voice-over character are consistent with the words spoken by the character in the setting of the second plot description, etc.
[0067] Assuming the second plot description identifies the character as Boy A, and Boy A says, "Ha, is that all you've got?", then the voice actor's identity information in the attribute information is "Boy A," and the dialogue description is "Ha, is that all you've got?". When the user plays the target storyboard video in the storyboard video generation interface, and the speaker is Boy A, saying "Ha, is that all you've got?", then it can be considered that the current voice actor in the target storyboard video matches the second plot description. However, if due to an error in generating the target storyboard video, the voice actor is identified as Boy B, resulting in the speaker in the generated target storyboard video being Boy B instead of Boy A, then it can be considered that the current voice actor in the target storyboard video does not match the second plot description. When verifying the match between the current voice actor in the target storyboard video and the second plot description, the target storyboard video can be used to generate a new plot video; otherwise, S304 can be re-executed to regenerate the target storyboard video.
[0068] One method for generating a narrative video from a target storyboard video is through video editing. For example, if multiple target storyboard videos are obtained, they can be edited together to create the narrative video. In one possible implementation, the video generation operation can be any operation that triggers the generation of the storyboard video, such as a click, double-click, or long-press. In some cases, the video generation operation may be, for example, clicking the "Video Generation" control on the storyboard video generation interface. See also... Figure 6 As shown, the "Video Generation" control can be found in Figure 603, and the generated target storyboard video can be found in Figure 604.
[0069] In one possible implementation, in response to a video generation operation performed on the storyboard video generation interface, the way to display the target storyboard video of the current storyboard generated based on the target storyboard image could be in response to the video generation operation, displaying the target storyboard video generated based on video prompts and the target storyboard image.
[0070] The video prompt can be an instruction or descriptive text that guides the AI to generate a storyboard video. In this embodiment, if the second plot description is "close-up of boy A holding a fire-tipped spear, with a provocative smile on his face, staring sharply ahead, with rolling giant waves behind him", in order to generate a storyboard image that matches the first plot description, the video prompt could be "boy A (short, neat hair) holding a fire-tipped spear, with a provocative smile on his face, staring sharply ahead, with rolling giant waves behind him, the camera slowly pulls away from boy A's entire body, showing his confrontation with the giant waves behind him."
[0071] In one possible implementation, video cues can be displayed in the storyboard video generation interface, facilitating intuitive correction of whether the generated storyboard video follows the guidance of the video cues. For example, if the video cues are "Boy A (short, neat hair) holds a flintlock pistol, a defiant smile on his face, his eyes sharply fixed ahead, with a giant wave rolling behind him; the camera slowly pulls back from Boy A's full body, showing his confrontation with the giant wave," then the video cues can be found in [reference needed]. Figure 8 As shown in Figure 801.
[0072] In one possible implementation, the generation of the target storyboard video can be achieved by an intelligent agent, which can be called a storyboard video generation intelligent agent.
[0073] It should be noted that there may be multiple storyboards obtained from the splitting process. In this embodiment, to improve the efficiency of storyboard video generation, storyboard videos can be generated in batches. In this case, the storyboard video generation interface can be found here. Figure 9 As shown, this includes the generation of storyboard videos for multiple storyboards, such as Storyboard 1, Storyboard 2, ... . 901 shows the generation of the storyboard video for Storyboard 1, including the second plot description of Storyboard 1 ("Close-up of boy A holding a fire-tipped spear, a provocative smile on his face, his sharp eyes staring ahead, with rolling giant waves behind him"), the attribute information of the voice actor for Storyboard 1, such as the voice actor's identity information being "boy A", the voice actor's dialogue description being "Heh, is that all you've got?", and the generated target storyboard video for Storyboard 1. 902 shows the generation of the storyboard video for Storyboard 2, including the attribute information of the voice actor for Storyboard 2, such as the voice actor's identity information and dialogue description. Since boy B may not speak in Storyboard 2, the voice actor's identity information and dialogue description here can be empty, and the generated target storyboard video for Storyboard 1.
[0074] It's important to note that storyboard images or videos may contain multiple subjects. Maintaining consistency across these subjects is crucial when generating the target storyboard image or video. Multi-subject consistency refers to ensuring the continuity and uniformity of multiple subjects across different frames during storyboard video generation. A subject can be an independent entity in an image or video, such as a character, object, scene, or style. To achieve multi-subject consistency, multi-subject driven generation technology can be used. This technology can simultaneously receive multi-dimensional inputs such as characters, scenes, and prompts, generating an image generation model that satisfies subject consistency (e.g., character appearance, scene style). This can also be called a multi-subject consistency model, ensuring consistency in character appearance, scene style, and narrative logic throughout the entire process. The pass rate of the generated results is significantly improved, for example, to over 95%, ensuring the rationality and consistency of the results from the generation source.
[0075] It's understandable that target storyboard images can be used to generate target storyboard videos. Therefore, target storyboard images can also serve as a basis for verifying the reasonableness of target storyboard videos. Based on this, in one possible implementation, in response to the first confirmation operation of the target storyboard image, the way to enter the storyboard video generation interface could be that, in response to the first confirmation operation, the target storyboard image is imported into the image display area of the storyboard video generation interface. That is to say, the storyboard video generation interface not only needs to display the target storyboard image, but the target storyboard image displayed in the storyboard video generation stage can also be directly imported from the storyboard image generation stage. In other words, there can be a data interaction interface between the storyboard image generation stage and the storyboard video generation stage. The target storyboard image generated in the previous stage, such as the storyboard image generation stage, is automatically imported into the storyboard video generation stage for use, without requiring manual transfer by the user.
[0076] by Figure 7 Taking the generated target storyboard image as an example, the target storyboard image can be displayed in the storyboard video generation interface. See [link / reference]. Figure 10 As shown in Figure 1001, this allows users to visually verify the generated storyboard video based on the target storyboard image.
[0077] In this embodiment of the application, when generating a storyboard video, the target storyboard image on which the target storyboard video depends is displayed, allowing for a direct comparison between the target storyboard image and the target storyboard video, thus facilitating user verification of the target storyboard video. Furthermore, the target storyboard image displayed in the storyboard video generation interface is automatically imported from the storyboard image generation stage, eliminating the need for manual transfer by the user and preventing information loss or inaccuracies.
[0078] As can be seen from the above technical solution, the generation process of the narrative video provided in this application is visualized, and the intermediate results are transparent, allowing users to verify and adjust them. Specifically, when a narrative video needs to be generated, a storyboard image generation interface can be displayed, which includes the first plot description of the current storyboard. The first plot description can serve as reference information for storyboard image generation. Thus, in response to the image generation operation performed on the storyboard image generation interface, after displaying the target storyboard image of the current storyboard, the user can verify whether the target storyboard image conforms to the storyboard intent based on the first plot description. This facilitates the user's verification and adjustment of the rationality of the generated storyboard image, thereby obtaining a target storyboard image that conforms to the storyboard intent. If a first confirmation operation is performed on the target storyboard image, the storyboard video generation interface is entered in response to the first confirmation operation. The storyboard video generation interface includes the second plot description of the current storyboard and the attribute information of the voice actor. The second plot description and the attribute information of the voice-over characters can serve as reference information for generating storyboard videos. In response to the video generation operation performed on the storyboard video generation interface, after displaying the target storyboard video generated based on the target storyboard image, users can verify the matching degree between the current voice-over character and the second plot description in the generated storyboard video based on the attribute information. This facilitates users in verifying and adjusting the rationality of the generated storyboard video, avoiding common problems such as character misalignment, and obtaining a reasonable target storyboard video. Furthermore, a reasonable plot video can be generated based on the target storyboard video. Therefore, the generation process of storyboard images and storyboard videos in this application is visualized; the intermediate results, namely the target storyboard image and target storyboard video, are transparent. Reference information is provided to users during the generation process, allowing them to intuitively and accurately verify and adjust the intermediate results based on this information, ensuring the rationality of the intermediate results and thus improving the quality of the generated plot video.
[0079] It should be noted that there are multiple ways to generate target storyboard images for display. One way is to directly generate target storyboard images based on image prompts for display.
[0080] In another approach, since the generated storyboard images may include at least one subject with its own appearance and state, a subject reference image can be introduced during the storyboard image generation process to accurately and efficiently generate the target storyboard image. This subject reference image can then be displayed in the storyboard image generation interface. Thus, in response to the image generation operation performed in the storyboard image generation interface, the terminal can display the target storyboard image generated based on the subject reference image.
[0081] Taking a character as the main subject, if the character is boy A, a reference image of boy A can be obtained and displayed in the storyboard image generation interface. This reference image can then be used to generate the target storyboard image, ensuring that the character in the target storyboard image is derived from the reference image. The storyboard image generation interface can be found here. Figure 11 As shown, the main reference image can be found here. Figure 11 As shown in 1101.
[0082] This application embodiment uses a main reference image as the basis for generating the target storyboard image, thereby ensuring accurate and efficient generation of the target storyboard image. Simultaneously, the main reference image is displayed in the storyboard image generation interface, allowing users to intuitively compare the target storyboard image and the main reference image, thus facilitating user verification of the target storyboard image.
[0083] It should be noted that the sources of the main reference images are not limited in the embodiments of this application. In some cases, the process of generating a narrative video may include not only the storyboard image generation stage and the storyboard video generation stage, but also other stages. For example, a storyboard directing and planning stage may be included before the storyboard image generation stage. Storyboard directing and planning is a professional process that transforms a story script into a specific visual presentation plan, including storyboard breakdown and storyboard presentation design. The essence of storyboard directing and planning is to visually rehearse the story script, obtain the target planning result through storyboard directing and planning, and then transform abstract ideas into executable visual instructions so that the effect of the final generated narrative video can be predicted.
[0084] In this scenario, the main reference image can originate from the target planning result obtained during the storyboard director's planning phase. One possible implementation involves displaying the storyboard image generation interface as a storyboard director's planning interface, which includes the target planning result, comprising the main design result and a first plot description. If the user approves the target planning result, they can perform a second confirmation operation on the storyboard director's planning interface. The terminal responds to this second confirmation operation by entering the storyboard image generation interface. During this process, the first plot description is imported into the plot display area of the storyboard image generation interface, and the main reference image corresponding to the main design result is imported into the image display area of the storyboard image generation interface. The plot display area can refer to the area within the storyboard image generation interface that displays the plot description, and the image display area can refer to the area within the storyboard image generation interface that displays the main reference image.
[0085] It should be noted that the main design results displayed in the storyboard director planning interface may include at least one of the following: main name, main description, and main reference image. The main design results obtained during the storyboard director planning stage may include the main reference image; however, the main reference image may or may not be displayed in the storyboard director planning interface. This application embodiment does not limit this; therefore, the main reference image corresponding to the main design result here may refer to the main reference image generated together with the main design result during the storyboard director planning stage.
[0086] The second confirmation action can be any action performed by the user to confirm the target planning result, such as clicking, double-clicking, or long-pressing. If the storyboard director's planning interface includes a "Confirm Planning Result" control, the second confirmation action can be clicking the "Confirm Planning Result" control.
[0087] The storyboard director's planning interface can be found here. Figure 12 As shown in Figure 1201, the main design results can include the main character name (e.g., boy A, boy B, boy C) and the main character description (e.g., neat short hair). The first plot description can be seen in Figure 1202. Figure 12 The specific details of the first plot description are not shown. The "Confirm Planning Results" control can be found in 1203.
[0088] In one possible implementation, the generation of the target planning result can be achieved by an intelligent agent, which can be called the storyboard director planning intelligent agent.
[0089] In this embodiment of the application, the main reference images used in the storyboard image generation stage are automatically imported from the storyboard director planning stage, without the need for manual transfer by the user, thereby avoiding information loss or deviation.
[0090] In one possible implementation, the main design result can be generated during the main design phase. Therefore, the storyboard director's planning interface can be displayed as a main design interface, which includes the main design result. If the user approves the main design result, they can perform a third confirmation operation on the main design interface. The terminal responds to this third confirmation operation by entering the storyboard director's planning interface. During this process, the main design result is imported into the main display area of the storyboard director's planning interface. The main display area can be the area within the storyboard director's planning interface that displays the main design result.
[0091] It should be noted that in the embodiments of this application, there are many categories of subjects, such as characters, objects, scenes, styles, etc. When displaying the subject design results, the subject's number may be displayed. In order to distinguish different categories, in one possible implementation, different categories may be numbered separately. For example, if it is necessary to generate subject design results for two characters and subject design results for two scenes, then the characters can be numbered as Character 1 and Character 2, and the scenes can be numbered as Scene 1 and Scene 2.
[0092] In one possible implementation, the main design interface can also display a main description, which describes the relevant characteristics of the main body, thereby ensuring a more accurate main design result. The third confirmation operation can be any operation to confirm the main design result, such as a click, double-click, or long-press. If the main design interface includes a "Confirm Main Design Result" control, then the third confirmation operation can be the user clicking the "Confirm Main Design Result" control.
[0093] Based on the above introduction, in one possible implementation, the main design interface can be found here. Figure 13 As shown, the main design results can be found in [reference needed]. Figure 13 As shown in 1301 and 1302, 1301 is the main body name, and 1302 is the main body reference image. The control for "Generate Main Body Reference Image" can be found in [reference needed]. Figure 13 As shown in section 1303, the control for "Confirming the Main Design Result" can be found in [reference needed]. Figure 13 As shown in page 1304, the main description is in Figure 13 It is not shown in the image. It should be noted that... Figure 13 This is merely an example of the main design interface and does not constitute a limitation on the main design interface.
[0094] In one possible implementation, the subject design can be implemented by an intelligent agent, which can be called the subject design intelligent agent.
[0095] In this embodiment, the main design results used in the storyboard director planning stage are automatically imported from the main design stage, without requiring manual transfer by the user, thereby avoiding information loss or deviation.
[0096] It should be noted that the main design results may also include main reference images, which are obtained during the main design phase and displayed in the storyboard director's planning interface. The main reference images obtained during the main design phase can be generated in real-time, specifically based on the main description; alternatively, they can be external images uploaded by the user, specifically obtained in response to an image upload operation.
[0097] To facilitate the real-time generation of the main reference image, a design result generation operation can be performed within the main design interface, triggering the step of generating the main design result. This design result generation operation can be any action that triggers it, such as a click, double-click, or long-press. If the main design interface includes a "Generate Main Reference Image" control, then the design result generation operation can be initiated by the user clicking this control. See also... Figure 13 The control for "generating a reference image for the main subject" can be found in [link / reference]. Figure 13 As shown in Figure 1304.
[0098] Image upload can be any operation by the user to upload a reference image for the main body. In one possible implementation, the main design interface can also include an image upload control. When the user clicks the image upload control, they can select an image and upload it as the reference image for the main body.
[0099] This application embodiment allows users to obtain subject reference images during the subject design stage, including real-time generation or uploading of external images, thereby overcoming the limitations of relying on prompt words for generation in related technologies. By combining user-provided visual materials, it significantly improves the accuracy and diversity of generated subjects.
[0100] In one possible implementation, the target planning result also includes image prompts for the current storyboard. In this case, the terminal can also import the image prompts into the prompt display area of the storyboard image generation interface during the process of entering the storyboard image generation interface. Then, in response to the image generation operation performed in the storyboard image generation interface, the way to display the generated target storyboard image for the current storyboard can be to display the target storyboard image generated according to the image prompts, in response to the image generation operation. That is, when the storyboard director planning is completed, the image prompts generated in the storyboard director planning stage are directly transmitted to the storyboard image generation stage through a data interaction interface and displayed in the storyboard image generation interface. The prompt display area of the storyboard image generation interface can be the area in the storyboard image generation interface used to display image prompts. The displayed image prompts can be found in [reference needed]. Figure 4 As shown in Figure 404.
[0101] In one possible implementation, the image prompts can be derived from a professional shot language rule library built into the storyboard director planning AI, enabling non-professional users to generate storyboard schemes that conform to film and television creation standards and improve the professionalism of the content.
[0102] This application embodiment displays image prompts in the storyboard image generation interface, facilitating intuitive correction of whether the generated storyboard images follow the guidance of the image prompts. Furthermore, the image prompts are automatically imported from the storyboard director's planning stage, eliminating the need for manual input from the user and thus preventing information loss or deviation.
[0103] In one possible implementation, the target planning result also includes video prompts for the current storyboard. In this case, in response to the first confirmation operation of the target storyboard image, the way to enter the storyboard video generation interface could be that, in response to the first confirmation operation, the video prompts are imported into the prompt display area of the storyboard video generation interface. Then, in response to the video generation operation performed in the storyboard video generation interface, the way to display the target storyboard video generated based on the target storyboard image could be that, in response to the video generation operation, the target storyboard video generated based on the video prompts and the target storyboard image is displayed. That is, when the storyboard director planning is completed, the video prompts generated in the storyboard director planning stage are directly transmitted to the storyboard video generation stage through a data interaction interface and displayed in the storyboard video generation interface. The prompt display area of the storyboard video generation interface can be the area in the storyboard video generation interface used to display video prompts. The displayed video prompts can be found in [reference needed]. Figure 8 As shown in Figure 801.
[0104] In one possible implementation, video prompts can be derived from a professional shot language rule library built into the storyboard director planning AI, enabling non-professional users to generate storyboard schemes that conform to film and television creation standards and improve the professionalism of the content.
[0105] This application embodiment displays video prompts in the storyboard video generation interface, facilitating intuitive correction of whether the generated storyboard video follows the guidance of the video prompts. Furthermore, the video prompts are automatically imported from the storyboard director's planning stage, eliminating the need for manual input from the user and thus preventing information loss or deviation.
[0106] To improve the usability of the generated storyboard videos, this application optimizes the intermediate results generated at key stages of storyboard video generation. It is understood that during storyboard directing planning, if limited by templated prompts, some abstract planning rules may be imposed, leading to insufficient instantiation and various problems in the resulting plans. For example, the storyboard breakdown in the prompts may be unreasonable, with some scenes being excessively long; the prompts obtained from the storyboard planning may be unreasonable, with character restrictions in the prompts not conforming to established rules; and the selected characters in the storyboards may exceed the given range. Therefore, to avoid these problems, the intermediate results generated during the storyboard directing planning stage can be self-optimized to obtain high-quality target planning results. Self-optimization can be achieved by using an evaluator to analyze the intermediate results and automatically optimize the adaptive learning ability of the generation strategy.
[0107] Based on this, in one possible implementation, the target planning result can be generated by generating an initial planning result based on the initial planning prompts, script creation results, and main design results; evaluating the initial planning result using the director's planning evaluation criteria to obtain a first evaluation result; optimizing the initial planning prompts based on the first evaluation result to obtain optimized planning prompts; and generating the target planning result based on the optimized planning prompts, script creation results, and main design results, thereby achieving self-optimization.
[0108] The self-optimization process mainly involves optimizing the prompt words. Prompt word optimization refers to the process of iteratively improving the input prompt words (which can refer to planning prompt words) through an evaluation and feedback mechanism. Based on the optimized planning prompt words, an optimized planning result is generated until the target planning result is obtained.
[0109] The scriptwriting results may include research reports on the script story obtained during the scriptwriting process, the video style selected by the user, and the creative presentation techniques.
[0110] It should be noted that in the embodiments of this application, the initial planning result may undergo multiple iterations of optimization to reach the target planning result. Each iteration of optimization uses the director's planning evaluation criteria to evaluate the planning result obtained in the previous iteration to obtain a first evaluation result. Then, based on the first evaluation result, the planning prompts obtained in the previous iteration are optimized to obtain optimized planning prompts. Based on the optimized planning prompts, script creation results, and main design results, the optimized planning result for this iteration is generated until the iteration optimization is completed and the target planning result is obtained.
[0111] For any iteration of optimization, assuming that any iteration of optimization is represented as the i-th iteration, where i = 0, 1, 2, ..., the iteration optimization formula can be as follows:
[0112] in, For the script story, As the main design result, The video style selected by the user A research report on creative presentation techniques. For the first Planning prompts before each iteration of optimization For the first The optimized planning prompts after multiple iterations The AI agent is designed for the storyboard director. For the first The planning results before each iteration optimization For the first The first evaluation result of the planning results of the round of iterative optimization. For the evaluation function of the planning results, Evaluation criteria for the director Optimize the prompt word model for the AI agent designed by the storyboard director. For the first The planning results after rounds of iterative optimization.
[0113] When i=0 These could be initial planning prompts. It could be the initial planning result, after iterative optimization. It can be used as the result of target planning.
[0114] In one possible implementation, the outcome of goal planning can be defined as:
[0115] in, Indicates the result of the target planning. The number of scenes obtained from the division, For the first The plot description of each storyboard. For the first The narration dialogue in each scene, For the first The voice-over character for each storyboard. For the first The main elements of a storyboard selection can include characters and scenes. For the first Image prompts generated from each storyboard scene. for Video prompts generated from each storyboard.
[0116] In the storyboard planning stage of this application, the intermediate results generated are professionally evaluated and self-optimized, so as to use the optimized planning results as the target planning results, thereby controlling the rationality of the target planning results, avoiding various problems in the generated target planning results, and improving the rationality of the target planning results.
[0117] In other cases, to avoid the goal planning results obtained in the storyboard directing stage failing to meet user needs, embodiments of this application also provide another source of the main reference image. For example, external images can be inserted into the storyboard image generation interface. In this case, if the user wishes to insert an external image as the main reference image, they can also perform an image insertion operation in the storyboard image generation interface. The terminal responds to the image insertion operation performed in the storyboard image generation interface by displaying an image browsing interface; if a selection operation is performed on the main reference image in the image browsing interface, the main reference image is displayed in the storyboard image generation interface.
[0118] The image insertion operation can be any operation that inserts a main reference image into the storyboard image generation interface, such as a click, double-click, long-press, etc. If the storyboard image generation interface includes a "Insert Main Reference Image" control, then the image insertion operation can be the user clicking the "Insert Main Reference Image" control.
[0119] See Figure 11 As shown, the "Insert Main Reference Image" control can be found in [reference]. Figure 11 As shown in Figure 1102. After the user clicks the "Insert Main Reference Image" control, the image browsing interface can be displayed. Figure 11 Based on this, the interface diagram becomes Figure 14 The interface shown here is the image browsing interface, which can be found in [reference needed]. Figure 14 As shown in Figure 1401, if a user selects the first image in the image browsing interface, that image can be displayed as the main reference image. The selection operation can be any operation confirming the insertion of the main reference image. Figure 14 In the image browsing interface, the selected action can be that the user clicks on the first image and then clicks the "Confirm" control shown in 1402. Figure 14 This is merely an example of a change in the interface and an image browsing interface, and does not constitute a limitation on the change in the interface or the image browsing interface.
[0120] This application embodiment inserts external images into the storyboard image generation interface to obtain main reference images, allowing users to customize them. This ensures that the main reference images used can better meet user needs, improves the accuracy and diversity of the main reference images, and thus improves the accuracy and diversity of target storyboard image generation.
[0121] Understandably, in order to avoid generating unreasonable target storyboard videos that would affect the quality of the narrative video, this application provides interactive editing capabilities so that users can intervene in a timely manner if they find inaccuracies or unreasonable aspects during the generation of the target storyboard video, and ensure the rationality of the final generated target storyboard video through editing.
[0122] In this embodiment, generating the target storyboard video mainly involves a storyboard image generation stage and a storyboard video generation stage. In these two stages, the quality of the target storyboard video is primarily affected by image prompts, the target storyboard image, and video prompts. The target storyboard image and video prompts are used to generate the target storyboard video; therefore, their quality affects the overall quality of the target storyboard video. Conversely, image prompts are used to generate the target storyboard image; therefore, image prompts also affect the quality of the target storyboard video. Therefore, in this embodiment, the provided interactive editing capabilities mainly target image prompts, target storyboard images, and video prompts. The three editing methods are described below.
[0123] Understandably, if the target storyboard image is generated based on image cues, the storyboard image generation interface can also include image cues. To avoid inaccurate image cues that could hinder AI from generating a reasonable target storyboard image, this application provides an interactive editing interface for image cues. When a user wants to modify an image cue (e.g., finding it unreasonable), the user can perform a modification operation. The terminal responds to the modification operation and obtains the modified image cues. In this case, in response to the image generation operation performed on the storyboard image generation interface, the way to display the target storyboard image generated for the current storyboard can be to display the target storyboard image generated according to the modified image cues.
[0124] For example Figure 4 As shown in error 404, the image tooltip can be edited by the user by first performing an edit trigger operation. While in editable mode, the user can then modify the tooltip. The edit trigger operation can be any action that puts the tooltip into editable mode, such as a long press, a double tap, or clicking on the editing controls provided in the tooltip display area.
[0125] In this embodiment, users can freely modify the image prompts, supporting fine-tuning of local content and avoiding the inefficiency caused by global modifications in related technologies.
[0126] In some cases, if the generated storyboard image is unreasonable, besides modifying the image prompts to regenerate the target storyboard image, one can also directly modify the generated storyboard image to obtain a reasonable target storyboard image. In this situation, in response to the image generation operation performed on the storyboard image generation interface, the way to display the target storyboard image for the current storyboard can be either in response to the image generation operation, generating the initial storyboard image for the current storyboard; or in response to the input operation for the target storyboard image, displaying the target storyboard image on the storyboard image generation interface, where the target storyboard image is obtained by adjusting the initial storyboard image.
[0127] In other words, after generating storyboard images (such as initial storyboard images) based on image prompts, users can verify the generated storyboard images. If the storyboard images are not reasonable, users can use image processing software to adjust the generated storyboard images to obtain satisfactory storyboard images as target storyboard images. Then, input operations are performed on the target storyboard images, and the target storyboard images are displayed on the storyboard image generation interface.
[0128] The embodiments of this application do not limit the input operation; for example, it can be an upload operation or a drag operation.
[0129] The process of generating target storyboard images can be found in [reference needed]. Figure 15 As shown, after determining that adjustments to the initial storyboard images are needed, the storyboard image generation interface can be found here. Figure 15 As shown in 1501, the area indicated in 1502 of the storyboard image generation interface is used to display the target storyboard image. Here, the user is prompted to "drag and drop the image here or click upload" to obtain the adjusted target storyboard image for display. If the input operation is a drag-and-drop operation, an example image of the target storyboard image can be found here. Figure 15 As shown in 1503, the adjusted target storyboard image is saved in the location marked (1), such as the computer desktop. The user can drag the target storyboard image with the mouse to the location shown in (2). After the user releases the mouse, the storyboard image generation interface shown in 1504 is entered, and the target storyboard image is displayed.
[0130] The process of generating target storyboard images can be found in [reference needed]. Figure 16 As shown, after determining that adjustments to the initial storyboard images are needed, the storyboard image generation interface can be found here. Figure 16As shown in 1601, the area indicated in 1602 of the storyboard image generation interface is used to display the target storyboard image. Here, the user is prompted to "drag and drop the image here or click upload" to obtain the adjusted target storyboard image for display. If the input operation is an upload operation, the user can click in the area indicated in 1602 to enter... Figure 16 The interface shown in 1603. The image browsing interface shown in 1604 allows users to select a target storyboard image. For example, the user could tap the target storyboard image named "1.jpg" and then click the "Confirm" control shown in 1605, thus entering the storyboard image generation interface shown in 1606, which displays the target storyboard image.
[0131] In this embodiment, the user can upload externally generated results, that is, adjust the generated storyboard images outside the story video generation client, such as in image processing software, to obtain the adjusted storyboard images, and then input the adjusted storyboard images as the target storyboard images, thereby realizing fine-tuning of automatically generated storyboard images and avoiding the inefficiency caused by global modifications in related technologies.
[0132] It is understandable that if the target storyboard video is generated based on video cues and target storyboard images, the storyboard video generation interface may also include video cues. To avoid inaccurate video cues that could hinder AI from generating a reasonable target storyboard video, this application embodiment provides an interactive editing interface for video cues. When a user wants to modify the video cues (e.g., finding inappropriate parts), the user can perform a modification operation on the video cues. The terminal responds to the modification operation and obtains the modified video cues. In this case, in response to the video generation operation performed on the storyboard video generation interface, the way to display the target storyboard video generated based on the target storyboard image for the current storyboard can be to respond to the video generation operation by displaying the target storyboard video generated based on the modified video cues and target storyboard images.
[0133] For example Figure 8 As shown in section 801, the video prompt can be edited by the user by first performing an edit trigger operation. While the prompt is in editable mode, the user can then modify it. The edit trigger operation can be any action that puts the prompt into editable mode, such as a long press, a double tap, or clicking on the editing controls provided in the prompt's display area.
[0134] In this embodiment, users can freely modify the video prompts, supporting fine-tuning of local content and avoiding the inefficiency caused by global modifications in related technologies.
[0135] This application's embodiments can optimize intermediate results generated at crucial stages of generating narrative videos. The storyboard image generation stage is also a critical step in generating narrative videos; therefore, to avoid unstable results from the generated target storyboard images, self-optimization can be performed at this stage, achieving self-review and self-iterative optimization of the generated results.
[0136] Based on this, in one possible implementation, the image prompt is an initial image prompt. Responding to the image generation operation performed on the storyboard image generation interface, the method for displaying the target storyboard image generated according to the current storyboard's image prompt can be as follows: In response to the image generation operation, generate the initial storyboard image for the current storyboard according to the initial image prompt; evaluate the initial storyboard image using basic quality assessment criteria and instructions following the assessment criteria to obtain a second assessment result; optimize the initial image prompt based on the second assessment result to obtain an optimized image prompt; generate the target storyboard image according to the optimized image prompt and display the target storyboard image, thereby achieving self-optimization.
[0137] The self-optimization process mainly involves optimizing the cue words. Cue word optimization can refer to the process of iteratively improving the input cue words (which can refer to image cue words) through an evaluation feedback mechanism. Based on the optimized image cue words, optimized storyboard images are generated until the target storyboard image is obtained.
[0138] It should be noted that, in the embodiments of this application, the initial storyboard image may undergo multiple iterations of optimization to reach the target storyboard image. Each iteration of optimization uses the basic quality assessment criteria and the instruction to follow the assessment criteria to evaluate the storyboard image obtained from the previous iteration of optimization to obtain a second assessment result. Then, based on the second assessment result, the image prompts obtained from the previous iteration of optimization are optimized to obtain optimized image prompts. Based on the optimized image prompts, the storyboard image optimized for the current iteration of optimization is generated until the iteration optimization is completed and the target storyboard image is obtained.
[0139] For any iteration of optimization, assuming that any iteration of optimization is represented as the i-th iteration, where i = 0, 1, 2, ..., the iteration optimization formula can be as follows:
[0140] in, Generate intelligent agents for storyboard images. For the first Image prompts for each scene before the i-th iteration of optimization. For the first The image shows the storyboard image before the i-th iteration of optimization. The basic quality assessment criteria include some physical criteria for visual composition and assessment standards for collapse. The instruction must adhere to the evaluation criteria, namely whether the generated storyboard images match the description of the prompt words. This is an evaluation model for storyboard images. For the first The second evaluation result of the segment in the i-th iteration optimization. An optimized model for image-based prompts. For the first The first scene in the... Storyboard images after rounds of iteration and optimization.
[0141] First, the storyboard image generation AI is based on automatically selected character elements. and image prompts Multi-subject-driven, high-fidelity storyboard image generation Then evaluate the model The generated storyboard images will be evaluated under two assessment criteria: basic quality and instruction compliance. The basic quality criterion primarily assesses whether the composition meets basic physical rules, such as whether character animation is flawed or whether character composition violates physical rules. The instruction compliance criterion primarily evaluates whether the lighting, composition, and character positions and poses of the generated images conform to the descriptions in the image prompts. Next, an optimization model for the image prompts will be implemented. Based on the results of the second evaluation and some unreasonable aspects of the generated storyboard images, the image cues will be optimized and fine-tuned to obtain the optimized image cues. Finally, the storyboard image generation agent obtained the first... The optimized storyboard images.
[0142] When i=0 It could be an initial image prompt. It could be the initial storyboard image, which was used after iterative optimization. It can be used as a target storyboard image.
[0143] In the storyboard image generation stage, this application embodiment performs professional evaluation and self-optimization on the intermediate results generated by individual storyboards at a finer granular level. The optimized storyboard image is then used as the target storyboard image to control the rationality of the target storyboard image more finely, avoid various problems in the generated target storyboard image, and improve the rationality of the target storyboard image.
[0144] The storyboard video generation stage is also a crucial step in generating narrative videos. Therefore, to avoid instability in the generated target storyboard video, self-optimization can be performed during the storyboard video generation stage, enabling self-review and self-iterative optimization of the generated results.
[0145] Based on this, in one possible implementation, the video prompt is an initial video prompt. In response to the video generation operation performed on the storyboard video generation interface, the method for displaying the target storyboard video generated based on the current storyboard video prompt and the target storyboard image can be as follows: In response to the video generation operation, based on the initial video prompt and the target storyboard image, an initial storyboard video for the current storyboard is generated; the initial storyboard video is evaluated using basic quality assessment criteria and instruction-following assessment criteria to obtain a third assessment result; the initial video prompt is optimized based on the third assessment result to obtain an optimized video prompt; the target storyboard video is generated based on the optimized video prompt and the target storyboard image, and then displayed.
[0146] The self-optimization process mainly involves optimizing the prompt words. Prompt word optimization can refer to the process of iteratively improving the input prompt words (which can refer to video prompt words) through an evaluation feedback mechanism. Based on the optimized video prompt words, an optimized storyboard video is generated until the target storyboard video is obtained.
[0147] It should be noted that, in the embodiments of this application, the initial storyboard video to the target storyboard video may have undergone multiple iterations of optimization. Each iteration of optimization uses the basic quality assessment criteria and the instruction to follow the assessment criteria to evaluate the storyboard video obtained from the previous iteration of optimization to obtain a third assessment result. Then, based on the third assessment result, the video prompts obtained from the previous iteration of optimization are optimized to obtain optimized video prompts. Based on the optimized video prompts, the storyboard video optimized for the current iteration of optimization is generated until the iteration optimization is completed and the target storyboard video is obtained.
[0148] For any iteration of optimization, assuming that any iteration of optimization is represented as the i-th iteration, where i = 0, 1, 2, ..., the iteration optimization formula can be as follows:
[0149] in, For the target storyboard image, For the first The video prompts for each scene before the i-th iteration of optimization. Generate intelligent agents for storyboard videos. For the first The storyboard video before the i-th iteration optimization. An evaluation model for storyboard videos. and These are the basic quality assessment criteria and the instruction compliance assessment criteria, respectively. For the first The third evaluation result of the segment in the i-th iteration optimization. An optimized model for video prompts. For the first The video prompts for each scene after the i-th iteration optimization. For the first The first scene in the... The storyboard video obtained after rounds of iterative optimization.
[0150] When i=0 It could be the initial video prompt. It could be the initial storyboard video, after iterative optimization. It can be used as a target storyboard video.
[0151] In the storyboard video generation stage, this application embodiment performs professional evaluation and self-optimization on the intermediate results generated by individual storyboards at a finer granular level. The optimized storyboard video is then used as the target storyboard video to control the rationality of the target storyboard video at a finer granular level, avoid various problems in the generated target storyboard video, and improve the rationality and quality of the target storyboard video.
[0152] It should be noted that in some cases, the input received during storyboard video generation includes the video prompts generated by the driver and the target storyboard image. Due to the overall workflow design, the target storyboard image may not yet be available when planning the video prompts in the preceding storyboard director planning stage. Therefore, there may be misalignment issues between the video prompts and the target storyboard image. For example, the video prompts may include subjects that do not exist in the target storyboard image, or the subjects in the video prompts may refer to subjects in locations that do not match the actual target storyboard image. Therefore, this application embodiment first introduces an information alignment mechanism in the storyboard video generation agent. By aligning the storyboard descriptions in long-term memory, the target storyboard image, and the video prompts in short-term memory, information consistency is achieved during the video generation process, ensuring the generated storyboard video.
[0153] Based on this, in one possible implementation, in response to the video generation operation, the method for generating the initial storyboard video of the current storyboard based on the initial video cues and the target storyboard image could be to align the information based on the initial video cues and the target storyboard image to obtain aligned video cues; and then generate the initial storyboard video based on the aligned video cues and the target storyboard image. Correspondingly, the method for optimizing the initial video cues according to the third evaluation result to obtain optimized video cues could be to optimize the aligned video cues according to the third evaluation result to obtain optimized video cues.
[0154] In one possible implementation, the information alignment can be done as follows:
[0155] in, Some guidelines for information alignment, An alignment model is generated to provide the necessary information for the storyboard video. For the first The video cue words for each shot before information alignment (e.g., the initial video cue words). For the first Each scene has its own video cues after alignment.
[0156] If information alignment is performed, then during the self-optimization process in the storyboard video generation stage, when i=0, the used... These are the aligned video prompts.
[0157] It's important to note that short-term memory can be considered the agent's "workbench" for processing the current task or a single conversation, directly relying on the context window of the large language model. All information within the current context window, such as system commands, the history of the current conversation, and the direct results of tool calls, constitute its short-term memory. Long-term memory, on the other hand, can be seen as the agent's "personal knowledge base," used to persistently store key information across conversations, such as user preferences, important facts, or learned experiences. Its implementation does not depend on the model's context window but is accomplished through an external storage system. When needed, the agent uses techniques such as Retrieval-Augmented Generation (RAG) to retrieve the most relevant information fragments from the long-term memory on demand and dynamically inject them into the current short-term memory (context window). This allows the agent to respond based on historical knowledge, achieving truly personalized service and continuous learning.
[0158] To achieve both short-term and long-term memory, embodiments of this application provide context management engineering. Detailed design diagrams of the context management engineering can be found in [reference needed]. Figure 17 As shown. Short-term memory includes intermediate results generated by each agent, such as target planning results, target storyboard images, target storyboard videos, and prompts. Long-term memory includes historical generation results stored in the database, such as historical planning results, historical prompts, historical main design results, historical storyboard images, and historical storyboard videos. In each optimization iteration, the agent selectively searches, organizes, and processes information from both short-term and long-term memory. For example, in the storyboard image generation stage, this embodiment can automatically generate the required storyboard images based on the planning results. In self-optimization, prompts are optimized by retrieving historical generation results and planning results to obtain optimized prompts, thereby achieving self-optimization of storyboard image generation.
[0159] This application first introduces an information alignment mechanism in the storyboard video generation agent. Based on the initial video prompts and the target storyboard image, information is aligned to obtain aligned video prompts, which realizes information consistency in the video generation process. This allows for the use of more accurate video prompts in the subsequent generation of storyboard videos, ensuring the generation result of the storyboard videos.
[0160] Based on the foregoing embodiments, generating a narrative video mainly includes the following six stages: scriptwriting, main design, storyboard directing and planning, storyboard image generation, storyboard video generation, and video editing. (See [link to documentation]). Figure 18As shown in the diagram, the main design stage, storyboard directing and planning stage, storyboard image generation stage, and storyboard video generation stage are each implemented using corresponding intelligent agents. Similarly, the scriptwriting stage can also be implemented using an intelligent agent, which can be called a scriptwriting intelligent agent, and the video editing stage can also be implemented using an intelligent agent, which can be called a video editing intelligent agent.
[0161] This application's embodiments can automatically generate narrative videos, shortening the traditional production cycle of several weeks to hours (e.g., generating a 10-minute short video only takes 2-3 hours), reducing manual operation time by more than 90%. Simultaneously, through multi-agent collaboration and task scheduling optimization, it further improves process throughput and meets the needs of batch creation. Multi-agent collaboration refers to a working mode in which multiple specialized agents collaborate through storyboarding to jointly complete complex tasks.
[0162] In the scriptwriting stage, the scriptwriting interface can be found here. Figure 19 As shown, during script creation, users only need to input creative inspiration and style requirements. For example, at position 1901, users can input creative inspiration (Boy A and Boy B initially dislike each other, but later cooperate to defeat character C), and at position 1902, they can input style requirements, such as general, anime style, realistic style, or work style. Afterward, users can trigger script generation, thus automatically creating the script.
[0163] In this embodiment, users do not need to master professional skills such as scriptwriting, storyboard design, and video editing to complete high-quality film and television production; independent creators and small teams do not need to invest a lot of manpower and equipment costs to quickly produce professional-grade content, thus promoting the popularization of content creation.
[0164] To improve the usability of the generated narrative videos, this application's embodiments optimize the intermediate results generated in key stages of generating the narrative videos, such as self-optimization of the storyboard directing and planning stage, self-optimization of the storyboard video generation stage, and self-optimization of the storyboard video generation stage. See [link to relevant documentation]. Figure 18 As shown.
[0165] The self-optimization of the storyboard directing planning stage, the self-optimization of the storyboard video generation stage, and the self-optimization of the storyboard video generation stage form a hierarchical iterative optimization mechanism from front to back and from coarse to fine granularity. The flowchart of this hierarchical iterative optimization mechanism can be found in [link to flowchart]. Figure 20 As shown.
[0166] The first layer of iterative optimization involves optimizing the storyboard planning AI. During the storyboard planning process, the AI fills in prompt templates based on contextual information (such as theme, style, creative inspiration, and script) to obtain planning prompts. In generating the target planning result based on these prompts, self-optimization can occur. Specifically, this self-optimization can involve generating a planning result based on the prompts, then conducting a rationality assessment of the generated result based on the director's planning evaluation, resulting in a first evaluation result. If the first evaluation result indicates that the generated planning result is unusable, the prompt optimization model of the storyboard director planning AI optimizes the used prompts to obtain optimized prompts. The planning result is then regenerated, and the subsequent evaluation steps are repeated until a usable planning result is obtained, which is then used as the target planning result.
[0167] The second iterative optimization layer optimizes the storyboard image generation agent. During the storyboard image generation process, the agent self-optimizes the image cues for each storyboard (e.g., storyboard 1, storyboard 2, ..., storyboard N), thereby generating the target storyboard image using the optimized cues. Specifically, the self-optimization can be based on generating storyboard images from image cues, evaluating the generated storyboard images based on basic quality assessment criteria, and obtaining a second evaluation result. If the second evaluation result indicates that the generated storyboard image is unusable, the optimization model further optimizes the image cues to obtain optimized cues, regenerates the storyboard image, and repeats the subsequent evaluation steps until a usable storyboard image is obtained, which is then used as the target storyboard image.
[0168] The third layer of iterative optimization involves optimizing the storyboard video generation agent. During the storyboard video generation process, the agent aligns and self-optimizes the video cues for each storyboard (e.g., storyboard 1, storyboard 2, ..., storyboard N), thereby generating the target storyboard video using the optimized cues. First, the video cues in the target planning results are aligned to obtain aligned cues. Then, the aligned cues are self-optimized to obtain optimized cues, which are then used to generate the target storyboard video. The specific method of information alignment can be based on the video cues in the target planning results, the storyboard's plot description, the voice-over characters, the target storyboard images, and some information alignment criteria. This alignment is performed using an alignment model to obtain the aligned video cues.
[0169] This application's embodiments construct a hierarchical iterative optimization mechanism from front to back and from coarse to fine granularity during multi-agent collaboration. Through a three-level closed-loop feedback, it achieves a gradual improvement in the quality of generated content, thereby enhancing the usability of the generated narrative videos and ensuring narrative fluency across scenes and multiple shots. This provides a new solution for long-form video content production in fields such as film, education, and advertising. Simultaneously, automatic optimization during the generation of narrative videos reduces repetitive manual labor, optimizes resource allocation, frees up manpower to focus on creative conception and content review, avoids resource waste, and reduces operating costs.
[0170] It should be noted that, based on the implementation methods provided in the above aspects, this application can be further combined to provide more implementation methods.
[0171] Based on the narrative video generation method provided in the foregoing embodiments, this application also provides a narrative video generation apparatus 2100. See also... Figure 21 As shown, the plot video generation device 2100 includes a display unit 2101 and an entry unit 2102: The display unit 2101 is used to display the storyboard image generation interface, which includes the first plot description of the current storyboard. The display unit 2101 is also used to display the target storyboard image of the current storyboard in response to the image generation operation performed on the storyboard image generation interface, and the first plot description is used to verify whether the target storyboard image conforms to the storyboard intention. The entry unit 2102 is used to enter the storyboard video generation interface in response to the first confirmation operation of the target storyboard image. The storyboard video generation interface includes the second plot description of the current storyboard and the attribute information of the voice actor. The display unit 2101 is further configured to respond to the video generation operation performed on the storyboard video generation interface, and to display the target storyboard video of the current storyboard generated based on the target storyboard image. The attribute information is used to verify the matching degree between the current voice-over character in the target storyboard video and the second plot description. The target storyboard video is used to generate the plot video.
[0172] In one possible implementation, the storyboard image generation interface further includes a main reference image, and the display unit 2101 is used for: In response to the image generation operation, the target storyboard image generated based on the main reference image is displayed.
[0173] In one possible implementation, the display unit 2101 is used for: The storyboard director planning interface is displayed. The storyboard director planning interface includes the target planning result, which includes the main design result and the first plot description. In response to the second confirmation operation performed on the storyboard director planning interface, the storyboard image generation interface is entered. During the process of entering the storyboard image generation interface, the first plot description is imported into the plot display area of the storyboard image generation interface, and the main reference image corresponding to the main design result is imported into the image display area of the storyboard image generation interface.
[0174] In one possible implementation, the display unit 2101 is used for: The main design interface is displayed, and the main design interface includes the main design result; In response to the third confirmation operation performed on the main design interface, the storyboard director planning interface is entered. During the process of entering the storyboard director planning interface, the main design result is imported into the main display area of the storyboard director planning interface.
[0175] In one possible implementation, the main design result includes the main reference image, which is generated based on the main description; Alternatively, the main reference image may be obtained in response to an image upload operation.
[0176] In one possible implementation, the target planning result further includes the image prompt of the current storyboard, and the display unit 2101 is further used for: During the process of entering the storyboard image generation interface, the image prompts are imported into the prompt display area of the storyboard image generation interface; The display unit 2101, in response to the image generation operation performed on the storyboard image generation interface, displays the generated target storyboard image of the current storyboard, including: In response to the image generation operation, the target storyboard image generated according to the image prompt is displayed.
[0177] In one possible implementation, the target planning result further includes the video prompts for the current scene, and the entry unit 2102 is used for: In response to the first confirmation operation, the storyboard video generation interface is entered. During the process of entering the storyboard video generation interface, the video prompt words are imported into the prompt word display area of the storyboard video generation interface. The display unit 2101, in response to the video generation operation performed on the storyboard video generation interface, displays the target storyboard video of the current storyboard generated based on the target storyboard image, including: In response to the video generation operation, the target storyboard video generated based on the video prompt and the target storyboard image is displayed.
[0178] In one possible implementation, the generation of the target planning result includes: Based on the initial planning prompts, script creation results, and the aforementioned main design results, generate the initial planning results; The initial planning results were evaluated using the director's planning evaluation criteria to obtain the first evaluation result; The initial planning prompts are optimized based on the first evaluation results to obtain optimized planning prompts. Based on the optimized planning prompts, the script creation results, and the main design results, the target planning result is generated.
[0179] In one possible implementation, the display unit 2101 is further configured to: In response to the image insertion operation performed on the storyboard image generation interface, an image browsing interface is displayed; If a selection operation is performed on the main reference image in the image browsing interface, the main reference image will be displayed in the storyboard image generation interface.
[0180] In one possible implementation, the display unit 2101 is used for: In response to the image generation operation, the target storyboard image generated according to the image prompts of the current storyboard is displayed.
[0181] In one possible implementation, the storyboard image generation interface includes the image prompts, and the device further includes a modification unit: The modification unit is used to obtain the modified image prompt in response to the modification operation of the image prompt; The display unit 2101 is used for: In response to the image generation operation, the target storyboard image generated according to the modified image prompt is displayed.
[0182] In one possible implementation, the display unit 2101 is used for: In response to the image generation operation, an initial storyboard image for the current storyboard is generated; In response to an input operation for the target storyboard image, the target storyboard image is displayed on the storyboard image generation interface. The target storyboard image is obtained by adjusting the initial storyboard image.
[0183] In one possible implementation, the entry unit 2102 is used for: In response to the first confirmation operation, the storyboard video generation interface is entered. During the process of entering the storyboard video generation interface, the target storyboard image is imported into the image display area of the storyboard video generation interface.
[0184] In one possible implementation, the display unit 2101 is used for: In response to the video generation operation, the target storyboard video, generated based on the video prompts of the current storyboard and the target storyboard image, is displayed.
[0185] In one possible implementation, the storyboard video generation interface includes the video prompts, and the device further includes a modification unit: The modification unit is used to obtain the modified video prompt in response to the modification operation of the video prompt word; The display unit 2101 is used for: In response to the video generation operation, the target storyboard video generated based on the modified video prompt and the target storyboard image is displayed.
[0186] In one possible implementation, the image prompt is an initial image prompt, and the display unit 2101 is used for: In response to the image generation operation, the initial storyboard image for the current storyboard is generated according to the initial image prompt. The initial storyboard images are evaluated using basic quality assessment criteria and instructions, and a second assessment result is obtained. The initial image prompts are optimized based on the second evaluation results to obtain optimized image prompts; The target storyboard image is generated according to the optimized image prompts, and then displayed.
[0187] In one possible implementation, the video prompt is an initial video prompt, and the display unit 2101 is used for: In response to the video generation operation, an initial storyboard video for the current storyboard is generated based on the initial video prompt and the target storyboard image; The initial storyboard video is evaluated using basic quality assessment criteria and instructions, and a third assessment result is obtained. The initial video prompts are optimized based on the third evaluation results to obtain optimized video prompts. The target storyboard video is generated based on the optimized video prompts and the target storyboard image, and then displayed.
[0188] In one possible implementation, the display unit 2101 is used for: In response to the video generation operation, information alignment is performed based on the initial video prompt and the target storyboard image to obtain the aligned video prompt; The initial storyboard video is generated based on the aligned video cue words and the target storyboard image; The aligned video prompts are optimized based on the third evaluation results to obtain the optimized video prompts.
[0189] As can be seen from the above technical solution, the generation process of the narrative video provided in this application is visualized, and the intermediate results are transparent, allowing users to verify and adjust them. Specifically, when a narrative video needs to be generated, a storyboard image generation interface can be displayed, which includes the first plot description of the current storyboard. The first plot description can serve as reference information for storyboard image generation. Thus, in response to the image generation operation performed on the storyboard image generation interface, after displaying the target storyboard image of the current storyboard, the user can verify whether the target storyboard image conforms to the storyboard intent based on the first plot description. This facilitates the user's verification and adjustment of the rationality of the generated storyboard image, thereby obtaining a target storyboard image that conforms to the storyboard intent. If a first confirmation operation is performed on the target storyboard image, the storyboard video generation interface is entered in response to the first confirmation operation. The storyboard video generation interface includes the second plot description of the current storyboard and the attribute information of the voice actor. The second plot description and the attribute information of the voice-over characters can serve as reference information for generating storyboard videos. In response to the video generation operation performed on the storyboard video generation interface, after displaying the target storyboard video generated based on the target storyboard image, users can verify the matching degree between the current voice-over character and the second plot description in the generated storyboard video based on the attribute information. This facilitates users in verifying and adjusting the rationality of the generated storyboard video, avoiding common problems such as character misalignment, and obtaining a reasonable target storyboard video. Furthermore, a reasonable plot video can be generated based on the target storyboard video. Therefore, the generation process of storyboard images and storyboard videos in this application is visualized; the intermediate results, namely the target storyboard image and target storyboard video, are transparent. Reference information is provided to users during the generation process, allowing them to intuitively and accurately verify and adjust the intermediate results based on this information, ensuring the rationality of the intermediate results and thus improving the quality of the generated plot video.
[0190] This application also provides a computer device capable of executing a method for generating narrative videos. This computer device may be a terminal. Figure 22 This diagram illustrates the structure of a terminal according to an embodiment of this application. Figure 22 In this example, using a smartphone as the terminal: refer to Figure 22 A smartphone includes components such as: radio frequency (RF) circuitry 2210, memory 2220, input unit 2230, display unit 2240, sensor 2250, audio circuitry 2260, Wi-Fi module 2270, processor 2280, and power supply 2290. The input unit 2230 may include a touch panel 2231 and other input devices 2232, the display unit 2240 may include a display panel 2241, and the audio circuitry 2260 may include a speaker 2261 and a microphone 2262. It is understood that... Figure 22 The smartphone structure shown does not constitute a limitation on smartphones and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0191] The memory 2220 can be used to store software programs and modules. The processor 2280 executes various functions and data processing of the smartphone by running the software programs and modules stored in the memory 2220. The memory 2220 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, applications required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the smartphone (such as audio data, phonebook, etc.). In addition, the memory 2220 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0192] The processor 2280 is the control center of the smartphone, connecting various parts of the smartphone via various interfaces and lines. It performs various functions and processes data by running or executing software programs and / or modules stored in the memory 2220, and by accessing data stored in the memory 2220. Optionally, the processor 2280 may include one or more processing units; preferably, the processor 2280 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 2280.
[0193] In this embodiment, the processor 2280 in the smartphone can execute the methods provided in the various embodiments of this application.
[0194] The computer device provided in this application embodiment can also be a server. Please refer to [link / reference]. Figure 23 As shown, Figure 23 This is a structural diagram of the server 2300 provided in this application embodiment. The server 2300 can vary significantly due to different configurations or performance. It may include one or more processors, such as a central processing unit (CPU) 2322, and a memory 2332, and one or more storage media 2330 (e.g., one or more mass storage devices) for storing application programs 2342 or data 2344. The memory 2332 and storage media 2330 can be temporary or persistent storage. The program stored in the storage media 2330 may include one or more modules (not shown in the diagram), each module may include a series of instruction operations on the server. Furthermore, the CPU 2322 may be configured to communicate with the storage media 2330 and execute the series of instruction operations in the storage media 2330 on the server 2300.
[0195] Server 2300 may also include one or more power supplies 2326, one or more wired or wireless network interfaces 2350, one or more input / output interfaces 2358, and / or one or more operating systems 2341, such as Windows Server. TM Mac OS X TM Unix TM Linux TM FreeBSD TM etc.
[0196] In this embodiment, the central processing unit 2322 in the server 2300 can execute the methods provided in the various embodiments of this application.
[0197] According to one aspect of this application, a computer-readable storage medium is provided for storing a computer program for performing the methods described in the foregoing embodiments.
[0198] According to one aspect of this application, a computer program product is provided, comprising a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to perform the methods provided in various optional implementations of the above embodiments.
[0199] The descriptions of the processes or structures corresponding to the above figures each have their own emphasis. For parts of a process or structure that are not described in detail, please refer to the relevant descriptions of other processes or structures.
[0200] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0201] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.
[0202] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0203] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0204] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a terminal, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing computer programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0205] In the embodiments of this application, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0206] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for generating a narrative video, characterized in that, The method includes: The storyboard image generation interface is displayed, which includes the first plot description of the current storyboard. In response to the image generation operation performed on the storyboard image generation interface, the target storyboard image of the generated current storyboard is displayed, and the first plot description is used to verify whether the target storyboard image conforms to the storyboard intent; In response to the first confirmation operation of the target storyboard image, the storyboard video generation interface is entered, which includes a second plot description of the current storyboard and attribute information of the voice actor. In response to the video generation operation performed on the storyboard video generation interface, a target storyboard video of the current storyboard generated based on the target storyboard image is displayed. The attribute information is used to verify the matching degree between the current voice-over character in the target storyboard video and the second plot description. The target storyboard video is used to generate the plot video.
2. The method according to claim 1, characterized in that, The storyboard image generation interface also includes a main reference image. The step of responding to the image generation operation performed on the storyboard image generation interface by displaying the generated target storyboard image for the current storyboard includes: In response to the image generation operation, the target storyboard image generated based on the main reference image is displayed.
3. The method according to claim 2, characterized in that, The interface for generating storyboard images includes: The storyboard director planning interface is displayed. The storyboard director planning interface includes the target planning result, which includes the main design result and the first plot description. In response to the second confirmation operation performed on the storyboard director planning interface, the storyboard image generation interface is entered. During the process of entering the storyboard image generation interface, the first plot description is imported into the plot display area of the storyboard image generation interface, and the main reference image corresponding to the main design result is imported into the image display area of the storyboard image generation interface.
4. The method according to claim 3, characterized in that, The interface for displaying storyboard director planning includes: The main design interface is displayed, and the main design interface includes the main design result; In response to the third confirmation operation performed on the main design interface, the storyboard director planning interface is entered. During the process of entering the storyboard director planning interface, the main design result is imported into the main display area of the storyboard director planning interface.
5. The method according to claim 4, characterized in that, The main design result includes the main reference image, which is generated based on the main description. Alternatively, the main reference image may be obtained in response to an image upload operation.
6. The method according to claim 3, characterized in that, The target planning result also includes the image prompts for the current storyboard, and the method further includes: During the process of entering the storyboard image generation interface, the image prompts are imported into the prompt display area of the storyboard image generation interface; The step of responding to the image generation operation performed on the storyboard image generation interface and displaying the generated target storyboard image for the current storyboard includes: In response to the image generation operation, the target storyboard image generated according to the image prompt is displayed.
7. The method according to claim 3, characterized in that, The target planning result also includes the video prompt text for the current storyboard. The step of entering the storyboard video generation interface in response to the first confirmation operation of the target storyboard image includes: In response to the first confirmation operation, the storyboard video generation interface is entered. During the process of entering the storyboard video generation interface, the video prompt words are imported into the prompt word display area of the storyboard video generation interface. The step of responding to the video generation operation performed on the storyboard video generation interface by displaying the target storyboard video of the current storyboard generated based on the target storyboard image includes: In response to the video generation operation, the target storyboard video generated based on the video prompt and the target storyboard image is displayed.
8. The method according to claim 3, characterized in that, The methods for generating the target planning results include: Based on the initial planning prompts, script creation results, and the aforementioned main design results, generate the initial planning results; The initial planning results were evaluated using the director's planning evaluation criteria to obtain the first evaluation result; The initial planning prompts are optimized based on the first evaluation results to obtain optimized planning prompts. Based on the optimized planning prompts, the script creation results, and the main design results, the target planning result is generated.
9. The method according to claim 2, characterized in that, The method further includes: In response to the image insertion operation performed on the storyboard image generation interface, an image browsing interface is displayed; If a selection operation is performed on the main reference image in the image browsing interface, the main reference image will be displayed in the storyboard image generation interface.
10. The method according to claim 1, characterized in that, The step of responding to the image generation operation performed on the storyboard image generation interface and displaying the generated target storyboard image for the current storyboard includes: In response to the image generation operation, the target storyboard image generated according to the image prompts of the current storyboard is displayed.
11. The method according to claim 10, characterized in that, The storyboard image generation interface includes the image prompts, and the method further includes: In response to the modification operation of the image prompt, the modified image prompt is obtained; The step of responding to the image generation operation by displaying the target storyboard image generated according to the image prompts of the current storyboard includes: In response to the image generation operation, the target storyboard image generated according to the modified image prompt is displayed.
12. The method according to claim 10, characterized in that, The image prompt is the initial image prompt. The step of responding to the image generation operation by displaying the target storyboard image generated according to the current storyboard image prompt includes: In response to the image generation operation, the initial storyboard image for the current storyboard is generated according to the initial image prompt. The initial storyboard images are evaluated using basic quality assessment criteria and instructions, and a second assessment result is obtained. The initial image prompts are optimized based on the second evaluation results to obtain optimized image prompts; The target storyboard image is generated according to the optimized image prompts, and then displayed.
13. The method according to claim 1, characterized in that, The step of responding to the image generation operation performed on the storyboard image generation interface and displaying the generated target storyboard image for the current storyboard includes: In response to the image generation operation, an initial storyboard image for the current storyboard is generated; In response to an input operation for the target storyboard image, the target storyboard image is displayed on the storyboard image generation interface. The target storyboard image is obtained by adjusting the initial storyboard image.
14. The method according to claim 1, characterized in that, The step of entering the storyboard video generation interface in response to the first confirmation operation of the target storyboard image includes: In response to the first confirmation operation, the storyboard video generation interface is entered. During the process of entering the storyboard video generation interface, the target storyboard image is imported into the image display area of the storyboard video generation interface.
15. The method according to claim 1, characterized in that, The step of responding to the video generation operation performed on the storyboard video generation interface by displaying the target storyboard video of the current storyboard generated based on the target storyboard image includes: In response to the video generation operation, the target storyboard video, generated based on the video prompts of the current storyboard and the target storyboard image, is displayed.
16. The method according to claim 15, characterized in that, The storyboard video generation interface includes the video prompts, and the method further includes: In response to the modification operation of the video prompt, the modified video prompt is obtained; The step of displaying the target storyboard video generated based on the video cues of the current storyboard and the target storyboard image in response to the video generation operation includes: In response to the video generation operation, the target storyboard video generated based on the modified video prompt and the target storyboard image is displayed.
17. The method according to claim 15, characterized in that, The video prompt is the initial video prompt. The process of responding to the video generation operation by displaying the target storyboard video generated based on the current storyboard video prompt and the target storyboard image includes: In response to the video generation operation, an initial storyboard video for the current storyboard is generated based on the initial video prompt and the target storyboard image; The initial storyboard video is evaluated using basic quality assessment criteria and instructions, and a third assessment result is obtained. The initial video prompts are optimized based on the third evaluation results to obtain optimized video prompts. The target storyboard video is generated based on the optimized video prompts and the target storyboard image, and then displayed.
18. The method according to claim 17, characterized in that, In response to the video generation operation, the process of generating an initial storyboard video for the current storyboard based on the initial video prompt and the target storyboard image includes: In response to the video generation operation, information alignment is performed based on the initial video prompt and the target storyboard image to obtain the aligned video prompt; The initial storyboard video is generated based on the aligned video cue words and the target storyboard image; The optimization of the initial video prompts based on the third evaluation result to obtain optimized video prompts includes: The aligned video prompts are optimized based on the third evaluation results to obtain the optimized video prompts.
19. A device for generating a narrative video, characterized in that, The device includes a display unit and an entry unit: The display unit is used to display the storyboard image generation interface, which includes the first plot description of the current storyboard. The display unit is also used to display the target storyboard image of the current storyboard in response to the image generation operation performed on the storyboard image generation interface, and the first plot description is used to verify whether the target storyboard image conforms to the storyboard intention. The entry unit is used to enter the storyboard video generation interface in response to the first confirmation operation of the target storyboard image. The storyboard video generation interface includes the second plot description of the current storyboard and the attribute information of the voice actor. The display unit is also configured to respond to the video generation operation performed on the storyboard video generation interface, and to display the target storyboard video of the current storyboard generated based on the target storyboard image. The attribute information is used to verify the matching degree between the current voice-over character in the target storyboard video and the second plot description. The target storyboard video is used to generate the plot video.
20. A computer device, characterized in that, The computer device includes a processor and memory: The memory is used to store computer programs and to transfer the computer programs to the processor; The processor is configured to execute the method according to any one of claims 1-18 according to instructions in the computer program.
21. A computer-readable storage medium for storing a computer program that, when executed by a processor, causes the processor to perform the method of any one of claims 1-18.
22. A computer program product comprising a computer program that, when executed by a processor, implements the method of any one of claims 1-18.
Citation Information
Cited By
A knowledge-enhanced iterative self-optimizing text-to-video method and system
CN122138025A