Multi-modal reinforcement learning based multi-role interactive video generation method, electronic device, readable storage medium and computer program product
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-21
- Publication Date
- 2026-08-11
AI Technical Summary
[0007]本申请的主要目的在于提供一种基于多模态强化学习的多角色交互视频生成方法、电子设备、可读存储介质及计算机程序产品,旨在解决如何在利用AI生成效率的同时,赋予用户充分控制权以实现多角色交互视频生成的技术问题
[0043]通过强化学习叙事引擎获得动态叙事路径,实现自动化叙事分支决策,解决了传统方法叙事僵化的问题;而通过无线画布界面响应用户的交互编辑指令,实现了实时交互编辑能力,使得AI生成结果不再是一个不可修改的黑盒,用户可以灵活调整叙事逻辑和镜头顺序,从而实现了人机协同的高效创作,进而达到既具备AI的高效生成能力,又保留了创作者的艺术控制权的目的。
Smart Images

Figure CN122554700A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video generation technology, and in particular to a multi-role interactive video generation method based on multimodal reinforcement learning, an electronic device, a readable storage medium, and a computer program product. Background Technology
[0002] In the current field of AI video generation, especially in the automated production of multi-character narrative content such as short dramas and comics, the technical paths are mainly divided into two categories: one is the linear generation from script to video based on large language models, and the other is the optimization of storyboard sequences based on reinforcement learning.
[0003] For example, Chinese patent CN1210009 discloses a method for generating animation videos based on video scripts. It optimizes a predefined shot rule base using a reinforcement learning model to generate a narratively coherent sequence of shots. This solution achieves a mapping from structured scripts to dynamic shot sequences, but its technical approach is unidirectional and closed. Once the reinforcement learning model determines the shot sequence and shot parameters based on the initial script, subsequent generation processes do not allow user intervention or adjustment. This means that the narrative rhythm, shot logic, and character performance of the entire video are entirely decided by the pre-trained model in a black box. However, AI models have limited ability to understand advanced narrative logic such as emotional progression and suspense building in complex narratives, and the linear results they generate often deviate from the creator's expectations. When users are dissatisfied with the narrative order, character performance, or style of a shot, they can only modify the initial script parameters and re-execute the entire generation process, leading to a serious waste of computational resources and low creative efficiency.
[0004] For example, Chinese patent CN121126081A discloses a method that focuses on improving lip-sync and visual consistency in multi-role videos through a multimodal diffusion model, but it also does not provide a mechanism that allows users to intervene in and adjust intermediate results during the generation process.
[0005] Therefore, there is currently a lack of a video generation architecture that can integrate AI's automated narrative decision-making with users' refined and interactive control. This results in existing systems either relying entirely on rigid automated processes, producing video content that fails to meet creators' advanced narrative needs, or offering only extremely limited parameter adjustment interfaces, unable to achieve flexible, real-time, and non-linear editing and control of core generation elements such as narrative paths and storyboard sequences. How to provide a multi-role interactive video generation solution that can both utilize AI's generation efficiency and give users sufficient control to achieve "what you see is what you get" is a technical problem that urgently needs to be solved in this field.
[0006] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention
[0007] The main objective of this application is to provide a method, electronic device, readable storage medium, and computer program product for generating multi-role interactive videos based on multimodal reinforcement learning, aiming to solve the technical problem of how to give users full control while utilizing the efficiency of AI generation to achieve multi-role interactive video generation.
[0008] To achieve the above objectives, this application proposes a multi-role interactive video generation method based on multimodal reinforcement learning, wherein the multimodal reinforcement learning-based multi-role interactive video generation method includes:
[0009] Generate an initial narrative blueprint based on the user's multimodal input;
[0010] The initial narrative blueprint is passed to the reinforcement learning narrative engine to obtain a dynamic narrative path. The reinforcement learning narrative engine takes the narrative state as input, maximizes the preset reward as the decision objective, and outputs the next narrative node.
[0011] Based on the dynamic narrative path, generate multi-character dialogue scripts and dynamic storyboard sequences;
[0012] Input the dynamic storyboard sequence and the multi-role dialogue script into the multimodal video generation pipeline to generate visual and audio materials;
[0013] In response to user interactive editing commands on the dynamic storyboard sequence or the visual material, the dynamic narrative path or the dynamic storyboard sequence is adjusted in real time to generate adjusted video parameters;
[0014] Based on the visual materials, audio materials, and the adjusted video parameters, a target multi-role interactive video is generated by fusion.
[0015] In one embodiment, the step of passing the initial narrative blueprint to the reinforcement learning narrative engine to obtain a dynamic narrative path includes:
[0016] The narrative nodes in the initial narrative blueprint are defined as the state space, and the narrative branch actions are defined as the action space;
[0017] A narrative strategy model is trained based on a deep reinforcement model, and the state-action value function of the narrative strategy model is: ,in, For the current narrative state, For the execution of narrative branch actions, For the preset reward, As a discount factor, The next narrative state after the action is performed;
[0018] Based on the narrative strategy model, the optimal branch is selected from the narrative branch actions to generate the dynamic narrative path.
[0019] In one embodiment, the method further includes:
[0020] Construct a role consistency knowledge base, which is used to store the visual appearance information, clothing information and voice parameter information of each role in historically generated videos;
[0021] The steps of generating multi-character dialogue scripts and dynamic storyboard sequences based on the dynamic narrative path include:
[0022] In the aforementioned role consistency knowledge base, retrieve the historical dialogue style and voice parameters corresponding to the current role, and generate dialogue content with corresponding role consistency constraints; and,
[0023] The visual appearance information stored in the character consistency knowledge base is invoked to lock the visual consistency of the same character in different shots.
[0024] In one embodiment, the step of generating multi-character dialogue scripts and dynamic storyboard sequences based on the dynamic narrative path further includes:
[0025] Within the dynamic narrative path, obtain the core theme, target emotion, and list of involved characters for the current node;
[0026] Based on the character list, retrieve the historical dialogue style and voice parameters of each character from the character consistency knowledge base;
[0027] Based on the core theme, target emotion, character list, and corresponding historical dialogue styles, prompt words are generated and input into a pre-tuned large language model to obtain at least two rounds of dialogue content as the multi-character dialogue script.
[0028] In one embodiment, the step of generating multi-character dialogue scripts and dynamic storyboard sequences based on the dynamic narrative path further includes:
[0029] Analyze the emotion tags and character action descriptions in the multi-role dialogue script;
[0030] Load a predefined shot rule library, which defines the mapping relationship between emotion tags and shot type and camera movement method;
[0031] Input the emotional intensity of the current dialogue, the positional relationship of the characters, and the scene complexity into the reinforcement learning model, optimize the mapping relationship, and output the shot switching sequence.
[0032] In one embodiment, the step of responding to a user's interactive editing command on the dynamic storyboard sequence or the visual material, adjusting the dynamic narrative path or the dynamic storyboard sequence in real time, and generating adjusted video parameters includes:
[0033] In response to the user's drag operation on the selected shot, the order of the dynamic shot sequence is reconstructed, and the multimodal video generation pipeline is triggered to perform regeneration and transition frame interpolation operations on the visual materials of the selected shot and its adjacent shots.
[0034] In response to the user's command to switch the style of the selected shot, visual materials of the corresponding style are regenerated based on the diffusion model.
[0035] In one embodiment, the step of inputting the dynamic storyboard sequence and the multi-role dialogue script into a multimodal video generation pipeline to generate visual and audio materials includes:
[0036] Based on the text-based diffusion model, keyframe images are generated according to the shot descriptions and character appearance information in the dynamic storyboard sequence.
[0037] Based on the text-to-speech model, the dialogue audio of each character is generated according to the lines and voice parameters of the multi-role dialogue script, and each audio segment is marked with a corresponding timestamp and character ID.
[0038] Based on the timestamp and character ID, align the dialogue audio with the character's lip-sync region in the keyframe image.
[0039] In addition, to achieve the above objectives, this application also proposes an electronic device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the multi-role interactive video generation method based on multimodal reinforcement learning as described above.
[0040] In addition, to achieve the above objectives, this application also proposes a readable storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the multi-role interactive video generation method based on multimodal reinforcement learning as described above.
[0041] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the multi-role interactive video generation method based on multimodal reinforcement learning as described above.
[0042] One or more technical solutions proposed in this application have at least the following technical effects:
[0043] By using a reinforcement learning narrative engine to obtain dynamic narrative paths and achieve automated narrative branch decisions, the problem of rigid narrative in traditional methods is solved. Furthermore, by responding to user interactive editing commands through a wireless canvas interface, real-time interactive editing capabilities are achieved. This makes the AI-generated results no longer an unmodifiable black box, allowing users to flexibly adjust the narrative logic and shot order. This enables efficient human-computer collaboration and achieves the goal of possessing both the efficient generation capabilities of AI and retaining the creator's artistic control.
[0044] By constructing a unified character consistency knowledge base, visual, clothing, and voice information of characters are centrally stored and consistently invoked in stages such as dialogue generation, storyboard generation, and video generation, forming a closed loop. Compared to using a fixed seed only in a single stage, such as image generation, this application's embodiment achieves cross-modal and cross-step consistency constraints, further enhancing the professional look of multi-character narrative videos. It avoids inconsistencies in facial features, clothing, and voice for the same character across multiple shots and episodes. Attached Figure Description
[0045] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0046] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0047] Figure 1 This is a flowchart illustrating the first embodiment of the multi-role interactive video generation method based on multimodal reinforcement learning in this application;
[0048] Figure 2 This is a flowchart illustrating step S300 of an embodiment of the multi-role interactive video generation method based on multimodal reinforcement learning in this application.
[0049] Figure 3 This is a flowchart illustrating another embodiment of step S300 in the multi-role interactive video generation method based on multimodal reinforcement learning in this application.
[0050] Figure 4 This is a flowchart illustrating another embodiment of step S300 in the multi-role interactive video generation method based on multimodal reinforcement learning in this application.
[0051] Figure 5This is a schematic diagram of the module structure of the multi-role interactive video generation system based on multimodal reinforcement learning in this application;
[0052] Figure 6 This is a schematic diagram of the device structure of the hardware operating environment involved in the multi-role interactive video generation method based on multimodal reinforcement learning in the embodiments of this application.
[0053] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0054] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.
[0055] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0056] It should be noted that the execution subject of the multi-role interactive video generation method based on multimodal reinforcement learning in various embodiments of this application can be a multi-role interactive video generation system based on multimodal reinforcement learning, or a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device capable of realizing the above functions. This embodiment does not specifically limit it in this way. The following uses a multi-role interactive video generation system based on multimodal reinforcement learning as the execution subject as an example to describe this embodiment and the following embodiments.
[0057] Based on this, this application proposes a first embodiment of a multi-role interactive video generation method based on multimodal reinforcement learning. Please refer to [link / reference]. Figure 1 The multi-role interactive video generation method based on multimodal reinforcement learning includes steps S100 to S600.
[0058] Step S100: Generate an initial narrative blueprint based on the multimodal content input by the user.
[0059] It should be noted that users can input multimodal content through a front-end interface, such as a web page or client. Multimodal content includes at least one or more of the following types: textual ideas, such as "a suspenseful short drama about an ancient swordsman searching for a lost sword in a modern city"; image materials, such as reference images of the main character and scene atmosphere images; character settings, such as character names and personality descriptions; style instructions, such as "cyberpunk and martial arts" or "fast-paced, suspenseful".
[0060] The system is internally configured with a pre-trained large language model. This model performs semantic parsing and structuring processing on the aforementioned multimodal content. Specifically, the pre-trained large language model first inputs non-textual content, such as images, and then uses a multimodal encoder to perform text transformation operations on the non-textual content to obtain corresponding text descriptions. Then, based on the text descriptions corresponding to the non-textual content and the user-inputted text ideas, it performs natural language understanding to extract core narrative elements. These core narrative elements include, but are not limited to, the protagonist, the goal, the conflict, key events, and the emotional tone. Finally, these elements are organized according to chronological order or logical relationships into an initial narrative blueprint containing multiple narrative nodes. For example, the initial narrative blueprint can be represented as: Node 1 "Traversal and Confusion"; Node 2 "Initial Encounter with Clues"; Node 3 "Conflict and Battle"; Node 4 "Revelation of the Truth," etc. Each node also contains metadata such as the node's core theme, target emotion, and a list of involved characters.
[0061] Step S200: The initial narrative blueprint is passed to the reinforcement learning narrative engine to obtain a dynamic narrative path. The reinforcement learning narrative engine takes the narrative state as input, maximizes the preset reward as the decision objective, and outputs the next narrative node.
[0062] It's important to note that the system takes the generated initial narrative blueprint as input and feeds it into a pre-trained reinforcement learning narrative engine. The core of this engine is a deep reinforcement learning model whose function is to autonomously decide, based on the current narrative state, what the next narrative node should be, thereby generating a non-linear, dynamically adjusted narrative path. This current narrative state includes, for example, which narrative node the system is currently at, as well as the user interactions or predicted audience behavior at that node.
[0063] For example, in the initial narrative blueprint, after node 2 "Initial Encounter with Clues," there are two branches: branch one, winning in direct combat; branch two, escaping through clever tactics. The reinforcement learning narrative engine selects the optimal branch based on preset rewards, such as predicting the increase in audience dwell time or the intensity of emotional resonance that the branch will bring. If the system detects that the pacing of the preceding nodes is too slow, it can choose "winning in direct combat" to speed up the pace; conversely, it can choose "escaping through clever tactics" to increase suspense.
[0064] The output of step S200 is a dynamic narrative path, which may differ from the linear order of the initial narrative blueprint. For example, it may follow the order of node 1, node 2, node 4, skipping node 3 and supplementing it later through flashbacks.
[0065] Step S300: Generate a multi-character dialogue script and a dynamic storyboard sequence based on the dynamic narrative path.
[0066] Once the dynamic narrative path is determined, the system calls two parallel or sequential sub-modules: the multi-role dialogue generation module and the dynamic storyboard generation module.
[0067] In the multi-character dialogue generation module, for the current narrative node, such as the "science museum dialogue," based on the core theme, target emotion, and character list of the current narrative node, and combined with the historical dialogue style and voice parameters of each character retrieved from the character consistency knowledge base, a large language model is invoked to generate multi-turn dialogue content that matches the character's personality. For example, lines with an ancient Chinese tone are generated for the knight-errant character, while mechanical and rapid lines are generated for the AI administrator character.
[0068] In the dynamic storyboard generation module, emotional tags such as "anger," "sadness," and "calm" and character action descriptions are parsed from the generated multi-character dialogue scripts. Combined with a predefined shot rule library (e.g., anger corresponds to close-ups and rapid zooms), and another reinforcement learning model, shot transitions are smoothly optimized. This ultimately generates a dynamic storyboard sequence that includes shot type, camera movement, shot duration, and the corresponding character and scene for each shot. Shot types can include close-ups, medium shots, and wide shots; camera movements can include push-in, pull-out, pan, and tilt.
[0069] Step S400: Input the dynamic storyboard sequence and the multi-role dialogue script into the multimodal video generation pipeline to generate visual and audio materials.
[0070] In this embodiment, the multimodal video generation pipeline includes a visual generation task, an audio generation task, and an audio-visual synchronization task.
[0071] In the visual generation task, textual graph diffusion models, such as Stable Diffusion, are used to generate keyframe images based on shot descriptions in the storyboard sequence, such as "close-up, a swordsman in a green robe, with a determined look, against a background of a science museum full of holographic projections," and character appearance information extracted from a character consistency knowledge base.
[0072] In the audio generation task, a text-to-speech (TTS) model is used to generate dialogue audio for each character, based on lines from a multi-role dialogue script and character vocal parameters such as pitch, speech rate, and timbre extracted from a character consistency knowledge base. Simultaneously, each audio segment is tagged with a timestamp and its corresponding character ID.
[0073] In the audio-visual synchronization task, the audio-visual synchronization module is used to perform pixel-level alignment between the dialogue audio and the lip-sync region of the corresponding character in the keyframe image based on the audio timestamp and character ID, ensuring that the lip-sync matches the speech.
[0074] Step S500: In response to the user's interactive editing instructions on the dynamic storyboard sequence or the visual material, adjust the dynamic narrative path or the dynamic storyboard sequence in real time and generate the adjusted video parameters.
[0075] This application provides a wireless canvas interface that presents multiple shots in a dynamic storyboard sequence as visual cards. Users can directly drag, copy, delete, or adjust the style parameters of the visual materials on the canvas.
[0076] For example, if a user feels the second scene (battle scene) is too fast-paced and wants to swap its position with the fourth scene (suspense reveal), they simply need to drag the fourth scene card before the second scene card. Upon detecting this drag operation, the system reconstructs the dynamic scene sequence order and triggers the multimodal video generation pipeline to regenerate and interpolate transition frames for the visual materials of the affected scene (i.e., the moved shot and its adjacent shots) to ensure a smooth transition. Simultaneously, the system saves the adjusted scene sequence order, shot parameters, etc., as adjusted video parameters.
[0077] Step S600: Based on the visual materials, audio materials, and the adjusted video parameters, a target multi-role interactive video is generated by fusing them together.
[0078] Finally, the system calls a video compositing module (such as the FFmpeg library) to synchronously encapsulate the visual and audio materials generated in step S400 according to the timeline based on the adjusted video parameters, and renders and outputs a complete target multi-character interactive video file. The adjusted video parameters may include the adjusted scene order, the start and end times of each shot, transition effects, etc.
[0079] In the technical solution provided in this embodiment, a dynamic narrative path is obtained through a reinforcement learning narrative engine, and automated narrative branch decision-making is achieved, which solves the problem of rigid narrative in traditional methods. Furthermore, by responding to the user's interactive editing commands through a wireless canvas interface, real-time interactive editing capabilities are realized, so that the AI-generated results are no longer an unmodifiable black box. Users can flexibly adjust the narrative logic and shot order, thereby achieving efficient creation through human-computer collaboration. In this way, the goal of having both the efficient generation capabilities of AI and retaining the creator's artistic control is achieved.
[0080] In one feasible implementation, step S200 may include steps S210 to S230:
[0081] Step S210: Define the narrative nodes in the initial narrative blueprint as the state space, and define the narrative branch actions as the action space;
[0082] Step S220: Train a narrative strategy model based on a deep reinforcement model, wherein the state-action value function of the narrative strategy model is: ,in, For the current narrative state, For the execution of narrative branch actions, For the preset reward, As a discount factor, The next narrative state after the action is performed;
[0083] Step S230: Based on the narrative strategy model, select the optimal branch from the narrative branch actions to generate the dynamic narrative path.
[0084] It should be noted that when defining the state space, each narrative node in the initial narrative blueprint (such as "navigating confusion," "initial encounter with clues," and "victory in battle") is defined as a discrete or continuous state. Each state also includes additional characteristics of that node, such as its expected duration, emotional intensity, and audience dwell time feedback from previous nodes. When defining the action space, multiple optional narrative branch actions are defined under each state node. For example, under the "initial encounter with clues" node, the actions could be "Branch 1: Direct combat," "Branch 2: Outsmarting and escaping," or "Branch 3: Seeking allies." Each action corresponds to a definite next narrative node.
[0085] Next, the narrative policy model is trained using deep reinforcement learning algorithms, such as Deep Q-Network (DQN). The core of this model is the state-action value function Q(s,a).
[0086] in, It is an instant reward function. In this embodiment, The calculation is based on the predicted audience engagement metrics at the current narrative node. Specifically, the system has a built-in engagement prediction model that can estimate the expected viewing time and emotional resonance intensity of viewers after the broadcast of a given node, based on node characteristics such as emotional tags and conflict intensity. For example, a regression model trained on historical data outputs a score between 0 and 1. The reward value is a linear weighted sum of these scores. This is a discount factor, with a value between 0.9 and 0.99, used to balance the importance of current rewards and future rewards. To perform the action The next narrative state that is reached later.
[0087] For example, the training process includes iteratively training the DQN network using historical video viewing data or simulated interaction data as training samples. The goal of DQN is to learn an optimal policy π(s) such that the action selected in each state maximizes the cumulative discount reward.
[0088] During the inference phase, i.e. when the video is actually generated, the current narrative state s is input into the trained narrative strategy model. The model calculates the Q-value of each possible action and selects the action with the highest Q-value as the output, thereby determining the next narrative node. This process is repeated until a terminal node (such as the ending) is reached, thus generating a complete dynamic narrative path.
[0089] During the inference phase, i.e. when the video is actually generated, the current narrative state s is input into the trained narrative strategy model. The model calculates the Q-value of each possible action and selects the action with the highest Q-value as the output, thereby determining the next narrative node. This process is repeated until a terminal node, such as the ending, is reached, thus generating a complete dynamic narrative path.
[0090] Thus, compared to simple rule-based branching or random selection, this embodiment of the application learns successful narrative patterns from historical data through a deep Q-network and adaptively selects the branch most likely to attract the audience based on the current context. Instant reward function. By directly incorporating predicted audience engagement into the decision-making process, the generated narrative path is not static but dynamically optimized, thereby enhancing the appeal and retention rate of video content.
[0091] Reference Figure 2 As an optional implementation, step S300 may further include steps S310 to S330:
[0092] Step S310: parse the emotion tags and character action descriptions in the multi-role dialogue script;
[0093] Step S320: Load the predefined shot rule library, which defines the mapping relationship between emotion tags and shot type and camera movement method;
[0094] Step S330: Input the emotional intensity, character positional relationship, and scene complexity of the current dialogue into the reinforcement learning model, optimize the mapping relationship, and output the shot switching sequence.
[0095] For example, suppose the emotional label for the first line of dialogue is "deep and doubtful", and the action description is "frowning"; the emotional label for the second line of dialogue is "mechanical and fast", and the action description is "expressionless scanning".
[0096] Understandably, a predefined shot rule base is a knowledge table that defines the mapping from basic emotional tags to shot types. This predefined shot rule base can be pre-built based on prior knowledge in the film and television production field, as shown in Table 1.
[0097] Table 1: Emotional Tag - Shot Type Mapping Table
[0098] low, doubtful Close-up Slowly advance 2-4 seconds surprise close up Urgent push 1-2 seconds Mechanical, calm Mid-range fixed 2-3 seconds
[0099] It should be noted that although the predefined shot rule base provides a basic mapping, directly applying it may result in abrupt shot transitions; for example, strictly switching based on emotion may lead to overly fragmented shots. Therefore, this application introduces an additional reinforcement learning model to optimize the smoothness of shot transitions and narrative coherence.
[0100] The state of this reinforcement learning model includes the sentiment intensity of the current dialogue, the positional relationships between the characters currently participating in the dialogue, and the scene complexity. The sentiment intensity of the current dialogue is a continuous value, output by the sentiment classification model; the positional relationships between the characters currently participating in the dialogue can be calculated from the distance and orientation of the character bounding boxes detected in the previous frame; scene complexity includes, but is not limited to, the number of objects / characters in the scene and the richness of textures.
[0101] In this embodiment, the reinforcement learning model's action is to fine-tune the type, camera movement, and duration of the next shot. For example, it might adjust whether to extend the current shot, switch earlier, or insert a reaction shot. The reward function for this reinforcement learning model is designed as a comprehensive score of the smoothness of the switched shot sequence compared to adjacent shots (e.g., avoiding jump cuts) and the accuracy of emotional expression (e.g., whether emotional high points match close-ups).
[0102] This reinforcement learning model enables the system to adaptively adjust mapping relationships. For example, even if the current emotional label is "depressed, confused," if the characters' positions indicate a rapid firefight (high scene complexity), the reinforcement learning model can decide to use a "medium shot, rapid panning" instead of a "close-up, zoom in" to better serve the narrative rhythm. The final output shot transition sequence is a timeline list, for example: [0.0s-2.5s: medium shot, fixed; 2.5s-4.0s: close-up, zoom in; 4.0s-5.5s: medium shot, fixed].
[0103] Thus, this embodiment of the application, based on a static predefined shot rule base, achieves dynamic fine-tuning of shot parameters through a reinforcement learning model, avoiding rigidity in the shot sequence. Adaptive optimization is performed based on real-time dialogue content and scene characteristics, thereby improving the smoothness and professionalism of shot transitions while maintaining the accuracy of emotional expression.
[0104] In an optional implementation, step S500 may include steps S510 and S520:
[0105] Step S510: In response to the user's drag operation on the selected shot, reconstruct the order of the dynamic shot sequence and trigger the multimodal video generation pipeline to perform regeneration and transition frame interpolation operations on the visual materials of the selected shot and its adjacent shots.
[0106] Step S520: In response to the user's style switching command for the selected shot, regenerate the corresponding style visual material based on the diffusion model.
[0107] In this embodiment, each scene on the wireless canvas interface is represented by a card, which displays a thumbnail of the scene, the shot type, and a brief description. Users can drag and drop any card to any position using a mouse or touch. When the system detects a drag-and-release event—that is, when the user places the selected shot card in a new position—it executes a reconstruction sequence step, triggers a local regeneration action, and provides real-time feedback.
[0108] Specifically, the sequence reconstruction action includes automatically updating the order of the internally maintained list of storyboard sequences.
[0109] Triggering local regeneration involves identifying the affected shot range, encompassing all shots from the original position of the moved shot to its new position, as well as the moving shot itself. For these shots, the system re-invokes the multimodal video generation pipeline, but does not completely regenerate them. For the moved shot, its original script and character descriptions remain unchanged, but transition frames are regenerated using a frame interpolation model based on the new context to ensure smooth motion between adjacent shots. For adjacent shots, if their content creates a new logical relationship with the moved shot (such as a question-and-answer correspondence), the dialogue timeline is fine-tuned.
[0110] Real-time feedback includes animated movement of dragged cards to new positions on the wireless canvas interface, while other cards automatically rearrange. The generation progress is displayed as a progress bar on the card.
[0111] In this embodiment, users can issue style switching commands to selected shots through the right-click menu or property panel, such as switching the current "cyberpunk" style to "ink wash style" or "Japanese anime style".
[0112] The system's response to style switching commands includes: extracting the foreground subject (characters and key props) from the current visual material using an image segmentation model, preserving their outline and pose information to lock the main content; replacing the style words in the original storyboard description with user-specified new style words. For example, the original description "cyberpunk-style science museum" becomes "ink-wash style science museum" to redirect the cue words. Subsequently, using ControlNet's canny edge or depth map, the locked subject outline is used as a condition, and the new cue words are input into a diffusion model for image-to-image conversion. This aims to maintain the layout structure of characters and scenes while changing only the rendering style. Finally, new keyframe images are generated to replace the images in the original cards, triggering subsequent audio-visual synchronization fine-tuning to update the material.
[0113] Understandably, style switching does not affect the audio, so there is no need to re-synthesize the audio.
[0114] Thus, this embodiment of the application provides a drag-and-drop reconstruction function, allowing users to adjust the narrative order, and the system only regenerates the affected parts locally, avoiding the waste of computing power and waiting time caused by global regeneration. The style switching function, through subject locking technology, achieves fine-grained control of "changing style without changing content," further enriching creative freedom.
[0115] Optionally, step S400 may include steps S410 to S430:
[0116] Step S410: Based on the text-based diffusion model, keyframe images are generated according to the shot descriptions and character appearance information in the dynamic storyboard sequence.
[0117] Step S420: Based on the text-to-speech model, generate dialogue audio for each character according to the lines and character voice parameters in the multi-role dialogue script, and mark each audio segment with a corresponding timestamp and character ID;
[0118] Step S430: Based on the timestamp and character ID, align the dialogue audio with the character's lip-sync region in the keyframe image.
[0119] In this embodiment, a text-based image diffusion model (such as Stable Diffusion XL) is employed. For each shot in the dynamic shot sequence, generation conditions are constructed, including text descriptions and character appearance conditions. The text description refers to the text description of the shot, such as "close-up, the swordsman is speaking directly to the camera with a determined expression"; the character appearance conditions are obtained from the character consistency knowledge base, acquiring the visual feature vectors of all characters appearing in the current frame, and injecting them into the diffusion model through a cross-attention mechanism. The text-based image diffusion model generates high-resolution keyframe images and saves them as a PNG sequence.
[0120] Optionally, the above generation conditions may also include pose / depth conditions. If the previous frame has been generated, the pose skeleton of the previous frame is extracted as ControlNet input to ensure the continuity of the character's movements.
[0121] A text-to-speech model (such as VITS or Bark) is employed. For each line of dialogue in a multi-role dialogue script, the character's vocal parameters (timbre embedding vector) are retrieved from the character consistency knowledge base. The dialogue text and vocal parameters are then input into the TTS model as conditions to generate the corresponding WAV format audio segment. Simultaneously, the start and end times of the audio segment (relative to the entire video timeline), as well as the corresponding character ID, are recorded. This information is packaged into a metadata file (such as JSON format).
[0122] Then, the audio-visual synchronization module aligns the dialogue audio with the lip-sync regions of the characters in the keyframe images.
[0123] Specifically, for each line of dialogue, its audio sequence and corresponding timestamp are extracted. A pre-trained lip-sync model (such as Wav2Lip) is used as input. This model takes the facial region and audio features of the corresponding character in the keyframe image as input and outputs the displacement of the character's lip shape key points in each frame. Then, the modified lip shape region (lip shape image) is pasted back into the original keyframe image to generate a new frame. It should be noted that this process only modifies the mouth region and does not change the facial expression or background. Finally, all processed frames are combined in timestamp order to form an H.264 encoded video stream.
[0124] Thus, this embodiment of the application achieves high-precision audio-visual matching by independently generating visuals and audio, and then using a synchronization mechanism based on timestamps and character IDs for lip-syncing. Since lip-syncing is performed only in a local area, it avoids full-image redrawing and is computationally efficient. Simultaneously, the introduction of character IDs ensures that in multi-character dialogue scenarios, only the currently speaking character's lip movements are driven, while other characters' lip movements remain static or display default expressions, conforming to real-world interaction logic.
[0125] Reference Figure 3 Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in the first embodiment described above can be referred to the above description and will not be repeated hereafter. On this basis, the multi-role interactive video generation based on multimodal reinforcement learning may further include step S700:
[0126] Step S700: Construct a character consistency knowledge base, which is used to store the visual appearance information, clothing information and voice parameter information of each character in the historically generated videos;
[0127] The step of generating multi-character dialogue scripts and dynamic storyboard sequences based on the dynamic narrative path may further include steps S340 to S350:
[0128] Step S340: Retrieve the historical dialogue style and voice parameters corresponding to the current role from the role consistency knowledge base, and generate dialogue content corresponding to the role consistency constraints; and,
[0129] Step S350: Call the visual appearance information stored in the character consistency knowledge base to lock the visual consistency of the same character in different shots.
[0130] In this embodiment, upon system initialization or the first receipt of a user-created role, the system assigns a unique ID to the role. The role consistency knowledge base is a structured database, such as a relational database or a vector database, used to store role information for each role. Role information includes, but is not limited to, visual appearance information, clothing information, and voice parameter information.
[0131] The visual appearance information includes the feature vector of the user-uploaded character reference image, and the image photo generated by the Wensheng Image Model and confirmed by the user. The feature vector of the character reference image can be obtained through facial features extracted by a face recognition model such as FaceNet. If the user does not provide a reference image, the system generates an initial image based on the character description, such as "a thirty-year-old ancient swordsman, square face, sword-like eyebrows," and stores its features in the knowledge base.
[0132] The clothing information storage includes descriptions of a character's attire and corresponding image features in different scenarios. For example, a "knight-errant" wears a "black trench coat" in a "modern scene" and a "blue robe" in an "ancient memory." The system retrieves information by scene tag.
[0133] Vocal parameters include the character's timbre vector, speech rate range, pitch baseline, and other parameters. The timbre vector can be extracted from user-provided speech samples or default timbre using a speech encoder.
[0134] When generating multi-character dialogue scripts, the system first retrieves a list of character IDs involved in the current narrative node. Then, it queries the character consistency knowledge base. For each character ID, it retrieves the latest or most relevant "historical dialogue style" record. Historical dialogue style refers to the tone, word choice, and emotional inclination exhibited by the corresponding character in previously generated dialogues. For example, if the knowledge base records that the "knight-errant" frequently used archaic phrases like "that's enough" and "don't be rude" in previous episodes, then this style will be maintained in subsequent dialogue generation. Subsequently, these retrieved style descriptions, such as "knight-errant: classical, resolute, occasionally humorous," are used as part of the prompts and input into the large language model. The dialogue content generated by the large language model will automatically match the character's historical style, thus ensuring the consistency of the character's personality.
[0135] When generating storyboard sequences, especially when generating shot descriptions involving corresponding characters, the system utilizes visual appearance information stored in the character consistency knowledge base. For example, when a storyboard requires generating a close-up shot of a swordsman, the system extracts the swordsman's visual feature vector from the knowledge base, including facial features and clothing features corresponding to the current scene, and uses these features as conditional inputs to the text-based image diffusion model. In this way, regardless of the shot or episode, the generated swordsman image maintains a high degree of consistency in facial structure and clothing details, avoiding the "face-swapping" phenomenon.
[0136] In the technical solution provided in this embodiment, a unified character consistency knowledge base is constructed to centrally store the visual, clothing, and voice information of characters, which is then consistently invoked in stages such as dialogue generation, storyboard generation, and video generation, forming a closed loop. Compared to using a fixed seed only in a single stage, such as image generation, this embodiment achieves cross-modal and cross-step consistency constraints, further enhancing the professional look of multi-character narrative videos. It avoids inconsistencies in facial features, clothing, and voice for the same character across multiple shots and episodes.
[0137] Reference Figure 4 Furthermore, step S300 also includes steps S360 to S380:
[0138] Step S360: In the dynamic narrative path, obtain the core theme, target emotion, and list of involved characters of the current node;
[0139] Step S370: Based on the character list, retrieve the historical dialogue style and voice parameters of each character from the character consistency knowledge base;
[0140] Step S380: Based on the core theme, target emotion, character list and corresponding historical dialogue style, generate prompt words and input them into the pre-tuned large language model to obtain at least two rounds of dialogue content as the multi-role dialogue script.
[0141] In this embodiment, when generating a multi-role dialogue script, the core theme, target emotion, and list of involved characters for the current node are first obtained. For example, the current node is "Confrontation at the Science Museum", the core theme is "Explaining the Origin of the Sword", the target emotion is "Surprise and Doubt", and the list of involved characters is ["Knight-errant", "AI Administrator"].
[0142] Then, based on the character list, the historical dialogue style and voice parameters of each character are retrieved from the character consistency knowledge base. For example, suppose the search results are ["Knight-errant": dialogue style = "mixed use of classical Chinese and modern Chinese, firm tone, frequently uses words such as 'Your Excellency' and 'How dare you'"; voice parameters = "deep, magnetic, medium speaking speed"]; ["AI Administrator": dialogue style = "clear logic, emotionless, short sentences"; voice parameters = "mechanical synthesized voice, fast speaking speed, no pitch variation"].
[0143] Finally, based on the core theme, target emotion, character list, and corresponding historical dialogue styles, prompt words are generated and input into a pre-tuned large language model to obtain at least two rounds of dialogue content as a multi-character dialogue script. Specifically, a structured prompt word template is constructed, for example: [You are a professional short play writer. Please generate a multi-character dialogue based on the following information. Current scene core theme: {core theme} Target emotion: {target emotion} Character 1: {character A name}, style: {style A}, voice: {voice A} Character 2: {character B name}, style: {style B}, voice: {voice B} Please generate 3-5 rounds of dialogue, which must conform to the character's style and drive the plot forward.] After replacing the placeholders with actual values, the prompt words are input into the pre-tuned large language model. Understandably, the pre-tuned large language model has been fine-tuned using a large number of scripts and dialogues, and has the ability to generate dialogue content that conforms to the character settings.
[0144] Each dialogue segment output by the pre-tuned large language model is also accompanied by emotional tags, such as "deep," "mechanical," and "surprised," for use by the subsequent storyboard generation module.
[0145] Thus, this embodiment of the application further improves the stability and character fit of dialogue generation by explicitly incorporating historical style information from the character consistency knowledge base into the prompt words, rather than relying solely on the model's implicit memory. The pre-tuned large language model ensures the professionalism and naturalness of the generated content, while the retrieval-enhanced prompt word construction method gives each character's dialogue a unique and traceable style identifier.
[0146] Please see Figure 5 The multi-role interactive video generation system based on multimodal reinforcement learning provided in this application embodiment may include:
[0147] The input parsing module 10 is used to receive multimodal content input by the user and generate an initial narrative blueprint based on the multimodal content input by the user.
[0148] Reinforcement learning narrative engine 20 is used to generate dynamic narrative paths based on an initial narrative blueprint;
[0149] Multi-role dialogue generation module 30 is used to generate multi-role dialogue scripts based on dynamic narrative paths;
[0150] The dynamic storyboard generation module 40 is used to generate a dynamic storyboard sequence based on multi-character dialogue scripts and dynamic narrative paths.
[0151] A multimodal video generation pipeline 50 is used to generate visual and audio materials based on dynamic storyboard sequences and multi-role dialogue scripts;
[0152] The interactive canvas editing module 60 provides a wireless canvas interface, responding to user-interactive editing commands to adjust dynamic narrative paths or dynamic scene sequences in real time; and...
[0153] The video compositing module 70 is used to fuse and generate target multi-role interactive videos.
[0154] The multi-role interactive video generation system based on multimodal reinforcement learning provided in this application adopts the multi-role interactive video generation method based on multimodal reinforcement learning in the above embodiments. Compared with the prior art, the beneficial effects of the multi-role interactive video generation system based on multimodal reinforcement learning provided in this application are the same as those of the multi-role interactive video generation method based on multimodal reinforcement learning provided in the above embodiments, and other technical features of the multi-role interactive video generation system based on multimodal reinforcement learning are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0155] This application provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the multi-role interactive video generation method based on multimodal reinforcement learning in the above embodiments.
[0156] The following is for reference. Figure 6 The diagram illustrates a structural schematic of an electronic device suitable for implementing embodiments of this application. The electronic devices in these embodiments may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 6 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0157] like Figure 6 As shown, the electronic device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory 1002 or a program loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the electronic device. The processing unit 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. The communication device 1009 allows the electronic device to communicate wirelessly or wiredly with other devices to exchange data. Although the diagrams show electronic devices with various systems, it should be understood that it is not required to implement or have all of the systems shown. More or fewer systems may be implemented alternatively.
[0158] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0159] The electronic device provided in this application employs the multi-role interactive video generation method based on multimodal reinforcement learning in the above embodiments, which solves the technical problem of how to leverage AI generation efficiency while granting users sufficient control to achieve multi-role interactive video generation. Compared with the prior art, the beneficial effects of the electronic device provided in this application are the same as those of the multi-role interactive video generation method based on multimodal reinforcement learning provided in the above embodiments, and other technical features of this electronic device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0160] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0161] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0162] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the multi-role interactive video generation method based on multimodal reinforcement learning in the above embodiments.
[0163] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0164] The aforementioned computer-readable storage medium may be included in an electronic device or may exist independently without being assembled into an electronic device.
[0165] The aforementioned computer-readable storage medium carries one or more programs that, when executed by an electronic device, cause the electronic device to: generate an initial narrative blueprint based on multimodal content input by a user; transmit the initial narrative blueprint to a reinforcement learning narrative engine to obtain a dynamic narrative path, wherein the reinforcement learning narrative engine takes the narrative state as input, maximizes a preset reward as its decision objective, and outputs the next narrative node; generate a multi-character dialogue script and a dynamic storyboard sequence based on the dynamic narrative path; input the dynamic storyboard sequence and the multi-character dialogue script to a multimodal video generation pipeline to generate visual and audio materials; respond to user interactive editing commands on the dynamic storyboard sequence or the visual materials, adjust the dynamic narrative path or the dynamic storyboard sequence in real time, and generate adjusted video parameters; and fuse and generate a target multi-character interactive video based on the visual materials, audio materials, and the adjusted video parameters.
[0166] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0167] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0168] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0169] The readable storage medium provided in this application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., a computer program) for executing the above-described multi-role interactive video generation method based on multimodal reinforcement learning. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the multi-role interactive video generation method based on multimodal reinforcement learning provided in the above embodiments, and will not be repeated here.
[0170] This application provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the multi-role interactive video generation method based on multimodal reinforcement learning as described above.
[0171] Compared with the prior art, the computer program product provided in this application has the same beneficial effects as the multi-role interactive video generation method based on multimodal reinforcement learning provided in the above embodiments, and will not be repeated here.
[0172] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent scope of this application.
Claims
1. A method for generating multi-role interactive videos based on multimodal reinforcement learning, characterized in that, The multi-role interactive video generation method based on multimodal reinforcement learning includes: Generate an initial narrative blueprint based on the user's multimodal input; The initial narrative blueprint is passed to the reinforcement learning narrative engine to obtain a dynamic narrative path. The reinforcement learning narrative engine takes the narrative state as input, maximizes the preset reward as the decision objective, and outputs the next narrative node. Based on the dynamic narrative path, generate multi-character dialogue scripts and dynamic storyboard sequences; Input the dynamic storyboard sequence and the multi-role dialogue script into the multimodal video generation pipeline to generate visual and audio materials; In response to user interactive editing commands on the dynamic storyboard sequence or the visual material, the dynamic narrative path or the dynamic storyboard sequence is adjusted in real time to generate adjusted video parameters; Based on the visual materials, audio materials, and the adjusted video parameters, a target multi-role interactive video is generated by fusing them together.
2. The multi-modal reinforcement learning based multi-role interactive video generation method of claim 1, wherein, The step of passing the initial narrative blueprint to the reinforcement learning narrative engine to obtain a dynamic narrative path includes: The narrative nodes in the initial narrative blueprint are defined as the state space, and the narrative branch actions are defined as the action space; A narrative strategy model is trained based on a deep reinforcement model, and the state-action value function of the narrative strategy model is: ,in, As the current narrative state, For the execution of narrative branch actions, For the preset reward, As a discount factor, The next narrative state after the action is performed; Based on the narrative strategy model, the optimal branch is selected from the narrative branch actions to generate the dynamic narrative path.
3. The multi-role interactive video generation method based on multimodal reinforcement learning as described in claim 1, characterized in that, The method further includes: Construct a role consistency knowledge base, which is used to store the visual appearance information, clothing information and voice parameter information of each role in historically generated videos; The steps of generating multi-character dialogue scripts and dynamic storyboard sequences based on the dynamic narrative path include: In the aforementioned role consistency knowledge base, retrieve the historical dialogue style and voice parameters corresponding to the current role, and generate dialogue content with corresponding role consistency constraints; and, The visual appearance information stored in the character consistency knowledge base is invoked to lock the visual consistency of the same character in different shots.
4. The multi-modal reinforcement learning based multi-role interactive video generation method of claim 3, wherein, The step of generating multi-character dialogue scripts and dynamic storyboard sequences based on the dynamic narrative path further includes: Within the dynamic narrative path, obtain the core theme, target emotion, and list of involved characters for the current node; Based on the character list, retrieve the historical dialogue style and voice parameters of each character from the character consistency knowledge base; Based on the core theme, target emotion, character list, and corresponding historical dialogue styles, prompt words are generated and input into a pre-tuned large language model to obtain at least two rounds of dialogue content as the multi-character dialogue script.
5. The multi-modal reinforcement learning based multi-role interactive video generation method of claim 1, wherein, The step of generating multi-character dialogue scripts and dynamic storyboard sequences based on the dynamic narrative path further includes: Analyze the emotion tags and character action descriptions in the multi-role dialogue script; Load a predefined shot rule library, which defines the mapping relationship between emotion tags and shot types and camera movement methods; Input the emotional intensity of the current dialogue, the positional relationship of the characters, and the scene complexity into the reinforcement learning model, optimize the mapping relationship, and output the shot switching sequence.
6. The multi-role interactive video generation method based on multimodal reinforcement learning as described in claim 1, characterized in that, The step of responding to user interactive editing commands on the dynamic storyboard sequence or the visual material, adjusting the dynamic narrative path or the dynamic storyboard sequence in real time, and generating adjusted video parameters includes: In response to the user's drag operation on the selected shot, the order of the dynamic shot sequence is reconstructed, and the multimodal video generation pipeline is triggered to perform regeneration and transition frame interpolation operations on the visual materials of the selected shot and its adjacent shots. In response to the user's command to switch the style of the selected shot, visual materials of the corresponding style are regenerated based on the diffusion model.
7. The multi-role interactive video generation method based on multimodal reinforcement learning as described in claim 1, characterized in that, The step of inputting the dynamic storyboard sequence and the multi-role dialogue script into the multimodal video generation pipeline to generate visual and audio materials includes: Based on the text-based diffusion model, keyframe images are generated according to the shot descriptions and character appearance information in the dynamic storyboard sequence. Based on the text-to-speech model, the dialogue audio of each character is generated according to the lines and voice parameters of the multi-role dialogue script, and each audio segment is marked with a corresponding timestamp and character ID. Based on the timestamp and character ID, align the dialogue audio with the character's lip-sync region in the keyframe image.
8. An electronic device, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the multi-role interactive video generation method based on multimodal reinforcement learning as described in any one of claims 1 to 7.
9. A readable storage medium, characterized in that, The readable storage medium is a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the multi-role interactive video generation method based on multimodal reinforcement learning as described in any one of claims 1 to 7.
10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the multi-role interactive video generation method based on multimodal reinforcement learning as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Video generation method and device, computer equipment and storage medium
CN121126081A