Video generation method and apparatus, electronic device, storage medium, and program product

CN122845883APending Publication Date: 2026-09-29BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610817762.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-08
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0003]但是,目前基于视频生成模型所生成的视频存在一些缺陷,比如,未实现对多个镜头各自对应的音视频片段的独立生成控制,导致无法在镜头粒度下进行精细化的视频生成,导致出现跨镜头音画不同步、跨镜头剧情混串、分镜难管控等问题

Benefits of technology

本公开提出视频生成方法、装置、电子设备、存储介质及程序产品,该视频生成方法能够通过目标脚本中携带的各镜头专属镜头约束信息,为每个镜头的音视频生成过程提供精准的生成方向指引,避免不同镜头的生成信息互相干扰,解决了多镜头视频生成中剧情割裂、信息混乱的问题,保障多个镜头生成后能够稳定联合表达完整统一的目标剧情;同时,通过同步音视频跨模态联合生成的模式,能够让音频和视频信息在生成过程中同步对齐,避免出现音画不同步、内容不匹配的问题,有效提升了生成出的目标视频的内容一致性与生成质量,也能够更好地满足多镜头视频的生成需求,解决了跨镜头音画不同步、跨镜头剧情混串、分镜难管控等诸多问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122845883A_ABST
    Figure CN122845883A_ABST
Patent Text Reader

Abstract

The present disclosure relates to a video generation method, device, electronic equipment, storage medium and program product, the method comprising: obtaining a target script, the target script comprising shot layer information, the shot layer information being used to indicate shot constraint information corresponding to each of a plurality of shots of a target video, the plurality of shots of the target video being used to jointly express a plot corresponding to the target script; based on the target script, controlling a video generation model to generate audio-video segments corresponding to each of the plurality of shots under the constraint of the shot constraint information corresponding to each of the plurality of shots, and obtaining the target video based on each of the audio-video segments; wherein the video generation model is a model for generating audio-video segments in a mode of synchronous audio-video cross-modal joint. The present disclosure provides precise generation direction guidance for the audio-video generation process of each shot, and avoids interference between the generation information of different shots.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to a video generation method, apparatus, electronic device, storage medium, and program product. Background Technology

[0002] With the development of artificial intelligence technology, video generation models have become increasingly powerful, and they now have the ability to automatically generate complete videos with audio based on input text information.

[0003] However, the videos generated by the current video generation model have some shortcomings. For example, they do not achieve independent generation control of the audio and video segments corresponding to multiple shots, which makes it impossible to perform fine-grained video generation at the shot level. This leads to problems such as audio and video desynchronization across shots, mixed plots across shots, and difficulty in managing shot composition. Summary of the Invention

[0004] This disclosure provides a video generation method, apparatus, electronic device, storage medium, and program product to solve at least one of the aforementioned technical problems. The technical solution of this disclosure is as follows: According to a first aspect of the present disclosure, a video generation method is provided, comprising: Obtain the target script, which includes shot layer information. The shot layer information is used to indicate the shot constraint information corresponding to each of the multiple shots of the target video. The multiple shots of the target video are used to jointly express the plot corresponding to the target script. Based on the target script, the video generation model is controlled to generate audio and video segments corresponding to each of the multiple shots under the constraints of the shot constraint information corresponding to each of the multiple shots, and the target video is obtained based on each of the audio and video segments; The video generation model is a model that generates audio and video segments through a synchronous audio and video cross-modal joint mode.

[0005] In one exemplary implementation, the video generation model is a model that performs simultaneous audio and video cross-modal joint generation based on an attention mechanism; The step of controlling the video generation model based on the target script to generate audio and video segments corresponding to each of the multiple shots under the constraints of the shot constraint information corresponding to each of the multiple shots includes: In the target script, the lens constraint information corresponding to each of the lenses is extracted; Based on the lens constraint information, attention boundary constraint information is generated. The attention boundary constraint information is used to constrain the generation of audio and video segments corresponding to the target lens during the generation process. The attention mechanism only selects the lens constraint information corresponding to the target lens to generate the corresponding audio and video segments. The target lens is any one of the multiple lenses. By inputting the attention boundary constraint information into the video generation model, audio and video segments corresponding to each of the multiple shots are generated.

[0006] In one exemplary implementation, the shot constraint information includes the corresponding shot's time constraint and plot constraint; The time constraints include at least one of the following: the time interval corresponding to the shot, the dialogue time interval in the shot, and the time interval in which the character exists within the shot. The plot constraints are used to indicate the plot expressed by the audio and video segments corresponding to the corresponding shots.

[0007] In one exemplary implementation, the step of inputting the attention boundary constraint information into the video generation model to generate audio and video segments corresponding to each of the multiple shots includes: The boundary-aware routing matrix generated based on the attention boundary constraint information is input into the text attention layer of the video generation model, so that the video generation model generates audio and video segments corresponding to the multiple shots. The boundary-aware routing matrix is ​​used to control the visibility of different text terms to different audio and video latent variable terms.

[0008] In one exemplary embodiment, the target script further includes global layer information, which is used to control the overall information of the target video; The step of controlling the video generation model based on the target script to generate audio and video segments corresponding to each of the multiple shots under the constraints of the shot constraint information corresponding to each of the multiple shots includes: Global constraint information is generated based on the global layer information, and the global constraint information is applied to the generation process of audio and video segments corresponding to all shots of the target video. By inputting the overall constraint information into the video generation model, audio and video segments corresponding to each of the multiple shots are generated.

[0009] In one exemplary embodiment, the target script further includes role layer information, which is used to constrain the role characteristics corresponding to a single role; The step of controlling the video generation model based on the target script to generate audio and video segments corresponding to each of the multiple shots under the constraints of the shot constraint information corresponding to each of the multiple shots includes: In the target script, extract the role identifier information corresponding to each role; Based on the character identification information, attention character constraint information is generated. The attention character constraint information is used to constrain the generation of audio and video clips of the same character in different shots. The attention mechanism generates corresponding audio and video clips based on the character characteristics of the same character. By inputting the attention role constraint information into the video generation model, audio and video segments corresponding to each of the multiple shots are generated.

[0010] In one exemplary implementation, the step of inputting the attention role constraint information into the video generation model to generate audio and video segments corresponding to each of the multiple shots includes: By inputting a character mask matrix generated based on the attention character constraint information into at least one attention layer of the video generation model, the video generation model is controlled to generate audio and video clips of the same character corresponding to different shots. The character mask matrix is ​​used to control the character features used in the audio and video generation process.

[0011] In one exemplary embodiment, the at least one attention layer includes at least one of the following: a video self-attention layer, an audio self-attention layer, a text attention layer, and an audio-video cross-attention layer; The process of inputting the character mask matrix generated based on the attention character constraint information into multiple attention layers of the video generation model to control the generation of audio and video clips of the same character corresponding to different shots includes performing at least one of the following operations: By inputting the character mask matrix into the video self-attention layer, the interaction between the corresponding video latent variables and the corresponding character features is controlled; By inputting the character mask matrix into the audio self-attention layer, the interaction between the corresponding audio latent variables and the corresponding character features is controlled; By inputting the character mask matrix into the text attention layer, the interaction between the corresponding text words and the corresponding character features is controlled; By inputting the character mask matrix into the audio-video cross-attention layer, the interaction between the audio latent variables and video latent variables for the same character is controlled.

[0012] In one exemplary embodiment, the method further includes: acquiring media reference information, the media reference information including at least one of the following: a reference image corresponding to the character, a reference sound corresponding to the character, and a posture reference corresponding to the character; The step of inputting the attention role constraint information into the video generation model to generate audio and video segments corresponding to each of the multiple shots includes: The media reference information is input into the video generation model to obtain the character features corresponding to the character; By inputting the attention role constraint information into the video generation model, the video generation model generates audio and video clips corresponding to the role in multiple shots based on the role's corresponding role features.

[0013] In one exemplary implementation, obtaining the target script includes: Obtain text prompt information, which is used to indicate the basic plot of the target video to be generated; The text prompt information is input into the script generation agent for hierarchical script generation processing to obtain the target script, which also includes global layer information and role layer information.

[0014] In one exemplary embodiment, the script generation agent includes: a verification agent, a structure processing agent, an identity alignment agent, and a narrative refinement agent; The step of inputting the text prompt information into a script generation agent for hierarchical script generation processing to obtain the target script includes: The text prompt information is input into the verification agent to perform common sense contradiction detection; After correcting the common-sense contradictions in the text prompt information, the corrected text prompt information is input into the structured processing agent for structured processing of global layer information and camera layer information to obtain the first processing result. The first processing result is input into the identity-aligned agent to perform structured processing of role-layer information, resulting in the second processing result. The second processing result is input into the narrative refinement agent to supplement the details of the global layer information, the shot layer information, and the character layer information, and the target script is output.

[0015] According to a second aspect of the present disclosure, a video generation apparatus is provided, comprising: The script acquisition module is configured to acquire a target script, which includes shot layer information. The shot layer information is used to indicate the shot constraint information corresponding to each of the multiple shots of the target video. The multiple shots of the target video are used to jointly express the plot corresponding to the target script. The video generation module is configured to execute a video generation model based on the target script, under the constraints of the shot constraint information corresponding to each of the multiple shots, to generate audio and video segments corresponding to each of the multiple shots, and to obtain the target video based on each of the audio and video segments; The video generation model is a model that generates audio and video segments through a synchronous audio and video cross-modal joint mode.

[0016] In one exemplary implementation, the video generation model is a model that performs simultaneous audio and video cross-modal joint generation based on an attention mechanism; The video generation module is configured to execute: In the target script, the lens constraint information corresponding to each of the lenses is extracted; Based on the lens constraint information, attention boundary constraint information is generated. The attention boundary constraint information is used to constrain the generation of audio and video segments corresponding to the target lens during the generation process. The attention mechanism only selects the lens constraint information corresponding to the target lens to generate the corresponding audio and video segments. The target lens is any one of the multiple lenses. By inputting the attention boundary constraint information into the video generation model, audio and video segments corresponding to each of the multiple shots are generated.

[0017] In one exemplary implementation, the shot constraint information includes the corresponding shot's time constraint and plot constraint; The time constraints include at least one of the following: the time interval corresponding to the shot, the dialogue time interval in the shot, and the time interval in which the character exists within the shot. The plot constraints are used to indicate the plot expressed by the audio and video segments corresponding to the corresponding shots.

[0018] In one exemplary implementation, the video generation module is configured to perform: The boundary-aware routing matrix generated based on the attention boundary constraint information is input into the text attention layer of the video generation model, so that the video generation model generates audio and video segments corresponding to the multiple shots. The boundary-aware routing matrix is ​​used to control the visibility of different text terms to different audio and video latent variable terms.

[0019] In one exemplary embodiment, the target script further includes global layer information, which is used to control the overall information of the target video; The video generation module is configured to execute: Global constraint information is generated based on the global layer information, and the global constraint information is applied to the generation process of audio and video segments corresponding to all shots of the target video. By inputting the overall constraint information into the video generation model, audio and video segments corresponding to each of the multiple shots are generated.

[0020] In one exemplary embodiment, the target script further includes role layer information, which is used to constrain the role characteristics corresponding to a single role; The video generation module is configured to execute: In the target script, extract the role identifier information corresponding to each role; Based on the character identification information, attention character constraint information is generated. The attention character constraint information is used to constrain the generation of audio and video clips of the same character in different shots. The attention mechanism generates corresponding audio and video clips based on the character characteristics of the same character. By inputting the attention role constraint information into the video generation model, audio and video segments corresponding to each of the multiple shots are generated.

[0021] In one exemplary implementation, the video generation module is configured to perform: By inputting a character mask matrix generated based on the attention character constraint information into at least one attention layer of the video generation model, the video generation model is controlled to generate audio and video clips of the same character corresponding to different shots. The character mask matrix is ​​used to control the character features used in the audio and video generation process.

[0022] In one exemplary embodiment, the at least one attention layer includes at least one of the following: a video self-attention layer, an audio self-attention layer, a text attention layer, and an audio-video cross-attention layer; the video generation module is configured to perform at least one of the following operations: By inputting the character mask matrix into the video self-attention layer, the interaction between the corresponding video latent variables and the corresponding character features is controlled; By inputting the character mask matrix into the audio self-attention layer, the interaction between the corresponding audio latent variables and the corresponding character features is controlled; By inputting the character mask matrix into the text attention layer, the interaction between the corresponding text words and the corresponding character features is controlled; By inputting the character mask matrix into the audio-video cross-attention layer, the interaction between the audio latent variables and video latent variables for the same character is controlled.

[0023] In one exemplary implementation, the video generation module is configured to perform: Obtain media reference information, which includes at least one of the following: a reference image corresponding to the character, a reference sound corresponding to the character, and a posture reference corresponding to the character; The media reference information is input into the video generation model to obtain the character features corresponding to the character; By inputting the attention role constraint information into the video generation model, the video generation model generates audio and video clips corresponding to the role in multiple shots based on the role's corresponding role features.

[0024] In one exemplary implementation, the script acquisition module is configured to execute: Obtain text prompt information, which is used to indicate the basic plot of the target video to be generated; The text prompt information is input into the script generation agent for hierarchical script generation processing to obtain the target script, which also includes global layer information and role layer information.

[0025] In one exemplary embodiment, the script generation agent includes: a verification agent, a structure processing agent, an identity alignment agent, and a narrative refinement agent; The script acquisition module is configured to execute: The text prompt information is input into the verification agent to perform common sense contradiction detection; After correcting the common-sense contradictions in the text prompt information, the corrected text prompt information is input into the structured processing agent for structured processing of global layer information and camera layer information to obtain the first processing result. The first processing result is input into the identity-aligned agent to perform structured processing of role-layer information, resulting in the second processing result. The second processing result is input into the narrative refinement agent to supplement the details of the global layer information, the shot layer information, and the character layer information, and the target script is output.

[0026] According to a third aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute the instructions to implement the video generation method as described above.

[0027] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided that, when instructions in the computer-readable storage medium are executed by a processor of an electronic device, enables the electronic device to perform the video generation method as described above.

[0028] According to a fifth aspect of the present disclosure, a computer program product is provided, the computer program product including a computer program stored in a readable storage medium, wherein at least one processor of a computer device reads from the readable storage medium and executes the computer program, causing the device to perform the video generation method described above.

[0029] The technical solutions provided by the embodiments of this disclosure have at least the following beneficial effects: This disclosure proposes a video generation method, apparatus, electronic device, storage medium, and program product. The video generation method can provide precise generation direction guidance for the audio and video generation process of each shot through the shot-specific constraint information carried in the target script, avoiding mutual interference between the generation information of different shots. This solves the problems of plot fragmentation and information confusion in multi-shot video generation, ensuring that multiple shots can stably and jointly express a complete and unified target plot after generation. At the same time, through the synchronous audio and video cross-modal joint generation mode, audio and video information can be synchronously aligned during the generation process, avoiding problems such as audio and video desynchronization and content mismatch. This effectively improves the content consistency and generation quality of the generated target video, and can better meet the generation needs of multi-shot videos, solving many problems such as cross-shot audio and video desynchronization, cross-shot plot mixing, and difficulty in controlling scene division.

[0030] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0031] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.

[0032] Figure 1 This is a schematic diagram of an implementation environment according to an exemplary embodiment.

[0033] Figure 2 This is a flowchart illustrating a video generation method according to an exemplary embodiment.

[0034] Figure 3 This is a schematic diagram of a video generation process according to an exemplary embodiment. Figure 1 .

[0035] Figure 4 This is a schematic diagram of a video generation process according to an exemplary embodiment. Figure 2 .

[0036] Figure 5 This is a schematic diagram of a video generation process according to an exemplary embodiment. Figure 3 .

[0037] Figure 6 This is a schematic diagram of a video generation scheme framework according to an exemplary embodiment.

[0038] Figure 7 This is a block diagram of a video generation apparatus according to an exemplary embodiment.

[0039] Figure 8 This is a block diagram illustrating an electronic device for video generation according to an exemplary embodiment.

[0040] Figure 9 This is another block diagram of an electronic device for video generation according to an exemplary embodiment. Detailed Implementation

[0041] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0042] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0043] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data used for analysis, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties.

[0044] With the development of artificial intelligence technology, there are currently two main video generation solutions: Option 1: Cascaded generation (generate video first, then add audio) Implementation logic: First, use a video generation model similar to Wensheng Video to create multi-camera footage. Then, use a V2A model to generate audio based on the finished video. V2A stands for Video-to-Audio, and the core task of this type of model is to generate corresponding audio that matches the scene, actions, and emotions of the video based on the input video content.

[0045] The solution has several drawbacks: audio and video belong to two separate models and there is no unified temporal constraint, which can lead to various problems such as misalignment of lip movements and dialogue, interruption of sound effects at camera transitions, inability to constrain dialogue during specific time periods, and easy crosstalk between multiple characters. Option 2: Global prompt word joint generation (a single sentence controls the generation of the entire video) Implementation logic: The video generation model is based on a single global prompt word, which controls the entire generation process. The solution has the following drawbacks: there is no definition of shot boundaries, semantic leakage between shots, inability to precisely control the generation process of a single shot, and the inability to limit the time interval of shots and dialogue within shots. It is evident that existing technologies cannot achieve independent generation control of audio and video segments corresponding to multiple shots in a synchronized audio-video joint generation mode. Therefore, this disclosure provides a video generation scheme that can perform independent audio-video joint generation for each shot based on the shot constraint information corresponding to multiple shots. This addresses both the problems caused by Solution 1 through the synchronized audio-video joint generation mode and the problems of Solution 2 by supporting fine-grained generation control of audio and video segments at the shot level.

[0046] The technical solution provided in this disclosure will be described in detail below: Please see Figure 1 The illustration shows an implementation environment provided by an embodiment of the present disclosure. The implementation environment may include at least one video generation terminal 110 and an interactive server 120, wherein the video generation terminal 110 and the interactive server 120 can communicate with each other via a network.

[0047] Specifically, the video generation terminal 110 interacts with the user through interaction with the interaction server 120. Specifically, the video generation terminal 110 can upload a target script to the interaction server 120, and the interaction server 120 can obtain the target script. The target script includes shot layer information, which indicates the shot constraint information corresponding to each shot in the target video. The multiple shots in the target video are used to jointly express the plot corresponding to the target script. Based on the target script, the video generation model is controlled to generate audio and video segments corresponding to each of the multiple shots under the constraints of the shot constraint information. Based on each audio and video segment, the target video is obtained. The video generation model is a model that generates audio and video segments through a synchronous audio and video cross-modal joint mode.

[0048] After the interactive server 120 completes the generation of the target video, it will return the generation result to the video generation terminal 110, where users can view and export the target video that meets their needs.

[0049] The interactive server 120 can deploy a video generation model. Alternatively, some of the inference capabilities of the video generation model can be deployed on the video generation terminal 110, completing the generation process through edge-cloud collaboration and reducing the computing power pressure on different devices.

[0050] The video generation terminal 110 can communicate with the interactive server 120 based on a browser / server (B / S) mode or a client / server (C / S) mode. The video generation terminal 110 may include physical devices such as smartphones, tablets, laptops, digital assistants, smart wearable devices, in-vehicle terminals, and servers, and may also include software running on the physical device, such as applications. The operating system running on the video generation terminal 110 in this embodiment may include, but is not limited to, Android, iOS, Linux, and Windows.

[0051] The interactive server 120 and the video generation terminal 110 can establish a communication connection via wired or wireless means. The interactive server 120 may include a stand-alone server, a distributed server, or a server cluster consisting of multiple servers, wherein the server may be a cloud server.

[0052] Please refer to Figure 2 The diagram illustrates a flowchart of a video generation method in an exemplary embodiment of this disclosure. The execution entity of this method can be the aforementioned video generation terminal, interactive server, or a combination of both. Please refer to the following for details. Figure 2 The method includes: S210. Obtain the target script, which includes shot layer information. The shot layer information is used to indicate the shot constraint information corresponding to each of the multiple shots of the target video. The multiple shots of the target video are used to jointly express the plot corresponding to the target script.

[0053] The target script records structured, hierarchical script information. It is not a collection of fragmented and vague creative requirements, but rather fine-grained information that provides a precise basis for the generation and control of independent audio-visual segments from multiple shots. Therefore, it differs from the single, complete sentence-based global prompts of related technologies, such as Solution 2. Because the shot-level information in this target script does not affect the entire system, but rather provides corresponding constraints for each independent shot, the generation of each shot only needs to follow the constraints set within its own shot layer; the constraints set for other shots within their shot layers have no effect on it.

[0054] The lens layer information matches exclusive constraints to each independent lens, defining the boundaries of different lenses from the input level and preventing the generation process of different lenses from interfering with each other.

[0055] The shot constraint information carried by the shot layer information can contain a wealth of content depending on the specific generation requirements. For example, it can include shot duration, shot type, camera movement, shooting angle, main content of the scene, details of the main action, composition requirements, lighting and color style, background environment information, transition effect requirements, the type of sound effects to be matched, voice-over content, subtitle text content, and subtitle presentation style. These dimensions of information can all be included as independent constraints in the shot layer information, providing a comprehensive and precise control basis for the generation of a single shot, without affecting other shots.

[0056] In this disclosure, multiple shots of the target video are used to jointly express the plot corresponding to the target script. Instead of generating multiple independent and unrelated video segments and forcibly splicing them together, all shots are organized with the core objective of completing a complete and unified target plot. The constraints on all shot-level information serve the expression of the overall plot, ensuring the independence of each shot while maintaining the integrity of the plot. Unlike traditional methods of generating and splicing segmented videos, in this disclosure, all shots are generated based on the same complete target script. The global plot logic has already been outlined during the script construction phase. The constraints of each shot build upon the content of the previous shot and initiate the progression of the next shot. This achieves precise control over shot segmentation while ensuring the overall plot's coherence and unity.

[0057] S220. Based on the target script, the video generation model is controlled to generate audio and video segments corresponding to each of the multiple shots under the constraints of the shot constraint information corresponding to each of the multiple shots, and the target video is obtained based on each of the audio and video segments; wherein, the video generation model is a model that generates audio and video segments through a synchronous audio and video cross-modal joint mode.

[0058] Each shot constraint information only applies to the corresponding single shot and will not cause additional interference to the generation of other shots. At the same time, it naturally follows the overall plot progression logic. This allows the generation of audio and video segments corresponding to each of the multiple shots under the constraints of the shot constraint information of each of the multiple shots. Based on each of the audio and video segments, the resulting target video has both precise customization of the shot breakdown and maintains the smooth continuity of the overall plot.

[0059] The synchronous audio-video cross-modal joint generation model differs from the models used in traditional cascade generation schemes. It does not follow the logic of "video first, audio later" but instead generates audio and video content with corresponding relationships synchronously. It jointly learns and generates audio and video features, ensuring the temporal alignment and content matching of audio and video from the initial stage of generation, thus avoiding audio-video adaptation problems.

[0060] The process of obtaining the target video based on the aforementioned audio and video segments ensures that the overall narrative logic remains consistent and smooth. This disclosure does not limit the specific operations performed; for example, transitions can be made between adjacent audio and video segments: the ending scene of the previous shot and the beginning scene of the next shot will be matched with appropriate transition effects. For instance, a fade-in / fade-out transition is used for a calm, everyday scene, while a hard-cut transition is used for a scene of intense conflict. Simultaneously, the volume transition of the audio in adjacent segments is adjusted to avoid sudden volume changes that could affect the viewing experience. After completing the single-segment transition, the basic parameters of the entire video can be uniformly calibrated to ensure consistency in resolution, frame rate, and color space.

[0061] The video generation model disclosed herein generates audio and video segments through a synchronous audio-video cross-modal joint model. It directly and synchronously generates matching audio content while generating video content, eliminating the need for two separate generation stages. This disclosure does not limit the video generation model; for example, it can be trained based on a Transformer architecture or a Diffusion Model, without posing an obstacle to implementation. For instance, it can be trained based on the model architecture of Scheme 2.

[0062] The Transformer architecture, proposed in 2017, is a fundamental deep learning architecture that relies on a self-attention mechanism to model the contextual dependencies of sequential data, efficiently capturing the relationships between different input units. Transformers support parallel computation, offering superior training speed and the ability to capture long-range dependencies, making them suitable for processing multimodal long-sequence data such as audio and video. The Diffusion Model is a widely used generative model architecture in the current video generation field. Its core logic is to generate target content that conforms to the data distribution by progressively removing random noise. Originally used primarily for image generation, it has been extended to adapt to the generation of multimodal content such as video and audio.

[0063] The video generation method proposed in this disclosure can provide precise generation direction guidance for the audio and video generation process of each shot through the shot-specific constraint information carried in the target script, avoiding mutual interference between the generation information of different shots. This solves the problems of plot fragmentation and information confusion in multi-shot video generation, and ensures that multiple shots can stably and jointly express a complete and unified target plot after generation. At the same time, through the synchronous audio and video cross-modal joint generation mode, audio and video information can be synchronously aligned during the generation process, avoiding problems such as audio and video desynchronization and content mismatch. This effectively improves the content consistency and generation quality of the generated target video, and can better meet the generation needs of multi-shot videos, solving many problems such as cross-shot audio and video desynchronization, cross-shot plot mixing, and difficulty in controlling shot division.

[0064] In one exemplary implementation, please refer to Figure 3 It illustrates the video generation process of this disclosure. Figure 1 The video generation model is a model that performs synchronous audio and video cross-modal joint generation based on an attention mechanism; the step of controlling the video generation model to generate audio and video segments corresponding to each of the multiple shots under the constraints of the shot constraint information corresponding to each of the multiple shots, based on the target script, includes: S310. Extract the lens constraint information corresponding to each of the lenses from the target script.

[0065] In practice, the target script can be structured and parsed to obtain each shot unit, and then the constraint information of the corresponding dimension can be extracted from each shot unit. For example, if it is the image dimension, the scene type, main character, action requirements, lighting style, etc., which are clearly marked in the script can be extracted. For example, if the script is marked "In the early morning forest, a fox runs quickly out from behind a tree, warm golden sidelight, natural realistic style", this description will be converted into the image constraint features corresponding to that shot.

[0066] In the audio dimension, constraint features corresponding to information such as ambient sounds, dialogue content, sound effect style, and background music mood marked in the script can be extracted. For example, the above shot is marked with "the rustling sound of a fox stepping on fallen leaves, birds chirping in the distance, and light and natural background music", which can be converted into the audio constraint features corresponding to that shot.

[0067] If it's the time dimension of the shot, you can extract the single shot duration, shot entry and exit times, and time rhythm requirements of shot movement that are clearly marked in the script. For example, if the shot duration is marked as 2.5 seconds, the cut occurs at the 8th second after the opening, and the zoom movement from the panoramic view to the fox's face is completed in 0.5 seconds, these values ​​will be organized into the corresponding time constraint features.

[0068] All extracted constraint features can be combined as the constraint basis for generating the audio and video clip of that shot, i.e., the corresponding shot constraint information.

[0069] S320. Based on the lens constraint information, generate attention boundary constraint information. The attention boundary constraint information is used to constrain the generation of audio and video segments corresponding to the target lens during the generation process. The attention mechanism only selects the lens constraint information corresponding to the target lens to generate the corresponding audio and video segments. The target lens is any one of the multiple lenses.

[0070] The purpose of attention boundary constraint information is to define the scope of the model's attention calculation, so that when the model generates an audio and video clip of a target shot, it only focuses on the shot constraint information of that shot itself and will not mistakenly read the shot constraint information of other shots, thereby avoiding the mutual interference of semantic information of different shots at the model calculation level.

[0071] In one specific implementation, an attention mask matrix can be generated. This mask matrix is ​​equivalent to adding an "access lock" to the shot constraint information of each shot. When generating audio and video data for the target shot, the mask matrix allows the shot constraint information corresponding to the current target shot to participate in the attention calculation, while the shot constraint information of other shots does not participate in the attention calculation.

[0072] For example, when generating the third shot, the attention mask will only open the attention weights of the shot constraint information of the third shot. The shot constraint information of the first, second and fourth shots will not be read by the model's attention mechanism. This isolates the shot constraint information of different shots from the root of the computational logic, ensuring that the generation process of each shot only follows its own requirements and there will be no problem of semantic leakage and crosstalk between shots.

[0073] S330. By inputting the attention boundary constraint information into the video generation model, audio and video segments corresponding to each of the multiple shots are generated.

[0074] Attention boundary constraint information is added to the model's attention calculation process as an explicit hard boundary constraint. Unlike the implicit constraint method that relies solely on the model to learn and summarize the shot boundaries, it directly divides the scope of the constraint information of different shots at the computational logic level. This avoids the cross-interference of constraint information of different shots from the root, achieving precise control over the content generated by each shot, while maintaining the logical coherence of the overall video content by relying on a unified global computational framework.

[0075] In one exemplary implementation, the shot constraint information includes the corresponding shot's time constraint and plot constraint; The time constraints include at least one of the following: the time interval corresponding to the shot, the dialogue time interval in the shot, and the time interval in which the character exists within the shot. The plot constraints are used to indicate the plot expressed by the audio and video segments corresponding to the corresponding shots.

[0076] Time constraints are used to define the time range of shots and various elements within shots, providing precise time dimension control for the generation process and avoiding issues such as inconsistent durations and misaligned audio-visual timing. The time interval corresponding to a shot refers to the total duration of a single shot, as well as the start and end times of that shot within the entire target video. For example, if the total duration of the target video is 120 seconds, a shot introducing the main character would have a time interval of "from the 15th to the 22nd second of the total video duration, with a single shot duration of 7 seconds."

[0077] The dialogue time interval within a shot refers to the start and end points of the dialogue content within that shot. For example, if the protagonist's line in a shot is "The weather is really nice today," the corresponding dialogue time interval is from the 2nd to the 4th second after the shot begins playing. The character's presence time interval refers to the range of time a specific character appears in the frame of that shot. For example, if the protagonist appears in the frame for the first 5 seconds and then leaves the frame after the 5th second, the corresponding character's presence time interval is from 0 to 5 seconds after the shot begins playing.

[0078] Plot constraints are the core content requirements for each shot, used to clarify the plot expression task that the shot needs to accomplish. For example, the plot constraint for the shot of the protagonist's appearance mentioned above can be defined as "the protagonist walks in the park, greets the audience, and introduces the theme of this trip."

[0079] This approach, which breaks down shot constraints into time constraints and narrative constraints, allows for multi-dimensional and refined control over the generated content, meeting the needs of generating complex narrative videos. Time constraints provide the model with a clear generation benchmark from a temporal perspective. Whether it's the overall shot duration, audio-visual timing alignment, or the order in which elements appear, there are clear guidelines to follow. This prevents issues such as shot duration discrepancies, mismatched dialogue and lip-syncing, and disordered element appearance due to ambiguous timeframes, significantly improving the controllability of the generated results. Narrative constraints, on the other hand, clarify the core task of each shot at the content level, ensuring that the generated content of each shot revolves around its corresponding storyline. This complements the time constraints, providing a clear and accurate control direction for model generation.

[0080] This setup, combined with the attention boundary constraint mechanism, further enhances the accuracy of shot generation and avoids crosstalk between different shots in terms of time and content information. Each shot only needs to be generated according to its own time and plot constraints, ensuring the accuracy of individual shot content without interfering with the generation process of other shots. Simultaneously, since all constraints are contained within a unified target script, the time settings and plot content of all shots are already organized around the overall target plot. The time intervals of different shots connect naturally, and the plot logic progresses layer by layer, avoiding the overall logical break caused by shot control issues. This balances precise control of individual shots with the overall narrative coherence of the video.

[0081] In one exemplary implementation, the step of inputting the attention boundary constraint information into the video generation model to generate audio and video segments corresponding to each of the multiple shots includes: The boundary-aware routing matrix generated based on the attention boundary constraint information is input into the text attention layer of the video generation model, so that the video generation model generates audio and video segments corresponding to the multiple shots. The boundary-aware routing matrix is ​​used to control the visibility of different text terms to different audio and video latent variable terms.

[0082] The boundary-aware routing matrix is ​​a computable matrix generated based on attention boundary constraint information. Essentially, it's a mapping table used to control access permissions between text terms and audio / video latent variable terms. Text terms are the smallest semantic units obtained by encoding the shot constraint information extracted from the script. Audio / video latent variable terms are the basic units in the latent space used to represent the features of audio / video frame segments during the video generation model's audio / video content generation process. The text attention layer is the core computational layer in the video generation model responsible for establishing the association between text constraints and audio / video generation features. This implementation converts the attention boundary constraint information into a boundary-aware routing matrix that the model's text attention layer can directly recognize and call. This matrix directly intervenes in the attention layer's computation process, ensuring that the association between text constraints and audio / video generation strictly follows the shot boundaries.

[0083] This setup achieves hard boundary partitioning directly through explicit control of the routing matrix, avoiding crosstalk between text constraints from different shots at the computational root. This significantly improves the accuracy of the content generated for each shot, preventing crosstalk errors such as "the subject of the previous shot appearing in the next shot" or "the style of the next shot appearing prematurely in the previous content." It doesn't require modifying the overall architecture of the video generation model; the functionality can be achieved simply by injecting a boundary-aware routing matrix into the text attention layer. It is compatible with various commonly used generative model architectures such as Transformer and diffusion models, making it easy to modify and highly compatible.

[0084] In one exemplary embodiment, the target script further includes global layer information for controlling the overall information of the target video.

[0085] Global layer information refers to the information in the hierarchical script of the target script that provides overall guidance for the creation of the target video. For example, it may include information indicating the creative direction, style tone, and macro scene layout. It may be used to constrain: a unified narrative theme, a consistent visual style, and a coherent emotional rhythm.

[0086] For example, if a creator wants to create a travel video of a "springtime city stroll," the global layer information could be set to "an overall warm-toned film style, a relaxed and soothing narrative pace, and a core expression of the healing atmosphere of springtime in the city." Here, the consistent warm-toned film style is the overall visual information that needs to be maintained, while the relaxed and soothing pace and the core healing atmosphere are the overall content information that needs to be maintained. Without the constraints of global layer information, even if each shot is precisely controlled, it is easy to have individual shots that are of acceptable quality, but the overall tone is fragmented: for example, the beginning is a warm-toned spring scene, but the middle suddenly changes to a cold and rainy style, or the first half has a relaxed pace and the second half suddenly cuts into a fast-paced rhythm, which destroys the overall unity of the video. The existence of global layer information, on the other hand, defines a unified creative framework for all shots from the top level, ensuring that the generation of all shots revolves around the overall goal.

[0087] Global layer information and attention boundary constraints of shot composition are complementary and mutually reinforcing. Attention boundary constraints ensure that the customized content of each shot does not interfere with each other, while global layer information ensures that the creation of all shots does not deviate from the overall direction. The two work together to ensure that the final video can not only meet the creator's precise customization requirements for each shot, but also maintain the overall tone unity and narrative coherence of the work.

[0088] Please refer to Figure 4 It illustrates the video generation process of this disclosure. Figure 2 The step of controlling the video generation model to generate audio and video segments corresponding to each of the multiple shots, based on the target script and under the constraints of the shot constraint information corresponding to each of the multiple shots, includes: S410. Generate overall constraint information based on the global layer information, and the overall constraint information is applied to the generation process of audio and video segments corresponding to all shots of the target video.

[0089] The overall constraint information participates in the feature calculation of all shots throughout the entire process. When generating its own audio and video content, each shot needs to simultaneously read the local constraints of the shot and the overall constraints of the global layer to ensure that each shot meets the requirements set globally. Unlike shot constraints, which only apply to a single shot, the scope of overall constraints covers all stages of the target video generation. It does not become invalid with the switching of shot boundaries, nor is it isolated by attention boundary constraints. It always serves as a unified top-level requirement to guide the entire generation process.

[0090] For example, when a creator wants to generate a graduation commemorative short video with a "retro campus" theme, after filling in information such as "the overall style adopts a 90s retro DV shooting style, the overall color tone is yellowish, the picture has a slight graininess, and the narrative tone is nostalgic and warm" at the global layer of the target script, the model will first convert this natural language description into a computationally comprehensible overall constraint code to obtain the final overall constraint information. When generating each shot, whether it is the opening empty shot of the school gate, the interactive shot between classes in the classroom, or the group photo shot on the playground at the end, this overall constraint will participate in the generation calculation of each shot. Each shot, while matching its own storyline and time requirements, will also align with the style and tone requirements of the overall constraint, ensuring that all shots conform to the retro DV visual style and nostalgic emotional tone.

[0091] S420. By inputting the overall constraint information into the video generation model, audio and video segments corresponding to each of the multiple shots are generated.

[0092] In this implementation, global layer information is converted into overall constraint information that takes effect throughout the entire process. Technically, this achieves a collaborative division of labor between global constraints and shot-by-shot constraints: overall constraints provide a unified creative framework, while shot-by-shot constraints provide precise control requirements for individual shots. The two do not conflict and work together in the generation process. Attention boundary constraints only isolate local constraints between different shots, without masking the overall constraints of the global layer. Therefore, it ensures that local information in shot-by-shots does not interfere with each other, while allowing overall constraints to take effect throughout all shots. This adapts to the control logic of hierarchical scripts, technically guaranteeing that the final generated video possesses both the precision of individual shots and the unity of the overall work.

[0093] In one exemplary implementation, the target script further includes role layer information, which is used to constrain the role characteristics corresponding to a single role.

[0094] When the generated multi-camera video needs to feature multiple different characters, the character layer information can set independent constraint boundaries for each character, avoiding interference between the features of different characters during the generation process. At the same time, it can ensure that the core features such as appearance and personality of the same character remain consistent in different shots, and there will be no problem of the same character's features being contradictory or style being inconsistent in different shots, which further improves the controllability of the generated video content.

[0095] Please refer to Figure 5 It illustrates the video generation process of this disclosure. Figure 3 The step of controlling the video generation model to generate audio and video segments corresponding to each of the multiple shots, based on the target script and under the constraints of the shot constraint information corresponding to each of the multiple shots, includes: S510. Extract the role identification information corresponding to each role from the target script.

[0096] Role identification information is a unique identifier extracted from the role layer information of the target script to distinguish different roles. Each role corresponds to an independent identifier and will not overlap with the identifiers of other roles. Therefore, role identification information can be regarded as the identity anchor of a role. S510 is essentially the construction of the identity anchor of a role.

[0097] For example, this step can sort out and count all the roles recorded in the role layer information of the target script. For each independent role identified, a unique and non-repeating identity code is assigned to it. At the same time, all the features corresponding to the role are bound to this unique code to complete the extraction of role identification information.

[0098] For example, when creating a short video featuring two main characters, the script character layer records "Protagonist A: A young woman around 20 years old, with shoulder-length black hair, usually wears a blue denim jacket, and has a lively and outgoing personality," and "Protagonist B: A young man around 25 years old, wears black-rimmed glasses, usually wears a gray cotton sweatshirt, and has a calm and reserved personality." When the model sorts out the character information, it first identifies the existence of two independent characters and assigns each character a unique code, such as ID001 and ID002. Then, it binds the characteristics of each character to the corresponding code, thus completing the extraction of the character identification information. This extraction process is equivalent to giving each character a unique identity tag, clearly distinguishing the characteristics of different characters, and laying the foundation for avoiding character feature crosstalk and maintaining consistency of characteristics for the same character in the subsequent generation process.

[0099] S520. Based on the character identification information, generate attention character constraint information. The attention character constraint information is used to constrain the generation of audio and video clips of the same character in different shots. The attention mechanism generates corresponding audio and video clips based on the character characteristics of the same character.

[0100] Attention role constraint information is essentially a constraint condition built upon independent role identifiers to regulate the attention calculation process. It applies to the same character appearing in different shots. Its core objective is to ensure that when generating content for any shot featuring that character, the model can consistently access the unique character features bound to the corresponding role identifier without feature shift. Unlike shot boundary constraints, which define information boundaries between different shots, attention role constraints define information boundaries between different characters. They constrain the model so that when generating content for a particular character in any shot, it only needs to read the features bound to the corresponding role identifier, without needing to read feature information from other characters.

[0101] S530. By inputting the attention role constraint information into the video generation model, audio and video segments corresponding to each of the multiple shots are generated.

[0102] For example, in the short video featuring protagonist A (ID001) and protagonist B (ID002) mentioned earlier, the first shot shows protagonist A browsing a bookstore alone, the second shot shows the two protagonists meeting in a coffee shop, and the third shot shows protagonist B walking alone on the street with an umbrella. When generating protagonist A in the first shot, only the traits associated with ID001—short hair, denim jacket, and lively personality—are used. When generating the two protagonists appearing simultaneously in the second shot, protagonist A's traits are used, while protagonist B's traits are used, and the two do not interfere with each other. When generating protagonist B in the third shot, only the traits associated with ID002—black-rimmed glasses, gray hoodie, and calm personality—are used.

[0103] This constraint mechanism avoids the mixing of features from other characters when generating the current character, preventing errors such as protagonist A suddenly growing black-rimmed glasses like protagonist B. It also ensures that the same character uses the same set of features regardless of which shot it appears in, avoiding inconsistencies such as protagonist A having shoulder-length short hair in the first shot and suddenly having long hair in the third shot. From the perspective of computational logic, it guarantees the consistency and independence of character features, further improving the controllability of multi-character video generation.

[0104] In one exemplary implementation, the step of inputting the attention role constraint information into the video generation model to generate audio and video segments corresponding to each of the multiple shots includes: By inputting a character mask matrix generated based on the attention character constraint information into at least one attention layer of the video generation model, the process of the video generation model generating audio and video clips of the same character corresponding to different shots is controlled. The character mask matrix is ​​used to control the character features used in the audio and video generation process.

[0105] In this implementation, a character mask matrix is ​​generated and input into the model's attention layer. Matrix operations on the mask matrix directly control the range of features invoked during attention calculations, ensuring that the generation calculation for each character only reads the features corresponding to that character. This achieves isolation and unification of character features from the underlying computational logic. Each position in the mask matrix corresponds to the access permission of a character feature. For the target character currently being generated, only the features bound to that character are granted access permission, allowing the model to invoke and calculate them. All other non-target character features are masked and prohibited from participating in the generation calculation of the current character, fundamentally eliminating the possibility of crosstalk between different character features.

[0106] In practice, a character mask matrix of corresponding dimensions can be constructed based on the already generated attention character constraint information and the list of characters appearing in the currently generated shot. The dimensions of the matrix will be aligned with the feature dimensions of the model's attention layer input. The positions of the target character features in the matrix will be set to valid values ​​that can be masked, while the positions of other irrelevant character features will be set to masked values.

[0107] The constructed character mask matrix is ​​input into the corresponding attention layer and plays a role in the attention calculation process: the attention scores of irrelevant character features that are masked are reset to extremely low values, so that only the features of the target character can retain normal attention weights and participate in content generation.

[0108] This implementation method is naturally compatible with the attention calculation framework of video generation models, requiring no large-scale modification to the overall model structure. It only requires inserting a masking operation into the attention calculation process to achieve constraint control of character features, resulting in low implementation costs and strong compatibility. By achieving isolation directly at the feature calculation level, it does not affect the normal calculation of other generation stages. It can accurately isolate the features of different characters without affecting the normal effectiveness of overall constraint information and shot constraint information. Therefore, overall constraint information, shot constraint information, and attention character constraint information can work synergistically. Specifically, the attention character constraint information ensures that during multi-character video generation, features of different characters do not interfere with each other, and the features of the same character remain consistent across shots, improving the content controllability of AI-generated videos.

[0109] In one exemplary embodiment, the at least one attention layer includes at least one of the following: a video self-attention layer, an audio self-attention layer, a text attention layer, and an audio-video cross-attention layer; The process of inputting the character mask matrix generated based on the attention character constraint information into multiple attention layers of the video generation model to control the generation of audio and video clips of the same character corresponding to different shots includes performing at least one of the following operations: (1) By inputting the character mask matrix into the video self-attention layer, the corresponding video latent variables are controlled to interact with the corresponding character features.

[0110] The video self-attention layer is a module in the video generation model specifically responsible for calculating the attention between video features. Its core function is to calculate the dependencies between different latent video variables, thereby generating video frame features that meet the requirements. This step is used for the visual computation stage of video generation, applying the character mask matrix to the video self-attention layer to constrain the invocation of character features during the video feature generation process.

[0111] During the generation of video content for the current shot, when the calculation progresses to the feature interaction stage of the video self-attention layer, a character mask matrix is ​​input into this layer and combined with the self-attention calculation process. The mask matrix opens the video feature interaction path of the target character being generated in the current shot, while blocking the video feature interaction paths of other non-target characters. Only the features of the target character can interact normally with the video latent variables to generate video frame features that conform to the characteristics of that character.

[0112] This step, by applying the character mask matrix to the video self-attention layer, constrains and controls the character features, directly ensuring the correctness of character features in the video content. Since the video self-attention layer directly determines the final output visual content, adding constraints here can prevent crosstalk or inconsistencies in character visual features from the source, directly affecting the final output, thus ensuring greater accuracy and effectiveness of the constraints.

[0113] (2) By inputting the character mask matrix into the audio self-attention layer, the corresponding audio latent variables are controlled to interact with the corresponding character features.

[0114] The audio self-attention layer is the module in the video generation model responsible for calculating the attention between audio features. Its core function is to sort out the dependencies between different audio latent variables and generate audio features that match the video. For example, the timbre and tone of voice corresponding to different characters need to be calculated here. This step applies role constraints to the audio generation calculation process by applying a role mask matrix to the audio self-attention layer. This constrains the invocation of audio features corresponding to the corresponding characters during the audio feature generation process, ensuring that the audio features of different characters meet the requirements.

[0115] During the generation of audio content for the current shot, when the calculation progresses to the feature interaction stage of the audio self-attention layer, a character mask matrix is ​​input into this layer and integrated into the self-attention calculation process. The mask matrix opens the interaction path of the audio features corresponding to the target character appearing in the current shot, and blocks the audio feature interaction path of all non-target characters. Only the audio features of the target character can interact normally with the audio latent variables, ultimately generating audio content that conforms to the characteristics of that character.

[0116] This step applies a character mask matrix to the audio self-attention layer, constraining and controlling the character's audio features at the core of audio generation. This effectively avoids voice mismatch issues when multiple characters are in dialogue—for example, the voice of protagonist B is output when protagonist A is supposed to speak. Simultaneously, it ensures that the voice and tone of the same character remain consistent across different shots, preventing a complete break in the character's vocal style. Combined with visual character constraints, this further enhances the overall consistency of multi-character video generation.

[0117] (3) By inputting the character mask matrix into the text attention layer, the corresponding text words and corresponding character features are controlled to interact.

[0118] The text attention layer is a core module in the video generation model responsible for the attention interaction between text features and various features generated during the model's process. Its role is to provide text-level guidance for video and audio generation. This operation applies role constraints at the text level by applying a role mask matrix to the text attention layer. This constrains the text lexical feature invocation process to only match features corresponding to the target role, avoiding confusion of role information at the text description level.

[0119] This operation avoids interference between text information from different characters, reducing the probability of errors in character characteristics.

[0120] (4) By inputting the character mask matrix into the audio-video cross-attention layer, the interaction between the audio latent variables and video latent variables for the same character is controlled.

[0121] The audio-video cross-attention layer is a core module in the video generation model responsible for connecting audio and video features and enabling their alignment and interaction. It ensures that the generation logic of audio content and video visuals matches each other, guaranteeing synchronous correspondence between sound and visual content. For example, when a character speaks, the audio content aligns with the corresponding character's lip movements. This operation applies character constraints during the audio-video feature cross-alignment process, applying a character mask matrix to the audio-video cross-attention layer to constrain the interaction and matching process of the audio and video latent variables corresponding to the same character, thus preventing cross-mismatches in audio and video features from different characters.

[0122] In practice, when the model calculation progresses to the feature interaction stage of audio and video cross-attention, the mask matrix will control the interaction range of audio and video features: only the audio latent variables and video latent variables corresponding to the same role will have open interaction paths, allowing the two to complete attention matching calculations. The audio and video feature interaction paths between different roles will be completely blocked, and the audio and video features of different roles will not be allowed to complete incorrect matching.

[0123] This operation implements cross-constraints at the character level during the audio-visual alignment process, preventing mismatch issues from the audio-visual matching level. For example, it avoids situations where protagonist A speaks in the video, but the corresponding audio is incorrectly bound to protagonist B's lip movements, or misalignment issues where sound and visual characters are not aligned. This constraint further completes the character control logic throughout the entire process. Combined with constraints from video self-attention, audio self-attention, and text attention, it achieves isolated control of character features across the entire generation chain, including text input, visual generation, audio generation, and audio-visual content. This comprehensively ensures character consistency in multi-character video generation and significantly reduces the probability of various character mismatch issues.

[0124] In one exemplary embodiment, the method further includes: acquiring media reference information, the media reference information including at least one of the following: a reference image corresponding to the character, a reference sound corresponding to the character, and a posture reference corresponding to the character. The media reference information is provided as a reference when generating audio or visual images related to a specific character, thereby controlling the appearance, voiceprint, posture, and other unique characteristics of a specific character in each shot, and improving the controllability of the characters in the generated content.

[0125] The step of inputting the attention role constraint information into the video generation model to generate audio and video segments corresponding to each of the multiple shots includes: The media reference information is input into the video generation model to obtain the character features corresponding to the character; by inputting the attention role constraint information into the video generation model, the video generation model generates audio and video clips corresponding to the character in multiple shots based on the character features corresponding to the character.

[0126] By inputting media reference information into the video generation model, the model extracts unique features specific to each character. For example, it extracts facial features and clothing style from reference images, timbre and voiceprint characteristics from reference audio, and signature movement habits from posture references. These features are integrated into a fixed set of unique features (character features) stored for use throughout the entire generation process. By managing these extracted unique features through attention-based character constraint information, whenever the model generates content for the corresponding character, it calls upon the pre-stored unique features, ensuring that the content generated for all shots of the character is based on the same set of unique features.

[0127] Media reference information provides a controllable source of character features, while attention-based character constraint information provides computational assurance for feature isolation. Together, they form a complete character control chain of custom feature input + full-process isolated generation, which not only meets the user's need to create their own exclusive characters, but also ensures from a computational perspective that there is no crosstalk when multiple characters coexist and that the same character is consistent across shots.

[0128] In one exemplary implementation, obtaining the target script includes: Obtain text prompt information, which is used to indicate the basic plot of the target video to be generated; The text prompt information is input into the script generation agent for hierarchical script generation processing to obtain the target script, which also includes global layer information and role layer information.

[0129] Users only need to input text prompts describing the basic plot; without manually breaking down the structured, layered information, the script generation agent can automatically complete plot expansion, storyboarding, and hierarchical script generation.

[0130] For example, the script generation agent first expands the complete global plot outline based on the basic plot text prompts input by the user, sorting out the core theme, story direction, and overall scene setting of the entire video. This content will be organized into the global layer information of the target script. Then, the agent will further break down the content corresponding to each shot to obtain the shot layer information, sort out the basic settings of all the characters, and clarify the characters corresponding to each shot. This content will be organized into the character layer information of the target script, and finally form a clearly structured and layered target script.

[0131] This automated hierarchical script generation method simplifies the user's creative process and provides a standardized hierarchical script. The global layer information of the script can provide unified global guidance for the entire video generation process, ensuring that the generation of all shots conforms to the overall plot setting and that there will be no plot deviation. The character layer information can ensure that there is no mixing of different shots, and the character layer ensures the consistency of the same character in different shots, thereby ultimately guaranteeing the quality of the generated video.

[0132] In one exemplary embodiment, the script generation agent includes: a verification agent, a structure processing agent, an identity alignment agent, and a narrative refinement agent; The step of inputting the text prompt information into a script generation agent for hierarchical script generation processing to obtain the target script includes: The text prompt information is input into the verification agent to perform common sense contradiction detection; After correcting the common-sense contradictions in the text prompt information, the corrected text prompt information is input into the structured processing agent for structured processing of global layer information and camera layer information to obtain the first processing result. The first processing result is input into the identity-aligned agent to perform structured processing of role-layer information, resulting in the second processing result. The second processing result is input into the narrative refinement agent to supplement the details of the global layer information, the shot layer information, and the character layer information, and the target script is output.

[0133] The script generation agent, verification agent, structure processing agent, identity alignment agent, and narrative refinement agent disclosed herein can all be trained based on a large language model and do not pose an obstacle to implementation.

[0134] This implementation provides a script generation architecture with multi-agent division of labor and cooperation. It can verify, sort and optimize the original input from different levels, improve the quality of the target script from the source. The agents with different functions perform their respective duties and complete the processing work of different stages. This avoids the logical omissions and information loss problems that are easy to occur when processing with a single model. It makes the entire hierarchical script generation process more standardized and the output script information more accurate, which can ensure the quality of hierarchical scripts from the root.

[0135] Please refer to Figure 6 The diagram illustrates a framework of a video generation scheme in an exemplary implementation.

[0136] Figure 6The entire technical process of video generation is presented according to a three-level architecture of input layer, processing layer and output layer. Free text prompts are input in the left link and media reference information can be input in the right link. The two links can introduce various information to the model in parallel to guide video generation.

[0137] Users can input two types of materials: one is free text prompts in natural language form, i.e., text prompts in this disclosure; the other is media reference information, which can include visual reference information, audio reference information, etc., covering various materials such as character reference images, reference voiceprints, and posture materials.

[0138] Free text prompts are fed into a multi-agent script generation pipeline. Relying on the division of labor among four types of agents—verification, structured processing, identity alignment, and narrative refinement—hierarchical target scripts are generated layer by layer. Shot constraint information is then extracted from the scripts. Media reference information, through role constraint information, influences the calculation process of the attention layer of the video generation model via the role mask matrix. This role constraint information is also obtained by parsing the target scripts.

[0139] By leveraging shot constraint information and character mask matrices, the video generation model can be collaboratively controlled to achieve independent generation of storyboards and unified character representation across shots. The shot constraint information output from the left link can be used to define the access boundaries of single-shot information through explicit mask matrices, avoiding semantic leakage and plot interference between shots. The character mask matrices generated by the right link are then fed into four levels: video self-attention, audio self-attention, text attention, and audio-video cross-attention. These levels respectively manage the interaction processes of four types of features: image features, audio voiceprints, text semantics, and audio-visual pairing, achieving feature isolation for different characters and feature unification for the same character across shots.

[0140] The output layer receives the latent audio and video variables generated by the model and relies on the VAE decoder to complete feature restoration and output the final video clip. The VAE decoder refers to the feature decoding module based on the variational autoencoder architecture, which is responsible for restoring the latent variables output by the model into high-resolution video frames and corresponding audio signals that meet the requirements, and finally integrating and outputting a complete audio and video clip.

[0141] The video generation model in this scheme employs a dual-tower diffusion Transformer for joint audio and video generation. This dual-tower structure includes a video tower and an audio tower, used to model latent video and audio variables, respectively. Each video and audio tower can include a self-attention layer, a hierarchical text attention layer, an audio-video cross-attention layer, and a feedforward network layer. Specifically, the self-attention layer extracts the internal latent variables of the audio and video modalities, the hierarchical text attention layer injects narrative text information, and the audio-video cross-attention layer achieves feature alignment between the audio and video modalities.

[0142] During the training phase, a diffusion model or flow matching objective can be used for optimization, enabling the video tower and audio tower to learn the generation process from noise latent variables to real video latent variables and real audio latent variables, respectively. In one implementation, the real video latent variables and real audio latent variables can be linearly interpolated with Gaussian noise to obtain intermediate states, and then the video tower and audio tower can be trained to predict the corresponding velocity fields. This training method allows the model to learn the joint generation trajectory of video and audio within a unified framework.

[0143] During the inference phase, based on the hierarchical script generated from the user's input, the Twin Towers Diffusion Transformer generates the video according to the constraints at the global, camera, and role levels corresponding to the script.

[0144] Figure 7 This is a block diagram illustrating a video generation apparatus according to an exemplary embodiment. (Refer to...) Figure 7 The device includes: The script acquisition module 710 is configured to execute the acquisition of a target script, the target script including shot layer information, the shot layer information being used to indicate the shot constraint information corresponding to each of the multiple shots of the target video, the multiple shots of the target video being used to jointly express the plot corresponding to the target script; The video generation module 720 is configured to execute a video generation model based on the target script, under the constraints of the shot constraint information corresponding to each of the multiple shots, to generate audio and video segments corresponding to each of the multiple shots, and to obtain the target video based on each of the audio and video segments; The video generation model is a model that generates audio and video segments through a synchronous audio and video cross-modal joint mode.

[0145] In one exemplary implementation, the video generation model is a model that performs simultaneous audio and video cross-modal joint generation based on an attention mechanism; The video generation module 720 is configured to perform: In the target script, the lens constraint information corresponding to each of the lenses is extracted; Based on the lens constraint information, attention boundary constraint information is generated. The attention boundary constraint information is used to constrain the generation of audio and video segments corresponding to the target lens during the generation process. The attention mechanism only selects the lens constraint information corresponding to the target lens to generate the corresponding audio and video segments. The target lens is any one of the multiple lenses. By inputting the attention boundary constraint information into the video generation model, audio and video segments corresponding to each of the multiple shots are generated.

[0146] In one exemplary implementation, the shot constraint information includes the corresponding shot's time constraint and plot constraint; The time constraints include at least one of the following: the time interval corresponding to the shot, the dialogue time interval in the shot, and the time interval in which the character exists within the shot. The plot constraints are used to indicate the plot expressed by the audio and video segments corresponding to the corresponding shots.

[0147] In one exemplary embodiment, the video generation module 720 is configured to perform: The boundary-aware routing matrix generated based on the attention boundary constraint information is input into the text attention layer of the video generation model, so that the video generation model generates audio and video segments corresponding to the multiple shots. The boundary-aware routing matrix is ​​used to control the visibility of different text terms to different audio and video latent variable terms.

[0148] In one exemplary embodiment, the target script further includes global layer information, which is used to control the overall information of the target video; The video generation module 720 is configured to perform: Global constraint information is generated based on the global layer information, and the global constraint information is applied to the generation process of audio and video segments corresponding to all shots of the target video. By inputting the overall constraint information into the video generation model, audio and video segments corresponding to each of the multiple shots are generated.

[0149] In one exemplary embodiment, the target script further includes role layer information, which is used to constrain the role characteristics corresponding to a single role; The video generation module 720 is configured to perform: In the target script, extract the role identifier information corresponding to each role; Based on the character identification information, attention character constraint information is generated. The attention character constraint information is used to constrain the generation of audio and video clips of the same character in different shots. The attention mechanism generates corresponding audio and video clips based on the character characteristics of the same character. By inputting the attention role constraint information into the video generation model, audio and video segments corresponding to each of the multiple shots are generated.

[0150] In one exemplary embodiment, the video generation module 720 is configured to perform: By inputting a character mask matrix generated based on the attention character constraint information into at least one attention layer of the video generation model, the video generation model is controlled to generate audio and video clips of the same character corresponding to different shots. The character mask matrix is ​​used to control the character features used in the audio and video generation process.

[0151] In one exemplary embodiment, the at least one attention layer includes at least one of the following: a video self-attention layer, an audio self-attention layer, a text attention layer, and an audio-video cross-attention layer; the video generation module 720 is configured to perform at least one of the following operations: By inputting the character mask matrix into the video self-attention layer, the interaction between the corresponding video latent variables and the corresponding character features is controlled; By inputting the character mask matrix into the audio self-attention layer, the interaction between the corresponding audio latent variables and the corresponding character features is controlled; By inputting the character mask matrix into the text attention layer, the interaction between the corresponding text words and the corresponding character features is controlled; By inputting the character mask matrix into the audio-video cross-attention layer, the interaction between the audio latent variables and video latent variables for the same character is controlled.

[0152] In one exemplary embodiment, the video generation module 720 is configured to perform: Obtain media reference information, which includes at least one of the following: a reference image corresponding to the character, a reference sound corresponding to the character, and a posture reference corresponding to the character; The media reference information is input into the video generation model to obtain the character features corresponding to the character; By inputting the attention role constraint information into the video generation model, the video generation model generates audio and video clips corresponding to the role in multiple shots based on the role's corresponding role features.

[0153] In one exemplary embodiment, the script acquisition module 710 is configured to execute: Obtain text prompt information, which is used to indicate the basic plot of the target video to be generated; The text prompt information is input into the script generation agent for hierarchical script generation processing to obtain the target script, which also includes global layer information and role layer information.

[0154] In one exemplary embodiment, the script generation agent includes: a verification agent, a structure processing agent, an identity alignment agent, and a narrative refinement agent; The script acquisition module 710 is configured to execute: The text prompt information is input into the verification agent to perform common sense contradiction detection; After correcting the common-sense contradictions in the text prompt information, the corrected text prompt information is input into the structured processing agent for structured processing of global layer information and camera layer information to obtain the first processing result. The first processing result is input into the identity-aligned agent to perform structured processing of role-layer information, resulting in the second processing result. The second processing result is input into the narrative refinement agent to supplement the details of the global layer information, the shot layer information, and the character layer information, and the target script is output.

[0155] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0156] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided. Optionally, the computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc. When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is able to perform any of the methods described above.

[0157] In an exemplary embodiment, a computer program product is also provided, the computer program product including a computer program stored in a readable storage medium, wherein at least one processor of a computer device reads from the readable storage medium and executes the computer program, causing the device to perform any of the methods described above.

[0158] Figure 8 This is a block diagram illustrating an electronic device for video generation according to an exemplary embodiment. The electronic device may be a terminal, and its internal structure diagram may be as follows: Figure 8 As shown, the device may include an RF (Radio Frequency) circuit 810, a memory 820 including one or more computer-readable storage media, an input unit 830, a display unit 840, a sensor 850, an audio circuit 860, a WiFi (Wireless Fidelity) module 870, a processor 880 including one or more processing cores, and a power supply 890, among other components. Those skilled in the art will understand that... Figure 8The terminal structure shown does not constitute a limitation on the terminal and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein: The RF circuit 810 can be used for receiving and transmitting signals during information transmission or calls. Specifically, it receives downlink information from the base station and hands it over to one or more processors 880 for processing; additionally, it transmits uplink data to the base station. Typically, the RF circuit 810 includes, but is not limited to, an antenna, at least one amplifier, a tuner, one or more oscillators, a Subscriber Identity Module (SIM) card, a transceiver, a coupler, an LNA (Low Noise Amplifier), a duplexer, etc. Furthermore, the RF circuit 810 can also communicate wirelessly with networks and other terminals. Wireless communication can use any communication standard or protocol, including but not limited to GSM (Global System for Mobile communication), GPRS (General Packet Radio Service), CDMA (Code Division Multiple Access), WCDMA (Wideband Code Division Multiple Access), LTE (Long Term Evolution), email, SMS (Short Messaging Service), etc.

[0159] The memory 820 can be used to store software programs and modules. The processor 880 executes various functional applications and data processing by running the software programs and modules stored in the memory 820. The memory 820 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for the functions, etc.; the data storage area may store data created according to the use of the terminal, etc. In addition, the memory 820 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 820 may also include a memory controller to provide access to the memory 820 for the processor 880 and the input unit 830.

[0160] The input unit 830 can be used to receive input digital or character information, and to generate keyboard, mouse, joystick, optical, or trackball signal inputs related to user settings and function control. Specifically, the input unit 830 may include a touch-sensitive surface 831 and other input devices 832. The touch-sensitive surface 831, also known as a touch display screen or touchpad, can collect touch operations performed by the user on or near it (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near the touch-sensitive surface 831), and drive the corresponding connected devices according to a pre-set program. Optionally, the touch-sensitive surface 831 may include two parts: a touch detection device and a touch controller. The touch detection device detects the user's touch position and the signal generated by the touch operation, and transmits the signal to the touch controller; the touch controller receives touch information from the touch detection device, converts it into touch point coordinates, and sends it to the processor 880, and can also receive and execute commands sent by the processor 880. In addition, the touch-sensitive surface 831 can be implemented using various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch-sensitive surface 831, the input unit 830 may also include other input devices 832. Specifically, other input devices 832 may include, but are not limited to, one or more of the following: physical keyboard, function keys (such as volume control buttons, power buttons, etc.), trackball, mouse, joystick, etc. The display unit 840 can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces of the terminal. These graphical user interfaces can be composed of graphics, text, icons, video, and any combination thereof. The display unit 840 may include a display panel 841, which may optionally be configured as an LCD (Liquid Crystal Display), OLED (Organic Light-Emitting Diode), or similar display panel 841. Further, a touch-sensitive surface 831 may cover the display panel 841. When the touch-sensitive surface 831 detects a touch operation on or near it, it transmits the information to the processor 880 to determine the type of touch event. Subsequently, the processor 880 provides corresponding visual output on the display panel 841 according to the type of touch event. The touch-sensitive surface 831 and the display panel 841 can be two independent components to implement input and output functions. However, in some embodiments, the touch-sensitive surface 831 and the display panel 841 can be integrated to achieve input and output functions.

[0161] The terminal may also include at least one sensor 850, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. The ambient light sensor can adjust the brightness of the display panel 841 according to the ambient light level, and the proximity sensor can turn off the display panel 841 and / or the backlight when the terminal is moved to the ear. As a type of motion sensor, a gravity acceleration sensor can detect the magnitude of acceleration in various directions (generally three axes). When stationary, it can detect the magnitude and direction of gravity and can be used for applications that identify the terminal's posture (such as landscape / portrait switching, related games, magnetometer posture calibration), vibration recognition-related functions (such as pedometer, tapping), etc. Other sensors that may be configured on the terminal, such as gyroscopes, barometers, hygrometers, thermometers, and infrared sensors, will not be described in detail here.

[0162] Audio circuitry 860, speaker 861, and microphone 862 provide an audio interface between the user and the terminal. Audio circuitry 860 converts received audio data into electrical signals, which are then transmitted to speaker 861, where they are converted into sound signals for output. Conversely, microphone 862 collects sound signals, converts them into electrical signals, which are then received by audio circuitry 860, converted back into audio data, processed by processor 880, and transmitted via RF circuitry 810 to, for example, another terminal, or output to memory 820 for further processing. Audio circuitry 860 may also include an earphone jack to facilitate communication between a peripheral headset and the terminal.

[0163] WiFi is a short-range wireless transmission technology. This terminal, through the WiFi module 870, can help users send and receive emails, browse web pages, and access streaming media, providing users with wireless broadband internet access. Although Figure 8 WiFi module 870 is shown, but it is understood that it is not a necessary component of the terminal and can be omitted as needed without changing the nature of the invention.

[0164] The processor 880 is the control center of the terminal, connecting various parts of the terminal through various interfaces and lines. It executes software programs and / or modules stored in the memory 820, and calls data stored in the memory 820 to perform various functions and process data, thereby enabling overall monitoring of the terminal. Optionally, the processor 880 may include one or more processing cores; preferably, the processor 880 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interaction area, and applications, while the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 880.

[0165] The terminal also includes a power supply 890 (such as a battery) to power various components. Preferably, the power supply can be logically connected to the processor 880 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 890 may also include one or more DC or AC power supplies, a recharging system, a power fault detection circuit, a power converter or inverter, a power status indicator, and other arbitrary components.

[0166] Although not shown, the terminal may also include a camera, Bluetooth module, etc., which will not be described in detail here. Specifically, in this embodiment, the display unit of the terminal is a touch screen display, and the terminal also includes a memory and one or more programs, wherein one or more programs are stored in the memory and configured to be executed by one or more processors of the instructions in the method embodiment of the present invention.

[0167] Please refer to Figure 9 This illustrates another block diagram of an electronic device for video generation, provided in another exemplary embodiment of this disclosure. The computer device may be a server for performing the video generation method described above. Specifically: Computer device 900 includes a Central Processing Unit (CPU) 901, a system memory 904 including Random Access Memory (RAM) 902 and Read Only Memory (ROM) 903, and a system bus 905 connecting the system memory 904 and the CPU 901. Computer device 900 also includes a basic input / output system (I / O system) 906 that facilitates information transfer between various devices within the computer, and a mass storage device 907 for storing the operating system 913, application programs 914, and other program modules 911.

[0168] The basic input / output system 906 includes a display 908 for displaying information and an input device 909 for user input, such as a mouse or keyboard. Both the display 908 and the input device 909 are connected to the central processing unit 901 via an input / output controller 190 connected to the system bus 905. The basic input / output system 906 may also include the input / output controller 190 for receiving and processing input from multiple other devices such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 190 also provides output to a display screen, printer, or other types of output devices.

[0169] Mass storage device 907 is connected to central processing unit 901 via a mass storage controller (not shown) connected to system bus 905. Mass storage device 907 and its associated computer-readable media provide non-volatile storage for computer device 900. That is, mass storage device 907 may include computer-readable media (not shown) such as hard disk or CD-ROM (CompactDisc Read-Only Memory) drive.

[0170] Without loss of generality, computer-readable media can include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented using any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes RAM, ROM, EPROM (Erasable Programmable Read Only Memory), EEPROM (Electrically Erasable Programmable Read Only Memory), flash memory or other solid-state storage technologies, CD-ROM, DVD (Digital Video Disc) or other optical storage, magnetic tape cassettes, magnetic tape, disk storage, or other magnetic storage devices. Of course, those skilled in the art will recognize that computer storage media are not limited to the above-mentioned types. The system memory 904 and mass storage device 907 described above can be collectively referred to as memory.

[0171] According to various embodiments of this disclosure, the computer device 900 can also be connected to a remote computer on a network, such as the Internet. That is, the computer device 900 can be connected to a network 912 via a network interface unit 911 connected to a system bus 905, or the network interface unit 911 can be used to connect to other types of networks or remote computer systems (not shown).

[0172] The aforementioned memory also includes a computer program stored in the memory and configured to be executed by one or more processors to implement the aforementioned video generation method.

[0173] In an exemplary embodiment, a computer-readable storage medium is also provided, which stores at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, at least one program, code set, or instruction set is executed by a processor to implement the video generation method.

[0174] Optionally, the computer-readable storage medium may include: ROM (Read Only Memory), RAM (Random Access Memory), SSD (Solid State Drives), or optical disc, etc. The random access memory may include ReRAM (Resistance Random Access Memory) and DRAM (Dynamic Random Access Memory).

[0175] In an exemplary embodiment, a computer-readable storage medium including program code is also provided, such as a memory including program code, which can be executed by a processor to complete the video generation method described above. Optionally, the computer-readable storage medium may be read-only memory (ROM), random access memory (RAM), compact-disc read-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.

[0176] In an exemplary embodiment, a computer program product is also provided, including a computer program that, when executed by a processor, implements the video generation method described above.

[0177] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.

[0178] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. A video generation method, characterized in that, The method includes: Obtain the target script, which includes shot layer information. The shot layer information is used to indicate the shot constraint information corresponding to each of the multiple shots of the target video. The multiple shots of the target video are used to jointly express the plot corresponding to the target script. Based on the target script, the video generation model is controlled to generate audio and video segments corresponding to each of the multiple shots under the constraints of the shot constraint information corresponding to each of the multiple shots, and the target video is obtained based on each of the audio and video segments; The video generation model is a model that generates audio and video segments through a synchronous audio and video cross-modal joint mode.

2. The video generation method according to claim 1, characterized in that, The video generation model is a model that performs synchronous audio and video cross-modal joint generation based on an attention mechanism; The step of controlling the video generation model based on the target script to generate audio and video segments corresponding to each of the multiple shots under the constraints of the shot constraint information corresponding to each of the multiple shots includes: In the target script, the lens constraint information corresponding to each of the lenses is extracted; Based on the lens constraint information, attention boundary constraint information is generated. The attention boundary constraint information is used to constrain the generation of audio and video segments corresponding to the target lens during the generation process. The attention mechanism only selects the lens constraint information corresponding to the target lens to generate the corresponding audio and video segments. The target lens is any one of the multiple lenses. By inputting the attention boundary constraint information into the video generation model, audio and video segments corresponding to each of the multiple shots are generated.

3. The video generation method according to claim 2, characterized in that, The shot constraint information includes the time constraint and plot constraint of the corresponding shot; The time constraints include at least one of the following: the time interval corresponding to the shot, the dialogue time interval in the shot, and the time interval in which the character exists within the shot. The plot constraints are used to indicate the plot expressed by the audio and video segments corresponding to the corresponding shots.

4. The video generation method according to claim 3, characterized in that, The step of inputting the attention boundary constraint information into the video generation model to generate audio and video segments corresponding to each of the multiple shots includes: The boundary-aware routing matrix generated based on the attention boundary constraint information is input into the text attention layer of the video generation model, so that the video generation model generates audio and video segments corresponding to the multiple shots. The boundary-aware routing matrix is ​​used to control the visibility of different text terms to different audio and video latent variable terms.

5. A video generation method according to any one of claims 2 to 4, characterized in that, The target script also includes global layer information, which is used to control the overall information of the target video; The step of controlling the video generation model based on the target script to generate audio and video segments corresponding to each of the multiple shots under the constraints of the shot constraint information corresponding to each of the multiple shots includes: Global constraint information is generated based on the global layer information, and the global constraint information is applied to the generation process of audio and video segments corresponding to all shots of the target video. By inputting the overall constraint information into the video generation model, audio and video segments corresponding to each of the multiple shots are generated.

6. The video generation method according to claim 5, characterized in that, The target script also includes role layer information, which is used to constrain the role characteristics corresponding to a single role. The step of controlling the video generation model based on the target script to generate audio and video segments corresponding to each of the multiple shots under the constraints of the shot constraint information corresponding to each of the multiple shots includes: In the target script, extract the role identifier information corresponding to each role; Based on the character identification information, attention character constraint information is generated. The attention character constraint information is used to constrain the generation of audio and video clips of the same character in different shots. The attention mechanism generates corresponding audio and video clips based on the character characteristics of the same character. By inputting the attention role constraint information into the video generation model, audio and video segments corresponding to each of the multiple shots are generated.

7. A video generation method according to claim 6, characterized in that, The step of inputting the attention role constraint information into the video generation model to generate audio and video segments corresponding to each of the multiple shots includes: By inputting a character mask matrix generated based on the attention character constraint information into at least one attention layer of the video generation model, the video generation model is controlled to generate audio and video clips of the same character corresponding to different shots. The character mask matrix is ​​used to control the character features used in the audio and video generation process.

8. A video generation method according to claim 7, characterized in that, The at least one attention layer includes at least one of the following: video self-attention layer, audio self-attention layer, text attention layer, and audio-video cross-attention layer; The process of inputting the character mask matrix generated based on the attention character constraint information into multiple attention layers of the video generation model to control the generation of audio and video clips of the same character corresponding to different shots includes performing at least one of the following operations: By inputting the character mask matrix into the video self-attention layer, the interaction between the corresponding video latent variables and the corresponding character features is controlled; By inputting the character mask matrix into the audio self-attention layer, the interaction between the corresponding audio latent variables and the corresponding character features is controlled; By inputting the character mask matrix into the text attention layer, the interaction between the corresponding text words and the corresponding character features is controlled; By inputting the character mask matrix into the audio-video cross-attention layer, the interaction between the audio latent variables and video latent variables for the same character is controlled.

9. A video generation method according to claim 6, characterized in that, The method further includes: acquiring media reference information, the media reference information including at least one of the following: a reference image corresponding to the character, a reference sound corresponding to the character, and a posture reference corresponding to the character; The step of inputting the attention role constraint information into the video generation model to generate audio and video segments corresponding to each of the multiple shots includes: The media reference information is input into the video generation model to obtain the character features corresponding to the character; By inputting the attention role constraint information into the video generation model, the video generation model generates audio and video clips corresponding to the role in multiple shots based on the role's corresponding role features.

10. A video generation method according to claim 1, characterized in that, The acquisition of the target script includes: Obtain text prompt information, which is used to indicate the basic plot of the target video to be generated; The text prompt information is input into the script generation agent for hierarchical script generation processing to obtain the target script, which also includes global layer information and role layer information.

11. A video generation method according to claim 10, characterized in that, The script-generated intelligent agents include: a verification intelligent agent, a structured processing intelligent agent, an identity alignment intelligent agent, and a narrative refinement intelligent agent; The step of inputting the text prompt information into a script generation agent for hierarchical script generation processing to obtain the target script includes: The text prompt information is input into the verification agent to perform common sense contradiction detection; After correcting the common-sense contradictions in the text prompt information, the corrected text prompt information is input into the structured processing agent for structured processing of global layer information and camera layer information to obtain the first processing result. The first processing result is input into the identity-aligned agent to perform structured processing of role-layer information, resulting in the second processing result. The second processing result is input into the narrative refinement agent to supplement the details of the global layer information, the shot layer information, and the character layer information, and the target script is output.

12. A video generation apparatus, characterized in that, The device includes: The script acquisition module is configured to acquire a target script, which includes shot layer information. The shot layer information is used to indicate the shot constraint information corresponding to each of the multiple shots of the target video. The multiple shots of the target video are used to jointly express the plot corresponding to the target script. The video generation module is configured to execute a video generation model based on the target script, under the constraints of the shot constraint information corresponding to each of the multiple shots, to generate audio and video segments corresponding to each of the multiple shots, and to obtain the target video based on each of the audio and video segments; The video generation model is a model that generates audio and video segments through a synchronous audio and video cross-modal joint mode.

13. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the instructions to implement the video generation method as described in any one of claims 1 to 11.

14. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is enabled to perform the video generation method as described in any one of claims 1 to 11.

15. A computer program product, characterized in that, The computer program product includes a computer program stored in a readable storage medium, wherein at least one processor of a computer device reads from and executes the computer program, causing the device to perform the video generation method as described in any one of claims 1 to 11.