Video generation method and device, electronic equipment, storage medium and program product

Generating videos with transparent areas through the target video generation model solves the problem that users need to additionally process extract part of the video content and improves the user experience.

CN120264035APending Publication Date: 2025-07-04NETEASE (HANGZHOU) NETWORK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510308206.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In the prior art, after generating video, users need to extract part of the video content through the application, resulting in a decrease in user experience.

Method used

By obtaining the target text, the trained target video generation model is called to generate the target video. The first area content of the target video frame is generated based on the text, and the second area includes transparent elements to realize the RGB-transparency channel video generation.

Benefits of technology

No additional processing is required to output videos containing transparent areas, improving user experience and simplifying post-operation processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120264035A_ABST
    Figure CN120264035A_ABST
Patent Text Reader

Abstract

The invention provides a video generation method and device, electronic equipment, a storage medium and a program product, and relates to the technical field of computer vision, and the method comprises the steps: obtaining a target text which is used for indicating the generation of a video; the trained target video generation model is called to generate a target video based on the target text, the target video comprises multiple frames of target video frames, the content of the first area of any target video frame is generated according to the target text, and the second area of any target video frame comprises transparent elements, so that after a user inputs the text, the content of the target video frame can be obtained. And outputting the video with the content in the first area and the transparent element in the second area, so that a user does not need to use some application programs to extract part of the content of the video, thereby improving the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and in particular, to a video generation method, apparatus, electronic device, storage medium, and program product. Background Art

[0002] In the related art, a user can input text to guide a large language model (LLM) to generate a video.

[0003] However, after the video is generated, the user may only want to extract some content of the video. In this case, the user needs to use some applications (Apps) to extract some content of the video, which will reduce the user experience. Summary of the Invention

[0004] In view of this, the purpose of this application is to provide a video generation method, apparatus, electronic device, storage medium, and program product, which can output a video in which the first area has content but the second area includes transparent elements after the user inputs text, so that the user does not need to use some applications to extract some content of the video, thereby improving the user experience.

[0005] In a first aspect, this application provides a video generation method, which includes: obtaining a target text, where the target text is used to indicate the generation of a video; calling a trained target video generation model to generate a target video based on the target text, where the target video includes multiple target video frames, and the content of the first area of any target video frame is generated according to the target text, and the second area of any target video frame includes transparent elements.

[0006] In a possible implementation, the step of calling a trained target video generation model to generate a target video based on the target text includes: obtaining a first initial noise video, where the first initial noise video is a video including multiple color channels and an alpha channel, the first initial noise video is generated based on the target text, the first initial noise video includes multiple first initial noise video frames, the content of the first area of any first initial noise video frame is generated according to the target text, and the second area of any first initial noise video frame is filled with content; calling a trained target video generation model to generate a target video based on the target text and the first initial noise video, where the target video is a video including multiple color channels and an alpha channel.

[0007] In a possible implementation, generating a target video based on a target text and a first initial noisy video includes: extracting a feature vector of the target text to obtain a target text feature vector; predicting a noise residual of the first initial noisy video to obtain a first noise residual, where the first noise residual is used to indicate the difference between the first initial noisy video and the target video; determining a gating coefficient based on the target text feature vector, and adjusting a first attention to the alpha channel of a first region of any first initial noisy video frame based on the gating coefficient, and adjusting the transparency of the content of the first region of any first initial noisy video frame based on the first attention, where the gating coefficient is used to represent the influence degree of the target text on the alpha channel; performing denoising processing based on the first noise residual to set the content filled in a second region of any first initial noisy video frame as a transparent element; and outputting the first initial noisy video with the transparency of the content of the first region adjusted and the transparent element of the second region set as the target video.

[0008] In a possible implementation, the gating coefficient includes a first gating coefficient or a second gating coefficient. Determining the gating coefficient based on the target text feature vector includes: if it is determined based on the target text feature vector that the target text includes a preset keyword, determining the gating coefficient as the first gating coefficient, where the preset keyword includes a word related to transparency; if it is determined based on the target text feature vector that the target text does not include the preset keyword, determining the gating coefficient as the second gating coefficient, and the influence degree represented by the first gating coefficient is greater than the influence degree represented by the second gating coefficient.

[0009] In a possible implementation, predicting the noise residual of the first initial noisy video to obtain the first noise residual includes: performing downsampling processing on the first initial noisy video to obtain the downsampled first initial noisy video; performing upsampling processing on the downsampled first initial noisy video to obtain the upsampled first initial noisy video; and predicting the noise residual of the upsampled first initial noisy video to obtain the first noise residual.

[0010] In a possible implementation, when performing downsampling or upsampling processing, at least one of the following is executed: capturing intra-frame dependencies and inter-frame dependencies of the first initial noisy video, where the intra-frame dependencies are used to indicate the relationships between different regions in the same frame, and the inter-frame dependencies are used to indicate the relationships between different video frames, and the intra-frame dependencies and inter-frame dependencies are used as references for upsampling or downsampling processing; fusing image feature information of different resolutions of any first initial noisy video frame, and the image feature information of different resolutions is used as a reference for upsampling or downsampling processing; introducing a marker embedding for the transparency channel, where the marker embedding is a learnable domain embedding to distinguish multiple color channels and the transparency channel through the marker embedding, and the marker embedding is used as a reference for upsampling or downsampling processing; fusing low-rank adaptation parameters with the frozen base parameters of the target video generation model to perform upsampling or downsampling processing on the first initial noisy video through the fused parameters obtained by fusing the low-rank adaptation parameters and the base parameters.

[0011] In a possible implementation, the method further includes: obtaining at least one adjusted target video frame; retraining the target video generation model using the at least one target video frame.

[0012] In a possible implementation, the target video generation model is trained as follows: obtaining a plurality of training samples and an initial video generation model, where any training sample includes a second initial noisy video and a text sample corresponding to the second initial noisy video, the second initial noisy video is a video including multiple color channels and a transparency channel, the second initial noisy video includes multiple frames of second initial noisy video frames, the content of the first region of any second initial noisy video frame is generated according to the text sample, and the second region of any second initial noisy video frame is filled with content; iteratively training the initial video generation model using the plurality of training samples until the training end condition is met to obtain the target video generation model.

[0013] In a possible implementation, the plurality of training samples are obtained as follows: obtaining a plurality of reference videos, each reference video corresponding to a description text, the reference video including multiple frames of reference video frames, for any reference video frame, the content of the first region of any reference video frame is generated according to the text sample, and the second region of any reference video frame includes a transparent element; for any reference video, performing noise addition processing on the reference video to obtain a second initial noisy video, and determining the description text corresponding to the reference video as the text sample corresponding to the second initial noisy video.

[0014] In a possible implementation, an initial video generation model is iteratively trained using multiple training samples until a training end condition is met to obtain a target video generation model, including: training the initial video generation model using a first part of the multiple training samples to obtain a predicted video output by the initial video generation model based on a second initial noise video and a corresponding text sample in the first part of the samples; calculating a training loss based on the predicted video and a reference video corresponding to the first part of the samples; if it is determined that the training end condition is met based on the training loss, determining the initial video generation model as the target video generation model; if it is determined that the training end condition is not met based on the training loss, after updating the model parameters of the initial video generation model, training the initial video generation model using a second part of the multiple training samples.

[0015] In a possible implementation, the predicted video output based on the second initial noise video and the corresponding text sample in the first part of the samples includes: extracting a feature vector of the text sample to obtain a text sample feature vector; predicting a noise residual of the second initial noise video to obtain a second noise residual, where the second noise residual is used to indicate the difference between the second initial noise video and the reference video; determining a gating coefficient based on the text sample feature vector, and adjusting a second attention to the alpha channel of a first region of any second initial noise video frame based on the gating coefficient, and adjusting the transparency of the content of the first region of any second initial noise video frame based on the second attention, where the gating coefficient is used to represent the influence degree of the text sample on the alpha channel; performing denoising processing based on the second noise residual to set the content filled in a second region of any second initial noise video frame as a transparent element; outputting the second initial noise video with the transparency of the content of the first region adjusted and the transparent element of the second region set as the predicted video.

[0016] In a possible implementation, the gating coefficient includes a third gating coefficient or a fourth gating coefficient. Determining the gating coefficient based on the text sample feature vector includes: if it is determined based on the text sample feature vector that the text sample includes a preset keyword, determining the gating coefficient as the third gating coefficient, where the preset keyword includes a word related to transparency; if it is determined based on the text sample feature vector that the text sample does not include the preset keyword, determining the gating coefficient as the fourth gating coefficient, and the influence degree represented by the third gating coefficient is greater than the influence degree represented by the fourth gating coefficient.

[0017] In a possible implementation, predicting the noise residual of the second initial noise video to obtain a second noise residual includes: performing downsampling on the second initial noise video to obtain the downsampled second initial noise video; performing upsampling on the downsampled second initial noise video to obtain the upsampled second initial noise video; predicting the noise residual of the upsampled second initial noise video to obtain the second noise residual.

[0018] In a possible implementation, when performing downsampling or upsampling, at least one of the following is performed: capturing the intra-frame dependency and inter-frame dependency of the second initial noise video, where the intra-frame dependency is used to indicate the relationship between different regions in the same frame, and the inter-frame dependency is used to indicate the relationship between different video frames, and the intra-frame dependency and inter-frame dependency are used as references for upsampling or downsampling; fusing the image feature information of different resolutions of any second initial noise video frame, and the image feature information of different resolutions is used as a reference for upsampling or downsampling; introducing a marker embedding for the transparency channel, where the marker embedding is a learnable domain embedding to distinguish multiple color channels and the transparency channel through the marker embedding, and the marker embedding is used as a reference for upsampling or downsampling; fusing the low-rank adaptation parameters with the frozen base parameters of the initial video generation model to perform upsampling or downsampling on the second initial noise video through the fused parameters after fusing the low-rank adaptation parameters and the base parameters.

[0019] In a possible implementation, the rank of the low-rank adaptation parameters corresponding to downsampling is less than the rank of the low-rank adaptation parameters corresponding to upsampling.

[0020] In a second aspect, an embodiment of the present application provides a video generation device, including: a first acquisition module, configured to acquire a target text, where the target text is used to indicate the generation of a video; a video generation module, configured to call a trained target video generation model to generate a target video based on the target text, where the target video includes multiple target video frames, and the content of the first region of any target video frame is generated according to the target text, and the second region of any target video frame includes a transparent element.

[0021] In a third aspect, an embodiment of the present application provides an electronic device, including a processor and a memory, where the memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the above method.

[0022] In a fourth aspect, an embodiment of the present application provides a machine-readable storage medium, where the machine-readable storage medium stores machine-executable instructions, and when the machine-executable instructions are called and executed by a processor, the machine-executable instructions cause the processor to implement the above method.

[0023] Fifth aspect, an embodiment of the present application provides a computer program product, which includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions that, when executed by a computer device, cause the computer device to execute the above-mentioned method.

[0024] The embodiments of the present application bring the following beneficial effects: By obtaining a target text, which is used to indicate the generation of a video; calling a trained target video generation model to generate a target video based on the target text. The target video includes multiple frames of target video frames. The content of the first region of any target video frame is generated according to the target text, and the second region of any target video frame includes a transparent element. In this way, after the user inputs the text, a video with content in the first region but including a transparent element in the second region can be output, so that the user does not need to use some application programs to extract part of the content of the video, thereby improving the user experience.

[0025] Other features and advantages of the present application will be described in the following specification, and, in part, will be obvious from the specification, or will be understood by implementing the present application. The objectives and other advantages of the present application are achieved and obtained by the structures specifically pointed out in the specification, claims, and drawings.

[0026] To make the above objectives, features, and advantages of the present application more obvious and understandable, the following specific preferred embodiments are given, and in conjunction with the accompanying drawings, the detailed description is as follows. Description of the Drawings

[0027] In order to more clearly illustrate the specific implementation manners of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific implementation manners or the prior art. Obviously, the following drawings are some implementation manners of the present application. For those skilled in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0028] Figure 1 It is a schematic flowchart of a video generation method provided by an embodiment of the present application; Figure 2 It is a schematic structural diagram of a target video generation model provided by an embodiment of the present application; Figure 3 It is a schematic flowchart of a model training method provided by an embodiment of the present application; Figure 4 It is a schematic diagram of a target video frame provided by an embodiment of the present application; Figure 5 It is a schematic structural diagram of a video generation device provided by an embodiment of the present application; Figure 6Schematic structural diagram of a model training device provided by an embodiment of the present application; Figure 7 Schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0029] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions of the present application will be clearly and completely described below with reference to the accompanying drawings. Apparently, the described embodiments are some but not all of the embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0030] In the related art, a user can input text to guide a large language model (LLM) to generate a video.

[0031] However, after the video is generated, the user may only want to extract some parts of the video. At this time, the user needs to use some applications to extract the parts of the video, which will reduce the user experience.

[0032] Specifically, in the related art, a video in the red green blue color mode (RGB) can be generated first, and then the video matting algorithm of an application can be used to obtain a transparent channel, so as to extract some parts of the video. This process is scattered and has high requirements for temporal consistency; if the model can output a video with RGB - alpha channel (Alpha, RGBA), the later cost will be greatly reduced.

[0033] Therefore, the embodiments of the present application provide a video generation method, device, electronic device, storage medium, and program product, which can output a video with content in a first area but including transparent elements in a second area after the user inputs text, so that the user does not need to use some applications to extract some parts of the video, thereby improving the user experience.

[0034] Specifically, the technical solution of this embodiment can be applied to scenarios such as film and television special effects, augmented reality (AR), virtual reality (VR), and game animation.

[0035] Next, the video generation method will be described first.

[0036] The video generation method in one embodiment of the present application can run on a local terminal device or a server. When the video generation method runs on the server, the method can be implemented and executed based on a cloud interaction system, where the cloud interaction system includes a server and client devices (also referred to as local terminal devices or terminal devices).

[0037] In an alternative embodiment, various cloud applications and models can run under the cloud interaction system. The models can be deployed on the client devices or on the server. The cloud applications can be, for example, cloud games. Taking cloud games as an example, cloud games refer to a game mode based on cloud computing. In the operation mode of cloud games, the running entity of the game program and the presenting entity of the game screen are separated. The storage and running of the image editing method are completed on the cloud game server. The role of the client device is to receive, send, and present the game screen. For example, the client device can be a display device with data transmission function near the user side, such as a mobile terminal, a television, a computer, a palm computer, etc.; however, the information processing is performed by the cloud game server in the cloud. When playing a game, the player operates the client device to send an operation instruction to the cloud game server. The cloud game server runs the game according to the operation instruction, encodes and compresses data such as the game screen, returns it to the client device through the network, and finally, the client device decodes and outputs the game screen.

[0038] In an alternative embodiment, taking a game as an example, the local terminal device stores a game program and is used to present the game screen. The local terminal device is used to interact with the player through a graphical user interface, that is, conventionally, the game program is downloaded and installed on an electronic device and run. The way the local terminal device provides the graphical user interface to the player can include various methods. For example, it can be rendered and displayed on the display screen of the terminal, or provided to the player through holographic projection. For example, the local terminal device can include a display screen and a processor. The display screen is used to present the graphical user interface, which includes the game screen. The processor is used to run the game, generate the graphical user interface, and control the display of the graphical user interface on the display screen.

[0039] For ease of understanding of this embodiment, first, an exemplary description of a video generation method disclosed in the embodiments of the present application is given. The video generation method of this embodiment can be executed by the server, or by the client device, or jointly executed by the server and the client device. Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a video generation method provided by the embodiments of the present application. As Figure 1 shown, the video generation method can include: S110. Obtain a target text, where the target text is used to indicate the generation of a video.

[0040] Among them, the target text can be used to indicate or describe the video content to be generated. The target text can be a short description, a story outline, a series of keywords, or any text information that can convey the video theme or content, and there is no limitation here.

[0041] S120. Call the trained target video generation model to generate a target video based on the target text. The target video includes multiple frames of target video frames. The content of the first region of any target video frame is generated according to the target text, and the second region of any target video frame includes a transparent element.

[0042] Among them, the target video generation model can receive text input and generate a corresponding video according to this input. The target video generation model may be a deep learning model, such as generative adversarial networks (GANs), transformers, or other types of neural networks, which have been trained with a large amount of video and text data and have learned how to convert text content into videos. Any target video frame is a static image. The content of the first region of each video frame is generated according to the target text, so the first region may contain images related to the target text. These images can be visual elements generated according to the text description, such as people, animals, objects, or scenes, for intuitively expressing the content described in the text. The content of the second region of each video frame includes a transparent element. The transparent element can refer to an element with a certain transparency. In this embodiment, the transparency of the transparent element can be 0%, indicating completely transparent. Among them, the first region and the second region can be different regions. In this embodiment, the first region may include the foreground, and the second region may include the background, or the first region includes the background and the second region includes the foreground, and there is no limitation here.

[0043] It can be understood that the content of the first region of the target video frame is generated according to the target text, and the second region of any target video frame includes a transparent element, that is, in the generated video, except for the content of the first region, there are no additional visual elements in the second region. In this way, the target video output by the model is a video including the transparent second region, and there is no need for the user to perform further processing to obtain a video with a transparent second region.

[0044] In a possible implementation manner, the step of calling the trained target video generation model to generate a target video based on the target text includes: Obtain a first initial noise video, where the first initial noise video is a video containing multiple color channels and an alpha channel. The first initial noise video is generated based on the target text. The first initial noise video includes multiple frames of first initial noise video frames. The content of the first region of any first initial noise video frame is generated according to the target text, and the second region of any first initial noise video frame is filled with content. Call the trained target video generation model to generate a target video based on the target text and the first initial noise video. The target video is a video containing multiple color channels and an alpha channel.

[0045] Among them, the multiple color channels can be, for example, the red green blue (RGB) color channels. The higher the value of the alpha channel, the less transparent the pixel, and the smaller the value, the more transparent. With the alpha channel, the content of the first region of the video can be seamlessly superimposed on other regions later without a complex matte extraction process. Among them, the first initial noise video can be generated based on the target text through related technologies. In this embodiment, the target video generation model has the ability to set the content filled in the second region of the first initial noise video as a transparent element based on the input text. In other words, it has the ability to adjust the transparency of the content filled in the second region of the first initial noise video based on the input text. In this embodiment, the first initial noise video and the target video can be RGB-alpha channel (RGBA) videos.

[0046] In this embodiment, by obtaining the first initial noise video, where the first initial noise video is a video containing multiple color channels and an alpha channel, the first initial noise video is generated based on the target text, the first initial noise video includes multiple frames of first initial noise video frames, the content of the first region of any first initial noise video frame is generated according to the target text, and the second region of any first initial noise video frame is filled with content; call the trained target video generation model to generate a target video based on the target text and the first initial noise video, where the target video is a video containing multiple color channels and an alpha channel, that is, the target video generation model only needs to have the ability to adjust the transparency of the content filled in the second region of the first initial noise video based on the input text, which can reduce the model architecture of the target video generation model.

[0047] In another possible implementation, the target video generation model can also generate the target video based on the target video guidance, so that it is not necessary to pre-generate the first initial noise video, thereby improving the generation efficiency of the target video.

[0048] In one possible implementation, generating the target video based on the target text and the first initial noise video includes: Extract the feature vector of the target text to obtain the target text feature vector; predict the noise residual of the first initial noise video to obtain the first noise residual, where the first noise residual is used to indicate the difference between the first initial noise video and the target video; determine the gating coefficient based on the target text feature vector, and adjust the first attention to the transparency channel of the first region of any first initial noise video frame based on the gating coefficient, and adjust the transparency of the content of the first region of any first initial noise video frame based on the first attention, where the gating coefficient is used to represent the influence degree of the target text on the transparency channel; perform denoising processing based on the first noise residual to set the content filled in the second region of any first initial noise video frame as a transparent element; output the first initial noise video with the transparency of the content of the first region adjusted and the transparent element of the second region set as the target video.

[0049] It should be noted that the above steps can be implemented by one or more modules in the model for generating the target video, which can be set as needed and are not limited here.

[0050] In a possible implementation manner, the above steps can be implemented by multiple modules in the model for generating the target video. For the convenience of understanding, the following embodiments will exemplarily illustrate with the model architecture of one of the target video generation models. Please refer to Figure 2 , Figure 2 which is a schematic diagram of the architecture of a target video generation model provided by an embodiment of the present application.

[0051] As Figure 2 shown, the target video generation model includes a text feature extraction module, a diffusion transformation module, an inverse diffusion module, and an output module. Generating a target video based on the target text and the first initial noise video includes: Extract the feature vector of the target text through the text feature extraction module to obtain the target text feature vector; predict the noise residual of the first initial noise video through the diffusion transformation module to obtain the first noise residual, where the first noise residual is used to indicate the difference between the first initial noise video and the target video; determine the gating coefficient based on the target text feature vector, and adjust the first attention to the transparency channel of the first region of any first initial noise video frame based on the gating coefficient, and adjust the transparency of the content of the first region of any first initial noise video frame based on the first attention, where the gating coefficient is used to represent the influence degree of the target text on the transparency channel; perform denoising processing based on the first noise residual through the inverse diffusion module to set the content filled in the second region of any first initial noise video frame as a transparent element; output the first initial noise video with the transparency of the content of the first region adjusted and the transparent element of the second region set through the output module as the target video.

[0052] Among them, the text feature vector can be a mathematical representation obtained after being processed by the feature extraction module. The target text is transformed into a feature vector through the text feature extraction module, that is, the target text feature vector. This vector contains the key information of the text and can be used to understand the content and meaning of the text.

[0053] The diffusion transformation module predicts the noise residual between the first initial noisy video and the target video through a deep learning network (such as U-Net, etc.). This prediction process is based on the model's understanding of the target text feature vector and its previously learned ability to restore useful information from noise. The obtained first noise residual can then be used to guide the reverse diffusion process, that is, the process of gradually restoring the target video from noise. In this process, the model will use the noise residual to adjust the noise level of each frame and gradually approach the state of the target video. Specifically, the first noise residual is used to indicate the difference between the first initial noisy video and the target video. Specifically, it can be used to indicate the difference between the content filled in the second region of the first initial noisy video and the transparent element in the second region of the target video. In this way, the first noise residual can be used to guide the content filled in the second region of the first initial noisy video to be set as a transparent element, thereby obtaining the target video.

[0054] Among them, the gating coefficient is calculated based on the target text feature vector, and it reflects the influence degree of the text on the transparency of the first region of the video frame. The gating coefficient can be regarded as a weight used to adjust the transparency of the first region in the video frame. In image processing, the transparency channel is used to control the transparency of an image. In the RGBA color model, A represents the Alpha channel, which determines the transparency of the pixel.

[0055] In this embodiment, with the Alpha channel, in the later stage, the content of the first region of the video can be seamlessly superimposed on any region without a complex matte extraction process. For the Alpha channel, a "transparent token (Alpha Token)" is introduced, and the attention flow between the text and Alpha is determined through a learnable gating network. When there is a lack of relevant semantics, unnecessary text interference is blocked; when the text clearly describes a "transparent / translucent" object, the transparency adjustment of the first region is enabled.

[0056] Specifically, "unnecessary text interference" refers to the negative impact of text information on the generation process of the transparency channel when the text description has no relation to the transparency characteristics that the transparency channel in the video should present. For example: Suppose the text description is "A cat sitting on a mat". This sentence mainly describes the RGB information such as the color, texture, and posture of the cat and the mat, and has no direct relation to whether the objects in the video should have transparency. In this case, if the model still forces the text information to affect the generation of the transparency channel, it may cause unnecessary interference. For example, it may result in unnatural semi-transparent effects at the edges of the cat or the mat, or noise unrelated to the text semantics in the transparency channel. The reason is that: The transparency channel is mainly responsible for describing the outline and transparency of objects, and its information characteristics are essentially different from the RGB channels (color, texture, etc.). When the focus of the text description has nothing to do with transparency, the text information is limited in helping to accurately generate the transparency channel, and may instead introduce noise or mislead the model.

[0057] Enabling transparency adjustment in the first region: This can refer to the information exchange and influence between text information and the transparency channel generation process. Specifically, it can achieve adaptive adjustment between text information and transparency requirements, so as to more precisely control the generation of the transparency channel. Specific goals: Avoiding unnecessary interference: When the text description has nothing to do with transparency, minimize the impact of text information on the transparency channel generation process as much as possible, so that the model can focus more on learning the transparency characteristics from the video content itself. Utilizing relevant semantic information: When the text description clearly indicates the need to generate transparent or semi-transparent objects (for example, containing keywords such as "glass", "transparent", etc.), then appropriately allow the text information to guide the generation of the transparency channel. For example, it helps the model identify the outline of the object and predict a reasonable transparency distribution.

[0058] In this embodiment, the target video can be an RGBA video. This embodiment can split the RGBA video into multiple blocks (video frames), and each Patch corresponds to an RGB marker + an Alpha marker, where the Alpha marker carries a learnable domain embedding , enabling the model to distinguish transparency information from color information. Specifically, the role of the Alpha marker is to explicitly inform the model which markers are used to represent transparency information and which markers are used to represent color information, thereby helping the model distinguish and decouple these two different types of information, and more specifically learn their respective feature representations and generation rules. Specific manifestations of the role: Feature representation decoupling: Through domain embedding (dα), the model can map RGB tokens and Alpha tokens to different feature spaces and learn their respective independent feature representations. The feature representation of RGB tokens can focus more on information such as color and texture, while the feature representation of Alpha tokens can focus more on information such as contours and transparency changes.

[0059] Differentiated information processing: In subsequent Transformer Blocks and attention calculations, the model can adopt different processing methods according to the type of tokens. For example, for Alpha tokens, the model can pay more attention to their spatial relationships with surrounding tokens and their temporal relationships with tokens at the same position in the time dimension, so as to better capture transparent edges and temporal consistency.

[0060] Parameter-efficient fine-tuning: Combining with the hierarchical LoRA fine-tuning strategy, the model parameters related to Alpha tokens can be more targeted adjusted. For example, only the Query / Key / Value matrices related to Alpha tokens are fine-tuned, while the parameters related to RGB are kept unchanged, so as to more efficiently learn the generation of the transparent channel on a limited dataset. Update domain embedding and color information: During training, the parameters of the learnable domain embedding (dα) will be updated to better complete the task of distinguishing RGB and Alpha channel information. At the same time, the parameters related to RGB channel generation (color information) in the model will also be updated so that the model can continuously improve the generation quality of the RGB channel. The purpose of differentiation is not to only update the domain embedding without updating the color information, but to more effectively update the model parameters so that it can better handle the RGBA video generation task.

[0061] The gating network can be a learnable mask. Specifically, let the gating network Input text features and then output a scalar or a small-dimensional vector to control the release or shielding of Text→Alpha attention. If the text contains keywords such as transparent / glass / semitransparent, then more attention is activated; if the text has nothing to do with transparency, then the "Text→Alpha" attention path is restricted to avoid useless interference.

[0062] The mask formula can be expressed as:

[0063] Specifically, Aij represents the i×j-th element of the mask matrix, which is used to control the weight relationship between features. i∈Text indicates that feature i belongs to text-related features (such as the text vector after CLIP encoding). j∈Alpha indicates that feature j belongs to the features of the alpha channel. Γ(z) represents the gating coefficient, which is calculated from the text feature z (Γ(z)∈[0, 1]).

[0064] The first attention in this embodiment can refer to the attention degree or importance score of the model to the first region of the video frame. This attention is calculated based on the gating coefficient and is used to guide how to adjust the transparency of the first region. Based on the first attention, the model will adjust the transparency of the first region of the video frame. This usually involves modifying the value of the alpha channel, thereby changing the visibility of the first region part. If the first attention is high, it means that the text has a greater impact on the transparency of the first region, and the model may increase the transparency of the first region; conversely, if the first attention is low, the model may reduce the transparency of the first region.

[0065] In this embodiment, the input text description is analyzed by the gating network to determine whether the text contains semantic information related to transparency, and a gating coefficient g is output. The adaptive attention mask of the gating network can adaptively adjust the attention weight of "text-attention-transparency channel" dynamically according to the gating coefficient g. When g≈0 (the text description is irrelevant to transparency): the attention mask will block the attention of "text-attention-transparency channel", that is, prevent the text information from directly affecting the generation of the transparency channel. When g≈1 (the text description contains transparency-related semantics): the attention mask allows the attention of "text-attention-transparency channel" to flow partially or completely, that is, allows the text information to guide the generation of the transparency channel.

[0066] In the technical solution of this embodiment, a feature vector of the target text is extracted by a text feature extraction module to obtain a target text feature vector; the noise residual of the first initial noise video is predicted by a diffusion transformation module to obtain a first noise residual, and the first noise residual is used to indicate the difference between the first initial noise video and the target video; a gating coefficient is determined based on the target text feature vector, and the first attention of the transparency channel of the first region of any first initial noise video frame is adjusted based on the gating coefficient, and the transparency of the content of the first region of any first initial noise video frame is adjusted based on the first attention. The gating coefficient is used to represent the influence degree of the target text on the transparency channel; denoising processing is performed based on the first noise residual by an inverse diffusion module to set the content filled in the second region of any first initial noise video frame as a transparent element; the first initial noise video with the transparency of the content of the first region adjusted and the transparent element of the second region set is output by an output module as the target video. In this way, not only can a video with a transparent element in the second region be obtained, but also the transparency of the visual elements in the first region can be adjusted according to the text input by the user, so that the content of the first region part of the generated target video is more realistic.

[0067] It should be understood that although the model architecture of the target video generation model shown in the figure includes multiple modules, a single module can also be set, which is not limited here.

[0068] Exemplarily, the gating coefficient includes a first gating coefficient or a second gating coefficient. Determining the gating coefficient based on the target text feature vector includes: If it is determined based on the target text feature vector that the target text includes a preset keyword, the gating coefficient is determined to be the first gating coefficient, and the preset keyword includes words related to transparency; if it is determined based on the target text feature vector that the target text does not include a preset keyword, the gating coefficient is determined to be the second gating coefficient, and the influence degree represented by the first gating coefficient is greater than the influence degree represented by the second gating coefficient.

[0069] It should be understood that this embodiment is not limited to two gating coefficients, but different situations may have different gating coefficients, so as to guide the model to generate the target video and thus improve the accuracy of the generated target video.

[0070] In this embodiment, if it is determined that the target text includes a preset keyword based on the target text feature vector, the gating coefficient is determined to be the first gating coefficient, and the preset keyword includes a word related to transparency; if it is determined that the target text does not include the preset keyword based on the target text feature vector, the gating coefficient is determined to be the second gating coefficient, and the degree of influence represented by the first gating coefficient is greater than the degree of influence represented by the second gating coefficient. In this way, when the target text includes a keyword related to transparency, the first region part of the generated target video also has a certain transparency, making the generated video content more realistic.

[0071] In a possible implementation manner, predicting the noise residual of the first initial noise video to obtain the first noise residual includes: Performing downsampling processing on the first initial noise video to obtain the downsampled first initial noise video; performing upsampling processing on the downsampled first initial noise video to obtain the upsampled first initial noise video; predicting the noise residual of the upsampled first initial noise video to obtain the first noise residual.

[0072] It should be noted that the above steps may be implemented by one or more modules in the diffusion transformation module, which can be set according to needs and are not limited here. In a possible implementation manner, the above steps may be implemented by multiple modules in the diffusion transformation module.

[0073] Please continue to refer to Figure 2 , the diffusion transformation module includes a downsampling network, an upsampling network, and a gating network. Predicting the noise residual of the first initial noise video to obtain the first noise residual includes: Performing downsampling processing on the first initial noise video through the downsampling network to obtain the downsampled first initial noise video; performing upsampling processing on the downsampled first initial noise video through the upsampling network to obtain the upsampled first initial noise video, and the upsampled first initial noise video is used as the basis for predicting the first noise residual.

[0074] Among them, downsampling is the process of downsampling image or video data, aiming to reduce the resolution of the data. In this process, the downsampling network performs downsampling on the first initial noisy video to obtain the first initial noisy video after downsampling with a lower resolution. Downsampling can be achieved by methods such as average pooling, max pooling, etc., depending on the design of the downsampling network. Upsampling is the process of upsampling image or video data, aiming to increase the resolution of the data. In this process, the upsampling network performs upsampling on the first initial noisy video after downsampling to restore its resolution to the original video or close to the resolution of the original video. Upsampling can be achieved by methods such as bilinear interpolation, nearest neighbor interpolation, transposed convolution, etc., depending on the design of the upsampling network.

[0075] In this embodiment, the downsampling network performs downsampling on the first initial noisy video to obtain the first initial noisy video after downsampling; the upsampling network performs upsampling on the first initial noisy video after downsampling to obtain the first initial noisy video after upsampling. The first initial noisy video after upsampling is used as the basis for predicting the first noise residual. The downsampled first initial noisy video can focus on its global information, while the upsampled first initial noisy video can focus on local information. That is to say, the global information and local information of the first initial noisy video can be used as the basis for predicting the first noise residual, thereby guiding denoising, which can improve the accuracy of setting the content of the second region as a transparent element.

[0076] Determining a gating coefficient based on the target text feature vector, and adjusting the first attention to the transparency channel of the first region of any first initial noisy video frame based on the gating coefficient, and adjusting the transparency of the content of the first region of any first initial noisy video frame based on the first attention, including: determining the gating coefficient based on the target text feature vector through a gating network, and adjusting the first attention to the transparency channel of the first region of any first initial noisy video frame based on the gating coefficient, and adjusting the transparency of the content of the first region of any first initial noisy video frame based on the first attention.

[0077] In this embodiment, the gating coefficient is determined based on the target text feature vector through the gating coefficient, and the first attention to the transparency channel of the first region of any first initial noisy video frame is adjusted based on the gating coefficient, and the transparency of the content of the first region of any first initial noisy video frame is adjusted based on the first attention. This can make the adjustment of the first noise residual and the transparency of the content of the first region proceed synchronously, which can improve the efficiency of video generation.

[0078] In a possible implementation manner, when performing downsampling or upsampling, at least one of the following is executed: Capture the intra-frame dependencies and inter-frame dependencies of the first initial noisy video. The intra-frame dependencies are used to indicate the relationships between different regions within the same frame, and the inter-frame dependencies are used to indicate the relationships between different video frames. The intra-frame dependencies and inter-frame dependencies serve as references for upsampling or downsampling processing; Fuse the image feature information of different resolutions of any first initial noisy video frame. The image feature information of different resolutions serves as a reference for upsampling or downsampling processing; Introduce a token embedding for the transparency channel. The token embedding is a learnable domain embedding to distinguish multiple color channels and the transparency channel through the token embedding. The token embedding serves as a reference for upsampling or downsampling processing; Fuse the low-rank adaptation parameters with the frozen base parameters of the target video generation model to perform upsampling or downsampling processing on the first initial noisy video through the fused parameters after fusing the low-rank adaptation parameters and the base parameters.

[0079] It should be noted that the above steps can be implemented by one or more modules in the downsampling network or the upsampling network, which can be set according to needs and are not limited here. In one possible implementation, the above steps can be implemented by multiple modules in the downsampling network or the upsampling network.

[0080] In one possible implementation, at least one of the downsampling network or the upsampling network includes: Transformer blocks, and each Transformer block includes at least one of the following: A spatio-temporal attention sub-module for capturing the intra-frame dependencies and inter-frame dependencies of the first initial noisy video. The intra-frame dependencies are used to indicate the relationships between different regions within the same frame, and the inter-frame dependencies are used to indicate the relationships between different video frames; a multi-scale feature fusion sub-module for fusing the image feature information of different resolutions of any first initial noisy video frame; an embedding sub-module for introducing a token embedding for the transparency channel. The token embedding is a learnable domain embedding to distinguish multiple color channels and the transparency channel through the token embedding; a low-rank adaptation (LoRA) sub-module including low-rank adaptation parameters, and the low-rank adaptation parameters are used to fuse with the frozen base parameters of the target video generation model to process the first initial noisy video through the fused parameters after fusing the low-rank adaptation parameters and the base parameters.

[0081] Among them, the spatio-temporal attention sub-module can adopt multi-head attention division. Within a Transformer Block, let the total number of multi-heads be . Among them heads focus on the space (dependencies between Patches within the same frame), For the head focusing on time (the evolution of adjacent frames or the same pixel position over time), the remaining heads can perform hybrid attention. Specifically, the Transformer Block can adopt the U-Net structure of image diffusion, but each U-Net Block contains a spatio-temporal integrated Transformer module to capture both intra-frame and inter-frame dependencies simultaneously. Input: (Batch size B, number of frames T, RGBA channel number 4, height H, width W), and then split into patches or directly perform convolutional mapping to intermediate features; Output: The noise predicted by the model or directly predict . During training, it is compared with the real noise to calculate the mean squared error.

[0082] Multi-scale feature fusion sub-module: When the intermediate hidden representation of the Transformer is output at each layer, perform multi-level downsampling (e.g., 1 / 2, 1 / 4 resolution), and fuse them through Cross-Attention or Skip-Connection. In practice, the U-Net or pyramid method can be referred to. In order not to over-blur the Alpha channel at high resolution, more attention heads can also be specifically allocated at the shallow high resolution. Specifically, in the U-Net structure, after each downsampling layer, multi-resolution features are obtained; in the upsampling stage, they are fused with the features of the previous layer through skip-connection. Keep the transparent channel at a higher resolution in the shallow features to ensure that the model can better learn the transparency of fine edges.

[0083] Principle of LoRA: Traditional fine-tuning modifies the large matrix of the Transformer . LoRA adds a small low-rank matrix , and only adds the increment during the derivation of Q, K, V, which can effectively reduce the training parameters and prevent overfitting when the data is limited.

[0084] In this embodiment, the spatial dependencies within the video frame and the temporal dependencies between frames are combined. To address issues such as jitter that are prone to occur at the transparent edges, multi-scale features are used to stably generate the channels. In the application of hierarchical LoRA in the Alpha path, it focuses on the detailed contours (edge shapes) in the shallow layer and the temporal consistency in the deep layer. Only perform low-rank fine-tuning on the Query / Key / Value matrices related to Alpha, significantly reducing the number of training parameters and effectively solving the problem of insufficient data.

[0085] In this embodiment, the meaning of multi-scale features: In image or video processing, "multi-scale features" refer to features extracted from different scaled versions (different resolutions) of the same image or video frame. For example, for an original image, it can be scaled to 1 / 2 size, 1 / 4 size, etc., and then features are extracted from the original size and the images of different scaled sizes respectively. The purpose of multi-scale feature fusion: The purpose of multi-scale feature fusion is to enable the model to capture information at different scales of the image or video content simultaneously. Low-resolution features: Capture global and macroscopic information, such as the overall shape of an object, the layout of a scene, etc. High-resolution features: Capture local and fine information, such as the detailed texture of an object, the edge contour, etc.

[0086] In transparent video generation, multi-scale feature fusion helps to stabilize the generation of the transparency channel and improve the quality of the transparent edges. Stabilize the transparency channel: Low-resolution features can provide global and macroscopic scene information, helping the model to better understand the overall structure of the video content, thereby stabilizing the generation of the transparency channel and reducing inter-frame jitter. Improve edge quality: High-resolution features can capture more detailed edge contour information, helping the model to more accurately predict the transparency of the transparent edges and improve the clarity and quality of the transparent edges. Spatiotemporal integrated attention and multi-scale feature fusion are two different technical modules, which work together to jointly improve the generation quality of transparent videos. Spatiotemporal integrated attention: Responsible for modeling the information dependence relationships within and between video frames to ensure the spatiotemporal consistency of the generated video. Multi-scale feature fusion: Responsible for extracting and fusing features at different scales, enabling the model to capture information at different scales simultaneously and improving the generation quality of the transparency channel. Spatiotemporal integrated attention can act on multi-scale features, that is, applying the spatiotemporal attention mechanism on feature maps of different resolutions, so as to ensure spatiotemporal consistency at different scales. And multi-scale feature fusion provides richer input information for spatiotemporal attention, enabling spatiotemporal attention to play a better role. In this embodiment, the spatial dependence within the video frame and the temporal dependence between frames are combined. Aiming at problems such as jitter that are prone to occur in the transparent edges, multi-scale features are used to stabilize the channel generation.

[0087] It should be noted that assuming that when performing downsampling processing or upsampling processing, the intra-frame dependence and inter-frame dependence of the first initial noise video are captured, the image feature information of different resolutions of any first initial noise video frame is fused, a marker embedding is introduced into the transparency channel, and the low-rank adaptation parameters are fused with the frozen basic parameters of the target video generation model, then when performing upsampling processing or downsampling processing, upsampling processing or downsampling processing is performed based on the fused parameters, intra-frame dependence and inter-frame dependence, image feature information of different resolutions, and marker embedding.

[0088] In a possible implementation, the rank of the low-rank adaptation parameter corresponding to the downsampling process is less than the rank of the low-rank adaptation parameter corresponding to the upsampling process. In other words, the rank of the low-rank adaptation parameter in the downsampling network is less than the rank of the low-rank adaptation parameter in the upsampling network.

[0089] Among them, the low-rank adaptation (LoRA) parameter in the downsampling network can be a shallow LoRA, and the downsampling network can be a deep LoRA, which can also be understood as hierarchical LoRA.

[0090] Hierarchical LoRA: Divide the Transformer into the front layers (shallow layers) and the back layers (deep layers). Set a larger rank (such as 32) in the shallow LoRA to specifically adapt to local edge deformations; set a smaller rank (such as 8) in the deep layer to focus on global temporal relationships. Only perform LoRA on the projections related to the Alpha marker, such as: , and the original pre-trained weights of the remaining RGB channels or text channel paths can remain unchanged, maximizing the retention of the original model's ability in RGB generation. Assume that the entire model has 12 layers (4 layers of Down sample + 4 layers of Middle + 4 layers of Up sample). Set the rank = 32 in the first 4 layers (shallow layers) of LoRA, mainly focusing on local shapes and Alpha edges; set the rank = 8 in the last 4 layers (deep layers) of LoRA, and learn more about global temporal consistency.

[0091] Specifically, "shallow layer" and "deep layer" refer to different levels in the diffusion transformer model (U-Net structure + Transformer Blocks). In a typical U-Net structure, it usually includes a downsampling (Down sample) stage, a middle layer (Middle Layers) stage, and an upsampling (Up sample) stage. Shallow Layers: Refer to the layers close to the input end in the U-Net structure, usually referring to the Transformer Blocks in the downsampling stage. Characteristics: Process high-resolution feature maps, have a small receptive field, and mainly capture local and detailed information, such as edge contours, texture details, etc. Role in transparent video generation: Focusing on detail contours (edge morphology) in the shallow layer means that in the shallow layer of the U-Net, the model mainly learns how to accurately predict the edge contours of transparent objects and local transparency changes.

[0092] A relatively large rank is set for the shallow LoRA (e.g., 32) so that the model can have sufficient parameters in the shallow layer to capture diverse local deformations and edge details. Deep Layers: Refer to the layers in the U-Net structure closer to the output end, usually referring to the Transformer Blocks in the upsampling stage and the intermediate layer stage. Characteristics: Deal with low-resolution feature maps, have a large receptive field, and mainly capture global and macroscopic information, such as the overall shape of an object, the layout of a scene, and the temporal dependencies between frames. Role in transparent video generation: Paying attention to temporal consistency in the deep layer means that in the deep layer of the U-Net, the model mainly learns how to ensure the temporal coherence and smoothness of the generated video. For example, ensure the smooth movement trajectory of transparent objects between different frames and avoid inter-frame jitter of transparent edges. A relatively small rank is set for the deep LoRA (e.g., 8) so that the model can focus more on learning global temporal features in the deep layer without overly concerning local detail changes. Hierarchical LoRA fine-tuning strategy: The hierarchical LoRA fine-tuning strategy refers to applying LoRA adapters with different ranks at different levels (shallow and deep) of the U-Net. Only performing low-rank fine-tuning on the Query / Key / Value matrices related to Alpha is for parameter efficiency and preventing overfitting. Since the training data containing transparent channels is limited, to avoid overfitting caused by training the entire model from scratch on a small amount of data, we only fine-tune these small amounts of parameters of the LoRA adapter while keeping most of the weights of the pre-trained model unchanged, which can significantly reduce the number of training parameters and effectively solve the problem of insufficient data, while also maintaining the powerful ability of the pre-trained model in RGB generation.

[0093] In one possible implementation, the method further includes: Obtain at least one adjusted target video frame; retrain the target video generation model using the at least one target video frame.

[0094] In this embodiment, if the user is not satisfied with the output target video frame, the at least one target video frame can be adjusted, and then the target video generation model can be retrained using the at least one target video frame, thereby improving the ability of the target video generation model to generate videos.

[0095] In this embodiment, only perform attention update on the Alpha markers within the mask area; keep the rest frozen; perform several steps (e.g., 10 - 20 steps) of diffusion reverse iteration to correct. This can both retain the generated complete timings and fine-tune the local contours.

[0096] Next, an actual inference process will be used as an example for illustration.

[0097] Generate a text description: "A floating glass sphere with swirling semi-transparent smoke inside."

[0098] At the text encoding stage, input the above sentence into the CLIP text encoder to obtain the text vector. . At the same time, pass through the gating network. Obtain value. If the system detects keywords such as "glass", "semi-transparent", "sphere", etc., usually , it means allowing the text to have a significant impact on Alpha.

[0099] Then, sample a purely random noise video (T = 16 frames, 4 channels RGBA).

[0100] In the reverse diffusion iteration process, from to , gradually execute the denoising steps: . In each step: The improved diffusion transformer uses spatio-temporal attention to generate intermediate features; based on gate the Text→Alpha attention; output the noise residual , and calculate according to the diffusion formula.

[0101] The final output can be , including 16 frames, each frame being a 4-channel image of 256×256. It can be combined into a 2-second video (frame rate 16 FPS). When viewing the Alpha channel, it will be found that the sphere and the internal smoke are semi-transparent, and the Alpha of the second region is almost 0.

[0102] For ease of understanding, an embodiment of the present application provides one frame of the target video frame for illustration. As Figure 3 shown, Figure 3 is a schematic diagram of one frame of the target video frame provided by an embodiment of the present application. As Figure 3 shown in the schematic diagram, the content of the first region matches "A floating glass sphere with swirling semi-transparent smoke inside.", while the second region includes transparent elements.

[0103] Generally speaking, this embodiment proposes a text-to-transparent video generation scheme based on an improved diffusion transformer, which can directly output a short video with an Alpha channel (RGBA format). The core lies in introducing "transparent markers" and adaptive attention masks, combining spatio-temporal integrated attention with multi-scale feature fusion to accurately process semi-transparent edges and dynamic changes. In addition, through hierarchical LoRA fine-tuning, only the network parameters related to Alpha are low-rank adapted, greatly reducing the training cost and ensuring edge consistency and jitter stability. Compared with traditional methods of "generating RGB first and then matting" or simple multi-channel expansion, the present invention has significantly improved transparency channel quality, temporal consistency, and efficiency, and is widely applicable to scenarios such as film and television special effects, AR / VR, and game animations, greatly simplifying the post-processing process; at the same time, partial regions can be quickly corrected through local resampling, improving flexibility and controllability.

[0104] If the user finds during preview that: the edge of a certain frame of the sphere is overly transparent and appears "broken"; or the internal smoke effect is too light; an interactive annotation tool can be used to draw a mask in the 5th to 6th frames to enclose the area that needs to be corrected. The system unfreezes the Alpha Tokens in the masked area and performs 10 to 20 steps of local reverse diffusion: the attention of other areas and RGB channels remains frozen; only the Alpha features at the mask are denoised iteratively several times to make the transparency of the area more in line with expectations. Finally, the corrected video segment is obtained, and the generation results of other frames or regions are not damaged.

[0105] The above embodiment illustrates how to generate a video from the input text after the model training is completed. In the following embodiment, an exemplary illustration of how to train the model is provided.

[0106] Next, an explanation of how the target video generation model is trained is given.

[0107] Please refer to Figure 4 , Figure 4 which is a schematic flowchart of a model training method provided by an embodiment of the present application. The model training method of this embodiment can be executed by a server, or by a client device, or jointly by a server and a client device. As Figure 4 shown, the model training method includes: S410. Obtain a plurality of training samples and an initial video generation model. Any training sample includes a second initial noise video and a text sample corresponding to the second initial noise video. The second initial noise video is a video including a plurality of color channels and a transparency channel. The second initial noise video includes multiple frames of second initial noise video frames. The content of the first region of any second initial noise video frame is generated according to the text sample, and the second region of any second initial noise video frame is filled with content.

[0108] Among them, the initial video generation model can be a predefined neural network model. The content of the first region is generated according to the text sample, and it can be understood that the content of the first region includes an image (also known as a visual element) generated according to the text sample.

[0109] S420. Iteratively train the initial video generation model using multiple training samples until the training end condition is met to obtain the target video generation model.

[0110] In this embodiment, the training process is carried out through multiple iterations. In each iteration, a batch of training samples is used to update the weights of the model. In each iteration, the model receives a second initial noise video and the corresponding text sample as inputs. The model attempts to at least adjust at least part of the second region of the noise video according to the content of the text sample to generate a video frame that matches the text description and has a transparent second region. The training end condition can be a condition for determining whether the model has finished training, for example, including a condition for determining whether the model has converged.

[0111] In this embodiment, by obtaining multiple training samples and the initial video generation model, any training sample includes a second initial noise video and the text sample corresponding to the second initial noise video. The second initial noise video is a video containing multiple color channels and a transparency channel. The second initial noise video includes multiple frames of second initial noise video frames. The first region of any second initial noise video frame includes an image related to the text sample, and the second region of any second initial noise video frame is filled with content; iteratively train the initial video generation model using multiple training samples until the training end condition is met to obtain the target video generation model. In this way, when inferring using the trained target video generation model, the output target video is a video including a transparent second region, and there is no need for the user to perform further processing to obtain a video with a transparent second region.

[0112] In a possible implementation manner, the multiple training samples are obtained through the following method: Obtain multiple reference videos. Each reference video corresponds to a description text. The reference video includes multiple frames of reference video frames. For any reference video frame, the content of the first region of any reference video frame is generated according to the text sample, and the second region of any reference video frame includes a transparent element; for any reference video, perform noise addition processing on any reference video to obtain a second initial noise video, and determine the description text corresponding to any reference video as the text sample corresponding to the second initial noise video.

[0113] In this embodiment, performing noise addition processing on any reference video can be adding content to the second region of the reference video.

[0114] In this embodiment, by obtaining a plurality of reference videos, each reference video is associated with a description text. The reference video includes multiple frames of reference video frames. For any reference video frame, the content of the first region of any reference video frame is generated according to a text sample, and the second region of any reference video frame includes a transparent element. For any reference video, the reference video is subjected to noise addition processing to obtain a second initial noisy video, and the description text corresponding to any reference video is determined as the text sample corresponding to the second initial noisy video. In this way, the simplicity of obtaining the reference video and the second initial noisy video can be reduced.

[0115] In another possible implementation manner, it may also be to first obtain the second initial noisy video, and then denoise the second initial noisy video to obtain the reference video.

[0116] In one possible implementation manner, an initial video generation model is iteratively trained using a plurality of training samples until a training end condition is satisfied to obtain a target video generation model, including: Using a first part of the plurality of training samples to train the initial video generation model to obtain a predicted video output by the initial video generation model based on the second initial noisy video and the corresponding text sample in the first part of the samples; calculating a training loss based on the predicted video and the reference video corresponding to the first part of the samples; if it is determined based on the training loss that the training end condition is satisfied, then determining the initial video generation model as the target video generation model; if it is determined based on the training loss that the training end condition is not satisfied, then after updating the model parameters of the initial video generation model, using a second part of the plurality of training samples to train the initial video generation model.

[0117] In this embodiment, the training loss may be the similarity between the predicted video and the reference video. Specifically, the similarity may include a first similarity between the second regions of the predicted video and the reference video. In addition, the similarity may further include a second similarity between the first regions of the predicted video and the reference video. The model parameters may be adjustable values inside the model for capturing the patterns in the data. For example, in a neural network, the model parameters are weights and biases, which affect how the input data is transformed into the output result. In this embodiment, the model parameters affect the ability to adjust the transparency of the second region of the video. In addition, the model parameters may also affect the ability to adjust the transparency of the first region of the video.

[0118] In this embodiment, not only is the second initial noisy video generated using the reference video, but also the reference video is used to determine whether the model can end the training. In this way, the convenience of model training can be improved.

[0119] It should be noted that the model architecture of the initial video generation model and the target video generation model can be the same, but the specific model parameters are different. After the initial video generation model undergoes multiple iterations of training until the training end condition is met, the model parameters of the initial video generation model are consistent with those of the target generation model.

[0120] In a possible implementation, the predicted video output based on the second initial noise video and the corresponding text sample in the first part of the samples includes: Extract the feature vector of the text sample to obtain the text sample feature vector; predict the noise residual of the second initial noise video to obtain the second noise residual, where the second noise residual is used to indicate the difference between the second initial noise video and the reference video; determine the gating coefficient based on the text sample feature vector, and adjust the second attention to the alpha channel of the first region of any second initial noise video frame based on the gating coefficient, and adjust the transparency of the content of the first region of any second initial noise video frame based on the second attention. The gating coefficient is used to represent the influence degree of the text sample on the alpha channel; perform denoising processing based on the second noise residual to set the content filled in the second region of any second initial noise video frame as a transparent element; output the second initial noise video with the transparency of the content of the first region adjusted and the transparent element of the second region set as the predicted video.

[0121] For ease of understanding, an exemplary description of how to output a predicted video based on the second initial noise video and the corresponding text sample in the first part of the samples is given in combination with the model architecture of the initial video generation model.

[0122] Please continue to refer to Figure 2 ... In a possible implementation, the initial video generation model includes a text feature extraction module, a diffusion transformation module, an inverse diffusion module, and an output module; The predicted video output based on the second initial noise video and the corresponding text sample in the first part of the samples includes: Extract the feature vector of the text sample through the text feature extraction module to obtain the text sample feature vector; predict the noise residual of the second initial noise video through the diffusion transformation module to obtain the second noise residual, and the second noise residual is used to indicate the difference between the second initial noise video and the reference video; determine the gating coefficient based on the text sample feature vector, and adjust the second attention of the transparency channel of the first region of any second initial noise video frame based on the gating coefficient, and adjust the transparency of the content of the first region of any second initial noise video frame based on the second attention, and the gating coefficient is used to represent the influence degree of the text sample on the transparency channel; perform denoising processing based on the second noise residual through the reverse diffusion module to set the content filled in the second region of any second initial noise video frame as a transparent element; the output module is used to output the second initial noise video with the transparency of the content of the first region adjusted and the transparent element of the second region set as the predicted video.

[0123] Exemplarily, in combination with Figure 2 it can be said that the model parameters updated in this application may include but are not limited to the parameters of the diffusion transformer and / or the parameters of the reverse diffusion module.

[0124] In a possible implementation manner, the gating coefficient includes a third gating coefficient or a fourth gating coefficient. Determining the gating coefficient based on the text sample feature vector includes: If it is determined based on the text sample feature vector that the text sample includes a preset keyword, determine the gating coefficient as the third gating coefficient, and the preset keyword includes words related to transparency; if it is determined based on the text sample feature vector that the text sample does not include a preset keyword, determine the gating coefficient as the fourth gating coefficient, and the influence degree represented by the third gating coefficient is greater than the influence degree represented by the fourth gating coefficient.

[0125] Among them, this embodiment may refer to the relevant descriptions of the above embodiments and will not be elaborated here.

[0126] In a possible implementation manner, predicting the noise residual of the second initial noise video to obtain the second noise residual includes: Perform downsampling processing on the second initial noise video to obtain the downsampled second initial noise video; perform upsampling processing on the downsampled second initial noise video to obtain the upsampled second initial noise video; predict the noise residual of the upsampled second initial noise video to obtain the second noise residual.

[0127] For ease of understanding, the following will give an exemplary description of how to predict the noise residual of the second initial noise video in combination with the architecture of the diffusion transformation module.

[0128] In a possible implementation, the diffusion transformation module includes a downsampling network, an upsampling network, and a gating network; Predicting the noise residual of the second initial noise video to obtain a second noise residual includes: performing downsampling processing on the second initial noise video through the downsampling network to obtain the second initial noise video after downsampling processing; performing upsampling processing on the second initial noise video after downsampling processing through the upsampling network to obtain the second initial noise video after upsampling processing, and the second initial noise video after upsampling processing is used as the basis for predicting the second noise residual.

[0129] This embodiment can refer to the description of the above embodiment and will not be elaborated here.

[0130] Determining a gating coefficient based on the text sample feature vector, and adjusting the second attention of the alpha channel of the first region of any second initial noise video frame based on the gating coefficient, and adjusting the transparency of the content of the first region of any second initial noise video frame based on the second attention, includes: Determining a gating coefficient through the gating network based on the text sample feature vector, and adjusting the second attention of the alpha channel of the first region of any second initial noise video frame based on the gating coefficient, and adjusting the transparency of the content of the first region of any second initial noise video frame based on the second attention.

[0131] This embodiment can refer to the description of the above embodiment and will not be elaborated here.

[0132] In a possible implementation, when performing downsampling processing or upsampling processing, at least one of the following is performed: Capturing the intra-frame dependency and inter-frame dependency of the second initial noise video, where the intra-frame dependency is used to indicate the relationship between different regions in the same frame, and the inter-frame dependency is used to indicate the relationship between different video frames, and the intra-frame dependency and inter-frame dependency are used as references for upsampling processing or downsampling processing; fusing the image feature information of different resolutions of any second initial noise video frame, and the image feature information of different resolutions is used as a reference for upsampling processing or downsampling processing; introducing a token embedding for the alpha channel, where the token embedding is a learnable domain embedding to distinguish multiple color channels and the alpha channel through the token embedding, and the token embedding is used as a reference for upsampling processing or downsampling processing; fusing the low-rank adaptation parameters with the frozen base parameters of the initial video generation model to perform upsampling processing or downsampling processing on the second initial noise video through the fused parameters after fusing the low-rank adaptation parameters and the base parameters.

[0133] For ease of understanding, the following embodiments will exemplarily illustrate how to perform upsampling processing and downsampling processing in combination with the architectures of the downsampling network or the upsampling network.

[0134] In a possible implementation, at least one of the downsampling network or the upsampling network includes: a Transformer block, and each Transformer block includes at least one of the following: A spatio-temporal attention sub-module for capturing intra-frame dependencies and inter-frame dependencies of the second initial noisy video, where the intra-frame dependencies are used to indicate the relationships between different regions in the same frame, and the inter-frame dependencies are used to indicate the relationships between different video frames; a multi-scale feature fusion sub-module for fusing image feature information of different resolutions of any second initial noisy video frame; an embedding sub-module for introducing token embeddings for the transparency channel, where the token embeddings are learnable domain embeddings to distinguish multiple color channels and the transparency channel through the token embeddings; a low-rank adaptation sub-module including low-rank adaptation parameters, where the low-rank adaptation parameters are used to fuse with the frozen base parameters of the initial video generation model to process the second initial noisy video through the fused parameters after fusing the low-rank adaptation parameters and the base parameters.

[0135] This embodiment can refer to the description of the above embodiment and will not be elaborated here.

[0136] In a possible implementation, the rank of the low-rank adaptation parameters in the downsampling network is less than the rank of the low-rank adaptation parameters in the upsampling network.

[0137] This embodiment can refer to the description of the above embodiment and will not be elaborated here.

[0138] It should be noted that although this embodiment provides descriptions that the model includes multiple modules, the modules include multiple networks, and the networks include multiple sub-modules, the number of modules, networks, and sub-modules can be set as needed. For example, the model includes one module, and this one module implements the functions of the model; or the module includes one network, and this one network implements the functions of the module; or the network includes one sub-module, and this one sub-module implements the functions of the network, and there is no limitation here.

[0139] The method embodiments are described above, and the product embodiments will be described below.

[0140] Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of a video generation device provided by an embodiment of the present application. As Figure 5 shown, the device may include: A first acquisition module 510 for acquiring a target text, where the target text is used to indicate the generation of a video; a video generation module 520 for calling a trained target video generation model to generate a target video based on the target text, where the target video includes multiple target video frames, and the content of the first region of any target video frame is generated according to the target text, and the second region of any target video frame includes transparent elements.

[0141] The video generation device of this embodiment can refer to the description of the video generation method, which will not be elaborated here.

[0142] Please refer to Figure 6 , Figure 6 , which is a schematic structural diagram of a model training device provided by an embodiment of this application. As Figure 6 shown, the device may include: A second acquisition module 610, configured to acquire a plurality of training samples and an initial video generation model. Any training sample includes a second initial noise video and a text sample corresponding to the second initial noise video. The second initial noise video is a video including a plurality of color channels and a transparency channel. The second initial noise video includes multiple frames of second initial noise video frames. The content of the first region of any second initial noise video frame is generated according to the text sample, and the second region of any second initial noise video frame is filled with content; a model training module 620, configured to iteratively train the initial video generation model using the plurality of training samples until a training end condition is satisfied to obtain a target video generation model.

[0143] The model training device of this embodiment can refer to the description of the model training method, which will not be elaborated here.

[0144] This embodiment further provides an electronic device, including a processor and a memory. The memory stores machine-executable instructions that can be executed by the processor. The processor executes the machine-executable instructions to implement the above image editing method. The electronic device can be a server or a terminal device.

[0145] Refer to Figure 7 shown. The electronic device includes a processor 100 and a memory 101. The memory 101 stores machine-executable instructions that can be executed by the processor 100. The processor 100 executes the machine-executable instructions to implement the steps of the above method.

[0146] Furthermore, Figure 7 the electronic device shown further includes a bus 102 and a communication interface 103. The processor 100, the communication interface 103, and the memory 101 are connected through the bus 102.

[0147] Among them, the memory 101 may include high-speed random access memory (RAM), and may also include non-volatile memory, such as at least one disk memory. The communication connection between this system network element and at least one other network element is realized through at least one communication interface 103 (which can be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. can be used. The bus 102 can be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 7 only a bidirectional arrow is used in Figure 7 , but it does not mean that there is only one bus or one type of bus.

[0148] The processor 100 may be an integrated circuit chip with signal processing capabilities. In the implementation process, the steps of the above method can be completed by the integrated logic circuit in the hardware of the processor 100 or the instructions in the form of software. The above-mentioned processor 100 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field-programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, registers, etc. This storage medium is located in the memory 101, and the processor 100 reads the information in the memory 101 and combines its hardware to complete the steps of the method in the foregoing embodiments.

[0149] This embodiment also provides a machine-readable storage medium. The machine-readable storage medium stores machine-executable instructions. When the machine-executable instructions are called and executed by the processor, the machine-executable instructions cause the processor to implement the steps of the above method.

[0150] This embodiment also provides a computer program product. The computer program product includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions that, when executed by a computer device, cause the computer device to perform the steps of the above-described method.

[0151] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems and devices described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0152] In addition, in the description of the embodiments of the present application, unless otherwise clearly specified and limited, the terms "installation", "connection", and "coupling" shall be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be a direct connection or an indirect connection through an intermediate medium, and it may be the communication inside two components. For those skilled in the art, the specific meanings of the above terms in the present application can be understood according to specific situations.

[0153] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0154] In the description of the present application, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be understood as a limitation to the present application. In addition, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.

[0155] Finally, it should be noted that the above embodiments are only specific implementation manners of the present application, used to illustrate the technical solutions of the present application, rather than limiting it. The protection scope of the present application is not limited thereto. Although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art within the technical scope disclosed by the present application can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. A video generation method, characterized in that, The method includes: Obtaining a target text, where the target text is used to indicate the generation of a video; Invoking a trained target video generation model to generate a target video based on the target text, where the target video includes multiple frames of target video frames, the content of the first region of any target video frame is generated according to the target text, and the second region of any target video frame includes a transparent element.

2. The method according to claim 1, wherein The step of invoking a trained target video generation model to generate a target video based on the target text includes: Obtaining a first initial noise video, where the first initial noise video is a video including multiple color channels and an alpha channel, the first initial noise video is generated based on the target text, the first initial noise video includes multiple frames of first initial noise video frames, the content of the first region of any first initial noise video frame is generated according to the target text, and the second region of any first initial noise video frame is filled with content; Invoking a trained target video generation model to generate the target video based on the target text and the first initial noise video, where the target video is a video including the multiple color channels and the alpha channel.

3. The method according to claim 2, characterized in that, The generating the target video based on the target text and the first initial noise video includes: Extracting a feature vector of the target text to obtain a target text feature vector; Predicting a noise residual of the first initial noise video to obtain a first noise residual, where the first noise residual is used to indicate the difference between the first initial noise video and the target video; determining a gating coefficient based on the target text feature vector, and adjusting a first attention to the alpha channel of the first region of any first initial noise video frame based on the gating coefficient, and adjusting the transparency of the content of the first region of any first initial noise video frame based on the first attention, where the gating coefficient is used to represent the influence degree of the target text on the alpha channel; Performing denoising processing based on the first noise residual to set the content filled in the second region of any first initial noise video frame as a transparent element; Outputting the first initial noise video with the transparency of the content of the first region adjusted and the transparent element of the second region set as the target video.

4. The method according to claim 3, characterized in that, The gating coefficient includes a first gating coefficient or a second gating coefficient, and the determining the gating coefficient based on the target text feature vector includes: If it is determined based on the target text feature vector that the target text includes a preset keyword, determining the gating coefficient as the first gating coefficient, where the preset keyword includes a word related to transparency; If it is determined based on the target text feature vector that the target text does not include a preset keyword, determining the gating coefficient as the second gating coefficient, where the influence degree represented by the first gating coefficient is greater than the influence degree represented by the second gating coefficient.

5. The method according to claim 3, wherein The predicting the noise residual of the first initial noise video to obtain a first noise residual includes: Performing downsampling processing on the first initial noise video to obtain a downsampled first initial noise video; Perform upsampling on the first initial noise video after the downsampling process to obtain the first initial noise video after the upsampling process; Predict the noise residual of the first initial noise video after the upsampling process to obtain the first noise residual.

6. The method according to claim 5, wherein When performing the downsampling process or the upsampling process, perform at least one of the following: Capture the intra-frame dependency and inter-frame dependency of the first initial noise video, where the intra-frame dependency is used to indicate the relationship between different regions in the same frame, the inter-frame dependency is used to indicate the relationship between different video frames, and the intra-frame dependency and inter-frame dependency are used as references for the upsampling process or the downsampling process; Fuse the image feature information of different resolutions of any first initial noise video frame, and the image feature information of different resolutions is used as a reference for the upsampling process or the downsampling process; Introduce a marker embedding for the transparency channel, where the marker embedding is a learnable domain embedding to distinguish the multiple color channels and the transparency channel through the marker embedding, and the marker embedding is used as a reference for the upsampling process or the downsampling process; Fuse the low-rank adaptation parameters with the frozen basic parameters of the target video generation model, so as to perform upsampling or downsampling on the first initial noise video through the fusion parameters after fusing the low-rank adaptation parameters and the basic parameters.

7. The method according to any one of claims 1-6, characterized in that, The method further includes: Obtain at least one adjusted target video frame; Retrain the target video generation model by using the at least one target video frame.

8. The method according to claim 1, wherein The target video generation model is trained in the following manner: Obtain a plurality of training samples and an initial video generation model. Any training sample includes a second initial noise video and the text sample corresponding to the second initial noise video. The second initial noise video is a video including multiple color channels and a transparency channel. The second initial noise video includes multiple frames of second initial noise video frames. The content of the first region of any second initial noise video frame is generated according to the text sample, and the second region of any second initial noise video frame is filled with content; Iteratively train the initial video generation model by using the plurality of training samples until the training end condition is satisfied to obtain the target video generation model.

9. The method according to claim 8, wherein The plurality of training samples are obtained in the following manner: Obtain a plurality of reference videos, each reference video corresponding to a description text. The reference video includes multiple frames of reference video frames. For any reference video frame, the content of the first region of the any reference video frame is generated according to the text sample, and the second region of the any reference video frame includes a transparent element; For any reference video, perform noise addition processing on the any reference video to obtain the second initial noise video, and determine the description text corresponding to the any reference video as the text sample corresponding to the second initial noise video.

10. The method according to claim 9, wherein The iteratively training the initial video generation model by using the plurality of training samples until the training end condition is satisfied to obtain the target video generation model includes: Training the initial video generation model using the first part of the multiple training samples to obtain a predicted video output by the initial video generation model based on the second initial noise video and the corresponding text sample in the first part of the samples; Calculating a training loss based on the predicted video and the reference video corresponding to the first part of the samples; If it is determined that the training end condition is satisfied based on the training loss, determining the initial video generation model as the target video generation model; If it is determined that the training end condition is not satisfied based on the training loss, after updating the model parameters of the initial video generation model, training the initial video generation model using the second part of the multiple training samples.

11. The method according to any one of claims 8-10, characterized in that, The predicted video output based on the second initial noise video and the corresponding text sample in the first part of the samples includes: Extracting the feature vector of the text sample to obtain a text sample feature vector; Predicting the noise residual of the second initial noise video to obtain a second noise residual, where the second noise residual is used to indicate the difference between the second initial noise video and the reference video; determining a gating coefficient based on the text sample feature vector, and adjusting the second attention to the alpha channel of the first region of any second initial noise video frame based on the gating coefficient, and adjusting the transparency of the content of the first region of any second initial noise video frame based on the second attention, where the gating coefficient is used to represent the influence degree of the text sample on the alpha channel; Performing denoising processing based on the second noise residual to set the content filled in the second region of any second initial noise video frame as a transparent element; Outputting the second initial noise video with the transparency of the content of the first region adjusted and the transparent element of the second region set as the predicted video.

12. The method according to claim 11, wherein The gating coefficient includes a third gating coefficient or a fourth gating coefficient, and determining the gating coefficient based on the text sample feature vector includes: If it is determined based on the text sample feature vector that the text sample includes a preset keyword, determining the gating coefficient as the third gating coefficient, where the preset keyword includes a word related to transparency; If it is determined based on the text sample feature vector that the text sample does not include a preset keyword, determining the gating coefficient as the fourth gating coefficient, where the influence degree represented by the third gating coefficient is greater than the influence degree represented by the fourth gating coefficient.

13. The method according to claim 11, wherein Predicting the noise residual of the second initial noise video to obtain a second noise residual includes: Performing downsampling processing on the second initial noise video to obtain a downsampled second initial noise video; Performing upsampling processing on the downsampled second initial noise video to obtain an upsampled second initial noise video; Predicting the noise residual of the upsampled second initial noise video to obtain a second noise residual.

14. The method according to claim 13, wherein When performing downsampling processing or upsampling processing, performing at least one of the following: Capture the intra-frame dependencies and inter-frame dependencies in the second initial noisy video, where the intra-frame dependencies are used to indicate the relationships between different regions within the same frame, the inter-frame dependencies are used to indicate the relationships between different video frames, and the intra-frame dependencies and inter-frame dependencies are used as references for upsampling or downsampling processes; Fuse the image feature information with different resolutions in any of the second initial noisy video frames, and the image feature information with different resolutions is used as a reference for upsampling or downsampling processes; Introduce a token embedding for the transparency channel, where the token embedding is a learnable domain embedding to distinguish the multiple color channels and the transparency channel through the token embedding, and the token embedding is used as a reference for upsampling or downsampling processes; Fuse the low-rank adaptation parameters with the frozen base parameters of the initial video generation model, and perform upsampling or downsampling processes on the second initial noisy video using the fused parameters obtained by fusing the low-rank adaptation parameters and the base parameters.

15. The method according to claim 6 or 14, characterized in that The rank of the low-rank adaptation parameters corresponding to the downsampling process is less than the rank of the low-rank adaptation parameters corresponding to the upsampling process.

16. A video generation device, characterized in that, The apparatus includes: A first acquisition module for acquiring a target text, where the target text is used to indicate the generation of a video; A video generation module for calling a trained target video generation model to generate a target video based on the target text, where the target video includes multiple target video frames, the content of the first region of any target video frame is generated according to the target text, and the second region of any target video frame includes a transparent element.

17. An electronic device, characterized in that, It includes a processor and a memory, where the memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the video generation method according to any one of claims 1-15.

18. A machine-readable storage medium, characterized in that, The machine-readable storage medium stores machine-executable instructions, and when the machine-executable instructions are called and executed by a processor, the machine-executable instructions cause the processor to implement the video generation method according to any one of claims 1-15.

19. A computer program product, characterized in that, The computer program product includes a computer program stored on a computer-readable storage medium, where the computer program includes program instructions, and when the program instructions are executed by a computer device, the computer device is caused to execute the video generation method according to any one of claims 1-15.