Video generation method and device, electronic equipment and storage medium
By extracting and updating the features of the initial noise sequence and integrating the constraints of the time dimension, the problem of inter-frame incoherence in the DiT model is solved, high-quality target videos are generated, and the viewing experience and practicality of the video are improved.
Patent Information
- Application Number
- CN202511003873.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-10-03
AI Technical Summary
When generating videos, the existing DiT model lacks effective temporal constraints between frames, resulting in random fluctuations in content, texture, color and other elements between adjacent frames, severe flickering in specific areas, and incoherent motion, which affects the viewing experience and practicality.
By obtaining the video description text input by the user, an initial noise sequence is generated, and its features are gradually extracted and updated using a preset update strategy. The constraints of the time dimension are integrated, and differentiated processing is performed on regions with different spatiotemporal characteristics to generate a target noise sequence that meets the preset requirements and finally convert it into a target video.
It effectively reduces flickering, enhances motion continuity and video viewing experience, improves the clarity and smoothness of generated videos, and ensures the continuity and consistency of video content.
Smart Images

Figure CN120751213A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of video processing technology, and in particular to a video generation method, device, electronic device and storage medium. Background Art
[0002] With the rapid development of artificial intelligence technology, text-to-video generation technology has made significant progress. Among them, the video generation method based on diffusion models has attracted widespread attention due to its high generation quality and strong diversity.
[0003] However, when generating videos, the existing DiT model lacks effective temporal dimension constraints between frames, and the existing attention mechanism mainly focuses on spatial dimension modeling, ignoring the consistency requirements of the temporal dimension. It does not perform differentiated processing on regions with different spatiotemporal characteristics, resulting in random fluctuations in content, texture, color and other elements between adjacent frames, more serious flickering in specific areas and incoherent motion, which seriously affects the appearance and practicality of the generated video and restricts the widespread application of text video technology.
[0004] Therefore, there is an urgent need to develop a video generation method, device, electronic device and storage medium to solve one or more of the above-mentioned problems. Summary of the Invention
[0005] In view of this, in order to solve the above technical problems or part of the technical problems, the embodiments of the present invention provide a video generation method, device, electronic device and storage medium.
[0006] In a first aspect, the present application provides a video generation method, the method comprising:
[0007] Get the video description text entered by the user;
[0008] generating an initial noise sequence according to the video description text;
[0009] Using a preset updating strategy, the initial noise sequence is updated to obtain a target noise sequence;
[0010] The target noise sequence is converted into a corresponding pixel value set, and the video description text corresponding to the target video is generated based on the pixel value set.
[0011] In one possible implementation, the updating of the initial noise sequence using a preset updating strategy to obtain a target noise sequence includes:
[0012] Acquiring initial feature information of the initial noise sequence, and performing an updating process on the initial noise sequence according to the initial feature information to obtain a first noise sequence, wherein the initial feature information is used to characterize a regional attention weight distribution of the initial noise sequence;
[0013] Acquiring first feature information of the first noise sequence, and updating the first noise sequence based on the first feature information to obtain a second noise sequence, wherein the first feature information is used to characterize a motion attention weight distribution of the first noise sequence;
[0014] Obtaining second feature information of the second noise sequence, and updating the second noise sequence based on the second feature information to obtain a third noise sequence, wherein the second feature information is used to characterize a fused attention weight distribution of the second noise sequence;
[0015] Determining whether the third noise sequence meets preset requirements;
[0016] In a case where the third noise sequence meets a preset requirement, converting the third noise sequence into a corresponding pixel value set, and generating the video description text corresponding to the target video based on the pixel value set;
[0017] If the third noise sequence does not meet the preset requirements, the third noise sequence is used as the initial noise sequence, the steps of obtaining initial feature information of the initial noise sequence are re-executed, and the initial noise sequence is updated according to the initial feature information to obtain the first noise sequence.
[0018] In a possible implementation, the obtaining of initial feature information of the initial noise sequence includes:
[0019] Performing motion detection on the initial noise sequence to determine motion regions in the initial noise sequence and regional motion data of each motion region;
[0020] Performing detail detection on the initial noise sequence to determine detail regions in the initial noise sequence and region detail data of each detail region;
[0021] Initial feature information of the initial noise sequence is obtained according to the regional motion data of each motion region and the regional detail data of each detail region in the initial noise sequence.
[0022] In a possible implementation, updating the initial noise sequence according to the initial feature information to obtain a first noise sequence includes:
[0023] generating an attention mask of the initial noise sequence according to the initial feature information;
[0024] Adjusting the attention weights of each motion region and each detail region in the initial noise sequence according to the attention mask;
[0025] Based on the adjusted attention weights, the initial noise sequence is updated to obtain the first noise sequence.
[0026] In one possible implementation, obtaining first feature information of the first noise sequence includes:
[0027] Acquire, by a feature extraction unit, a moving direction and a motion intensity of any noise frame in the first noise sequence, wherein the moving direction is a moving direction of the noise frame relative to a previous adjacent noise frame;
[0028] For any noise frame in the first noise sequence, determining a motion vector of the noise frame according to a running direction and motion intensity of the noise frame;
[0029] A running vector of each frame in the first noise sequence is obtained to obtain first feature information of the first noise sequence.
[0030] In a possible implementation, updating the first noise sequence according to the first feature information to obtain the second noise sequence includes:
[0031] Inputting the second feature information and the video description text into a cross-attention model so that the cross-attention model outputs a weight adjustment parameter for each frame in the first noise sequence, thereby obtaining a weight update strategy for the first noise sequence;
[0032] According to the weight updating strategy, the first noise sequence is updated to obtain a second noise sequence.
[0033] In one possible implementation, obtaining second feature information of the second noise sequence includes:
[0034] For any noise frame in the second noise sequence, obtaining first association information of the noise frame, where the first association information includes an attention association feature between the noise frame and an adjacent previous noise frame;
[0035] acquiring second association information of the noise frame, wherein the second association information includes an attention association feature between the noise frame and a corresponding noise frame in the first noise sequence;
[0036] Determining a fusion attention weight of the noise frame according to the first association relationship and the second association relationship of the noise frame;
[0037] Second feature information of the second noise sequence is obtained according to the fused attention weight of each noise frame in the second noise sequence.
[0038] In a possible implementation, the updating of the second noise sequence according to the second characteristic information includes:
[0039] The second noise sequence is updated according to the fused attention weight of each noise frame in the second noise sequence to obtain an updated noise sequence.
[0040] In a second aspect, the present application provides a video generation device, the device comprising:
[0041] The acquisition module is used to obtain the video description text input by the user;
[0042] A sequence generation module, configured to generate an initial noise sequence according to the video description text;
[0043] An updating module, configured to update the initial noise sequence using a preset updating strategy to obtain a target noise sequence;
[0044] The video generation module is used to convert the target noise sequence into a corresponding pixel value set, and generate the target video corresponding to the video description text based on the pixel value set.
[0045] In a third aspect, the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of the video generation method described in any one of the embodiments of the first aspect are implemented.
[0046] In a fourth aspect, the present application further provides a computer storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the video generation method described in any one of the embodiments of the first aspect.
[0047] The above-mentioned technical solution provided by the embodiment of the present application has the following advantages over the prior art: the method provided by the embodiment of the present application first generates a rough initial noise sequence based on the video description text input by the user, and then gradually updates the initial noise sequence according to the characteristics of the initial noise sequence. In the process of updating the noise sequence, the constraints of the time dimension are effectively integrated, and differentiated processing is performed on regions with different spatiotemporal characteristics, and the continuity of the motion is enhanced, so as to obtain an updated noise sequence that meets the preset requirements, and then generate the target video, thereby improving the appearance and practicality of the generated video. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0049] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0050] One or more embodiments are exemplarily illustrated by pictures in the corresponding drawings. These exemplifications do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements. Unless otherwise stated, the figures in the drawings do not constitute proportional limitations.
[0051] Figure 1 A flowchart of a video generation method provided in an embodiment of the present application;
[0052] Figure 2 A flowchart of a noise sequence updating method provided in an embodiment of the present application;
[0053] Figure 3 A flowchart of a method for determining initial feature information provided in an embodiment of the present application;
[0054] Figure 4 A flowchart of a method for updating an initial noise sequence provided in an embodiment of the present application;
[0055] Figure 5 A schematic diagram of the steps of a method for updating an initial noise sequence provided in an embodiment of the present application;
[0056] Figure 6 A flowchart of a method for determining first characteristic information provided in an embodiment of the present application;
[0057] Figure 7 A schematic flow chart of a first noise sequence updating method provided in an embodiment of the present application;
[0058] Figure 8 A schematic diagram of the steps of a first noise sequence updating method provided in an embodiment of the present application;
[0059] Figure 9 A flowchart of a method for determining second characteristic information provided in an embodiment of the present application;
[0060] Figure 10 A schematic diagram of the steps of a second noise sequence updating method provided in an embodiment of the present application;
[0061] Figure 11 A schematic diagram of the steps of a video generation method provided in an embodiment of the present application;
[0062] Figure 12 An architectural diagram of a video generation system provided in an embodiment of the present application;
[0063] Figure 13 A schematic diagram of the structure of a video generation device provided in an embodiment of the present application;
[0064] Figure 14 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0065] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0066] The disclosure below provides many different embodiments or examples for implementing different configurations of the present invention. To simplify the disclosure of the present invention, the components and configurations of specific examples are described below. Of course, these are merely examples and are not intended to limit the present invention. In addition, the present invention may repeat reference numerals and / or letters in different examples. Such repetition is for the purpose of simplicity and clarity and does not in itself indicate the relationship between the various embodiments and / or configurations discussed.
[0067] In order to solve the technical problems in the prior art that the content, texture, color and other elements of adjacent frames are inconsistent in time and show random fluctuations, the flickering phenomenon in specific areas (such as complex textures and areas rich in details) is more serious, and the object movement trajectory is not smooth and shows jump-like changes, the present application provides a video generation method, device, electronic device and storage medium, which generates a rough initial noise sequence by obtaining the video description text input by the user, and then gradually updates the initial noise sequence according to the characteristics of the initial noise sequence. In the process of updating the noise sequence, the constraints of the time dimension are effectively integrated to achieve differentiated processing for areas with different spatiotemporal characteristics, and enhance the continuity of motion, so as to obtain an updated noise sequence that meets the preset requirements, and then generate a target video, thereby improving the appearance and practicality of the generated video.
[0068] Figure 1 A flow chart of a video generation method provided in an embodiment of the present application is shown as follows: Figure 1 As shown, the method specifically includes:
[0069] S101: Obtain video description text input by a user.
[0070] Video description text refers to the textual information entered by the user to describe the target video content. This video description text can include the video's theme, setting, characters, actions, emotions, and other relevant details. For example, a video description text might be "An elephant strolls leisurely across the grassland, the sun shining down on it, surrounded by green grass and sparse trees."
[0071] S102: Generate an initial noise sequence according to the video description text.
[0072] The initial noise sequence refers to a preliminary random data sequence generated based on the video description text and used to represent the video content. The initial noise sequence can only roughly restore the information in the video description text, but with low details and accuracy. It can generally be regarded as the original material or basic framework of the video.
[0073] In this embodiment, based on the descriptive text content provided by the video, the video text information is accurately interpreted and converted, the key elements and features in the video description are determined, and an initial noise sequence that is fuzzy but has the video structure is preliminarily generated as the basis for subsequent further processing and optimization.
[0074] S103: Using a preset updating strategy, update the initial noise sequence to obtain a target noise sequence.
[0075] The preset update strategy refers to the denoising strategy for the initial noise sequence during the video generation process, which aims to improve the matching degree between the noise sequence and the video description text through gradual iterative optimization until the target noise sequence that meets the preset requirements is obtained.
[0076] In this embodiment, differentiated processing is performed on regions with different spatiotemporal characteristics to reduce flickering and enhance the smoothness of motion trajectories. At the same time, the constraints of the time dimension are integrated to ensure the coherence and consistency of the content between adjacent frames.
[0077] S104: Convert the target noise sequence into a corresponding pixel value set, and generate the video description text corresponding to the target video based on the pixel value set.
[0078] The target noise sequence refers to a noise sequence that has been optimized using a preset update strategy and is highly matched with the video description text. This target noise sequence has integrated the constraints of the time dimension and performs differentiated processing on regions with different spatiotemporal characteristics, effectively reducing flickering and enhancing the smoothness of the motion trajectory.
[0079] In this embodiment, the obtained target noise sequence is converted into a pixel value set in frame order, and then compressed into a continuous video stream through a video encoding algorithm (such as H.264, H.265). During the encoding process, the video file size is reduced through intra-frame compression (single-frame pixel redundancy removal) and inter-frame compression (storage of adjacent frame difference information), and finally a target video is generated that contains a visual frame sequence and timestamp information, and parameters such as frame rate and resolution are consistent with the preset dimensions of the pixel value set and conform to the text description.
[0080] The video generation method provided in the embodiment of the present application, after obtaining the video description text input by the user, first generates a rough initial noise sequence as a preliminary representation of the video, and then improves the details and accuracy of the initial noise sequence. Specifically, different spatiotemporal characteristic regions in the initial noise sequence are differentiated according to a pre-set update strategy. During the processing, not only the information within each frame is considered, but also the constraints of the time dimension are integrated to ensure the coherence and consistency of the content between adjacent frames, effectively reduce the flickering phenomenon, and enhance the smoothness of the motion trajectory; after multiple rounds of iterative optimization, a target noise sequence that meets the preset requirements and highly matches the video description text is obtained; finally, the target noise sequence is converted into a corresponding set of pixel values, which is compressed into a continuous video stream through a video encoding algorithm to generate a target video that is completely corresponding to the video description text; the target video finally obtained has higher clarity and smoothness, can more accurately restore the scenes and actions in the video description text, and improves the viewing experience and practicality of the video.
[0081] Figure 2 A flow chart of a noise sequence updating method provided in an embodiment of the present application is shown as follows: Figure 2 As shown, constructing a storyboard script set based on the video description information includes:
[0082] S201, obtaining initial feature information of the initial noise sequence, and updating the initial noise sequence according to the initial feature information to obtain a first noise sequence, wherein the initial feature information is used to characterize the regional attention weight distribution of the initial noise sequence;
[0083] The initial feature information refers to the key areas contained in the initial noise sequence and their corresponding attention weight distribution. These key areas may include motion areas, detail-rich areas, etc., and their attention weight distribution reflects the importance of the key areas in the video generation process.
[0084] In this embodiment, feature extraction and analysis are performed on the initial noise sequence. Through algorithms and technical means, motion features, detail features, and other key information in the sequence are accurately identified. Then, based on the analysis results, the importance of each region is determined, and the initial noise sequence is updated in a targeted manner, thereby obtaining a more refined and accurate first noise sequence. This ensures that the updated first noise sequence can more accurately reflect the content of the video description text, while reducing flicker and enhancing the smoothness of the motion trajectory.
[0085] S202. Obtain first feature information of the first noise sequence, and update the first noise sequence based on the first feature information to obtain a second noise sequence, where the first feature information is used to characterize a motion attention weight distribution of the first noise sequence.
[0086] The first feature information refers to the motion vectors and their associated attention weight distributions in the first noise sequence. These motion vectors reflect the motion relationship and direction between noise frames.
[0087] In this embodiment, the motion characteristics of the first noise sequence are analyzed, and the motion direction and intensity of each frame in the sequence are accurately identified through a motion vector detection algorithm, thereby obtaining the motion vector of each frame; then, the first noise sequence is updated based on the motion vector and the corresponding attention weight distribution to further reduce the unevenness of the motion trajectory, while enhancing the coherence and consistency of the video content, and obtaining a more refined and accurate second noise sequence.
[0088] S203. Obtain second feature information of the second noise sequence, and update the second noise sequence based on the second feature information to obtain a third noise sequence, where the second feature information is used to characterize the fused attention weight distribution of the second noise sequence.
[0089] The second feature information refers to the fused attention weight distribution of each noise frame in the second noise sequence, which comprehensively considers the correlation between the noise frames in the time dimension and the spatial dimension; the time dimension refers to the position of the noise frame in the sequence and its relationship with the temporally adjacent frames (i.e., with the historical frames), while the spatial dimension refers to the correlation between different pixels within the noise frame.
[0090] In this embodiment, by conducting an in-depth analysis of the second noise sequence and utilizing the attention mechanism model, the fused attention weight of each frame is calculated. This weight not only reflects the strength of the association between frames, but also reflects the importance of each frame in the video generation process. Then, based on the fused attention weight, the second noise sequence is refined and updated to further optimize the details and accuracy of the noise sequence, thereby obtaining a higher quality and more stable third noise sequence.
[0091] S204: Determine whether the third noise sequence meets a preset requirement.
[0092] The preset requirements generally include indicators such as video clarity, smoothness, detail restoration, and matching with the target video description text.
[0093] In this embodiment, after the third noise sequence is updated, it is necessary to perform evaluation processing on the third noise sequence to determine whether the third noise sequence satisfies the preset picture or fluency and restoration requirements.
[0094] S205 : When the third noise sequence meets a preset requirement, convert the third noise sequence into a corresponding pixel value set, and generate the video description text corresponding to the target video based on the pixel value set.
[0095] In this embodiment, if the third noise sequence meets the preset requirements of clarity, smoothness, detail restoration, and matching with the video description text, it is regarded as the final target noise sequence. Furthermore, each frame of data in the target noise sequence is converted into a corresponding set of pixel values, thereby mapping the noise data to specific pixel values, thereby constructing a target video that is highly matched with the target video description text and has high-quality visual effects.
[0096] S206. If the third noise sequence does not meet the preset requirements, use the third noise sequence as the initial noise sequence, re-execute the step of obtaining initial feature information of the initial noise sequence, and update the initial noise sequence according to the initial feature information to obtain a first noise sequence.
[0097] In this embodiment, if the third noise sequence fails to meet the preset standards in terms of clarity, smoothness, detail restoration, or matching with the target video description text, it indicates that the current noise sequence still needs to be further optimized. In this case, the third noise sequence that does not meet the preset requirements is backtracked to a new initial noise sequence, and the entire update process is restarted. The iterative process is repeated until a noise sequence that meets all the preset requirements is obtained. Through the cyclic optimization mechanism, it is ensured that the final generated target video reaches the optimal state in terms of both visual effects and practicality.
[0098] The noise sequence updating method provided in the embodiment of the present application refines and optimizes the noise sequence through three step-by-step iterative update processes, thereby ensuring that the target video finally generated reaches the best state in terms of visual effects and practicality. During the iteration, the attention mechanism model is fully utilized to perform differentiated processing on regions with different spatiotemporal characteristics, effectively reducing flickering and enhancing the smoothness of motion trajectories. At the same time, the constraints of the time dimension are integrated to ensure the coherence and consistency of content between adjacent frames. Through this cyclic optimization mechanism, a target video is generated that not only highly matches the video description text, but also performs well in clarity, smoothness, and detail restoration, providing users with a smoother visual experience.
[0099] Figure 3 A flow chart of a method for determining initial feature information provided in an embodiment of the present application is shown as follows: Figure 3 As shown, the obtaining of initial feature information of the initial noise sequence includes:
[0100] S301: Perform motion detection on the initial noise sequence to determine motion regions in the initial noise sequence and regional motion data of each motion region.
[0101] Motion detection refers to the frame-by-frame analysis of the initial noise sequence, and the identification of the motion areas in the sequence and the specific motion data of each motion area through the motion detection algorithm; the motion data includes but is not limited to key information such as motion speed and motion intensity, providing a data basis for subsequent video generation and update processing.
[0102] In this embodiment, each frame of data in the initial noise sequence is analyzed by a motion detection algorithm to accurately identify the motion area, and the motion intensity of each motion area is calculated as the regional motion data of each motion area.
[0103] S302: Perform detail detection on the initial noise sequence to determine detail regions in the initial noise sequence and region detail data of each detail region.
[0104] Detail detection refers to the frame-by-frame detail analysis of the initial noise sequence. Through the detail detection algorithm, the detail-rich areas in the sequence and the specific detail data of each detail area are identified; the detail data includes but is not limited to key information such as texture complexity and color changes. The regional detail data is used for subsequent update processing to enhance the clarity and detail restoration of the video.
[0105] In this embodiment, a detail detection algorithm is used to perform in-depth analysis on each frame of the initial noise sequence, accurately identify detail-rich areas, and extract detail data such as texture complexity and color changes of each detail area, providing key information support for subsequent video generation.
[0106] S303 : Obtain initial feature information of the initial noise sequence according to the regional motion data of each motion region and the regional detail data of each detail region in the initial noise sequence.
[0107] In this embodiment, the motion areas, detail areas and their corresponding motion data and detail data in the initial noise sequence are integrated to form initial feature information; this initial feature information not only includes key information such as the motion speed and motion intensity of the motion area, but also includes detail data such as texture complexity and color changes of the detail area, providing comprehensive data support for subsequent targeted update processing of the initial noise sequence.
[0108] The initial feature information determination method provided in the embodiment of the present application uses a motion detection algorithm to perform frame-by-frame analysis on the initial noise sequence to identify the motion areas in the sequence and the specific motion data of each motion area. Secondly, using a detail detection algorithm, the initial noise sequence is analyzed frame by frame in detail to identify the detail-rich areas in the sequence and the specific detail data of each detail area. Finally, the motion areas, detail areas and their corresponding motion data and detail data are integrated to form initial feature information, which provides comprehensive data support for subsequent targeted update processing of the initial noise sequence, so that the update processing can perform differentiated processing on areas with different spatiotemporal characteristics, further optimizing the details and accuracy of the noise sequence.
[0109] Figure 4 A flow chart of a method for updating an initial noise sequence provided in an embodiment of the present application is shown as follows: Figure 4 As shown, the updating process of the initial noise sequence according to the initial feature information to obtain a first noise sequence includes:
[0110] S401. Generate an attention mask for the initial noise sequence based on the initial feature information.
[0111] An attention mask is a mask used to indicate the importance of different regions in the initial noise sequence. The attention mask assigns higher attention weights to key regions (such as motion regions and detail-rich regions), so that in subsequent update processing, regions with different spatiotemporal characteristics can be treated differently, thereby optimizing the details and accuracy of the noise sequence.
[0112] In this embodiment, based on the motion data and detail data in the initial feature information, the attention mechanism model is used to calculate the attention weight of each frame and generate a corresponding attention mask; this mask reflects the importance of different areas in the video generation process and provides key information support for subsequent targeted update processing of the initial noise sequence.
[0113] S402: Adjust the attention weights of each motion region and each detail region in the initial noise sequence according to the attention mask.
[0114] In this embodiment, targeted adjustments are made to key areas in the initial noise sequence based on the generated attention mask; for motion areas, their attention weights are increased so that the updated noise sequence can more accurately reflect the motion trajectory and dynamic features in the video; for detail-rich areas, their attention weights are also increased to improve the clarity and detail restoration of the video; by adjusting the attention weights, the updated first noise sequence is significantly improved in details and accuracy, providing a higher quality data foundation for subsequent video generation.
[0115] S403: Based on the adjusted attention weights, the initial noise sequence is updated to obtain a first noise sequence.
[0116] In this embodiment, after adjusting the attention weights of each region in the initial noise sequence, advanced video generation algorithms are used to perform differentiated update processing on regions with different spatiotemporal characteristics, thereby optimizing the details and accuracy of the noise sequence. After the update processing, the first noise sequence is obtained, which is significantly improved in terms of detail richness and motion trajectory smoothness.
[0117] The initial feature information updating method provided in the embodiment of the present application first calculates the attention weight of each frame according to the initial feature information using the attention mechanism model, and generates a corresponding attention mask, which accurately reflects the importance of different areas in the video generation process. Then, based on the attention mask, the attention weights of key areas in the initial noise sequence, such as motion areas and detail-rich areas, are adjusted; for motion areas, their attention weights are increased to accurately reflect the motion trajectory and dynamic features in the video; for detail-rich areas, their attention weights are also increased to improve the clarity and detail restoration of the video; finally, based on the adjusted attention weights, the initial noise sequence is updated using an advanced video generation algorithm to obtain a first noise sequence that is significantly improved in terms of detail richness, motion trajectory smoothness, etc.
[0118] Figure 5 A schematic diagram of the steps of an initial noise sequence updating method provided in an embodiment of the present application is shown as follows: Figure 5 As shown, the initial noise sequence update mainly includes the following steps:
[0119] Step 1: Regional feature analysis to detect high-motion areas and detail-rich areas in the video.
[0120] Step 2: Attention mask generation, generate adaptive attention mask based on regional characteristics.
[0121] Step 3: Mask modulates self-attention and uses the mask to adjust the attention calculation weights.
[0122] Step 4: Differentiated processing, applying different attention calculation strategies to different specific areas.
[0123] Figure 6 A flowchart of a method for determining first characteristic information provided in an embodiment of the present application is shown as follows: Figure 6 As shown, the obtaining of first feature information of the first noise sequence includes:
[0124] S601: Obtain, by a feature extraction unit, a moving direction and a motion intensity of any noise frame in the first noise sequence, where the moving direction is a moving direction of the noise frame relative to a previous adjacent noise frame.
[0125] The feature extraction unit refers to a functional module specially designed to extract key features from noise sequences. It is generally used to analyze the motion relationship between noise frames. It can efficiently extract the core features that reflect the dynamic characteristics of the video from the noise sequence, so that the update processing can more accurately simulate and reflect the motion trajectory and dynamic characteristics in the real world.
[0126] In this embodiment, the feature extraction unit can perform a detailed analysis of each frame in the first noise sequence, accurately capturing the motion direction and intensity of the noise frame relative to its adjacent previous frame. This provides important data support for subsequent video generation and update processing, enabling the update processing to more accurately reflect the motion trajectory and dynamic features in the video, thereby generating smoother and more coherent video content.
[0127] S602: For any noise frame in the first noise sequence, determine a motion vector of the noise frame according to a running direction and motion intensity of the noise frame.
[0128] In this embodiment, the motion vector of each noise frame is calculated by a motion vector detection algorithm. The motion vector not only includes motion direction and intensity information, but also reflects the motion relationship and dynamic characteristics between noise frames.
[0129] The determination of motion vectors provides key data support for subsequent video generation and update processing, enabling the update processing to perform differentiated processing on noise frames with different motion characteristics, further optimizing the coherence and consistency of video content.
[0130] S603: Acquire a running vector of each frame in the first noise sequence to obtain first feature information of the first noise sequence.
[0131] The second feature information is the motion direction and intensity of each frame in the noise sequence, as well as the motion relationship and dynamic characteristics between frames, which provides a crucial data basis for subsequent video generation and further noise sequence updates.
[0132] In this embodiment, the above-mentioned analysis of running direction and motion intensity is performed on each noise frame in the first noise sequence, and after the motion vector is determined, the motion vector information of all frames is integrated to obtain the second feature information of the first noise sequence. By accurately capturing and analyzing the motion characteristics in the noise sequence, a smoother, more coherent and detailed target video can be generated, providing users with an excellent visual experience.
[0133] The first feature information determination method provided in the embodiment of the present application performs an in-depth analysis of the first noise sequence through a feature extraction unit, accurately captures the running direction and motion intensity of each noise frame, and provides key data support for subsequent video generation and further noise sequence updates.
[0134] Figure 7 A flowchart of a first noise sequence updating method provided in an embodiment of the present application is shown as follows: Figure 7 As shown, the updating process of the first noise sequence according to the first feature information to obtain a second noise sequence includes:
[0135] S701. Input the second feature information and the video description text into a cross-attention model so that the cross-attention model outputs a weight adjustment parameter for each frame in the first noise sequence, and obtains a weight update strategy for the first noise sequence.
[0136] The cross-attention model refers to a deep learning model that can comprehensively consider the feature information of video description text and noise sequence; the cross-attention model calculates the importance of each noise frame in the process of generating the target video by analyzing the key information in the video description text and the feature information in the noise sequence, and outputs the corresponding weight adjustment parameters.
[0137] In this embodiment, the second feature information and the video description text are simultaneously input into the cross-attention model, so that the cross-attention model comprehensively considers the text information and the noise sequence feature information through a complex calculation process, generates a weight adjustment parameter for each noise frame, and adjusts the importance of each frame in the process of generating the target video, thereby obtaining the weight update strategy of the first noise sequence.
[0138] S702: Update the first noise sequence according to the weight update strategy to obtain a second noise sequence.
[0139] In this embodiment, after the cross-attention model outputs the weight update strategy, the update strategy is executed to adjust the weight of each frame in the first noise sequence; for frames with higher weights, it means that they play a more important role in the process of generating the target video; on the contrary, for frames with lower weights, their importance in video generation can be appropriately reduced, so that the updated second noise sequence is more in line with the requirements of the video description text in terms of content expression and video quality.
[0140] The first noise sequence updating method provided in the embodiment of the present application introduces a cross-attention model, comprehensively considers the characteristic information of the video description text and the noise sequence, calculates the importance of each noise frame in the process of generating the target video, and adjusts the weight of the first noise sequence accordingly, so that the updated second noise sequence is more in line with the requirements of the video description text in terms of content expression and video quality; not only improves the accuracy of video generation, but also enhances the coherence and consistency of the video, providing users with a more refined visual experience.
[0141] Figure 8 A schematic diagram of the steps of a first noise sequence updating method provided in an embodiment of the present application is shown as follows: Figure 8 As shown, the first noise sequence update mainly includes the following steps:
[0142] Step 1: Motion feature extraction, analyze adjacent frames and extract motion features (intensity, direction).
[0143] Step 2: Motion conditional embedding, encoding the motion features into conditional embedding vectors.
[0144] Step 3: Cross-attention fusion, which fuses motion conditions with text conditions through the cross-attention mechanism.
[0145] Step 4: Conditional guidance, using enhanced conditional information to guide the diffusion model to generate stable frames.
[0146] Figure 9 A flow chart of a method for determining second characteristic information provided in an embodiment of the present application is shown as follows: Figure 9 As shown, the obtaining of second feature information of the second noise sequence includes:
[0147] S901. For any noise frame in a second noise sequence, obtain first association information of the noise frame, where the first association information includes attention association features between the noise frame and an adjacent previous noise frame.
[0148] The first correlation information refers to the spatial attention feature of the second noise sequence, which describes the spatial correlation and mutual influence between the noise frame and its adjacent frames.
[0149] In this embodiment, by obtaining the attention association between the noise frame and its previous frame, the dynamic dependency between frames is captured, which helps to more accurately simulate the transition and coherence between frames in the subsequent video generation process.
[0150] S902. Obtain second association information of the noise frame, where the second association information includes attention association features between the noise frame and a corresponding noise frame in the first noise sequence.
[0151] The second correlation information refers to the temporal attention feature of the noise frame, which reflects the dynamic correlation and mutual influence between the noise frame and its corresponding frame on the time axis.
[0152] In this embodiment, by obtaining the attention association between the noise frame and the corresponding frame in the first noise sequence, the temporal dependency between frames is captured, which helps to more accurately simulate and maintain the temporal coherence of the video in the subsequent video generation process.
[0153] S903: Determine a fusion attention weight of the noise frame according to the first association relationship and the second association relationship of the noise frame.
[0154] The fused attention weight is a complex calculation process that comprehensively considers the spatial and temporal attention characteristics of noise frames and generates a fused attention weight for each noise frame. This weight reflects the overall importance of the noise frame in the video generation process, providing key data support for subsequent video generation and further noise sequence updates.
[0155] By determining the fusion attention weight of the noise frame, it is possible to perform differentiated processing on noise frames with different characteristics, further optimize the coherence and consistency of the video content, and provide users with a smoother, more coherent and detail-rich target video, thereby enhancing the viewing and practicality of the video.
[0156] In this embodiment, based on the first correlation relationship and the second correlation relationship of the noise frame, the mutual influence and effect between the two are comprehensively analyzed and evaluated to determine the fusion attention weight of the noise frame. The fusion attention weight not only takes into account the spatial correlation between frames, but also takes into account the temporal coherence, providing comprehensive data support for subsequent video generation and further noise sequence updates.
[0157] S904: Obtain second feature information of the second noise sequence according to the fused attention weight of each noise frame in the second noise sequence.
[0158] The second feature information contains the spatial correlation features of noise frames and the temporal dependencies, providing comprehensive data support for subsequent video generation and further noise sequence optimization. By in-depth analysis and utilization of the correlation features in the noise sequence, a more coherent, stable and detail-rich target video can be generated to meet users' demand for high-quality video content.
[0159] In this embodiment, after performing association analysis and determining the fused attention weight on each frame in the second noise sequence, the fused attention weight information of all frames is integrated to obtain the second feature information of the second noise sequence.
[0160] The second feature information determination method provided in the embodiment of the present application captures the comprehensive importance of each noise frame by deeply analyzing the spatial and temporal attention characteristics of the noise frame, providing key data support for subsequent video generation and further noise sequence optimization; it not only improves the accuracy and coherence of video generation, but also enhances the detail richness and viewing experience of the video, providing users with a richer visual experience.
[0161] In an optional solution of the embodiment of the present invention, the updating of the second noise sequence according to the second feature information includes:
[0162] The second noise sequence is updated according to the fused attention weight of each noise frame in the second noise sequence to obtain an updated noise sequence.
[0163] In this embodiment, refined update processing is performed based on the fused attention weight of each frame in the second noise sequence; frames with higher weights are optimized to improve key details and overall quality of the video; moderate adjustments are made to frames with lower weights to maintain the smoothness and coherence of the video; and then the second noise sequence is differentially updated by utilizing the fused attention weights to obtain an updated noise sequence, which is significantly improved in detail richness, motion trajectory smoothness and overall video quality, providing a more accurate data basis for subsequent video generation.
[0164] The adaptive spatiotemporal attention module introduces a temporal attention calculation mechanism into the Transformer block of DiT. Its principle is as follows:
[0165] The standard self-attention calculation formula is:
[0166]
[0167] This embodiment expands the above standard self-attention formula and introduces attention calculation in the time dimension. The expanded formula is:
[0168]
[0169] Among them, Q t represents the query matrix of the current frame, K t-n:t and V t-n:t Represents a key-value matrix containing history frames.
[0170] Figure 10 A schematic diagram of the steps of a second noise sequence updating method provided in an embodiment of the present application is shown as follows: Figure 10 As shown, the second noise sequence update mainly includes the following steps:
[0171] Step 1: Feature mapping, mapping the input features into query (Q), key (K), and value (V) spaces.
[0172] Step 2: Spatial self-attention, calculating the attention relationship between pixels in the current frame.
[0173] Step 3: Temporal self-attention, calculating the attention relationship between the current frame and historical frames.
[0174] Step 4: Adaptive fusion, dynamically fuse spatial and temporal attention results according to content characteristics.
[0175] Figure 11 A schematic diagram of the steps of a video generation method provided in an embodiment of the present application is shown as follows: Figure 11 As shown in Figure 2, the video generation process mainly includes the following steps:
[0176] Step 1: Input and initialization. First, "input conditional information (text description, reference image, etc.)", then "initialize random noise sequence" to start the diffusion reverse process.
[0177] Step 2: Diffusion reverse iteration, enter the loop, first "get the current denoising step number", and then proceed in sequence:
[0178] Step 3: Region-adaptive attention processing, analyzing the characteristics of the current sequence region, identifying and processing detail-rich regions, and generating an attention mask.
[0179] Step 4: Motion-aware conditional control, extract motion features, calculate motion intensity map, generate motion-conditional embedding and fuse it with text.
[0180] Step 5: Enhanced DIT processing generates feature maps Q, K, and V, calculates spatiotemporal attention, obtains adaptive fusion results, and then applies them to the feedforward network to predict noise, update the noise sequence, and determine whether the final step has been reached. If not, the cycle continues.
[0181] Steps 3 to 5 are loop steps;
[0182] Step 6: Output the results, the loop ends, and "generate the final video sequence", completing the video generation process based on the diffusion model.
[0183] Figure 12 This is an architecture diagram of a video generation system provided in an embodiment of the present application, such as Figure 11 As shown in the figure, the video generation system consists of a conditional information module, a motion perception condition control module, an original DIT Transformer processing module, an adaptive spatiotemporal attention module, a regional adaptive attention module, a noise prediction network, a noise initialization module, a diffusion step controller and a video frame generation module.
[0184] The matching process of each module of the video generation system is as follows:
[0185] Input condition information: The process starts with "condition information", which serves as the input content of the entire system.
[0186] Motion perception condition control: The condition information enters the "motion perception condition control module" in the "enhanced DIT module" for preliminary processing.
[0187] Raw DIT Transformer processing: After being processed by the motion-aware conditional control module, the data flows to the “Raw DIT Transformer block”.
[0188] Attention module branch: The original DIT Transformer block output is divided into three paths, entering the "adaptive spatiotemporal attention module" and "regional adaptive attention module" respectively, and connected to the "noise prediction network" at the same time.
[0189] Noise-related processing: The "noise initialization" provides input to the "diffusion step controller" and combines with the output of the noise prediction network to act on the "diffusion step controller".
[0190] Generate video frames: The diffusion step controller finally outputs "generate video frames", completing the entire adaptive spatiotemporal attention mechanism system process based on the DIT model.
[0191] Figure 13 A schematic diagram of the structure of a video generation device provided in an embodiment of the present application is shown in FIG. Figure 13 As shown, the device specifically includes:
[0192] An acquisition module 1301 is used to acquire a video description text input by a user;
[0193] A sequence generation module 1302 is configured to generate an initial noise sequence according to the video description text;
[0194] An updating module 1303 is configured to update the initial noise sequence using a preset updating strategy to obtain a target noise sequence;
[0195] The video generation module 1304 is configured to convert the target noise sequence into a corresponding pixel value set, and generate a target video corresponding to the video description text based on the pixel value set.
[0196] In one possible implementation, the updating module 1303 is further configured to obtain initial feature information of the initial noise sequence, perform an updating process on the initial noise sequence based on the initial feature information, and obtain a first noise sequence, where the initial feature information is used to characterize the regional attention weight distribution of the initial noise sequence; obtain first feature information of the first noise sequence, perform an updating process on the first noise sequence based on the first feature information, and obtain a second noise sequence, where the first feature information is used to characterize the motion attention weight distribution of the first noise sequence; obtain second feature information of the second noise sequence, perform an updating process on the second noise sequence based on the second feature information, and obtain a third noise sequence, where the second feature information is used to characterize the fusion attention weight distribution of the second noise sequence; determine whether the third noise sequence meets preset requirements; if the third noise sequence meets the preset requirements, convert the third noise sequence into a corresponding pixel value set, and generate the target video corresponding to the video description text based on the pixel value set; if the third noise sequence does not meet the preset requirements, use the third noise sequence as the initial noise sequence and re-execute the steps of obtaining the initial feature information of the initial noise sequence and performing an updating process on the initial noise sequence based on the initial feature information to obtain the first noise sequence.
[0197] In one possible embodiment, the update module 1303 is further used to perform motion detection on the initial noise sequence to determine the motion area in the initial noise sequence and the regional motion data of each motion area; perform detail detection on the initial noise sequence to determine the detail area in the initial noise sequence and the regional detail data of each detail area; and obtain initial feature information of the initial noise sequence based on the regional motion data of each motion area and the regional detail data of each detail area in the initial noise sequence.
[0198] In one possible embodiment, the update module 1303 is further used to generate an attention mask of the initial noise sequence based on the initial feature information; adjust the attention weights of each motion area and each detail area in the initial noise sequence based on the attention mask; and update the initial noise sequence based on the adjusted attention weights to obtain a first noise sequence.
[0199] In one possible embodiment, the updating module 1303 is further configured to obtain, through a feature extraction unit, a running direction and motion intensity of any noise frame in the first noise sequence, where the running direction is the motion direction of the noise frame relative to the previous adjacent noise frame; determine, for any noise frame in the first noise sequence, a motion vector of the noise frame based on the running direction and motion intensity of the noise frame; and obtain the running vector of each frame in the first noise sequence to obtain second feature information of the first noise sequence.
[0200] In one possible embodiment, the update module 1303 is also used to input the second feature information and the video description text into a cross-attention model so that the cross-attention model outputs a weight adjustment parameter for each frame in the first noise sequence, thereby obtaining a weight update strategy for the first noise sequence; and according to the weight update strategy, the first noise sequence is updated to obtain a second noise sequence.
[0201] In one possible embodiment, the update module 1303 is also used to obtain, for any noise frame in the second noise sequence, first association information of the noise frame, the first association information including attention association features between the noise frame and the previous adjacent noise frame; obtain second association information of the noise frame, the second association information including attention association features between the noise frame and the corresponding noise frame in the first noise sequence; determine the fused attention weight of the noise frame based on the first association relationship and the second association relationship of the noise frame; and obtain second feature information of the second noise sequence based on the fused attention weight of each noise frame in the second noise sequence.
[0202] In a possible implementation, the updating module 1303 is further configured to update the second noise sequence according to the fused attention weight of each noise frame in the second noise sequence to obtain an updated noise sequence.
[0203] The video generation device provided in this embodiment can be as follows Figure 13 The video generating device shown in , can execute the following Figure 1-12 All steps of the video generation method in Figure 1-12 For technical effects of the video generation method shown, please refer to Figure 1-12 For the sake of brevity, the relevant description will not be repeated here.
[0204] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.
[0205] Figure 14 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application is shown in FIG. Figure 14 As shown, an embodiment of the present application provides an electronic device, including a processor 1401, a communication interface 1402, a memory 1403, and a communication bus 1404, wherein the processor 1401, the communication interface 1402, and the memory 1403 communicate with each other through the communication bus 1404; the memory 1403 is used to store computer programs; the processor 1401 is used to execute the program stored in the memory 1403, and implement the steps of the video generation method provided by any of the aforementioned method embodiments:
[0206] The method comprises the following steps: obtaining a video description text input by a user; generating an initial noise sequence based on the video description text; updating the initial noise sequence using a preset update strategy to obtain a target noise sequence; converting the target noise sequence into a corresponding pixel value set, and generating a target video corresponding to the video description text based on the pixel value set.
[0207] In one possible implementation, initial feature information of the initial noise sequence is obtained, and the initial noise sequence is updated based on the initial feature information to obtain a first noise sequence, where the initial feature information is used to characterize the regional attention weight distribution of the initial noise sequence; first feature information of the first noise sequence is obtained, and the first feature information is updated based on the first feature information to obtain a second noise sequence, where the first feature information is used to characterize the motion attention weight distribution of the first noise sequence; second feature information of the second noise sequence is obtained, and the second feature information is updated based on the second feature information to obtain a third noise sequence, where the second feature information is used to characterize the fusion attention weight distribution of the second noise sequence; whether the third noise sequence meets preset requirements is determined; if the third noise sequence meets the preset requirements, the third noise sequence is converted into a corresponding pixel value set, and the video description text corresponding to the target video is generated based on the pixel value set; if the third noise sequence does not meet the preset requirements, the third noise sequence is used as the initial noise sequence, and the steps of obtaining the initial feature information of the initial noise sequence and updating the initial noise sequence based on the initial feature information to obtain the first noise sequence are re-executed.
[0208] In one possible implementation, motion detection is performed on the initial noise sequence to determine the motion areas in the initial noise sequence and the regional motion data of each motion area; detail detection is performed on the initial noise sequence to determine the detail areas in the initial noise sequence and the regional detail data of each detail area; and initial feature information of the initial noise sequence is obtained based on the regional motion data of each motion area and the regional detail data of each detail area in the initial noise sequence.
[0209] In one possible implementation, an attention mask of the initial noise sequence is generated based on the initial feature information; the attention weights of each motion area and each detail area in the initial noise sequence are adjusted based on the attention mask; and the initial noise sequence is updated based on the adjusted attention weights to obtain a first noise sequence.
[0210] In one possible implementation, a feature extraction unit is used to obtain a running direction and motion intensity of any noise frame in the first noise sequence, where the running direction is the motion direction of the noise frame relative to the previous adjacent noise frame. For any noise frame in the first noise sequence, a motion vector of the noise frame is determined based on the running direction and motion intensity of the noise frame. The running vector of each frame in the first noise sequence is obtained to obtain second feature information of the first noise sequence.
[0211] In one possible embodiment, the second feature information and the video description text are input into a cross-attention model so that the cross-attention model outputs a weight adjustment parameter for each frame in the first noise sequence, thereby obtaining a weight update strategy for the first noise sequence; and according to the weight update strategy, the first noise sequence is updated to obtain a second noise sequence.
[0212] In one possible embodiment, it is also used to obtain, for any noise frame in the second noise sequence, first association information of the noise frame, the first association information including attention association features between the noise frame and the previous adjacent noise frame; obtain second association information of the noise frame, the second association information including attention association features between the noise frame and the corresponding noise frame in the first noise sequence; determine the fused attention weight of the noise frame based on the first association relationship and the second association relationship of the noise frame; and obtain second feature information of the second noise sequence based on the fused attention weight of each noise frame in the second noise sequence.
[0213] In a possible implementation, it is further configured to update the second noise sequence according to the fused attention weight of each noise frame in the second noise sequence to obtain an updated noise sequence.
[0214] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, or of course, by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the relevant technology, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiment.
[0215] It should be understood that the terms used herein are for the purpose of describing specific example embodiments only and are not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms "one", "an" and "said" as used herein may also be meant to include plural forms. The terms "comprise", "include", "contain" and "have" are inclusive and therefore specify the presence of stated features, steps, operations, elements and / or parts, but do not exclude the presence or addition of one or more other features, steps, operations, elements, parts, and / or combinations thereof. The method steps, processes, and operations described herein are not to be construed as necessarily requiring them to be performed in the specific order described or illustrated, unless the order of execution is clearly indicated. It should also be understood that additional or alternative steps may be used.
[0216] The foregoing description is intended only to provide specific embodiments of the present invention, which will enable those skilled in the art to understand and implement the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not intended to be limited to the embodiments shown herein, but is intended to be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A video generation method, characterized in that: include: Get the video description text entered by the user; generating an initial noise sequence according to the video description text; Using a preset updating strategy, the initial noise sequence is updated to obtain a target noise sequence; The target noise sequence is converted into a corresponding pixel value set, and the video description text corresponding to the target video is generated based on the pixel value set.
2. The method according to claim 1, characterized in that The method of updating the initial noise sequence using a preset updating strategy to obtain a target noise sequence includes: Acquiring initial feature information of the initial noise sequence, and performing an updating process on the initial noise sequence according to the initial feature information to obtain a first noise sequence, wherein the initial feature information is used to characterize a regional attention weight distribution of the initial noise sequence; Acquiring first feature information of the first noise sequence, and updating the first noise sequence based on the first feature information to obtain a second noise sequence, wherein the first feature information is used to characterize a motion attention weight distribution of the first noise sequence; Obtaining second feature information of the second noise sequence, and updating the second noise sequence based on the second feature information to obtain a third noise sequence, wherein the second feature information is used to characterize a fused attention weight distribution of the second noise sequence; Determining whether the third noise sequence meets preset requirements; In a case where the third noise sequence meets a preset requirement, converting the third noise sequence into a corresponding pixel value set, and generating the video description text corresponding to the target video based on the pixel value set; If the third noise sequence does not meet the preset requirements, the third noise sequence is used as the initial noise sequence, the steps of obtaining initial feature information of the initial noise sequence are re-executed, and the initial noise sequence is updated according to the initial feature information to obtain the first noise sequence.
3. The method according to claim 2, characterized in that The obtaining of the initial characteristic information of the initial noise sequence includes: Performing motion detection on the initial noise sequence to determine motion regions in the initial noise sequence and regional motion data of each motion region; Performing detail detection on the initial noise sequence to determine detail regions in the initial noise sequence and region detail data of each detail region; Initial feature information of the initial noise sequence is obtained according to the regional motion data of each motion region and the regional detail data of each detail region in the initial noise sequence.
4. The method according to claim 3, characterized in that The updating process of the initial noise sequence according to the initial feature information to obtain a first noise sequence includes: generating an attention mask of the initial noise sequence according to the initial feature information; Adjusting the attention weights of each motion region and each detail region in the initial noise sequence according to the attention mask; Based on the adjusted attention weights, the initial noise sequence is updated to obtain the first noise sequence.
5. The method according to claim 2, characterized in that The acquiring first feature information of the first noise sequence includes: Acquire, by a feature extraction unit, a moving direction and a motion intensity of any noise frame in the first noise sequence, wherein the moving direction is a moving direction of the noise frame relative to a previous adjacent noise frame; For any noise frame in the first noise sequence, determining a motion vector of the noise frame according to a running direction and motion intensity of the noise frame; A running vector of each frame in the first noise sequence is obtained to obtain first feature information of the first noise sequence.
6. The method according to claim 5, characterized in that The updating process of the first noise sequence according to the first feature information to obtain a second noise sequence includes: Inputting the second feature information and the video description text into a cross-attention model so that the cross-attention model outputs a weight adjustment parameter for each frame in the first noise sequence, thereby obtaining a weight update strategy for the first noise sequence; According to the weight updating strategy, the first noise sequence is updated to obtain a second noise sequence.
7. The method according to claim 2, characterized in that The acquiring second feature information of the second noise sequence includes: For any noise frame in the second noise sequence, obtaining first association information of the noise frame, where the first association information includes an attention association feature between the noise frame and an adjacent previous noise frame; acquiring second association information of the noise frame, wherein the second association information includes an attention association feature between the noise frame and a corresponding noise frame in the first noise sequence; Determining a fusion attention weight of the noise frame according to the first association relationship and the second association relationship of the noise frame; Second feature information of the second noise sequence is obtained according to the fused attention weight of each noise frame in the second noise sequence.
8. The method according to claim 7, characterized in that The updating of the second noise sequence according to the second characteristic information includes: The second noise sequence is updated according to the fused attention weight of each noise frame in the second noise sequence to obtain an updated noise sequence.
9. A video generating device, characterized in that: include: The acquisition module is used to obtain the video description text input by the user; A sequence generation module, configured to generate an initial noise sequence according to the video description text; An updating module, configured to update the initial noise sequence using a preset updating strategy to obtain a target noise sequence; The video generation module is used to convert the target noise sequence into a corresponding pixel value set, and generate the target video corresponding to the video description text based on the pixel value set.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the video generation method according to any one of claims 1 to 8 are implemented.
11. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the video generation method according to any one of claims 1 to 8 are implemented.
Citation Information
Cited By
Video generation method and related equipment
CN121284363A