An audio-driven video generation method, system and electronic device
Patent Information
- Application Number
- CN202610698940.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-20
- Publication Date
- 2026-09-22
AI Technical Summary
然而,这类自回归架构存在固有缺陷:自注意力机制的计算复杂度随输入序列长度呈平方级增长,导致长视频生成时间不可接受;内存占用随序列长度线性增长,所有历史帧的嵌入向量需持续积累,将处理方法集成在边缘设备上实施在性能上和响应速度上都会很快成为瓶颈;现有方法无法支持真正的实时流式推理,必须等待音频数据完整或达到一定长度才能开始生成,且每帧生成时间随历史增长而延长,导致视频会议、虚拟主播等场景中出现明显延迟,无法实现边说边动的自然效果
(1)通过设置第一缓存区存储音频注意力参数、第二和第三缓存区存储融合特征和视频注意力参数,在生成当前生成视频帧的过程中仅计算当前帧的注意力参数,历史帧的注意力参数可以直接从缓存读取复用,从而减少处理时间和减低处理复杂度,大幅减少冗余计算和内存占用,实现低延迟的实时音频驱动视频生成。
Smart Images

Figure CN122802739A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision, and more specifically, to an audio-driven video generation method, system, and electronic device. Background Technology
[0002] Audio-driven video refers to a technology that generates corresponding speaker actions in real time or offline by analyzing input audio data. Currently, audio-driven speaker action video, especially speaker facial action video, is widely used in scenarios such as virtual anchors, video conferencing, and digital human interaction.
[0003] Current mainstream voice-driven speaker action video generation methods generally adopt an autoregressive decoding architecture based on a self-attention mechanism. Specifically, the autoregressive decoding architecture takes historically generated facial action sequences, corresponding historical audio segments, and speaker identity information as inputs and outputs the facial actions of the current frame frame by frame. However, this type of autoregressive architecture has inherent drawbacks: the computational complexity of the self-attention mechanism increases quadratically with the length of the input sequence, making the generation time of long videos unacceptable; memory usage increases linearly with the sequence length, and the embedding vectors of all historical frames need to be continuously accumulated, which quickly becomes a bottleneck in terms of performance and response speed when integrating the processing method on edge devices; existing methods cannot support true real-time streaming inference, and must wait for the audio data to be complete or reach a certain length before generation can begin, and the generation time of each frame increases with the increase of history, resulting in significant delays in scenarios such as video conferencing and virtual anchors, and failing to achieve a natural effect of speaking and moving simultaneously. Summary of the Invention
[0004] This invention provides an audio-driven video generation method, system, and electronic device for realizing real-time streaming generation of audio-driven video.
[0005] According to a first aspect of this application, an audio-driven video generation method is provided, the method comprising: Based on the audio data, obtain the audio attention parameters corresponding to several frames of video to be generated and store them in the first buffer. The fusion features of the previous generated video frame are obtained from the second buffer, and combined with the video style information, the video attention parameters of the current video frame to be generated are obtained and stored in the third buffer. The video attention parameters of several generated video frames preceding the current video frame to be generated are obtained from the third buffer, and the video attention features of the current video frame to be generated are obtained by combining the video attention parameters of the current video frame to be generated. The audio attention parameters of the current video frame to be generated and several frames preceding the current video frame to be generated are obtained from the first buffer. Combined with the video attention features, the fusion features of the current video frame to be generated are obtained and stored in the second buffer. The fusion features of the current video frame to be generated are mapped to a preset video space to generate the video data of the current video frame.
[0006] Understandably, by setting the first buffer to store audio attention parameters and the second and third buffers to store fusion features and video attention parameters, only the attention parameters of the current frame are calculated during the generation of the current video frame. The attention parameters of historical frames can be directly read from the buffer and reused, thereby reducing processing time and complexity, significantly reducing redundant calculations and memory usage, and achieving low-latency real-time audio-driven video generation.
[0007] Optionally, obtaining audio attention parameters corresponding to several frames of the video to be generated based on the audio data includes: Collect audio data; When the collected audio data reaches a preset time length, the audio data of the preset time length is used as an audio processing block; Feature extraction is performed on the audio processing block to obtain audio features corresponding to several frames of video to be generated; Based on the audio features corresponding to the several video frames to be generated, obtain the audio attention parameters corresponding to the several video frames to be generated.
[0008] Understandably, by dividing continuous audio data into audio processing blocks according to a preset time length, audio features are extracted and corresponding audio attention parameters are calculated in real time after each audio processing block is acquired, without waiting for the complete audio sequence. This achieves audio-driven streaming processing, supports block-by-block input and frame-by-frame generation, significantly reduces the latency of the generated video, and meets the low latency requirements of real-time interactive scenarios.
[0009] Optionally, the step of obtaining the fusion features of the previous generated video frame from the second buffer, and combining them with video style information to obtain the video attention parameters of the current video frame to be generated includes: Based on the time sequence information of the current video frame to be generated and the preset absolute position encoding table, the absolute position encoding of the current video frame to be generated is obtained; wherein, the absolute position encoding table is encoded based on the preset global time sequence and absolute position; Based on the time sequence information of the current video frame to be generated and the preset periodic position encoding table, the periodic position encoding of the current video frame to be generated is obtained; wherein, the periodic position encoding table is encoded based on the periodic time sequence and the periodic relative position, and the periodic time sequence is obtained by dividing the global time sequence into periodic times; The fusion features of the previous video frame generated from the second buffer are obtained as historical fusion features; Based on the video style information and the historical fusion features, a first initial video feature is obtained; The absolute position code and the periodic position code are superimposed on the first initial video feature to obtain the second initial video feature; Based on the second initial video features, obtain the video attention parameters of the current video frame to be generated.
[0010] Understandably, by introducing absolute positional encoding and periodic positional encoding, it is possible to perceive the global temporal position of each frame and its phase within periodic actions. After superimposing the dual positional encoding, the dynamic changes of humans in different periodic phases can be learned, making the facial expressions and movements generated in the video more consistent with real physiological inertia, significantly improving the anthropomorphic realism of audio-driven video, and avoiding mechanical movement patterns in repetitive cycles.
[0011] Optionally, the video attention parameters are set with a first query vector, a first key vector, and a first value vector; The step of obtaining video attention parameters of several generated video frames preceding the current video frame from the third buffer, and combining these parameters with the video attention parameters of the current video frame to be generated to obtain the video attention features of the current video frame to be generated, includes: The video attention parameters of several frames generated before the current video frame to be generated are obtained from the third buffer and used as historical video attention parameters. The first key vector in the video attention parameters of the current video frame to be generated is concatenated with the first key vector in the historical video attention parameters to obtain the first concatenation key vector. The first value vector of the video attention parameters of the current video frame to be generated is concatenated with the first value vector of the historical video attention parameters to obtain the first concatenated value vector. Based on the first query vector, the first concatenation key vector, and the first concatenation value vector in the video attention parameters of the current video frame to be generated, the video attention features of the current video frame to be generated are obtained.
[0012] Understandably, by reading the key and value vectors from the historical video attention parameters in the third buffer and concatenating them with the key and value vectors of the current frame, the self-attention calculation only needs to process the concatenated key-value sequence without repeatedly calculating the historical frames and their corresponding key and value vectors. This effectively reduces the computational complexity of attention in video generation, avoids redundant extraction of historical features, ensures the complete transmission of temporal dependencies, and improves generation efficiency.
[0013] Optionally, the video attention parameters include overall attention parameters and several group attention parameters; The step of obtaining the fusion features of the previous generated video frame from the second buffer, and combining them with video style information to obtain the video attention parameters of the current video frame to be generated includes: The fusion features of the previous video frame generated from the second buffer are obtained as historical fusion features; Based on the video style information and the historical fusion features, a first initial video feature is obtained; Based on the first initial video features, obtain the overall attention parameters of the current video frame to be generated; Based on the first initial video features, obtain several group attention parameters for the current video frame to be generated.
[0014] Understandably, by simultaneously acquiring global attention parameters and several group attention parameters, global motion features and local region details can be captured separately. The synergy between global and group attention enhances the ability to differentiate between different spatial regions of the video during the video generation process, resulting in a significant improvement in both global coherence and local detail in the generated facial videos.
[0015] Optionally, the overall attention parameters are set with a second query vector, a second key vector, and a second value vector; the group attention parameters are set with a corresponding third query vector, a third key vector, and a third value vector. The step of obtaining video attention parameters of several generated video frames preceding the current video frame from the third buffer, and combining these parameters with the video attention parameters of the current video frame to be generated to obtain the video attention features of the current video frame to be generated, includes: The video attention parameters of several frames generated before the current video frame to be generated are obtained from the third buffer and used as historical video attention parameters. The second key vector and the second value vector in the overall attention parameters of the current video frame to be generated are concatenated with the corresponding second key vector and the second value vector in the historical video attention parameters to obtain the second concatenation key vector and the second concatenation value vector. The third key vector and third value vector in the group attention parameters of the current video frame to be generated are concatenated with the corresponding third key vector and third value vector in the historical video attention parameters to obtain a number of third concatenation key vectors and a number of third concatenation value vectors. Based on the second query vector, second concatenation key vector, and second concatenation value vector in the overall attention parameters of the current video frame to be generated, and the third query vector, several third concatenation key vectors, and several third concatenation value vectors in the grouped attention parameters of the current video frame to be generated, the attention features of the current video frame to be generated are obtained.
[0016] Understandably, by setting key-value pair caches for overall attention and group attention respectively, and concatenating the overall and group key-value vectors of historical video attention parameters with the overall and group key-value vectors of the current video frame to be generated, efficient multi-granularity temporal feature reuse is achieved. While reducing the computational complexity of self-attention, the temporal consistency of global motion and local details is ensured, significantly improving the efficiency and quality of video generation.
[0017] Optionally, obtaining the attention features of the current video frame to be generated based on the second query vector, second concatenation key vector, second concatenation value vector, and third query vector, several third concatenation key vectors, and several third concatenation value vectors from the overall attention parameters of the current video frame to be generated, includes: Based on the second query vector and the second concatenation key vector of the overall attention parameters of the current video frame to be generated, obtain the overall attention weight; Based on the third concatenation key vector and the third query vector among several group attention parameters of the current video frame to be generated, the group attention weight is obtained; All the group attention weights are weighted and aggregated to obtain the group aggregate weights; Based on the overall attention weight and the group aggregation weight, the second splicing value vector and the third splicing value vector are weighted and fused to obtain the video attention features of the current video frame to be generated.
[0018] Understandably, by calculating the overall attention features and group attention features separately, and then performing a second weighted fusion of all group features with the overall attention features, adaptive coordination between global motion and local details is achieved. This allows for a flexible balance between the expression weights of overall posture and fine regions, resulting in facial videos that achieve better results in both macroscopic coherence and microscopic realism.
[0019] Optionally, the audio attention parameters include a fourth key vector and a fourth value vector; The step of obtaining the audio attention parameters of the current video frame to be generated and several frames preceding the current video frame from the first buffer, and combining the video attention features to obtain the fusion features of the current video frame to be generated, includes: The audio attention parameters of the current video frame to be generated and several frames preceding the current video frame to be generated are obtained from the first buffer and used as the audio attention parameters to be fused. Based on the video attention features, obtain the total query vector of the current video frame to be generated; The total attention weight is obtained based on the total query vector and the fourth key vector in the attention parameters of the audio to be fused; The fusion features of the current video frame to be generated are obtained based on the total attention weight and the fourth value vector in the audio attention parameters to be fused.
[0020] Understandably, by directly obtaining the current and historical audio attention parameters from the first buffer as the audio attention parameters to be fused, and combining them with the total query vector generated by the video attention features to perform cross-modal attention calculation, the repeated extraction and calculation of historical audio features are avoided. Under the premise of ensuring the temporal alignment of audio and video, the computational overhead of cross-modal fusion is significantly reduced, supporting real-time streaming generation, and enabling the generated video frames to maintain high-precision lip-sync and facial expression matching with the input audio.
[0021] According to a second aspect of this application, an audio-driven video generation system is provided, the system comprising: The audio attention parameter acquisition module is used to acquire the audio attention parameters corresponding to several frames of video to be generated based on the audio data and store them in the first buffer. The video attention parameter acquisition module is used to obtain the fusion features of the previous generated video frame from the second buffer, combine them with video style information, obtain the video attention parameters of the current video frame to be generated, and store them in the third buffer. The video attention feature acquisition module is used to obtain video attention parameters of several generated video frames preceding the current video frame to be generated from the third buffer, and combine them with the video attention parameters of the current video frame to be generated to obtain the video attention features of the current video frame to be generated. The fusion feature acquisition module is used to acquire the audio attention parameters of the current video frame to be generated and several frames of video frames generated before the current video frame to be generated from the first buffer, and to acquire the fusion feature of the current video frame to be generated by combining the video attention features and storing it in the second buffer. The video data generation module is used to map the fusion features of the current video frame to be generated to a preset video space to generate the video data of the current video frame.
[0022] According to a third aspect of this application, an electronic device is provided, comprising: Memory, used to store one or more computer programs; A processor, when the one or more computer programs are executed by the processor, implements the audio-driven video generation method described in the first aspect above.
[0023] Based on any of the above aspects, the audio-driven video generation method, system, and electronic device provided in this application can achieve the following technical effects: (1) By setting the first buffer to store audio attention parameters and the second and third buffers to store fusion features and video attention parameters, only the attention parameters of the current frame are calculated during the generation of the current video frame. The attention parameters of historical frames can be directly read from the buffer and reused, thereby reducing processing time and processing complexity, greatly reducing redundant calculations and memory occupation, and realizing low-latency real-time audio-driven video generation.
[0024] (2) By dividing continuous audio data into audio processing blocks according to a preset time length, audio features are extracted and corresponding audio attention parameters are calculated in real time after each audio processing block is collected. There is no need to wait for the complete audio sequence, thus realizing audio-driven streaming processing, supporting block-by-block input and frame-by-frame generation, significantly reducing the latency of the generated video, and meeting the low latency requirements of real-time interactive scenarios.
[0025] (3) By introducing absolute position coding and periodic position coding, the global temporal position of each frame and its phase within periodic actions can be perceived during the video generation process. After the dual position coding is superimposed, the dynamic change pattern of humans in different periodic stages can be learned, so that the expressions and actions generated in the video are more in line with real physiological inertia, significantly improving the anthropomorphic realism of audio-driven video and avoiding mechanical movement patterns in repetitive cycles.
[0026] (4) By reading the key vector and value vector from the historical video attention parameters in the third buffer and concatenating them with the key and value vector of the current frame, the self-attention calculation only needs to process the concatenated key-value sequence without repeating the calculation of the historical frame and the corresponding key and value vector. This effectively reduces the attention calculation complexity in video generation, avoids redundant extraction of historical features, and ensures the complete transmission of temporal dependencies, thereby improving the generation efficiency.
[0027] (5) By simultaneously acquiring the overall attention parameters and several group attention parameters, global motion features and local area details can be captured separately during video generation. The overall and group attention work together to enhance the ability to differentiate the modeling of different spatial regions of the video during the video generation process, thereby significantly improving the global coherence and local fineness of the generated facial video.
[0028] (6) By setting key-value pair caches for overall attention and group attention respectively, and concatenating the overall and group key-value vectors of historical video attention parameters with the overall and group key-value vectors of the current video frame to be generated, efficient multi-granularity temporal feature reuse is achieved. While reducing the computational complexity of self-attention, the temporal consistency of global motion and local details is ensured, which significantly improves the efficiency and quality of video generation.
[0029] (7) By calculating the overall attention features and group attention features separately, and then weighting and fusing all group features before weighting and fusing them with the overall attention features, adaptive coordination between global motion and local details is achieved. During the video generation process, the expression weights of the overall posture and fine regions can be flexibly balanced, so that the generated facial video achieves better results in both macroscopic coherence and microscopic realism.
[0030] (8) By directly obtaining the current and historical audio attention parameters from the first buffer as the audio attention parameters to be fused, and combining the total query vector generated by the video attention features to perform cross-modal attention calculation, the repeated extraction and calculation of historical audio features are avoided. Under the premise of ensuring the temporal alignment of audio and video, the computational overhead of cross-modal fusion is significantly reduced, real-time streaming generation is supported, and the generated video frames maintain high-precision lip-sync and facial expression matching with the input audio. Attached Figure Description
[0031] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0032] Figure 1 This is an illustrative application scenario diagram of an audio-driven video generation method provided in this embodiment.
[0033] Figure 2 This is a flowchart of an audio-driven video generation method provided in this embodiment.
[0034] Figure 3 This is a flowchart for obtaining audio attention parameters provided in this embodiment.
[0035] Figure 4 The process for obtaining video attention parameters provided in this embodiment Figure 1 .
[0036] Figure 5 The process for obtaining video attention parameters provided in this embodiment Figure 2 .
[0037] Figure 6 The process for obtaining video attention features provided in this embodiment Figure 1 .
[0038] Figure 7 The process for obtaining video attention features provided in this embodiment Figure 2 .
[0039] Figure 8 The process for obtaining video attention features provided in this embodiment Figure 3 .
[0040] Figure 9 This is a flowchart for obtaining fusion features provided in this embodiment.
[0041] Figure 10 This is a schematic diagram of the functional modules of an audio-driven video generation system provided in this embodiment.
[0042] Figure 11 This is a schematic diagram of the structure of the electronic device provided in this embodiment. Detailed Implementation
[0043] The accompanying drawings are for illustrative purposes only and should not be construed as limiting the scope of this application. To better illustrate the following embodiments, some components in the drawings may be omitted, enlarged, or reduced, and do not represent the actual dimensions of the product; it is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings.
[0044] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0045] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0046] Existing audio-driven speaker action video generation methods generally adopt autoregressive decoding architecture. The computational complexity of its self-attention increases quadratically with the sequence length, and the memory usage expands linearly. This results in excessively long video generation time and difficulty in deployment on edge devices. More importantly, it cannot support true real-time streaming inference. It must wait for the audio to accumulate to a certain length before it can start generating, and the latency per frame continues to worsen with historical data. This makes it difficult to meet the low-latency requirements for speaking and moving simultaneously in scenarios such as video conferencing and virtual anchors.
[0047] This embodiment provides a technical solution that can solve the above problems. The specific implementation of this application will be described in detail below with reference to the accompanying drawings.
[0048] An exemplary diagram illustrating an application scenario of an audio-driven video generation method provided in this application embodiment. Figure 1 As shown, the application scenario includes at least a server 100 and a terminal 200 that can communicate with the server 100.
[0049] Understandably, the server 100 can be an independent electronic device or a cluster of multiple electronic devices; the terminal 200 can be a smartphone terminal, personal computer, tablet computer, vehicle terminal, etc., but is not limited to these.
[0050] In one possible implementation, server 100 and terminal 200 may each execute an audio-driven video generation method provided in the embodiments of this application. Alternatively, the audio-driven video generation method provided in the embodiments of this application may be executed partly in server 100 and partly in terminal 200.
[0051] This embodiment provides an audio-driven video generation method, which can be implemented based on a video generation model with an attention mechanism. The video generation model with an attention mechanism can be a model built using an autoregressive decoding architecture with an attention mechanism. The attention mechanism can include a self-attention mechanism and a cross-attention mechanism. The self-attention mechanism can focus on the relationships between elements within the input data itself, where the input data can be video data or audio data. The cross-attention mechanism can focus on the relationships between elements in two different input sequences, where the two different input sequences can be video data and audio data.
[0052] In this embodiment, the video generation model with attention mechanism used as audio-driven video can be a model that has already been trained and tested, so that it has learned and stored knowledge of processing video data and audio data, and can perform the processing task of audio-driven video generation well in practical applications.
[0053] like Figure 2 As shown, this embodiment provides an audio-driven video generation method that can be further divided into the following steps: S100. Based on the audio data, obtain the audio attention parameters corresponding to several frames of video to be generated and store them in the first buffer. In this embodiment, the audio data driving video generation can be either a pre-recorded fixed-length segment containing natural language information or a continuously output real-time audio stream. The continuously output real-time audio stream can be used to simulate natural speaking processes and scenarios where video generation can be driven while speaking, and is especially suitable for scenarios with high real-time requirements such as live streaming.
[0054] Understandably, the audio data processed by the attention-based video generation model can be pre-recorded audio data of a certain duration, or it can be audio data continuously input into the video generation model. For the video generation model, inputting pre-recorded audio data of a certain duration yields better training and learning results. Therefore, during training and testing, pre-recorded audio data of a certain duration is generally used as its sample set. In practical applications, either pre-recorded audio data of a certain duration or continuously input audio data can be used as input data to adapt to different application scenarios.
[0055] Specifically, such as Figure 3 As shown, obtaining audio attention parameters corresponding to several frames of the video to be generated based on audio data may include the following steps: S110, Acquire audio data; In this embodiment, to adapt to scenarios with high real-time requirements, such as live streaming, audio data can be continuously collected to simulate the scenario of continuous human speech. This allows for real-time streaming processing of the audio data, thereby providing a data foundation for subsequent streaming-driven video generation and reducing unnecessary delays in video generation.
[0056] S120. When the collected audio data reaches a preset time length, the audio data of the preset time length is used as an audio processing block. In this embodiment, when the acquired audio data reaches a preset time length, it can be processed as an audio processing block while the audio data is still being acquired synchronously. Once the audio processing block completes feature extraction and attention parameter calculation, the next acquired audio data block can be processed. That is, the acquisition and processing steps are complementary and can be performed synchronously, thereby achieving streaming processing of the audio data. For example, in this embodiment, the time length can be preset to 0.5 seconds and acquired at a sampling rate of 16kHz, or adjusted appropriately according to actual conditions.
[0057] S130. Perform feature extraction on the audio processing block to obtain audio features corresponding to several frames of video to be generated. In this embodiment, feature extraction is performed on the audio processing block. Specifically, the audio processing block can be input into a pre-trained audio processing model for feature extraction, converting the waveform corresponding to the original audio data into a high-level semantic feature sequence as the audio features. For example, in this embodiment, the audio features can have a dimension of 512, a frame rate of 25 FPS, and each frame of audio features corresponds to 40 milliseconds of audio data, with one frame of audio features corresponding to one video frame.
[0058] S140. Based on the audio features corresponding to the several video frames to be generated, obtain the audio attention parameters corresponding to the several video frames to be generated.
[0059] In this embodiment, audio features can first be mapped to a unified feature space so that they can be fused with video attention features in the same dimension later. For example, after the audio features are mapped to the unified feature space, the mapped audio features are obtained, and their dimension is 128.
[0060] In this embodiment, attention parameters are calculated on the mapped audio features based on an attention mechanism to obtain audio attention parameters corresponding to several frames of video to be generated. In this embodiment, the audio attention parameters are cached using a preset first buffer, allowing the calculation of audio attention parameters and subsequent reading and use of these parameters to occur simultaneously. The subsequent video generation driven by the audio attention parameters is not significantly affected by the audio data acquisition process; that is, it does not require waiting for all audio data used for driving to be acquired before video generation can begin, thus ensuring the smooth execution of video streaming generation.
[0061] S200: Obtain the fusion features of the previous generated video frame from the second buffer, combine them with video style information, obtain the video attention parameters of the current video frame to be generated, and store them in the third buffer. Understandably, when generating the previous video frame, the fusion features of the previously generated video frame are obtained. In this embodiment, a second buffer is preset, and the fusion features of the previously generated video frame are stored to provide crucial contextual information for generating the current video frame. Therefore, when generating the current video frame, the fusion features of the previously generated video frame can be used as one of the processing data of the video generation model, thus serving as the data basis for audio-driven video generation.
[0062] In this embodiment, preset video style information needs to be used as one of the processing data so that the video generation model can learn the personalized facial movement style of the speaker in the video when generating each video frame, thus avoiding the phenomenon of style shift in the video generation model after generating video frames for a long time.
[0063] Understandably, when the current video frame to be generated is the first video frame, and the second buffer has not yet cached any video frame fusion features, the video style information can be used as the processing data of the video generation model. This ensures that the video generation model can obtain the speaker's identity information in the video, laying the foundation for learning style information for subsequent video frame generation.
[0064] Specifically, such as Figure 4 As shown, obtaining the fusion features of the previous generated video frame from the second buffer, and combining them with video style information to obtain the video attention parameters of the current video frame to be generated, may include the following steps: S211. Obtain the absolute position code of the current video frame to be generated based on the time sequence information of the current video frame to be generated and the preset absolute position encoding table; wherein, the absolute position encoding table is encoded based on the preset global time sequence and absolute position; In this embodiment, the absolute position encoding table is based on a preset global time series and global position, that is, it records the overall evolution trend of the action phase under the preset global time series. Its function is to endow the video generation model with the ability to perceive the habitual patterns of human long-term movement, that is, the absolute position will gradually change with the length of time elapsed. Taking the mouth as an example again, although... and Their relative positions are consistent within their respective periods, but the absolute position encoding modulates this: in the initial stage, such as... to The mouth movements are small, indicating an initial adaptation state, meaning the absolute position is set to a relatively small amplitude; entering the habituation stage... to The amplitude naturally increases, meaning the absolute position is set to the normal amplitude; while in the later fatigue stage, such as to The amplitude will be appropriately reduced, meaning the absolute position is set to a smaller amplitude.
[0065] In this embodiment, the time position of the current video frame to be generated in the entire generated video can be obtained based on the time sequence information of the current video frame to be generated. Based on the time position and the absolute position encoding table, the absolute position encoding of the current video frame to be generated can be obtained. The absolute position encoding can carry the phase amplitude that the current video frame to be generated should hold in the entire generated video, so that the video attention parameters generated subsequently have the amplitude perception of the current video frame to be generated based on the entire video.
[0066] For example, taking the mouth as an example, if the time position of the current video frame to be generated is... By referring to the absolute encoding table, it is found that the current video frame to be generated is in the middle of human speech, so the relative motion amplitude is relatively large. Therefore, the obtained absolute encoding can include information about the large motion amplitude.
[0067] S212. Obtain the periodic position code of the current video frame to be generated based on the time sequence information of the current video frame to be generated and the preset periodic position encoding table; wherein, the periodic position encoding table is encoded based on the periodic time sequence and the periodic relative position, and the periodic time sequence is obtained by dividing the global time sequence into periodic times; In this embodiment, the periodic position encoding table is based on periodic time series and periodic relative positions; wherein the periodic time series is obtained by dividing the global time series. The global time series is divided into fixed-length periodic time series, and the relative position within each period is encoded. Its core function is to ensure that the action phase of the same relative position remains consistent in different periods, thereby injecting a stable periodic repetition pattern. Taking mouth movement as an example, if a period is set to 5 seconds, the periodic position encoding table will constrain the mouth movement in... and The phases of the mouths are the same (e.g., both are in a closed state), and the trajectory of the mouth opening and closing amplitude changes is also consistent in each cycle, simulating the natural habit of regular short pauses between sentences or after the exertion of movement when humans speak.
[0068] In this embodiment, the absolute position encoding table and the periodic position encoding table enable the generated video attention parameters to simultaneously capture short-term periodic patterns and long-term temporal evolution, thereby making the speaker's actions in the generated video more human-like and realistic. By incorporating the information from the absolute position encoding table and the periodic position encoding table into the calculation of the video attention parameters, the video generation model can remember the phase consistency within local periods and perceive the gradual changes in state on the global timeline, thus making the actions in the generated long video both regular and natural long-term evolution, more closely resembling real human behavior.
[0069] In this embodiment, the temporal position of the current video frame to be generated within the entire generated video can be obtained based on the time series information of the current video frame to be generated. It is understood that the periodic position encoding table includes a periodic time series. Based on the temporal position of the current video frame to be generated within the entire generated video and the periodic time series, the relative position of the current video frame to be generated within the period can be obtained. Based on the relative position and the periodic position encoding table, the periodic position encoding of the current video frame to be generated can be obtained. This periodic position encoding can carry the periodic amplitude that the current video frame to be generated should possess within the entire generated video, thereby enabling the subsequently generated video attention parameters to have a perception of the periodic amplitude of the generated video frame.
[0070] For example, taking the mouth as an example, if the time position of the video frame to be generated is... By referring to the periodic encoding table, it is found that the current video frame to be generated is located in the latter part of the second period, that is, in the latter part of the human speaking period. Therefore, its phase is similar to the phase of mouth closure. Thus, the obtained absolute encoding can include information similar to the phase of mouth closure.
[0071] S213. Obtain the fusion features of the previous video frame generated from the second buffer as historical fusion features; S214. Obtain the first initial video features based on the video style information and the historical fusion features; In this embodiment, feature fusion based on video style information and the historical fusion features can be performed to obtain a first initial video feature, which serves as the data basis for generating the video attention parameters of the current video frame to be generated.
[0072] S215. The absolute position code and the periodic position code are superimposed on the first initial video feature to obtain the second initial video feature; In this embodiment, absolute position encoding and periodic position encoding are superimposed on the first initial video feature to obtain the second initial video feature, so that the second initial video feature contains absolute position information and periodic relative position information. Based on the second initial video feature, the anthropomorphism and realism of the subsequently generated video data are improved.
[0073] S216. Based on the second initial video features, obtain the video attention parameters of the current video frame to be generated.
[0074] In this embodiment, attention parameters can be calculated on the second initial video features to obtain the video attention parameters of the current video frame to be generated.
[0075] In this embodiment, the video attention parameters are cached by a preset third buffer. This allows the historical video attention parameters to be directly obtained from the third buffer when generating the next video frame, thus obtaining the context information of the video. This avoids having to recalculate the corresponding video attention parameters based on historical data when generating the next video frame, thereby reducing a large amount of redundant calculations and significantly reducing the computational complexity of frame-by-frame generation. It ensures that the inference time for each frame does not increase with the length of the video, truly realizing low-latency, real-time streaming video generation that allows for simultaneous speaking and movement.
[0076] In another embodiment, the video attention parameters include overall attention parameters and several group attention parameters; In this embodiment, the video attention parameters may include overall attention parameters and several group attention parameters. The overall attention parameters focus on the overall key information of the features, while the group attention parameters focus on the key information of each group region. In this application, in addition to focusing on overall key information, attention is also paid to the detailed key information of each group region, enabling the model to learn the transformation rules of key information in the overall video and to refine the learning of transformation rules of key information in group regions, thereby improving the precision of the generated video. For example, the group regions may include key regions such as the eye region, nose region, and mouth region, which can be appropriately adjusted according to the actual situation.
[0077] like Figure 5 As shown, obtaining the fusion features of the previous generated video frame from the second buffer, and combining them with video style information to obtain the video attention parameters of the current video frame to be generated, may include the following steps: S221. Obtain the fusion features of the previous video frame generated from the second buffer as historical fusion features; S222. Obtain the first initial video features based on the video style information and the historical fusion features; In this embodiment, the description of the first initial video feature is similar to that of the first initial video feature in step S213, and can be referred to the description of the first initial video feature in step S213, which will not be repeated here.
[0078] S223. Based on the first initial video features, obtain the overall attention parameters of the current video frame to be generated; In this embodiment, attention parameters can be calculated from the first initial video features to obtain the overall attention parameters of the current video frame to be generated.
[0079] S224. Based on the first initial video features and the preset grouping regions, obtain several grouping attention parameters of the current video frame to be generated.
[0080] In this embodiment, attention parameters can be calculated on the first initial video features based on preset grouping regions to obtain several grouping attention parameters for the current video frame to be generated.
[0081] S300: Obtain video attention parameters of several generated video frames preceding the current video frame to be generated from the third buffer, and combine them with the video attention parameters of the current video frame to be generated to obtain the video attention features of the current video frame to be generated. In this embodiment, the video attention parameters of the generated video frame based on several frames preceding the current video frame and the video attention parameters of the current video frame are used to obtain the video attention features of the current video frame. These features are the video feature carriers into which the audio attention parameters are subsequently integrated, thereby obtaining the corresponding generated video frame. However, the video attention features require the acquisition of the corresponding historical context video information. Therefore, it is necessary to obtain the video attention parameters of the generated video frame based on several frames preceding the current video frame and the video attention parameters of the current video frame to improve the continuity of the entire generated video data.
[0082] In another embodiment, the video attention parameters include overall attention parameters and several group attention parameters; The step of obtaining the fusion features of the previous generated video frame from the second buffer and combining them with video style information to obtain the video attention parameters of the current video frame to be generated may further include: Based on the time sequence information of the current video frame to be generated and the preset absolute position encoding table, the absolute position encoding of the current video frame to be generated is obtained; wherein, the absolute position encoding table is encoded based on the preset global time sequence and absolute position; Based on the time sequence information of the current video frame to be generated and the preset periodic position encoding table, the periodic position encoding of the current video frame to be generated is obtained; wherein, the periodic position encoding table is encoded based on the periodic time sequence and the periodic relative position, and the periodic time sequence is obtained by dividing the global time sequence into periodic times; The fusion features of the previous video frame generated from the second buffer are obtained as historical fusion features; Based on the video style information and the historical fusion features, a first initial video feature is obtained; The absolute position code and the periodic position code are superimposed on the first initial video feature to obtain the second initial video feature; Based on the second initial video features, obtain the overall attention parameters of the current video frame to be generated; Based on the second initial video features and the preset grouping regions, obtain several grouping attention parameters for the current video frame to be generated.
[0083] Specifically, the video attention parameters are set with a first query vector, a first key vector, and a first value vector; like Figure 6 As shown, obtaining the video attention parameters of several frames of generated video frames preceding the current video frame from the third buffer, and combining them with the video attention parameters of the current video frame to be generated, to obtain the video attention features of the current video frame to be generated, may include the following steps: S311. Obtain the video attention parameters of several video frames generated before the current video frame to be generated from the third buffer, and use them as historical video attention parameters; In this embodiment, the third buffer includes a third key vector buffer and a third value vector buffer. When the video attention parameters of the current video frame to be generated are obtained, the key vectors in the video attention parameters of the current video frame to be generated are stored in the third key vector buffer according to time sequence. Similarly, the value vectors in the video attention parameters of the current video frame to be generated are stored in the third value vector buffer according to time sequence, providing video history basis for the subsequent generation of the next video frame.
[0084] In this embodiment, during the generation of the current video frame to be generated, it is necessary to extract the video attention parameters of several video frames generated before the current video frame to be generated from the third buffer, as a historical basis for the current video frame to be generated. It is understood that it is necessary to obtain the first key vector of the video attention parameters of the several video frames generated before the current video frame to be generated from the third key vector buffer, and the first value vector of the video attention parameters of the several video frames generated before the current video frame to be generated from the third value vector buffer.
[0085] S312. The first key vector in the video attention parameters of the current video frame to be generated is concatenated with the first key vector in the historical video attention parameters to obtain the first concatenation key vector. S313. The first value vector of the video attention parameters of the current video frame to be generated is concatenated with the first value vector of the historical video attention parameters to obtain the first concatenated value vector. S314. Obtain the video attention features of the current video frame to be generated based on the first query vector, the first concatenation key vector, and the first concatenation value vector in the video attention parameters of the current video frame to be generated.
[0086] In this embodiment, obtaining the video attention features of the current video frame to be generated requires not only the direct features of the video attention parameters of the current video frame to be generated, but also the historical context features of the video attention parameters of several previous video frames to be generated. Therefore, it is necessary to obtain the video attention features of the current video frame to be generated based on the first query vector, the first concatenation key vector, and the first concatenation value vector in the video attention parameters of the current video frame to be generated.
[0087] Specifically, attention can be calculated on the first query vector, the first concatenation key vector, and the first concatenation value vector in the video attention parameters of the current video frame to be generated, so as to obtain the video attention features of the current video frame to be generated.
[0088] In another implementation, the overall attention parameter is configured with a second query vector, a second key vector, and a second value vector; the group attention parameter is configured with a corresponding third query vector, a third key vector, and a third value vector; the historical video attention parameter also includes historical overall attention parameters and several historical group attention parameters. like Figure 7 As shown, obtaining the video attention parameters of several frames of generated video frames preceding the current video frame from the third buffer, and combining them with the video attention parameters of the current video frame to be generated, to obtain the video attention features of the current video frame to be generated, may include the following steps: S321. Obtain the video attention parameters of several frames of video frames generated before the current video frame to be generated from the third buffer, and use them as historical video attention parameters; In this embodiment, the third buffer includes an overall key-value buffer and several grouped key-value vector buffers. The overall key-value buffer further includes an overall key vector buffer and an overall value vector buffer; each grouped key-value vector buffer includes a corresponding grouped key vector buffer and a grouped value vector buffer. When the overall attention parameters of the current video frame to be generated are obtained, the key vectors in the overall attention parameters of the current video frame to be generated are stored in the overall key vector buffer according to time sequence. Similarly, the value vectors in the overall attention parameters of the current video frame to be generated are stored in the overall value vector buffer according to time sequence. Similarly, when the grouped attention parameters of the current video frame to be generated are obtained, the key vectors in the grouped attention parameters of the current video frame to be generated are stored in the corresponding grouped key vector buffer according to time sequence. Similarly, the value vectors in the grouped attention parameters of the current video frame to be generated are stored in the corresponding grouped value vector buffer according to time sequence, providing video history for the subsequent generation of the next video frame.
[0089] S322. The second key vector and the second value vector in the overall attention parameters of the current video frame to be generated are concatenated with the second key vector and the second value vector in the historical video attention parameters to obtain the second concatenation key vector and the second concatenation value vector. Specifically, in this embodiment, the second key vector in the overall attention parameters of the current video frame to be generated is concatenated with the second key vector in the historical video attention parameters to obtain a second concatenation key vector; the second value vector in the overall attention parameters of the current video frame to be generated is concatenated with the second value vector in the historical video attention parameters to obtain a second concatenation value vector.
[0090] S323. The third key vector and the third value vector in the group attention parameters of the current video frame to be generated are concatenated with the third key vector and the third value vector in the historical video attention parameters to obtain a number of third concatenation key vectors and a number of third concatenation value vectors. Specifically, in this embodiment, the third key vector in the group attention parameters of the current video frame to be generated is concatenated with the third key vector in the historical video attention parameters to obtain a plurality of third concatenation key vectors; the third value vector in the group attention parameters of the current video frame to be generated is concatenated with the third value vector in the historical video attention parameters to obtain a plurality of third concatenation value vectors.
[0091] S324. Based on the second query vector, second concatenation key vector, and second concatenation value vector in the overall attention parameters of the current video frame to be generated, and the third query vector, several third concatenation key vectors, and several third concatenation value vectors in the grouped attention parameters of the current video frame to be generated, obtain the attention features of the current video frame to be generated.
[0092] Specifically, such as Figure 8 As shown, obtaining the attention features of the current video frame to be generated based on the second query vector, second concatenation key vector, second concatenation value vector of the overall attention parameters of the current video frame to be generated, and the third query vector, several third concatenation key vectors, and several third concatenation value vectors of the grouped attention parameters of the current video frame to be generated, may include the following steps: A1. Obtain the overall attention weight based on the second query vector and the second concatenation key vector of the overall attention parameters of the current video frame to be generated; In this embodiment, attention is calculated on the second query vector and the second concatenation key vector of the overall attention parameters of the current video frame to be generated to obtain the overall attention weight.
[0093] A2. Based on the third concatenation key vector and the third query vector among several group attention parameters of the current video frame to be generated, obtain the group attention weight; In this embodiment, attention calculations are performed on the third query vector and the third query vector among several group attention parameters of the current video frame to be generated, respectively, with different third query vectors to obtain multiple group attention weights.
[0094] A3. Perform a weighted aggregation of all the group attention weights to obtain the group aggregation weight; In this embodiment, based on a preset first learnable parameter, all the group attention weights are weighted and aggregated to obtain the group aggregate weight.
[0095] A4. Based on the overall attention weight and the group aggregation weight, the second splicing value vector and the third splicing value vector are weighted and fused to obtain the video attention features of the current video frame to be generated.
[0096] In this embodiment, the second spliced value vector and the third spliced value vector are weighted and fused to obtain the video attention features of the current video frame to be generated; wherein the weight of the second spliced value vector is the overall attention weight, and the third spliced value vector is the group aggregation weight.
[0097] S400: Obtain the audio attention parameters of the current video frame to be generated and several frames preceding the current video frame from the first buffer, and combine the video attention features to obtain the fusion features of the current video frame to be generated and store them in the second buffer. Specifically, the audio attention parameters include a fourth key vector and a fourth value vector; like Figure 9 As shown, the step of obtaining the audio attention parameters of the current video frame to be generated and several frames preceding the current video frame from the first buffer, and combining the video attention features to obtain the fusion features of the current video frame to be generated, may include the following steps: S410. Obtain the audio attention parameters of the current video frame to be generated and several frames preceding the current video frame from the first buffer, and use them as audio attention parameters to be fused. In this embodiment, the first buffer includes a first key vector buffer and a first value vector buffer. It is understood that in step S140, audio attention parameters corresponding to several frames of video to be generated are obtained. Simultaneously, the key vectors from these audio attention parameters are stored in the first key vector buffer according to their chronological order. Similarly, the value vectors from these audio attention parameters are stored in the first value vector buffer according to their chronological order, providing audio history for the subsequent generation of the next video frame.
[0098] S420. Based on the video attention features, obtain the total query vector of the current video frame to be generated; In this embodiment, the video attention features are linearly projected to obtain the total query vector of the current video frame to be generated. This vector is used to dynamically retrieve the most relevant driving information from the audio attention parameters to be fused, based on the visual state of the video frame to be generated as the query condition, thereby achieving cross-modal temporal alignment and fusion of audio and video.
[0099] S430. Obtain the total attention weight based on the total query vector and the fourth key vector in the attention parameters of the audio to be fused; In this embodiment, attention is calculated based on the total query vector and the fourth key vector in the attention parameters of the audio to be fused to obtain the total attention weight.
[0100] S440. Obtain the fusion features of the current video frame to be generated based on the total attention weight and the fourth value vector in the audio attention parameters to be fused.
[0101] In this embodiment, attention calculation is performed on the total attention weight and the fourth value vector in the audio attention parameters to be fused to obtain the fusion features of the current video frame to be generated.
[0102] S500: Map the fusion features of the current video frame to be generated to a preset video space to generate the video data of the current video frame.
[0103] In this embodiment, the fusion features of the current video frame to be generated can first be processed through a preset feedforward neural network to obtain the final fusion features. These final fusion features are then mapped onto a preset video space to obtain the video data of the currently generated video frame. It is understood that the video data of several generated video frames are aggregated to obtain the complete video data driven by the audio data. This complete video data drives the characters in the video to perform transformations based on the audio data. Preferably, the video space can be preset to have 1404 dimensions, corresponding to 468 vertices, and each vertex has 3 coordinates.
[0104] like Figure 10 As shown in the illustration, this application also provides an audio-driven video generation system. Optionally, the system includes: The system includes an audio attention parameter acquisition module 611, a video attention parameter acquisition module 612, a video attention feature acquisition module 613, a fusion feature acquisition module 614, and a video data generation module 615, wherein: The audio attention parameter acquisition module 611 is used to acquire audio attention parameters corresponding to several frames of video to be generated based on audio data and store them in the first buffer area. In this embodiment, the audio attention parameter acquisition module 611 can be used to perform... Figure 2 For a detailed description of the audio attention parameter acquisition module 611 shown in step S100, please refer to the description of step S100.
[0105] The audio attention parameter acquisition module 611 is also used to collect audio data; when the collected audio data reaches a preset time length, the audio data of the preset time length is used as an audio processing block; feature extraction is performed on the audio processing block to obtain audio features corresponding to several frames of video frames to be generated; and audio attention parameters corresponding to the several frames of video frames to be generated are obtained based on the audio features corresponding to the several frames of video frames to be generated.
[0106] In this embodiment, the audio attention parameter acquisition module 611 can also be used to perform... Figure 3For a detailed description of the audio attention parameter acquisition module 611, see steps S110-S140 shown below. For further details on steps S110-S140, please refer to the description of steps S110-S140.
[0107] The video attention parameter acquisition module 612 is used to obtain the fusion features of the previous generated video frame from the second buffer, combine the video style information, obtain the video attention parameters of the current video frame to be generated, and store them in the third buffer. In this embodiment, the video attention parameter acquisition module 612 can be used to perform... Figure 2 For a detailed description of the video attention parameter acquisition module 612, please refer to the description of step S200 shown.
[0108] The video attention parameter acquisition module 612 is further configured to: acquire the absolute position encoding of the current video frame to be generated based on the time sequence information of the current video frame to be generated and a preset absolute position encoding table; wherein the absolute position encoding table is encoded based on a preset global time sequence and absolute position; acquire the periodic position encoding of the current video frame to be generated based on the time sequence information of the current video frame to be generated and a preset periodic position encoding table; wherein the periodic position encoding table is encoded based on a periodic time sequence and periodic relative position, and the periodic time sequence is obtained by dividing the global time sequence into periodic times; acquire the fusion features of the previous generated video frame from the second buffer as historical fusion features; acquire a first initial video feature based on the video style information and the historical fusion features; superimpose the absolute position encoding and the periodic position encoding on the first initial video feature to obtain a second initial video feature; and acquire the video attention parameters of the current video frame to be generated based on the second initial video feature.
[0109] In this embodiment, the video attention parameter acquisition module 612 can also be used to perform... Figure 4 For a detailed description of the video attention parameter acquisition module 612, see steps S211-S216 shown below.
[0110] The video attention parameters of the video attention parameter acquisition module 612 also include overall attention parameters and several group attention parameters; the video attention parameter acquisition module 612 is further configured to acquire the fusion features of the previous generated video frame from the second buffer as historical fusion features; acquire a first initial video feature based on the video style information and the historical fusion features; acquire the overall attention parameters of the current video frame to be generated based on the first initial video feature; and acquire several group attention parameters of the current video frame to be generated based on the first initial video feature and a preset grouping region.
[0111] In this embodiment, the video attention parameter acquisition module 612 can also be used to perform... Figure 5 For a detailed description of the video attention parameter acquisition module 612, see steps S221-S224 shown below. For further details, please refer to the description of steps S221-S224.
[0112] The video attention feature acquisition module 613 is used to acquire video attention parameters of several generated video frames before the current video frame to be generated from the third buffer, and to acquire video attention features of the current video frame to be generated by combining the video attention parameters of the current video frame to be generated. In this embodiment, the video attention feature acquisition module 613 can be used to perform... Figure 2 For a detailed description of the video attention feature acquisition module 613 shown in step S300, please refer to the description of step S300.
[0113] The video attention parameters of the video attention feature acquisition module 613 are further configured with a first query vector, a first key vector, and a first value vector. The video attention feature acquisition module 613 is also used to obtain video attention parameters of several generated video frames preceding the current video frame to be generated from the third buffer, as historical video attention parameters; concatenate the first key vector in the video attention parameters of the current video frame to be generated with the first key vector in the historical video attention parameters to obtain a first concatenation key vector; concatenate the first value vector in the video attention parameters of the current video frame to be generated with the first value vector in the historical video attention parameters to obtain a first concatenation value vector; and acquire the video attention features of the current video frame to be generated based on the first query vector, the first concatenation key vector, and the first concatenation value vector in the video attention parameters of the current video frame to be generated.
[0114] In this embodiment, the video attention feature acquisition module 613 can also be used to perform... Figure 6For a detailed description of the video attention feature acquisition module 613, see steps S311-S314 shown below. For further details on steps S311-S314, please refer to the description of steps S311-S314.
[0115] The overall attention parameters of the video attention feature acquisition module 613 are set with a second query vector, a second key vector, and a second value vector; the group attention parameters of the video attention feature acquisition module 613 are set with corresponding third query vector, third key vector, and third value vector; the video attention feature acquisition module 613 is further used to obtain the video attention parameters of several frames of generated video frames before the current video frame to be generated from the third buffer, as historical video attention parameters; and to compare the second key vector and second value vector in the overall attention parameters of the current video frame to be generated with the second key vector and second value vector in the historical video attention parameters, respectively. The two concatenation vectors are concatenated to obtain a second concatenation key vector and a second concatenation value vector. The third key vector and third value vector of the grouped attention parameters of the current video frame to be generated are concatenated with the third key vector and third value vector of the historical video attention parameters, respectively, to obtain several third concatenation key vectors and several third concatenation value vectors. Based on the second query vector, second concatenation key vector, and second concatenation value vector of the overall attention parameters of the current video frame to be generated, and the third query vector, several third concatenation key vectors, and several third concatenation value vectors of the grouped attention parameters of the current video frame to be generated, the attention features of the current video frame to be generated are obtained.
[0116] In this embodiment, the video attention feature acquisition module 613 can also be used to perform... Figure 7 For a detailed description of the video attention feature acquisition module 613, see steps S321-S324 shown below. For further details, please refer to the description of steps S321-S324.
[0117] The video attention feature acquisition module 613 further includes: acquiring overall attention weights based on a second query vector and a second concatenation key vector of the overall attention parameters of the current video frame to be generated; acquiring group attention weights based on a third concatenation key vector and a third query vector among several group attention parameters of the current video frame to be generated; weighting and aggregating all the group attention weights to obtain group aggregation weights; and weighting and fusing the second concatenation value vector and the third concatenation value vector based on the overall attention weights and the group aggregation weights to obtain the video attention features of the current video frame to be generated.
[0118] In this embodiment, the video attention feature acquisition module 613 can also be used to perform... Figure 8For a more detailed description of the video attention feature acquisition module 613, see steps A1-A4 shown below. Further details regarding steps A1-A4 can be found in the description of steps A1-A4.
[0119] The fusion feature acquisition module 614 is used to acquire the audio attention parameters of the current video frame to be generated and several frames before the current video frame to be generated from the first buffer, and combine the video attention features to acquire the fusion feature of the current video frame to be generated and store it in the second buffer. In this embodiment, the fusion feature acquisition module 614 can be used to perform... Figure 2 For a detailed description of the fusion feature acquisition module 614 shown in step S400, please refer to the description of step S400.
[0120] The audio attention parameters of the fusion feature acquisition module 614 include a fourth key vector and a fourth value vector. The fusion feature acquisition module 614 is further configured to acquire the audio attention parameters of the current video frame to be generated and several video frames generated before the current video frame to be generated from the first buffer as audio attention parameters to be fused; acquire the total query vector of the current video frame to be generated based on the video attention features; acquire the total attention weight based on the total query vector and the fourth key vector in the audio attention parameters to be fused; and acquire the fusion features of the current video frame to be generated based on the total attention weight and the fourth value vector in the audio attention parameters to be fused.
[0121] In this embodiment, the fusion feature acquisition module 614 can also be used to perform... Figure 9 For a detailed description of the fusion feature acquisition module 614, see steps S410-S440 shown below. For further details on steps S410-S440, please refer to the description of steps S410-S440.
[0122] The video data generation module 615 is used to map the fusion features of the current video frame to be generated to a preset video space to generate the video data of the current video frame.
[0123] In this embodiment, the video data generation module 615 can be used to perform... Figure 2 For a detailed description of the video data generation module 615, please refer to the description of step S500 shown.
[0124] This application also provides an electronic device, the structure of which is as follows: Figure 11As shown, the electronic device includes a memory 711, a processor 712, a communication module 713, and an input / output interface 714, etc. Optionally, the memory 711, the processor 712, the communication module 713, and the input / output interface 714 can be connected and communicate with each other through a bus 715.
[0125] The memory 711 is used to store one or more computer programs and to transfer the code of the computer programs to the processor 712; when the one or more computer programs are executed by the processor 712, an audio-driven video generation method according to an embodiment of this application is implemented.
[0126] Optionally, the electronic device can be connected to a network via the communication module 713 to communicate with other devices, such as terminals or servers, and to interact with data. The electronic device can be various forms of digital computers, exemplarily such as desktop computers, servers, workbenches, mainframes, or other types of computers. The electronic device can also be various forms of mobile terminals, exemplarily such as smartphones, tablets, wearable devices (such as helmets, glasses, watches, etc.), and other similar mobile terminals.
[0127] Optionally, the electronic device can connect to desired input / output devices, such as a keyboard or display device, via the input / output interface 714. The electronic device itself may have a display device, and other display devices can also be connected externally via the input / output interface 714. Optionally, a storage device, such as a hard disk, can also be connected via the input / output interface 714 to store data from the electronic device, read data from the storage device, or store data from the storage device in the memory 711. It is understood that the input / output interface 714 can be a wired interface or a wireless interface. Depending on the actual application scenario, the device connected to the input / output interface 714 can be a component of the electronic device or an external device connected to the electronic device when needed.
[0128] Optionally, the memory 711 may be a volatile memory and / or a non-volatile memory. The volatile memory may be a random access memory, etc., and the non-volatile memory may be a read-only memory, a programmable read-only memory, an erasable programmable read-only memory, an electrically erasable programmable read-only memory, or a flash memory, etc.
[0129] Optionally, the computer program stored in the memory 711 can be divided into one or more modules, which are stored in the memory 711 and executed by the processor 712 to perform the method provided in this embodiment. The one or more modules can be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program in the electronic device.
[0130] Optionally, the processor 712 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 712 include, but are not limited to, a central processing unit, a graphics processing unit, a digital signal processor, various special-purpose artificial intelligence computing chips, various processors running machine learning model algorithms, and can also be any suitable controller, microcontroller, processor, etc. The processor 712 executes the various methods and processes of this embodiment, exemplarily, such as an audio-driven video generation method according to an embodiment of this application.
[0131] Optionally, the bus 715 may include a path for transmitting information. Depending on its function, the bus 715 may be classified as an address bus, a data bus, a control bus, etc.
[0132] In an optional implementation, this application embodiment also provides a computer storage medium storing a computer program thereon. When the computer program is executed by a computer, it enables the computer to perform the methods described in the above-described method embodiments. Part or all of the computer program can be loaded and / or installed on the memory 711 of an electronic device. When the computer program is executed by the processor 712, one or more steps of an audio-driven video generation method according to an embodiment of this application can be performed.
[0133] Optionally, the computer-readable storage medium may be a random access memory, a read-only memory, a programmable read-only memory, an erasable programmable read-only memory, an electrically erasable programmable read-only memory, etc.
[0134] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the technical solution of the present invention, and are not intended to limit the specific implementation of the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the claims of the present invention should be included within the protection scope of the claims of the present invention.
Claims
1. An audio-driven video generation method, characterized in that, The method includes: Based on the audio data, obtain the audio attention parameters corresponding to several frames of video to be generated and store them in the first buffer. The fusion features of the previous generated video frame are obtained from the second buffer, and combined with the video style information, the video attention parameters of the current video frame to be generated are obtained and stored in the third buffer. The video attention parameters of several frames generated before the current video frame to be generated are obtained from the third buffer, and the video attention features of the current video frame to be generated are obtained by combining the video attention parameters of the current video frame to be generated. The audio attention parameters of the current video frame to be generated and several frames preceding the current video frame to be generated are obtained from the first buffer. Combined with the video attention features, the fusion features of the current video frame to be generated are obtained and stored in the second buffer. The fusion features of the current video frame to be generated are mapped to a preset video space to generate the video data of the current video frame.
2. The method according to claim 1, characterized in that, The step of obtaining audio attention parameters corresponding to several frames of the video to be generated based on audio data includes: Collect audio data; When the collected audio data reaches a preset time length, the audio data of the preset time length is used as an audio processing block; Feature extraction is performed on the audio processing block to obtain audio features corresponding to several frames of video to be generated; Based on the audio features corresponding to the several video frames to be generated, obtain the audio attention parameters corresponding to the several video frames to be generated.
3. The method according to claim 1, characterized in that, The step of obtaining the fusion features of the previous generated video frame from the second buffer, and combining them with video style information to obtain the video attention parameters of the current video frame to be generated includes: Based on the time sequence information of the current video frame to be generated and the preset absolute position encoding table, the absolute position encoding of the current video frame to be generated is obtained; wherein, the absolute position encoding table is encoded based on the preset global time sequence and absolute position; Based on the time sequence information of the current video frame to be generated and the preset periodic position encoding table, the periodic position encoding of the current video frame to be generated is obtained; wherein, the periodic position encoding table is encoded based on the periodic time sequence and the periodic relative position, and the periodic time sequence is obtained by dividing the global time sequence into periodic times; The fusion features of the previous video frame generated from the second buffer are obtained as historical fusion features; Based on the video style information and the historical fusion features, a first initial video feature is obtained; The absolute position code and the periodic position code are superimposed on the first initial video feature to obtain the second initial video feature; Based on the second initial video features, obtain the video attention parameters of the current video frame to be generated.
4. The method according to claim 1, characterized in that, The video attention parameters are set with a first query vector, a first key vector, and a first value vector; The step of obtaining video attention parameters of several generated video frames preceding the current video frame from the third buffer, and combining these parameters with the video attention parameters of the current video frame to be generated to obtain the video attention features of the current video frame to be generated, includes: The video attention parameters of several frames generated before the current video frame to be generated are obtained from the third buffer and used as historical video attention parameters. The first key vector in the video attention parameters of the current video frame to be generated is concatenated with the first key vector in the historical video attention parameters to obtain the first concatenation key vector. The first value vector of the video attention parameters of the current video frame to be generated is concatenated with the first value vector of the historical video attention parameters to obtain the first concatenated value vector. Based on the first query vector, the first concatenation key vector, and the first concatenation value vector in the video attention parameters of the current video frame to be generated, the video attention features of the current video frame to be generated are obtained.
5. The method according to claim 1, characterized in that, The video attention parameters include overall attention parameters and several grouped attention parameters; The step of obtaining the fusion features of the previous generated video frame from the second buffer, and combining them with video style information to obtain the video attention parameters of the current video frame to be generated includes: The fusion features of the previous video frame generated from the second buffer are obtained as historical fusion features; Based on the video style information and the historical fusion features, a first initial video feature is obtained; Based on the first initial video features, obtain the overall attention parameters of the current video frame to be generated; Based on the first initial video features and the preset grouping regions, obtain several grouping attention parameters for the current video frame to be generated.
6. The method according to claim 5, characterized in that, The overall attention parameters are set with a second query vector, a second key vector, and a second value vector; the group attention parameters are set with a corresponding third query vector, a third key vector, and a third value vector. The step of obtaining video attention parameters of several generated video frames preceding the current video frame from the third buffer, and combining these parameters with the video attention parameters of the current video frame to be generated to obtain the video attention features of the current video frame to be generated, includes: The video attention parameters of several frames generated before the current video frame to be generated are obtained from the third buffer and used as historical video attention parameters. The second key vector and the second value vector in the overall attention parameters of the current video frame to be generated are concatenated with the second key vector and the second value vector in the attention parameters of the historical video to obtain the second concatenation key vector and the second concatenation value vector. The third key vector and third value vector in the group attention parameters of the current video frame to be generated are concatenated with the third key vector and third value vector in the historical video attention parameters to obtain several third concatenation key vectors and several third concatenation value vectors. Based on the second query vector, second concatenation key vector, and second concatenation value vector in the overall attention parameters of the current video frame to be generated, and the third query vector, several third concatenation key vectors, and several third concatenation value vectors in the grouped attention parameters of the current video frame to be generated, the attention features of the current video frame to be generated are obtained.
7. The method according to claim 6, characterized in that, The step of obtaining the attention features of the current video frame to be generated based on the second query vector, second concatenation key vector, second concatenation value vector, and third query vector, several third concatenation key vectors, and several third concatenation value vectors from the overall attention parameters of the current video frame to be generated, includes: Based on the second query vector and the second concatenation key vector of the overall attention parameters of the current video frame to be generated, obtain the overall attention weight; Based on the third concatenation key vector and the third query vector among several group attention parameters of the current video frame to be generated, the group attention weight is obtained; All the group attention weights are weighted and aggregated to obtain the group aggregate weights; Based on the overall attention weight and the group fusion weight, the second splicing value vector and the third splicing value vector are weighted and fused to obtain the video attention features of the current video frame to be generated.
8. The method according to any one of claims 1-7, characterized in that, The audio attention parameters include a fourth key vector and a fourth value vector; The step of obtaining the audio attention parameters of the current video frame to be generated and several frames preceding the current video frame from the first buffer, and combining the video attention features to obtain the fusion features of the current video frame to be generated, includes: The audio attention parameters of the current video frame to be generated and several frames preceding the current video frame to be generated are obtained from the first buffer and used as the audio attention parameters to be fused. Based on the video attention features, obtain the total query vector of the current video frame to be generated; The total attention weight is obtained based on the total query vector and the fourth key vector in the attention parameters of the audio to be fused; The fusion features of the current video frame to be generated are obtained based on the total attention weight and the fourth value vector in the audio attention parameters to be fused.
9. An audio-driven video generation system, characterized in that, The system includes: The audio attention parameter acquisition module is used to acquire the audio attention parameters corresponding to several frames of video to be generated based on the audio data and store them in the first buffer. The video attention parameter acquisition module is used to obtain the fusion features of the previous generated video frame from the second buffer, combine them with video style information, obtain the video attention parameters of the current video frame to be generated, and store them in the third buffer. The video attention feature acquisition module is used to obtain video attention parameters of several generated video frames preceding the current video frame to be generated from the third buffer, and combine them with the video attention parameters of the current video frame to be generated to obtain the video attention features of the current video frame to be generated. The fusion feature acquisition module is used to acquire the audio attention parameters of the current video frame to be generated and several frames of video frames generated before the current video frame to be generated from the first buffer, and to acquire the fusion feature of the current video frame to be generated by combining the video attention features and storing it in the second buffer. The video data generation module is used to map the fusion features of the current video frame to be generated to a preset video space to generate the video data of the current video frame.
10. An electronic device, characterized in that, include: Memory, used to store one or more computer programs; A processor, when the one or more computer programs are executed by the processor, implements an audio-driven video generation method as described in any one of claims 1-8.