An audio-driven video generation method and system
By acquiring audio and images, determining the motion control parameters of the target object, and generating video, the problem of unnatural motion in audio-driven video generation is solved, and high-quality video generation with matching motion and audio emotion is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHEJIANG TONGHUASHUN INTELLIGENT TECH CO LTD
- Filing Date
- 2026-03-18
- Publication Date
- 2026-06-16
AI Technical Summary
In existing audio-driven video generation solutions, the mouth shape, facial expressions, or body movements of the subjects are often unnatural, and the range of motion does not match the emotional content of the audio, resulting in poor realism and naturalness.
By acquiring audio and images, the motion control parameters of the target object when performing the audio are determined, and the target video is generated using a video generation model. A user-adjustable motion intensity coefficient is introduced to control the motion amplitude. A trained machine learning model is used to predict the motion control parameters, and the target video is generated by combining audio and image features.
It achieves the matching of the target object's motion amplitude with the audio content, generating high-quality videos with natural motion and consistent with the audio's emotion, thus improving the controllability and realism of audio-driven video generation.
Smart Images

Figure CN122227012A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of video generation technology, and in particular to an audio-driven video generation method and system. Background Technology
[0002] With the development of deep learning technology, audio-driven video generation technology has made significant progress, showing broad application prospects in fields such as digital human interaction, film and television production, and virtual anchors. This type of technology is usually based on generative networks such as diffusion models, which automatically synthesize a target video of the object performing audio content by taking a single image of the input object and driving audio.
[0003] However, in the videos generated by existing audio-driven video generation solutions, the objects' mouth shapes, facial expressions, or body movements often exhibit unnatural and exaggerated effects. That is, the range of motion is too large or too small compared to the real speaking scene, or it does not match the emotional intensity expressed by the audio, resulting in poor realism and naturalness of the generated video.
[0004] Therefore, there is a need to provide an audio-driven video generation method and system that can precisely control the motion amplitude of objects in the generated video based on audio, improve the matching degree between the video generation result and the audio content in terms of motion intensity, and thus generate high-quality videos with natural motion amplitude and consistent with the emotion of the audio. Summary of the Invention
[0005] This specification provides one or more embodiments of an audio-driven video generation method, the method comprising: acquiring audio and an image, the image including a target object; determining target motion control parameters of the target object performing the audio based on the audio; and generating a target video of the target object performing the audio using a video generation model based on the target motion control parameters, the audio, and the image.
[0006] This specification provides one or more embodiments of an audio-driven video generation system, the system including an acquisition module, a parameter determination module, and a video generation module. The acquisition module is configured to acquire audio and an image, the image including a target object. The parameter determination module is configured to determine target motion control parameters based on the audio when the target object performs the audio. The video generation module is configured to generate a target video of the target object performing the audio using a video generation model, based on the target motion control parameters, the audio, and the image.
[0007] This specification provides one or more embodiments of a computer-readable storage medium that stores computer instructions which are executed by a processor to implement an audio-driven video generation method.
[0008] This specification provides one or more embodiments of a computer program product, including a computer program that, when executed by a processor, implements an audio-driven video generation method. Attached Figure Description
[0009] This specification will be further described by way of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting; in these embodiments, the same reference numerals denote the same structures, wherein: Figure 1 This is a schematic diagram illustrating an application scenario of the audio-driven video generation method according to some embodiments of this specification; Figure 2 This is an exemplary flowchart of an audio-driven video generation method according to some embodiments of this specification; Figure 3 This is an exemplary schematic diagram illustrating the determination of motion amplitude characterization parameters corresponding to each sample video frame according to some embodiments of this specification; Figure 4 This is another exemplary schematic diagram illustrating the determination of motion amplitude characterization parameters corresponding to each sample video frame according to other embodiments of this specification; Figure 5 This is yet another exemplary schematic diagram illustrating the determination of motion amplitude characterization parameters corresponding to each sample video frame according to some embodiments of this specification; Figure 6 These are exemplary schematic diagrams illustrating the generation of a target video using a video generation model according to some embodiments of this specification; Figure 7 This is another exemplary schematic diagram illustrating the generation of a target video using a video generation model according to other embodiments of this specification; Figure 8 This is an exemplary schematic diagram illustrating the generation of a target video through multiple blocks included in a video generation module, according to some embodiments of this specification; Figure 9 This is an exemplary schematic diagram illustrating the generation of a target video through multiple audio modulation units included in the audio modulation module and multiple blocks included in the video generation module, according to some embodiments of this specification. Figure 10 This is another exemplary schematic diagram illustrating the generation of a target video through a plurality of audio modulation units included in an audio modulation module and a plurality of blocks included in a video generation module, according to other embodiments of this specification. Figure 11 This is an exemplary block diagram of an audio-driven video generation system according to some embodiments of this specification; and Figure 12 This is a schematic diagram of a computing device according to some embodiments of this specification. Detailed Implementation
[0010] To more clearly illustrate the technical solutions of the embodiments in this specification, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are merely some examples or embodiments of this specification. For those skilled in the art, these drawings can be applied to other similar scenarios without creative effort. Unless obvious from the context or otherwise specified, the same reference numerals in the drawings represent the same structures or operations.
[0011] It should be understood that the terms “system,” “device,” “unit,” and / or “module” used herein are one way to distinguish different components, elements, parts, sections, or assemblies at different levels. However, if other terms can achieve the same purpose, they may be replaced by other expressions.
[0012] As indicated in this specification and claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" do not specifically refer to the singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of expressly identified steps and elements, which do not constitute an exclusive list, and the method or apparatus may also include other steps or elements.
[0013] Flowcharts are used in this specification to illustrate the operations performed by the system according to embodiments of this specification. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, the steps can be processed in reverse order or simultaneously. Furthermore, other operations can be added to these processes, or one or more steps can be removed from them.
[0014] With the rapid development of deep learning and multimodal generation technologies, audio-driven video generation has shown broad application potential in fields such as digital human interaction, virtual anchors, film and television production, and online education. This technology aims to automatically synthesize dynamic videos of a given target object image and driving audio, requiring the generated content to possess a high degree of realism and naturalness in terms of lip-sync, facial expressions, and emotional expression. However, current audio-driven video generation schemes produce videos with poor realism and naturalness. The main reasons for this problem are: firstly, model training relies on large-scale real-world speech video datasets, and existing technologies lack reasonable measurement and quality control of motion amplitude in training samples during data construction, causing the model to learn incorrect generation preferences; secondly, most existing solutions do not specifically design explicit control structures for motion amplitude, but instead rely on the model to implicitly learn motion patterns from audio through cross-attention mechanisms, making it difficult to accurately guide and constrain the motion intensity of the generated object.
[0015] To address the aforementioned issues, a reasonable control mechanism needs to be designed to precisely guide the motion amplitude of the object during the generation process to match the audio content. Therefore, an audio-driven video generation method is proposed.
[0016] Figure 1 These are schematic diagrams illustrating application scenarios of the audio-driven video generation method according to some embodiments of this specification. For example... Figure 1 As shown, the application scenario 100 of the audio-driven video generation method may include a processor 110, a network 120, a user terminal 130, and a storage device 140.
[0017] Processor 110 can process data and / or information obtained from other devices or other components of application scenario 100. Processor 110 can execute program instructions based on this data, information, and / or processing results to perform one or more functions described in the embodiments of this specification. For example, processor 110 can process audio and images to generate video by executing the audio-driven video generation method disclosed in this specification. Exemplarily, processor 110 can acquire audio and images; determine target motion control parameters when a target object performs audio based on the audio; and generate a target video of the target object performing audio using a video generation model based on the target motion control parameters, audio, and images. For more information on this section, please refer to [link to relevant documentation]. Figure 2 Related descriptions.
[0018] In some embodiments, the processor 110 can communicate with the user terminal 130 and the storage device 140 via the network 120 to provide various functions of online services.
[0019] In some embodiments, processor 110 may include one or more sub-processing devices (e.g., a single-core processing device or a multi-core multi-chip processing device). By way of example only, processor 110 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), an application-specific instruction processor (ASIP), a graphics processing unit (GPU), a digital signal processor (DSP), a microprocessor, or any combination thereof. In some embodiments, processor 110 may be implemented on a cloud platform. By way of example only, the cloud platform may include a private cloud, a public cloud, a hybrid cloud, a community cloud, a distributed cloud, an internal cloud, a multi-tiered cloud, or any combination thereof.
[0020] Network 120 may include any suitable network capable of facilitating information and / or data exchange within application scenario 100. Network 120 enables communication between components and with other external components, facilitating the exchange of data and / or information. For example, sample video stored in storage device 140 can be transmitted via network 120 to processor 110 for processing. As another example, processor 110 can transmit target video to storage device 140 via network 120.
[0021] In some embodiments, network 120 can be any one or more of wired or wireless networks. For example, network 120 may include fiber optic networks, the Internet, local area networks (LANs), wide area networks (WANs), wireless local area networks (WLANs), metropolitan area networks (MANs), Bluetooth networks, ZigBee networks, cable connections, and any combination thereof. In some embodiments, the network can be various topologies such as point-to-point, shared, and centralized, or a combination of multiple topologies.
[0022] User terminal 130 refers to one or more terminal devices or software used by a user. For example, user terminal 130 may include mobile device 130-1, tablet computer 130-2, laptop computer 130-3, etc., or any combination thereof. Users can interact with other components of application scenario 100 via user terminal 130. For example, user terminal 130 may include an interactive interface through which users can interact with the processor, such as selecting audio and images, and motion intensity coefficients.
[0023] Storage device 140 can be used to store data, instructions, and / or any other information. In some embodiments, storage device 140 can store data and / or information obtained from at least one component of application scenario 100 or an external data source. In some embodiments, storage device 140 can also store data and / or instructions related to an audio-driven video generation method. For example, storage device 140 can store computer instructions for performing video generation. As another example, storage device 140 can store sample videos, target videos, etc.
[0024] In some embodiments, storage device 140 may be connected to network 120 to communicate with processor 110.
[0025] In some embodiments, storage device 140 may include random access memory (RAM), read-only memory (ROM), mass storage, removable memory, volatile read-write memory, or any combination thereof. For example, mass storage may include a hard disk, optical disk, solid-state drive, etc. In some embodiments, storage device 140 may be implemented on a cloud platform.
[0026] For more details regarding the sample videos, target videos, motion intensity coefficients, etc. mentioned above, please refer to [link / reference]. Figures 2-7 And its related descriptions.
[0027] It should be noted that the application scenario 100 of the audio-driven video generation method is provided for illustrative purposes only and is not intended to limit the scope of this application. Those skilled in the art can make various modifications or variations based on the description in this specification. For example, the application scenario 100 of the audio-driven video generation method can implement similar or different functions on other devices. However, these variations and modifications will not depart from the scope of this application.
[0028] Figure 2 This is an exemplary flowchart illustrating an audio-driven video generation method according to some embodiments of this specification. Figure 2 As shown, process 200 may include the following steps. In some embodiments, process 200 may be executed by a processor.
[0029] Step 210: Acquire audio and images, including the target object.
[0030] Audio refers to audio data used to drive the performance of a target object in video. For example, audio can be a user-recorded voice, a musical clip, or a human voice containing specific emotional expressions.
[0031] An image refers to a static image containing the target object, serving as the visual basis for the target object in the target video.
[0032] The target object refers to the main subject that needs to perform the audio in the generated target video. For example, the target object can be a person, a virtual avatar (e.g., a cartoon character), an animal (e.g., a pet), or any object with a movable face or body part.
[0033] In some embodiments, the processor can acquire audio and images input by the user from the user terminal. For example, a user uploads a photo of themselves and an audio recording of a text being read aloud via their mobile phone.
[0034] A target video refers to a dynamic video generated using a video generation model, in which a target object performs audio content. For example, a target video could be a video of a person in an image speaking based on a spoken text. Another example is a video of a virtual avatar in an image singing based on music.
[0035] Step 220: Based on the audio, determine the target motion control parameters when the target object performs the audio.
[0036] Target motion control parameters refer to the motion control parameters ultimately used to generate the target video. Motion control parameters are parameters used to control the range of motion of the target object in the target video (e.g., the range of mouth opening and closing, eye opening and closing, eyebrow raising, and limb movement).
[0037] In some embodiments, the target motion control parameters can be in the form of a sequence, i.e., a target motion control parameter sequence. Multiple target motion control parameters are arranged in time to form a target motion control parameter sequence. Multiple target motion control parameters can correspond to video frames in the final generated target video or audio frames in the audio. The correspondence can be one-to-one or one-to-many (e.g., two video frames correspond to one target motion control parameter).
[0038] In some embodiments, the processor may determine the target motion control parameters based on audio using a motion control parameter determination model, whereby the motion control parameter determination model is a trained machine learning model.
[0039] A motion control parameter determination model is a machine learning model that, after training, can predict the corresponding motion control parameters based on input audio. Examples of motion control parameter determination models include trained Transformer models, Convolutional Neural Network (CNN) models, and Recurrent Neural Network (RNN) models.
[0040] In some embodiments, the input to the motion control parameter determination model can be audio, and the output can be the target motion control parameters.
[0041] In some embodiments, the input to the motion control parameter determination model can be audio, and the output can be initial motion control parameters.
[0042] In some embodiments, the processor may determine initial motion control parameters based on audio and a motion control parameter determination model; and determine target motion control parameters based on the initial motion control parameters and a motion intensity coefficient, wherein the motion intensity coefficient is determined based on user input.
[0043] Initial motion control parameters refer to the unprocessed motion control parameters directly output by the motion control parameter determination model. In some embodiments, the processor can input audio into the motion control parameter determination model to output initial motion control parameters; the initial motion control parameters are then multiplied by a motion intensity coefficient to obtain the target motion control parameters.
[0044] The motion intensity coefficient is an adjustment factor for the amplitude of motion, used to adjust the initial motion control parameters. In some embodiments, the processor can obtain the motion intensity coefficient input by the user from the user terminal. The user can set the motion intensity coefficient according to actual needs. For example, if the user wants more exaggerated facial expressions in the target video, the motion intensity coefficient can be set to a higher value (e.g., 1.5); if the user wants smoother motion in the target video, the motion intensity coefficient can be set to a lower value (e.g., 0.8).
[0045] In the embodiments described in this specification, by introducing a user-adjustable motion intensity coefficient, flexible control of motion amplitude is achieved, meeting the user's personalized needs for the motion intensity of target objects in the generated target video.
[0046] In some embodiments, the motion control parameter determination model can be obtained by training a first initial machine learning model based on multiple first training samples. Each first training sample includes sample audio, and the first training label corresponding to each first training sample is the sample motion control parameter.
[0047] The first initial machine learning model refers to the basic machine learning model to be trained. The trained first initial machine learning model becomes the motion control parameter determination model. For example, the first initial machine learning model can be a neural network that has not yet been trained or has only been partially trained. For instance, the first initial machine learning model can be an untrained Transformer, CNN, etc.
[0048] In some embodiments, the processor can acquire multiple sample videos; and based on each sample video, determine each first training sample and a first training label corresponding to each first training sample. Specifically, the processor can separate sample audio and sample video frame sequences from a sample video; extract key points from each sample video frame in the sample video frame sequence; determine motion amplitude representation parameters corresponding to each sample video frame based on the key points; and determine sample motion control parameters for the sample video based on the motion amplitude representation parameters corresponding to each sample video frame.
[0049] Key points refer to feature points of a target object in an image, which can describe the action of the target object in the image. For example, for a face, key points may include the corners of the eyes, the tip of the nose, the corners of the mouth, and the outlines of the upper and lower lips; for limbs, key points may include joints. The embodiments in this specification mainly use facial key points as examples for illustration, but do not limit the scope of this specification.
[0050] Motion amplitude representation parameters refer to parameters used to characterize the motion amplitude of an object in a sample video frame. For example, motion amplitude representation parameters may include angles, distances, etc. Taking the eye region as an example, the fan-shaped angle formed by the upper and lower eyelids (the angle formed by the line connecting the inner corner of the eye to the midpoint of the upper eyelid and the line connecting the inner corner of the eye to the midpoint of the lower eyelid) is used as the motion amplitude representation parameter. In some embodiments, each sample video frame has a corresponding motion amplitude representation parameter.
[0051] In some embodiments, the processor can acquire multiple real videos from a storage device as multiple sample videos and perform separation processing. For each sample video: a sample audio and a sequence of sample video frames including multiple sample video frames are obtained, and the separation method includes, but is not limited to, using FFmpeg for audio-video separation; key points are extracted from each sample video frame, and the extraction method includes, but is not limited to, key point detection algorithms such as MediaPipe and OpenPose; based on the key points, the motion amplitude representation parameters corresponding to each sample video frame are determined, and this part can be found in [link to relevant documentation]. Figures 3-5 The relevant description; then, based on the motion amplitude representation parameters corresponding to each sample video frame, the sample motion control parameters corresponding to the sample video are determined.
[0052] For example, the processor directly uses the motion amplitude representation parameters corresponding to each sample video frame in the sample video as the sample motion control parameters for each sample video frame, and then arranges them in frame order to form a sequence of sample motion control parameters for the sample video.
[0053] In some embodiments, the processor may also determine the distribution statistics of multiple motion amplitude characterization parameters corresponding to multiple sample video frames; and based on the distribution statistics, normalize the multiple motion amplitude characterization parameters corresponding to multiple sample video frames to determine the sample motion control parameters corresponding to each sample video.
[0054] Distribution statistics can include mean, standard deviation, maximum value, minimum value, etc.
[0055] Normalization methods include, but are not limited to, Z-score normalization and min-max normalization. In some embodiments, the standard distance between a person's eyes can be used as a reference in the normalization process. In some embodiments, the processor arranges the corresponding motion amplitude characterization parameters of multiple normalized sample video frames in the sequential order of the multiple sample video frames to form the sample motion control parameters of the sample video.
[0056] In the embodiments of this specification, by normalizing the motion amplitude representation parameters, the scale effect caused by differences in target objects, shooting distances, etc., between different sample videos can be eliminated, so that the trained motion control parameters have better generalization ability.
[0057] In some embodiments, the processor can input a first training sample (i.e., sample audio) into a first initial machine learning model, construct a loss function based on the output of the first initial machine learning model and a first training label (i.e., sample motion control parameters), and iteratively update the parameters of the first initial machine learning model based on multiple first training samples until preset conditions are met (e.g., loss function convergence, loss function value less than a preset value, number of iterations reaching a preset threshold, etc.). Training is then complete to obtain the trained machine learning model (i.e., the motion control parameter determination model). Methods for updating the parameters of the first initial machine learning model may include, but are not limited to, Batch Gradient Descent (BGD) and Stochastic Gradient Descent (SGD).
[0058] In the embodiments described in this specification, the first training label is constructed by extracting key points from sample videos and calculating motion amplitude representation parameters. This enables the motion control parameter determination model to learn realistic and natural motion patterns, thereby improving the model's prediction accuracy.
[0059] In the embodiments described in this specification, motion control parameters are predicted directly from audio using a trained machine learning model. This model can automatically learn the mapping relationship between audio and motion amplitude, thereby improving the accuracy and naturalness of motion amplitude control.
[0060] Step 230: Based on the target motion control parameters, audio, and images, generate a target video of the target object performing the audio using a video generation model.
[0061] A video generation model is a machine learning model used to generate a target video based on audio and an image containing the target object, in which the target object performs the audio.
[0062] In some embodiments, the processor can acquire audio features of the audio; modulate the audio features using an audio modulation model based on target motion control parameters to obtain modulated audio features; and generate a target video using a video generation model based on the modulated audio features and image features of the image. For more information on this topic, please refer to [link to relevant documentation]. Figure 6 Related descriptions.
[0063] In some embodiments, the video generation model can be a comprehensive model integrating audio encoding, audio modulation, image encoding, and video generation functions, such as a multimodal video generation model based on the Transformer architecture. The input to the video generation model can be target motion control parameters, audio, and images, and the output can be the target video.
[0064] For example, a video generation model may include an audio encoding module, an audio modulation module, an image encoding module, and a video generation module. The processor can input audio into the audio encoding module to obtain audio features; input target motion control parameters and audio features into the audio modulation module to obtain modulated audio features; input an image into the image encoding module to obtain image features; and input the modulated audio features and image features into the video generation module to generate the target video. For more information on this topic, please refer to [link to relevant documentation]. Figure 7 Related descriptions.
[0065] In the embodiments of this specification, by determining the target motion control parameters of the target object when performing audio based on the audio, and generating a target video based on the target motion control parameters, audio, and images, explicit control of the motion amplitude of the target object in the generated video is achieved. This method enables the motion amplitude of the target object in the generated video to match the audio content, avoiding problems of excessive or insufficient motion amplitude, thereby generating a high-quality video with natural motion that matches the emotion of the audio, and improving the controllability and realism of audio-driven video generation.
[0066] Figure 3 This is an exemplary schematic diagram illustrating the determination of motion amplitude characterization parameters corresponding to each sample video frame according to some embodiments of this specification.
[0067] In some embodiments, such as Figure 3 As shown, the processor, based on key points, determines the motion amplitude representation parameters corresponding to each sample video frame, which may include: determining multiple pairs of same-frame key points 320 located within the target region from the key points 310 extracted from the same sample video frame (e.g., ...). Figure 3 The keypoint pairs shown are 320-1, ..., 320-n. Each keypoint pair in a frame includes two keypoints (e.g., ...). Figure 3The keypoint pair 320-1 shown includes two keypoints 310-1-1 and 310-1-2, and the keypoint pair 320-n includes two keypoints 310-n-1 and 310-n-2. The two keypoints are symmetrically distributed with respect to the axis of the target area. The distance 330 between the two keypoints included in each keypoint pair is determined. The average distance 340 of multiple keypoint pairs in the target area is determined as the motion amplitude representation parameter 350 of the target area in the sample video frame.
[0068] A target region refers to a local area on a target object that needs to be represented in terms of motion amplitude. For example, a target region may include the mouth region, the left eye region, the right eye region, etc. Methods for identifying target regions include, but are not limited to, dividing based on the coordinate range of key points (such as determining the mouth region based on the minimum bounding rectangle of the mouth key points) and using pre-trained region segmentation models (such as face parsing models).
[0069] The target area may include two parts symmetrically distributed along an axis, and these two parts may move in a direction perpendicular to the axis.
[0070] A frame-specific keypoint pair refers to a pair of keypoints selected from the same sample video frame, where the two keypoints are located within two parts of the same target area and are symmetrically distributed along the axis of the target area. The axis direction is perpendicular to the main direction of movement of the target area. For example, for the mouth area, the axis direction can be a horizontal direction along the human body from left to right, in which case the line connecting the two keypoints in each frame-specific keypoint pair is perpendicular to the axis direction. When the target object performs audio, the two keypoints in a frame-specific keypoint pair move along a direction perpendicular to the axis. The direction of movement can be consistent or inconsistent. Taking a frame-specific keypoint pair in the mouth area (such as the upper lip keypoint and the lower lip keypoint) as an example, when the mouth is open, the upper lip keypoint moves upward, and the lower lip keypoint moves downward.
[0071] In some embodiments, the processor can determine multiple pairs of keypoints located within a target region from keypoints extracted from the same sample video frame using various methods (such as pairing based on semantic labels of keypoints, pairing based on geometric relationships of keypoint coordinates, etc.). Pairing based on semantic labels of keypoints refers to using a keypoint detection model (such as MediaPipe Face Mesh, OpenFace, etc.) to predefine semantic indexes or labels for each keypoint, pairing upper and lower keypoints with corresponding relationships. Pairing based on geometric relationships of keypoint coordinates can include, but is not limited to, vertical projection pairing methods, convex hull-based or contour-based pairing methods, etc. Keypoint extraction methods can be found in [reference needed]. Figure 2 Related descriptions in Chinese.
[0072] The distance between the two keypoints in each frame keypoint pair can be the straight-line distance between the two keypoints. The methods for determining the distance include, but are not limited to, Euclidean distance calculation, Manhattan distance calculation, etc.
[0073] Taking the mouth region of sample video frame 1 as an example, the keypoints in the mouth region include three keypoints A1, A2, and A3 on the upper lip and three keypoints B1, B2, and B3 on the lower lip. Therefore, the mouth region of sample video frame 1 includes three pairs of keypoints within the same frame (A1 and B1, A2 and B2, and A3 and B3). The processor determines the distance between the two keypoints in each pair (D1 for A1 and B1, D2 for A2 and B2, and D3 for A3 and B3), and calculates the average distance of the three pairs of keypoints within the same frame (i.e., the average of D1, D2, and D3). The parameters representing the motion amplitude of the mouth region in sample video frame 1 are used as the parameters for representing the motion amplitude of the mouth region. It should be noted that sample video frame 1 may also include other target regions such as the left eye region and the right eye region, and their corresponding motion amplitude parameters in sample video frame 1 can be obtained in the same way.
[0074] In the embodiments of this specification, by constructing keypoint pairs symmetrically distributed relative to the axis direction within the same video frame and calculating the average distance between multiple keypoint pairs as a motion amplitude representation parameter, the instantaneous motion amplitude of the target region in the current frame (such as the degree of mouth opening and closing) can be accurately quantified. This method does not rely on information from adjacent frames, is computationally simple and efficient, and can provide accurate label data for training subsequent motion control parameter determination models, enabling the model to learn the correspondence between audio and instantaneous motion amplitude, thereby improving the accuracy of motion amplitude control in the generated video.
[0075] Figure 4 This is another exemplary schematic diagram illustrating the determination of motion amplitude characterization parameters corresponding to each sample video frame according to other embodiments of this specification.
[0076] In some embodiments, such as Figure 4 As shown, the processor, based on key points, determines the motion amplitude representation parameters corresponding to each sample video frame, which may include: determining multiple pairs of neighboring frame key points 420 located within the target region from the key points 410 extracted from two adjacent sample video frames (e.g., ...). Figure 4 The neighboring frame keypoint pairs shown are 420-1, ..., 420-n. Each neighboring frame keypoint pair includes two keypoints located in two adjacent sample video frames (e.g., ...). Figure 4The adjacent frame keypoint pair 420-1 includes two keypoints 410-1-1 and 410-1-2, and the adjacent frame keypoint pair 420-n includes two keypoints 410-n-1 and 410-n-2. The two keypoints correspond to the same part or position of the object in the sample video. The distance 430 between the two keypoints included in each adjacent frame keypoint pair is determined. The average distance 440 of multiple adjacent frame keypoint pairs in the target area is determined as the motion amplitude representation parameter 450 of the target area in the subsequent sample video frame in two adjacent sample video frames.
[0077] Neighboring frame keypoint pairs refer to pairs of keypoints selected from two adjacent sample video frames. These two keypoints are located within the same target area and correspond to the same part (i.e., the same physical location) of the target object in the sample video. For example, for the mouth region, keypoint A1 of the upper lip in sample video frame 1 and keypoint A2 of the upper lip in sample video frame 2... To form a keypoint pair between adjacent frames, A1 and A1 It corresponds to the same physical location on the upper lip of the target object.
[0078] In some embodiments, the processor can determine multiple pairs of neighboring frame keypoints located within the target region from keypoints extracted from two adjacent sample video frames using various methods (such as cross-frame matching based on semantic tags of keypoints, correspondence based on keypoint indexes, etc.). Keypoint extraction methods can be found in [reference needed]. Figure 2 The relevant description is as follows. The same part can refer to the same semantic point or the same index point. Since key point detection models usually assign the same semantic label or index to the same part of the target object in each sample video frame, the processor can establish a cross-frame correspondence based on the semantic label or index to determine the two key points corresponding to the same part, and thus form a key point pair of adjacent frames.
[0079] In two adjacent sample video frames, the later sample video frame refers to the sample video frame that is later in time.
[0080] The distance between the two keypoints in each adjacent frame keypoint pair can be considered as the displacement distance between the two keypoints in the image coordinate system, used to measure the change in position of the same keypoint between two adjacent frames. Methods for determining this distance include, but are not limited to, Euclidean distance calculation and Manhattan distance calculation. The image coordinate system can be a unified coordinate system for the video frames.
[0081] Taking the mouth region of the sample video as an example, assume the sample video includes two adjacent frames, sample video frame 1 and sample video frame 2, where sample video frame 2 follows sample video frame 1. The key points in the mouth region of sample video frame 1 include three key points A1, A2, and A3 on the upper lip and three key points B1, B2, and B3 on the lower lip. The corresponding key points in sample video frame 2 include three key points A1 on the upper lip. A2 A3 Three key points of the lower lip B1 B2 B3 The processor identifies multiple neighboring keypoint pairs between sample video frame 1 and sample video frame 2, including (A1, A1) keypoints in the upper lip region. (A2, A2) (A3, A3) ), and the lower lip area (B1, B1) (B2, B2) (B3, B3) There are a total of 6 neighboring frame keypoint pairs. The processor calculates the distance between the two keypoints in each neighboring frame keypoint pair, for example, A1 and A1... The distance is D1 A2 and A2 The distance is D2 A3 and A3 The distance is D3 B1 and B1 The distance is D4 B2 and B2 The distance is D5 , B3 and B3 The distance is D6 Then calculate the average of these 6 distances, i.e., (D1) +D2 +D3 +D4 +D5 +D6 ) / 6, which serves as the motion amplitude representation parameter for the mouth region in sample video frame 2 (i.e., the next sample video frame).
[0082] It should be noted that for the first sample video frame (e.g., sample video frame 1), since it does not have a preceding frame, the motion amplitude representation parameter cannot be calculated using adjacent frame keypoint pairs. In some embodiments, the processor can set the motion amplitude representation parameter of the first sample video frame to a preset value (e.g., 0).
[0083] In the embodiments of this specification, by constructing keypoint pairs between adjacent frames and calculating the displacement distance of the keypoints, the amount of motion change of the target region in the time dimension can be effectively characterized (such as the degree of opening and closing of the mouth from sample video frame 1 to sample video frame 2). This method can capture the dynamic process of motion, providing labeled data reflecting the continuity and magnitude of motion for training the motion control parameter determination model, enabling the model to learn the correspondence between audio and motion change, thereby improving the naturalness and smoothness of motion processes in the generated video.
[0084] In some embodiments, the processor may also select a sample video frame as a reference frame, and... Figure 4 The method shown calculates motion amplitude representation parameters for frames other than the reference frame. Specifically, the processor can determine multiple keypoint pairs located within the target region from keypoints extracted from other frames and the reference frame. Each keypoint pair includes two keypoints located in one of the other frames and the reference frame, with the two keypoints corresponding to the same part of the object in the sample video. The processor determines the distance between the two keypoints in each keypoint pair and the average distance of multiple keypoint pairs within the target region, which serves as the motion amplitude representation parameter for the target region in other frames.
[0085] Figure 5 This is yet another exemplary schematic diagram illustrating the determination of motion amplitude characterization parameters corresponding to each sample video frame according to some embodiments of this specification.
[0086] In some embodiments, such as Figure 5 As shown, the processor, based on key points, determines the motion amplitude representation parameters corresponding to each sample video frame, which may include: determining multiple pairs of same-frame key points 320 located within the target region from the key points 310 extracted from the same sample video frame (e.g., ...). Figure 5 As shown in 320-1, ..., 320-n), each of the multiple same-frame keypoint pairs includes two keypoints (e.g., ...). Figure 5 The key point pair 320-n shown includes two key points 310-n-1 and 310-n-2, which are symmetrically distributed with respect to the axis of the target area. A reference key point pair 510 is determined based on multiple key point pairs in the same frame. The distance 520 between the two reference key points included in the reference key point pair is determined as the motion amplitude representation parameter 350 of the target area in the sample video frame.
[0087] For an explanation of keypoint pairs within the same frame, please refer to [link / reference]. Figure 3 The relevant descriptions will not be repeated here.
[0088] A reference keypoint pair refers to a pair of keypoints determined based on multiple keypoint pairs within a target area in the same frame, used to represent the overall motion amplitude of that target area. A reference keypoint pair consists of two reference keypoints, each obtained by centering multiple keypoints located on the same side (e.g., the upper lip side or the lower lip side). Centering methods include, but are not limited to, calculating average coordinates. For example, for the mouth area, the average position of multiple keypoints on the upper lip (e.g., A1, A2, A3) can be used as the upper lip reference keypoint A, and the average position of multiple keypoints on the lower lip (e.g., B1, B2, B3) can be used as the lower lip reference keypoint B, forming a reference keypoint pair (A, B).
[0089] The distance between two reference keypoints in a reference keypoint pair refers to the straight-line distance between the two reference keypoints, representing the overall motion amplitude of the target region in the current frame. Methods for determining this distance include, but are not limited to, Euclidean distance calculation and Manhattan distance calculation.
[0090] Taking the mouth region of sample video frame 1 as an example, the key points in the mouth region include three key points A1, A2, and A3 on the upper lip and three key points B1, B2, and B3 on the lower lip. The processor first determines the average coordinate position of the three key points A1, A2, and A3 on the upper lip as reference key point A, and determines the average coordinate position of the three key points B1, B2, and B3 on the lower lip as reference key point B. A and B form a reference key point pair (A, B). Then, the distance D_AB between reference key point A and reference key point B is calculated as the motion amplitude representation parameter of the mouth region in sample video frame 1. Sample video frame 1 may also include other target regions such as the left eye region and the right eye region. The motion amplitude representation parameters corresponding to these regions in sample video frame 1 can be obtained similarly (for example, for the left eye region, two reference key points can be determined based on the upper eyelid key points and the lower eyelid key points, and the distance can be calculated).
[0091] In the embodiments of this specification, reference keypoint pairs are obtained by centering multiple keypoint pairs within the target area in the same frame. Motion amplitude representation parameters are then determined based on the distance between these reference keypoint pairs. This approach more robustly reflects the overall motion amplitude of the target area and reduces the interference of individual keypoint detection errors or local abnormal motion on the motion amplitude representation. This method provides more stable and representative labeled data for training the motion control parameter determination model, helping the model learn the mapping relationship between audio and the overall motion amplitude of the area, thereby improving the stability and naturalness of motion control in the generated video.
[0092] It should be noted that, through the above... Figures 3 to 5After determining the motion amplitude representation parameters corresponding to each sample video frame using the three methods shown, the processor can determine the sample motion control parameters of the sample video based on these parameters. Specifically, a general approach can be used: directly using the motion amplitude representation parameters corresponding to each sample video frame as the sample motion control sub-parameters for each sample video frame, arranging them in frame order to form the sample motion control parameters of the sample video; alternatively, a normalization approach can be used: determining the distribution statistics of multiple motion amplitude representation parameters corresponding to multiple sample video frames, normalizing these parameters based on the distribution statistics, and arranging the normalized parameters in frame order to form the sample motion control parameters of the sample video. This part has been explained above; please refer to [link to relevant documentation]. Figure 2 Related descriptions in Chinese.
[0093] Figure 6 This is an exemplary schematic diagram illustrating the generation of a target video using a video generation model according to some embodiments of this specification.
[0094] In some embodiments, such as Figure 6 As shown, the processor generates a target video of the target object performing the audio based on the target motion control parameters, audio, and image through a video generation model, which may include: acquiring the audio features 610; modulating the audio features 610 through an audio modulation model 630 based on the target motion control parameters 620 to obtain modulated audio features 640; and generating the target video 670 through a video generation model 660 based on the modulated audio features 640 and the image features 650 of the image.
[0095] For explanations of parameters such as target motion control parameters, audio, image, target object, and target video, please refer to [link / reference needed]. Figure 2 Related descriptions.
[0096] Audio features refer to feature vectors that digitally represent audio data, used to characterize the acoustic properties (such as timbre, pitch, rhythm, etc.) and / or semantic content of audio.
[0097] Processors can acquire audio features in various ways. For example, they can input audio into a pre-trained audio encoder (such as Wav2Vec2.0, HuBERT, etc.) to extract hidden layer features as audio features, or directly use signal processing methods (such as short-time Fourier transform) to generate Mel spectrograms as audio features.
[0098] The audio modulation model can be a model independent of the video generation model. It is used to modulate audio features according to the target motion control parameters, so that the modulated audio features carry motion amplitude control information, thereby guiding the video generation model to generate the target video.
[0099] In some embodiments, the audio modulation model can be a machine learning model trained to modulate audio features based on target motion control parameters. For example, the audio modulation model can be a Multilayer Perceptron (MLP) model, an attention-based Transformer model, etc. The input to the audio modulation model can include target motion control parameters and audio features, and the output can be modulated audio features. In some embodiments, the audio modulation model can be trained by: acquiring multiple sample videos; for each sample video, determining the sample motion control parameters based on that sample video (e.g., through...). Figures 2-5 The method described above is used to extract sample audio; calculate the matching degree between sample audio and sample motion control parameters in each sample video, which can be determined by calculating the Pearson correlation coefficient between the audio energy sequence and the motion amplitude characterization parameter sequence; select sample videos with a matching degree greater than a preset threshold as high-quality samples; for each high-quality sample, extract the audio features of its sample audio as sample audio features, and use the sample audio features as the third training label; combine the sample motion control parameters and the sample audio features extracted from the same high-quality sample (as input audio features) to form the third training sample, that is, the third training sample includes sample motion control parameters and sample audio features; input the third training sample into the initial audio modulation model, construct a loss function based on the model output and the third training label, iteratively update the model parameters until the training conditions are met, and obtain the trained audio modulation model.
[0100] Modulated audio features refer to audio features that have been processed by an audio modulation model and carry motion amplitude control information. While retaining the acoustic and semantic information of the original audio features, the modulated audio features incorporate motion amplitude information indicated by the target motion control parameters, thereby guiding the motion amplitude of the target object during subsequent video generation.
[0101] Image features refer to the feature vectors that digitally represent an image containing a target object, used to characterize information such as the visual appearance, posture, and identity of the target object.
[0102] Processors can acquire image features in various ways. For example, a processor can input an image into a pre-trained image encoder (such as ResNet, Vision Transformer, etc.) to extract image features.
[0103] In some embodiments, the video generation model can be a basic generation model that is only responsible for video generation, such as the Stable Video Diffusion model, the Make-Your-Video model, etc. The input to the video generation model can be modulated audio features and image features of the image, and the output can be the target video.
[0104] In this case, the video generation model can be an existing pre-trained model, which does not need to be retrained for this solution; or it can be fine-tuned as needed. Fine-tuning can be based on multiple fourth training samples (including the audio features of the modulated samples and the image features of the samples) and the corresponding fourth training labels (sample videos). The fourth training samples can be obtained based on the third training labels and historical data, and the sample videos can be historical real videos.
[0105] In the embodiments described in this specification, an audio modulation model independent of the video generation model is introduced. Audio features are explicitly modulated based on target motion control parameters, enabling motion amplitude control information to be injected into the video generation process through feature modulation. This method does not require modification of the internal structure of the video generation model, exhibits good versatility and transferability, and can be flexibly combined with various existing video generation models to achieve effective control of the motion amplitude of the target object in the generated video.
[0106] Figure 7 This is another exemplary schematic diagram illustrating the generation of a target video using a video generation model according to other embodiments of this specification.
[0107] In some embodiments, such as Figure 7 As shown, the video generation model 710 includes an audio encoding module 710-1, an audio modulation module 710-2, an image encoding module 710-3, and a video generation module 710-4. The processor, based on target motion control parameters, audio, and an image, generates a target video of the target object performing audio through the video generation model. This can include: inputting audio 720 to the audio encoding module 710-1 to obtain audio features 610; inputting the target motion control parameters 620 and audio features 610 to the audio modulation module 710-2 to obtain modulated audio features 640; inputting an image 730 to the image encoding module 710-3 to obtain image features 650; and inputting the modulated audio features 640 and image features 650 to the video generation module 710-4 to generate the target video 670.
[0108] For explanations of parameters such as target motion control parameters, audio, image, target object, target video, audio features, image features, and modulated audio features, please refer to [link to relevant documentation]. Figures 2-6 Related descriptions.
[0109] An audio encoding module is a module used to extract features from audio data and obtain audio features. In some embodiments, the audio encoding module can be a machine learning model, a signal processing algorithm module, or any component capable of performing audio feature extraction. For example, the audio encoding module can be a pre-trained Wav2Vec2.0 model, a HuberT model, a VGGish model, or a feature extraction algorithm module based on Mel-Frequency Cepstral Coefficients (MFCC), etc.
[0110] In some embodiments, the audio encoding module takes audio as input and outputs audio features of the audio.
[0111] An audio modulation module is a module used to modulate audio features to obtain modulated audio features. In some embodiments, the audio modulation module can be a machine learning model, a mathematical operation unit, or any component capable of modulating audio features. For example, the audio modulation module can be an MLP, an attention-based Transformer network, a feature affine transformation unit, or a processing module that implements feature concatenation and mapping, etc.
[0112] In some embodiments, the audio modulation module takes target motion control parameters and audio features as input and outputs modulated audio features as output. For details on the specific implementation of the audio modulation module, please refer to the following related descriptions.
[0113] An image encoding module is a module used to extract features from an image containing a target object, thereby obtaining the image's features. In some embodiments, the image encoding module can be a machine learning model, an image processing algorithm module, or any component capable of performing image feature extraction. For example, the image encoding module can be a pre-trained ResNet model, a Vision Transformer (ViT) model, a CLIP image encoder, etc.
[0114] In some embodiments, the input to the image encoding module is an image, and the output is the image features of the image.
[0115] A video generation module is a module that generates a target video based on modulated audio features and image features, depicting the audio of a target object. In some embodiments, the video generation module can be a machine learning model, a video synthesis algorithm module, or any component capable of video generation. For example, the video generation module can be a video generation network based on a diffusion model, a video generation network based on a Transformer architecture, or a video generation network based on a Generative Adversarial Network (GAN), etc.
[0116] In some embodiments, the input to the video generation module may include modulated audio features and image features, and the output is the target video.
[0117] In some embodiments, the video generation model can be obtained by training a second initial machine learning model based on multiple second training samples. Each of the multiple second training samples includes sample motion control parameters, sample audio, and sample image. The second training label corresponding to each second training sample is a sample video, and the sample image is a reference image frame extracted from the sample video.
[0118] The second initial machine learning model refers to the base machine learning model to be trained. The trained second initial machine learning model is the video generation model. For example, the second initial machine learning model can be a neural network model that has not been trained or has been partially trained, and includes an audio encoding module, an audio modulation module, an image encoding module, and a video generation module.
[0119] A reference image frame is a frame selected from a sample video and used as input to the image encoding module in the video generation model or the second initial machine learning model to provide the visual basis for the target object. For example, the reference image frame can be any sample video frame from the sample video.
[0120] In some embodiments, the second training samples and the second training labels can be constructed simultaneously based on sample videos: for each sample video (as the second training label), the processor can extract sample audio and a reference image frame (as the sample image); simultaneously, based on the sample video, according to... Figures 3-5 The method described above determines the motion control parameters of the sample. Thus, based on the same sample video, a first training sample (sample audio, sample motion control parameters) for determining the motion control parameters and a second training sample (sample motion control parameters, sample audio, sample image) for the video generation model can be constructed simultaneously, along with their corresponding second training labels (sample video). The processor can input the second training samples (sample motion control parameters, sample audio, sample image) into the second initial machine learning model, construct a loss function based on the output of the second initial machine learning model and the second training labels (i.e., sample video), and iteratively update the parameters of the second initial machine learning model based on multiple second training samples until preset conditions are met, thus completing the training and obtaining the video generation model.
[0121] In some embodiments, the audio encoding module, audio modulation module, image encoding module, and video generation module in the video generation model can be jointly trained. The processor can input sample audio into the audio encoding module to obtain sample audio features; input sample motion control parameters and sample audio features into the audio modulation module to obtain sample modulated audio features; input sample images into the image encoding module to obtain sample image features; and input the sample modulated audio features and sample image features into the video generation module to obtain a predicted video. A loss function is constructed based on the predicted video and the sample video (i.e., the second training label), and the parameters of the audio encoding module, audio modulation module, image encoding module, and video generation module are jointly updated using a backpropagation algorithm until preset conditions are met, resulting in a trained video generation model. Through joint training, the modules can collaboratively optimize, enabling audio feature extraction, motion amplitude modulation, and video generation processes to adapt to each other, thus improving the overall quality of the generated video.
[0122] In the embodiments of this specification, by simultaneously constructing training samples for both the motion control parameter determination model and the video generation model based on the same sample video, the consistency between the motion control parameters and the video content is ensured. This enables the video generation model to learn the mapping relationship for generating videos with corresponding motion amplitudes under given motion control parameters, thereby improving the efficiency of model training and the controllability of generation results.
[0123] In some embodiments, such as Figure 7 As shown, the processor can determine the mask image 740 of the target region of the target object in the image; and input the modulated audio features 640, image features 650 and mask image 740 into the video generation module 710-4 to generate the target video 670.
[0124] A mask image is a binary image or weighted image used to identify the target region of a target object in an image. It guides the video generation module to apply the modulated audio features primarily to the target region. For example, a mask image can be a binary image with pixel values of 1 for the face region and 0 for the background region.
[0125] In some embodiments, the processor can determine the mask image of the target region of a target object in an image in a variety of ways. For example, the processor can use an image recognition algorithm to perform face detection in the image to determine the face location, locate facial key points, and generate a mask image of the face region based on the key points; or it can directly generate the mask image of the target region using a pre-trained face parsing model (such as the FaceParsing model).
[0126] In some embodiments, the second training samples may further include sample mask images, which can be determined based on the sample images in the same manner as described above. When training the video generation model, sample motion control parameters, sample audio, sample images, and sample mask images are input into the second initial machine learning model, and a loss function is constructed based on the output of the second initial machine learning model and the sample videos for training.
[0127] In the embodiments described in the specification, a mask image is introduced to guide the modulated audio features to act primarily on the target area (such as the face area), avoiding unnecessary interference with non-target areas (such as the background, body, etc.), making the motion amplitude control more precise, and improving the matching degree between the motion of the target area and the audio content in the generated video.
[0128] In some embodiments, the audio modulation module is configured to perform positional multiplication of the audio features and the target motion control parameters to obtain the modulated audio features.
[0129] Positional multiplication refers to performing element-wise multiplication of audio features and target motion control parameters at the same time dimension. In actual model processing, data is usually input in batches, so audio features and target motion control parameters often contain a batch dimension. For example, if the audio features are feature tensors of shape [B, T, C] (B is the batch size, T is the number of time frames, and C is the number of feature channels), and the target motion control parameters are tensors of shape [B, T, 1], then the audio modulation module can broadcast the target motion control parameters along the dimension of the number of feature channels (i.e., copy and expand to C channels), making its shape [B, T, C], and then multiply it element-wise with the audio features to obtain the modulated audio features, whose shape remains [B, T, C]. For simplicity, when batch processing is not involved, it can also be represented as broadcasting the audio features of [T, C] and the target motion control parameters of [T, 1] and then performing positional multiplication. The broadcast mechanism is implemented as follows: when two tensors are used in element-wise operations, if their dimensions do not perfectly match, the computational framework (such as PyTorch or TensorFlow) will automatically copy and expand the dimension of size 1 to match the corresponding dimension of the other tensor. In this example, the last dimension (number of feature channels) of the target motion control parameter is 1. Through the broadcast mechanism, this dimension will be copied C times, expanding the target motion control parameter to C identical values in the dimension of the number of feature channels, thus enabling element-wise multiplication with each feature channel of the audio feature.
[0130] In the embodiments of this specification, the modulation of audio features is achieved through a simple operation method of bitwise multiplication, which is computationally efficient, easy to implement, and can quickly inject motion amplitude control information into audio features.
[0131] In some embodiments, the audio modulation module includes a plurality of audio modulation units, each configured to generate modulated audio features based on audio features and target motion control parameters; and the video generation module includes a plurality of blocks, each of which is configured to process the intermediate video generated by the previous block and the modulated audio features to generate its own intermediate video.
[0132] An audio modulation unit (AMU) is a sub-unit within an audio modulation module. Each AMU can independently modulate audio features. Each AMU can be an independent MLP or other type of network model. Multiple AMUs can share parameters (i.e., use the same network structure and weights) or operate independently (i.e., each AMU has its own network structure and weights).
[0133] A block refers to a basic processing unit within a video generation module. Each block can progressively process the input video features to generate a higher-level video representation. For example, a block could be the Denoising Block in a diffusion model or the Decoder Layer in a Transformer.
[0134] Figure 8 This is an exemplary schematic diagram illustrating the generation of a target video through multiple blocks included in a video generation module, according to some embodiments of this specification.
[0135] like Figure 8As shown, multiple blocks in the video generation module 710-4 (the first block 710-4-1, the second block 710-4-2, ..., the m-th block 710-4-m) are arranged in series in a pre-set order, with the output of the previous block serving as the input of the next block. The first block 710-4-1 receives image features 650 from the image encoding module as input and processes them in conjunction with modulated audio features 640 to output an intermediate video (which can be denoted as the first intermediate video 810-1); the second block 710-4-2 receives the first intermediate video 810-1 output by the first block 710-4-1 as input and processes it in conjunction with modulated audio features 640 to output an intermediate video (which can be denoted as the second intermediate video 810-2); and so on, with the last block (the m-th block 710-4-m) outputting the final target video 670. A higher level of video representation refers to the process of abstracting video features from shallow details (such as edges and textures) to deeper semantic information (such as object structure and motion patterns) as blocks are processed, while gradually recovering the spatial and temporal details of video frames. The intermediate video generated by the previous block serves as the input for the next block, and the next block further optimizes and improves the video content based on the previous block, making the generated result more refined and natural.
[0136] Intermediate video refers to the intermediate representation output after each block of the video generation module has been processed. Depending on the specific architecture of the video generation module, intermediate video can be the hidden features of the video frame sequence (such as feature map tensors) or a pre-generated video frame sequence with a certain resolution. For example, in a diffusion model, intermediate video can be the hidden representations of different denoising stages; in a Transformer-based generative model, intermediate video can be a feature sequence after self-attention processing.
[0137] Each audio modulation unit can generate modulated audio features in several ways. For example, each audio modulation unit can determine scaling and translation coefficients based on target motion control parameters and modulate the audio features; or it can encode the target motion control parameters, concatenate them with the audio features, and then obtain the modulated audio features through an MLP. See the description below for details.
[0138] In some embodiments, when each block processes the intermediate video generated by the previous block and the modulated audio features, the modulated audio features can be fused into the intermediate video through methods such as splicing, addition, and cross-attention to generate a new intermediate video containing motion amplitude control information. The modulated audio features used here can be the same modulated audio feature generated by a single audio modulation unit and uniformly input into all blocks, or multiple modulated audio features generated independently by audio modulation units connected to each block and input into their respective blocks. See the relevant description below for details.
[0139] In the embodiments described in this specification, by injecting modulated audio features layer by layer into multiple blocks of the video generation module, motion amplitude control information can continue to play a role in various stages of video generation, thereby improving the transmission efficiency of control information and the generation effect.
[0140] In some embodiments, each of the plurality of audio modulation units is configured to: determine a scaling factor and a translation factor that matches the size of the audio features based on target motion control parameters; and process the audio features based on the scaling factor and the translation factor to obtain modulated audio features.
[0141] Scaling and translation factors are parameters used to perform affine transformations on audio features. Scaling factors control the magnification or reduction of audio features, while translation factors control their offset. Scaling and translation factors can be vectors or tensors that match the size of the audio features. For example, if the shape of an audio feature is [B, T, C] (B is the batch size, T is the number of time frames, and C is the number of feature channels), then the shapes of the scaling and translation factors should also be [B, T, C], perfectly corresponding to the audio feature in the batch, time, and channel dimensions, ensuring that each feature channel in each time frame can be subjected to an independent affine transformation.
[0142] In some embodiments, the audio modulation unit can map the target motion control parameters using a multilayer perceptron (MLP) to predict scaling and translation coefficients. For example, the MLP processes the target motion control parameters and outputs two tensors that match the size of the audio features, which serve as the scaling and translation coefficients, respectively. The MLP in each audio modulation unit is trained independently; therefore, even if multiple audio modulation units receive the same audio features and target motion control parameters, the MLPs in different audio modulation units can learn different mapping relationships, thereby predicting different scaling and translation coefficients and generating diverse modulation effects.
[0143] In some embodiments, the modulated audio features can be calculated using the formula "Modulated audio features = Audio features × Scaling factor + Translation factor". This formula is an element-wise operation, where "×" represents element-wise multiplication and "+" represents element-wise addition.
[0144] Taking an audio feature with B=1 and T=1 as an example, assuming its feature channel number C=3, the audio feature is [0.5, 1.0, -0.2], the corresponding scaling factor is [0.8, 1.2, -0.5], and the translation factor is [0.1, -0.1, 0.3]. The modulated audio feature is calculated as follows: First channel: 0.5×0.8+0.1=0.4+0.1=0.5; Second channel: 1.0×1.2+(-0.1)=1.2-0.1=1.1; Third channel: -0.2×(-0.5)+0.3=0.1+0.3=0.4. It can be seen that the scaling factor controls the scaling degree of each feature channel value: a value greater than 1 enhances the feature channel value (e.g., the second channel), a value between 0 and 1 weakens the feature channel value (e.g., the first channel), and a negative value inverts the sign of the feature channel value (e.g., the third channel). The translation factor adds an offset to the scaled value. Through the combined effect of scaling and translation factors, audio features are modulated into new features carrying motion amplitude control information, thereby guiding the video generation process.
[0145] It should be noted that in an audio modulation module architecture containing multiple audio modulation units, each audio modulation unit can process the same audio features (i.e., the same audio features from the audio coding module). However, since the MLP networks in each audio modulation unit are trained independently, they will predict different scaling and translation coefficients based on the same input. Therefore, the modulated audio features output by each audio modulation unit are different. These different modulated audio features are then input into different blocks of the video generation module to achieve multi-level, differentiated motion amplitude control.
[0146] In the embodiments described in this specification, affine transformations of audio features are performed using learnable scaling and translation coefficients, which can flexibly adjust the intensity distribution of audio features at different times and in different channels, thereby achieving fine-grained control over motion amplitude.
[0147] In some embodiments, each of the plurality of audio modulation units is configured to: encode the target motion control parameters to obtain motion control parameter features; concatenate the audio features and motion control parameter features to obtain audio-motion control parameter features; and determine the modulated audio features based on the audio-motion control parameter features.
[0148] Motion control parameter features refer to the feature representation obtained after encoding the target motion control parameters, which are used to map the scalar form of motion control parameters to a feature dimension space suitable for interacting with audio features.
[0149] In some embodiments, the processor can encode the target motion control parameters into motion control parameter features using methods such as trigonometric function position encoding and linear layer mapping. The purpose of encoding the target motion control parameters into motion control parameter features is to elevate them to the same semantic level and temporal resolution as the audio features, enabling effective information fusion. Therefore, the motion control parameter features and audio features must be identical in the batch dimension B and temporal dimension T (to ensure time alignment), but they can be equal or unequal in the channel dimension. For example, the shape of the motion control parameter features can be [B, T, C]. ], C and C They can be equal or unequal.
[0150] Audio-motion control parameter features refer to the fused features obtained by concatenating audio features and motion control parameter features along the channel dimension.
[0151] Channel concatenation refers to the operation of merging two feature tensors along the feature dimension (i.e., the channel dimension). For example, if the shape of the audio feature is [B, T, C], and the shape of the motion control parameter feature is [B, T, C]... If the shape of the audio-motion control parameter feature obtained after splicing is [B, T, C+C], then the feature shape is [B, T, C+C]. The splicing operation merges motion control parameter features and audio features along the channel dimension, so that the fused features (i.e., audio-motion control parameter features) simultaneously contain audio content and motion amplitude information.
[0152] In some embodiments, the audio modulation unit can input audio-motion control parameter features into the MLP for processing, and the MLP can change the number of channels from C+C. Mapping back to C, the modulated audio features are output with a shape of [B, T, C]. Each audio modulation unit can have an independent MLP, so even with the same input audio-motion control parameters, different units can output different modulated audio features, thus achieving diverse modulation effects.
[0153] In the embodiments described in this specification, by splicing and fusing audio features with motion control parameter features, the audio modulation unit can fully learn the interaction between audio content and motion amplitude, and generate more expressive modulated audio features.
[0154] In some embodiments, each of the plurality of audio modulation units is a trained machine learning model.
[0155] A trained machine learning model refers to an audio modulation unit (EMU) that has been trained and whose parameters are determined. Each EMU can be an independent neural network model, such as an MLP or a one-dimensional convolutional network. During the training of the video generation model, the EMU is jointly trained end-to-end with the video generation module, audio encoding module, and image encoding module to learn how to optimally modulate audio features based on target motion control parameters. The joint training process can be found in the relevant description above.
[0156] In the embodiments described in this specification, each audio modulation unit is designed as a trainable machine learning model, enabling it to adaptively learn the fusion method of motion amplitude control information and audio features, thereby improving the modulation effect and the overall performance of the model.
[0157] Figure 9 This is an exemplary schematic diagram illustrating the generation of a target video through multiple audio modulation units included in the audio modulation module and multiple blocks included in the video generation module, according to some embodiments of this specification.
[0158] In some embodiments, such as Figure 9 As shown, the audio modulation module 710-2 includes multiple audio modulation units (audio modulation unit 710-2-1, ..., audio modulation unit 710-2-j, ..., audio modulation unit 710-2-m, where j is any value from 2, ..., m-1). The processor can input the modulated audio features and image features to the video generation module to generate the target video, including: using the modulated audio features generated by one of the multiple audio modulation units (such as...) Figure 9 The modulated audio feature 640-j generated by the audio modulation unit 710-2-j shown can also be the modulated audio feature generated by any other audio modulation unit. This is input to each of the multiple blocks (e.g., ...). Figure 9 The first block 710-4-1, ..., the j-th block 710-4-j, ..., the m-th block 710-4-m shown are used to generate the target video 670.
[0159] In some embodiments, the processor may first perform trigonometric function position encoding on the target motion control parameters, expanding them into motion control parameter features (shape [T, C]) that match the size of the audio features; then, the audio features and the motion control parameter features are concatenated along the channel dimension to obtain audio-motion control parameter features (shape [T, 2C]); the concatenated audio-motion control parameter features are input into an audio modulation unit for adaptive modulation to obtain modulated audio features. These modulated audio features are then input into each block of the video generation module. Each block, while processing the intermediate video generated by the previous block, fuses the modulated audio features through cross-attention, concatenation, or addition to generate its own intermediate video. Finally, the target video is output by the last block.
[0160] In the embodiments described in this specification, a unified modulated audio feature is generated by a single audio modulation unit and injected into all blocks, which ensures the consistency of motion amplitude control information in the processing of each layer, while reducing the number of model parameters and computational complexity.
[0161] Figure 10 This is another exemplary schematic diagram illustrating the generation of a target video through a plurality of audio modulation units included in an audio modulation module and a plurality of blocks included in a video generation module, according to other embodiments of this specification.
[0162] In some embodiments, each of the plurality of audio modulation units is connected to one of the plurality of blocks (e.g., ...). Figure 10 The audio modulation unit 710-2-1 shown is connected to the first block 710-4-1, ..., audio modulation unit 710-2-j is connected to the j-th block 710-4-j, ..., audio modulation unit 710-2-m is connected to the m-th block 710-4-m. The processor inputs the modulated audio features and image features to the video generation module to generate the target video, which may include: using the modulated audio features generated by multiple audio modulation units (such as...) Figure 10 The modulated audio features 640-1, ..., modulated audio features 640-j, ..., modulated audio features 640-m shown are respectively input to the block connected to the audio modulation unit (e.g., ...). Figure 10 The first block 710-4-1, ..., the j-th block 710-4-j, ..., the m-th block 710-4-m shown are used to generate the target video 670.
[0163] In some embodiments, the processor may first perform trigonometric function position encoding on the target motion control parameters, expanding them into motion control parameter features that match the audio feature size. Then, the audio features and motion control parameter features are input together into m independent audio modulation units (e.g., m MLPs). Each audio modulation unit independently modulates the audio features, generating m corresponding modulated audio features. The m audio modulation units are connected to m blocks in the video generation module. The processor inputs the modulated audio features generated by each audio modulation unit into its corresponding connected block. When processing the intermediate video generated by the previous block, each block fuses its corresponding input modulated audio features to generate its own intermediate video. Finally, the target video is output by the last block.
[0164] In the embodiments described in this specification, by configuring an independent audio modulation unit for each block of the video generation module, blocks at different levels can learn motion amplitude control information at different abstract levels, thereby improving the precision of control and the diversity of generation effects.
[0165] In the embodiments described in this specification, an audio encoding module, an audio modulation module, an image encoding module, and a video generation module are integrated into a unified video generation model. Audio features are modulated based on target motion control parameters and then injected into the video generation process, achieving end-to-end generation from audio to video. This structure can precisely control the motion amplitude of the target object in the generated video, ensuring a high degree of match between the generated result and the audio content in terms of motion intensity, resulting in high-quality videos with natural motion and emotional consistency. Furthermore, the introduction of mask images and multi-level modulation units further enhances the accuracy of control and the flexibility of the model.
[0166] Figure 11 This is an exemplary block diagram of an audio-driven video generation system according to some embodiments of this specification. Figure 11 As shown, the audio-driven video generation system 1100 includes an acquisition module 1110, a parameter determination module 1120, and a video generation module 1130.
[0167] An acquisition module refers to a module used to acquire data. In some embodiments, the acquisition module is configured to acquire audio and images, where the images include the target object.
[0168] The parameter determination module refers to the module used for target motion control parameters. In some embodiments, the parameter determination module is configured to determine the target motion control parameters when the target object performs audio, based on the audio.
[0169] A video generation module refers to a module used to generate video. In some embodiments, the video generation module can be configured to generate a target video of the target object performing audio based on target motion control parameters, audio, and images using a video generation model.
[0170] In some embodiments, the acquisition module, parameter determination module, and video generation module may be stored in memory and executed by a processor to perform their functions. For a related description of the execution content, please refer to [link to relevant documentation]. Figures 2-10 Related descriptions.
[0171] Figure 12 These are schematic diagrams of computing devices illustrated according to some embodiments of this specification. For example... Figure 12 As shown, the computing device 1200 includes a memory 1210, a processor 1220, and a computer program stored in the memory and executable on the processor, wherein the processor 1220 executes the program to implement the steps of audio-driven video generation.
[0172] In some embodiments, Figure 12 The system also includes a bus architecture (represented by bus 1230), which may include any number of interconnected buses and bridges. Bus 1230 links together various circuits including one or more processors represented by processor 1220 and memory represented by memory 1210. Bus 1230 may also link together various other circuits such as peripherals, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. Bus interface 1240 provides an interface between bus 1230 and receiver 1250 and transmitter 1260. Receiver 1250 and transmitter 1260 may be the same element, i.e., a transceiver, providing a unit for communicating with various other devices over a transmission medium. Processor 1220 is responsible for managing bus 1230 and general processing, while memory 1210 may be used to store data used by processor 1220 during operation.
[0173] This embodiment provides a computer-readable storage medium having a computer program stored thereon, which is executed by a processor to implement audio-driven video generation.
[0174] Those skilled in the art will understand that embodiments of this specification can be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this specification may take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. This specification is described with reference to flowchart illustrations and / or block diagrams of methods, systems, and computer program products according to embodiments of this specification. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0175] Those skilled in the art will understand that Figure 12 The structure shown is merely a block diagram of a portion of the structure related to the scheme described in this specification, and does not constitute a limitation on the computing device 1200 to which the scheme described in this specification is applied. The specific computing device 1200 may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0176] The basic concepts have been described above. Obviously, for those skilled in the art, the detailed disclosure above is merely illustrative and does not constitute a limitation of this specification. Although not explicitly stated herein, those skilled in the art may make various modifications, improvements, and corrections to this specification. Such modifications, improvements, and corrections are suggested in this specification and therefore remain within the spirit and scope of the exemplary embodiments described herein.
[0177] Furthermore, this specification uses specific terms to describe embodiments thereof. For example, "an embodiment," "one embodiment," and / or "some embodiments" refer to a particular feature, structure, or characteristic associated with at least one embodiment of this specification. Therefore, it should be emphasized and noted that references to "an embodiment," "one embodiment," or "an alternative embodiment" in different locations throughout this specification do not necessarily refer to the same embodiment. Moreover, certain features, structures, or characteristics in one or more embodiments of this specification can be appropriately combined.
[0178] Furthermore, unless expressly stated in the claims, the order of processing elements and sequences, the use of numbers and letters, or other names described in this specification are not intended to limit the order of the processes and methods described herein. Although various examples have been discussed in the foregoing disclosure of some embodiments of the invention that are currently considered useful, it should be understood that such details are for illustrative purposes only, and the appended claims are not limited to the disclosed embodiments; rather, the claims are intended to cover all modifications and equivalent combinations that conform to the spirit and scope of the embodiments described herein. For example, while the system components described above can be implemented using hardware devices, they can also be implemented solely using software solutions, such as installing the described system on existing servers or mobile devices.
[0179] Similarly, it should be noted that, in order to simplify the description disclosed herein and thus aid in the understanding of one or more embodiments of the invention, the foregoing description of embodiments in this specification may sometimes combine multiple features into a single embodiment, drawing, or description thereof. However, this method of disclosure does not imply that the subject matter of this specification requires more features than those mentioned in the claims. In fact, the embodiments contain fewer features than all the features of a single embodiment disclosed above.
[0180] In some embodiments, numbers describing the quantity of components and attributes are used. It should be understood that such numbers used in the description of embodiments are modified in some examples with the terms "approximately," "approximately," or "generally." Unless otherwise stated, "approximately," "approximately," or "generally" indicates that the numbers are allowed to vary by ±20%. Accordingly, in some embodiments, the numerical parameters used in the specification and claims are approximate values, which may be changed depending on the characteristics required by individual embodiments. In some embodiments, numerical parameters should take into account specified significant digits and employ a general method of digit reservation. Although the numerical ranges and parameters used to confirm their breadth of range in some embodiments of this specification are approximate values, in specific embodiments, such values are set as precisely as feasible.
[0181] For each patent, patent application, patent application publication, and other material such as articles, books, specifications, publications, and documents referenced in this specification, the entire contents of which are incorporated herein by reference. This excludes historical application documents that are inconsistent with or conflict with the content of this specification, as well as documents that limit the broadest scope of the claims in this specification (currently or subsequently appended to this specification). It should be noted that in the event of any inconsistency or conflict between the descriptions, definitions, and / or terminology used in the supplementary materials to this specification and the content of this specification, the descriptions, definitions, and / or terminology used in this specification shall prevail.
[0182] Finally, it should be understood that the embodiments described in this specification are merely illustrative of the principles of the embodiments described herein. Other variations may also fall within the scope of this specification. Therefore, alternative configurations of the embodiments described herein are intended to be illustrative rather than limiting, and should be considered consistent with the teachings of this specification. Accordingly, the embodiments described herein are not limited to those explicitly introduced and described herein.
Claims
1. An audio-driven video generation method, characterized in that, The method includes: Acquire audio and images, wherein the images include the target object; Based on the audio, determine the target motion control parameters when the target object performs the audio; and Based on the target motion control parameters, the audio, and the image, a target video of the target object performing the audio is generated using a video generation model.
2. The method according to claim 1, characterized in that, The determination of target motion control parameters when the target object performs the audio based on the audio includes: Based on the audio, the target motion control parameters are determined by a motion control parameter determination model, which is a trained machine learning model.
3. The method according to claim 2, characterized in that, The determination of the target motion control parameters based on the audio through a motion control parameter determination model includes: Based on the audio, the initial motion control parameters are determined using the motion control parameter determination model; and Based on the initial motion control parameters and motion intensity coefficient, the target motion control parameters are determined, wherein the motion intensity coefficient is determined based on user input.
4. The method according to claim 2, characterized in that, The motion control parameter determination model is obtained by training a first initial machine learning model based on multiple first training samples. Each of the multiple first training samples includes sample audio, and the first training label corresponding to each first training sample is the sample motion control parameter. The method further includes: Acquire multiple sample videos; and Based on each of the plurality of sample videos, determine each first training sample and the first training label corresponding to each first training sample, including: Separate the sample audio and sample video frame sequence from a sample video; Extract key points from each sample video frame in the sample video frame sequence; Based on the aforementioned key points, determine the motion amplitude representation parameters corresponding to each sample video frame; and Based on the motion amplitude characterization parameters corresponding to each sample video frame, the sample motion control parameters of the sample video are determined.
5. The method according to claim 4, characterized in that, The determination of motion amplitude representation parameters corresponding to each sample video frame based on the key points includes: Multiple pairs of key points located within the target area are determined from the key points extracted from the same sample video frame. Each pair of key points includes two key points, which are symmetrically distributed with respect to the axial direction of the target area. Determine the distance between the two keypoints included in each frame keypoint pair; and The average distance between the multiple key point pairs in the same frame within the target area is determined as the motion amplitude representation parameter of the target area in the sample video frame.
6. The method according to claim 4, characterized in that, The determination of motion amplitude representation parameters corresponding to each sample video frame based on the key points includes: From the key points extracted from two adjacent sample video frames, multiple pairs of neighboring frame key points located within the target area are determined. Each pair of neighboring frame key points includes two key points located in the two adjacent sample video frames. The two key points correspond to the same part or position of an object in the sample video. Determine the distance between the two keypoints included in each neighboring frame keypoint pair; and The average distance between the multiple neighboring keypoint pairs within the target region is determined and used as the motion amplitude representation parameter of the target region in the subsequent sample video frame of the two adjacent sample video frames.
7. The method according to claim 4, characterized in that, The determination of motion amplitude representation parameters corresponding to each sample video frame based on the key points includes: Multiple pairs of key points located within the target area are determined from the key points extracted from the same sample video frame. Each pair of key points includes two key points, which are symmetrically distributed with respect to the axial direction of the target area. Based on the multiple same-frame key point pairs, a reference key point pair is determined; and The distance between the two reference key points included in the reference key point pair is determined as the motion amplitude representation parameter of the target region in the sample video frame.
8. The method according to claim 4, characterized in that, The step of determining the sample motion control parameters of the sample video based on the motion amplitude representation parameters corresponding to each sample video frame includes: Determine the distribution statistics of multiple motion amplitude characterization parameters corresponding to the multiple sample video frames; and Based on the distribution statistics, the multiple motion amplitude representation parameters corresponding to the multiple sample video frames are normalized to determine the sample motion control parameters corresponding to each sample video.
9. The method according to claim 1, characterized in that, The step of generating a target video of the target object performing the audio based on the target motion control parameters, the audio, and the image through a video generation model includes: Obtain the audio features of the audio; Based on the target motion control parameters, the audio features are modulated using an audio modulation model to obtain modulated audio features; and The target video is generated based on the modulated audio features and the image features of the image, using the video generation model.
10. The method according to claim 1, characterized in that, The video generation model includes an audio encoding module, an audio modulation module, an image encoding module, and a video generation module. The step of generating a target video of the target object performing the audio based on the target motion control parameters, the audio, and the image through the video generation model includes: The audio is input to the audio encoding module to obtain the audio features of the audio; The target motion control parameters and the audio features are input into the audio modulation module to obtain modulated audio features; The image is input to the image encoding module to obtain the image features; and The modulated audio features and the image features are input into the video generation module to generate the target video.
11. The method according to claim 10, characterized in that, The step of inputting the modulated audio features and the image features into the video generation module to generate the target video includes: Determine the mask image of the target region of the target object in the image; and The modulated audio features, the image features, and the mask image are input into the video generation module to generate the target video.
12. The method according to claim 10, characterized in that, The audio modulation module is configured to perform positional multiplication of the audio features and the target motion control parameters to obtain the modulated audio features.
13. The method according to claim 10, characterized in that, The audio modulation module includes multiple audio modulation units, which are configured to generate the modulated audio features based on the audio features and the target motion control parameters; and The video generation module includes multiple blocks, each of which is configured to process the intermediate video generated by the previous block and the modulated audio features to generate its own intermediate video.
14. The method according to claim 13, characterized in that, The process of inputting the modulated audio features and the image features into the video generation module to generate the target video includes: The modulated audio features generated by one of the plurality of audio modulation units are input into each of the plurality of blocks to generate the target video.
15. The method according to claim 13, characterized in that, Each of the plurality of audio modulation units is connected to one of the plurality of blocks, and the step of inputting the modulated audio features and the image features into the video generation module to generate the target video includes: The modulated audio features generated by each of the plurality of audio modulation units are respectively input into each block connected to each of the audio modulation units to generate the target video.
16. The method according to claim 13, characterized in that, Each of the plurality of audio modulation units is configured as follows: Based on the target motion control parameters, a scaling factor and translation factor matching the size of the audio feature are determined; and The modulated audio features are obtained by processing the audio features based on the scaling factor and the translation factor.
17. The method according to claim 13, characterized in that, Each of the plurality of audio modulation units is configured as follows: The target motion control parameters are encoded to obtain motion control parameter features; The audio features and the motion control parameter features are concatenated to obtain audio-motion control parameter features; and The modulated audio features are determined based on the audio-motion control parameter features.
18. The method according to claim 13, characterized in that, Each of the plurality of audio modulation units is a trained machine learning model.
19. The method according to claim 1, characterized in that, The video generation model is obtained by training a second initial machine learning model based on multiple second training samples. Each of the multiple second training samples includes sample motion control parameters, sample audio, and sample image. The second training label corresponding to each second training sample is a sample video, and the sample image is a reference image frame extracted from the sample video.
20. An audio-driven video generation system, characterized in that, The system includes: The acquisition module is configured to acquire audio and images, the images including the target object; The parameter determination module is configured to determine target motion control parameters of the target object when it performs the audio, based on the audio. The video generation module is configured to generate a target video of the target object performing the audio based on the target motion control parameters, the audio, and the image, using a video generation model.
21. A computer-readable storage medium, characterized in that, The storage medium stores computer instructions that are executed by a processor to implement the method as described in any one of claims 1-19.
22. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method as described in any one of claims 1 to 19.