Video generation method, video live broadcast method and training method of video generation model
By extracting the identification features of the sound-producing object from the reference image and combining the target audio to determine the motion features, and processing the appearance and motion information separately, the problem of low video generation speed and efficiency in the existing technology is solved, and the requirement for real-time playback is realized.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-13
- Publication Date
- 2026-03-10
AI Technical Summary
Existing technologies require the generation of the overall features of the sound-producing object for each frame, resulting in low video generation speed and efficiency, which cannot meet the needs of real-time playback, especially in practical application scenarios such as live video streaming.
By extracting the identification features of the sound-producing object from the reference image and combining them with the target audio to determine the target motion features, the appearance and motion information are processed separately to generate target video frames, avoiding the need to generate overall features for each frame.
It improves the speed and efficiency of video generation, meets the real-time playback needs of live video streaming and other practical application scenarios, and brings users a smooth viewing experience.
Smart Images

Figure CN121644922A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video generation technology, specifically to a video generation method, a live video streaming method, a training method for a video generation model, an apparatus, an electronic device, and a computer-readable storage medium. Background Technology
[0002] With the rapid development of computer and artificial intelligence technologies, video synthesis technology has been increasingly applied in various scenarios. Video synthesis usually involves first acquiring an audio clip, then generating a corresponding video based on that audio. The generated video can be made to sound the audio from a specific person. Video synthesis can be applied to various scenarios such as voice broadcasting, movie and TV clip generation, animation generation, user interaction, and live streaming.
[0003] When generating videos corresponding to audio, related technologies typically determine the overall characteristics of the sound-producing object in each corresponding video frame directly based on the audio, thereby generating the corresponding video frame.
[0004] However, this method of video generation is slow and inefficient because each frame requires the generation of the overall features of the sound-producing object. In practical applications, such as live video streaming, this method may not meet the requirements for real-time playback. Summary of the Invention
[0005] This application provides a video generation method, a live video streaming method, a training method for a video generation model, an apparatus, an electronic device, and a computer-readable storage medium, which can improve the speed and efficiency of video generation and meet the real-time playback requirements in live video streaming or other practical application scenarios. The specific solution is as follows: In a first aspect, this application provides a video generation method, the method comprising: Obtain the target audio and reference image for generating the video, wherein the reference image includes the sound-producing object; The identification features of the sound-producing object are extracted from the reference image, and the identification features are used to represent the feature information of each part of the sound-producing object; Determine the target motion features that match the target audio, wherein the target motion features are used to represent the motion information of each part of the sound-producing object in each target video frame corresponding to the target audio; Based on the recognition features and the target motion features, the overall features of the sound-producing object in the target video frame are obtained, and the target video frame is obtained based on the overall features, so as to generate a corresponding video based on each target video frame and the target audio.
[0006] Secondly, this application provides a video live streaming method, including: Obtain the target audio and reference image corresponding to the video to be live-streamed, wherein the reference image includes the broadcaster object; A live video is generated using the video generation method described in the first aspect; The live video is played in real time via the client.
[0007] Thirdly, this application provides a training method for a video generation model, wherein the video generation model includes a motion feature generation model, and the method includes: Obtain a first training sample, which includes a sample video, a sample audio that matches the sample video, and a sample reference image. The sample video and the sample reference image include the same first sample sound source. Extract the first sample motion feature of the first sample sound-emitting object from the sample video frame of the sample video, and extract the second sample motion feature of the first sample sound-emitting object from the sample reference image; Based on the second sample motion features and the sample audio, the first sample input information is determined, and the first sample input information is input into the motion feature generation model to be trained to obtain the first output motion feature. Based on the difference between the first output motion feature and the first sample motion feature, the parameters of the motion feature generation model to be trained are adjusted to obtain the trained motion feature generation model.
[0008] Fourthly, this application provides a video generation apparatus, the apparatus comprising: An information acquisition unit is used to acquire target audio and reference images for generating video, wherein the reference images include the sound-emitting object; A feature extraction unit is used to extract the identification features of the sound-producing object from the reference image, wherein the identification features are used to represent the feature information of each part of the sound-producing object; A feature determination unit is used to determine target motion features that match the target audio, wherein the target motion features are used to represent the motion information of each part of the sound-producing object in each target video frame corresponding to the target audio; The video generation unit is configured to obtain the overall features of the sound-producing object in the target video frame based on the recognition features and the target motion features, and to obtain the target video frame based on the overall features, so as to generate a corresponding video based on each target video frame and the target audio.
[0009] Fifthly, this application provides a training apparatus for a video generation model, comprising: The sample acquisition unit is used to acquire a first training sample, which includes a sample video, a sample audio that matches the sample video, and a sample reference image. The sample video and the sample reference image include the same first sample sound source. The sample feature extraction unit is used to extract the first sample motion feature of the first sample sound-producing object from the sample video frame of the sample video, and to extract the second sample motion feature of the first sample sound-producing object from the sample reference image. The output feature unit is used to determine the first sample input information based on the second sample motion features and the sample audio, and input the first sample input information into the motion feature generation model to be trained to obtain the first output motion features. The training unit is used to adjust the parameters of the motion feature generation model to be trained based on the difference between the first output motion feature and the first sample motion feature, so as to obtain the trained motion feature generation model.
[0010] In a sixth aspect, this application also provides an electronic device, comprising: a processor, a memory, and computer program instructions stored in the memory and executable on the processor; wherein the processor, when executing the computer program instructions, implements the method as described in any one of the first to third aspects.
[0011] In a seventh aspect, this application provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the method described in any one of the first to third aspects.
[0012] Eighthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the method as described in any one of the first to third aspects.
[0013] Compared with the prior art, this application has the following advantages: The video generation method provided in this application embodiment obtains target audio and reference images for generating video. The reference images include a sound-producing object. Identification features of the sound-producing object are extracted from the reference images. Since the identification features represent the feature information of each part of the sound-producing object, they can determine the appearance of the sound-producing object. Next, target motion features matching the target audio are determined. These target motion features represent the motion information of each part of the sound-producing object in a target video frame corresponding to the target audio. Based on the identification features and the target motion features, the overall features of the sound-producing object in the target video frame are obtained. The target video frame is then obtained based on the overall features. A corresponding video can be generated based on each target video frame and the target audio.
[0014] The solution provided in this application determines the target motion features of various parts of the sound-producing object that match the target audio. These motion features effectively reflect the motion state of each part of the sound-producing object within the target video frame. Combined with the recognition features of the sound-producing object extracted from the reference image, the overall features of the sound-producing object in the target video frame can be determined. This avoids generating the overall features of the sound-producing object for each frame; instead, it processes the appearance and motion information of the sound-producing object separately. The recognition features representing the appearance only need to be determined once from the reference image, and each frame only needs to generate the motion features of each part of the sound-producing object. This significantly improves the speed and efficiency of video generation. For practical applications such as live video streaming and video production, it can meet the needs of real-time playback and rapid production, providing users with a smooth viewing experience. Attached Figure Description
[0015] Figure 1 This is a schematic diagram illustrating the application scenario of the solution provided in this application.
[0016] Figure 2 This is a flowchart illustrating an example of the video generation method provided in this application embodiment.
[0017] Figure 3 This is a flowchart illustrating another example of the video generation method provided in the embodiments of this application.
[0018] Figure 4 This is a schematic flowchart of an example of the training method for the video generation model provided in the embodiments of this application.
[0019] Figure 5 This is a schematic diagram illustrating the training of the motion feature generation model in an embodiment of this application.
[0020] Figure 6 This is a schematic diagram of training the motion feature database in an embodiment of this application.
[0021] Figure 7This is a schematic diagram of the training of the emotion projection model in an embodiment of this application.
[0022] Figure 8 This is a structural block diagram of an example of the video generation apparatus provided in the embodiments of this application.
[0023] Figure 9 This is a structural block diagram of the electronic device provided in this application. Detailed Implementation
[0024] To enable those skilled in the art to better understand the technical solutions of this application, the application will be clearly and completely described below with reference to the accompanying drawings of the embodiments. However, this application can be implemented in many other ways different from those described below. Therefore, based on the embodiments provided in this application, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this application.
[0025] It should be noted that the terms "first," "second," "third," etc., in the claims, specification, and drawings of this application are used to distinguish similar objects and are not used to describe a specific order or sequence. Such data are interchangeable where appropriate so that the embodiments of this application described herein can be implemented in a sequence other than that shown or described in this application. Furthermore, the terms "comprising," "having," and their variations are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or apparatuses.
[0026] It should be understood that in the embodiments of this application, "at least one" means one or more, and "more than one" means two or more. "And / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. The character " / " generally indicates that the related objects before and after it are in an "or" relationship. "Contains A, B and / or C" means containing any one, two, or three of A, B, and C.
[0027] It should be understood that in the embodiments of this application, "B corresponding to A", "B corresponding to A", "A corresponds to B" or "B corresponds to A" means that B is associated with A, and B can be determined based on A. Determining B based on A does not mean that B is determined solely based on A; B can also be determined based on A and / or other information.
[0028] To facilitate understanding of the various embodiments of this application, the application background of the embodiments will be explained.
[0029] With the rapid development of computer and artificial intelligence technologies, video synthesis technology is increasingly being applied in various scenarios. Video synthesis typically involves first acquiring an audio clip, then generating a corresponding video based on that audio. The generated video can then be made to sound like a specific person uttering the audio. Video synthesis can be applied to various scenarios such as voice broadcasting, film and television clip generation, animation generation, user interaction, and live streaming. When generating videos corresponding to audio, related technologies usually directly determine the overall features of the speaker in each corresponding video frame based on the audio, thereby generating the corresponding video frame. However, this method results in low speed and efficiency in video generation because the overall features of the speaker need to be generated for each frame. In practical applications, such as live video streaming, this video generation method may not meet the requirements for real-time playback.
[0030] To address the above issues, embodiments of this application provide a video generation method, a video generation model training method, an apparatus, an electronic device, a computer-readable storage medium, and a computer program product. The aim is to improve the speed and efficiency of video generation, meeting the real-time playback requirements of live video streaming or other practical application scenarios.
[0031] The video generation method provided in this application can be applied to video generation in various professional fields. Specifically, it can be applied to voice broadcasting, film and television clip generation, animation generation, video creation, user interaction, virtual assistants, etc., but is not limited to these.
[0032] To facilitate understanding of the method embodiments of this application, their application scenarios are described. Please refer to... Figure 1 , Figure 1 This is a schematic diagram illustrating an application scenario of the solution provided in the embodiments of this application. This application scenario is merely an illustrative example and is not intended to limit the specific application scenario. Figure 1 As shown, in this application scenario, a server 102 and a client 101 are provided. In this embodiment, the client 101 and the server 102 establish a connection through network communication to transmit data.
[0033] Client 101 can be an electronic device with display and data processing functions, such as a mobile phone, tablet, smartwatch, desktop computer, smart TV, VR device, in-vehicle device, wearable device, or laptop. Client 101 is used to obtain the target audio and reference image input by the user, and send the target audio and reference image to server 102, so that server 102 generates video frames that are synchronized with the target audio and based on the reference image, with sound emanating from the sound-emitting object in the reference image. Client 101 also obtains each video frame from server 102 to display the generated video to the user. Client 101 can also be used to send access requests, interactive information, etc. to server 102, so that server 102 sends the corresponding request data to client 101 for display.
[0034] Server 102 possesses high computing power. Server 102 can be a server, featuring high-speed central processing unit (CPU) computing power, long-term reliable operation, powerful input / output (I / O) external data throughput, and better scalability. Server 102 can be a single server or a server cluster. Server 102 is used to obtain target audio and reference images input by the user from client 101, and generate video frames that are synchronized with the target audio and based on the reference images, with sound emanating from the sound-emitting objects in the reference images. The generated video frames are then sent to client 101. Server 102 can also provide other specific services to client 101, such as user information access, website access, and application access, which are not specifically limited in this application.
[0035] Client 101 and server 102 can communicate using various communication systems, such as wired or wireless communication systems. Wireless communication systems can include, for example, Global System for Mobile Communications (GSM), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), General Packet Radio Service (GPRS), Long Term Evolution (LTE), LTE Frequency Division Duplex (FDD), LTE Time Division Duplex (TDD), Universal Mobile Telecommunication System (UMTS), Worldwide Interoperability for Microwave Access (WiMAX), future 5th generation (5G) systems or new radio (NR), and satellite communication systems.
[0036] In other application scenarios, only a client application is required, where users can select target audio and reference images, and the client can then generate a video based on these elements. The solution provided in this application can also be applied to other application scenarios; this application does not specifically limit its scope.
[0037] Example 1 The first embodiment of this application provides a video generation method, which can be applied to electronic devices, such as servers, desktop computers, laptops, mobile phones, tablets, smartwatches, smart TVs, VR devices, in-vehicle devices, wearable devices, and other electronic devices with data processing functions.
[0038] like Figure 2 As shown, the video generation method provided in the first embodiment of this application includes the following steps S110 to S140.
[0039] Step S110: Obtain the target audio and reference image for generating the video, wherein the reference image includes the sound source.
[0040] The target audio can be audio selected by the user, audio input by the user, audio downloaded from the network by the electronic device, or audio selected by the electronic device from pre-stored audio, etc. The target audio can include audio containing speech, audio containing singing, etc., and this application does not specifically limit it. It is understood that the target audio includes speech information such as speaking and singing, and not just audio without speech information such as light music.
[0041] The aforementioned reference images can be images selected or input by the user, or images downloaded from the network or selected from pre-stored images by the electronic device. Reference images can be photographs, composite images, video screenshots, cartoon character images, etc. The sound-producing objects included in the reference images can be people, cartoon characters, animals, game virtual characters, etc., but are not limited to these. Those skilled in the art can flexibly set the specific forms of the target audio and reference images; this application does not specifically limit them.
[0042] Step S120: Extract the identification features of the sound-producing object from the above reference image. The identification features are used to represent the feature information of each part of the sound-producing object.
[0043] The identification features can specifically be used to represent the shape and contour features of various parts of the voice-producing object, skin color features, and other features that can identify the voice-producing object.
[0044] Recognition features can include the facial features of the speaker. Since the speaker may also have corresponding body movements while speaking, recognition features can also include the body features to comprehensively represent the speaker's overall appearance. Facial features can include the speaker's facial contours, facial features, skin tone, etc., while body features can include the speaker's body proportions, clothing, limb shape, hand shape, etc. By extracting these recognition features, the speaker's appearance can be accurately determined, providing a foundation for subsequent generation of target video frames.
[0045] Specifically, a pre-trained shape feature recognition model can be used to extract the identification features of the sound-producing object from a reference image. This shape feature recognition model can be trained on a large number of images containing different sound-producing objects and can accurately capture the feature information of each part of the sound-producing object from the image. During training, a labeled image set can be used, in which the shape features of each part of the sound-producing object in the image are labeled features. Then, a deep learning algorithm is used to train and optimize the model to obtain a shape feature recognition model that can accurately extract the identification features from the input reference image. Those skilled in the art can use supervised training methods, unsupervised training methods, etc., from related technologies to train the shape feature recognition model; this application is not specifically limited to this.
[0046] Step S130: Determine the target motion features that match the target audio. These target motion features are used to represent the motion information of each part of the sound-emitting object in each target video frame corresponding to the target audio.
[0047] Step S130 specifically involves determining the target motion features that match each audio frame of the target audio. The target motion features corresponding to each audio frame represent the motion information of various parts of the sound-producing object when emitting the sound of that audio frame. The target motion features may include the degree of opening and closing of the lips, the degree of opening of the eyes, the angle of tilt of the corners of the mouth, the state of raising or lowering the eyebrows, the direction and angle of head rotation, the amplitude and direction of hand swing, the twisting posture of the body, etc. Those skilled in the art can flexibly define the specific content included in the target motion features according to actual needs.
[0048] Target motion features can be determined using a pre-trained motion feature generation model. This model can be trained on a large amount of audio and corresponding video data, containing the motion features of various parts of different voice-producing objects when emitting various audio sounds. During training, sample audio and corresponding sample video are first collected. The sample video is labeled with the motion features of various parts of the voice-producing object in each frame. Then, the sample audio and corresponding sample video are input into the motion feature generation model to be trained. Through deep learning algorithms, the model's parameters are continuously adjusted, enabling it to learn the mapping relationship between audio and the motion features of various parts of the voice-producing object. The resulting motion feature generation model can accurately output the matching target motion features based on the input target audio.
[0049] Target motion features can also be determined using a pre-defined audio feature and motion feature retrieval database, which stores the correspondence between various audio features and motion features. When determining target motion features, one can first extract the audio features of the target audio, such as pitch, duration, timbre, and intonation. Then, the database is used to search for motion features that match the audio features of the target audio. This method fully utilizes existing motion feature data to determine the target motion features corresponding to the target audio, is simple and efficient, and reduces the cost and time of model training.
[0050] In one implementation, step S120a may be included before step S130.
[0051] Step S120a: Extract the original motion features of the sound-producing object from the reference image.
[0052] The original motion features are used to represent the motion information of various parts of the sound-producing object in the reference image. These original motion features can include the degree of lip opening and closing, eye opening and closing, the angle of the corners of the mouth, the raised or lowered state of the eyebrows, the direction and angle of head rotation, the amplitude and direction of hand movements, and the twisting posture of the body. This motion state information provides a basic reference for subsequently generating target motion features that match the target audio.
[0053] Extracting the original motion features can be achieved through a pre-trained motion feature recognition model. This model can be trained on a large number of sample images containing different voice-producing objects in various motion states. During training, a set of labeled sample images can be used, where the motion features of each part of the voice-producing object in the sample images are labeled features. The sample images are input into the model to be trained, and the model parameters are adjusted and optimized based on the motion features output by the model and the labeled motion features of the sample images. This results in the motion feature recognition model, which can accurately extract the original motion features of the voice-producing object from the input reference images.
[0054] In one specific embodiment, the original motion features of the sound-emitting object in the reference image can be extracted through the following steps S121 to S122.
[0055] Step S121: Determine the motion coefficients of the sound-producing object for each basic motion unit from the reference image. The motion coefficients are used to represent the weight of the sound-producing object on the corresponding motion unit.
[0056] Basic motor units represent fundamental, atomic-level movement patterns used to construct various motor features (including at least one of facial expressions, lip movements, body movements, and trunk movements). For example, basic motor units may include: slightly upturned corners of the mouth, raised left upper lip by 1 mm, lowered right eyebrow tip, slight forward movement of the chin, slight leftward tilt of the head, naturally hanging arms, and slightly bent fingers, among other basic motor units.
[0057] Motion coefficients represent the weight of a vocal subject in a corresponding motion unit, that is, the degree of the vocal subject's performance in that basic motion unit. For example, if the vocal subject's mouth is slightly upturned in the reference image, the motion coefficient corresponding to the basic motion unit "slightly upturned mouth" will be relatively large, for example, 0.8; if the vocal subject's arm is naturally hanging down in the reference image without any other obvious movements, the motion coefficient corresponding to the basic motion unit "naturally hanging arm" is 1, while the motion coefficients corresponding to other arm-related basic motion units are 0; if the vocal subject has no expression related to "right eyebrow down" in the reference image, the motion coefficient corresponding to the basic motion unit "right eyebrow down" is 0; if the vocal subject's chin moves slightly forward in the reference image, the motion coefficient corresponding to the basic motion unit "slight forward chin movement" will be greater than 0, for example, 0.5. By determining the motion coefficients corresponding to each basic motion unit, the motion state of the vocal subject in the reference image can be quantified.
[0058] The motion coefficients corresponding to each basic motion unit of a sound-producing object differ under different motion states. These coefficients can be determined by performing feature analysis and quantification on a reference image. For example, using image processing techniques, keypoint detection can be performed on the face and body parts of the sound-producing object in the reference image. Based on the position and changes of these keypoints, the motion coefficients corresponding to each basic motion unit can be determined. By determining these motion coefficients, the motion state of the sound-producing object in the reference image can be represented quantitatively, facilitating subsequent processing and analysis of motion features.
[0059] In one specific embodiment, a pre-trained motion coefficient determination model can also be used to determine the motion coefficients of the sound-producing object for each basic motion unit from a reference image. This motion coefficient determination model can be trained on a large number of sample images containing different sound-producing objects in various motion states. During training, a set of sample images labeled with the motion coefficients corresponding to each basic motion unit is used. The sample images are input into the model to be trained, and the model parameters are adjusted and optimized based on the motion coefficients output by the model and the motion coefficients labeled in the sample images, resulting in a motion coefficient determination model capable of accurately determining the motion coefficients of the sound-producing object for each basic motion unit from the input reference image. Alternatively, the motion coefficient determination model can be trained in other ways, such as through the model training method described in the second embodiment below.
[0060] Optionally, when determining the motion coefficients of the sound-producing object for each basic motion unit from the reference image using a pre-trained motion coefficient determination model, the reference image can be encoded using a pre-trained image coding model to obtain image coding features, and then the motion coefficients of the sound-producing object in each motion dimension can be determined from the image coding features using the pre-trained motion coefficient determination model.
[0061] Encoding a reference image converts it into a feature representation in the latent space, reducing redundancy and noise in the image data and improving the accuracy of motion coefficient determination. Image encoding features can be image encoding vectors, image encoding matrices, etc. After obtaining the image encoding features, the motion coefficient determination model can more accurately analyze the motion coefficients of the sound-generating object on each basic motion unit based on these features.
[0062] Image coding models can be trained using deep learning models such as autoencoders. An autoencoder consists of an encoder and a decoder. The encoder compresses the input reference image into low-dimensional image coding features, while the decoder attempts to reconstruct the original reference image from these features. During training, the autoencoder parameters can be optimized by minimizing the reconstruction error, enabling the encoder to encode effective feature representations of the reference image. Specifically, during training, sample images containing different sound-producing objects can be acquired as training data and input into the autoencoder. The encoder encodes the sample images to obtain image coding features, and the decoder decodes and reconstructs the image based on these features. The error between the reconstructed image and the original sample image is calculated, for example, using mean squared error (MSE) as the loss function. The autoencoder parameters are continuously adjusted using the backpropagation algorithm, gradually reducing the value of the loss function. After multiple iterations of training, the trained image coding model is obtained. Alternatively, the image coding model can be trained using the training method described in the second embodiment below.
[0063] In one specific embodiment, the identification features of the sound-producing object in the reference image are extracted by determining the identification features of the sound-producing object from the image encoding features.
[0064] Image coding features contain rich feature information about the speaker. By filtering and processing this feature information, the identification features of the speaker can be accurately determined. For example, a recognition feature extraction model can extract information related to facial contours, facial features, skin color, and other facial shape features of the speaker, as well as information related to body proportions, clothing, limb shape, hand shape, and other body shape features from the image coding features.
[0065] The training methods for some of the models used in the first embodiment will be explained in the second embodiment. The first embodiment mainly provides a brief description of the model training process.
[0066] The motion coefficients corresponding to each basic motion unit of the sound-emitting object described in step S121 can be represented in the form of a coefficient matrix. The coefficient matrix clearly presents the quantitative information of the motion state of the sound-emitting object in the reference image, facilitating subsequent calculations and processing.
[0067] Step S122: Based on a pre-trained motion feature library, determine the motion features corresponding to the motion coefficients as the original motion features of the sound-emitting object in the reference image, wherein the motion feature library includes various basic motion units.
[0068] A motion feature library, also known as a motion dictionary, is a pre-built database that stores various basic motion units. The various basic motion units in the motion feature library can be combined to form various motion features of the vocal object.
[0069] The motion feature library can be in matrix form, namely a feature library matrix or a motion dictionary matrix. The feature library matrix can include a set of orthogonal basis vectors, where each basis vector represents a basic motion unit.
[0070] In step S122, the weights of the corresponding basic motion units in the motion feature library can be determined based on the motion coefficients mentioned above, and the motion features corresponding to the motion coefficients can be determined based on the weights of each basic motion unit. For example, as shown in (1) below, the coefficient matrix composed of each motion coefficient can be multiplied with the feature library matrix to obtain the original motion features of the sound-producing object. In this case, the original motion features of the sound-producing object can be understood as being composed of a linear combination of a set of orthogonal basis vectors. Since the motion coefficients in the coefficient matrix represent the weights of the sound-producing object on each basic motion unit, and the basis vectors in the feature library matrix represent the basic motion units, the multiplication of the two can comprehensively reflect the actual motion state of the sound-producing object in the reference image.
[0071] (1) In formula (1), Let M represent the original motion features, and M represent the feature library matrix. The feature library matrix includes N basis vectors, and the N basis vectors correspond to N basic motion features. Let m represent the coefficient matrix. i Let i represent the i-th basis vector. This represents the motion coefficient corresponding to the i-th basic motion unit.
[0072] This method of determining the original motion features by multiplying the coefficient matrix with the feature library matrix can efficiently and accurately quantify and represent the motion state of the sound-producing object in the reference image. Furthermore, the matrix operation method facilitates rapid calculation and processing in a computer system, providing an important data foundation for generating target motion features that match the target audio in subsequent steps. At the same time, the construction and storage method of the motion feature library also gives the entire system good scalability and flexibility; when new basic motion units need to be added, the feature library matrix can be easily updated and maintained.
[0073] Motion feature libraries can be constructed by acquiring video data from different sound-producing objects. The acquired video data should cover various scenes and different motion states to ensure the comprehensiveness and diversity of the motion feature library. For the collected video data, frame-by-frame analysis can be performed to identify the basic motion units of the sound-producing object in each frame and to label these basic motion units. After labeling, the motion feature library is formed based on these labeled basic motion units. During storage, a matrix format can be used, representing each basic motion unit as a basis vector and storing it in the feature library matrix.
[0074] Alternatively, the motion database can be trained using the method described in the second embodiment below.
[0075] This embodiment determines the motion coefficients of the sound-producing object for each basic motion unit and, based on a pre-trained motion feature library, can accurately and efficiently extract the original motion features of the sound-producing object from a reference image.
[0076] After extracting the original motion features of the sound-generating object from the reference image, step S130 above can proceed to step S131 to determine the target motion features that match the target audio.
[0077] Step S131: Adjust the original motion features according to the target audio to obtain the target motion features of the sound-producing object that match the target audio.
[0078] Adjusting the original motion features can involve modifying their weights, adding, or deleting certain features to match the pronunciation, rhythm, emotion, and semantics of the target audio. For example, when the target audio has a cheerful rhythm, the weights of motion features such as smiling and light body movements can be increased, while motion features associated with sadness, such as frowning and head-bowing, can be decreased or deleted. When the target audio expresses anger, the weights of motion features such as frowning, glaring, and leaning forward can be increased. The specific adjustment method can be determined based on actual needs and a large amount of experimental data.
[0079] After extracting the original motion features of the emitting object from the reference image, these features can be adjusted by combining them with the speech information of the target audio. For example, if the target audio frame is pronounced "ah," and the lips of the emitting object in the original motion features of the reference image are in a closed state, the degree of lip opening and closing in the original motion features can be adjusted to match the open state of the emitting "ah" based on the lip movement information corresponding to the pronunciation of "ah."
[0080] Specifically, the original motion features can be adjusted based on the pre-defined correspondence between different audio features (such as phonemes, pitch, intonation, timbre, etc.) and motion features. For example, the phoneme "a" can be set to correspond to a larger opening of the lips, and the phoneme "i" to correspond to a smaller opening of the lips and the corners of the mouth stretching to the sides. This can be represented by motion parameters for each part, such as the lip opening amplitude parameter and the corners of the mouth stretching amplitude parameter. When the audio frame of the target audio is the phoneme "a", the degree of lip opening and closing in the original motion features is adjusted to a larger opening. When the audio frame of the target audio is the phoneme "i", the degree of lip opening and closing in the original motion features is adjusted to a smaller opening and the corners of the mouth stretching. When the pitch of the audio frame of the target audio is high, the degree of eyebrow raising and eye opening in the original motion features can be increased to reflect the facial changes during high-pitched sounds.
[0081] Alternatively, a pre-trained motion feature generation model can be used, combined with the aforementioned original motion features and target audio, to output target motion features. Specifically, when training the motion feature generation model, in addition to inputting sample audio and corresponding sample video, the original motion features of the sound-producing object in the sample image can also be used as input information. In this way, when learning the mapping relationship between audio and the motion features of various parts of the sound-producing object, the motion feature generation model can consider the initial motion state of the sound-producing object in the reference image. When determining the target motion features, the target audio and the original motion features extracted from the reference image are input together into the trained motion feature generation model. The model can then integrate the information from both and output target motion features that match the target audio and are adjusted based on the original motion features.
[0082] This implementation extracts the original motion features of the sound-producing object from a reference image and adjusts them in conjunction with the target audio to generate target motion features that highly match the target audio. This approach fully considers the initial motion state of the sound-producing object in the reference image, making the generated target motion features more natural, reasonable, and consistent with motion behavior in real-world scenarios. Because the original motion features already reveal most motion characteristics, target motion features can be generated efficiently and quickly. This method also reduces unnecessary feature generation, improving video generation efficiency.
[0083] In one specific embodiment, step S131 can be performed by following steps S131a to S131c to obtain the target motion characteristics of the sound-producing object that match the target audio.
[0084] Step S131a: Obtain the target emotion, which is used to indicate the emotion of the speaker when emitting the target audio.
[0085] Optionally, the euphony corresponding to the target audio can be used as the target emotion, or the target emotion can be determined for each part of the target audio based on its content. Different parts of the target audio may correspond to different emotions; therefore, the target audio can be segmented, and the target emotion for each segment can be determined. For example, when the target audio is a dialogue containing transitions between sadness and joy, the sadness and joy parts can be identified, and their corresponding target emotions can be determined separately. Sentiment analysis techniques can be used to extract and analyze the speech features of the target audio (such as intonation, speech rate, and volume), combined with semantic understanding, to determine the emotion expressed in each audio segment. Alternatively, a pre-trained emotion recognition model can be used, where the target audio is input into the model, and the model outputs the target emotion corresponding to each part.
[0086] Optionally, the text content corresponding to the target audio can be determined, and the corresponding emotion can be identified from the text content as the target emotion. Specifically, natural language processing technology can be used to perform semantic understanding and sentiment analysis on the text content corresponding to the target audio to identify the emotional tendency of the text corresponding to the target audio, such as positive, negative, neutral, happy, excited, surprised, angry, joyful, sad, and furious. Alternatively, the corresponding emotion can be determined as the target emotion based on the acoustic features of the target audio. Acoustic features can include pitch, volume, speech rate, and intonation. For example, high pitch, fast speech rate, and large volume usually indicate excitement or anger, while low pitch, slow speech rate, and small volume may indicate sadness or calmness.
[0087] Optionally, in some application scenarios, users may have a specific need to have the voice emit a target audio with a specific emotion. In this case, the user can input or select the corresponding target emotion in the system, and the electronic device can use the emotion input by the user as the target emotion.
[0088] Once the target emotion is identified, different emotions often correspond to different body postures, facial expressions, and body movements of the speaker. For example, happiness may correspond to a raised corner of the mouth, bright eyes, and a slight swaying of the body; sadness may correspond to a lowered head, a furrowed brow, and a curled-up body.
[0089] Step S131b: Based on the above original motion features, determine the target emotion-related motion features associated with the above target emotion.
[0090] Target emotion-related motion features refer to motion features that reflect the target emotion and match the original motion features of the vocal subject. Target emotion-related motion features can be adjusted based on the vocal subject's original motion features, or supplemented by motion features associated with the target emotion. First, feature information related to body posture, facial expressions and appearance, and limb movements can be extracted from the original motion features. Then, based on the characteristics of the target emotion, this feature information can be adjusted and expanded.
[0091] For example, if the target emotion is anger, the vocalization subject in the original motor characteristics may be in a normal standing position. Then, the motor characteristics associated with anger may be leaning forward, clenching fists, and frowning.
[0092] Since different emotions manifest specific movement characteristics in various parts of the speaker's body, the target emotion-related movement characteristics can be determined based on the movement characteristics corresponding to each emotion. For example, an emotion-movement feature mapping table can be constructed, recording the body posture, facial expressions, and limb movements corresponding to various emotions. The original movement characteristics are compared with the movement characteristics corresponding to the target emotion in the mapping table. Features in the original movement characteristics that do not match the target emotion are adjusted or supplemented, while features in the original movement characteristics that match the target emotion are retained, resulting in the target emotion-related movement characteristics. For example, if the speaker's gaze is calm in the original movement characteristics, but the target emotion is anger, and the gaze characteristic corresponding to anger in the mapping table is wide-open glare, then the gaze characteristic in the original movement characteristics is adjusted to wide-open glare. If the speaker's feet are naturally standing in the original movement characteristics, and the feet characteristic corresponding to anger in the mapping table is also standing with feet apart, then the feet characteristic in the original movement characteristics is retained. In this way, the target emotion-related movement characteristics can be accurately determined.
[0093] Alternatively, the above-mentioned target emotion-related motor characteristics can be determined through the following steps A to B.
[0094] Step A: Obtain a pre-trained target projection model corresponding to the target emotion. The target projection model is used to map motion features to the corresponding motion features in the target emotion subspace.
[0095] Training the target projection model can be accomplished by collecting a large amount of motion feature data of voice-speaking objects under different emotions. For example, video data containing multiple emotions can be collected, and the voice-speaking objects in the videos can be analyzed frame by frame to extract the motion features of each frame. Then, these motion features are classified according to emotion categories, such as classifying the motion features of different emotions like anger, happiness, and sadness. For the target emotion, video frames and corresponding motion features under that emotion are selected as training samples. During training, the model learns how to map any motion feature to the corresponding motion feature in the target emotion subspace, so that the input motion features, after being processed by the model, can reflect the features corresponding to the target emotion. Other methods can also be used to train the target projection model; other specific training methods can be found in Example 2.
[0096] The target projection model can take the form of a target projection matrix, a target projection vector, etc. When the target projection model is a target projection matrix, the matrix can project the original motion features onto the target emotion subspace to obtain a motion feature vector associated with the target emotion.
[0097] Step B: Project the original motion features onto the subspace of the target emotion according to the target projection model to obtain the target emotion-related motion features associated with the target emotion.
[0098] Specifically, the original motion features can be input into the target projection model to obtain target emotion-related motion features associated with the target emotion. Alternatively, the original motion features can be represented as vectors or matrices and then multiplied with the target projection matrix to obtain target emotion-related motion features projected onto the target emotion subspace. This projection method using a target projection model can more accurately convert the original motion features into motion features that conform to the target emotion, fully considering the complex mapping relationship between emotion and motion features, and improving the accuracy of the determined target emotion-related motion features.
[0099] In one specific embodiment, step B can be performed according to steps a to e to obtain the target emotion-related motion features associated with the target emotion.
[0100] Step a: Determine the reference emotion of the voice-speaking object in the reference image.
[0101] The reference emotion of a speaker can be determined from aspects such as body posture, facial expression, and body movements in a reference image. For example, if the speaker's mouth is upturned and they are smiling in a reference image, their reference emotion can be judged as happiness; if the speaker is looking down and frowning, their reference emotion may be sadness. Image recognition technology can be used to analyze the reference image, identify the speaker's facial features, body posture, and other information, and then combine this with a pre-trained emotion recognition model to determine the speaker's reference emotion.
[0102] Step b: Obtain a pre-trained reference projection model corresponding to the reference emotion, the reference projection model being used to map the parent motion features to the corresponding motion features in the reference emotion subspace.
[0103] The training method and related content of the reference projection model are similar to those of the target projection model, and will not be repeated here.
[0104] Step c: Based on the reference projection model, filter out the motion features in the subspace of the reference emotion from the original motion features to obtain emotion-independent motion features.
[0105] Specifically, the original motion features can be input into a reference projection model, which will project the original motion features onto a reference emotion subspace to obtain the motion features in the reference emotion subspace. Then, the motion features in the reference emotion subspace are subtracted from the original motion features to obtain the emotion-independent motion features. The purpose of this step is to remove the part of the original motion features that is related to the reference emotion of the speaking object in the reference image, so as to more accurately add target emotion-related features later. Specifically, the emotion-independent motion features can be obtained through the following formula (2).
[0106] (2) In formula (2), w The above-mentioned emotion-independent motion features are represented by I, which represents the identity matrix. This represents the reference projection matrix corresponding to the reference projection model. This represents the original motion features. This formula can accurately filter out the parts of the original motion features that are related to the reference emotion.
[0107] Step d: Project the original motion features onto the subspace of the target emotion according to the target projection model to obtain the target emotion motion features.
[0108] Specifically, the original motion features can be input into the target projection model. If the target projection model is a target projection matrix, the original motion features are represented as vectors or matrices and then multiplied with the target projection matrix to obtain the target emotion motion features projected into the target emotion subspace. The target emotion motion features can reflect the motion characteristics that the voice-over should have under the target emotion.
[0109] Step e: Combine the emotion-independent motion features with the target emotion motion features to obtain target emotion-related motion features associated with the target emotion.
[0110] Specifically, the emotion-independent motion features can be added to the target emotion motion features to obtain the target emotion-related motion features. For example, the target emotion-related motion features can be obtained using the following formula (3).
[0111] + (3) in, Indicates the target emotion-related motor characteristics, This indicates the above-mentioned emotion-independent motor characteristics. Represents the reference projection matrix. Indicates the original motion characteristics, This represents the target emotional motion characteristics.
[0112] This embodiment, by first filtering reference emotion-related features and then adding target emotion-related features, enables the generated target emotion-related motion features to retain the emotion-irrelevant parts of the original motion features while accurately reflecting the features corresponding to the target emotion. This further improves the matching degree between the generated motion features and the target emotion, thereby making the performance of the speaking object in the final video more consistent with the expected emotional setting, making the video more vivid and realistic.
[0113] Step S131c: Adjust the target emotion-related motion features according to the target audio to obtain the target motion features of the vocal object that match the target audio.
[0114] After obtaining the target emotion-related motor characteristics, it is necessary to consider various features of the target audio to make adjustments. These features include, but are not limited to, pronunciation, rhythm, and intensity. For example, different pronunciations may lead to differences in the speaker's mouth shape and facial muscle movements. If the target audio contains plosives, the speaker's lips may need to make rapid opening and closing movements; if it contains long vowels, the mouth shape needs to remain relatively stable.
[0115] Specifically, an audio feature-motion feature adjustment rule base can be constructed, which records the motion features of the vocal object corresponding to different audio features. For example, for the pronunciation of "ah," the rule base can specify that the vocal object's mouth should open wide and the tongue should retract. The features of the target audio are compared with the rules in the rule base, and the motion features related to the target emotion are adjusted according to the rules. Alternatively, a machine learning model can be used to adjust the motion features related to the target emotion. An audio-to-motion feature adjustment model is trained, whose inputs are the features of the target audio and the target emotion-related motion features, and whose output is the adjusted target motion features. The specific process of generating the target motion features can be found in step S131.
[0116] This implementation first determines the target emotion-related motion features, and then adjusts them based on the target audio. This fully combines the characteristics of the target audio with the expression of the target emotion, ensuring that the generated target motion features not only reflect the target emotion but also match the pronunciation, rhythm, intensity, and other characteristics of the target audio. In this way, when generating the video, the movement of the speaker can be synchronized with the audio and naturally display the appropriate emotion and actions.
[0117] In one implementation, step S131c can be performed by following step S131c-1 to obtain the target motion characteristics of the sound-producing object that match the target audio.
[0118] Step S131c-1: Based on the pre-trained motion feature generation model, and according to the target audio and the target emotion, adjust the motion features related to the target emotion to obtain the target motion features of the voice-producing object that match the target audio.
[0119] Specifically, the target audio, the target emotion, and the motion features related to the target emotion can be used to determine the input information for the motion feature generation model. This input information is then fed into the motion feature generation model to obtain the target motion features of the vocal object that match the target audio and correspond to the target emotion. The motion feature generation model is used to generate target motion features of the vocal object that match the target audio and target emotion based on the target audio and target emotion, and on the motion features related to the target emotion. The motion feature generation model has been trained on a large amount of data, learning the complex relationships between audio features, emotions, and motion features. Upon receiving the target audio, target emotion, and target emotion-related motion features, the model will make targeted adjustments to the target emotion-related motion features based on the patterns and rules it has learned to obtain the target motion features. A well-trained motion feature generation model can efficiently and accurately generate target motion features that match the target audio and target emotion, reducing the workload of manual intervention and adjustment, and improving generation efficiency and accuracy.
[0120] Optionally, the input information of the motion feature generation model may also include the motion features of the previous target video frame. In this way, the model can take into account the coherence and continuity of the motion. The motion information of the previous target video frame can provide context for the generation of motion features of the current frame, making the generated target motion features more coordinated with the motion state of the previous frame and avoiding abrupt and discontinuous actions.
[0121] The specific training process of the motion feature generation model will be described in detail in the second embodiment.
[0122] In one specific embodiment, the motion feature generation model obtains the target motion features of the sound-producing object in the target video frame by performing multi-step denoising on the noisy motion features after adding noise. In other words, the motion feature generation model is used to gradually remove noise according to a preset multi-step denoising strategy, continuously approximating the target motion features that the sound-producing object in the target video frame should have. At each denoising step, the motion feature generation model adjusts the noisy motion features based on the relationships between previously learned audio features, emotions, and motion features, combined with information about the target audio and target emotion. As the denoising steps progress, the motion features gradually transform from a noisy initial state to target motion features that conform to the target audio and target emotion. This multi-step denoising process can be seen as a continuous optimization and refinement process, with each step capturing the motion feature information corresponding to the target audio and target emotion more accurately.
[0123] Step S131c-1 can be implemented in the following steps: In the multi-step denoising process of the motion feature generation model, according to the condition injection order of the target emotion-related motion features, the target audio, and the target emotion, denoising reference conditions are injected into the motion feature generation model at different denoising stages, so that the motion feature generation model performs denoising in each denoising stage in sequence with the target emotion-related motion features, the target audio, and the target emotion as conditions, to obtain the target motion features of the voice object that match the target audio.
[0124] A denoising stage can include a single denoising step or multiple denoising steps. For example, in the first denoising stage of the motion feature generation model, target emotion-related motion features are injected into the model as denoising reference conditions. At this point, the model will initially adjust the noisy motion features based on these conditions, causing the motion features to shift towards those related to the target emotion. Since target emotion-related motion features can reflect the basic motion characteristics of the target emotion, in this stage, the motion features will initially possess some expression of the target emotion, allowing the model to focus on target emotion-related motion features and laying the foundation for subsequent adjustments. In the second denoising stage, the target audio is injected into the motion feature generation model as a denoising reference condition. Since the first stage has already established a certain correlation between the motion features and the target emotion-related motion features, in this stage, the model will further adjust the motion features by combining features such as pronunciation, rhythm, and intensity of the target audio, allowing the motion features to better match the audio features. In the third denoising stage, the target emotion is injected as a denoising reference condition. At this time, the model will combine the adjustment results of the previous two steps and refine the motion features again according to the target emotion to ensure that the motion features are not only synchronized in audio, but also more accurate and natural in emotional expression. For example, if the target emotion is excitement, the model will enhance the liveliness and intensity in the motion features, so that the actions of the speaker can better reflect the excited emotion.
[0125] By injecting denoising reference conditions in stages, the motion feature generation model can progressively and accurately denoise and adjust the noisy motion features, ultimately obtaining the target motion features of the sounding object that match the target audio. Furthermore, following the order of target emotion-related motion features, the target audio, and the target emotion condition, the model first uses target emotion-related motion features as denoising conditions to obtain motion features that initially possess the target emotion. Based on this, the target audio condition, which is weakly correlated with motion, is then injected, followed by the target emotion condition, which is strongly correlated with motion. This gradual adjustment method, injecting denoising conditions from weak to strong, helps improve the model's stability and adjustment effect, avoiding over-adjustment or under-adjustment during the denoising process. This results in more accurate determined motion features and more accurate and natural target video frames.
[0126] Optionally, such as Figure 5As shown, the motion feature generation model can include multiple model parts arranged sequentially. Each model part includes a condition adjustment layer, a denoising layer, and a gate control layer. The condition adjustment layer adjusts the injected denoising reference conditions to adapt them to the model's internal feature representation. For example, it can scale or translate the denoising reference conditions to ensure their dimensions and numerical range match the model's internal features. The denoising layer denoises the data input from the previous layer based on the injected denoising reference conditions. The gate control layer controls the output of the denoising layer. Depending on the denoising reference conditions, the gate control layer can dynamically adjust the influence of the denoising layer's output. For example, when injecting target audio as the denoising reference condition, the gate control layer can control the adjustment magnitude of the motion features based on the importance of the audio features. The gate control layer can generate a control gate through a gating mechanism to determine which feature information can pass through and which needs to be suppressed. This effectively filters out unnecessary noise and interference information, allowing the model to focus more on features related to the denoising reference conditions.
[0127] When the input information of the motion feature generation model can also include the motion features of the previous target video frame, the motion features of the previous target video frame can be fused with the target audio, the target emotion, and the target emotion-related motion features by a pre-trained multi-head fusion model. Then, the target emotion-related motion features, the target audio, and the target emotion fused with the motion features of the previous target video frame are used as the input information of the motion feature generation model until denoising is achieved.
[0128] Step S140: Based on the above-mentioned identification features and target motion features, obtain the overall features of the sound-emitting object in the target video frame, and obtain the target video frame based on the overall features, so as to generate a corresponding video based on each target video frame and the target audio.
[0129] After obtaining the identification features and target motion features, the two are combined to obtain the overall features of the sound-producing object. In other words, the overall features of the sound-producing object in the target video frame can be obtained by following the steps S141.
[0130] Step S141: Combine the recognition features and the target motion features to obtain the overall features of the sound-emitting object in the target video frame.
[0131] Recognition features may include static information such as the speaker's shape and appearance, while target motion features reflect its dynamic performance under the influence of target audio and target emotion. Feature concatenation can be used to combine recognition features and target motion features along the feature dimension, forming a feature vector containing more information. For example, if the recognition feature is a vector of length m and the target motion feature is a vector of length n, the overall feature vector obtained after concatenation will have a length of m + n.
[0132] Alternatively, a weighted fusion method can be used to obtain the overall features. Different weights are assigned to the recognition features and the target motion features, and then the two are linearly combined based on these weights. The weights can be adjusted according to the specific application scenario. For example, in video generation that emphasizes emotional expression, a higher weight can be assigned to the target motion features; while in scenarios that focus on identity recognition, the weight of the recognition features can be increased. After obtaining the overall features of the speaking object, the target video frame is generated based on these features. An image generation model can be used, taking the overall features as input and mapping the feature information to the image space, thereby generating a target video frame containing the speaking object.
[0133] In step S140, the target video frame can be obtained by following step S142.
[0134] Step S142: Decode the overall features using a pre-trained decoding model to obtain the target video frame.
[0135] During training, the decoding model learns the mapping relationship between overall features and image pixels, enabling it to transform abstract overall features into concrete image representations. The training process of the decoding model will be described in the second embodiment.
[0136] This implementation method can efficiently and accurately decode the overall features into target video frames through the decoding model, ensuring the quality and efficiency of video generation.
[0137] After determining the target video frames corresponding to the target audio, the electronic device can combine the target audio with each target video frame to generate a video in which the target audio is emitted by the aforementioned sound-emitting object. Specifically, the electronic device can align each target video frame with its corresponding target audio frame according to the target video frame corresponding to each audio frame of the target audio to obtain the corresponding video, ensuring that the actions of the sound-emitting object in the video match the pronunciation, rhythm, and other features of the audio.
[0138] The solution provided in this application determines the target motion features of various parts of the sound-producing object that match the target audio. These motion features effectively reflect the motion state of each part of the sound-producing object within the target video frame. Combined with the recognition features of the sound-producing object extracted from the reference image, the overall features of the sound-producing object in the target video frame can be determined. This avoids generating the overall features of the sound-producing object for each frame; instead, it processes the appearance and motion information of the sound-producing object separately. The recognition features representing the appearance only need to be determined once from the reference image, and each frame only needs to generate the motion features of each part of the sound-producing object. This significantly improves the speed and efficiency of video generation. For practical applications such as live video streaming and video production, it can meet the needs of real-time playback and rapid production, providing users with a smooth viewing experience.
[0139] The following example illustrates the process of the video generation method provided in this application.
[0140] like Figure 3 As shown, the video generation method in this example includes the following steps 10-19.
[0141] Step S10: Input the reference image into the image coding model to obtain the image coding features.
[0142] Step S11: Input the image encoding features into the motion coefficient determination model to obtain the coefficient matrix formed by the motion coefficients of the vocal object for each basic motion unit. .
[0143] Step S12: Multiply the motion coefficients corresponding to each basic motion unit of the sound-producing object with the motion dictionary matrix to obtain the original motion features of the sound-producing object.
[0144] Step S13: Based on the target projection matrix and the reference projection matrix, determine the target emotion-related motion features that are projected onto the target emotion subspace according to the original motion features.
[0145] Step S14: Fuse the target emotion-related motion features, target audio, and target emotion with the target motion features corresponding to the previous frame using a multi-head fusion model. w t-1 After fusion, the feature input is used to generate a motion feature model to obtain the target motion features of the sound-producing object that match the target audio. w t .
[0146] Step S15: Extract the recognition features of the sound-producing object from the image coding features.
[0147] Step S16: Combine the identified features and target motion featuresw t By combining these elements, we obtain the overall characteristics.
[0148] Step S17: Input the overall features into the decoding model to obtain the target video frame.
[0149] For details on this example, please refer to the detailed explanation and analysis of each step above. This section mainly describes the data processing flow, and further details will not be repeated here.
[0150] Example 2 The second embodiment of this application also provides a training method for a video generation model. This method is applied to an electronic device, which may be a server, desktop computer, laptop computer, mobile phone, tablet computer, smartwatch, smart TV, VR device, in-vehicle device, wearable device, or other electronic device with data processing capabilities. The video generation model includes an audio alignment model and may also include other models.
[0151] like Figure 4 , Figure 5 As shown, the training method includes the following steps S210 to S240.
[0152] Step S210: Obtain a first training sample, which includes a sample video, a sample audio that matches the sample video, and a sample reference image. The sample video and the sample reference image include the same first sample speaker.
[0153] The first training sample can be a pre-stored sample, a sample downloaded from the network, or a sample received from other devices, etc. This application does not specifically limit this.
[0154] Step S220: Extract the first sample motion feature of the first sample sound-emitting object from the sample video frame of the sample video, and extract the second sample motion feature of the first sample sound-emitting object from the sample reference image.
[0155] The extraction methods for the motion features of the first sample and the motion features of the second sample can refer to the extraction method of the original motion features in the first embodiment, and will not be described in detail here.
[0156] Step S230: Determine the first sample input information based on the second sample motion features and the sample audio, and input the first sample input information into the motion feature generation model to be trained to obtain the first output motion features.
[0157] Specifically, the motion features of the second sample and the audio of the sample can be determined as the input information of the first sample. Alternatively, other training-related information can be added to the input information of the first sample, such as sample emotion information and sample noise features. Or, the motion features of the second sample and the audio of the sample can be preprocessed (e.g., normalized, reduced in dimensionality) before being determined as the input information of the first sample.
[0158] Step S240: Based on the difference between the first output motion feature and the first sample motion feature, adjust the parameters of the motion feature generation model to be trained to obtain the trained motion feature generation model.
[0159] Specifically, a first loss function can be used to quantify the difference between the first output motion feature and the first sample motion feature. Examples of first loss functions include the mean squared error loss function and the cross-entropy loss function. By calculating the value of the first loss function, optimization algorithms (such as stochastic gradient descent) are used to adjust the parameters of the motion feature generation model to be trained, gradually decreasing the value of the loss function, thus allowing the first output motion feature to continuously approximate the first sample motion feature. Through multiple iterations of training, the model parameters are continuously adjusted, ultimately resulting in a trained motion feature generation model that can more accurately generate outputs that match the actual motion features based on the input information.
[0160] The motion feature generation model trained in this embodiment can generate motion features corresponding to the target audio, providing a more accurate motion feature foundation for subsequent video generation. Furthermore, because it generates features along the motion feature dimension, the generation speed is faster and more efficient, meeting the real-time requirements of video generation.
[0161] In one implementation, the first training sample further includes sample emotions corresponding to each frame of the sample video. Step S230 can determine the first sample input information according to the following step S231.
[0162] Step S231: Determine the first sample input information based on the second sample motion features, the sample audio, and the sample emotion.
[0163] Specifically, such as Figure 5 As shown, the motion features of the second sample, the audio features corresponding to the sample audio, and the emotion features corresponding to the sample emotion can be determined as the first sample input information. Alternatively, other information can be added to determine the first sample input information. In this embodiment, adding sample emotion to the first sample input information allows the model to consider the influence of emotional factors on motion features during training, enabling the trained motion feature generation model to generate motion features that are more realistic and expressive of emotions.
[0164] In this embodiment, sample audio features and sample emotion features can be extracted from the sample audio and the sample emotion, and the first sample input information can be determined based on the second sample motion features, the sample audio features, and the sample emotion features.
[0165] After determining the initial sample input information, it is fed into the motion feature generation model to be trained. The model adjusts and generates motion features more accurately based on the input information including the sample's emotion. During training, the model learns the changing patterns of motion features under different emotions, enabling it to generate corresponding target motion features based on the target emotion in practical applications. Once training is complete, the trained motion feature generation model can not only generate matching motion features based on the target audio but also further refine and adjust the motion features according to the target emotion, making the movement of the speaking object in the generated target video frame more natural and realistic, with stronger emotional impact. In scenarios such as live video streaming and video production, this can provide viewers with a better viewing experience and meet users' needs for video content with rich emotional expression. In one implementation, such as Figure 5 As shown, the motion feature generation model to be trained can obtain output motion features by performing multi-step denoising on the noisy motion features of the sample after adding noise. In step S230, the first output motion feature can be obtained by following step S232.
[0166] Step S232: In the multi-step denoising process of the motion feature generation model to be trained, sample denoising reference conditions are injected into the motion feature generation model to be trained in different denoising stages according to the sample condition injection order of the second sample motion feature, the sample audio, and the sample emotion, to obtain the first output motion feature.
[0167] The execution process of step S232 can refer to the process of injecting denoising reference conditions into the motion feature generation model at different denoising stages in step S131c-1 of the first embodiment, which will not be described in detail here. In this embodiment, sample denoising reference conditions are injected into the motion feature generation model to be trained at different denoising step stages according to the order of sample condition injection. This allows the model to gradually learn the influence of different conditions on motion features during the denoising process. Injection begins with conditions directly related to motion, such as the second sample motion feature, allowing the model to gain a preliminary understanding of basic motion features. Next, sample audio, a condition weakly related to motion, is injected, enabling the model to integrate audio information into the adjustment of motion features. Finally, sample emotion, a condition strongly related to motion, is injected to further refine the motion features and make them more consistent with emotional expression. This staged injection method helps the model gradually adapt to the influence of different conditions, avoiding the difficulty of model learning or instability caused by injecting too many complex conditions at once. In each denoising step stage, the model adjusts the noisy sample motion features after denoising according to the injected sample denoising reference conditions, gradually removing noise and making the output motion features closer to the real motion features. Through this multi-step denoising and phased condition injection training method, the trained motion feature generation model can more comprehensively and deeply learn the influence of various factors on motion features. The resulting trained model can generate more accurate, natural motion features with rich emotional expression in practical applications, providing strong support for high-quality video generation. In fields such as live video streaming and video production, videos generated using this model can better attract viewers' attention and enhance their viewing experience.
[0168] In one specific embodiment, such as Figure 5 As shown, the motion feature generation model to be trained may include multiple model parts arranged in sequence. Each model part includes a condition adjustment layer, a denoising layer, and a gate control layer. The condition adjustment layer is used to adjust the injected denoising reference conditions. The denoising layer is used to denoise the data input from the previous layer based on the injected denoising reference conditions. The gate control layer is used to control the output of the denoising layer. The model structure of the motion feature generation model to be trained is similar to that of the motion feature generation model in the first embodiment, and will not be described in detail here. In one specific embodiment, such as Figure 5 As shown, step S231 can determine the first sample input information according to the following steps S231a.
[0169] Step S231a: Using the multi-head fusion model to be trained, the motion features of the second sample, the audio of the sample, and the emotion of the sample are fused with the first output motion features corresponding to the previous sample video frame, and the corresponding sample fusion features are determined as the first sample input information.
[0170] Correspondingly, the training method for the above model may also include the following step S250.
[0171] Step S250: Adjust the parameters of the multi-head fusion model to be trained based on the difference between the first output motion feature corresponding to the sample video frame and the first sample motion feature to obtain the trained multi-head fusion model.
[0172] In this implementation, the video generation model may further include a multi-head fusion model. In this embodiment, the multi-head fusion model is trained simultaneously with the motion feature generation model, enabling the two models to cooperate and optimize collaboratively. The trained multi-head fusion model can more effectively fuse the second sample motion features, sample audio, and sample emotion with the first output motion features corresponding to the previous sample video frame, providing higher-quality input information for the motion feature generation model. This allows the motion feature generation model to generate outputs that better match the actual motion features based on more accurate input. The collaboratively trained models can better adapt to different video generation scenarios. This collaborative training method also enhances the model's generalization ability. During training, the model learns the complex relationships and mutual influences between different factors, enabling it to generate accurate and expressive motion features in more diverse scenarios, thereby improving the quality and efficiency of video generation.
[0173] In one embodiment, the video generation model may further include a motion feature library, which includes various basic motion units; the training method of the above model may further include the following steps S260 to S2100.
[0174] Step S260: Obtain a second training sample. The second training sample includes a first sample source image and a first sample target image. The first sample source image and the first sample target image contain the same second sample vocal object. The second sample vocal object in the first sample source image is in a first motion state, and the second sample vocal object in the sample target image is in a second motion state. The first motion state is different from the second motion state.
[0175] The second training sample is used to train and update the motion feature database. Using the first sample source image and the first sample target image in the second training sample, it is possible to obtain motion change information of the second sample sound-emitting object from the first motion state to the second motion state.
[0176] The second training sample can be a pre-stored sample, a sample downloaded from the network, or a sample received from other devices, etc., and this application does not specifically limit it. The second training sample can be a portion of the samples selected from the first training sample, or it can be a sample set completely independent of the first training sample.
[0177] Optionally, such as Figure 6 As shown, the second sample speaking object in the first sample source image can be in an initial motion state, where the speaking object is neither making a sound nor moving. In the first sample target image, the second sample speaking object can be in a motion state, where the speaking object is making a sound and displaying a specific facial expression. By comparing these two different states of the speaking object, the key feature changes of the speaking object during its motion can be clearly captured.
[0178] Step S270: Determine the first sample motion coefficients of the second sample sound-emitting object for each basic motion unit from the first sample target image, and determine the sample recognition features of the second sample sound-emitting object from the first sample source image.
[0179] The determination method and related content of the first sample motion coefficient and sample recognition features can be referred to the relevant content of motion coefficient and recognition features in the first embodiment, which will not be described in detail here.
[0180] Step S280: Determine the motion features of the third sample corresponding to the motion coefficients of the first sample using the motion feature library to be trained.
[0181] The method for determining the motion features of the third sample can refer to the method for determining the original motion features in the first embodiment, and will not be described in detail here.
[0182] Step S290: Obtain the output image based on the sample recognition features and the motion features of the third sample.
[0183] The method for determining the output image can refer to the method for determining the target video frame in the first embodiment, and will not be described in detail here.
[0184] Step S2100: Adjust the parameters in the motion feature library to be trained based on the difference between the output image and the sample target image to obtain the trained motion feature library.
[0185] Specifically, a second loss function can be used to quantify the difference between the output image and the target image. By calculating the value of the second loss function, optimization methods such as gradient descent are used to adjust the parameters in the motion feature library, gradually reducing the value of the loss function, thus making the output image continuously approximate the target image. During multiple iterations of training, the parameters of the motion feature library are continuously adjusted, ultimately resulting in a trained motion feature library. This library can more accurately store and represent each basic motion unit, providing richer and more accurate motion feature references for subsequent video generation.
[0186] The trained motion feature library, in conjunction with the trained motion feature generation model and multi-head fusion model, can significantly improve the quality and effect of video generation.
[0187] Optionally, such as Figure 6 As shown, in step S270, the first sample target image can be input into the image encoding model to be trained to obtain the first sample target image encoding features. The first sample target image encoding features are then input into the motion coefficient determination model to be trained to obtain the first sample motion coefficients λ corresponding to each basic motion unit of the second sample vocal object. The first sample source image is input into the aforementioned image encoding model to be trained to obtain the first sample source image encoding features. The sample recognition features of the second sample vocal object are then extracted from the first sample source image encoding features. In this case, the training method of the above model may also include the following step S2110.
[0188] Step S2110: Based on the difference between the output image and the sample target image, adjust the parameters of the image encoding model to be trained and the motion coefficient determination model to be trained, so as to obtain the trained image encoding model and the trained motion coefficient determination model.
[0189] The trained image coding model can more accurately encode features from the input image, while the trained motion coefficient determination model can more precisely determine the motion coefficients of the sound-producing object for each basic motion unit. These two models, working in conjunction with the trained motion feature library, motion feature generation model, and multi-head fusion model, further improve the performance of the overall video generation model. During video generation, they can more accurately capture the motion features and changes of the sound-producing object, thus generating more natural, vivid, and realistic videos. In this case, the aforementioned video generation model can also include the image coding model and motion coefficient determination model.
[0190] In one implementation, such as Figure 6As shown, step S290 combines the sample recognition features with the third sample motion features to obtain sample combination features, and inputs the sample combination features into the decoding model to be trained to obtain the output image. In this case, the training method of the above model may also include the following step S2120.
[0191] Step S2120: Adjust the parameters of the decoding model to be trained according to the difference between the output image and the sample target image to obtain the trained decoding model.
[0192] By adjusting the parameters of the decoding model to be trained, it can better generate output images based on the combined features of sample recognition features and third-party sample motion features. The trained decoding model can more accurately transform the combined features into output images that are closer to the sample target image, further improving the accuracy and quality of image generation in video generation. In this case, the video generation model described above can also include the aforementioned decoding model.
[0193] When the trained image encoding model, motion coefficient determination model, motion feature library, motion feature generation model, multi-head fusion model, and decoding model work together, the entire video generation model will exhibit more powerful performance. During the video generation process, these models can accurately analyze and process the motion features and image features of the video from multiple dimensions. The entire video generation model can generate more natural, vivid, realistic, and expressive videos, meeting users' needs for high-quality videos in various fields such as live video streaming and video production, and bringing users a better video viewing experience. Optionally, step S220 may extract the first sample motion features of the first sample sound-emitting object according to the following steps S221~S222, and extract the second sample motion features of the first sample sound-emitting object according to the following steps S223~S224.
[0194] Step S221: Determine the second sample motion coefficients of the first sample sound-emitting object for each basic motion unit from the sample video frames of the sample video.
[0195] Step S222: Determine the motion features of the first sample corresponding to the motion coefficients of the second sample using a pre-trained motion feature library.
[0196] Step S223: Determine the motion coefficients of the third sample corresponding to each motion dimension of the first sample sound-emitting object from the sample reference image; Step S224: Determine the motion features of the second sample corresponding to the motion coefficients of the third sample using a pre-trained motion feature library.
[0197] In one embodiment, the video generation model may further include projection models corresponding to each emotion, wherein the projection models are used to map motion features to the subspace of the corresponding emotion.
[0198] like Figure 7 As shown, the training method of the model may further include the following steps S2130 to S2160.
[0199] Step S2130: Obtain a third training sample, the third training sample including a second sample source image and a second sample target image, the second sample source image and the second sample target image include the same third sample speaking object, and the emotions of the third sample speaking object in the second sample source image and the second sample target image are different.
[0200] The third training sample can be obtained by filtering from professional video material libraries, or by searching for relevant videos on various video websites and social media platforms online. Image recognition and classification algorithms can then be used to extract image pairs containing the same third-sample speaker but with different emotions from the search results. Alternatively, the third training sample can be obtained through manual recording. The third training sample can be the same as the second or first training sample, or it can be completely independent of the second training sample, or it can be a part of the second training sample; this application does not impose specific limitations in this regard.
[0201] Step S2140: Extract the source sample motion features of the third sample sound-emitting object from the second sample source image, and extract the target sample motion features of the third sample sound-emitting object from the second sample target image.
[0202] The extraction methods for source sample motion features and target sample motion features can refer to the extraction methods for the first sample motion features and the second sample motion features in step S220, and will not be described in detail here.
[0203] Step S2150: Project the motion features of the source samples onto the target emotion space using the target emotion projection model to be trained and the reference emotion projection model to be trained, to obtain the output emotion features.
[0204] like Figure 7 As shown, the target emotion projection model to be trained can be achieved through... e tgt It is stated that the reference emotion projection model to be trained can be obtained through... e src express.
[0205] The method for determining the output emotion features can refer to the method for determining the target emotion-related motion features in the first embodiment, and will not be described in detail here.
[0206] Step S2160: Based on the difference between the output emotion features and the target sample motion features, adjust the reference emotion projection model to be trained and the target emotion projection model to be trained to obtain the trained target emotion projection model.
[0207] Specifically, a third loss function can be used to measure the difference between the output emotion features and the motion features of the target sample. By calculating the value of the third loss function, optimization algorithms such as stochastic gradient descent are used to adjust the parameters of the projection model, continuously reducing the value of the loss function, thus making the output emotion features increasingly approximate the motion features of the target sample. During multiple iterations of training, the parameters of the reference emotion projection model and the target emotion projection model are continuously adjusted, ultimately resulting in the trained target emotion projection model. The trained target emotion projection model can more accurately map motion features to the corresponding emotion subspace, enabling the video generation model to more accurately represent emotion-related motion features when processing videos with different emotions. When this trained target emotion projection model works in conjunction with the trained image encoding model, motion coefficient determination model, motion feature library, motion feature generation model, multi-head fusion model, and decoding model, the entire video generation model will exhibit superior performance in emotion expression.
[0208] Alternatively, please refer to Figure 7 As shown in the process, step S260 can be implemented by following steps S261 to S262.
[0209] Step S261: Obtain the output emotion image corresponding to the output emotion feature based on the pre-trained decoding model.
[0210] Step S262: Identify the corresponding output emotion features from the output emotion image.
[0211] For example, a pre-trained emotion recognition model can be used to identify the corresponding output emotion features from the output emotion image.
[0212] Step S263: Based on the difference between the output emotion features and the emotion of the third sample in the second sample target image, and the difference between the output emotion features and the motion features of the target sample, adjust the reference emotion projection model to be trained and the target emotion projection model to be trained to obtain the trained target emotion projection model.
[0213] Outputting emotional features can be done through Specifically, a fourth loss function can be used to measure the difference between the output emotion features and the emotion of the third sample object in the second sample target image.
[0214] This embodiment simultaneously adjusts the reference emotion projection model and the target emotion projection model to be trained based on the differences between the output emotion features and the emotion of the third sample in the second sample target image, as well as the differences between the output emotion features and the motion features of the target sample. This allows for more comprehensive and accurate optimization of the projection model parameters. Because it comprehensively considers the differences between the emotion features and the emotion in the target image, as well as the differences with the target motion features, it avoids the limitations of a single difference measure. The trained target emotion projection model, adjusted in this way, can more accurately capture and map motion features under different emotions. This enables the video generation model to not only accurately present emotion-related motion features when generating videos, but also better match the emotion expressed by the target image, further improving the realism and accuracy of the video's emotional expression.
[0215] Optionally, step S231 can determine the first sample input information according to the following steps S231b~S231c, including: Step S231b: Project the motion features of the second sample onto the subspace corresponding to the sample emotion using a pre-trained sample emotion projection model to obtain the motion features related to the sample emotion.
[0216] The method for determining the emotion-related motion features of the sample in this step can refer to the method for determining the target emotion-related motion features in step B of the first embodiment, and will not be described in detail here.
[0217] Step S231c: Determine the first sample input information based on the sample emotional motion features, the sample audio, and the sample emotion.
[0218] This embodiment determines the first sample input information through steps S231b and S231c, providing input content that better matches the sample's emotion for the subsequent video generation process. This input information enables the video generation model to better combine factors such as audio, emotion, and motion features to generate videos that better reflect the expected emotional expression.
[0219] Example 3 The third embodiment of this application also provides a video live streaming method, which can be applied to electronic devices, such as servers, desktop computers, laptops, mobile phones, tablets, smartwatches, smart TVs, VR devices, in-vehicle devices, wearable devices, and other electronic devices with data processing functions. The video live streaming method provided in the second embodiment of this application includes the following steps S310 to S330.
[0220] Step S310: Obtain the target audio and reference image corresponding to the video to be live-streamed, wherein the reference image includes the broadcaster.
[0221] The target audio for the live video can be a sound signal captured in real time from the live broadcast site, or a pre-recorded audio file prepared for playback during the live broadcast. Reference images can serve as the basis for determining the anchor's appearance and other identifying features during subsequent video generation.
[0222] Step S320: Generate a live video using the video generation method described in any one of the first embodiments.
[0223] Step S330: Play the live video in real time via the client.
[0224] When the execution entity of the second embodiment is a server, the server can send the generated live video to a client, which can be either a live video viewer or a broadcaster, allowing the client to play the live video. When the execution entity of the second embodiment is a broadcaster, the broadcaster can directly play the generated live video and simultaneously push the video stream to the server, which then distributes the video to other live video viewers. This live video method combines the video generation method of the first embodiment, enabling the rapid generation of live videos with realism and expressiveness. During the live broadcast, reference images are used to determine the broadcaster's identification features, avoiding frequent processing of the broadcaster's appearance, improving video generation efficiency, and ensuring the real-time nature of the live broadcast.
[0225] Example 4 The fourth embodiment of this application also provides a video generation apparatus corresponding to the video generation method embodiment provided in the first embodiment. Since the apparatus embodiment is basically similar to the method embodiment, it is described simply. For details of the relevant technical features and their effects, please refer to the corresponding descriptions of the video generation method embodiments provided above. Figure 8 As shown, the video generation device provided in this embodiment includes: Information acquisition unit 410 is used to acquire target audio and reference image for generating video, wherein the reference image includes the sound-emitting object; Feature extraction unit 420 is used to extract the identification features of the sound-producing object from the reference image, wherein the identification features are used to represent the feature information of each part of the sound-producing object; The feature determination unit 430 is used to determine the target motion features that match the target audio, wherein the target motion features are used to represent the motion information of each part of the sound-producing object in each target video frame corresponding to the target audio; The video generation unit 440 is configured to obtain the overall features of the sound-producing object in the target video frame based on the recognition features and the target motion features, and to obtain the target video frame based on the overall features, so as to generate a corresponding video based on each target video frame and the target audio.
[0226] Example 5 The fifth embodiment of this application also provides a training apparatus for a video generation model corresponding to the training method embodiment for the video generation model provided in the second embodiment. Since the apparatus embodiment is basically similar to the method embodiment, it is described simply. For details of the relevant technical features and the effects achieved, please refer to the corresponding descriptions of the video generation method embodiments provided above. The training apparatus for the video generation model provided in this embodiment includes: The sample acquisition unit is used to acquire a first training sample, which includes a sample video, a sample audio that matches the sample video, and a sample reference image. The sample video and the sample reference image include the same first sample sound source. The sample feature extraction unit is used to extract the first sample motion feature of the first sample sound-producing object from the sample video frame of the sample video, and to extract the second sample motion feature of the first sample sound-producing object from the sample reference image. The output feature unit is used to determine the first sample input information based on the second sample motion features and the sample audio, and input the first sample input information into the motion feature generation model to be trained to obtain the first output motion features. The training unit is used to adjust the parameters of the motion feature generation model to be trained based on the difference between the first output motion feature and the first sample motion feature, so as to obtain the trained motion feature generation model.
[0227] Example 6 The sixth embodiment of this application also provides an electronic device embodiment corresponding to the video generation method provided in the first embodiment. The following description of the electronic device embodiment is merely illustrative. The electronic device embodiment is as follows: Please refer to Figure 9 Understanding the above electronic devices, Figure 9 This is a schematic diagram of an electronic device. The electronic device provided in this embodiment includes: a processor 1001, a memory 1002, a communication bus 1003, and a communication interface 1004; The memory 1002 is used to store computer instructions for data processing. When these computer instructions are read and executed by the processor 1001, the following steps are performed: Obtain the target audio and reference image for generating the video, wherein the reference image includes the sound-producing object; The identification features of the sound-producing object are extracted from the reference image, and the identification features are used to represent the feature information of each part of the sound-producing object; Determine the target motion features that match the target audio, wherein the target motion features are used to represent the motion information of each part of the sound-producing object in each target video frame corresponding to the target audio; Based on the recognition features and the target motion features, the overall features of the sound-producing object in the target video frame are obtained, and the target video frame is obtained based on the overall features, so as to generate a corresponding video based on each target video frame and the target audio.
[0228] The seventh embodiment of this application also provides an electronic device embodiment corresponding to the training method of the video generation model provided in the second embodiment. The video generation model includes a motion feature generation model. The following description of the electronic device embodiment is merely illustrative. The electronic device embodiment is as follows: The electronic device provided in this embodiment includes: a processor, a memory, a communication bus, and a communication interface; This memory is used to store computer instructions for data processing. When these computer instructions are read and executed by the processor, the following steps are performed: Obtain a first training sample, which includes a sample video, a sample audio that matches the sample video, and a sample reference image. The sample video and the sample reference image include the same first sample sound source. Extract the first sample motion feature of the first sample sound-emitting object from the sample video frame of the sample video, and extract the second sample motion feature of the first sample sound-emitting object from the sample reference image; Based on the second sample motion features and the sample audio, the first sample input information is determined, and the first sample input information is input into the motion feature generation model to be trained to obtain the first output motion feature. Based on the relationship between the first output motion feature and the first sample motion feature.
[0229] The eighth embodiment of this application also provides a computer-readable storage medium for implementing the method described in the first embodiment. The embodiments of the computer-readable storage medium provided in this application are described in a relatively simple manner; relevant parts can be found in the corresponding descriptions of the above method embodiments. The embodiments described below are merely illustrative.
[0230] The computer-readable storage medium provided in this embodiment stores computer instructions, which, when executed by a processor, implement the steps described in any one of the first embodiments.
[0231] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0232] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0233] 1. Computer-readable media includes both permanent and non-permanent, removable and non-removable media, which can store information by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined in this application, computer-readable media does not include non-transitory computer-readable media, such as modulated data signals and carrier waves.
[0234] 2. Those skilled in the art will understand that embodiments of this application can provide methods, systems, or computer program products. Therefore, embodiments of this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, embodiments of this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0235] 3. This application embodiment may involve the use of user data. In practical applications, user-specific personal data may be used within the scope permitted by applicable laws and regulations of the country in which the application is located (e.g., with the user's explicit consent and effective notification to the user, etc.). Furthermore, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. The collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0236] Although this application discloses preferred embodiments as described above, it is not intended to limit this application. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of this application. Therefore, the scope of protection of this application should be determined by the scope defined in the claims of this application.
Claims
1. A method of video generation, the method comprising: The method comprises: acquiring target audio for generating a video and a reference picture, the reference picture comprising a sound-emitting object; extracting identification features of the sound-emitting object from the reference picture, the identification features being used to represent feature information of each part of the sound-emitting object; determining target motion features matched with the target audio, the target motion features being used to represent motion information of each part of the sound-emitting object in each target video frame corresponding to the target audio; obtaining overall features of the sound-emitting object in the target video frame according to the identification features and the target motion features, and obtaining the target video frame according to the overall features, so as to generate a corresponding video based on each target video frame and the target audio.
2. The video generation method of claim 1, wherein, Before the determining of the target motion features matched with the target audio, the method comprises: extracting original motion features of the sound-emitting object from the reference picture; the determining of the target motion features of the sound-emitting object matched with the target audio comprises: adjusting the original motion features according to the target audio to obtain the target motion features of the sound-emitting object matched with the target audio.
3. The video generation method of claim 2, wherein, The original motion features of the sound-emitting object in the reference image are extracted in the following way: determining motion coefficients corresponding to each basic motion unit of the sound-emitting object from the reference image, the motion coefficients being used to represent weights of the sound-emitting object on the corresponding motion unit; determining motion features corresponding to the motion coefficients as the original motion features of the sound-emitting object in the reference image based on a pre-trained motion feature library, wherein the motion feature library comprises each basic motion unit.
4. The video generation method of claim 3, wherein, The determining of the motion coefficients corresponding to each basic motion unit of the sound-emitting object from the reference image comprises: encoding the reference image by using a pre-trained image encoding model to obtain image encoding features; determining the motion coefficients corresponding to each basic motion unit of the sound-emitting object from the image encoding features by using a pre-trained motion coefficient determination model. The identification features of the sound-emitting object in the reference image are extracted in the following way: determining the identification features of the sound-emitting object from the image encoding features.
5. The video generation method of claim 2, wherein, The adjusting of the original motion features according to the target audio to obtain the target motion features of the sound-emitting object matched with the target audio comprises: acquiring a target emotion, the target emotion being used to indicate an emotion of the sound-emitting object when the sound-emitting object emits the target audio; determining target emotion-related motion features associated with the target emotion based on the original motion features; adjusting the target emotion-related motion features according to the target audio to obtain the target motion features of the sound-emitting object matched with the target audio.
6. The video generation method of claim 5, wherein, The determining of the target emotion-related motion features associated with the target emotion based on the original motion features comprises: acquiring a pre-trained target projection model corresponding to the target emotion, the target projection model being used to map motion features to corresponding motion features in a target emotion subspace; Projecting the original motion feature into a subspace of the target emotion according to the target projection model to obtain a target emotion-related motion feature associated with the target emotion.
7. The video generation method of claim 6, wherein, The projecting the original motion feature into a subspace of the target emotion according to the target projection model to obtain a target emotion-related motion feature associated with the target emotion comprises: determining a reference emotion of the sound-making object in the reference image; obtaining a pre-trained reference projection model corresponding to the reference emotion, the reference projection model being used to map a motion feature in a parent space to a corresponding motion feature in a subspace of the reference emotion; filtering out a motion feature in the subspace of the reference emotion from the original motion feature based on the reference projection model to obtain an emotion-independent motion feature; projecting the original motion feature into a subspace of the target emotion according to the target projection model to obtain a target emotion motion feature; combining the emotion-independent motion feature and the target emotion motion feature to obtain a target emotion-related motion feature associated with the target emotion.
8. The video generation method of claim 5, wherein, The adjusting the target emotion-related motion feature according to the target audio to obtain a target motion feature of the sound-making object matching the target audio comprises: adjusting the target emotion-related motion feature according to the target audio and the target emotion based on a pre-trained motion feature generation model to obtain a target motion feature of the sound-making object matching the target audio.
9. The video generation method of claim 8, wherein, The motion feature generation model obtains a target motion feature of a sound-making object in a target video frame by performing multi-step denoising on a noise-added noise motion feature; The adjusting the target emotion-related motion feature according to the target audio to obtain a target motion feature of the sound-making object matching the target audio comprises: in a multi-step denoising process of the motion feature generation model, injecting a denoising reference condition into the motion feature generation model at different denoising stages according to an injection sequence of the target emotion-related motion feature, the target audio, and the target emotion, so that the motion feature generation model performs denoising with the target emotion-related motion feature, the target audio, and the target emotion as conditions in turn at each denoising stage to obtain a target motion feature of the sound-making object matching the target audio.
10. The video generation method of claim 9, wherein, The motion feature generation model comprises a plurality of model parts arranged in sequence, the model parts comprising a condition adjustment layer, a denoising layer, and a gate control layer, the condition adjustment layer being configured to adjust the injected denoising reference condition, the denoising layer being configured to perform denoising on data input from a previous layer based on the injected denoising reference condition, and the gate control layer being configured to control output of the denoising layer.
11. The video generation method of any of claims 1-10, wherein, The obtaining the overall feature of the sound-making object in the target video frame according to the recognition feature and the target motion feature comprises: combining the recognition feature and the target motion feature to obtain the overall feature of the sound-making object in the target video frame; The target video frame is obtained according to the overall feature, and the method comprises: The overall feature is decoded by a pre-trained decoding model to obtain the target video frame.
12. The video generation method of claim 5, wherein, The target emotion is obtained, and the method comprises: According to the audio content of the target audio, the target emotion corresponding to each part of the target audio is determined.
13. A method of live video streaming, the method comprising: The method comprises: The target audio corresponding to the live video to be broadcast and the reference picture are obtained, and the reference picture comprises the anchor object; The live video is generated by the video generation method in any one of claims 1 to 12; The live video is played in real time by the client.
14. A method for training a video generation model, comprising: The video generation model comprises a motion feature generation model, and the method comprises: A first training sample is obtained, and the first training sample comprises a sample video, a sample audio matched with the sample video, and a sample reference picture, wherein the sample video and the sample reference picture comprise a same first sample sound object; First sample motion features of the first sample sound object are extracted from sample video frames of the sample video, and second sample motion features of the first sample sound object are extracted from the sample reference picture; First sample input information is determined according to the second sample motion features and the sample audio, and the first sample input information is input into a to-be-trained motion feature generation model to obtain first output motion features; The to-be-trained motion feature generation model is adjusted in parameters according to a gap between the first output motion features and the first sample motion features, to obtain a trained motion feature generation model.
15. The training method of claim 14, wherein, The first training sample further comprises sample emotions corresponding to each frame of the sample video; The first sample input information is determined according to the second sample motion features and the sample audio, and the method comprises: The first sample input information is determined according to the second sample motion features, the sample audio, and the sample emotions.
16. The training method of claim 14, wherein, The to-be-trained motion feature generation model obtains output motion features by performing multi-step denoising on a sample noise motion feature after adding noise; The first sample input information is input into the to-be-trained motion feature generation model to obtain the first output motion features, and the method comprises: In a multi-step denoising process of the to-be-trained motion feature generation model, sample denoising reference conditions are injected into the to-be-trained motion feature generation model at different denoising stages according to a sample condition injection order of the second sample motion features, the sample audio, and the sample emotions, to obtain the first output motion features.
17. The training method of claim 14, wherein, The video generation model further comprises a motion feature library, and the motion feature library comprises various basic motion units; the method further comprises: A second training sample is obtained, and the second training sample comprises a first sample source picture and a first sample target picture, wherein the first sample source picture and the first sample target picture comprise a same second sample sound object, the second sample sound object in the first sample source picture is in a first motion state, the second sample sound object in the first sample target picture is in a second motion state, and the first motion state is different from the second motion state; determine first sample motion coefficients corresponding to each basic motion unit of the second sample vocal object from the first sample target picture; determine third sample motion features corresponding to the first sample motion coefficients through a to-be-trained motion feature library; obtain an output picture according to the sample identification feature and the third sample motion feature; adjust parameters in the to-be-trained motion feature library according to a difference between the output picture and the sample target picture, to obtain a trained motion feature library; the first sample motion feature of the first sample vocal object is extracted from a sample video frame of the sample video, including: determine second sample motion coefficients corresponding to each basic motion unit of the first sample vocal object from the sample video frame of the sample video; determine first sample motion features corresponding to the second sample motion coefficients through a pre-trained motion feature library; the second sample motion feature of the first sample vocal object is extracted from the sample reference picture, including: determine third sample motion coefficients corresponding to each basic motion unit of the first sample vocal object from the sample reference picture; determine second sample motion features corresponding to the third sample motion coefficients through a pre-trained motion feature library.
18. The training method of claim 15, wherein, The video generation model further includes a projection model corresponding to each emotion, which is used to map the motion feature to a subspace corresponding to the emotion; the method further includes: obtain a third training sample, the third training sample including a second sample source image and a second sample target image, the second sample source image and the second sample target image including a same third sample vocal object, the second sample source image and the second sample target image having different emotions of the third sample vocal object; extract a source sample motion feature of the third sample vocal object from the second sample source image, and extract a target sample motion feature of the third sample vocal object from the second sample target image; project the source sample motion feature to the target emotion space through a to-be-trained target emotion projection model to obtain an output emotion feature; adjust the to-be-trained reference emotion projection model and the to-be-trained target emotion projection model according to a difference between the output emotion feature and the target sample motion feature, to obtain a trained target emotion projection model; the first sample input information is determined according to the second sample motion feature, the sample audio, and the sample emotion, including: project the second sample motion feature into a subspace corresponding to the sample emotion through a pre-trained sample emotion projection model to obtain a sample emotion motion feature; determine first sample input information according to the sample emotion motion feature, the sample audio, and the sample emotion.
19. A video generating apparatus, comprising: The device includes: an information acquisition unit configured to acquire a target audio for generating a video and a reference picture including a vocal object. The feature extraction unit is configured to extract identification features of the sound-producing object from the reference picture, the identification features being used to represent feature information of each part of the sound-producing object. The feature determination unit is configured to determine target motion features matched with the target audio, the target motion features being used to represent motion information of each part of the sound-producing object in each target video frame corresponding to the target audio. The video generation unit is configured to obtain overall features of the sound-producing object in the target video frame according to the identification features and the target motion features, and obtain the target video frame according to the overall features, so as to generate a corresponding video based on each target video frame and the target audio.
20. An electronic device, comprising: Comprise: a processor, a memory, and computer program instructions stored on the memory and executable on the processor; the processor executes the computer program instructions to implement the method of any one of claims 1-18.
21. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer execution instructions, and the computer execution instructions are executed by the processor to implement the method of any one of claims 1-18.