Video generation method, apparatus and computer-readable storage medium

By extracting anthropomorphic features from audio-video sequences and using a virtual prediction network to generate virtual object videos, the method enhances the realism and accuracy of virtual object videos by considering both speaking and listening behaviors.

JP7818106B2Active Publication Date: 2026-02-19BEIJING JINGDONG SHANGKE INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024569572
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-07-14
Filing Date
2023-02-13
Publication Date
2026-02-19
Estimated Expiration
2043-02-13

AI Technical Summary

Technical Problem

Existing technologies for generating virtual object videos based on speaker reference images and time-varying signals result in poor image quality due to inadequate consideration of both listening and speaking behaviors in human-machine interaction.

Method used

A method involving feature extraction from audio-video sequences to determine anthropomorphic features, using a virtual prediction network to generate a video sequence of a virtual object by predicting posture and expression features based on standard features and first features, and fusing these with identity features to enhance video quality.

Benefits of technology

The method generates a more realistic and accurate video sequence of a virtual object by incorporating listening behaviors, improving the vividness and accuracy of virtual object videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007818106000037
    Figure 0007818106000037
  • Figure 0007818106000038
    Figure 0007818106000038
  • Figure 0007818106000039
    Figure 0007818106000039
Patent Text Reader

Abstract

Embodiments of the present invention provide a video generation method, apparatus, and computer-readable storage medium. Among them, the method includes: collecting the audio-video sequence of a real object; extracting features from the audio-video sequence to determine anthropomorphic features; using a virtual prediction network, a preset standard feature that is a feature corresponding to a reference object, and a first feature representing different attitudes to predict the anthropomorphic features; generating a video sequence of a virtual object, which is a video sequence in which a corresponding reaction of the virtual object is generated based on the audio-video sequence of the real object; and presenting the video sequence of the virtual object. Embodiments of the present invention generate a video sequence of a virtual object based on the audio-video sequence of a real object, thereby making the presented video sequence of the virtual object more vivid and accurate.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] The present invention is based on and claims priority from a Chinese patent application bearing application number 202210834191.6 and filing date July 14, 2022, the entire contents of which are hereby incorporated by reference into the present invention.

[0002] The present invention relates to the field of human-machine interaction, and in particular to a video generation method, apparatus and computer-readable storage medium. [Background technology]

[0003] From the perspective of human behavior, good communication refers to a two-way communication process. It is not a one-way input or output of information. Along with information interaction, communication and exchange between real people is a process of constantly switching between two states: listening and speaking. Of these, listening and speaking are equally important. Both are essential for building anthropomorphic digital humans and engaging in human-machine interaction. While digital humans need to express their own perspectives as clearly, concisely, and articulately as possible in a language that others can understand, anthropomorphic digital humans also need to be able to listen to and understand the perspectives of others. Previous technologies primarily generate corresponding speaker videos based on speaker reference images and time-varying signals. These methods primarily involve parameterizing the speaker using facial keypoints, a 3D facial model, a human skeletal model, etc., and then fitting these parameters to a deep neural network to generate a rendering image based on these parameters. However, the resulting image quality is poor. Summary of the Invention [Problem to be solved by the invention]

[0004] Embodiments of the present invention provide a video generation method, apparatus, and computer-readable storage medium that can enhance the vividness and accuracy of the generated video sequence of a virtual object by generating a video sequence of a virtual object based on an audio-video sequence of a real object. [Means for solving the problem]

[0005] The technical aspects of the present invention are realized as follows.

[0006] An embodiment of the present invention comprises: Acquiring an audio-video sequence of a real object; performing feature extraction on the audio-video sequence to determine anthropomorphic features; performing predictions on the anthropomorphic features using a virtual prediction network, predetermined standard features corresponding to the reference object, and first features representing different attitudes, and generating a video sequence of a virtual object, which is a video sequence in which a corresponding reaction of the virtual object is generated based on an audio-video sequence of the real object; and presenting a video sequence of the virtual object.

[0007] In the above aspect, performing predictions on the anthropomorphic features using the virtual prediction network, the preset standard features, and the first features to generate a video sequence of a virtual object includes: Acquiring preset standard features including a first posture / expression feature and a first identity feature; determining posture and expression features of a plurality of frames of a virtual object by performing prediction and decoding using the virtual prediction network based on the first posture and expression features, the anthropomorphic features, and the first features; and generating a video sequence of the virtual object based on the posture and expression features of the virtual object and the first identity features of the plurality of frames.

[0008] In the above aspect, the acquiring of the preset standard feature includes: obtaining a standard image representing an image of a reference object; Extracting features from the standard image using a face reconstruction model to obtain the preset standard features.

[0009] In the above aspect, the anthropomorphic features include anthropomorphic features of corresponding multiple frames of the audio-video sequence, and the virtual prediction network includes a first processing module and a second processing module. The step of predicting and decoding the first posture-expression features, the anthropomorphic features, and the first features using the virtual prediction network to determine posture-expression features of multiple frames of the virtual object includes: According to the anthropomorphic features of a first frame among the plurality of frames, the first posture-expression feature, and the first feature, which is one of a positive attitude, a negative attitude, and a general attitude, performing prediction by the first processing module to obtain a next predicted video frame; Decoding the next predicted video frame by the second processing module and determining a next pose / expression feature of the virtual object corresponding to the next predicted video frame; continuing to predict and decode based on the next posture and expression feature and the anthropomorphic feature of a next frame among the plurality of frames of anthropomorphic features until obtaining a last posture and expression feature of the virtual object corresponding to a last predicted video frame, thereby obtaining posture and expression features of a plurality of frames of the virtual object, wherein the first posture and expression feature is the posture and expression feature of a first frame among the plurality of frames of anthropomorphic features.

[0010] In the above aspect, generating a video sequence of the virtual object based on the posture and expression features of a plurality of frames of the virtual object and the first identity feature includes: Fusing each of the posture and expression features of each frame among the posture and expression features of the plurality of frames with the first identity features including a first identity mark, a first material, and first light irradiation information, to obtain a plurality of second features representing the fusion results of the identity features and the posture and expression features; and generating a video sequence of the virtual object by a renderer for the plurality of second characteristics.

[0011] In the above aspect, the audio-video sequence includes an audio sequence of a real object and a video sequence of a real object, and performing feature extraction on the audio-video sequence to determine anthropomorphic features includes: pre-processing a video sequence of the real object by an encoder to obtain a plurality of video features; performing feature extraction on the audio sequence of the real object by an encoder to obtain a plurality of audio features including loudness, zero-crossing rate, and cepstral coefficients; and performing feature transformation using a feature fusion function based on the plurality of video features and the plurality of audio features to determine anthropomorphic features for the corresponding plurality of frames of the audio-video sequence, the anthropomorphic features including video features and audio features.

[0012] In the above aspect, the step of extracting features from a video sequence of the real object by an encoder to obtain a plurality of video features includes: extracting features for each video frame of the video sequence of the real object using a face reconstruction model to obtain a plurality of video frame features including second identity features and second pose-expression features; and taking all the corresponding second pose-expression features in a video sequence of the real object as the video features.

[0013] In the above aspect, before performing predictions on the anthropomorphic features using the virtual prediction network, the preset standard features, and the first features, and generating a video sequence of the virtual object, the method further comprises: acquiring audio-video sequence samples of real talking objects and corresponding facial images of real listening objects; performing feature extraction on the sample audio-video sequence by an initial encoder to determine anthropomorphic sample features; generating predicted facial features in the audio-video sequence samples of the training real object, including predicted pose features and predicted facial expression features, using the initial virtual prediction network and the anthropomorphic sample features; Extracting features from the face image of the real listening object using a face reconstruction model to determine real face features including real posture features and real expression features; continue optimizing the initial encoder with a first loss function and the anthropomorphic sample features until a first loss function value satisfies a first preset threshold, and determine the encoder; The method further includes continuing to optimize the initial virtual predictive network using the second loss function and the third loss function based on the actual facial features and the predicted facial features until the second loss function value and the third loss function value satisfy a second preset threshold, thereby determining the virtual predictive network.

[0014] In the above aspect, continuing to optimize the initial virtual predictive network by the second loss function and the third loss function based on the actual facial features and the predicted facial features until the second loss function value and the third loss function value satisfy a second preset threshold, and determining the virtual predictive network, determining a second loss function based on the actual facial features and the predicted facial features to ensure that the predicted facial expression and pose are similar to the actual facial expression and pose; determining a third loss function based on the change function corresponding to the actual facial feature and the change function corresponding to the predicted facial feature to ensure that the inter-frame continuity of the predicted facial feature resembles the actual facial feature; and continuing to optimize the initial virtual prediction network using the second loss function and the third loss function until the second loss function value and the third loss function value satisfy a second preset threshold, thereby determining the virtual prediction network.

[0015] An embodiment of the present invention comprises: comprising an obtaining portion, a determining portion, and a generating portion; the acquisition part is configured to capture an audio-video sequence of a real object; the determining portion is configured to perform feature extraction on the audio-video sequence to determine anthropomorphic features; The generation unit is configured to make predictions on the anthropomorphic features using a virtual prediction network, predetermined standard features that are features corresponding to a reference object, and first features that represent different attitudes, generate a video sequence of a virtual object that is a video sequence in which a corresponding reaction of a virtual object is generated based on an audio-video sequence of a real object, and present the video sequence of the virtual object.

[0016] An embodiment of the present invention comprises: a memory for storing executable instructions; a processor for executing executable instructions stored in the memory, wherein execution of the executable instructions causes the processor to perform the video generation method described above.

[0017] An embodiment of the present invention comprises: A computer-readable storage medium is provided having stored thereon executable instructions that, when executed by one or more processors, cause the processors to perform the video generation method described above. [Effects of the Invention]

[0018]

[0013] Embodiments of the present invention provide a video generation method, apparatus, and computer-readable storage medium, the method including: collecting an audio-video sequence of a real object; performing feature extraction on the audio-video sequence to determine anthropomorphic features; performing prediction on the anthropomorphic features using a virtual prediction network, predetermined standard features corresponding to features of a reference object, and first features representing different attitudes; generating a video sequence of a virtual object, which is a video sequence of a virtual object corresponding to the audio-video sequence of the real object; and presenting the video sequence of the virtual object. The embodiments of the present invention generate a video sequence of a virtual object based on the audio-video sequence of the real object, thereby making the presented video sequence of the virtual object more realistic and accurate. [Brief explanation of the drawings]

[0019] [Figure 1] FIG. 2 is a schematic diagram of one selectable terminal in the video generation method according to an embodiment of the present invention; [Figure 2] 1 is a flow diagram 1 of one alternative of a video generation method according to an embodiment of the present invention; [Figure 3a] 1 is a schematic diagram 1 of one selectable speaker video generation of a video generation method according to an embodiment of the present invention; [Figure 3b] 2 is a schematic diagram 2 of one selectable speaker video generation of the video generation method according to an embodiment of the present invention; [Figure 3c] 3 is a schematic diagram 3 of one selectable speaker video generation of the video generation method according to an embodiment of the present invention; [Figure 4] 2 is a flow diagram 2 showing one alternative of a video generating method according to an embodiment of the present invention; [Figure 5] 3 is a flow diagram 3 showing one alternative of a video generating method according to an embodiment of the present invention; [Figure 6] 4 is a flow diagram 4 showing one alternative of a video generating method according to an embodiment of the present invention; [Figure 7] 5 is a flow diagram 5 showing one alternative of a video generating method according to an embodiment of the present invention; [Figure 8a] FIG. 1 is a result of a video sequence of a virtual object of a video generation method according to an embodiment of the present invention. [Figure 8b] FIG. 2 is a result of a video sequence of a virtual object of a video generation method according to an embodiment of the present invention. [Figure 9] 6 is a flow diagram 6 showing one alternative of a video generating method according to an embodiment of the present invention; [Figure 10] FIG. 2 is an architecture diagram of one alternative model of a video generation method according to an embodiment of the present invention. [Figure 11] 1 is a structural schematic diagram 1 of a video generating device according to an embodiment of the present invention; [Figure 12] 2 is a structural schematic diagram 2 of a video generating device according to an embodiment of the present invention; DETAILED DESCRIPTION OF THE INVENTION

[0020] Hereinafter, the technical aspects of the embodiments of the present invention will be described clearly and completely with reference to the drawings in the embodiments of the present invention. It is clear that the described embodiments are only some embodiments of the present invention, and are not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without the need for creative work fall within the scope of protection of the present invention.

[0021] In order to allow those skilled in the art to better understand the aspects of the present invention, the present invention will be described in more detail below with reference to the drawings and specific embodiments. Figure 1 is a schematic diagram of one optional terminal operation of a video generation method according to an embodiment of the present invention. As shown in Figure 1, the terminal is equipped with a speaker encoder (corresponding to an encoder) to extract features from the audio / video of a real object, and inputs the extracted features, attitudes (corresponding to first features) and a reference image (corresponding to a preset standard image) into a listener decoder (corresponding to a virtual prediction network) for prediction, thereby generating head movements and facial expression changes of the listener arranged in a timeline, and obtaining a video sequence of a virtual object.

[0022] In some embodiments of the present invention, FIG. 2 is an alternative flow diagram 1 of a video generation method according to an embodiment of the present invention, which will be described with reference to the steps shown in FIG.

[0023] In S101, an audio-video sequence of a real object is collected.

[0024] In some embodiments of the present invention, based on concepts from social psychology and anthropology, "listening" is a functional act during communication. Among these, listening behavior styles can be divided into four types: non-listener, marginal listener, evaluative listener, and active listener. Actively responsive listening is the most effective and plays an important role in communication. It requires the listener to fully concentrate and listen attentively to what the speaker is saying, while also providing some visual feedback to the speaker. These responses, such as whether the listener is interested, understands, or agrees with what is being said, can be fed back to the speaker to adjust the rhythm and process of the conversation and promote smooth communication.

[0025] In active listening, listeners often have visual patterns to express their opinions. For example, symmetrical circular movements are used to indicate "yes," "no," or similar signals. Small linear movements are often combined with stressed syllables in the other person's speech, and larger linear movements often appear during pauses in the other person's speech. In human face-to-face interactions, even the listener's blinking can be considered an interaction signal. Therefore, it is important to generate a video sequence of a virtual object listening to an audio-video sequence based on several audio-video sequences.

[0026] For example, given a reference image of a speaker and a time-varying signal of one section, a simulated segment that can be matched to the time-varying signal of one speaker is generated. Figures 3a, 3b, and 3c respectively show selectable speaker video generation schemes 1, 2, and 3 of a video generation method according to an embodiment of the present invention. As shown in Figure 3a, the speaker video generation task includes generating a speaker's body pose. As shown in Figure 3b, the speaker video generation task includes generating a speaker's lip movement. As shown in Figure 3c, the speaker video generation task includes generating a speaker's head (including face) movement. In Figure 3a, the speaker's body pose is generated by processing a time-varying signal of one section input within a dashed line frame using a body pose generation model to obtain a body pose indicated within a dashed line frame. In Figure 3b, the generation of the speaker's lip movement involves processing a time-varying signal of one section input in the dashed frame and a general reference image using a lip movement generation model, which then outputs an image frame of the speaker's lip movement shown in the dashed-dotted frame. In Figure 3c, the generation of the speaker's head (including face) movement involves processing a time-varying signal of one section input in the dashed frame, a reference image of the speaker, and their mood using a head movement generation model, which then renders the processed result using a head rendering model, which then outputs an image frame of the speaker's head (including face) movement shown in the dashed-dotted frame.

[0027] In some embodiments of the invention, the terminal is capable of capturing audio and video sequences of real objects by means of a collection device.

[0028] For example, the collection device may be a device with the function of capturing video and audio, such as a camera head, but the present invention is not limited thereto; the real object may be a person speaking in a scene, and the audio and video sequences may be obtained in a scene where the spectator is at a tourist attraction, in the process of the tourist inquiring at a self-service information device (carrier of the virtual object).

[0029] In some embodiments of the present invention, the present invention is applied in situations where human-machine interaction is required, such as an intelligent information device in a department store that can create a corresponding video based on a video displayed by a shopper and guide the shopper.

[0030] At S102, feature extraction is performed on the audio-video sequence to determine anthropomorphic features.

[0031] In some embodiments of the present invention, the audio-video sequence comprises an audio sequence of a real object and a video sequence of a real object.

[0032] In some embodiments of the present invention, feature extraction can be achieved by a neural network model. The feature extraction process involves inputting each frame of a video sequence into a neural network, and then performing feature extraction using multiple convolutional layers and pooling layers to obtain multiple video features. An anthropomorphic feature is a feature obtained by characterizing a certain limb movement, which is accompanied by several limb movements, when a person speaks in a real-world scene. The anthropomorphic feature includes audio features and video features, and is the audio and video features obtained after feature extraction is performed on an audio-video sequence. For example, if a speaker in a video raises his or her hand while speaking, the anthropomorphic feature may be a feature corresponding to the hand raising movement.

[0033] In some embodiments of the present invention, the terminal may perform feature extraction on a video sequence of a real object by an encoder to obtain a plurality of video features, perform feature extraction on an audio sequence of the real object by an encoder to obtain a plurality of audio features including loudness, zero-crossing rate and cepstral coefficients, and perform feature transformation using a feature fusion function based on the plurality of video features and the plurality of audio features to determine an anthropomorphic feature.

[0034] In some embodiments of the present invention, FIG. 4 is one optional flow diagram 2 of a video generation method according to an embodiment of the present invention, and as shown in FIG. 4, S102 can be realized by S1021 to S1023 as follows:

[0035] In S1021, an encoder performs feature extraction on a video sequence of a real object to obtain a plurality of video features.

[0036] In some embodiments of the present invention, the video features are obtained by recording head rotations and facial expressions during human interaction, where the head rotations and facial expressions are recorded, and the video features are obtained after performing feature extraction on the video sequence, where the video features include posture and facial expressions.

[0037] In some embodiments of the present invention, the terminal can perform feature extraction for each video frame of the video sequence of the real object using a facial reconstruction model to obtain multiple video frame features, and take all corresponding second pose and expression features in the video sequence of the real object as video features.

[0038] In some embodiments of the present invention, Figure 5 is one optional flow diagram 3 of a video generation method according to an embodiment of the present invention, and as shown in Figure 5, S1021 can be realized by S10211 and S10212 as follows:

[0039] In S10211, feature extraction is performed for each video frame of the video sequence of the real object using a face reconstruction model to obtain a plurality of video frame features.

[0040] In some embodiments of the present invention, the video frame features are obtained by recording the head rotation, facial expression, and multiple captured elements of a person in each video frame of a video sequence. The video frame features include a second identity feature and a second pose / expression feature. The second identity feature is the result of recording the captured environmental elements of the video sequence and the identity information of the subject being photographed. The second identity feature includes the identity mark, material, and lighting of the real object. The second pose / expression feature is the result of recording the head rotation and facial expression changes that occur when the real object is speaking. The second pose / expression feature includes the head pose and facial expression of the real object. As for the facial reconstruction model, a 3D facial reconstruction model is generally selected, but the present invention is not limited thereto.

[0041] In some embodiments of the present invention, the terminal can use a face reconstruction model to extract features for the person and the background in each video frame of a video sequence of a real object, and each video frame can provide the person's identity indicator, the material of the video frame, the lighting during shooting, the person's head pose, and the person's facial expression. The person's identity indicator, the material of the video frame, and the lighting during shooting are used as second identity features, the person's head pose and the person's facial expression are used as second pose / expression features, and all the features (i.e., the person's identity indicator, the material of the video frame, the lighting during shooting, the person's head pose, and the person's facial expression) are used as one video frame feature.

[0042] Exemplarily, a video sequence

number

[0043] In S10212, all corresponding second pose and expression features in the video sequence of the real object are taken as video features.

[0044] In some embodiments of the present invention, the terminal determines all corresponding second pose-expression features in the video sequence of the real object as the video features corresponding to the video sequence.

[0045] For example, the parameters {α, β, δ, p, γ} are divided into two types. One is a relatively fixed feature that is closely linked to the identity information.

number

number

[0046] It can be understood that in some embodiments of the present invention, the terminal performs feature extraction for each video frame of the video sequence of the real object using a face reconstruction model to obtain multiple video frame features, and takes all corresponding second pose and expression features in the video sequence of the real object as video features, thereby removing identity-specific features (such as a person's identity mark, the material of the video frame, and the lighting during shooting) and leaving only common features, thereby improving the effectiveness of feature extraction and providing data support for the subsequent generation of the video sequence of the virtual object.

[0047] In S1022, the encoder performs feature extraction on the audio sequence of the real object to obtain a plurality of audio features.

[0048] In some embodiments of the present invention, the audio features are some features associated with a speaker when he / she is speaking in a human interaction process, which are obtained after performing feature extraction on an audio sequence, and the audio features include loudness, zero-crossing rate and cepstral coefficients corresponding to when a real object is speaking.

[0049] In some embodiments of the present invention, the terminal performs feature extraction on the spectrum of the audio sequence of the real object through an encoder, obtains the loudness, zero-crossing rate and cepstral coefficient at each time when the real object speaks, and regards the loudness, zero-crossing rate and cepstral coefficient as all audio features to obtain multiple audio features.

[0050] Exemplarily, an audio sequence of a real object

number

number

number

[0051] It can be understood that the terminal can perform feature extraction based on the audio-video sequence generated when the real object is speaking, obtain a corresponding number of audio features and a corresponding number of video features, quickly extract effective audio features and effective video features, and improve the speed of feature extraction and the effectiveness of the features.

[0052] In S1023, based on the plurality of video features and the plurality of audio features, a feature transformation is performed by a feature fusion function to determine an anthropomorphic feature.

[0053] In some embodiments of the present invention, the anthropomorphic features are anthropomorphic features of corresponding multiple frames of an audio-video sequence, the anthropomorphic features include video features and audio features, and the feature fusion function can convert non-linear features into linear features.

[0054] In some embodiments of the present invention, the terminal performs nonlinear feature transformation on multiple audio features and multiple video features through a multi-modal feature fusion function in the encoder to obtain feature representations of real objects, i.e., determine anthropomorphic features.

[0055] It can be understood that by extracting features from an audio-video sequence, the terminal can increase the speed at which the audio-video sequence is processed and quickly obtain multiple video features and multiple audio features; and by converting multiple video features and multiple audio features into a first feature using a feature fusion function, the representation method of the features can be converted, which can facilitate the terminal's processing thereof and improve the feasibility of processing the video features and audio features.

[0056] In S103, prediction is performed on the anthropomorphic features using the virtual prediction network, the preset standard features, and the first features to generate a video sequence of the virtual object.

[0057] In some embodiments of the present invention, the video sequence of the virtual object is a video sequence in which a corresponding reaction of the virtual object is generated based on an audio-video sequence of a real object, and the predetermined standard features are features corresponding to a reference object, and the predetermined standard features include a first posture-expression feature and a first identity feature corresponding to the reference object, the first posture-expression feature being a head posture and a facial expression of the reference object, and the first identity feature being information related to the identity of the reference object. The first feature is an attitude feature having an emotional color when a person is speaking. The first feature represents different attitudes, and may be a positive attitude, a negative attitude, or a general attitude.

[0058] In some embodiments of the present invention, the terminal can obtain preset standard features, and perform prediction and decoding using a virtual prediction network based on the first posture-expression feature, the anthropomorphic feature and the first feature to determine posture-expression features of multiple frames of a virtual object, and generate a video sequence of the virtual object according to the posture-expression features of the multiple frames and the first identity feature of the virtual object.

[0059] For example, an input video sequence of real objects with time lengths from 1 to t

number

number

number

number

[0060] In some embodiments of the present invention, FIG. 6 is one optional flow diagram 4 of a video generation method according to an embodiment of the present invention, and as shown in FIG. 6, S103 can be realized by S1031 to S1033 as follows:

[0061] In S1031, a preset standard feature is acquired.

[0062] In some embodiments of the present invention, the terminal can obtain a standard image, extract features from the standard image using a face reconstruction model, obtain first posture and expression features and first identity features, and set the first posture and expression features and first identity features as pre-set standard features.

[0063] In some embodiments of the present invention, S1031 can be realized by S10311 and S10312 as follows:

[0064] In S10311, a standard image is acquired.

[0065] In some embodiments of the present invention, the standard image is an image of a reference object, which is any facial image different from the real object randomly taken in an image library.

[0066] In S10312, features are extracted from the standard image using a face reconstruction model to obtain preset standard features.

[0067] In some embodiments of the present invention, the terminal extracts features from the standard image using a face reconstruction model to obtain the person's identity mark, the material of the video frame, the lighting during shooting, the person's head pose, and the person's facial expression in the standard image. The person's identity mark, the material of the video frame, and the lighting during shooting in the standard image are defined as first identity features, the person's head pose and the person's facial expression in the standard image are defined as first pose-expression features, and the first identity feature and the first pose-expression feature are defined as pre-set standard features.

[0068] It can be understood that the terminal can obtain a standard image, and by extracting features from the standard image using a facial reconstruction model, can quickly and accurately obtain a first pose expression feature, thereby improving the accuracy and efficiency of processing the standard image, and the first pose expression feature can be used to generate a video sequence of a virtual object, ensuring the accuracy of the video sequence of the generated virtual object.

[0069] In S1032, prediction and decoding are performed by a virtual prediction network based on the first posture / expression feature, the anthropomorphic feature, and the first feature to determine posture / expression features of a plurality of frames of the virtual object.

[0070] In some embodiments of the present invention, the terminal performs prediction using the anthropomorphic features, the first posture / expression features, and the first features of a first frame among the anthropomorphic features of the multiple frames using a first processing module to obtain a next predicted video frame, decodes the next predicted video frame using a second processing module to determine a next posture / expression feature of the corresponding virtual object in the next predicted video frame, and continues predicting and decoding using the next posture / expression feature and the anthropomorphic features of the next frame among the anthropomorphic features of the multiple frames until a last posture / expression feature of the corresponding virtual object in the last predicted video frame is obtained, thereby obtaining posture / expression features of the virtual object in multiple frames.

[0071] In some embodiments of the present invention, S1032 can be realized by S10321, S10322, and S10323 as follows:

[0072] In S10321, a first processing module performs prediction based on the anthropomorphic features of a first frame, a first posture-expression feature, and a first feature of the anthropomorphic features of a plurality of frames, to obtain a next predicted video frame.

[0073] In some embodiments of the present invention, the first characteristic is one of a positive attitude, a negative attitude, and a general attitude, and the first processing module is an encoding module in a virtual predictive network, and the function is to realize encoding in predictive processing.

[0074] In some embodiments of the present invention, the terminal can select any one of a positive attitude, a negative attitude, and a general attitude as the first feature. In some embodiments of the present invention, the terminal selects a general attitude, and in the general attitude, the anthropomorphic features of a first frame among the anthropomorphic features of multiple frames and the first posture and facial expression feature are input into a first processing module in a virtual prediction network to perform prediction and obtain a next video predicted frame.

[0075] In S10322, the second processing module decodes the next predicted video frame and determines the next posture and expression feature of the corresponding virtual object in the next predicted video frame.

[0076] In some embodiments of the present invention, the second processing module is a decoding module in a virtual prediction network, the function of which is to implement decoding in the prediction process.

[0077] In some embodiments of the present invention, the terminal can decode the next predicted video frame by a second processing module in the virtual prediction network, and obtain a next posture and facial expression feature of the corresponding virtual object in the next predicted video frame, where the next posture and facial expression feature includes a next posture feature and a next facial expression feature.

[0078] In S10323, prediction and decoding are continued based on the next posture and expression feature and the anthropomorphic feature of the next frame among the anthropomorphic features of the multiple frames until the last posture and expression feature of the corresponding virtual object of the last predicted video frame is obtained, thereby obtaining posture and expression features of the virtual object of the multiple frames.

[0079] In some embodiments of the present invention, the first posture-expression feature is a first frame of posture-expression features for a plurality of frames.

[0080] In some embodiments of the present invention, the terminal inputs the next posture and facial expression feature and the anthropomorphic feature of the second frame among the anthropomorphic features of the multiple frames into a first processing module in a virtual prediction network to perform prediction, obtains the next predicted video frame, decodes it using a second processing module in the virtual prediction network, determines the next posture and facial expression feature of the corresponding virtual object in the next predicted video frame, and continues predicting and decoding using the virtual prediction network until obtaining the last posture and facial expression feature of the corresponding virtual object in the last predicted video frame, thereby determining the posture and facial expression features of the multiple frames of the virtual object.

[0081] It can be understood that the terminal can make predictions based on the anthropomorphic features, the first posture and facial expression features, and the first features of the first frame; the first features can increase the diversity of the predicted video frames to be generated; the anthropomorphic features and the first posture and facial expression features of the first frame can increase the accuracy of the predicted video frames to be generated; by generating a next predicted video frame based on the first predicted video frame and predicting the next video frame using the predicted video frame updated in real time, consecutive video frames can be generated, the continuity between video frames can be increased, and the completeness of the video sequence can be ensured; and by decoding multiple predicted video frames to obtain multiple frames of posture and facial expression features, the virtual object can fully embody the head and facial reactions when listening to audio / video of a real object, and the accuracy of the video sequence of the generated virtual object can be increased.

[0082] In S1033, a video sequence of the virtual object is generated based on the posture and expression features and the first identity features of the virtual object in multiple frames.

[0083] In some embodiments of the present invention, virtual objects are simulated objects with simple communication capabilities, typically present on an interactive device.

[0084] In some embodiments of the present invention, the terminal can obtain multiple second features by fusing each of the posture and expression features of each frame of multiple frames of the virtual object with the first identity feature, and generate a video sequence of the virtual object using a renderer for the multiple second features.

[0085] It can be understood that the terminal obtains preset standard features, and performs prediction and decoding using a virtual prediction network based on the first posture and expression features, anthropomorphic features and first features in the preset standard features, to determine posture and expression features of multiple frames of the virtual object, thereby improving the accuracy of the posture and expression features of the multiple frames, and generating a video sequence of the virtual object based on the posture and expression features and first identity features of the virtual object in multiple frames, thereby improving the accuracy and vividness of the video sequence of the virtual object.

[0086] In some embodiments of the present invention, S1033 can be realized by S10331 and S10332 as follows:

[0087] In S10331, each of the posture and expression features of each frame among the posture and expression features of the virtual object in a plurality of frames is fused with the first identity feature to obtain a plurality of second features.

[0088] In some embodiments of the present invention, the second feature is a fusion feature formed after applying the identity feature to the posture / expression feature, and the second feature is a fusion result of the identity feature and the posture / expression feature.

[0089] In some embodiments of the present invention, the terminal combines each of the posture and expression features of each frame of the posture and expression features of the virtual object with the first identity feature to obtain multiple second features corresponding to the posture and expression features of the multiple frames.

[0090] In S10332, a video sequence of the virtual object is generated by a renderer for the plurality of second features.

[0091] In some embodiments of the present invention, the first identity feature includes a first identity mark, a first material, and first light illumination information, where the first identity mark corresponds to the identity mark of the person in the standard image, the first material corresponds to the material of the video frame, and the first light illumination information corresponds to the light illumination during shooting.

[0092] In some embodiments of the present invention, the terminal is capable of generating a video sequence of the virtual object by a renderer for a plurality of second characteristics.

[0093] Exemplarily, the video sequence generation task of a virtual object can be expressed in Equations (2) and (3).

number

number

[0094] It can be understood that the terminal can obtain multiple second features by fusing each of the posture and expression features of each frame of the posture and expression features of the virtual object in multiple frames with the first identity feature, and by giving the second features identity attributes, the distinctiveness of the second features can be increased. It can be understood that the terminal can generate a video sequence of the virtual object using a renderer for the multiple second features, and by generating the video sequence of the virtual object with a focus, the video sequence can be made more vivid and accurate.

[0095] At S104, a video sequence of the virtual object is presented.

[0096] In some embodiments of the present invention, the terminal is capable of presenting a video sequence of a virtual object to a virtual human interface.

[0097] It can be understood that the terminal collects an audio-video sequence of the real object, extracts features from the audio-video sequence, determines anthropomorphic features, predicts the anthropomorphic features using a virtual prediction network, preset standard features that are features corresponding to the reference object, and first features that represent different attitudes, generates a video sequence of a virtual object, which is a video sequence of a virtual object corresponding to the audio-video sequence of the real object, and presents the video sequence of the virtual object. When the video sequence of the virtual object is generated based on the audio-video sequence of the real object, the presented video sequence of the virtual object is more vivid and accurate.

[0098] In some embodiments of the present invention, FIG. 7 is one optional flow diagram 5 of a video generation method according to an embodiment of the present invention. As shown in FIG. 7, before performing S103, steps S105 to S1010 as follows are further performed:

[0099] In S105, audio-video sequence samples of real talking objects and corresponding face images of real listening objects are collected.

[0100] In some embodiments of the present invention, the audio-video sequence samples are generated by recording the speech of a real talking object and its corresponding movements and facial expressions while speaking.

[0101] In some embodiments of the present invention, the terminal can use a sampling device to capture audio-video sequence samples of real speaking objects and corresponding facial images of real listening objects, where the captured facial images of the real listening objects are captured with different first features, and the first features include positive attitude, negative attitude and general attitude.

[0102] At S106, feature extraction is performed on the audio-video sequence samples by the initial encoder to determine anthropomorphic sample features.

[0103] In some embodiments of the present invention, the anthropomorphic sample features are features obtained by characterizing some body movements that accompany a person speaking in a scene corresponding to the audio-video sequence sample, and the anthropomorphic sample features are obtained after performing feature extraction on the audio-video sequence sample, and include video sample features and audio sample features.

[0104] In some embodiments of the present invention, the terminal may perform feature extraction on the audio-video sequence samples by the initial encoder to obtain anthropomorphic sample features that include anthropomorphic sample features for multiple frames.

[0105] In S107, the initial virtual prediction network and the anthropomorphic sample features are used to generate predicted facial features for the sample audio-video sequences of the real object being trained.

[0106] In some embodiments of the present invention, the predicted facial features include predicted pose features and predicted expression features.

[0107] In some embodiments of the present invention, the terminal generates predicted posture features including predicted posture features for multiple frames and predicted facial features including predicted facial features for multiple frames in an audio-video sequence sample of a real object to be trained using an initial virtual prediction network and anthropomorphic sample features, and the predicted posture features and predicted facial features can be predicted facial features including predicted facial features for multiple frames.

[0108] In S108, based on the face image of the real listening object, features are extracted using a face reconstruction model to determine real face features.

[0109] In some embodiments of the present invention, the real facial features include real posture features and real expression features; In some embodiments of the present invention, the terminal can perform feature extraction on the facial image of the real listening object through a facial reconstruction model to obtain real posture features and real expression features, and the real posture features and real expression features are used as real facial features, wherein the number of facial images of the real listening object and the number of anthropomorphic features of multiple frames obtained by audio-video sequence samples are consistent, the real posture features include real posture features of multiple frames, the real expression features include real expression features of multiple frames, and the real facial features also include real facial features of multiple frames, and the number of real facial features is consistent with the number of predicted facial features.

[0110] In S109, the initial encoder is continuously optimized by the first loss function and the anthropomorphic sample features until the first loss function value satisfies a first preset threshold, and an encoder is determined.

[0111] In some embodiments of the present invention, the terminal continues to optimize the initial encoder using the first loss function and the anthropomorphic sample features, and if the first loss function value is greater than or equal to the first preset threshold, the terminal may consider that the encoder has been trained and determine the encoder; if the first loss function value is less than the first preset threshold, the terminal may continue to train the encoder until the first loss function value is greater than or equal to the first preset threshold, and determine the encoder.

[0112] In S1010, the initial virtual predictive network is continued to be optimized by the second loss function and the third loss function based on the actual facial features and the predicted facial features until both the second loss function value and the third loss function value satisfy the second preset threshold, and a virtual predictive network is determined.

[0113] In some embodiments of the present invention, the terminal may continue to optimize the initial virtual prediction network using the actual facial features and the predicted facial features with the second loss function and the third loss function, and determine the virtual prediction network if the sum of the second loss function value and the third loss function value is greater than or equal to a second preset threshold; and continue training the virtual prediction network until the sum of the second loss function value and the third loss function value is greater than or equal to the second preset threshold, and determine the virtual prediction network if the sum of the second loss function value and the third loss function value is less than the second preset threshold.

[0114] 8a and 8b are, respectively, FIG. 1 and FIG. 2, which show the results of a virtual object video sequence obtained by a video generation method according to an embodiment of the present invention. As shown in FIGS. 8a and 8b, the abscissa indicates a series of video frames, including frames 0 to 32. FIG. 8a shows the results of testing within the domain of the generated virtual object video sequence results (referring to the training data including the speaker or listener in the test data set), while FIG. 8b shows the results of testing outside the domain of the generated virtual object video sequence results (referring to the fact that neither the speaker nor the listener's facial data has appeared in the training set, primarily to verify the model's generalization ability to unseen faces). In FIG. 8a, the real listener's attitude is positive, while in FIG. 8b, the real listener's attitude is neutral. The fourth, fifth, and sixth lines in Figure 8a and the fourth, fifth, and sixth lines in Figure 8b show the results of the video sequences of the virtual object generated under three different attitudes, respectively. The generated listener in Figures 8a and 8b is the virtual object. The 0th frame is the reference frame. The different triangles in the figures indicate significant changes, among which the triangles in the upper right corner indicate significant changes.

number

number

number

number

[0115] The virtual prediction network can capture the patterns (e.g., eye, mouth, and head movements) of common listeners (corresponding to virtual objects), and although these patterns may differ from those of real listeners, they are still meaningful. The virtual prediction network can also present the visual patterns of virtual objects with different attitudes. From the results shown in Figure 8a, we can see that even those with a neutral attitude smile (frames 2-8), but this smile lasts for a shorter period than those with an active attitude (frames 2-16). For those with a passive attitude, the virtual objects do not pay attention to the content of the conversation, and instead direct their eyes to the bottom of the screen in frames 10, 16, 22, and 30.

[0116] Finally, Figure 8b shows that the out-of-domain data also has relatively good generation results. The virtual object with a positive attitude smiles in frames 6-14, while the audience with a negative attitude frowns, showing negative mouth shapes throughout the process. The virtual object with a negative attitude has small changes in movement and a floating gaze, while the neutral virtual object maintains a relatively calm expression and regular head movements.

[0117] We evaluate the results on video sequences of the generated virtual objects with 10 volunteers in two tests:

[0118] We conduct a best-matching test: Given an attitude, an audio sequence of a real object, a video sequence of the real object, a video of a real listener, and a video sequence of the generated virtual object, volunteers must choose the listener that is most perceptually appropriate and best matches the given attitude.

[0119] We conduct an attitude classification test: given a video sequence of a generated virtual object, volunteers must determine its mood (positive, negative, or neutral), where neutral corresponds to a general attitude.

[0120] Both tests were conducted in a double-blind format, and the results are shown in Tables 1 and 2.

[0121] [Table 1]

[0122] Table 1 presents the mean and variance of the numbers of the two types of "best listeners." For the in-domain data, nearly 20% of the generated virtual objects were deemed more reasonable than the real listeners by volunteers, which verifies that the model can generate responsive listeners that match human subjective perception. Furthermore, the results generated for the out-of-domain data were preferred by more volunteers.

[0123] [Table 2]

[0124] As can be seen from Table 2, for each attitude, the model obtained by calculating the mean and variance of the classification accuracy of all volunteers can generate videos of a given attitude to some extent.

[0125] It can be understood that the terminal can ensure the diversity of training samples by collecting audio-video sequence samples of real speaking objects and collecting corresponding facial images of real listening objects with different second features, and can optimize training for the initial encoder and initial virtual prediction network using the audio-video sequence samples and the corresponding facial images of the real listening objects, and determine the encoder and virtual prediction network, thereby improving the accuracy of the results output by the encoder and virtual prediction network.

[0126] In some embodiments of the present invention, Figure 9 is one optional flow schematic diagram 6 of a video generation method according to an embodiment of the present invention, and as shown in Figure 9, before executing S1010, further executes S1011 to S1013 as follows:

[0127] In S1011, a second loss function is determined based on the actual facial features and the predicted facial features.

[0128] In some embodiments of the present invention, a second loss function is used to ensure that the predicted expressions and poses are similar to the real expressions and poses.

[0129] In some embodiments of the present invention, the terminal may determine the second loss function based on the actual facial features and the predicted facial features by performing a subtraction to obtain a norm.

[0130] For example, the second loss function is obtained by the following equation (4).

number

number

number

number

number

[0131] In S1012, a third loss function is determined based on the change function corresponding to the actual facial features and the change function corresponding to the predicted facial features.

[0132] In some embodiments of the present invention, a third loss function is used to ensure that the frame-to-frame continuity of the predicted facial features resembles the real facial features.

[0133] In some embodiments of the present invention, the terminal may determine the third loss function by performing a subtraction operation to obtain a norm using a change function corresponding to the actual facial feature and a change function corresponding to the predicted facial feature.

[0134] For example, the third loss function is obtained by the following equation (5).

number

number

number

number

number

[0135] In S1013, the initial virtual prediction network is continuously optimized using the second loss function and the third loss function until the sum of the second loss function value and the third loss function value satisfies a second preset threshold, thereby determining the virtual prediction network.

[0136] In some embodiments of the present invention, the terminal can determine the loss function of the virtual prediction network by adding the second loss function and the third loss function, and continue to optimize the initial virtual prediction network using the loss function until the loss function value (corresponding to the sum of the second loss function value and the third loss function value) satisfies a second preset threshold, thereby determining the virtual prediction network.

[0137] For example, the loss function of the virtual prediction network is given by the following equation (6):

number

[0138] It can be understood that the terminal determines a second loss function based on the actual facial features and the predicted facial features, and determines a third loss function using a change function corresponding to the actual facial features and a change function corresponding to the predicted facial features, thereby improving the effectiveness of the loss function of the virtual prediction network, and continues to optimize the initial virtual prediction network using the second loss function and the third loss function until the sum of the second loss function value and the third loss function value meets the second preset threshold, and by determining the virtual prediction network, the accuracy of the virtual prediction network and the prediction effect of the virtual prediction network can be improved.

[0139] An exemplary application of an embodiment of the present invention in one practical application scenario will be described below.

[0140] In some embodiments of the present invention, FIG. 10 is an alternative model architecture diagram of a video generation method according to an embodiment of the present invention. As shown in FIG. 10, the terminal

number

number

number

number

number

number

[0141] Illustratively, for a speaker encoder, at each time step t, we first generate audio features S t and speaker audio features m S t and then extract one multimodal feature fusion function f am We use this to perform nonlinear feature transformation and obtain anthropomorphic features.

[0142] To ensure that the virtual object responds with a certain attitude and can produce more natural head movements and facial expression changes, the attitude e and the features m of the listener's reference image are l Let 1 (corresponding to the first posture / expression feature) be the first frame of the video sequence of the virtual object. Then, at each time step t, the speaker's fusion feature f am (s t ,m S t ) (corresponding to the anthropomorphic features) is used as input to generate a predicted video frame for step t+1. Finally, the predicted video frame is decoded using the listener decoder as m l t+1 It contains two feature vectors, namely, β l t+1 indicates facial expression, and p l t+1 indicates the pose (rotation and translation). The terminal supports speaker input of any length. The flow can be expressed by equation (7) as follows:

number

[0143] It can be understood that the terminal can process the captured audio-video sequence of the real object with a speaker encoder and a listener decoder to generate a video sequence of the virtual object so that the video sequence of the virtual object becomes more vivid and accurate.

[0144] Based on the video generating method of the above embodiment, the embodiment of the present invention further provides a video generating device as shown in FIG. 11, which is a structural schematic diagram 1 of video generating according to the embodiment of the present invention, and the device 11 includes: an acquiring part 1101, a determining part 1102, and a generating part 1103; said acquisition part 1101 being adapted to capture an audio-video sequence of a real object; the determining portion 1102 is configured to perform feature extraction on the audio-video sequence to determine anthropomorphic features; The generating part 1103 is configured to perform prediction on the anthropomorphic features using a virtual prediction network, pre-set standard features which are features corresponding to a reference object, and first features which represent different attitudes, generate a video sequence of a virtual object which is a video sequence in which a corresponding reaction of a virtual object is generated based on an audio-video sequence of a real object, and present the video sequence of the virtual object.

[0145] In some embodiments of the present invention, the acquiring unit 1101 is configured to acquire preset standard features including a first posture-expression feature and a first identity feature; the determining unit 1102 is configured to perform prediction and decoding using the virtual prediction network based on the first posture-expression feature, the anthropomorphic feature, and the first feature to determine posture-expression features of multiple frames of a virtual object; The generating part 1103 is configured to generate a video sequence of the virtual object based on the posture and expression features of multiple frames of the virtual object and the first identity feature.

[0146] In some embodiments of the present invention, the acquisition unit 1101 is configured to acquire a standard image representing an image of a reference object, and perform feature extraction on the standard image using a face reconstruction model to obtain the predetermined standard features.

[0147] In some embodiments of the present invention, the anthropomorphic features include anthropomorphic features of corresponding multiple frames of the audio-video sequence, and the virtual prediction network includes a first processing module and a second processing module. The determination unit 1102 is configured to perform prediction using the first processing module based on the anthropomorphic features of a first frame among the multiple frames, the first posture-expression feature, and the first feature, which is one of a positive attitude, a negative attitude, and a general attitude, to obtain a next predicted video frame; decode the next predicted video frame using the second processing module to determine a next posture-expression feature of the virtual object corresponding to the next predicted video frame; and continue predicting and decoding based on the next posture-expression feature and the anthropomorphic features of the next frame among the multiple frames of anthropomorphic features until a last posture-expression feature of the virtual object corresponding to the last predicted video frame is obtained, thereby obtaining posture-expression features of multiple frames of the virtual object, and the first posture-expression feature is set to the first frame among the multiple frames of posture-expression features.

[0148] In some embodiments of the present invention, the acquiring unit 1101 is configured to combine each of the posture and expression features of each frame among the posture and expression features of the plurality of frames with the first identity features including a first identity indicator, a first material, and first light irradiation information, and obtain a plurality of second features representing the combination results of the identity features and the posture and expression features; The generating portion 1103 is configured to generate, for the plurality of second characteristics, a video sequence of the virtual object corresponding to the virtual object by a renderer.

[0149] In some embodiments of the present invention, the audio-video sequence comprises an audio sequence of a real object and a video sequence of a real object, and the acquisition portion 1101 is configured to perform feature extraction on the video sequence of the real object by an encoder to obtain a plurality of video features, and to perform feature extraction on the audio sequence of the real object by an encoder to obtain a plurality of audio features including loudness, zero-crossing rate and cepstral coefficients; The determining portion 1102 is configured to perform feature transformation using a feature fusion function based on the plurality of video features and the plurality of audio features to determine anthropomorphic features of corresponding frames of an audio-video sequence, the anthropomorphic features including video features and audio features.

[0150] In some embodiments of the present invention, the acquiring unit 1101 is configured to perform feature extraction for each video frame of the video sequence of the real object using a face reconstruction model, and obtain the plurality of video frame features including second identity features and second pose-expression features; The determining portion 1102 is configured to take all the corresponding second pose-expression features in a video sequence of the real object as the video features.

[0151] In some embodiments of the present invention, the acquisition part 1101 is configured to acquire audio-video sequence samples of real talking objects and corresponding face images of real listening objects; the determining portion 1102 is configured to perform feature extraction on the audio-video sequence samples by an initial encoder to determine anthropomorphic sample features; the generating unit 1103 is configured to generate predicted facial features for a sample audio-video sequence of a real object to be trained, including predicted pose features and predicted facial expression features, using the initial virtual predictive network and the anthropomorphic sample features; The determination unit 1102 is configured to: extract features using a face reconstruction model based on a face image of the real listening object, determine real face features including real posture features and real expression features; continue to optimize the initial encoder using a first loss function and the anthropomorphic sample features until a first loss function value satisfies a first preset threshold, thereby determining the encoder; continue to optimize an initial virtual prediction network using a second loss function and a third loss function based on the real face features and predicted face features until a sum of a second loss function value and a third loss function value satisfies a second preset threshold, thereby determining the virtual prediction network.

[0152] In some embodiments of the present invention, the determining unit 1102 is configured to determine a second loss function based on the real facial features and the predicted facial features to ensure that the predicted expression and posture are similar to the real expression and posture, determine a third loss function based on a change function corresponding to the real facial features and a change function corresponding to the predicted facial features to ensure that the inter-frame continuity of the predicted facial features is similar to the real facial features, and continue to optimize the initial virtual prediction network using the second loss function and the third loss function until the sum of the second loss function value and the third loss function value satisfies a second preset threshold, and determine the virtual prediction network.

[0153] Although the above-described division of each program module is merely an example of video generation, in actual application, the above processes can be allocated and completed by different program modules according to needs, i.e., the internal structure of the device can be divided into different program modules to complete all or part of the above-described processes. Furthermore, the video generation device according to the above-described embodiment belongs to the same concept as the video generation method embodiment, and its specific implementation process and beneficial effects are detailed in the method embodiment, and will not be described again here. For technical details not disclosed in the device embodiment, please refer to the description of the method embodiment of the present invention.

[0154] Based on the video generation method of the above embodiment, an embodiment of the present invention further provides a video generation device as shown in FIG. 12. FIG. 12 is a structural schematic diagram 2 of a video generation device according to an embodiment of the present invention, where the device 12 comprises a processor 1201 and a memory 1202. The memory 1202 stores one or more programs executable by the processor. When the one or more programs are executed, the processor 1201 performs any one of the video generation methods of the above embodiment.

[0155] Those skilled in the art will appreciate that embodiments of the present invention may be provided as a method, a system, or a computer program product. Therefore, the present invention may take the form of a hardware embodiment, a software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product embodied in a computer-usable storage medium (including, but not limited to, a magnetic disk memory, an optical memory, etc.) having one or more computer-usable program codes thereon.

[0156] The present invention will be described with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, whereby the instructions executed by the processor of the computer or other programmable data processing device generate an apparatus for implementing the function(s) specified in one or more flows in the flowcharts and / or one or more blocks in the block diagrams.

[0157] These computer program instructions may be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture that includes an instruction apparatus that implements the functions specified in one or more flows in the flow diagrams and / or one or more blocks in the block diagrams.

[0158] These computer program instructions may be loaded onto a computer or other programmable data processing device, causing the computer or other programmable device to perform a series of operational steps to generate a computer-implemented process, where the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more flows in the flow diagrams and / or one or more blocks in the block diagrams.

[0159] The above are only preferred embodiments of the present invention, and are not intended to limit the protection scope of the present invention. [Industrial Applicability]

[0160]

[0013] Embodiments of the present invention provide a video generation method, apparatus, and computer-readable storage medium, in which the method includes: collecting an audio-video sequence of a real object; performing feature extraction on the audio-video sequence to determine anthropomorphic features; performing prediction on the anthropomorphic features using a virtual prediction network, predetermined standard features corresponding to features of a reference object, and first features representing different attitudes; generating a video sequence of a virtual object, which is a video sequence of a virtual object corresponding to the audio-video sequence of the real object; and presenting the video sequence of the virtual object. The embodiments of the present invention generate a video sequence of a virtual object based on the audio-video sequence of the real object, thereby making the presented video sequence of the virtual object more realistic and accurate.

Claims

1. Acquiring an audio-video sequence of a real object; performing feature extraction on the audio-video sequence to determine anthropomorphic features; performing predictions on the anthropomorphic features using a virtual prediction network, predetermined standard features corresponding to the reference object, and first features representing different attitudes, and generating a video sequence of a virtual object, which is a video sequence in which a corresponding reaction of the virtual object is generated based on an audio-video sequence of the real object; presenting a video sequence of the virtual object; the audio-video sequence includes an audio sequence of a real object and a video sequence of a real object; performing feature extraction on the audio-video sequence to determine anthropomorphic features, performing feature extraction on a video sequence of the real object by an encoder to obtain a plurality of video features; performing feature extraction on the audio sequence of the real object by the encoder to obtain a plurality of audio features including loudness, zero-crossing rate, and cepstral coefficients; and performing feature transformation according to the plurality of video features and the plurality of audio features using a feature fusion function to determine anthropomorphic features for corresponding frames of an audio-video sequence, the anthropomorphic features including video features and audio features. Video generation method.

2. generating a video sequence of a virtual object by performing predictions on the anthropomorphic features using a virtual prediction network, predetermined standard features, and a first feature; Acquiring preset standard features including a first posture / expression feature and a first identity feature; determining posture and expression features of a virtual object for a plurality of frames by performing prediction and decoding using the virtual prediction network based on the first posture and expression features, the anthropomorphic features, and the first features; generating a video sequence of the virtual object based on the posture-and-expression features of the virtual object and the first identity features of the plurality of frames; The method of claim 1.

3. Obtaining preset standard features is obtaining a standard image representing an image of a reference object; extracting features from the standard image using a face reconstruction model to obtain the predetermined standard features; The method of claim 2.

4. the anthropomorphic features include anthropomorphic features for corresponding frames of the audio-video sequence, the virtual prediction network comprising a first processing module and a second processing module; determining posture and expression features of a plurality of frames of a virtual object by predicting and decoding using the virtual prediction network based on the first posture and expression features, the anthropomorphic features, and the first features; According to the anthropomorphic features of a first frame among the plurality of frames, the first posture-expression feature, and the first feature, which is one of a positive attitude, a negative attitude, and a general attitude, performing prediction by the first processing module to obtain a next predicted video frame; decoding the next predicted video frame by the second processing module and determining a next pose / expression feature of the virtual object corresponding to the next predicted video frame; continuing to predict and decode based on the next posture and expression feature and the anthropomorphic feature of a next frame among the anthropomorphic features of the plurality of frames until obtaining a last posture and expression feature of the virtual object corresponding to a last predicted video frame, thereby obtaining posture and expression features of the plurality of frames of the virtual object; the first posture / expression feature is a first frame among the posture / expression features of the plurality of frames; The method of claim 2.

5. generating a video sequence of the virtual object based on the posture and expression features of a plurality of frames of the virtual object and the first identity feature, fusing each of the posture and expression features of each frame among the posture and expression features of the plurality of frames with the first identity feature including a first identity mark, a first material, and first light irradiation information, to obtain a plurality of second features representing a fusion result of the first identity feature and the posture and expression features of the plurality of frames; generating, by a renderer, a video sequence of the virtual object for the plurality of second characteristics; The method of claim 2.

6. performing feature extraction on a video sequence of the real object by an encoder to obtain a plurality of video features; extracting features for each video frame of the video sequence of the real object using a face reconstruction model to obtain a plurality of video frame features including second identity features and second pose-expression features; and determining all corresponding second pose-expression features in a video sequence of the real object as the video features. The method of claim 1.

7. and performing predictions on the anthropomorphic features using a virtual prediction network, predetermined standard features, and a first feature to generate a video sequence of the virtual object. acquiring audio-video sequence samples of real talking objects and corresponding facial images of real listening objects; performing feature extraction on the audio-video sequence samples by an initial encoder to determine anthropomorphic sample features; generating predicted facial features in the audio-video sequence samples of the training real object, including predicted pose features and predicted facial expression features, using the initial virtual prediction network and the anthropomorphic sample features; Extracting features from the face image of the real listening object using a face reconstruction model to determine real face features including real posture features and real expression features; continue optimizing the initial encoder with a first loss function and the anthropomorphic sample features until a first loss function value satisfies a first preset threshold, and determine the encoder; and continuing to optimize the initial virtual predictive network by the second loss function and the third loss function based on the actual facial features and the predicted facial features until both the second loss function value and the third loss function value satisfy a second preset threshold, thereby determining the virtual predictive network. The method of claim 1.

8. Continuing to optimize the initial virtual predictive network by the second loss function and the third loss function based on the actual facial features and the predicted facial features until the second loss function value and the third loss function value satisfy a second preset threshold, and determining the virtual predictive network, determining a second loss function based on the actual facial features and the predicted facial features to ensure that the predicted facial expression and predicted pose are similar to the actual facial expression and actual pose; determining a third loss function based on the change function corresponding to the actual facial feature and the change function corresponding to the predicted facial feature to ensure that the frame-to-frame continuity of the predicted facial feature resembles the actual facial feature; and continuing to optimize the initial virtual prediction network using the second loss function and the third loss function until the second loss function value and the third loss function value satisfy a second preset threshold, thereby determining the virtual prediction network. The method of claim 7.

9. an acquisition part configured to capture an audio-video sequence of a real object; a determining portion configured to perform feature extraction on the audio-video sequence to determine anthropomorphic features; a generating unit configured to perform predictions on the anthropomorphic features using a virtual prediction network, predetermined standard features, which are features corresponding to a reference object, and first features, which represent different attitudes, to generate a video sequence of a virtual object, which is a video sequence in which a corresponding reaction of a virtual object is generated based on an audio-video sequence of a real object, and to present the video sequence of the virtual object; the audio-video sequence includes an audio sequence of a real object and a video sequence of a real object; the obtaining portion is further configured to perform feature extraction on a video sequence of the real object by an encoder to obtain a plurality of video features, and to perform feature extraction on an audio sequence of the real object by the encoder to obtain a plurality of audio features including loudness, zero-crossing rate, and cepstral coefficients; The determining unit is further configured to perform feature transformation according to the plurality of video features and the plurality of audio features using a feature fusion function to determine anthropomorphic features for corresponding frames of an audio-video sequence, the anthropomorphic features including video features and audio features. Video generation device.

10. The acquisition unit is further configured to acquire preset standard features including a first posture / expression feature and a first identity feature; the determining unit is further configured to determine posture-expression features of a plurality of frames of a virtual object by predicting and decoding using the virtual prediction network based on the first posture-expression feature, the anthropomorphic feature, and the first feature; The generating unit is further configured to generate a video sequence of the virtual object based on posture-expression features of a plurality of frames of the virtual object and the first identity feature.

10. The apparatus of claim 9.

11. The acquisition unit is further configured to acquire a standard image representing an image of a reference object, and perform feature extraction on the standard image using a face reconstruction model to obtain the predetermined standard features.

11. The apparatus of claim 10.

12. the anthropomorphic features include anthropomorphic features for corresponding frames of the audio-video sequence, the virtual prediction network comprising a first processing module and a second processing module; The acquisition unit is further configured to perform prediction by the first processing module based on the anthropomorphic features of a first frame among the anthropomorphic features of the plurality of frames, the first posture-expression feature, and the first feature, which is one of a positive attitude, a negative attitude, and a general attitude, to obtain a next predicted video frame; the determining unit is further configured to decode the next predicted video frame by the second processing module, and determine a next posture / expression feature of the virtual object corresponding to the next predicted video frame; the acquisition unit is further configured to continue predicting and decoding based on the next posture-expression feature and the anthropomorphic feature of a next frame among the anthropomorphic features of the plurality of frames until obtaining a last posture-expression feature of the virtual object corresponding to a last predicted video frame, thereby obtaining posture-expression features of the plurality of frames of the virtual object; the first posture / expression feature is a first frame among the posture / expression features of the plurality of frames; 11. The apparatus of claim 10.

13. the acquiring unit is further configured to combine each of the posture and expression features of each frame among the posture and expression features of the plurality of frames with the first identity feature including a first identity mark, a first material, and first light irradiation information, and obtain a plurality of second features representing a combination result of the first identity feature and the posture and expression features of the plurality of frames; the generating portion is further configured to generate, with a renderer, a video sequence of the virtual object for the plurality of second characteristics.

11. The apparatus of claim 10.

14. the acquiring unit is further configured to perform feature extraction for each video frame of the video sequence of the real object using a face reconstruction model to obtain a plurality of video frame features including second identity features and second pose-expression features; the determining unit is further configured to determine all corresponding second pose-expression features in a video sequence of the real object as the video features.

10. The apparatus of claim 9.

15. the acquiring part is further configured to acquire audio-video sequence samples of real talking objects and corresponding facial images of real listening objects; the determining portion is further configured to perform feature extraction on the audio-video sequence samples by an initial encoder to determine anthropomorphic sample features; the generating unit is further configured to generate predicted facial features in the audio-video sequence sample of the real object to be trained, including predicted pose features and predicted facial expression features, using the initial virtual predictive network and the anthropomorphic sample features; The determining unit is further configured to: extract features by a face reconstruction model based on a face image of the real listening object, determine real face features including real posture features and real expression features; continue to optimize the initial encoder by a first loss function and the anthropomorphic sample features until a first loss function value satisfies a first preset threshold, thereby determining the encoder; continue to optimize an initial virtual prediction network by a second loss function and a third loss function based on the real face features and predicted face features until a second loss function value and a third loss function value both satisfy a second preset threshold, thereby determining the virtual prediction network. An apparatus according to any one of claims 9 to 13.

16. The determining unit is further configured to determine a second loss function based on the real facial features and the predicted facial features to ensure that the predicted facial expression and the predicted posture resemble the real facial expression and the real posture, determine a third loss function based on a change function corresponding to the real facial features and a change function corresponding to the predicted facial features to ensure that the inter-frame continuity of the predicted facial features resembles the real facial features, and continue to optimize the initial virtual prediction network by the second loss function and the third loss function until the second loss function value and the third loss function value satisfy a second preset threshold, and determine the virtual prediction network.

16. The apparatus of claim 15.

17. a memory for storing executable instructions; a processor for implementing the video generation method of any one of claims 1 to 8 when executing executable instructions stored in said memory. Video generation device.

18. Executable instructions stored thereon are used, when executed, to cause a processor to perform the video generation method of any one of claims 1 to 8. A computer-readable storage medium.

Citation Information

Patent Citations

  • Head posture estimation device, head posture estimation method and program for making computer execute head posture estimation method

    JP2014093006A

  • Nonverbal information generation device, nonverbal information generation model learning device, method, and program

    WO2019160100A1