Video generation method, apparatus, and computer-readable storage medium

The video generation method addresses the challenge of creating vivid and accurate video sequences for virtual objects by using a virtual prediction network to process audio-video sequences of real objects, resulting in enhanced human-machine interaction capabilities.

JP2025516979AActive Publication Date: 2025-05-30BEIJING JINGDONG SHANGKE INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024569572
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-07-14
Filing Date
2023-02-13
Publication Date
2025-05-30
Estimated Expiration
2043-02-13

AI Technical Summary

Technical Problem

Existing methods for generating video sequences of virtual objects struggle to achieve vividness and accuracy, particularly in human-machine interaction scenarios, as they primarily rely on parameterizing speakers using face key points and 3D models without effectively capturing the nuances of real object interactions.

Method used

A video generation method that acquires an audio-video sequence of a real object, extracts anthropomorphic features, and uses a virtual prediction network along with preset standard features and attitude features to predict and generate a video sequence of a virtual object, thereby simulating a corresponding reaction to the real object's audio-video sequence.

Benefits of technology

The method enhances the vividness and accuracy of generated video sequences by effectively capturing and replicating the reactions and interactions of virtual objects based on real object audio-video sequences, improving human-machine interaction experiences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025516979000001_ABST
    Figure 2025516979000001_ABST
Patent Text Reader

Abstract

Embodiments of the present invention provide a video generation method, apparatus, and computer-readable storage medium. Among them, the method includes: collecting the audio-video sequence of a real object; extracting features from the audio-video sequence to determine anthropomorphic features; using a virtual prediction network, a preset standard feature that is a feature corresponding to a reference object, and a first feature representing different attitudes to predict the anthropomorphic features; generating a video sequence of a virtual object, which is a video sequence in which a corresponding reaction of the virtual object is generated based on the audio-video sequence of the real object; and presenting the video sequence of the virtual object. Embodiments of the present invention generate a video sequence of a virtual object based on the audio-video sequence of a real object, thereby making the presented video sequence of the virtual object more vivid and accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] [Cross - reference to Related Applications] The present invention is filed based on a Chinese patent application with an application number of 202210834191.6 and an application date of July 14, 2022, claims the priority of the Chinese patent application, and here, all the contents of the Chinese patent application are incorporated into the present invention by reference.

[0002] The present invention relates to the field of human - machine interaction, and particularly to a video generation method, apparatus, and computer - readable storage medium.

Background Art

[0003] From the perspective of human kinetics, good communication means a two - way communication process, not a one - way information input or output. Along with the information interaction, the communication and interaction between real people are a process that continuously switches and circulates between two states of listening and speaking. Among them, listening and speaking are equally important. Both of them are indispensable for constructing an anthropomorphic digital human to perform human - machine interaction. While it is required that the digital human expresses its own perspective in a language that the other party can understand as clearly, concisely, and plainly as possible, the anthropomorphic digital human also needs to be good at listening to and understanding the perspectives of others. The prior art mainly generates corresponding speaker videos for the reference image of the speaker and the time - varying signal. Regarding the main methods, after parameterizing the speaker using face key points, face 3D models, human body skeleton models, etc., these parameters are fitted by a deep neural network, and the rendering image based on these parameters is used as the generation result, and the effect of the generated image is poor.

Summary of the Invention

Problems to be Solved by the Invention

[0004] Embodiments of the present invention provide a video generation method, apparatus, and computer-readable storage medium that can enhance the vividness and accuracy of a generated video sequence of a virtual object by generating the video sequence of the virtual object based on the audio-video sequence of a real object. **Means for Solving the Problem**

[0005] The technical aspect of the present invention is realized as follows.

[0006] Embodiments of the present invention acquire the audio-video sequence of a real object, extract features from the audio-video sequence and determine anthropomorphic features, use a virtual prediction network, a preset standard feature that is a feature corresponding to a reference object, and a first feature representing different attitudes to perform a prediction on the anthropomorphic features, and generate a video sequence of a virtual object that is a video sequence in which a corresponding reaction of the virtual object is generated based on the audio-video sequence of the real object, and present the video sequence of the virtual object, and provide a video generation method including the above.

[0007] In the above aspect, the above-mentioned performing a prediction on the anthropomorphic features using the virtual prediction network, the preset standard feature, and the first feature to generate a video sequence of a virtual object includes: acquiring a preset standard feature including a first posture-expression feature and a first identity feature, performing prediction and decoding by the virtual prediction network based on the first posture-expression feature, the anthropomorphic features, and the first feature to determine the posture-expression features of multiple frames of the virtual object, and generating a video sequence of the virtual object based on the posture-expression features of multiple frames of the virtual object and the first identity feature.

[0008] In the above aspect, obtaining the preset standard features described above includes: obtaining a standard image representing an image of a reference object; and extracting features from the standard image by a face reconstruction model to obtain the preset standard features.

[0009] In the above aspect, the anthropomorphic features include anthropomorphic features of a plurality of corresponding frames of the audio-video sequence. The virtual prediction network includes a first processing module and a second processing module. Based on the above-described first pose-expression feature, anthropomorphic features, and first feature, performing prediction and decoding by the virtual prediction network to determine pose-expression features of a plurality of frames of a virtual object includes: performing prediction by the first processing module based on the anthropomorphic feature of the first frame among the anthropomorphic features of the plurality of frames, the first pose-expression feature, and the first feature, which is one of a positive attitude, a negative attitude, and a general attitude, to obtain the next predicted video frame; decoding the next predicted video frame by the second processing module to determine the next pose-expression feature of the corresponding virtual object of the next predicted video frame; continuing to perform prediction and decoding based on the next pose-expression feature and the anthropomorphic feature of the next frame among the anthropomorphic features of the plurality of frames until obtaining the last pose-expression feature of the corresponding virtual object of the last predicted video frame, thereby obtaining pose-expression features of a plurality of frames of the virtual object. The first pose-expression feature is the first frame among the pose-expression features of the plurality of frames.

[0010] In the above aspect, generating a video sequence of the virtual object based on the pose-expression features of a plurality of frames of the virtual object and the first identity feature described above includes: Fusing each of the posture and expression features of each frame among the posture and expression features of the plurality of frames with the first identity feature including the first identity label, the first material, and the first light irradiation information, to obtain a plurality of second features representing the fusion result of the identity feature and the posture and expression feature; Generating, by a renderer, a video sequence of the virtual object for the plurality of second features.

[0011] In the above aspect, the audio-video sequence includes an audio sequence of a real object and a video sequence of the real object. Performing feature extraction on the above-described audio-video sequence to determine anthropomorphic features includes: Preprocessing, by an encoder, the video sequence of the real object to obtain a plurality of video features; Extracting features from the audio sequence of the real object by an encoder to obtain a plurality of audio features including loudness, zero-crossing rate, and cepstrum coefficient; Based on the plurality of video features and the plurality of audio features, performing feature conversion by a feature fusion function to determine anthropomorphic features of corresponding multiple frames of the audio-video sequence, the anthropomorphic features including video features and audio features.

[0012] In the above aspect, the above-described extracting features from the video sequence of the real object by an encoder to obtain a plurality of video features includes: Extracting features from each video frame of the video sequence of the real object by a face reconstruction model to obtain a plurality of video frame features including a second identity feature and a second posture and expression feature; Using all the corresponding second posture and expression features in the video sequence of the real object as the video features.

[0013] In the above aspect, before making a prediction on the anthropomorphic feature using the virtual prediction network, the preset standard feature, and the first feature described above, and generating a video sequence of the virtual object, the method includes: collecting an audio-video sequence sample of the real object and a face image of the corresponding real listening object; extracting features from the sample audio-video sequence by an initial encoder to determine anthropomorphic sample features; generating predicted face features in the audio-video sequence sample of the real object to be trained, including predicted pose features and predicted expression features, by an initial virtual prediction network and the anthropomorphic sample features; extracting features from the face image of the real listening object by a face reconstruction model based on the face image of the real listening object to determine real face features including real pose features and real expression features; continuing to optimize the initial encoder by a first loss function and the anthropomorphic sample features until the first loss function value satisfies a first preset threshold, and determining the encoder; continuing to optimize the initial virtual prediction network by a second loss function and a third loss function based on the real face features and the predicted face features until the second loss function value and the third loss function value satisfy a second preset threshold, and determining the virtual prediction network.

[0014] In the above aspect, the step of continuing to optimize the initial virtual prediction network by a second loss function and a third loss function based on the real face features and the predicted face features until the second loss function value and the third loss function value satisfy a second preset threshold, and determining the virtual prediction network includes: determining a second loss function for ensuring that the predicted expression and the predicted pose are similar to the real expression and the real pose based on the real face features and the predicted face features; determining a third loss function for ensuring that the continuity between frames of the predicted face features is similar to that of the real face features based on the change function corresponding to the real face features and the change function corresponding to the predicted face features. Continuously optimize the initial virtual prediction network using the second loss function and the third loss function until the second loss function value and the third loss function value satisfy a second preset threshold, and determine the virtual prediction network.

[0015] Embodiments of the present invention Comprise an acquisition part, a determination part, and a generation part. The acquisition part is configured to collect an audio-visual sequence of a real object. The determination part is configured to extract features from the audio-visual sequence and determine anthropomorphic features. The generation part uses a virtual prediction network, a preset standard feature that is a feature corresponding to a reference object, and a first feature representing different attitudes to perform a prediction on the anthropomorphic features, and generates a video sequence of a virtual object that is a video sequence in which a corresponding reaction of the virtual object is generated based on the audio-visual sequence of the real object, and is configured to present the video sequence of the virtual object. A video generation device is provided.

[0016] Embodiments of the present invention A memory for storing executable instructions, And a processor for executing the executable instructions stored in the memory. When the executable instructions are executed, the processor executes the video generation method described above. A video generation device is provided.

[0017] Embodiments of the present invention A computer-readable storage medium storing executable instructions that, when executed by one or more processors, cause the processor to execute the video generation method described above.

Advantages of the Invention

[0018] Embodiments of the present invention provide a video generation method, apparatus, and computer-readable storage medium. Among them, the method includes: capturing an audio-video sequence of a real object; extracting features from the audio-video sequence to determine anthropomorphic features; using a virtual prediction network, a preset standard feature that is a feature corresponding to a reference object, and a first feature representing different attitudes to perform a prediction on the anthropomorphic features, and generating a video sequence of a virtual object that is a video sequence in which a corresponding reaction of the virtual object is generated based on the audio-video sequence of the real object; and presenting the video sequence of the virtual object. Embodiments of the present invention generate a video sequence of a virtual object based on the audio-video sequence of a real object, thereby making the presented video sequence of the virtual object more vivid and accurate.

Brief Description of the Drawings

[0019]

Figure 1

Figure 2

Figure 3a

Figure 3b

Figure 3c

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8a

Figure 8b

Figure 9

Figure 10

Figure 11

Figure 12

DETAILED DESCRIPTION OF THE INVENTION

[0020] Hereinafter, with reference to the drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. It is obvious that the described embodiments are only a part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative labor belong to the protection scope of the present invention.

[0021] To enable those skilled in the art to better understand the embodiments of the present invention, the present invention will be described in more detail below with reference to the drawings and specific embodiments. FIG. 1 is a working schematic diagram of one selectable terminal of the video generation method according to an embodiment of the present invention. As shown in FIG. 1, the terminal is provided with a speaker encoder, a virtual prediction network, and a virtual human interface (not shown). The terminal extracts features from the audio-video of the real object by the speaker encoder (equivalent to the encoder), and inputs the extracted features, attitude (equivalent to the first feature), and reference image (equivalent to a preset standard image) to the listener decoder (equivalent to the virtual prediction network) for prediction, and generates the head movement and expression changes of the listener arranged on the timeline, and can obtain the video sequence of the virtual object.

[0022] In some embodiments of the present invention, FIG. 2 is a selectable flow schematic diagram 1 of the video generation method according to an embodiment of the present invention, and it will be described with reference to the steps shown in FIG. 2.

[0023] In S101, an audio-video sequence of a real object is collected.

[0024] In some embodiments of the present invention, according to the concepts of social psychology and anthropology, "listening" is also a functional behavior during communication. Among them, the listening behavior styles can be divided into four types: non-listener, marginal listener, evaluative listener, and active listener. Among them, actively responding listening is the most effective and also plays an important role in communication. It requires the listener to fully concentrate on what a person says, listen carefully, and show some visual reactions to the speaker. By feeding back to the speaker these reactions regarding whether the listener is interested, understands, or agrees with the content of the conversation, the rhythm and process of the conversation can be adjusted, and the smooth progress of communication can be promoted.

[0025] Regarding active listening responses, listeners often have common visual patterns when expressing their opinions. For example, symmetrically circulating movements are used to indicate "yes," "no," or similar signals, small linear movements are combined with emphasized syllables in the speaker's speech, and wider linear movements often appear during the pauses in the speaker's speech. In human face-to-face interactions, even the blinking time of the listener can be regarded as an interaction signal. Therefore, based on some audio-video sequences, it is of great significance to generate video sequences when listening to the audio-video sequences of virtual objects.

[0026] Exemplarily, under the condition that a reference image of the speaker and a time-varying signal of one section are given, a simulated segment that can match the time-varying signal of one speaker is generated. FIGS. 3a, 3b, and 3c are respectively one selectable speaker video generation schematic diagram 1, speaker video generation schematic diagram 2, and speaker video generation schematic diagram 3 of the video generation method according to an embodiment of the present invention. As shown in FIG. 3a, the speaker video generation task includes generating the body posture of the speaker. As shown in FIG. 3b, the speaker video generation task includes generating the lip movement of the speaker. As shown in FIG. 3c, the speaker video generation task includes generating the movement of the head (including the face) of the speaker. In FIG. 3a, the generation of the body posture of the speaker is to process a time-varying signal of one section input in a dashed frame by a body posture generation model to obtain the body posture indicated by a dashed-dotted frame. In FIG. 3b, the generation of the lip movement of the speaker is to process a time-varying signal of one section input in a dashed frame and a general reference image by a lip movement generation model to output an image frame of the lip movement of the speaker indicated by a dashed-dotted frame. In FIG. 3c, the generation of the movement of the head (including the face) of the speaker mainly processes a time-varying signal of one section input in a dashed frame, the reference image of the speaker, and the mood by a head movement generation model, renders the processed result by a head rendering model, and outputs an image frame of the movement of the head (including the face) of the speaker indicated by a dashed-dotted frame.

[0027] In some embodiments of the present invention, the terminal can collect the audio sequence and video sequence of the real object by a collection device.

[0028] Exemplarily, the collection device may be a device having functions of collecting video and audio such as a camera head, but the present invention is not limited thereto. The real object may be a person speaking in one scene. The audio sequence and video sequence may be obtained in the process that a tourist inquires a self-service information device (the carrier of the virtual object) in a scene at a tourist destination.

[0029] In some embodiments of the present invention, the present invention is applicable when human-machine interaction is required, such as an intelligent information device in a department store that can create a corresponding video based on the video exhibited by a shopper and guide the shopper.

[0030] In S102, extract features from the audio-video sequence and determine anthropomorphic features.

[0031] In some embodiments of the present invention, the audio-video sequence includes the audio sequence of the real object and the video sequence of the real object.

[0032] In some embodiments of the present invention, feature extraction can be realized by a neural network model. The process of feature extraction is to input each frame of the video sequence into the neural network, perform feature extraction by a plurality of convolutional layers and pooling layers, and obtain a plurality of video features. The anthropomorphic feature is a feature obtained by characterizing the limb movements that accompany a person's speech in a real scene. The anthropomorphic feature is a feature that includes audio characteristics and video characteristics, and is the audio feature and video feature obtained after performing feature extraction on the audio-video sequence. Exemplarily, in a certain video, if there is an action of raising the hand when the speaker is speaking, the anthropomorphic feature may be the feature corresponding to the action of raising the hand.

[0033] In some embodiments of the present invention, the terminal performs feature extraction on the video sequence of the real object by an encoder to obtain a plurality of video features, performs feature extraction on the audio sequence of the real object by the encoder, and obtains a plurality of audio features including loudness, zero-crossing rate, and cepstrum coefficient. Based on the plurality of video features and the plurality of audio features, feature conversion is performed by a feature fusion function to determine anthropomorphic features.

[0034] In some embodiments of the present invention, FIG. 4 is one selectable flow schematic diagram 2 of the video generation method according to the embodiment of the present invention. As shown in FIG. 4, S102 can be realized by S1021 to S1023 as follows.

[0035] In S1021, the encoder performs feature extraction on the video sequence of the real object to obtain a plurality of video features.

[0036] In some embodiments of the present invention, video features are features obtained by recording the rotation of the head and several facial expression changes during the process of people communicating, including the rotation of the head and facial expressions. Video features are features obtained after extracting features from a video sequence, and video features include postures and expressions.

[0037] In some embodiments of the present invention, the terminal extracts features from each video frame of the video sequence of the actual object by means of a face reconstruction model, obtains a plurality of video frame features, and can use all corresponding second posture-expression features in the video sequence of the actual object as video features.

[0038] In some embodiments of the present invention, FIG. 5 is one selectable flow schematic diagram 3 of the video generation method according to the embodiment of the present invention. As shown in FIG. 5, S1021 can be realized by S10211 and S10212 as follows.

[0039] In S10211, features are extracted from each video frame of the video sequence of the actual object by means of a face reconstruction model, and a plurality of video frame features are obtained.

[0040] In some embodiments of the present invention, video frame features are obtained by recording the rotation of the head, facial expressions, and a plurality of captured elements of the person in each video frame in the video sequence. Video frame features include a second identity feature and a second posture-expression feature. The second identity feature is the result of recording the environmental elements captured in the video sequence and the identity information of the object to be captured. The second identity feature includes the identity label of the actual object, the material, and the light irradiation. The second posture-expression feature is the result obtained by recording the rotation of the head and facial expression changes when the actual object is speaking. The second posture-expression feature includes the posture of the head and facial expressions of the actual object. Generally, a 3D face reconstruction model is selected for the face reconstruction model, but the present invention is not limited thereto.

[0041] In some embodiments of the present invention, the terminal can extract features for the person and the background in each video frame of the video sequence of the real object by using a face reconstruction model. In each video frame, the identity label of the person, the material of the video frame, the light irradiation during shooting, the posture of the person's head, and the expression of the person's face can all be obtained. The identity label of the person, the material of the video frame, and the light irradiation during shooting are used as the second identity features, and the posture of the person's head and the expression of the person's face are used as the second posture-expression features. All the features (i.e., the identity label of the person, the material of the video frame, the light irradiation during shooting, the posture of the person's head, and the expression of the person's face) are used as the features of one video frame.

[0042] Exemplarily, for the video sequence

Number

[0043] In S10212, all the corresponding second posture-expression features in the video sequence of the real object are used as video features.

[0044] In some embodiments of the present invention, the terminal determines all the corresponding second posture-expression features in the video sequence of the real object as the video features corresponding to the video sequence.

[0045] Exemplarily, the parameters {α, β, δ, p, γ} are divided into two types. One is the relatively fixed features closely related to the identity label information

Number

Number

[0046] In some embodiments of the present invention, the terminal extracts features for each video frame of the video sequence of the real object by means of a face reconstruction model, obtains a plurality of video frame features, and uses all corresponding second posture expression features in the video sequence of the real object as video features, thereby removing identity-specific features (such as the identity marking of a person, the material of the video frame, and the light irradiation during shooting), leaving only common features, enhancing the effectiveness of feature extraction, and providing data support for the subsequent generation of the video sequence of the virtual object, which can be understood.

[0047] In S1022, the encoder extracts features for the audio sequence of the real object to obtain a plurality of audio features.

[0048] In some embodiments of the present invention, the audio features are some features accompanying the speaker during the human communication process. The audio features are the features obtained after extracting features for the audio sequence, and the audio features include the loudness, zero-crossing rate, and cepstrum coefficient corresponding to the time when the real object is speaking.

[0049] In some embodiments of the present invention, the terminal extracts features from the spectrum of the audio sequence of the real object by an encoder, obtains the loudness, zero-crossing rate, and cepstrum coefficient at each moment when the real object is speaking, regards all of the loudness, zero-crossing rate, and cepstrum coefficient as audio features, and can obtain a plurality of audio features.

[0050] Exemplarily, for the audio sequence of the real object

Number

Number

Number

[0051] Based on the audio-video sequence generated when the real object is speaking, the terminal can perform feature extraction and obtain corresponding multiple audio features and multiple video features. It can be understood that the terminal can quickly extract the effective features of audio and video, and improve the speed and effectiveness of feature extraction.

[0052] In S1023, based on a plurality of video features and a plurality of audio features, feature conversion is performed by a feature fusion function to determine anthropomorphic features.

[0053] In some embodiments of the present invention, the anthropomorphic feature is the anthropomorphic feature of a plurality of corresponding frames of an audio - video sequence. The anthropomorphic feature includes a video feature and an audio feature, and the feature fusion function can convert a non - linear feature into a linear feature.

[0054] In some embodiments of the present invention, the terminal performs a non - linear feature conversion on a plurality of audio features and a plurality of video features by means of a multi - modal feature fusion function in the encoder, obtains a feature representation of the real object, that is, determines the anthropomorphic feature.

[0055] It can be understood that by extracting features from the audio - video sequence, the terminal can increase the speed of processing the audio - video sequence, quickly obtain a plurality of video features and a plurality of audio features, and by converting a plurality of video features and a plurality of audio features into a first feature by the feature fusion function, the feature representation method can be converted, facilitating the terminal's processing of it, and enhancing the feasibility of processing video features and audio features.

[0056] In S103, a virtual prediction network, a preset standard feature, and the first feature are used to predict the anthropomorphic feature, and generate a video sequence of a virtual object.

[0057] In some embodiments of the present invention, the video sequence of the virtual object is a video sequence in which a corresponding reaction of the virtual object is generated based on the audio - video sequence of the real object. The preset standard feature is the feature corresponding to the reference object. The preset standard feature includes a first posture - expression feature and a first identity feature corresponding to the reference object. The first posture - expression feature is the head posture and facial expression of the reference object, and the first identity feature is information related to the identity corresponding to the reference object. The first feature is an attitude feature with an emotional color when a person is speaking. The first feature represents different attitudes, which may be a positive attitude, a negative attitude, or a general attitude.

[0058] In some embodiments of the present invention, the terminal acquires preset standard features, performs prediction and decoding by means of a virtual prediction network based on the first posture expression features, anthropomorphic features and the first features, determines the posture expression features of multiple frames of the virtual object, and can generate a video sequence of the virtual object according to the posture expression features of multiple frames of the virtual object and the first identity features.

[0059] Exemplarily, an input video sequence of a real object with a time length from 1 to t

Number

Number

Number

Number

[0060] In some embodiments of the present invention, FIG. 6 is one selectable flow schematic diagram 4 of the video generation method according to the embodiment of the present invention. As shown in FIG. 6, S103 can be realized by S1031 to S1033 as follows.

[0061] In S1031, preset standard features are acquired.

[0062] In some embodiments of the present invention, the terminal may obtain a standard image, extract features from the standard image by means of a face reconstruction model, obtain a first pose-expression feature and a first identity feature, and use the first pose-expression feature and the first identity feature as preset standard features.

[0063] In some embodiments of the present invention, S1031 can be realized by S10311 and S10312 as follows.

[0064] In S10311, obtain a standard image.

[0065] In some embodiments of the present invention, the standard image is an image of a reference object, and the standard image is an arbitrary face image different from the actual object randomly obtained from the image library.

[0066] In S10312, extract features from the standard image by means of a face reconstruction model to obtain preset standard features.

[0067] In some embodiments of the present invention, the terminal may extract features from the standard image by means of a face reconstruction model, and obtain the identity label of the person in the standard image, the material of the video frame, the light irradiation during shooting, the pose of the person's head, and the expression of the person's face. The identity label of the person in the standard image, the material of the video frame, and the light irradiation during shooting are used as the first identity feature, and the pose of the person's head and the expression of the person's face in the standard image are used as the first pose-expression feature. The first identity feature and the first pose-expression feature are used as preset standard features.

[0068] The terminal can obtain a standard image, and by extracting features from the standard image by means of a face reconstruction model, quickly and accurately obtain the first pose-expression feature, improve the accuracy and efficiency of the processing of the standard image. It can be understood that the first pose-expression feature can be used for generating the video sequence of the virtual object, and the accuracy of the generated video sequence of the virtual object is guaranteed.

[0069] In S1032, based on the first posture expression feature, anthropomorphic feature, and first feature, prediction and decoding are performed by a virtual prediction network to determine the posture expression features of multiple frames of the virtual object.

[0070] In some embodiments of the present invention, the terminal uses the anthropomorphic feature of the first frame among the anthropomorphic features of multiple frames, the first posture expression feature, and the first feature to perform prediction by the first processing module, obtains the next one predicted video frame, decodes the next one predicted video frame by the second processing module, determines the next one posture expression feature of the corresponding virtual object of the next one predicted video frame, and continues to perform prediction and decoding by the next one posture expression feature and the anthropomorphic feature of the next one frame among the anthropomorphic features of multiple frames until the last one posture expression feature of the corresponding virtual object of the last one predicted video frame is obtained, thereby obtaining the posture expression features of multiple frames of the virtual object.

[0071] In some embodiments of the present invention, S1032 can be realized by S10321, S10322, and S10323 as follows.

[0072] In S10321, based on the anthropomorphic feature of the first frame among the anthropomorphic features of multiple frames, the first posture expression feature, and the first feature, prediction is performed by the first processing module to obtain the next one predicted video frame.

[0073] In some embodiments of the present invention, the first feature is any one of a positive attitude, a negative attitude, and a general attitude, and the first processing module is an encoding module in the virtual prediction network, and its function is to realize encoding in prediction processing.

[0074] In some embodiments of the present invention, the terminal can select any one of a positive attitude, a negative attitude, and a general attitude as the first feature. In the embodiments of the present invention, the general attitude is selected. In the general attitude, the anthropomorphic feature of the first frame and the first posture-expression feature among the anthropomorphic features of multiple frames are input into the first processing module in the virtual prediction network for prediction, and the next one video prediction frame is obtained.

[0075] In S10322, the second processing module decodes the next one prediction video frame, and determines the next one posture-expression feature of the corresponding virtual object of the next one prediction video frame.

[0076] In some embodiments of the present invention, the second processing module is a decoding module in the virtual prediction network, and its function is to realize decoding in the prediction process.

[0077] In some embodiments of the present invention, the terminal can decode the next one prediction video frame by the second processing module in the virtual prediction network, and obtain the next one posture-expression feature of the corresponding virtual object of the next one prediction video frame. The next one posture-expression feature includes the next one posture feature and the next one expression feature.

[0078] In S10323, until the last one posture-expression feature of the corresponding virtual object of the last one prediction video frame is obtained, continue to perform prediction and decoding based on the next one posture-expression feature and the anthropomorphic feature of the next one frame among the anthropomorphic features of multiple frames, thereby obtaining the posture-expression features of multiple frames of the virtual object.

[0079] In some embodiments of the present invention, the first posture-expression feature is the first frame among the posture-expression features of multiple frames.

[0080] In some embodiments of the present invention, the terminal inputs the anthropomorphic feature of the second frame among the following one posture-expression feature and the anthropomorphic features of multiple frames into the first processing module in the virtual prediction network for prediction, obtains the following one predicted video frame, decodes it by the second processing module in the virtual prediction network, determines the following one posture-expression feature of the corresponding virtual object of the following one predicted video frame, and continues to perform prediction and decoding by the virtual prediction network until obtaining the last one posture-expression feature of the corresponding virtual object of the last one predicted video frame, thereby determining the posture-expression features of multiple frames of the virtual object.

[0081] The terminal can perform prediction based on the anthropomorphic feature of the first frame, the first posture-expression feature, and the first feature. The first feature can enhance the diversity of the generated predicted video frames. The anthropomorphic feature of the first frame and the first posture-expression feature can enhance the accuracy of the generated predicted video frames. Based on the first predicted video frame, the following one predicted video frame is generated, and by predicting the following one video frame with the real-time updated predicted video frame, continuous video frames can be generated, enhancing the continuity between the video frames and ensuring the integrity of the video sequence. By decoding multiple predicted video frames to obtain the posture-expression features of multiple frames, the head reaction and facial reaction of the virtual object when listening to the audio-video of the real object can be fully reflected, and it can be understood that the accuracy of the generated video sequence of the virtual object can be enhanced.

[0082] In S1033, based on the posture-expression features of multiple frames of the virtual object and the first identity feature, a video sequence of the virtual object is generated.

[0083] In some embodiments of the present invention, the virtual object is a simulated object with a simple communication function and generally exists in an interactive device.

[0084] In some embodiments of the present invention, the terminal can obtain a plurality of second features by fusing the posture and expression features of each frame among the posture and expression features of multiple frames of the virtual object with the first identity feature, and can generate a video sequence of the virtual object for the plurality of second features by a renderer.

[0085] The terminal can obtain a preset standard feature, perform prediction and decoding by a virtual prediction network based on the first posture and expression feature, anthropomorphic feature and first feature in the preset standard feature, determine the posture and expression features of multiple frames of the virtual object, improve the accuracy of the posture and expression features of multiple frames, and based on the posture and expression features of multiple frames of the virtual object and the first identity feature, generate a video sequence of the virtual object, so as to improve the accuracy and vividness of the video sequence of the virtual object, which can be understood.

[0086] In some embodiments of the present invention, S1033 can be realized by S10331 and S10332 as follows.

[0087] In S10331, fuse each of the posture and expression features of each frame among the posture and expression features of multiple frames of the virtual object with the first identity feature to obtain a plurality of second features.

[0088] In some embodiments of the present invention, the second feature is a fusion feature formed after assigning an identity feature to the posture and expression feature. The second feature is the fusion result of the identity feature and the posture and expression feature.

[0089] In some embodiments of the present invention, the terminal can fuse each of the posture and expression features of each frame among the posture and expression features of multiple frames of the virtual object with the first identity feature to obtain a plurality of second features corresponding to the posture and expression features of multiple frames.

[0090] In S10332, for a plurality of second features, a renderer generates a video sequence of virtual objects.

[0091] In some embodiments of the present invention, the first identity feature includes a first identity label, a first material, and first light irradiation information. The first identity label corresponds to the identity label of a person in a standard image. The first material corresponds to the material of a video frame. The first light irradiation information corresponds to the light irradiation during shooting.

[0092] In some embodiments of the present invention, the terminal can generate a video sequence of virtual objects by a renderer for a plurality of second features.

[0093] Exemplarily, the video sequence generation task of virtual objects can be represented by Formula (2) and Formula (3).

Number

Number

[0094] The terminal can obtain a plurality of second features by fusing the pose and expression features of each frame among the pose and expression features of multiple frames of the virtual object with the first identity feature. By endowing the second features with identity attributes, the distinctiveness of the second features can be enhanced. For the plurality of second features, a renderer generates a video sequence of the virtual object, and by focusing on generating the video sequence of the virtual object, it can be understood that the video sequence can be made more vivid and accurate.

[0095] In S104, a video sequence of the virtual object is presented.

[0096] In some embodiments of the present invention, the terminal can present the video sequence of the virtual object to the virtual human interface.

[0097] It can be understood that the terminal captures the audio - video sequence of the real object, extracts features from the audio - video sequence, determines anthropomorphic features, makes a prediction for the anthropomorphic features using a virtual prediction network, a preset standard feature which is the feature corresponding to the reference object, and a first feature representing different attitudes, generates a video sequence of the virtual object which is a video sequence in which a corresponding reaction of the virtual object is generated based on the audio - video sequence of the real object, and presents the video sequence of the virtual object. Based on the audio - video sequence of the real object, when generating the video sequence of the virtual object, the presented video sequence of the virtual object is more vivid and accurate.

[0098] In some embodiments of the present invention, FIG. 7 is one selectable flow schematic diagram 5 of the video generation method according to the embodiment of the present invention. As shown in FIG. 7, before executing S103, the following S105 - S1010 are further executed.

[0099] In S105, an audio - video sequence sample of the real - speaking object and a face image of the corresponding real - listening object are captured.

[0100] In some embodiments of the present invention, the audio - video sequence sample is formed by recording the speech of the real - speaking object and the corresponding actions and colors when speaking.

[0101] In some embodiments of the present invention, the terminal can collect the audio - video sequence sample of the speaking object and the face image of the corresponding listening object by a collection device. The collected face images of the listening objects are collected with different first features, and the first features include positive attitudes, negative attitudes, and general attitudes.

[0102] In S106, the initial encoder extracts features from the audio - video sequence sample to determine the anthropomorphic sample features.

[0103] In some embodiments of the present invention, the anthropomorphic sample features are features obtained by characterizing some limb movements when a person is speaking in a scene corresponding to the audio - video sequence sample. The anthropomorphic sample features are obtained after extracting features from the audio - video sequence sample, and include video sample features and audio sample features.

[0104] In some embodiments of the present invention, the terminal can extract features from the audio - video sequence sample by the initial encoder to obtain anthropomorphic sample features including anthropomorphic sample features of multiple frames.

[0105] In S107, the initial virtual prediction network and the anthropomorphic sample features are used to generate predicted face features in the sample audio - video sequence of the object to be trained.

[0106] In some embodiments of the present invention, the predicted face features include predicted posture features and predicted expression features.

[0107] In some embodiments of the present invention, the terminal can generate a predicted pose feature including predicted pose features of multiple frames in an audio-visual sequence sample of an object to be trained and a predicted expression feature including predicted expression features of multiple frames based on an initial virtual prediction network and anthropomorphic sample features, and use the predicted pose feature and the predicted expression feature as a predicted face feature including predicted face features of multiple frames.

[0108] In S108, based on the face image of the actual listening object, the face reconstruction model extracts features to determine the actual face features.

[0109] In some embodiments of the present invention, the actual face features include an actual pose feature and an actual expression feature. In some embodiments of the present invention, the terminal extracts features from the face image of the actual listening object by using the face reconstruction model to obtain an actual pose feature and an actual expression feature, and can use the actual pose feature and the actual expression feature as actual face features. Among them, the number of face images of the actual listening object is the same as the number of anthropomorphic features of multiple frames obtained from the audio-visual sequence sample. The actual pose feature includes actual pose features of multiple frames, the actual expression feature includes actual expression features of multiple frames, the actual face feature also includes actual face features of multiple frames, and the number of actual face features is the same as the number of predicted face features.

[0110] In S109, continue to optimize the initial encoder by using the first loss function and anthropomorphic sample features until the first loss function value satisfies the first preset threshold, and determine the encoder.

[0111] In some embodiments of the present invention, the terminal continues to optimize the initial encoder by using the first loss function and anthropomorphic sample features. When the first loss function value is greater than or equal to the first preset threshold, it is considered that the training of the encoder has been completed, and the encoder can be determined. When the first loss function value is less than the first preset threshold, continue to train the encoder until the first loss function value is greater than or equal to the first preset threshold, and the encoder can be determined.

[0112] In S1010, based on the actual face feature and the predicted face feature, continue to optimize the initial virtual prediction network with the second loss function and the third loss function until both the second loss function value and the third loss function value satisfy the second preset threshold, and determine the virtual prediction network.

[0113] In some embodiments of the present invention, the terminal continues to optimize the initial virtual prediction network with the second loss function and the third loss function using the actual face feature and the predicted face feature. When the sum of the second loss function value and the third loss function value is greater than or equal to the second preset threshold, determine the virtual prediction network. When the sum of the second loss function value and the third loss function value is less than the second preset threshold, continue training the virtual prediction network until the sum of the second loss function value and the third loss function value is greater than or equal to the second preset threshold, and the virtual prediction network can be determined.

[0114] Exemplarily, FIGS. 8a and 8b are respectively the result diagram 1 of the video sequence of the virtual object of the video generation method according to an embodiment of the present invention and the result diagram 2 of the video sequence of the virtual object of the video generation method according to an embodiment of the present invention. As shown in FIGS. 8a and 8b, the abscissa represents continuous video frames including 0 to 32 frames. In FIG. 8a, the results of testing within the region of the result of the generated video sequence of the virtual object (indicating that the training data includes the speaker or listener in the test data set) are shown. In FIG. 8b, the results of testing outside the region of the result of the generated video sequence of the virtual object (indicating that neither the face data of the speaker nor the listener has appeared in the training set, mainly verifying the generalization ability of the model with unseen faces) are shown. In FIG. 8a, the actual listener has a positive attitude, and in FIG. 8b, the actual listener has a natural attitude. In the 4th, 5th, and 6th rows in FIG. 8a, and the 4th, 5th, and 6th rows in FIG. 8b, the results of the video sequences of the virtual objects generated under three different attitudes are respectively shown. The generated listeners in FIGS. 8a and 8b are virtual objects. Among them, the 0th frame is the reference frame. Different triangles in the figure indicate significant changes. Among them, the

Number

Number

Number

Number

[0115] The virtual prediction network can capture the patterns of a common listener (equivalent to a virtual object), such as patterns of eyes, mouth, and head movements. Although these patterns may be different from those of a real listener, it can still be seen that they are still meaningful. Also, the virtual prediction network can present visual patterns of virtual objects with different attitudes. Regarding the results shown in Fig. 8a, it can be seen that even with a neutral attitude, it is smiling (frames 2 - 8), but the time it is maintained is shorter than that with a positive attitude (frames 2 - 16). Also, for a person with a negative attitude, the virtual object does not pay attention to the content of the conversation, and they turn their eyes towards the lower side of the screen at frames 10, 16, 22, and 30.

[0116] Finally, in Fig. 8b, there are also relatively good generation results for out - of - domain data. The virtual object with a positive attitude smiles from frames 6 - 14, and the audience with a negative attitude frowns and shows a negative mouth shape throughout the process. The virtual object in a negative attitude has small changes in movements and has a floating look in the eyes, while the neutral virtual object has regular head movements while maintaining a relatively calm expression.

[0117] The results of the video sequence of the generated virtual object are evaluated, and 10 volunteers conduct the following two tests.

[0118] Conduct an optimal matching test. When given an attitude, the audio sequence of the real object, the video sequence of the real object, the video of the real listener, and the video sequence of the generated virtual object, the volunteer needs to sensibly select the listener that is the most appropriate and most consistent with the given attitude.

[0119] Conduct an attitude classification test. When given the video sequence of the generated virtual object, the volunteer needs to determine its mood (positive, negative, natural). Note that natural corresponds to a general attitude.

[0120] Both of the two tests were conducted in a double-blind format, and the results are as shown in Table 1 and Table 2.

[0121]

Table 1

[0122] In Table 1, the average values and variances of the numbers of two types of "best listeners" are statistically calculated. In the in-region data, it is considered that, through the voting of volunteers, nearly 20% of the generated virtual objects look more reasonable than the real listeners, which verifies that the model can generate a response-style listener that is consistent with human subjective perception. Furthermore, the results generated from the out-of-region data are preferred by more volunteers.

[0123]

Table 2

[0124] As can be seen from Table 2, for each attitude, the model obtained by calculating the average values and variances of the classification accuracies of all volunteers can, to a certain extent, generate videos of a predetermined attitude.

[0125] The terminal can ensure the diversity of training samples by collecting the audio-video sequence samples of the actual talking objects and collecting the face images of the corresponding actual listening objects with different second features. It can be understood that by using the audio-video sequence samples and the corresponding face images of the actual listening objects to optimize the training of the initial encoder and the initial virtual prediction network and determine the encoder and the virtual prediction network, the accuracy of the results output from the encoder and the virtual prediction network can be improved.

[0126] In some embodiments of the present invention, FIG. 9 is one selectable flow schematic diagram 6 of the video generation method according to the embodiment of the present invention. As shown in FIG. 9, before the execution of S1010, the following S1011 to S1013 are further executed.

[0127] In S1011, based on the real face feature and the predicted face feature, a second loss function is determined.

[0128] In some embodiments of the present invention, the second loss function is used to ensure that the predicted expression and the predicted posture are similar to the real expression and the real posture.

[0129] In some embodiments of the present invention, the terminal can perform an operation of subtraction to obtain a norm based on the real face feature and the predicted face feature, and determine the second loss function.

[0130] Exemplarily, the second loss function is obtained by the following formula (4).

Number

Number

Number

Number

Number

[0131] In S1012, based on the change function corresponding to the real face feature and the change function corresponding to the predicted face feature, a third loss function is determined.

[0132] In some embodiments of the present invention, the third loss function is used to ensure that the continuity between frames of the predicted face feature is similar to the real face feature.

[0133] In some embodiments of the present invention, the terminal can perform an operation of subtracting and obtaining a norm by using the change function corresponding to the real face feature and the change function corresponding to the predicted face feature, and determine the third loss function.

[0134] Exemplarily, the third loss function is obtained by the following formula (5).

Number

Number

Number

Number

[0135] In S1013, continue to optimize the initial virtual prediction network with the second loss function and the third loss function until the sum of the second loss function value and the third loss function value satisfies the second preset threshold, and determine the virtual prediction network.

[0136] In some embodiments of the present invention, the terminal determines the loss function of the virtual prediction network by adding the second loss function and the third loss function, and continues to optimize the initial virtual prediction network with the loss function until the loss function value (equivalent to the sum of the second loss function value and the third loss function value) satisfies the second preset threshold, and the virtual prediction network can be determined.

[0137] Exemplarily, the loss function of the virtual prediction network is obtained by the following mathematical formula (6). [Number] Among them, L total represents the loss function of the virtual prediction network, L gen represents the second loss function, L mot represents the third loss function, and W is a scale for balancing these two loss functions.

[0138] The terminal determines a second loss function based on the real face features and the predicted face features, and determines a third loss function by using the change function corresponding to the real face features and the change function corresponding to the predicted face features, so as to enhance the effectiveness of the loss function of the virtual prediction network. The terminal continuously optimizes the initial virtual prediction network by using the second loss function and the third loss function until the sum of the second loss function value and the third loss function value satisfies a second preset threshold, and determines the virtual prediction network, thereby enhancing the accuracy of the virtual prediction network and the prediction effect of the virtual prediction network. This can be understood.

[0139] Next, an exemplary application in one actual application scenario of an embodiment of the present invention will be described.

[0140] In some embodiments of the present invention, FIG. 10 is a selectable model architecture diagram of a video generation method according to an embodiment of the present invention. As shown in FIG. 10, the terminal obtains a speaker video

Number

Number

Number

Number

Number

Number

[0141] Exemplarily, for the speaker encoder, at each time step t, first, the audio feature S t and the speaker's audio feature m S tExtract it and perform non-linear feature transformation using a single multi-modal feature fusion function f am to obtain anthropomorphic features.

[0142] To ensure that the virtual object can react with a certain attitude and produce more natural head movements and facial expression changes, the attitude e and the features m of the reference image of the listener l 1 (corresponding to the first pose expression feature) are used as the first frame of the video sequence of the virtual object. Then, at each time step t, the fusion feature f of the speaker am (s t ,m S t )(corresponding to the anthropomorphic feature) is used as the input to generate the predicted video frame at the t+1 step. Finally, the predicted video frame is decoded into m using the listener decoder l t+1 which contains two feature vectors, namely, β l t+1 shows the expression, and p l t+1 shows the pose (rotation and translation). The terminal supports the input of the speaker with an arbitrary length. The flow can be expressed by the following formula (7),

Number

[0143] It can be understood that, in order to make the video sequence of the virtual object more vivid and accurate, the terminal can process it by a speaker encoder and a listener decoder based on the audio-video sequence of the captured real object, and generate a video sequence of the virtual object.

[0144] Based on the video generation method of the above embodiment, an embodiment of the present invention further provides a video generation device as shown in FIG. 11. FIG. 11 is a structural schematic diagram 1 of video generation according to an embodiment of the present invention. The device 11 includes an acquisition part 1101, a determination part 1102, and a generation part 1103. The acquisition part 1101 is configured to capture an audio-video sequence of a real object. The determination part 1102 is configured to extract features from the audio-video sequence and determine anthropomorphic features. The generation part 1103 is configured to perform prediction on the anthropomorphic features by using a virtual prediction network, a preset standard feature that is a feature corresponding to a reference object, and a first feature representing different attitudes, and generate a video sequence of the virtual object, which is a video sequence in which corresponding reactions of the virtual object are generated based on the audio-video sequence of the real object, and present the video sequence of the virtual object.

[0145] In some embodiments of the present invention, the acquisition part 1101 is configured to acquire a preset standard feature including a first posture and expression feature and a first identity feature. The determination part 1102 is configured to perform prediction and decoding by the virtual prediction network based on the first posture and expression feature, the anthropomorphic feature, and the first feature, and determine posture and expression features of multiple frames of the virtual object. The generation part 1103 is configured to generate a video sequence of the virtual object based on the posture and expression features of multiple frames of the virtual object and the first identity feature.

[0146] In some embodiments of the present invention, the acquisition part 1101 is configured to acquire a standard image representing an image of a reference object, extract features from the standard image by means of a face reconstruction model, and obtain the preset standard features.

[0147] In some embodiments of the present invention, the anthropomorphic features include the anthropomorphic features of a plurality of corresponding frames of the audio-video sequence. The virtual prediction network includes a first processing module and a second processing module. The determination part 1102 is configured to perform a prediction by means of the first processing module based on the anthropomorphic features of the first frame among the anthropomorphic features of the plurality of frames, the first posture-expression feature, and one of the positive attitude, negative attitude, and general attitude, that is, the first feature, so as to obtain the next predicted video frame. The second processing module decodes the next predicted video frame, determines the next posture-expression feature of the corresponding virtual object of the next predicted video frame, and continues to perform prediction and decoding based on the next posture-expression feature and the anthropomorphic features of the next frame among the anthropomorphic features of the plurality of frames until the last posture-expression feature of the corresponding virtual object of the last predicted video frame is obtained, thereby obtaining the posture-expression features of a plurality of frames of the virtual object, and the first posture-expression feature is the first frame among the posture-expression features of the plurality of frames.

[0148] In some embodiments of the present invention, the acquisition part 1101 is configured to fuse each of the posture-expression features of each frame among the posture-expression features of the plurality of frames with the first identity feature including the first identity identifier, the first material, and the first light irradiation information, so as to obtain a plurality of second features representing the fusion result of the identity feature and the posture-expression feature. The generation part 1103 is configured to generate a video sequence of the virtual object corresponding to the virtual object by means of a renderer for the plurality of second features.

[0149] In some embodiments of the present invention, the audio-video sequence includes an audio sequence of a real object and a video sequence of the real object. The obtaining part 1101 is configured to extract features from the video sequence of the real object by an encoder to obtain a plurality of video features, and extract features from the audio sequence of the real object by the encoder to obtain a plurality of audio features including loudness, zero-crossing rate, and cepstrum coefficients. The determining part 1102 is configured to perform feature conversion by a feature fusion function based on the plurality of video features and the plurality of audio features, and determine anthropomorphic features corresponding to a plurality of frames of the audio-video sequence, the anthropomorphic features including video features and audio features.

[0150] In some embodiments of the present invention, the obtaining part 1101 is configured to extract features from each video frame of the video sequence of the real object by a face reconstruction model to obtain the plurality of video frame features including a second identity feature and a second pose-expression feature. The determining part 1102 is configured to use all the corresponding second pose-expression features in the video sequence of the real object as the video features.

[0151] In some embodiments of the present invention, the obtaining part 1101 is configured to collect an audio-video sequence sample of a real speaking object and a face image of a corresponding real listening object. The determining part 1102 is configured to extract features from the audio-video sequence sample by an initial encoder to determine anthropomorphic sample features. The generating part 1103 is configured to generate predicted face features in a sample audio-video sequence of a real object to be trained, including predicted pose features and predicted expression features, based on the initial virtual prediction network and the anthropomorphic sample features. The determination part 1102 extracts features by means of a face reconstruction model based on the face image of the actual listening object, determines actual face features including actual pose features and actual expression features, and continuously optimizes the initial encoder by means of a first loss function and the anthropomorphic sample features until the value of the first loss function satisfies a first preset threshold, determines the encoder, and continuously optimizes an initial virtual prediction network by means of a second loss function and a third loss function based on the actual face features and predicted face features until the sum of the value of the second loss function and the value of the third loss function satisfies a second preset threshold, and is configured to determine the virtual prediction network.

[0152] In some embodiments of the present invention, the determination part 1102 determines a second loss function for ensuring that a predicted expression and a predicted pose are similar to an actual expression and an actual pose based on the actual face features and the predicted face features, determines a third loss function for ensuring that the continuity between frames of the predicted face features is similar to that of the actual face features based on a change function corresponding to the actual face features and a change function corresponding to the predicted face features, and continuously optimizes the initial virtual prediction network by means of the second loss function and the third loss function until the sum of the value of the second loss function and the value of the third loss function satisfies a second preset threshold, and is configured to determine the virtual prediction network.

[0153] It should be noted that when generating a video, the division of the above program modules is only taken as an example for description. In actual applications, the above processing can be allocated according to needs and completed by different program modules, that is, the internal structure of the device can be divided into different program modules to complete all or part of the processing described above. In addition, the video generation device according to the above embodiments belongs to the same concept as the embodiments of the video generation method. For the specific implementation process and details of the beneficial effects, reference is made to the embodiments of the method, and the description is not repeated here. For the technical details not disclosed in the embodiments of the present device, reference is made to the description of the embodiments of the method of the present invention for understanding.

[0154] Based on the video generation method of the above embodiments, an embodiment of the present invention further provides a video generation device as shown in FIG. 12. FIG. 12 is a schematic structural diagram 2 of the video generation device according to an embodiment of the present invention. The device 12 includes a processor 1201 and a memory 1202. One or more programs executable by the processor are stored in the memory 1202. When one or more programs are executed, the processor 1201 executes any one of the video generation methods of the above embodiments.

[0155] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can adopt the form of an embodiment in hardware, an embodiment in software, or an embodiment in which software and hardware aspects are combined. Furthermore, the present invention can adopt the form of a computer program product implemented on a computer-usable storage medium (including but not limited to magnetic disk memory, optical memory, etc.) provided with one or more computer-usable program codes.

[0156] The present invention will be described with reference to the flowcharts and / or block diagrams of the method, apparatus (system), and computer program product according to the embodiments of the present invention. It should be understood that computer program instructions can implement each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate an apparatus for realizing the functions specified by one or more flows in the flowchart and / or one or more blocks in the block diagram.

[0157] These computer program instructions may be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, whereby the instructions stored in the computer-readable memory produce an article of manufacture including an instruction device, the instruction device realizing the functions specified in one or more flows of a flow diagram and / or one or more blocks of a block diagram.

[0158] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus, whereby a series of operational steps are executed on the computer or other programmable apparatus to produce a process implemented by the computer, and the instructions executed on the computer or other programmable apparatus provide steps for realizing the functions specified in one or more flows of a flow diagram and / or one or more blocks of a block diagram.

[0159] The above are only preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention.

Industrial Applicability

[0160] Embodiments of the present invention provide a video generation method, apparatus, and computer-readable storage medium. Among them, the method includes: collecting the audio-video sequence of a real object; extracting features from the audio-video sequence to determine anthropomorphic features; using a virtual prediction network, a preset standard feature that is a feature corresponding to a reference object, and a first feature representing different attitudes to predict the anthropomorphic features; generating a video sequence of a virtual object, which is a video sequence in which a corresponding reaction of the virtual object is generated based on the audio-video sequence of the real object; and presenting the video sequence of the virtual object. Embodiments of the present invention generate a video sequence of a virtual object based on the audio-video sequence of a real object, thereby making the presented video sequence of the virtual object more vivid and accurate.

Claims

1. Collecting the audio - video sequence of the real object, extracting features from the audio - video sequence and determining anthropomorphic features, using a virtual prediction network, a preset standard feature that is a feature corresponding to a reference object, and a first feature representing different attitudes to perform a prediction on the anthropomorphic features, and generating a video sequence of a virtual object that is a video sequence in which a corresponding reaction of the virtual object is generated based on the audio - video sequence of the real object, presenting the video sequence of the virtual object, including: A video generation method.

2. The above - mentioned performing a prediction on the anthropomorphic features using a virtual prediction network, a preset standard feature, and a first feature, and generating a video sequence of a virtual object includes: obtaining a preset standard feature including a first posture - expression feature and a first identity feature, performing prediction and decoding by the virtual prediction network based on the first posture - expression feature, the anthropomorphic features, and the first feature, and determining the posture - expression features of multiple frames of the virtual object, generating a video sequence of the virtual object based on the posture - expression features of multiple frames of the virtual object and the first identity feature, including: The method according to claim 1.

3. The above - mentioned obtaining a preset standard feature includes: obtaining a standard image representing an image of a reference object, extracting features from the standard image by a face reconstruction model to obtain the preset standard feature, including: The method according to claim 2.

4. The anthropomorphic features include anthropomorphic features of corresponding multiple frames of the audio - video sequence, and the virtual prediction network includes a first processing module and a second processing module. The above - mentioned performing prediction and decoding by the virtual prediction network based on the first posture - expression feature, the anthropomorphic features, and the first feature, and determining the posture - expression features of multiple frames of the virtual object is: Based on the anthropomorphic feature of the first frame among the anthropomorphic features of the plurality of frames, the first posture-expression feature, and the first feature which is one of a positive attitude, a negative attitude, and a general attitude, the first processing module makes a prediction to obtain the next one predicted video frame. The second processing module decodes the next one predicted video frame and determines the next one posture-expression feature of the corresponding virtual object of the next one predicted video frame. Until obtaining the last one posture-expression feature of the corresponding virtual object of the last one predicted video frame, continue to make predictions and decode based on the next one posture-expression feature and the anthropomorphic feature of the next one frame among the anthropomorphic features of the plurality of frames, thereby obtaining the posture-expression features of a plurality of frames of the virtual object. The first posture-expression feature is the first frame among the posture-expression features of the plurality of frames. The method according to claim 2.

5. Generating the video sequence of the virtual object based on the posture-expression features of a plurality of frames of the virtual object and the first identity feature described above includes: Fusing each of the posture-expression features of each frame among the posture-expression features of the plurality of frames with the first identity feature including a first identity label, a first material, and first light irradiation information to obtain a plurality of second features representing the fusion result of the identity feature and the posture-expression feature. Generating the video sequence of the virtual object by a renderer for the plurality of second features. The method according to claim 2.

6. The audio-video sequence includes the audio sequence of the real object and the video sequence of the real object. Performing feature extraction on the audio-video sequence described above to determine anthropomorphic features includes: Performing feature extraction on the video sequence of the real object by an encoder to obtain a plurality of video features. Performing feature extraction on the audio sequence of the real object by an encoder to obtain a plurality of audio features including loudness, zero-crossing rate, and cepstrum coefficient. Based on the plurality of video features and the plurality of audio features, perform feature conversion by a feature fusion function, which is the anthropomorphic features of corresponding multiple frames of an audio-video sequence, and determine the anthropomorphic features including video features and audio features. The method according to any one of claims 1 to 5.

7. The above-mentioned extraction of features from the video sequence of the actual object by the encoder to obtain a plurality of video features includes: Performing feature extraction on each video frame of the video sequence of the actual object by a face reconstruction model to obtain a plurality of video frame features including a second identity feature and a second pose-expression feature. Using all the corresponding second pose-expression features in the video sequence of the actual object as the video features. The method according to claim 6.

8. Before performing prediction on the anthropomorphic features using the virtual prediction network, the preset standard features, and the first feature to generate a video sequence of a virtual object, Collecting an audio-video sequence sample of the actual speaking object and a face image of the corresponding actual listening object. Performing feature extraction on the audio-video sequence sample by an initial encoder to determine anthropomorphic sample features. Generating predicted face features in the audio-video sequence sample of the actual object to be trained, including a predicted pose feature and a predicted expression feature, by an initial virtual prediction network and the anthropomorphic sample features. Performing feature extraction on the face image of the actual listening object based on the face reconstruction model to determine actual face features including an actual pose feature and an actual expression feature. Continuously optimizing the initial encoder by a first loss function and the anthropomorphic sample features until the first loss function value satisfies a first preset threshold, and determining the encoder. Continuously optimizing the initial virtual prediction network by a second loss function and a third loss function based on the actual face features and the predicted face features until both the second loss function value and the third loss function value satisfy a second preset threshold, and determining the virtual prediction network. The method according to any one of claims 1 to 5.

9. Continuing to optimize the initial virtual prediction network using the second loss function and the third loss function based on the actual face features and the predicted face features until the above-mentioned second loss function value and third loss function value satisfy the second preset threshold, and determining the virtual prediction network is determining a second loss function for ensuring that the predicted expression and predicted pose are similar to the actual expression and actual pose based on the actual face features and the predicted face features; determining a third loss function for ensuring that the continuity between frames of the predicted face features is similar to that of the actual face features based on the change function corresponding to the actual face features and the change function corresponding to the predicted face features; continuing to optimize the initial virtual prediction network using the second loss function and the third loss function until the second loss function value and the third loss function value satisfy the second preset threshold, and determining the virtual prediction network, including The method according to claim 8.

10. An acquisition part configured to acquire an audio-visual sequence of an actual object; a determination part configured to extract features from the audio-visual sequence and determine anthropomorphic features; a generation part configured to generate a video sequence of a virtual object, which is a video sequence in which a prediction is made on the anthropomorphic features using a virtual prediction network, a preset standard feature that is a feature corresponding to a reference object, and a first feature representing different attitudes, and a corresponding reaction of the virtual object is generated based on the audio-visual sequence of the actual object, and to present the video sequence of the virtual object A video generation device.

11. The acquisition part is further configured to acquire a preset standard feature including a first pose-expression feature and a first identity feature; The determination part is further configured to perform prediction and decoding by the virtual prediction network based on the first pose-expression feature, the anthropomorphic features, and the first feature, and determine pose-expression features of multiple frames of the virtual object; The generation part is further configured to generate the video sequence of the virtual object based on the pose-expression features of multiple frames of the virtual object and the first identity feature; The device according to claim 10.

12. The acquisition part is further configured to acquire a standard image representing an image of a reference object, extract features from the standard image by a face reconstruction model, and obtain the preset standard features. The apparatus according to claim 11.

13. The anthropomorphic features include anthropomorphic features of a plurality of corresponding frames of the audio-video sequence. The virtual prediction network includes a first processing module and a second processing module. The acquisition part is further configured to perform prediction by the first processing module based on the anthropomorphic feature of the first frame among the anthropomorphic features of the plurality of frames, the first posture and expression feature, and the first feature which is one of a positive attitude, a negative attitude, and a general attitude, so as to obtain the next predicted video frame. The determination part is further configured to decode the next predicted video frame by the second processing module and determine the next posture and expression feature of the corresponding virtual object of the next predicted video frame. The acquisition part is further configured to continue prediction and decoding based on the next posture and expression feature and the anthropomorphic feature of the next frame among the anthropomorphic features of the plurality of frames until the last posture and expression feature of the corresponding virtual object of the last predicted video frame is obtained, thereby obtaining the posture and expression features of a plurality of frames of the virtual object. The first posture and expression feature is the first frame among the posture and expression features of the plurality of frames. The apparatus according to claim 11.

14. The acquisition part is further configured to fuse each of the posture and expression features of each frame among the posture and expression features of the plurality of frames with the first identity feature including a first identity identifier, a first material, and first light irradiation information, so as to obtain a plurality of second features representing the fusion result of the identity feature and the posture and expression feature. The generation part is further configured to generate a video sequence of the virtual object by a renderer for the plurality of second features. The apparatus according to claim 11.

15. The audio-video sequence includes an audio sequence of a real object and a video sequence of the real object. The acquisition part further extracts features from the video sequence of the actual object by an encoder to obtain a plurality of video features, and extracts features from the audio sequence of the actual object by an encoder to obtain a plurality of audio features including loudness, zero-crossing rate, and cepstral coefficients. The determination part further performs feature conversion by a feature fusion function based on the plurality of video features and the plurality of audio features to determine anthropomorphic features of corresponding multiple frames of the audio-video sequence, and the anthropomorphic features including video features and audio features. The apparatus according to any one of claims 10 to 14.

16. The acquisition part further extracts features from each video frame of the video sequence of the actual object by a face reconstruction model to obtain a plurality of video frame features including second identity features and second pose-expression features. The determination part is further configured to use all the corresponding second pose-expression features in the video sequence of the actual object as the video features. The apparatus according to claim 15.

17. The acquisition part is further configured to collect an audio-video sequence sample of the actual speaking object and a face image of the corresponding actual listening object. The determination part is further configured to extract features from the audio-video sequence sample by an initial encoder to determine anthropomorphic sample features. The generation part is further configured to generate predicted face features in the audio-video sequence sample of the actual object to be trained including predicted pose features and predicted expression features by an initial virtual prediction network and the anthropomorphic sample features. The determining part further extracts features by a face reconstruction model based on the face image of the actual listening object, determines actual face features including actual pose features and actual expression features, and continues to optimize the initial encoder by the first loss function and the anthropomorphic sample features until the first loss function value satisfies a first preset threshold, determines the encoder, and continues to optimize the initial virtual prediction network by a second loss function and a third loss function based on the actual face features and predicted face features until both the second loss function value and the third loss function value satisfy a second preset threshold, and is configured to determine the virtual prediction network. The apparatus according to any one of claims 10 to 14.

18. The determining part further determines a second loss function for ensuring that the predicted expression and predicted pose are similar to the actual expression and actual pose based on the actual face features and the predicted face features, determines a third loss function for ensuring that the continuity between frames of the predicted face features is similar to that of the actual face features based on the change function corresponding to the actual face features and the change function corresponding to the predicted face features, and continues to optimize the initial virtual prediction network by the second loss function and the third loss function until both the second loss function value and the third loss function value satisfy a second preset threshold, and is configured to determine the virtual prediction network. The method according to claim 17.

19. A memory for storing executable instructions, A processor that, when executing the executable instructions stored in the memory, realizes the video generation method according to any one of claims 1 to 9, and includes: A video generation device.

20. When executed, it stores executable instructions used to cause a processor to execute the video generation method according to any one of claims 1 to 9. A computer-readable storage medium.

Citation Information

Patent Citations

  • Head posture estimation device, head posture estimation method and program for making computer execute head posture estimation method

    JP2014093006A

  • Nonverbal information generation device, nonverbal information generation model learning device, method, and program

    WO2019160100A1