Video generation method and device and related equipment

By acquiring user-input images and voice, and combining them with emotional information to generate target voice and images, this technology solves the problem of insufficient flexibility and editability in existing voice-driven photo virtual digital human technology. It achieves emotion matching and rich motion video generation, background replacement, and improves the display effect.

CN122073633APending Publication Date: 2026-05-22HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
Filing Date
2024-11-22
Publication Date
2026-05-22

Smart Images

  • Figure CN122073633A_ABST
    Figure CN122073633A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a video generation method and device and related equipment, and aims to generate a flexible virtual character image according to the demand of a user. The video generation method comprises the steps that an input image, original voice and emotion information are acquired, the input image comprises a face image of a target object, and the emotion information is used for indicating the emotion of a to-be-generated target video; performing emotion synthesis on the original voice based on the emotion information to obtain a target voice, the target voice expressing the content of the original voice through the emotion; a target image set is generated based on the input image, the emotion information and the target voice, the target image set comprises a plurality of target images, and the plurality of target images are used for displaying an image of the target voice sent by the target object with the emotion; and combining the target image set and the target voice to obtain a target video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a video generation method, apparatus and related equipment. Background Technology

[0002] With the continuous evolution and development of computer technology, the application of virtual digital human technology is becoming increasingly widespread. Virtual digital human technology can generate video data that allows virtual avatars to simulate certain behaviors of real objects. Among these, voice-driven photo-based virtual digital human technology is an important development direction. This technology can generate videos based on photos and audio data. The generated videos are used to demonstrate the process of a person in a photo speaking the audio data.

[0003] Specifically, when generating video data, photos of real people and audio recordings to be played can be acquired. By analyzing the photos and audio, the facial movement features of the real person speaking can be determined, and a corresponding set of image frames can be generated by combining the photos. The image frame set is then combined with the audio recordings to be played, and the resulting video data can be used to demonstrate the effect of the real person speaking.

[0004] However, traditional voice-driven photo-based virtual digital human technology can only adjust the virtual digital human's facial movements according to the voice to be played, which is not very flexible. Summary of the Invention

[0005] In view of this, embodiments of this application provide a video generation method aimed at improving the editability and flexibility of voice-driven photo virtual digital humans. This application also provides corresponding apparatus, computing device clusters, computer-readable storage media, and computer program products.

[0006] Firstly, this application provides a video generation method. When generating a video corresponding to a target object, a user can input an image corresponding to the target object (hereinafter referred to as the input image) and original speech. Additionally, the user can select the emotion of a virtual character (hereinafter referred to as emotion information). After obtaining the input image, input speech, and emotion information, the speech-related parts of the video can be determined first, and then the image-related parts of the video can be determined based on the speech parts. First, emotion synthesis can be performed on the original speech based on the emotion information to obtain target speech expressing the content of the original speech with the emotion indicated by the emotion information. After determining the target speech, the target speech and input image can be combined to simulate the image of the target object uttering the target speech with the emotion, resulting in a target image set including multiple target images. Playing the target images in the target image set in sequence can show the audience the image of the target object uttering the target speech with the emotion. Then, merging the target image set and the target speech yields a target video used to display content information in the image of the target object according to the emotion indicated by the emotion information. In other words, when generating a video corresponding to a virtual character, the speech part of the target video is not the original input speech, but rather the target speech generated based on the content of the input speech and the emotion information. As can be seen, the audio portion of the target video is not limited by the emotion of the input voice, and the emotion of the audio portion in the target video can be freely adjusted. Furthermore, the image portion of the target video is generated based on the target audio, ensuring that the emotion displayed in the target video matches the emotional information. Thus, by combining emotional information to generate the audio portion of the target video, and combining the audio portion to obtain the image portion of the target video, the resulting video data accurately expresses emotion. Therefore, the emotion of the voice spoken by the virtual character is selectable, rather than being inherent to the input voice. This is equivalent to being able to edit the emotion of the virtual character, improving the editability of voice-driven photo-realistic digital human technology and enabling it to meet complex business needs.

[0007] In some possible implementations, the target speech can be obtained by adjusting the audio attributes of the original speech. Specifically, the original speech can first be analyzed to determine its audio attributes. Then, based on the emotion indicated by the emotional information, the audio attributes of the original speech can be adjusted to obtain the target speech whose audio attributes match the emotion indicated by the emotional information. That is, speech with these adjusted audio attributes can express the emotion indicated by the emotional information. Thus, by adjusting the audio attributes of the original speech, a target speech with audio attributes that match the target emotion can be obtained, achieving the effect of displaying the target emotion.

[0008] In some possible implementations, the target speech can be obtained through feature extraction and feature fusion. Specifically, emotional features can be determined based on emotional information; these emotional features are feature vectors that match the emotion indicated by the emotional information. Alternatively, features can be extracted from the original speech to determine its characteristics. Then, the features of the original speech and the emotional features can be fused to obtain the target speech. In this way, the original speech can be deconstructed into feature vectors through feature extraction, filtering out emotion-related information. By then fusing the emotional features, the target speech corresponding to the emotional information can be obtained.

[0009] In some possible implementations, the features of the original speech include features related to the target object and features related to the content of the original speech. Thus, feature fusion is performed based on features related to the target object, features related to the content of the original speech, and emotional features. The resulting target speech matches the emotional information, has content consistent with the original speech, and has the same auditory effect as the speech directly spoken by the target object. This allows for a better simulation of the auditory effect of the target object expressing content with emotion, improving the realism of the target video.

[0010] In some possible implementations, emotion features can be extracted by a first encoder. Features of the original speech are extracted by a second encoder, and the target speech is obtained by a speech synthesis model. The first encoder, second encoder, and speech synthesis model can be trained based on a model training process. During model training, a first source speech, a second source speech, and a second reference speech are first acquired. The emotion corresponding to the first source speech is different from the emotion corresponding to the second source speech, but the emotion corresponding to the first source speech is the same as the emotion corresponding to the second reference speech. The content of the second reference speech is the same as the content of the second source speech. The second reference speech and the second source speech are spoken by the same person. Next, the first encoder can extract features from the first source speech to obtain the emotion features corresponding to the emotion of the first source speech. Furthermore, the second encoder can also extract features from the second source speech to obtain its features. Finally, the speech synthesis model combines the emotion features corresponding to the emotion of the first source speech and the features of the second source speech to generate the first reference speech. By comparing the first and second reference speech, the differences between the speech obtained through feature extraction and fusion and the real speech can be determined. Based on these differences, the first encoder, the second encoder, and the speech synthesis model can be adjusted. Thus, based on the difference between the real and generated values, the encoder and model can be adjusted to make their output values ​​closer to the real values.

[0011] In some possible implementations, the emotional features of the target speech can be obtained by extracting emotional audio samples from the target sample audio. Specifically, after acquiring the emotional information, the target sample emotional audio can be determined from multiple sample emotional audio samples based on this information. Different sample emotional audio samples can correspond to different emotions, and the target sample emotional audio corresponds to the emotional information. By extracting features from the target sample emotional audio, the emotional features can be obtained. In this way, the emotional features are obtained by extracting features from pre-recorded audio, which can better reflect the emotions corresponding to the emotional information.

[0012] In some possible implementations, the facial motion effects of the target object can be simulated. Specifically, when generating the target image set, the facial motion information of the target object can be determined based on the input images and the target speech. The facial motion information of the target object refers to the movement information of the facial tissues of the target object when uttering the target speech. Based on the facial motion information of the target object, a facial driving effect can be generated, and the target image set is obtained based on the facial driving effect. In the multiple images corresponding to the target image set, the face of the target object moves according to the facial motion information. In this way, the target video generated based on the target image set can demonstrate the facial motion of the target object during the process of uttering the target speech.

[0013] In some possible implementations, in addition to facial motion effects, limb motion effects of the target object can also be simulated. Specifically, user-configured instructions can be obtained. If the instructions specify generating only facial motion effects, then a target image can be generated based on the facial motion information. If the instructions specify generating both facial motion effects and limb-driven effects, then the limb motion information of the target object can be determined based on the input image, emotional information, and target speech, and a set of target images can be generated based on the limb motion information and facial motion information. In this way, not only will the face of the target object move with the target speech, but the limbs will also move with the target speech, improving the display effect of the target video.

[0014] In some possible implementations, background replacement can be performed on the input image. Specifically, when generating the target video, a background image can be acquired and used to replace the background of the target image. Furthermore, illumination information can be calculated based on the background image, and the target image can be relit based on this illumination information to obtain a target image with a replaced background, where the illumination information of the target image matches that of the Beijing portion. Finally, the processed set of target images and the target speech can be merged to obtain the target video. In this way, by recalculating the illumination information, the consistency between the target information and the background in the target video is ensured.

[0015] Secondly, this application provides a video generation apparatus, the apparatus comprising: an acquisition unit, configured to acquire an input image, original speech, and emotion information, wherein the input image includes a facial image of a target object, and the emotion information is used to indicate the emotion of a target video to be generated; a speech synthesis unit, configured to synthesize an emotion from the original speech based on the emotion information to obtain target speech, wherein the target speech expresses the content of the original speech with the emotion; an image generation unit, configured to generate a target image set based on the input image, the emotion information, and the target speech, wherein the target image set includes multiple target images, and the multiple target images are used to display the image of the target object uttering the target speech with the emotion; and a video generation unit, configured to merge the target image set and the target speech to obtain a target video.

[0016] In some possible implementations, the speech synthesis unit is specifically configured to determine the audio attributes of the original speech based on the original speech; adjust the audio attributes of the original speech according to the emotion indicated by the emotion information, and generate the target speech, wherein the audio attributes of the target speech match the emotion indicated by the emotion information.

[0017] In some possible implementations, the original speech is the speech emitted by the target object. The speech synthesis unit is specifically used to extract features from the original speech, determine the features of the original speech, the features of the original speech include features related to the target object and features related to the content of the original speech; determine emotional features based on the emotional information; and perform feature fusion on the features of the original speech and the emotional features to obtain the target speech.

[0018] In some possible implementations, the acquisition unit is further configured to acquire instruction information; the image generation unit is specifically configured to determine the facial motion information of the target object based on the input image and the target speech; if the instruction information indicates the generation of a facial-driven effect, generate the target image set based on the facial motion information; if the instruction information indicates the generation of both the facial-driven effect and the limb-driven effect, determine the limb motion information of the target object based on the input image, the emotion information, and the target speech, and generate the target image set based on the limb motion information and the facial motion information.

[0019] In some possible implementations, the image generation unit is specifically used to acquire a background image; replace the background of the at least one target image with the background image, calculate illumination information based on the background image and relight it to obtain multiple processed target images; the video generation unit is specifically used to merge the processed target image set and the target speech to obtain a target video.

[0020] In some possible implementations, the emotion features are extracted by a first encoder, the features of the original speech are extracted by a second encoder, and the target speech is obtained by a speech synthesis model. The first encoder, the second encoder, and the speech synthesis model are trained as follows: A first source speech and a second source speech are acquired, where the emotions corresponding to the first source speech and the second source speech are different; features are extracted from the first source speech using the first encoder to obtain the emotion features corresponding to the first source speech; features are extracted from the second source speech using the second encoder to obtain the features of the second source speech; a first reference speech is generated by combining the emotion features corresponding to the first source speech and the features of the second source speech using the speech synthesis model; the first reference speech and the second reference speech are compared, and the first encoder, the second encoder, and the speech synthesis model are adjusted based on the comparison result; wherein the emotion corresponding to the second reference speech is the same as the emotion corresponding to the first source speech, the content of the second reference speech is the same as the content of the second source speech, and the second reference speech and the second source speech are spoken by the same person.

[0021] In some possible implementations, the speech synthesis unit is specifically used to determine a target sample emotional audio from multiple sample emotional audios based on the emotional information, wherein different sample emotional audios correspond to different emotions; and to extract features from the target sample emotional audio to obtain the emotional features.

[0022] Thirdly, this application provides a computing device cluster, the computing device including at least one computing device, the at least one computing device including at least one processor and at least one memory; the at least one memory is used to store instructions, and the at least one processor executes the instructions stored in the at least one memory to cause the computing device cluster to execute the multi-role multi-terminal real-time synchronized code inspection method in the second aspect or any possible implementation of the second aspect. It should be noted that the memory can be integrated into the processor or can be independent of the processor. The at least one computing device may also include a bus. The processor is connected to the memory via the bus. The memory may include readable storage and random access memory.

[0023] Fourthly, this application provides a computer-readable storage medium storing instructions that, when executed on at least one computing device, cause the at least one computing device to perform the method described in the first aspect or any implementation thereof.

[0024] Fifthly, this application provides a computer program product containing instructions that, when run on at least one computing device, cause the at least one computing device to perform the method described in the first aspect or any implementation thereof.

[0025] Based on the implementation methods provided in the above aspects, this application can be further combined to provide more implementation methods. Attached Figure Description

[0026] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings.

[0027] Figure 1 A schematic diagram illustrating an application scenario provided in this application embodiment;

[0028] Figure 2 This is a schematic flowchart of a video generation method provided in an embodiment of this application;

[0029] Figure 3 This is a schematic flowchart illustrating the video generation process provided in an embodiment of this application.

[0030] Figure 4a A schematic flowchart illustrating the model training process provided in this application embodiment;

[0031] Figure 4b Another schematic diagram of the model training process provided in the embodiments of this application;

[0032] Figure 5 This is a schematic diagram of a video generation apparatus provided in an embodiment of this application.

[0033] Figure 6 A schematic diagram of the structure of a computing device provided in an embodiment of this application;

[0034] Figure 7 This is a schematic diagram of a computing device cluster provided in an embodiment of this application;

[0035] Figure 8 This is a schematic diagram illustrating one implementation of a computing device cluster provided in an embodiment of this application. Detailed Implementation

[0036] The solutions in the embodiments provided in this application will now be described with reference to the accompanying drawings.

[0037] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application.

[0038] First, let me introduce the terms used in this application.

[0039] Virtual avatar: A virtual avatar refers to an image generated through virtual digital human technology. The external appearance of a virtual avatar can be identical to that of a real person. For example, in photo-realistic digital human technology, the input image is a photograph of a real person. The generated virtual avatar is the image of the real person, and its external appearance matches the image of the real person in the input image. In the embodiments of this application, if the virtual avatar is generated based on the image of a real person and its appearance is identical or similar to that of the real person, the virtual avatar can be called a virtual task avatar.

[0040] Input audio information data: In this embodiment of the application, input audio information data refers to the audio data input by the user using the technology of voice-driven photo virtual digital human.

[0041] In voice-driven photorealistic digital human technology, input images and audio are analyzed to determine the facial movements of a real person speaking the corresponding audio. Based on these facial movements and the real person's photograph, multiple frames are generated. Playing these frames sequentially with the corresponding audio allows viewers to experience the effect of the real person speaking the corresponding audio. This technology allows for the generation of videos of a person speaking, based on their voice and photograph, without the need for actual video recording.

[0042] However, current voice-driven photo-based virtual digital human technology has poor editability and cannot meet increasingly complex business needs. This problem manifests itself in at least the following three aspects.

[0043] Firstly, the audio portion of the generated video is determined based on the input and cannot be flexibly edited or adjusted.

[0044] Specifically, if a user needs to generate a video of someone speaking a certain phrase, the user needs to upload not only a photo of that person but also record and upload the corresponding audio. The uploaded audio is then overlaid with a set of image frames generated from the photo to obtain the target video. If the user needs different videos, different audio data needs to be recorded. This approach lacks flexibility.

[0045] In real-world scenarios, the sounds a person makes can be influenced by their emotions. That is, the input audio data may include the emotional information of the person recording it. The audio portion of the video data to be generated needs to simulate the emotional information of a real person speaking. However, there may be a difference between the emotional information of the person recording and the emotional information corresponding to the expected output audio portion. Thus, the resulting video may not accurately express the emotional information the user intends to convey.

[0046] For example, suppose user A wants to generate a video of themselves speaking in a state of anger. However, user A's anger during audio recording is not sufficiently conveyed, resulting in an audio recording that doesn't adequately express the emotion. Furthermore, because the audio data doesn't fully express "anger," the facial expressions of user A in the generated image frames also fail to adequately convey "anger." Consequently, the generated video will not adequately express the emotion of "anger" and will not meet user A's video generation requirements.

[0047] Secondly, the generated virtual character images lack sufficient motion information.

[0048] Currently, voice-driven photo-based virtual digital human technology mainly focuses on the facial expressions of virtual characters, aiming to make the facial expressions of virtual characters generated from photos closely resemble those of real people when speaking. However, most other parts of the virtual character are in a static state, resulting in an unnatural display effect.

[0049] For example, in real-life speaking scenarios, people might use body language to enhance their expressiveness, such as waving their hand or leaning forward to emphasize certain points. However, currently generated virtual avatars can only make facial movements and cannot perform other body movements, resulting in poor expressiveness.

[0050] Thirdly, the background of the generated video data is monotonous.

[0051] In photo-based virtual human technology, the virtual avatar is generated based on the input photograph. However, the background of the real person in the photograph is often fixed and cannot be changed. Furthermore, the lighting on the avatar in the photograph may differ from the lighting on the avatar in a new background. Therefore, cropping the virtual avatar from the background of the input photograph and pasting it onto a new background may result in an unnatural blending of the virtual avatar and background in the new image, leading to a poor display effect.

[0052] As can be seen from the above three aspects, the current voice-driven photo virtual digital human technology suffers from poor editability.

[0053] Based on this, embodiments of this application provide a video generation method aimed at generating flexible virtual character images according to user needs. Specifically, when generating a video corresponding to a target object, the user can input an image corresponding to the target object (hereinafter referred to as the input image) and the original voice. In addition, the user can also select the emotion of the virtual character image (hereinafter referred to as emotion information). After obtaining the input image, input voice, and emotion information, the voice-related parts in the video can be determined first, and then the image-related parts in the video can be determined based on the voice parts. First, the original voice can be synthesized based on the emotion information to obtain the target voice expressing the content of the original voice with the emotion indicated by the emotion information. After determining the target voice, the target voice and the input image can be combined to simulate the image of the target object when uttering the target voice with the emotion, resulting in a target image set including multiple target images. Playing the target images in the target image set in sequence can show the audience the image of the target object when uttering the target voice with the emotion. Then, merging the target image set and the target voice can obtain a target video used to display content information in the image of the target object according to the emotion indicated by the emotion information. In other words, when generating videos corresponding to virtual avatars, the audio portion of the target video is not the original input audio, but rather a target audio generated by combining the content and emotional information of the input audio. Therefore, the audio portion of the target video is not limited by the emotion of the input audio, and the emotion of the audio portion in the target video can be freely adjusted. Furthermore, the image portion of the target video is generated based on the target audio, ensuring that the emotion displayed in the target video matches the emotional information. Thus, by combining emotional information to generate the audio portion of the target video, and combining the audio portion to obtain the image portion of the target video, the resulting video data accurately expresses emotion. In this way, the emotion of the voice emitted by the virtual avatar is selectable, rather than being inherent to the input audio, essentially allowing for the editing of the virtual avatar's emotions. This improves the editability of voice-driven photo-realistic digital human technology and can meet complex business needs.

[0054] Next, various non-limiting specific implementations of the video generation process will be described in detail.

[0055] First, an exemplary application scenario is introduced. The video generation method provided in this application can be applied to a client or a server. The client can be software or a software module running on a terminal device. The terminal device can be a mobile terminal device such as a mobile phone or tablet, or a device such as a personal computer. The server can be software or a software module running on a server or server cluster. Alternatively, the video generation method can also be implemented by other physical or virtual devices with data processing capabilities. No limitation is made here. The following section combines… Figure 1 This paper introduces some implementation methods for server-side video generation.

[0056] See Figure 1 , Figure 1 This is a schematic diagram illustrating one application scenario of the video generation method provided in this application. Figure 1 In the application scenario shown, the video generation method is implemented by server 20. Figure 1 In the application scenario shown, user A can trigger a digital human video generation command through the client 10 of the virtual digital human software. The digital human video generation command includes an input image, input speech, and emotional information. The client 10 forwards the digital human video generation command to the server 20 via the network. The server 20 includes an acquisition unit 21, a speech synthesis unit 22, an image generation unit 23, and a video generation unit 24. The acquisition unit 21 can obtain the input image, input speech, and emotional information according to the digital human video generation command. The speech synthesis unit 22 can synthesize emotions from the original speech based on the emotional information to generate target speech that matches the emotions expressed by the emotional information. The image generation unit 23 can obtain the input image from the acquisition unit 21 and the target speech from the speech synthesis unit 22, thereby simulating the image of the target object when it speaks the target speech with the emotion, and generating a set of target images. The video generation unit 24 can obtain the target speech from the speech synthesis unit 22 and generate a set of target images from the image generation unit 23, thereby merging the set of target images and the target speech to obtain the target video. After obtaining the target video, the server 20 can return the target video to the client 10 via the network.

[0057] exist Figure 1 In the implementation shown, the target video corresponding to the virtual digital human is generated by the backend server 20, and user A triggers the generation of the target video through client 10. Server 20 can also be referred to as a video generation device. Client 10 can run on the device used by user A, such as a mobile terminal device or personal computer. Server 20 can run on a server or a server cluster. Server 20 can run on one server cluster or multiple server clusters. If server 20 runs on multiple server clusters, different functional units within server 20 can run on the same server cluster or on multiple server clusters.

[0058] It is understandable that in some other possible implementations, it is also possible to achieve this without using methods such as... Figure 1 The "server-client" architecture shown implements video generation. Specifically, it can be... Figure 1In the illustrated embodiment, the functions of client 10 and server 20 are integrated into the same software module. This software module can be installed on a terminal device. This allows for the generation of videos corresponding to virtual digital humans in application scenarios that do not rely on a server. It should be noted that the above application scenarios are merely examples, and the video generation method provided in this application embodiment can be applied to any virtual digital human generation scenario.

[0059] Next, various non-limiting specific implementations of the video generation process will be described in detail.

[0060] See Figure 2 , Figure 2 This is a flowchart illustrating the video generation method provided in this application. This method can be applied to... Figure 1 The application scenarios shown can also be applied to other applicable application scenarios.

[0061] Specifically, Figure 2 The video generation method shown may specifically include:

[0062] S201: Acquire the input image, raw speech, and emotion information.

[0063] Before generating the target video, the video generation device first needs to acquire the information used to generate the video, including the input image, the original audio, and emotional information.

[0064] The input image refers to the image used to generate the virtual avatar, corresponding to the target object. The input image can be an image obtained by capturing images of the target object. The target object is the object simulated by the virtual avatar, such as a real person. Accordingly, the input image can be an image taken of that real person. The target image generated based on the input image also includes that real person. The target video generated from the target image can demonstrate the effect of a real task uttering target speech with target emotions.

[0065] The original audio refers to the audio input from the target object, used to generate the audio portion of the target video. Optionally, the original audio can be pre-recorded audio from the target object.

[0066] Target emotion information refers to the emotion expressed in the output speech. Emotional information indicates the target emotion, which is the emotion the virtual character in the target video is expected to express. Target emotions can be, for example, happiness, anger, calmness, and sadness. Users can configure the target emotion through the client. For example, the client can display multiple emotion selection controls, each corresponding to a specific emotion. Users can trigger the emotion selection control corresponding to the target emotion to allow the client to determine the target emotion. Optionally, the target emotion can include one emotion or multiple emotions. The following explanation primarily uses the example of a target emotion consisting of one emotion.

[0067] For example, in Figure 1 In the illustrated application scenario, client 10 can display multiple emotion controls to user A. Each emotion control corresponds to a selectable emotion. User A can select an emotion control on client 10. The emotion corresponding to the selected emotion control is the target emotion. Client 10 can send the tag corresponding to the target emotion to server 20 so that server 20 can obtain the target emotion. In other words, users can select a target emotion through configuration operations when using the cloud service corresponding to the virtual digital human.

[0068] Optionally, in addition to the target emotion, the emotion information may also include other emotion-related information. For example, the target emotion information may also include the degree of emotional expression of the target emotion. The degree of emotional expression indicates the intensity of the emotion of the generated virtual digital human.

[0069] S202: Based on emotional information, perform emotion synthesis on the original speech to obtain the target speech.

[0070] After extracting content-related information from the original speech, the video generation device can generate target speech based on the target emotion and content information. Target speech is speech that expresses the content of the original speech with the target emotion. That is, the content of the target speech matches the content of the original speech, and the emotion expressed by the target speech is consistent with the target emotion. In this way, using the target speech as the audio portion of the target video achieves the effect of the target audience expressing content information according to the target emotion.

[0071] The following describes two methods for determining the target speech.

[0072] Method 1: Determine the target speech by adjusting audio attributes.

[0073] In the first implementation, the audio attributes of the original speech can be adjusted to obtain the target speech. Specifically, the audio attributes of the original speech can be determined first. Then, the audio of the original speech can be adjusted according to the emotion indicated by the target emotion information (i.e., the target emotion) to obtain the target speech. The audio attributes of the target speech correspond to the target emotion.

[0074] Audio attributes refer to the emotion-related properties of speech, such as speech rate, pitch, and volume. By adjusting the audio attributes of the original speech, these attributes can be matched to the target emotion, thus obtaining speech that expresses the target emotion, i.e., the target speech. Therefore, by adjusting the audio feature attributes of the original speech, audio with the target emotion can be obtained, i.e., the target speech.

[0075] Method 2: Determine the target speech through feature extraction and feature fusion.

[0076] In the second implementation, the emotional characteristics of the target emotion can be determined, and features can be extracted from the original speech to identify its features. By fusing and transforming the emotional characteristics of the target emotion and the features of the original speech, the corresponding target speech is obtained.

[0077] In this context, both the emotional features and the features of the original speech are feature vectors. The features of the original speech are feature vectors obtained through an encoder. The emotional features are feature vectors corresponding to the target emotion, obtained through an encoder. For example, feature extraction can be performed on the audio corresponding to the target emotion to obtain the emotional features of the target emotion. In this embodiment, the encoder used to obtain the emotional features can be referred to as a first encoder, and the encoder used to obtain the features of the speech can be referred to as a second encoder.

[0078] The features of the original speech can include feature vectors related to the target object and feature vectors related to the content of the original speech. The feature vectors related to the target object can be called identity features, and the feature vectors related to the content of the original speech can be called content features. Accordingly, the second encoder can include two sub-encoders. One sub-encoder is used to determine the identity features and can be called the identity sub-encoder. The other sub-encoder is used to determine the content features and can be called the content sub-encoder.

[0079] The following sections will introduce the emotional characteristics, identity characteristics, and content characteristics respectively.

[0080] First, let's introduce the characteristics of emotions.

[0081] The emotional characteristics of a target emotion are determined based on the target emotion and represent the features of the speech produced by an object under the target emotion. Furthermore, the emotional characteristics of a target emotion are not associated with a specific object. That is, different objects can produce speech with the same emotional characteristics under the target emotion. Therefore, target speech incorporating the emotional characteristics of the target emotion also possesses the characteristics corresponding to the target emotion, enabling listeners to perceive the target emotion of the virtual character.

[0082] The following describes two methods for determining the emotional characteristics of a target emotion.

[0083] Implementation Method 1: The emotional characteristics of the target emotion are selected from multiple preset emotional characteristics.

[0084] Specifically, multiple preset emotional features can be pre-configured. Different preset emotional features correspond to different emotions, representing the characteristics of the speech emitted by an object under the corresponding emotion. After the video generation device acquires the target emotion, it can select the emotional feature corresponding to the target emotion from multiple preset emotional features as the emotional feature of the target emotion.

[0085] Method 2: The emotional features of the target emotion are obtained by extracting features from the speech.

[0086] Specifically, multiple sample emotional speech can be pre-set. Different sample emotional speech can correspond to different emotions. When the target speech needs to be generated, the sample emotional speech corresponding to the target emotion can be selected from the multiple sample emotional speech based on the target emotion, and its features can be extracted through a first encoder to obtain the emotional features of the target emotion. The selected sample emotional speech can be called the target sample emotional speech. The target sample emotional audio corresponds to the target emotion.

[0087] The sample emotional speech can be pre-set speech. Each sample emotional speech is associated with a specific emotion. Different sample emotional speech can correspond to the same emotion or different emotions. The sample emotional speech can be pre-recorded. For example, the sample emotional speech can be pre-recorded by the speech acquisition subject. The speech acquisition subject can record multiple audio clips under different emotions, thus obtaining multiple sample emotional speech.

[0088] If a target emotion corresponds to multiple sample emotional speech words, the video generation device can determine one target emotional speech word or multiple target emotional speech words. That is, when determining the emotional features of the target emotion, one sample emotional speech word can be selected from the multiple sample emotional speech words corresponding to the target emotion as the target emotional speech word. Alternatively, multiple sample emotional speech words corresponding to the target emotion (e.g., all sample emotional speech words corresponding to the target emotion) can be selected as the target emotional speech word. When extracting the target emotion features, the first encoder can extract emotional features from the multiple target emotional speech words separately, and the obtained multiple emotional features can be fused to finally obtain the emotional features of the target emotion.

[0089] The following describes the characteristics of the identity.

[0090] Here, the identity features of the target object refer to the feature vector associated with the target object. For example, if the target object is a real person, then the identity features of the target object are feature vectors that reflect attributes such as the real person's pronunciation timbre and pronunciation habits. The speech generated based on the identity features of the target object has the same or matching auditory effect as the speech directly spoken by the target object. Optionally, the video generation device may include an identity sub-encoder. After acquiring the original speech, the original speech can be input into the identity sub-encoder. The identity sub-encoder extracts features from the original speech to obtain the identity features of the target object.

[0091] The following describes the characteristics of the content.

[0092] Content features refer to feature vectors related to the content of the original speech. The content of the original speech refers to the textual expression corresponding to the original speech. Content features can reflect attributes such as phoneme information and pitch information in the input speech, excluding emotional information. Speech generated based on content features allows listeners to understand the content corresponding to the content features. Optionally, the video generation device may include a content sub-encoder. After acquiring the original speech, the original speech can be input into the content sub-encoder, and features can be extracted through the content sub-encoder to obtain the content features of the target speech.

[0093] It should be noted that the aforementioned emotional features, identity features, and content features are feature vectors obtained by the encoder through feature extraction of speech. These feature vectors can reflect certain attributes of speech. For example, the emotional features of the original speech can reflect the emotional attributes of the target audience's original speech, the identity features of the original speech can reflect attributes such as the target audience's timbre and pronunciation habits, and the content features of the original speech can reflect the content attributes of the original speech.

[0094] After obtaining the emotional features of the target emotion, the identity features of the target object, and the content features of the original speech, the video generation device can perform feature fusion on these three elements to obtain the target speech. Optionally, the video generation device may include a speech synthesis model. When target speech needs to be generated, the video generation device can input the emotional features of the target emotion, the identity features of the target object, and the content features of the original speech into the speech synthesis model, which will then perform feature fusion and generate the target speech.

[0095] In this way, the generated target speech is obtained by feature fusion based on the emotional characteristics of the target emotion, the identity characteristics of the target object, and the content characteristics of the original speech. It can not only convey the content of the original speech, but also simulate the pronunciation characteristics of the target object and reflect the target emotion. It can better achieve the effect of the target object expressing the content of the original speech with the target emotion.

[0096] The first encoder, identity sub-encoder, content sub-encoder, and speech synthesis model can all be trained from a training speech set. A description of the training model can be found below and will not be repeated here.

[0097] S203: Generate a set of target images based on the input image, emotion information, and target speech.

[0098] Through the above steps, the target speech, representing the content of the original speech with the target emotion, can be obtained, which is equivalent to determining the speech portion of the target video. To achieve the effect of simulating the target image speaking, the video generation device can determine the image portion of the target video based on the speech portion. Specifically, the video generation device can generate a target image set based on the input image, emotion information, and target speech. The target image set is the image portion of the target video. The target image set includes multiple target images. Each target image corresponds to a frame in the target video. When playing the target video, multiple target images can be displayed sequentially, thus achieving the effect of displaying the target object speaking and showing the image of the target object uttering the target speech with the target emotion. Optionally, the target image can be an image from any frame in the target video, or it can be an image from a keyframe in the target video. If the target image is an image from a keyframe in the target video, in step S204 below, a non-keyframe in the target video can be generated based on the target image set.

[0099] Specifically, when generating the target image set, the motion information of the target object during the delivery of the target speech can be determined based on the target speech. The input image is then modified based on this motion information to determine the action of the target object at each time point, resulting in the image corresponding to each time point, i.e., the target image. Here, "time point" refers to the time point corresponding to the target image in the target video. If the target object corresponds to a keyframe in the target video, then "time point" refers to the specific location of the keyframe in the target video. In other words, the video generation device can determine when the appearance of the virtual character of the target object undergoes a significant change based on the motion information, and then use the virtual character before and after the change as keyframes in the target video to obtain the corresponding target image.

[0100] In this embodiment, motion information is information used to describe the motion of one (or more) body parts of a target object. Based on the motion information, the shape and position of one (or more) body parts of the target object at a certain point in time can be determined, thereby adaptively adjusting the shape and position of the corresponding body parts of the target object in the input image to obtain the target image corresponding to that point in time.

[0101] Specifically, the video generation device may include a motion feature extraction model, an appearance feature extraction model, a motion information estimation model, and an image generation model. The motion feature extraction model extracts motion features of the target object from an image; the appearance feature extraction model extracts appearance features of the target object from an image; the motion information estimation model estimates motion information by combining motion features and output speech; and the image generation model generates a corresponding image based on the motion information and appearance features. Motion features represent the characteristics of the target object's body parts during movement, such as the constraint information of the target object's muscles. Appearance features represent the characteristics of the appearance of the target object's body parts, such as the location of organs, bone shape, skin color, and skin texture.

[0102] When generating the target image set, the input images can first be fed into a motion feature extraction model and an appearance feature extraction model respectively to obtain the motion and appearance features of the target object. Next, the motion features of the target object and the target speech can be input into a motion information estimation model to estimate the motion information of the target object during the delivery of the target speech. Then, the motion information and the appearance features of the target object can be input into an image generation model to generate the corresponding image of the target object. In this way, by separating the motion and appearance features of the target object, the movement of body parts of the target object during the delivery of the target speech can be determined first by combining the target speech and motion features, and then the corresponding image can be determined by combining the appearance features, thus completing the generation of the target image. Because the motion and appearance features are separated, the position and shape of the body parts of the target object during the delivery of the target speech can be accurately estimated.

[0103] The motion feature extraction model, appearance feature extraction model, motion information estimation model, and image generation model mentioned above can all be pre-trained. Details on the model training implementation are provided below and will not be elaborated upon here.

[0104] In this embodiment, the generated target video is used to simulate the image of a target object speaking the target speech. When a real person speaks, the main muscles that move are their facial muscles. Therefore, when generating the target video, it is necessary to simulate the facial movements of the target object. Thus, when determining the target image set, the video generation device can determine the facial movements of the target object during the speaking process based on the target speech, thereby determining the appearance of the target object's face in the target image. Accordingly, the aforementioned "one (or more) body parts of the target object" includes the target object's face, and the aforementioned motion information includes facial motion information.

[0105] In other words, the video generation device can extract features from the input image using a facial motion feature extraction model to obtain the facial motion features of the target object. Then, it can determine the facial motion information of the target object based on the target speech and facial motion features. Next, it uses a facial motion information estimation model to determine the facial motion information of the target object. If the motion of other body parts is not involved, the facial motion information and the appearance features of the target object can be input into an image generation model to obtain the target image.

[0106] In some application scenarios, there may be a need to emphasize target speech through the body movements of a virtual avatar. Therefore, in some possible implementations, the aforementioned motion information may include not only facial motion information but also limb motion information. Accordingly, the phrase "one (or more) body parts of the target object" includes the target object's face and limbs. The target object's limbs refer to any part other than the target object's face, and may include at least one of the target object's four limbs, or the target object's torso or parts of the torso. For example, if the target object needs to wave to emphasize the target video's presentation, the limbs may include the target object's left arm and / or right arm; if the target object needs to lean forward to emphasize the target video's presentation, the limbs may include the target object's upper body.

[0107] Accordingly, users can instruct the generation of facial motion effects or "face + limb" motion effects via guidance information. Specifically, if the user does not require limb motion of the digital human, the guidance information can only instruct the generation of facial-driven effects. In this case, facial motion information can be determined and a set of target images can be generated based on the facial motion information. If the user requires limb motion of the digital human, the guidance information can instruct the generation of both facial and limb-driven effects. Accordingly, limb and facial motion information of the target image can be determined, and a set of target images can be generated by combining the limb and facial motion information.

[0108] In other words, when generating the target image, the facial motion information of the target object can be obtained first through the aforementioned implementation method. Additionally, it can be determined whether the instruction information indicates the generation of a limb-driven effect. If the instruction information only indicates the generation of a facial-driven effect, a set of target images can be generated based on the facial motion information. If the instruction information indicates both facial and limb-driven effects, then the entity motion information of the target object can be determined based on the input image, emotion information, and target speech, and a set of target images can be generated based on the limb and facial motion information.

[0109] Specifically, the video generation device may include a limb motion feature extraction model and a limb motion information estimation model. The limb motion feature extraction model can extract features from the input image to obtain the limb motion features of the target object. The limb motion information estimation model can combine emotional information, target speech, and the limb motion features of the target object to determine the limb motion information of the target object. After obtaining the facial motion information and limb motion information of the target object, the target image can be generated by combining the appearance features, facial motion information, and limb motion information of the target object through an image generation model.

[0110] In the implementation methods described above, facial motion features and limb motion features are obtained through different motion feature extraction models, facial motion information and limb motion information are obtained through different motion information estimation models, and the images of the target object's face and limbs are obtained through the same image generation model. It is understandable that in some implementation methods, the different facial and limb movements mentioned above can be implemented by the same model or by different models.

[0111] To enhance the editability of virtual character avatars, some possible implementations can incorporate external control information to determine the target image information. Specifically, the user who triggers video generation can also trigger control commands. These commands represent the user's requirements for the generated target video. Accordingly, when generating the target image set, motion information and target images can be adaptively adjusted based on the control commands.

[0112] For example, control instructions can be used to indicate information such as the direction of human eye gaze, virtual acquisition distance, and expression coefficient. The direction of human eye gaze refers to the direction the generated virtual character's glasses are looking in. The virtual acquisition distance refers to the distance from the virtual camera to the virtual character. A virtual camera is a virtual camera that captures images of a real person, ensuring the captured image matches the target image. The expression coefficient represents the strength of the virtual character's facial expression in conveying the target emotion. The direction of human eye gaze and the virtual acquisition distance can be input into the image generation model to adjust the generated target image. The expression coefficient can be input into the facial motion information estimation model to adjust the strength of facial motion information in conveying the target emotion. Optionally, the control instructions may also include the aforementioned indication information.

[0113] Optionally, in a cloud service scenario, the cloud service client can display multiple functional controls to the user. Different functional controls can correspond to different functions. Users can configure control commands by triggering the controls. For example, users can configure a target emotion through the functional control corresponding to the emotion function, configure the aforementioned instruction information through the functional control corresponding to the body movement function, and configure the aforementioned human eye gaze direction through the functional control corresponding to the gaze direction function.

[0114] S204: Merge the target image set and the target speech to obtain the target video.

[0115] After obtaining the target image set and target audio, they can be merged to obtain the target video. Specifically, the target images can be inserted at the corresponding time points of the target audio, according to the timeline. In this way, when the target video is played, the audience can hear the target video and see the corresponding target images at the corresponding time points. The appearance of the virtual character in the corresponding target image matches the target object, and the facial expressions and body movements (if present) also match the target audio, achieving the effect of showing the audience the image of the target object delivering the target audio with the target emotion.

[0116] Furthermore, the audio portion of the target video is not the input audio data, but rather generated by combining the content and emotional information of the original audio. The audio portion of the target video is not limited by the emotion of the original audio, allowing for free adjustment of the emotions expressed in the target video. Moreover, the image portion of the target video is generated based on the target audio, ensuring that the emotions displayed in the target video match the emotional information. Thus, by combining emotional information to generate the audio portion of the target video, and combining the audio portion to obtain the image portion of the target video, the resulting video data accurately expresses the target emotion. Therefore, the emotion of the voice spoken by the virtual character is selectable, rather than being inherent to the original audio, essentially allowing for editing of the virtual character's emotions. This enhances the editability of voice-driven photo-realistic digital human technology and can meet complex business needs.

[0117] Optionally, the process of generating the target video can be as follows: Figure 3 As shown.

[0118] In some possible implementations, it may be necessary to adjust the background of the virtual character. For example, the background in the input image may not meet the requirements for generating the target video, and the lighting information of the target object in the input image may also fail to meet the requirements of the target video. Therefore, the background of the target images in the target image set can be adjusted. As another example, different parts of the generated target video may require different backgrounds.

[0119] Specifically, before generating the target video, if adjustments to the target images are needed, the background image to be replaced can be obtained first. Then, at least one target image is replaced with its background image, and the lighting information is recalculated to obtain a processed set of target images. Finally, the target speech and the processed set of target images are superimposed to obtain the corresponding target video. The lighting information can be matched with the lighting information of the background image. In this way, recalculating the lighting information of the target image based on the lighting information of the background image ensures that the lighting conditions of the target image in the target image are consistent with the lighting conditions of the background, improving the image display effect.

[0120] In this way, by recalculating the lighting information, the blending effect between the virtual character and the background can be guaranteed, avoiding unnatural appearances between the character and the background. Combining background image replacement allows for flexible changes to the background of the target video, improving the flexibility of virtual digital human generation.

[0121] Optionally, if it is necessary to process the i-th target image in the target image set (i is a positive integer and less than the total number of target images in the target image set), the video generation device can perform image extraction on the target image to extract the part corresponding to the target object. Then, material estimation can be performed on the extracted part, the extracted part can be superimposed on the background image, and the lighting information can be recalculated to finally obtain the processed target image. The material estimation step is used to determine the reflectivity of the target object's surface when recalculating the lighting information.

[0122] Optionally, the input background image can also be processed. For example, to facilitate the recalculation of lighting information, the dynamic range of the background image can be expanded. Specifically, the background image can be processed using a Low Dynamic Range (LDR) to High Dynamic Range (HDR) module (LDR2HDR) to adapt to the display effects of high dynamic range.

[0123] The preceding text introduced some methods for generating target videos. In the process of generating target videos, models may be needed for operations such as feature extraction, speech synthesis, and image generation. Optionally, these models can be pre-trained. The following sections describe some methods for training these models.

[0124] This section first introduces the models involved in generating the target image set and their training methods. For ease of explanation, the models used in generating the target image set will be illustrated below, including an appearance feature extraction model, a facial motion feature extraction model, a facial motion information estimation model, and an image generation model. Please refer to [link / reference needed]. Figure 4a , Figure 4a This is a schematic flowchart illustrating the model training process provided in an embodiment of this application.

[0125] During training, a training image set is first acquired. This set can include multiple training images. Each training image set includes one source image and one sample image. Both the source and sample images are captured from the same object. Furthermore, the actions of the object in the source image differ from those in the sample image. That is, the source and sample images are images of the same object's facial expression. For ease of explanation, the objects corresponding to the source and sample images will be referred to as the reference object below.

[0126] After obtaining the reference training image set, appearance features can be extracted from the source images using an appearance feature extraction model, and facial motion features can be extracted from both the source and sample images using a facial motion feature extraction model. Next, the facial motion features from the source and sample images can be input into a facial motion information estimation model. This model determines the facial motion information from the source image to the sample image. This facial motion information represents the facial movements required to transform an object's expression from the corresponding facial expression in the source image to the corresponding facial expression in the sample image.

[0127] Next, motion information can be used to distort the appearance features, resulting in distorted appearance features. These distorted features are then input into the image generator to obtain a reference image. The reference image combines the appearance features of a reference object with facial motion information from the source image to the sample image. Since the source image and the sample image correspond to the same object, the reference object should theoretically match the sample image. Therefore, the sample image and the reference image can be compared to calculate the loss function in the image generation process. This loss function represents the difference between the image generated by the appearance feature extraction model, motion feature extraction model, motion information estimation model, and image generation model and the real image.

[0128] In this way, by optimizing the appearance feature extraction model, facial motion feature extraction model, facial motion information estimation model, and image generation model based on the loss function, the goal of training these models can be achieved. For example, an image difference threshold can be preset. During model training, each model can be iteratively optimized multiple times based on the loss function until the loss calculated based on the reference image and the sample image reaches or falls below the image difference threshold.

[0129] Understandably, if the video generation device also has the function of determining limb movement information, then a similar implementation method described above can be used to train the limb movement feature extraction model and the limb movement information estimation model. Furthermore, in application scenarios involving limb movement, the source image and the sample image can be images of different limb postures of the same object.

[0130] The training methods for the models involved in generating the target image set have been introduced above. The training methods for the models involved in generating the target speech (including the first encoder, identity sub-encoder, content sub-encoder, and speech synthesis model) are introduced below.

[0131] During training, a training audio set can first be acquired. The training audio set may include multiple sets of training audio. Each set of training audio may include a first source audio segment and a second source audio segment. The emotion corresponding to the first source audio segment is different from the emotion corresponding to the second source audio segment. Optionally, the object uttering the first source audio segment and the object uttering the second source audio segment can also be different, and the content of the first source audio segment and the content of the second source audio segment can also be different. See also... Figure 4b , Figure 4b This is another schematic diagram of the model training process provided in the embodiments of this application.

[0132] After obtaining the training speech set, emotion features can be extracted from the first source speech using a first encoder, and features from the second source speech using a second encoder. Specifically, the second encoder includes a content sub-encoder and an identity sub-encoder. Content features can be extracted from the second source speech using the content sub-encoder, and identity features can be extracted from the second source speech using the identity sub-encoder. Then, the emotion features of the first source speech, the identity features of the second source speech, and the emotion features of the second source speech can be input into the speech synthesis model to obtain the first reference speech.

[0133] As described above regarding the functions of each encoder and model, without considering errors in the encoders and models, the emotion of the first reference speech is identical to that of the first source speech, and the identity and content features of the first reference speech are identical to those of the second source speech. Therefore, the first reference speech can simulate the object uttering the second source speech, expressing the content of the second source speech with the corresponding emotion of the first source speech. In other words, without considering errors, the first reference speech can simulate the object corresponding to the second source speech expressing the content of the second source speech with the emotion of the first source speech. If the expression effect of the first reference speech is poor, it indicates that there are errors in the first encoder, identity sub-encoder, content sub-encoder, and speech synthesis model.

[0134] To this end, a second reference speech can be obtained, and the differences between the first and second reference speeches can be compared. Based on these differences, the first encoder, identity sub-encoder, content sub-encoder, and speech synthesis model can be adjusted. The emotion of the second reference subject is the same as that of the first source speech, and the content of the second reference speech is the same as that of the second source speech. The second reference speech and the second source speech are spoken by the same subject. That is, the second reference speech is a real speech in which the subject corresponding to the second source speech expresses the content of the second source speech with the emotion of the first source speech. By comparing the differences between the first and second reference speeches, the difference between the speech generated by the model and the real value can be determined. Based on this difference, one or more of the first encoder, identity sub-encoder, content sub-encoder, and speech synthesis model can be adjusted in a targeted manner.

[0135] Specifically, the loss function in the feature extraction and speech synthesis processes can be calculated by comparing the first reference speech with the second reference speech. This loss function represents the difference between the speech generated by the first encoder, identity sub-encoder, content sub-encoder, and speech synthesis model and the real speech.

[0136] In this way, by optimizing the first encoder, identity sub-encoder, content sub-encoder, and speech synthesis model based on the loss function, the goal of training the first encoder, identity sub-encoder, content sub-encoder, and speech synthesis model can be achieved. For example, a speech difference threshold can be preset. During model training, each model can be iteratively optimized multiple times based on the loss function until the loss calculated based on the first and second reference speech reaches or is less than the speech difference threshold.

[0137] This application also provides a video generation apparatus. The video generation apparatus can be applied to… Figure 1 The server 20 in the implementation shown. Specifically, as... Figure 5 As shown, the video generation apparatus 500 includes:

[0138] The acquisition unit 510 is used to acquire an input image, raw audio, and emotion information. The input image includes a face image of the target object, and the emotion information is used to indicate the emotion of the target video to be generated.

[0139] The speech synthesis unit 520 is used to synthesize an emotion from the original speech based on the emotion information to obtain a target speech, wherein the target speech expresses the content of the original speech with the emotion.

[0140] The image generation unit 530 is configured to generate a target image set based on the input image, the emotion information, and the target speech. The target image set includes multiple target images, which are used to display the image of the target object uttering the target speech with the emotion.

[0141] The video generation unit 540 is used to merge the target image set and the target speech to obtain the target video.

[0142] The acquisition unit 510, speech synthesis unit 520, image generation unit 530, and video generation unit 540 can all be implemented in software or in hardware. For example, the implementation of speech synthesis unit 520 will be described below. Similarly, the implementation of acquisition unit 510, image generation unit 530, and video generation unit 540 can refer to the implementation of speech synthesis unit 520.

[0143] As an example of a software functional unit, the speech synthesis unit 520 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, or a container. Further, the aforementioned computing instance may be one or more. For example, the speech synthesis unit 520 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Further, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one or more geographically proximate data centers. Typically, a region may include multiple AZs.

[0144] Similarly, multiple hosts / virtual machines / containers used to run this code can be distributed within the same Virtual Private Cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Communication between two VPCs within the same region, as well as between VPCs in different regions, requires a communication gateway to be set up within each VPC to enable interconnection between VPCs.

[0145] As an example of a hardware functional unit, the speech synthesis unit 520 may include at least one computing device, such as a server. Alternatively, the speech synthesis unit 530 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), generic array logic (GAL), or any combination thereof.

[0146] The speech synthesis unit 520 includes multiple computing devices that can be distributed in the same region or in different regions. Similarly, the speech synthesis unit 520 can be distributed in the same Availability Zone (AZ) or in different AZs. Likewise, the speech synthesis unit 520 can be distributed in the same Virtual Private Cloud (VPC) or in multiple VPCs. These multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0147] It should be noted that, in other embodiments, the acquisition unit 510 is used to execute any step in the video generation method, the speech synthesis unit 520 is used to execute any step in the video generation method, the image generation unit 530 is used to execute any step in the video generation method, and the video generation unit 540 can be used to execute any step in the video generation method. The steps implemented by the acquisition unit 510, speech synthesis unit 520, image generation unit 530, and video generation unit 540 can be specified as needed. Thus, the acquisition unit 510, speech synthesis unit 520, image generation unit 530, and video generation unit 540 respectively implement all the functions of the video generation device based on different steps in the video generation method.

[0148] This application also provides a computing device. For example... Figure 6 As shown, the computing device 100 includes a bus 102, a processor 104, a memory 106, and a communication interface 108. The processor 104, the memory 106, and the communication interface 108 communicate with each other via the bus 102. The computing device 100 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 100.

[0149] Bus 102 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, Figure 6 The bus 102 may be represented by a single line, but this does not mean that there is only one bus or one type of bus. The bus 102 may include a path for transmitting information between various components of the computing device 100 (e.g., memory 106, processor 104, communication interface 108).

[0150] The processor 104 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0151] Memory 106 may include volatile memory, such as random access memory (RAM). Processor 104 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).

[0152] The memory 106 stores executable program code, which the processor 104 executes to implement the functions of the aforementioned acquisition unit 510, speech synthesis unit 520, image generation unit 530, and video generation unit 540, thereby realizing the video generation method. In other words, the memory 106 stores instructions for executing this stored method.

[0153] Alternatively, the memory 106 may store executable code, which the processor 104 executes to implement the functions of the aforementioned video generation apparatus, thereby implementing the video generation method. That is, the memory 106 stores instructions for executing the video generation method.

[0154] The communication interface 108 uses transceiver modules such as, but not limited to, network interface cards and transceivers to enable communication between the computing device 100 and other devices or communication networks.

[0155] This application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.

[0156] like Figure 7 As shown, the computing device cluster includes at least one computing device 100. The memory 106 of one or more computing devices 100 in the computing device cluster may store the same instructions for executing the video generation method.

[0157] In some possible implementations, the memory 106 of one or more computing devices 100 in the computing device cluster may also store partial instructions for executing the video generation method. In other words, a combination of one or more computing devices 100 can jointly execute instructions for executing the base video generation method.

[0158] It should be noted that the memory 106 in different computing devices 100 within the computing device cluster can store different instructions, which are used to execute some functions of the video generation device 500. That is, the instructions stored in the memory 106 of different computing devices 100 can implement the functions of one or more units among the first determining module 510, the second determining module 520, the image generation unit 530, and the video generation unit 540.

[0159] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc. Figure 8 One possible implementation is shown. For example... Figure 8 As shown, the two computing devices 100A and 100B are connected via a network. Specifically, they are connected to the network through communication interfaces in each computing device. In this possible implementation, the memory 106 in computing device 100A stores instructions for executing the acquisition unit 510. Meanwhile, the memory 106 in computing device 100B stores instructions for executing the functions of the speech synthesis unit 520, the image generation unit 530, and the video generation unit 540.

[0160] Figure 8 The connection method between the computing device clusters shown can be as follows: considering that the video generation method provided in this application can be divided into two parts, namely calling the decision model and interacting with the computing cluster, the function of calling the decision model is to be executed by computing device 100A, and the function of interacting with the computing cluster is to be executed by computing device 100B.

[0161] It should be understood that Figure 8 The functions of the computing device 100A shown can also be performed by multiple computing devices 100. Similarly, the functions of the computing device 100B can also be performed by multiple computing devices 100.

[0162] This application also provides another computing device cluster. The connection relationships between the computing devices in this computing device cluster can be similarly referred to... Figure 7 and Figure 8 The connection method of the computing device cluster is different in that the memory 106 of one or more computing devices 100A in the computing device cluster can store the same instructions for executing the video generation method.

[0163] In some possible implementations, the memory of one or more computing devices 100B in the computing device cluster may also store partial instructions for executing the video generation method. In other words, a combination of one or more computing devices can jointly execute the instructions for executing the video generation method.

[0164] It should be noted that the memory 106 in different computing devices 100A within the computing device cluster can store different instructions for executing some functions of the video generation apparatus. That is, the instructions stored in the memory 106 of different computing devices 100A can implement the functions of one or more units within the video generation apparatus.

[0165] This application also provides a computer program product containing instructions. The computer program product may be a software or program product containing instructions, capable of running on a computing device or stored on any usable medium. When the computer program product is run on at least one computing device, it causes the at least one computing device to perform a video generation method.

[0166] This application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store, or a data storage device such as a data center that includes one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to perform a video generation method, or instruct the computing device to perform a video generation method.

[0167] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.

Claims

1. A video generation method, characterized in that, The method includes: The system acquires an input image, raw audio, and emotional information, wherein the input image includes a facial image of the target object, and the emotional information is used to indicate the emotion of the target video to be generated. Based on the emotional information, the original speech is processed to synthesize an emotion to obtain a target speech, which expresses the content of the original speech with the emotion. A set of target images is generated based on the input image, the emotion information, and the target speech. The set of target images includes multiple target images, which are used to display the image of the target object uttering the target speech with the emotion. The target image set and the target speech are merged to obtain the target video.

2. The method according to claim 1, characterized in that, The method further includes: The audio attributes of the original speech are determined based on the original speech. The step of synthesizing emotion based on the emotion information to obtain target speech includes: Based on the emotion indicated by the emotion information, the audio attributes of the original speech are adjusted to generate the target speech, and the audio attributes of the target speech are matched with the emotion indicated by the emotion information.

3. The method according to claim 1, characterized in that, The original speech is the speech emitted by the target object, and the process of synthesizing emotion from the original speech based on the emotion information to obtain the target speech includes: Feature extraction is performed on the original speech to determine the features of the original speech. The features of the original speech include features related to the target object and features related to the content of the original speech. Determine emotional characteristics based on the emotional information; The features of the original speech and the emotion features are fused to obtain the target speech.

4. The method according to any one of claims 1 to 3, characterized in that, The generation of a target image set based on the input image, the emotion information, and the target speech includes: Based on the input image and the target speech, determine the facial motion information of the target object; Obtain instruction information; If the indication information indicates the generation of a face-driven effect, the target image set is generated based on the face motion information; If the instruction information indicates the generation of the face-driven effect and the body-driven effect, the body movement information of the target object is determined based on the input image, the emotion information and the target speech, and the target image set is generated based on the body movement information and the face movement information.

5. The method according to any one of claims 1 to 4, characterized in that, The process of merging the target image set and the target speech to obtain the target video includes: Get the background image; The background of at least one target image is replaced with the background image, and the lighting information is calculated based on the background image and the lighting is re-updated to obtain multiple processed target images. The processed set of target images and the target speech are combined to obtain the target video.

6. The method according to claim 3, characterized in that, The emotion features are extracted by the first encoder, the features of the original speech are extracted by the second encoder, and the target speech is obtained by a speech synthesis model. The first encoder, the second encoder, and the speech synthesis model are trained in the following manner: Acquire a first source speech and a second source speech, wherein the emotion corresponding to the first source speech and the emotion corresponding to the second source speech are different; The first encoder is used to extract features from the first source audio to obtain the emotional features of the emotion corresponding to the first source speech. The second source speech is feature extracted by the second encoder to obtain the features of the second source speech; By combining the emotional features of the emotion corresponding to the first source speech and the features of the second source speech through the speech synthesis model, a first reference speech is generated. The first reference speech and the second reference speech are compared, and the first encoder, the second encoder, and the speech synthesis model are adjusted according to the comparison result. The second reference speech corresponds to the same emotion as the first source speech, the content of the second reference speech is the same as the content of the second source speech, and the second reference speech and the second source speech are spoken by the same person.

7. The method according to claim 3, characterized in that, Determining emotional characteristics based on the emotional information includes: Based on the emotional information, a target sample emotional audio is determined from multiple sample emotional audios, and different sample emotional audios correspond to different emotions. The target sample emotional audio is subjected to feature extraction to obtain the emotional features.

8. A video generation apparatus, characterized in that, The device includes: An acquisition unit is used to acquire an input image, raw audio, and emotional information. The input image includes a facial image of the target object, and the emotional information is used to indicate the emotion of the target video to be generated. A speech synthesis unit is used to synthesize an emotion from the original speech based on the emotion information to obtain a target speech, wherein the target speech expresses the content of the original speech with the emotion. An image generation unit is configured to generate a target image set based on the input image, the emotion information, and the target speech. The target image set includes multiple target images, which are used to display the image of the target object uttering the target speech with the emotion. A video generation unit is used to merge the target image set and the target speech to obtain a target video.

9. The apparatus according to claim 8, characterized in that, The speech synthesis unit is specifically used to determine the audio attributes of the original speech based on the original speech; adjust the audio attributes of the original speech according to the emotion indicated by the emotion information, and generate the target speech, wherein the audio attributes of the target speech match the emotion indicated by the emotion information.

10. The apparatus according to claim 8, characterized in that, The original speech is the speech emitted by the target object. The speech synthesis unit is specifically used to extract features from the original speech, determine the features of the original speech, the features of the original speech include features related to the target object and features related to the content of the original speech; and determine emotional features based on the emotional information. The features of the original speech and the emotion features are fused to obtain the target speech.

11. The apparatus according to any one of claims 8 to 10, characterized in that, The acquisition unit is also used to acquire indication information; The image generation unit is specifically used to determine the facial motion information of the target object based on the input image and the target speech; if the indication information indicates the generation of a facial driving effect, the unit generates the target image set based on the facial motion information. If the instruction information indicates the generation of the face-driven effect and the body-driven effect, the body movement information of the target object is determined based on the input image, the emotion information and the target speech, and the target image set is generated based on the body movement information and the face movement information.

12. The apparatus according to any one of claims 8 to 10, characterized in that, The image generation unit is specifically used to acquire a background image; replace the background of the at least one target image with the background image; calculate illumination information based on the background image and relight the image to obtain multiple processed target images. The video generation unit is specifically used to merge the processed target image set and the target speech to obtain the target video.

13. The apparatus according to claim 10, characterized in that, The emotion features are extracted by the first encoder, the features of the original speech are extracted by the second encoder, and the target speech is obtained by the speech synthesis model. The first encoder, the second encoder, and the speech synthesis model are trained in the following manner: Acquire a first source speech and a second source speech, wherein the emotion corresponding to the first source speech and the emotion corresponding to the second source speech are different; The first encoder is used to extract features from the first source audio to obtain the emotional features of the emotion corresponding to the first source speech. The second source speech is feature extracted by the second encoder to obtain the features of the second source speech; By combining the emotional features of the emotion corresponding to the first source speech and the features of the second source speech through the speech synthesis model, a first reference speech is generated. The first reference speech and the second reference speech are compared, and the first encoder, the second encoder, and the speech synthesis model are adjusted according to the comparison result. The second reference speech corresponds to the same emotion as the first source speech, the content of the second reference speech is the same as the content of the second source speech, and the second reference speech and the second source speech are spoken by the same person.

14. The apparatus according to claim 10, characterized in that, The speech synthesis unit is specifically used to determine a target sample emotional audio from multiple sample emotional audios based on the emotional information, wherein different sample emotional audios correspond to different emotions; and to extract features from the target sample emotional audio to obtain the emotional features.

15. A computing device cluster, characterized in that, The computing device cluster includes at least one computing device, each computing device including a processor and memory: The memory is used to store instructions; The processor is configured to, according to the instructions, cause the computing device cluster to perform the method according to any one of claims 1 to 7.

16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed on a computing device, cause the computing device to perform the method as described in any one of claims 1 to 7.

17. A computer program product comprising instructions that, when run on a computing device, cause the computing device to perform the method as claimed in any one of claims 1 to 7.