Data processing method and device, electronic equipment, storage medium and program product

By discretizing the motion and voice information of digital human videos, the problem of missing details in existing technologies is solved, and more natural and coherent digital human videos are generated.

CN120690193APending Publication Date: 2025-09-23MASHANG CONSUMER FINANCE CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510829836.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

When generating digital human videos, existing technologies tend to miss out on detailed information, leading to inaccurate understanding and affecting the generation effect of video content.

Method used

By discretizing the action information of the digital human speech video, the feature vectors of the action and voice information are extracted respectively, and a more natural and coherent video is generated through a large multimodal model.

Benefits of technology

It achieves accurate extraction and separation of different actions and voice information. The generated video retains multiple action details, ensures the logical connection between the reaction actions and the scene, and improves the generation effect of digital human videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120690193A_ABST
    Figure CN120690193A_ABST
Patent Text Reader

Abstract

The invention provides a data processing method and device, electronic equipment, a storage medium and a program product. The method comprises the following steps: for a first video of a first digital person, performing first coding on each type of action information to obtain a first feature vector of each type of action information, and performing second coding on a first voice of the first digital person to obtain a second feature vector; generating a third feature vector of a second digital person based on the first feature vector, the second feature vector and a first cue word; for each type of action information, performing first decoding corresponding to the action information on the third feature vector to obtain a fourth feature vector corresponding to the action information; and generating a second video of the second digital human based on the fourth feature vector. According to the method and the device, discretization coding based on the action information can be performed on the digital human speaking video, so that the generation effect of the digital human listening video is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a data processing method, device, electronic device, storage medium, and program product. Background Art

[0002] The process of communication between two people is the foundation of effective information exchange in social interaction. This conversation involves alternating roles between the speaker and listener. The speaker conveys information through verbal and non-verbal means such as voice, facial expressions, and head posture, while the listener provides real-time feedback through non-verbal behaviors such as nodding, smiling, and shaking their head. With the development of artificial intelligence technology, the demand for human-computer interaction scenarios has increased. Human-computer interaction based on digital humans has been widely used in various practical fields, including customer service and online education.

[0003] Related technologies combine the understanding and generation capabilities of large models to comprehensively understand the speaker's video and then further generate the listener's video. However, the holistic understanding of the video will lead to the omission of detailed information, which will lead to an incorrect understanding of the video content. Therefore, the generation effect of the listener's video is not good. Summary of the Invention

[0004] The embodiments of the present application provide a data processing method, device, electronic device, storage medium, and program product, which can optimize the generation effect of the digital human listening video by discretizing the digital human speaking video based on the information of each action.

[0005] The technical solution of the embodiment of the present application is implemented as follows:

[0006] This embodiment of the present application provides a data processing method, the method comprising:

[0007] For the first video of the first digital human, perform a first encoding on each type of action information to obtain a first feature vector for each type of action information, and perform a second encoding on the first voice of the first digital human to obtain a second feature vector;

[0008] generating a third feature vector of the second digital person based on the first feature vector, the second feature vector, and the first prompt word;

[0009] For each type of action information, performing a first decoding corresponding to the action information on the third eigenvector to obtain a fourth eigenvector corresponding to the action information;

[0010] A second video of the second digital human is generated based on the fourth feature vector.

[0011] An embodiment of the present application provides a data processing device, including:

[0012] an encoding module configured to perform a first encoding on each type of action information in a first video of a first digital human to obtain a first feature vector for each type of action information, perform a second encoding on the first speech of the first digital human to obtain a second feature vector, and generate a third feature vector for the second digital human based on the first feature vector, the second feature vector, and a first prompt word;

[0013] The video generation module is used to perform a first decoding of the third eigenvector corresponding to each type of action information to obtain a fourth eigenvector corresponding to the action information; and generate a second video of the second digital human based on the fourth eigenvector.

[0014] An embodiment of the present application provides an electronic device, comprising:

[0015] a memory for storing computer-executable instructions or computer programs;

[0016] The processor is used to implement the data processing method provided in the embodiment of the present application when executing the computer-executable instructions or computer programs stored in the memory.

[0017] An embodiment of the present application provides a computer-readable storage medium storing a computer program or computer-executable instructions for implementing the data processing method provided in the embodiment of the present application when executed by a processor.

[0018] An embodiment of the present application provides a computer program product, including a computer program or computer-executable instructions. When the computer program or computer-executable instructions are executed by a processor, the data processing method provided in the embodiment of the present application is implemented.

[0019] The embodiments of the present application have the following beneficial effects:

[0020] For the first video of the first digital human, a first encoding is performed for each action information, resulting in a first feature vector for each action information. Different actions are then discretized and encoded separately to accurately extract independent features between them. A second encoding is performed for the first speech of the first digital human, resulting in a second feature vector. The action information and speech information are discretized and encoded separately to avoid interference between different modal information, enabling separation and processing of the first digital human's multimodal features, providing a basis for subsequent feature sequence generation. Based on the first and second feature vectors corresponding to the various action information, as well as the first prompt word, multimodal information (action, speech, and prompt word) is jointly processed to generate a third feature vector for the second digital human. This third feature vector integrates the multiple features of action, speech, and prompt word, ensuring the logical association of the second digital human's reaction movements with the scene. The third feature vector is then decomposed and decoded for each action information, recombining the discrete action codes into a complete video output, preserving the details of the various actions and producing a more natural and coherent second video of the second digital human. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 This is a schematic diagram of an application mode of the data processing method provided in an embodiment of the present application;

[0022] Figure 2 is a structural diagram of an electronic device provided in an embodiment of the present application;

[0023] Figure 3A This is a first flow chart of the data processing method provided in an embodiment of the present application;

[0024] Figure 3B This is a second flow chart of the data processing method provided in an embodiment of the present application;

[0025] Figure 3C 3 is a schematic diagram of a third flow chart of the data processing method provided in an embodiment of the present application;

[0026] Figure 4A This is a schematic diagram of the digital human video generation process provided by an embodiment of the present application;

[0027] Figure 4B Schematic diagram of the coding side structure of the digital human generation model provided in the embodiment of the present application;

[0028] Figure 4C Schematic diagram of the decoding side structure of the digital human generation model provided in an embodiment of the present application;

[0029] Figure 5 4 is a schematic diagram of a fourth flow chart of a data processing method provided in an embodiment of the present application;

[0030] Figure 6It is a schematic diagram of the structure of the PD-FGC algorithm provided in the embodiment of the present application;

[0031] Figure 7 This is a schematic diagram of the discrete encoder structure of a specified dimension provided in an embodiment of the present application.

[0032] It should be pointed out that the above-mentioned "first" and "second" are only used to distinguish different solutions, and do not represent the degree of distinction between the advantages and disadvantages of the solutions or the priority in the implementation process. DETAILED DESCRIPTION

[0033] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0034] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0035] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0036] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0037] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meanings as those commonly understood by those skilled in the art. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0038] The collection and processing of relevant data (e.g., interactive information of digital people) in the embodiments of this application should be strictly in accordance with the requirements of relevant laws and regulations when applied in instances, and the informed consent or separate consent of the personal information subject should be obtained. Subsequent data use and processing should be carried out within the scope of authorization of laws and regulations and the personal information subject.

[0039] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.

[0040] 1) Digital Human: This refers to a virtual character constructed through technologies such as computer graphics, natural language processing, and deep learning. It possesses human appearance, language skills, and interactive behaviors. Digital Humans have multimodal interactive capabilities, including voice dialogue, facial expression feedback, and action response. Based on the technical implementation method, Digital Humans can be divided into real-person driven and artificial intelligence driven types. In the embodiments of this application, the Digital Human is an artificial intelligence-driven Digital Human, which autonomously generates interactive behaviors through a large multimodal model. In a conversation scenario, the Digital Human can play the role of either speaker or listener.

[0041] 2) Large Multimodal Models (LMMs): These are artificial intelligence models that can understand and process multiple input forms, including images, videos, audio, and other modalities. The core of large multimodal models lies in how to efficiently model and fuse data from different modalities.

[0042] 3) Vector Quantization (VQ): It is a data compression and encoding technology based on clustering. Vector quantization achieves efficient data representation and transmission by mapping a high-dimensional continuous signal space into a finite discrete codebook.

[0043] 4) Codebook: This is a collection of discrete vectors, called code vectors, which are used to discretize the continuous latent space representation. By selecting the closest code vector from the codebook to replace the continuous latent vector output by the vector quantized variational autoencoder, data compression and representation learning are achieved.

[0044] 5) Vector Quantized-Variational Autoencoder (VQ-VAE): This is a variational autoencoder that incorporates vector quantization technology. It is primarily used for generative modeling and representation learning tasks. By introducing discrete latent vector representations, the VQ-VAE can generate high-quality samples and learn a more structured latent space. This discretization process helps the model learn a more regular and structured latent space, improving the quality of generative tasks.

[0045] 6) PD-FGC (Progressive Disentangled Fine-Grained Control) algorithm: This is a general digital human algorithm based on a single image. Based on a facial image of a target person, it can generate a digital human video of that person, without the need for additional model training for the target person. The PD-FGC algorithm achieves control of multiple fine motion attributes by decoupling the latent representation.

[0046] 7) Text-to-Speech (TTS): It is an artificial intelligence technology that automatically converts written text into audible speech signals. The goal of text-to-speech is to simulate the human speech generation process through algorithms and realize machine speech output of natural language.

[0047] Conversational communication, in which speakers and listeners alternate roles, is the foundation of effective information exchange in social interactions. Speakers convey information through verbal and nonverbal means such as voice, facial expressions, and head posture, while listeners provide real-time feedback through nonverbal behaviors such as nodding, smiling, and shaking their heads. With the increasing demand for human-computer interaction scenarios, human-computer interaction based on digital humans has been widely applied in multiple practical fields.

[0048] Related technologies combine the understanding and generation capabilities of large models to segment interactive videos, extract interactive information from the images, convert the extracted interactive information into a discrete sequence, analyze and predict the converted discrete sequence based on the model, determine the characteristic sequence of the digital human video, and achieve the generation of digital human videos in conversational scenarios. The overall conversion of images or videos will contain a large amount of redundant information. The extracted discrete sequence contains redundant information that has little impact on the interaction. Important information is coupled with non-important information, which is not conducive to the use of models and algorithms. It is impossible to highlight the more important information in the characteristics of the interactive scene, resulting in low performance and interactive capabilities of the synthesized digital human videos.

[0049] The present invention provides a data processing method, a data processing device, an electronic device, a computer-readable storage medium, and a computer program product, which can optimize the generation effect of a digital human listening video by discretizing the digital human's speech video based on the information of each action.

[0050] The following describes exemplary applications of the electronic devices provided in the embodiments of the present application. The devices provided in the embodiments of the present application can be implemented as various types of terminals, such as laptops, tablet computers, desktop computers, set-top boxes, smartphones, smart speakers, smart watches, smart TVs, and in-vehicle terminals. They can also be implemented as servers. The following describes exemplary applications when the devices are implemented as terminals or servers.

[0051] See also Figure 1 , Figure 1 This is a schematic diagram of an application mode of the data processing method provided in an embodiment of the present application, for example, to support a data processing application. Figure 1 The server 200, network 300, terminal device 400 and database 500 are involved. The terminal device 400 is connected to the server 200 via the network 300. The network 300 can be a wide area network or a local area network, or a combination of the two.

[0052] In some embodiments, the user may be a person skilled in the art, the server 200 is a server for data processing, the terminal device 400 is a terminal operated by the user, and the terminal device 400 is installed with an application capable of displaying the audio and video interaction of the digital human, and displays the video of the digital human on the human-computer interaction interface 100. The database 500 stores historical interaction data. The terminal device 400 sends the video of the speaker in the current round to the server 200 via the network 300. The server 200 generates the video of the listener in the current round based on the video of the speaker in the current round and the historical interaction data through the data processing method provided in the embodiments of the present application, and displays it on the human-computer interaction interface 100 on the terminal device 400.

[0053] The data processing method provided in the embodiments of the present application can be applied to various scenarios that require digital human video generation, such as customer service scenarios, game interaction scenarios, online education interaction scenarios, etc., as illustrated below.

[0054] 1) Customer service scenarios: For example, a terminal device receives a video of the customer in the current round, and the server generates a video of the customer service staff in the current round through data processing methods, returns it to the terminal device and presents it to the customer, supporting dialogue interaction between the customer and the customer service staff to solve user problems.

[0055] 2) Game interaction scenarios, such as when a terminal device receives a player's video of the current round, the server generates a video of the virtual character of the current round through data processing methods, returns it to the terminal device and presents it to the player, and makes different dialogue responses based on the player's different choices, thereby increasing the immersion of the game.

[0056] 3) Online education interaction scenarios, for example, the terminal device receives the student video of the current round, the server generates the teacher video of the current round through data processing methods, returns it to the terminal device and presents it to the student to assist in completing the teaching task.

[0057] See also Figure 2 , Figure 2 is a structural diagram of an electronic device provided in an embodiment of the present application, Figure 2 The terminal device 400 shown includes: at least one processor 410, a memory 450, at least one network interface 420, and a user interface 430. The various components in the terminal device 400 are coupled together via a bus system 440. It is understood that the bus system 440 is used to achieve connection and communication between these components. In addition to including a data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, the bus system 440 is not shown in FIG. Figure 2 Various buses are labeled as bus system 440 .

[0058] The processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0059] The user interface 430 includes one or more output devices 431 that enable presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.

[0060] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 450 may optionally include one or more storage devices that are physically remote from the processor 410.

[0061] The memory 450 includes volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.

[0062] In some embodiments, the memory 450 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplified below.

[0063] Operating system 451, including system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, and driver layer, which are used to implement various basic services and process hardware-based tasks;

[0064] A network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420. Exemplary network interfaces 420 include Bluetooth, Wi-Fi, and Universal Serial Bus (USB);

[0065] a presentation module 453 for enabling presentation of information via one or more output devices 431 (e.g., a display screen, a speaker, etc.) associated with the user interface 430 (e.g., a user interface for operating peripheral devices and displaying content and information);

[0066] The input processing module 454 is configured to detect one or more user inputs or interactions from one of the one or more input devices 432 and to translate the detected inputs or interactions.

[0067] In some embodiments, the apparatus provided in the embodiments of the present application may be implemented in software. Figure 2 The data processing device 455 stored in the memory 450 is shown. This device can be software in the form of a program or plug-in, and includes the following software modules: an encoding module 4551 and a video generation module 4552. These modules are logical and can be arbitrarily combined or further separated according to the functions they implement. The functions of each module will be described below.

[0068] In some embodiments, the terminal or server can implement the data processing method provided by the embodiment of the present application by running various computer executable instructions or computer programs. For example, computer executable instructions can be commands, machine instructions or software instructions at the microprogram level. The computer program can be a native program or software module in an operating system; it can be a local (Native) application (APPlication, APP); it can also be a small program that can be embedded in any APP, that is, a program that can be run only by downloading it to a browser environment. In short, the above-mentioned computer executable instructions can be instructions in any form, and the above-mentioned computer program can be an application, module or plug-in in any form.

[0069] The data processing method provided in the embodiment of the present application will be described in conjunction with the exemplary application and implementation of the terminal provided in the embodiment of the present application.

[0070] The following describes the data processing method provided by the embodiment of the present application. As mentioned above, the electronic device that implements the data processing method of the embodiment of the present application can be a terminal device, a server, or a combination of the two. Therefore, the execution entity of each step will not be repeated below.

[0071] See also Figure 3A , Figure 3A This is a first flow chart of the data processing method provided in the embodiment of the present application, which will be combined with Figure 3A The steps shown are explained.

[0072] In step 301, for a first video of a first digital human, a first encoding is performed on each action information to obtain a first feature vector for each action information, and a second encoding is performed on a first voice of the first digital human to obtain a second feature vector.

[0073] In some embodiments, see Figure 3B , Figure 3B This is a second flow chart of the data processing method provided in the embodiment of the present application. The first feature vector in step 301 can be obtained by executing Figure 3B Steps 3011 to 3013 in are implemented as described below.

[0074] In step 3011, the fifth eigenvector of each action information in the first video is extracted.

[0075] As an example, a digital human is an artificial intelligence-driven digital human that autonomously generates interactive behaviors through a multimodal large model. In a conversation scenario, the role of the first digital human can be a speaker or a listener. For the first video in which the first digital human is a speaker, the encoder in the PD-FGC algorithm performs feature extraction processing on each type of action information (such as eye movement, head movement, and expression change) in the first video, and extracts the fifth eigenvector corresponding to each type of action information, including: eye movement feature vector, head movement feature vector, and expression change feature vector. These fifth eigenvectors are the basis for subsequent discretization coding and can effectively characterize the various action information of the first digital human in the video. Action information may include eye movement, head movement, expression movement, body movement, etc. The PD-FGC algorithm is a general digital human algorithm based on a single image. Based on a facial image of a target person, a digital human video of the person can be generated without the need for additional model training for the target person. The PD-FGC algorithm achieves control of multiple fine motion attributes by decoupling potential representations.

[0076] In some embodiments, see Figure 4B , Figure 4B It is a schematic diagram of the encoding side structure of the digital human generation model provided in an embodiment of the present application; in the encoding module 4021, first, the action information in the first video 401 is feature extracted by the PD-FGC algorithm encoder 421 to obtain an eye movement feature sequence 422, a head movement feature sequence 423 and an expression change feature sequence 424 (corresponding to the fifth feature vector corresponding to each action information in the above step 3011).

[0077] In step 3012 , for each type of action information, a first dimension transformation is performed on the fifth eigenvector of the action information to obtain a sixth eigenvector of the action information.

[0078] Here, the third eigenvector is generated by a language model, and the sum of the dimensions of the first eigenvectors of multiple action information is the same as the dimensionality requirement of the input data of the language model. The language model is used to be called to generate the third eigenvector corresponding to the second digital person as the listener. The language model is an artificial intelligence model built based on deep learning technology. In the embodiment of the present application, the language model is a large multimodal model that can understand and process multiple input forms and model and fuse data of different modalities.

[0079] As an example, for each type of action information, the fifth eigenvector corresponding to the action information is transformed into the first dimension. Since the dimension of the fifth eigenvector may be inconsistent with the input dimension of the language model, the system adjusts the dimension of the fifth eigenvector through a reversible layer. A reversible layer is added after the vector quantized variational autoencoder of the specified dimension. The reversible layer is implemented by a reversible neural network. The fifth eigenvector is processed through a series of reversible transformations (for example, reversible linear transformations and nonlinear transformations) to ensure that the information is not lost during the dimensional conversion process. After processing by the reversible layer, the fifth eigenvector of each type of action information is converted into a sixth eigenvector. The dimension of the sixth eigenvector is consistent with the input dimension of the language model, including: the sixth eigenvector corresponding to eye movement, the sixth eigenvector corresponding to head movement, and the sixth eigenvector corresponding to facial expression changes. These sixth eigenvectors will be used as input for subsequent language model processing and generation tasks.

[0080] See also Figure 4B The eye reversible layer 4222 performs a first-dimensional transformation on the fifth eigenvector of eye movement, the head reversible layer 4232 performs a first-dimensional transformation on the fifth eigenvector of head movement, and the expression reversible layer 4242 performs a first-dimensional transformation on the fifth eigenvector of expression change to obtain the sixth eigenvector corresponding to eye movement, the sixth eigenvector corresponding to head movement, and the sixth eigenvector corresponding to expression change, thereby ensuring that the encoded information is completely decoded and that information will not be lost when the vector dimension changes.

[0081] In step 3013, for each type of action information, a first preset vector matching the sixth eigenvector is determined from a plurality of first preset vectors included in the first vector codebook of the action information as the first eigenvector of the action information.

[0082] As an example, the first vector codebook is a set containing a large number of discrete vectors, which are called code vectors and are used to discretize the continuous latent space representation. By selecting the closest code vector from the first vector codebook to replace the continuous latent vector output by the vector quantization variational autoencoder, data compression and representation learning are achieved. The first vector codebook is a codebook corresponding to each type of action information in the first video, and can effectively capture the main changes in the sixth eigenvector of the first video. The first vector codebook is determined by a pre-trained vector quantization variational autoencoder and includes multiple first preset vectors for matching the sixth eigenvector of the action information. For each type of action information, the sixth eigenvector closest to the first preset vector is matched as the first eigenvector of the action information.

[0083] For each action information, the sixth eigenvector of the action information is first discretized and encoded using a vector quantized variational autoencoder (VQ-VAE) of a specified dimension. The pre-trained vector quantized variational autoencoders include eye VQ-VAE model, head VQ-VAE model and expression VQ-VAE model. V For training, video collection D V It is a collection of video clips from all samples, regardless of whether they are speaking or listening videos. During the training process of the pre-trained vector quantized variational autoencoder, the eye VQ-VAE model, the head VQ-VAE model, and the expression VQ-VAE model respectively map the input eye movement feature vector, head movement feature vector, and expression change feature vector to the latent space to obtain continuous latent vectors. The vector quantization layer quantizes these continuous latent vectors to the nearest codebook vector to achieve discretization representation. Each vector in the first vector codebook represents a typical action feature pattern (eye movement pattern, head movement pattern, expression change pattern). By minimizing the reconstruction loss and vector quantization loss, the parameters of the eye VQ-VAE model, the head VQ-VAE model, and the expression VQ-VAE model are optimized to obtain the pre-trained eye VQ-VAE model, the head VQ-VAE model, the expression VQ-VAE model, and the final first vector codebook.

[0084] For each type of action information (eye movement, head movement, and expression change), the corresponding vector quantization variational autoencoder model is used to map the sixth eigenvector to a predefined first vector codebook. The vector quantization variational autoencoder uses vector quantization technology to map the continuous fifth eigenvector to a discrete vector representation. The eye movement feature vector is mapped to the first vector codebook for eye movement to obtain the first preset vector for eye movement; the head movement feature vector is mapped to the first vector codebook for head movement to obtain the first preset vector for head movement; the expression change feature vector is mapped to the first vector codebook for expression change to obtain the first preset vector for expression change. These discretized first preset vectors can effectively represent the action information and facilitate subsequent processing and generation.

[0085] For example, a video collection D VIt contains video clips of multiple different people in various scenes. In these video clips, the characters have different eye movements (rapid glances, blinks, staring, etc.), head movements (turning, nodding, shaking heads, etc.) and facial expressions (smiling, frowning, surprise, etc.). The first vector codebook is a dictionary of action features such as eyes, head, and expressions. For eye movements, each first preset vector in this dictionary represents a typical eye movement pattern. For example, a first preset vector represents a quick glance to the left, and another first preset vector represents a slow blink; for head movements, the code vector may correspond to action patterns such as a large turn of the head and a slight nod; for facial expression changes, the code vector may correspond to expression patterns such as grinning and pouting. These code vectors have specific coordinate positions in multidimensional space and can cover various common action feature changes.

[0086] Video Collection D V The system includes multiple video clips. For each video clip, the sixth eigenvector corresponding to eye movement, head movement, and facial expression change is extracted. Taking eye movement as an example, an eye tracking algorithm is used to obtain characteristic parameters such as eye position, movement speed, and open / closed state in each frame of the video, forming an eye movement feature sequence. Based on the eye movement feature vectors, head movement feature vectors, and facial expression change feature vectors, encoders for the eye VQ-VAE model, head VQ-VAE model, and facial expression VQ-VAE model are constructed respectively.

[0087] Taking the eye VQ-VAE model as an example, the encoder of the eye VQ-VAE model is a neural network that receives as input a sequence of eye movement features. Assuming the input is a sequence of eye movement features of length n, where each eye movement feature is a vector, the encoder processes the eye movement feature vectors and maps them into a latent space, resulting in a continuous latent vector. This latent vector can be understood as a more compact, higher-level representation of eye movement features, integrating various characteristic features of eye movement. The continuous latent vector is input into the vector quantization layer, which calculates the similarity between the latent vector and each code vector in the first vector codebook of eye movement. The latent vector is then quantized to obtain the closest first preset vector in the first vector codebook. This is equivalent to finding the first preset vector in the dictionary that best represents the current eye movement feature, achieving a discretized representation. If a latent vector for an eye movement is closer to the code vector for a quick left saccade in the first vector codebook, it is quantized to that code vector. The quantized code vector is then fed into the decoder of the VQ-VAE model. The decoder of the eye VQ-VAE model is also a neural network structure, which reconstructs the original eye movement feature vector based on the received code vector so that the reconstructed eye movement feature vector is as close as possible to the original input sequence.

[0088] By comparing the difference between the eye movement feature vector reconstructed by the decoder and the eye movement feature vector of the original input, the reconstruction loss can be calculated by the mean square error. The difference between the latent vector and the first preset vector in the first vector codebook during the vector quantization process is calculated to determine the vector quantization loss. Combining the reconstruction loss and vector quantization loss, the parameters of the eye VQ-VAE model are updated through the backpropagation algorithm, including the parameters of the encoder, vector quantization layer and decoder in the eye VQ-VAE model. By continuously repeating the process of updating parameters, the eye VQ-VAE model is iteratively trained for multiple rounds until the reconstruction loss and vector quantization loss of the eye VQ-VAE model reach the preset number of iterations. At this time, the pre-trained eye VQ-VAE model and the first vector codebook corresponding to the eye movement are obtained. The above training steps are also applicable to the pre-training of the head VQ-VAE model and the expression VQ-VAE model.

[0089] See also Figure 4B , input the eye movement feature sequence 422 into the eye discrete encoder 4221 for mapping processing, input the head movement feature sequence 423 into the head discrete encoder 4231 for mapping processing, and input the expression change feature sequence 424 into the expression discrete encoder 4241 for mapping processing. The eye discrete encoder 4221, the head discrete encoder 4231, and the expression discrete encoder 4241 are vector quantized variational autoencoders of specified dimensions. The eye discrete encoder 4221, the head discrete encoder 4231, and the expression discrete encoder 4241 respectively perform mapping processing based on the first vector codebook on the input eye movement feature sequence 422, the head movement feature sequence 423, and the expression change feature sequence 424 to obtain a discrete latent vector representation. The first preset vector corresponding to the sixth feature vector in the first vector codebook includes: the first preset vector for eye movement, the first preset vector for head movement, and the first preset vector for expression change (corresponding to the corresponding first preset vector in the above step 3013).

[0090] Through the embodiment of the present application, the action information in the first video is feature extracted, and the various action information of the digital human in the video is accurately captured, providing a basis for subsequent discretization coding. A vector quantization variational autoencoder of a specified dimension is used to discretize the extracted feature vector, map it to a predefined first vector codebook, and obtain a first preset vector corresponding to the action information. The discretization process realizes data compression and representation learning. The reversible layer can perform a first dimension transformation on the fifth feature vector of the extracted action information to obtain a sixth feature vector consistent with the input dimension of the language model. The dimensional transformation process ensures that information is not lost during the dimensional conversion process, so that the sixth feature vector can be used as input for subsequent language model processing and generation tasks, and the first feature vector as the action information that matches the sixth feature vector is determined according to the first vector codebook. The first feature vector can more accurately reflect the characteristics of the action information.

[0091] In some embodiments, the second feature vector in step 301 can be implemented by the following method: extracting the seventh feature vector in the first speech; performing a second dimensional transformation on the seventh feature vector to obtain an eighth feature vector; and determining a second preset vector matching the eighth feature vector from a plurality of second preset vectors included in the second vector codebook of the speech information as the second feature vector.

[0092] As an example, the dimension of the second eigenvector is the same as the input dimension of the language model, and the language model is used to be called to generate a third eigenvector corresponding to the second digital person as the listener. For the first speech of the first digital person as the speaker, the first speech is subjected to feature extraction processing through a pre-trained emotional speech embedding model to obtain a seventh eigenvector. The pre-trained emotional speech embedding model can be an emotion2vec model. The emotion2vec model extracts low-level features of the speech, including Mel-frequency cepstral coefficients (MFCC), pitch, speech rate, and energy. These low-level features of the speech can capture the basic acoustic characteristics of the speech. The low-level features are then nonlinearly transformed through the hidden layer of the emotion2vec model to extract the emotional features of the speech. Finally, a vector of fixed dimension is generated through the output layer of the emotion2vec model as the seventh eigenvector. A speech reversible layer is added after the speech VQ-VAE model. The speech reversible layer is a reversible neural network. The seventh eigenvector is transformed into the second dimension through the reversible neural network to ensure that information is not lost during the dimensional conversion process. After processing by the speech reversible layer, the seventh eigenvector is converted into the eighth eigenvector. The dimension of the eighth eigenvector is consistent with the input dimension of the language model.

[0093] The second vector codebook is the codebook corresponding to the speech feature, which can effectively capture the main changes in the eighth feature vector. The second vector codebook is determined by the pre-trained speech VQ-VAE model. A Train the speech VQ-VAE model, speech set D A The speech VQ-VAE model consists of the first speech from all samples. During speech VQ-VAE model training, the encoder maps the input eighth feature vector to the latent space, obtaining a continuous latent vector. The vector quantization layer quantizes these continuous latent vectors to the nearest code vector for discretization. Each vector in the resulting second vector codebook represents a typical speech feature pattern. The parameters of the speech VQ-VAE model are optimized by minimizing the reconstruction loss and the vector quantization loss to obtain the pre-trained speech VQ-VAE model and the final second vector codebook. The speech VQ-VAE model training process is the same as that of the eye VQ-VAE model described above.

[0094] The eighth eigenvector is mapped based on the second vector codebook of speech information, and the continuous eighth eigenvector is mapped to a discrete vector representation through vector quantization. The eighth eigenvector is mapped using a vector quantization variational autoencoder (speech VQ-VAE model) of a specified dimension, and the continuous speech feature sequence is mapped to multiple second preset vectors included in the second vector codebook through vector quantization, and a second preset vector matching the eighth eigenvector is determined as the second eigenvector.

[0095] See also Figure 4B In the encoding module 4021, the first speech 4311 is feature extracted by the speech coding model 4321 to obtain a speech feature sequence 433 (corresponding to the seventh feature vector mentioned above), and then the seventh feature vector is transformed into the second dimension by the speech reversible layer 435 to obtain the eighth feature vector. The speech coding model 4321 is a pre-trained emotional speech embedding model. The speech feature sequence 433 is mapped to the corresponding second vector codebook by the speech discrete encoder 434 to obtain the second feature vector that matches the second preset vector after the dimension conversion of the corresponding speech feature sequence 433. The speech discrete encoder 434 is a pre-trained speech VQ-VAE model.

[0096] Through the embodiments of the present application, the first speech is efficiently converted into a second feature vector compatible with the language model, significantly improving the processing efficiency and quality of speech features, enhancing the naturalness and vividness of multimodal interaction, and providing a basis for digital human interaction in complex dialogue scenarios.

[0097] Continue to see Figure 3AIn step 302, a third feature vector of the second digital person is generated based on the first feature vector, the second feature vector and the first prompt word.

[0098] In some embodiments, see Figure 3C , Figure 3C This is a third flow chart of the data processing method provided in the embodiment of the present application. Figure 3A In step 302, the Figure 3C Steps 3021 to 3025 in are implemented as described below.

[0099] In step 3021, text encoding is performed on the first prompt word to obtain a ninth feature vector.

[0100] As an example, the first prompt word is a text description of the current dialogue scene, which provides contextual information for the language model. The first prompt word is text-encoded by the word segmenter, and the first prompt word is decomposed into a series of text tags, and the text tags are converted into numerical representations that the language model can understand. The word segmenter maps each text tag to a unique numerical identifier (numerical ID) based on the vocabulary of the pre-trained model. These numerical IDs are further converted into vectors of fixed dimensions through the embedding matrix to form the ninth eigenvector. The ninth eigenvector is a numerical representation of the first prompt word, which can be directly processed by the language model to provide contextual support for subsequent multimodal generation tasks.

[0101] join Figure 4B The word segmenter 412 performs text encoding processing on the first prompt word 411, decomposes the first prompt word 411 into a series of text tokens, maps each text token into a unique numerical identifier through the text model projection layer 413, and further converts it into a vector of fixed dimension through the embedding matrix in the text model projection layer 413 to form a text latent space vector 414.

[0102] Continue to see Figure 3C In step 3022, the first feature vectors corresponding to the multiple action information are spliced ​​to obtain a first splicing result.

[0103] As an example, action information includes eye movement, head movement, and facial expression changes. The first feature vectors corresponding to the various action information include eye movement feature vectors, head movement feature vectors, and facial expression change feature vectors. The first feature vectors of each action information are combined according to the Cartesian product and spliced ​​along a specific dimension to form a unified feature representation to form a first splicing result. For example, there are two sets A and B, A includes {a, b, c}, and B includes {1, 2}. The Cartesian product of A and B is to pair each element in A with each element in B to form a tuple. The Cartesian product of A and B is {(a, 1), (a, 2), (b, 1), (b, 2), (c, 1), (c, 2)}, and all possible result combinations are obtained. The first splicing result is a feature vector that integrates multiple action information and can comprehensively characterize the behavioral characteristics of the first digital human in the video.

[0104] For example, the first eigenvectors corresponding to the three types of action information, namely eye movement, head movement and expression change, are combined according to the Cartesian product and then concatenated. The eye movement eigenvector C eye There are N eigenvectors in total, and the i-th eigenvector is recorded as Head motion feature vector C head There are M eigenvectors in total, and the jth eigenvector is recorded as Expression change feature vector C exp There are K eigenvectors in total, and the kth eigenvector is recorded as By splicing along the dimension, N*M*K feature vectors can be obtained to form the first splicing result.

[0105] In step 3023, the first concatenation result is projected and aligned to the feature space of the language model to obtain a first latent space feature vector.

[0106] As an example, since the dimension of the first splicing result may be inconsistent with the dimension of the feature space of the language model, the first splicing result is projected and aligned to the feature space of the language model through the video adaptation layer for dimensionality adjustment. The video adaptation layer is a module for projecting the video feature vector to the feature space of the language model. By performing dimensionality adjustment, it is ensured that the feature vector from the video processing module can be consistent with the input feature vector dimension of the language model. The video adaptation layer implements linear transformation through the fully connected layer, mapping the feature vector of the first splicing result to the feature space of the language model, ensuring that the first splicing result is compatible with the input feature vector of the language model to form the first latent space feature vector. The weight matrix and bias vector in the fully connected layer of the video adaptation layer are trainable parameters of the language model and can be optimized through the training process. The first latent space feature vector is a feature representation after dimensionality adjustment, which can be directly processed by the language model to provide input for subsequent prediction tasks. See Figure 4B , the first feature vectors corresponding to the multiple action information are spliced ​​to obtain a first splicing result. The alignment splicing is performed based on the Cartesian product combination. The first splicing result is projected and aligned through the video adaptation layer 425 to obtain a video latent space vector 426, which corresponds to the first latent space feature vector in step 3023.

[0107] Continue to see Figure 3C In step 3024, the second feature vector is projected and aligned to the feature space of the language model to obtain a second latent space feature vector.

[0108] As an example, the second feature vector is projected and aligned, and the second feature vector is mapped to the feature space of the language model through the linear transformation of the speech adaptation layer to ensure that the second feature vector is compatible with the input feature vector of the language model, thereby forming the second latent space feature vector. Figure 4B The second feature vector is projected and aligned through the speech adaptation layer 436 to obtain a speech latent space vector 437, which corresponds to the second latent space feature vector in step 3024.

[0109] In step 3025, the ninth eigenvector, the first latent space eigenvector, and the second latent space eigenvector are predicted using a language model to obtain a third eigenvector of the second digital human.

[0110] As an example, the language model is derived from an extension of a pre-trained large text model and can simultaneously process data from multiple modalities, including text, video, and speech. The language model uses the fifth eigenvector (representing the contextual information of the first prompt word), the first latent space eigenvector (representing the various actions of the first digital person), and the second latent space eigenvector (representing the speech characteristics of the first digital person) as input for prediction. The fifth eigenvector, the first latent space eigenvector, and the second latent space eigenvector are concatenated along the time dimension or the feature dimension to form a comprehensive feature vector. This comprehensive feature vector is then fused by the language model encoder to generate a contextual representation. The language model encoder encodes the comprehensive feature vector using a multi-head self-attention mechanism and a feedforward neural network. The language model decoder uses this mechanism to gradually generate an output sequence, generating an output vector at each step, representing the prediction result at the current moment. The decoder uses an attention mechanism during the generation process to dynamically focus on key information in the contextual representation. The attention mechanism calculates the similarity between the contextual representation and the current decoding state to generate a weighted context vector, which is used to assist in generating the output at the current moment. During the generation process, the generated content is dynamically adjusted based on the input feature vector and contextual information. The output vector generated by the decoder is mapped to the target feature space through a linear transformation layer. The linearly transformed vector is then nonlinearly transformed through an activation function to generate the final output vector. The output vectors generated at each step are concatenated in chronological order to obtain the third eigenvector of the second digital person. The third eigenvector represents the behavioral characteristics of the second digital person in the listening state, including eye movements, head movements, and changes in facial expressions. Figure 4C The language model 4022 performs prediction processing on the ninth eigenvector, the first latent space eigenvector, and the second latent space eigenvector to obtain the third eigenvector of the second digital person. The third eigenvector is the listener encoding sequence 441.

[0111] Through the embodiments of the present application, the first feature vector, the second feature vector and the first prompt word corresponding to the various action information are predicted and processed. The language model can comprehensively consider information from multiple modalities such as text, video and voice, and generate a second digital human behavior feature vector that is consistent with the input information, thereby realizing the fusion of multimodal information, providing a basis for subsequent video rendering, ensuring the naturalness and coherence of the generated results, and enhancing the vividness and realism of subsequent digital human interactions.

[0112] Continue to see Figure 3A In step 303, for each type of action information, the third eigenvector is first decoded to obtain a fourth eigenvector corresponding to the action information.

[0113] In some embodiments, step 303 can be implemented by the following method: performing mapping processing on the third feature vector corresponding to the video generation task to obtain a feature mapping result, and decomposing the feature mapping result to obtain a feature mapping result for each action information; for each action information, performing an inverse dimensional transformation processing on the feature mapping result corresponding to the action information corresponding to the first dimensional transformation processing to obtain a first dimensional transformation result of the action information; for each action information, decoding the first dimensional transformation result of the action information to obtain a fourth feature vector of the action information.

[0114] As an example, while the language model structure remains unchanged, a video head model is added. The video head model is used to map the third eigenvector to the corresponding video generation task. The video head model is used to handle tasks related to video generation. Through mapping and inverse transformation, feature mapping results related to the listener's video are generated. These feature mapping results are used to generate the video. The video head model maps the third eigenvector to the video feature space, obtaining mapping results for eye movement, head movement, and facial expression changes. The feature mapping results are then decomposed into points for each type of action information to obtain feature mapping results for each action information.

[0115] Then, for each action, the feature map corresponding to the action is transformed inversely to the first dimension through a reversible layer added after the video head model. This transforms the feature map from the latent space of the language model back to the original dimension associated with the action. This inverse transformation is accomplished through a reversible layer (implemented by a reversible neural network). This transform converts the mapping results for eye movement, head movement, and facial expression back to their respective original dimensions, resulting in the first dimension transformation for each action. For each action, the first dimension transformation is decoded, and the low-resolution feature map is gradually expanded to a high-resolution feature map through deconvolution operations. Ultimately, the high-resolution feature map is restored to the same spatial dimensions as the original input data, resulting in a high-resolution feature map containing more detailed action information. After each deconvolution layer, an inverse residual connection is applied, adding lower-level feature information to the higher-level feature map via skip connections. This feature map, processed through the inverse residual connection, contains richer details. The normalized feature vectors are restored to their original scale through an inverse normalization layer, ensuring that the generated feature map matches the scale of the original data. This results in a fourth feature vector corresponding to each action, including eye movement feature vectors, head movement feature vectors, and expression change feature vectors. This fourth feature vector is used as the feature vector for subsequent video rendering, and the decoder gradually restores the feature vector to a specific action.

[0116] See also Figure 4CThe video head model 4401 maps the listener code sequence 441 (corresponding to the third eigenvector) to obtain a mapping result corresponding to the video generation task. The mapping result is then subjected to a bitwise decomposition process to obtain mapping results for eye movement, head movement, and expression change features. These are then input into the eye reversible layer 4223, the head reversible layer 4233, and the expression reversible layer 4243 for inverse dimensional transformation of the corresponding first dimensional transformation, thereby obtaining a first dimensional transformation result corresponding to each type of action information. The first dimensional transformation result corresponding to each type of action information is then input into the eye discrete decoder 4224, the head discrete decoder 4234, and the expression discrete decoder 4244 for decoding, and restored to the original scale to ensure that the generated feature map is consistent with the scale of the original data. This results in a fourth eigenvector for the action information, including an eye movement feature sequence 427, a head movement feature sequence 428, and an expression change feature sequence 429.

[0117] Through the embodiment of the present application, on the basis of the original model structure, a new video head model is added to process the corresponding video generation task to achieve lightweight model structure. Through inverse dimensionality transformation processing, the feature vector in the third feature vector is accurately converted from the latent space dimension of the language model back to the original dimension related to the action information, ensuring the integrity and accuracy of the information during the dimensionality conversion process. Through the decoding process, the spatial dimension is restored to the same as the original input data, and the detailed information of the feature sequence is enhanced to ensure the coordination and consistency between different action information. The efficient decomposition and restoration of the third feature vector improves the extraction accuracy and generation quality of the action information.

[0118] In step 304 , a second video of the second digital human is generated based on the fourth feature vector.

[0119] As an example, the fourth eigenvectors corresponding to various motion information are fused, and the fused fourth eigenvector is decoded through the PD-FGC algorithm decoder. The features at different levels are fused to restore the complete video feature vector, including information such as eye movement feature vectors, head movement feature vectors, and expression change feature vectors, to obtain the video feature vector of the second digital person as the listener. The identity encoder in the PD-FGC algorithm extracts the identity features of the second digital person and obtains the identity features of the corresponding facial image. The identity features are fused with the video feature vector of the second digital person as the listener to ensure that the generated video has the correct identity information. The PD-FGC algorithm decoder renders the video feature vector of the second digital person as the listener into specific video frames, using a deep learning network as a generation module to generate the image content of each frame based on the input feature representation. The generation network maps the input feature map to specific pixel values ​​through learned weights and biases to generate high-quality video frames. The generated video frames are sorted according to chronological order. The sorted video frames are usually encoded into a video format through a video encoding tool (such as FFmpeg). The video content is generated frame by frame, ensuring that the features of each frame such as head posture, eye movement and expression changes are consistent with the input video feature vector. The generated video frames are synthesized into a complete video as the second video of the second digital human.

[0120] See also Figure 4C The head video model 4401 performs mapping processing on the listener coding sequence 441 (the above-mentioned third feature vector) to generate the corresponding video to obtain a mapping result, and performs bit decomposition and inverse processing on the mapping result and the corresponding motion information to obtain the eye movement feature sequence 427, head movement feature sequence 428 and expression change feature sequence 429 of the second digital person as a listener. The eye movement feature sequence 427, head movement feature sequence 428, expression change feature sequence 429 of the second digital person as a listener are decoded by the PD-FGC algorithm decoder 460, and the eye movement feature sequence 427, head movement feature sequence 428, expression change feature sequence 429 and mouth feature sequence 456 are rendered into specific video frames, and the generated video frames are synthesized into a complete video as the listener video 403 of the second digital person (the above-mentioned second video).

[0121] In some embodiments, audio and video of the second digital person as the speaker are also generated by the following method: when generating the third feature vector corresponding to the second digital person, generating the speech content corresponding to the second digital person through the language model; and generating audio and video of the second digital person based on the speech content.

[0122] As an example, the third feature vector of the second digital person is generated by the language model based on the first feature vector, the second feature vector and the first prompt word corresponding to multiple action information. While generating the third feature vector corresponding to the second digital person as a listener, the language model also generates the speech content corresponding to the second digital person as a speaker, and generates audio and video of the second digital person as a speaker based on the speech content.

[0123] When the second digital person is speaking, the text latent space vector generated by the language model is used to map the speech content to the corresponding speech generation task using the text header model. The mapping result corresponding to the speech information is input into the inverse word segmenter for reverse conversion of the text tokens. The inverse word segmenter converts the mapping result back into readable text by searching the language model's vocabulary. The inverse word segmenter is the inverse operation of the word segmenter, mapping the numerical ID sequence output by the text header model back to the original text tokens, and ultimately combining them into the complete speech content of the second digital person. When the second digital person is speaking, the speech synthesis module performs text-to-speech processing based on the speech content to generate the speech of the second digital person as the speaker. The speaker's speech is then feature extracted to generate a mouth feature vector. The mouth feature vector and the speech content are then processed through the PD-FGC algorithm decoder for video rendering to generate a video of the second digital person as the speaker. The video and speech are combined to obtain the audio and video of the second digital person as the speaker.

[0124] The speech synthesis module performs speech synthesis processing on the speech of the second digital person, removing extra spaces and special characters from the speech. Numbers, dates, abbreviations, and other information in the speech are converted into a readable format. The speech of the second digital person is then segmented into words and sentences, providing structured input for subsequent speech synthesis. The speech synthesis module converts each word in the speech of the second digital person into a corresponding phoneme sequence. Based on language rules, hyphens, accents, and other factors are processed. A pre-trained acoustic model is used to convert the phoneme sequence into speech features. A vocoder is used to convert the speech feature sequence into an actual speech signal, i.e., the speech of the second digital person. A speech-driven model is used to extract features from the speech of the second digital person. A pre-trained deep learning model is used to map the speech feature sequence to mouth movements, generating a mouth feature sequence synchronized with the speech signal to drive the mouth movements of the digital person. The fourth feature sequence and the mouth feature sequence are decoded through the PD-FGC algorithm decoder, and the video feature sequence and the mouth feature sequence are rendered into specific video frames. The video content is generated frame by frame to ensure that the head posture, eye movement, expression change and mouth change features of each frame are consistent with the input video feature sequence. The generated video frames are synthesized into a complete video, and the complete video and the speaker's voice are combined to serve as the audio and video of the second digital human as the speaker.

[0125] See also Figure 4CThe head video model 4401 maps the speaker's encoded sequence 442 to the corresponding video generated sequence, obtaining a mapping result. The mapping result is then subjected to bitwise decomposition and inverse processing of the corresponding motion information to obtain an eye movement feature sequence 427, a head movement feature sequence 428, and an expression change feature sequence 429 of the second digital person acting as the speaker. The speaker and listener use the same video head model for processing. The text head model 4501 maps the text latent space vector 414. The mapping result is input into the inverse tokenizer 4511 and mapped back to text tokens, obtaining the spoken content text 4521 (the aforementioned spoken content). The speech synthesis module 4531 performs text-to-speech synthesis on the spoken content text 4521 to obtain the speaker's speech 4541. The speech drive model 4550 extracts features from the second digital person's speech 4541. A pre-trained deep learning model is used to map the speech feature sequence to mouth movements, generating a mouth feature sequence 456 synchronized with the speech signal to drive the digital person's mouth movements. The eye movement feature sequence 427, head movement feature sequence 428, expression change feature sequence 429 and mouth feature sequence 456 of the second digital person as the speaker are decoded through the PD-FGC algorithm decoder 460, and the eye movement feature sequence 427, head movement feature sequence 428, expression change feature sequence 429 and mouth feature sequence 456 are rendered into specific video frames. The generated video frames are synthesized into a complete video, and the complete video and the speaker's voice are combined as the speaker audio and video 404 of the second digital person.

[0126] Through the embodiments of the present application, audio and video content with a second digital human as the listener and speaker is generated simultaneously, realizing the fusion and generation of multimodal information. Through speech synthesis and mouth movement generation technology, accurate synchronization of mouth movements and speech content is achieved. The generated audio and video content is temporally coherent, and the movements match the speech content, enhancing the user experience. Through identity feature fusion, it is ensured that the generated video has correct identity information, thereby improving the realism of the generated video.

[0127] In some embodiments, a language model is obtained by concatenating first vector codebooks corresponding to a plurality of action information to obtain a video codebook; and expanding an initial language model based on the video codebook and the second vector codebook to obtain a language model.

[0128] As an example, the third feature vector corresponding to the second digital person as a listener is generated through a language model, and the first encoding is performed based on the first vector codebook corresponding to each action information, and the second encoding is performed based on the second vector codebook corresponding to the voice information. The first vector codebooks corresponding to a variety of action information are spliced, and the first vector codebooks for each action information are combined according to the Cartesian product and spliced ​​along a specific dimension to form a comprehensive video codebook. The video codebook is a set containing discrete representations of all action information, which can comprehensively characterize the behavioral characteristics of the digital person. Based on the video codebook and the discrete vectors in the second vector codebook as a new vocabulary, the vocabulary of the initial language model is expanded to enhance the language model's ability to process video features. Through training data, the parameters of the language model are adjusted to obtain an optimized language model.

[0129] As an example, since the text model (language model) cannot understand and generate video and audio, it is necessary to expand the vocabulary to enable the text model to have multimodal understanding and generation capabilities. The video codebook obtained after splicing is expanded to the vocabulary (token vocabulary) of the text model, and the eye codebook is A total of N vectors, head codebook A total of M vectors, expression codebook A total of K vectors, along the dimension splicing can get N*M*K vectors, a total of NMK vectors, are added as tokens<video_1> ~<video_NMK> , use the video codebook to expand the input encoding layer of the text model (language model) so that token<video_1> Can be mapped to the corresponding vector in the video codebook; speech codebook A total of L vectors are added as tokens<speech_1> ~<speech_L> , use the second vector codebook to expand the input encoding layer of the text model (language model) so that token<speech_1> The input coding layer can be mapped to the corresponding vector in the speech codebook. The original input coding layer includes multiple text embedding vectors. After expansion, the input coding layer also includes the second preset vector in the second vector codebook and the first preset vector in the video codebook.

[0130] Through the embodiments of this application, the fusion and generation of multimodal information is achieved. Through vocabulary expansion and model development, the language model's processing capabilities for video and voice information are enhanced. The language model can simultaneously process data in multiple modalities, such as text, video, and voice, and generate feature sequences consistent with the input information. Through vocabulary expansion and model development, feature sequences related to the behavioral characteristics of the second digital human can be more accurately generated, improving the quality of generated content. The generated feature sequences can represent the natural behavioral characteristics of the second digital human in both listening and speaking states, enhancing the realism and vividness of digital human interactions.

[0131] The model training method provided by the embodiments of the present application has the following beneficial effects:

[0132] Feature extraction is performed on the motion information in the first video, accurately capturing the various movements of the digital human in the video and providing a foundation for subsequent discretization encoding. A vector quantization variational autoencoder with a specified dimension is used to discretize the extracted feature vectors, achieving data compression and representation learning. A reversible layer resizes the vector dimension to obtain a feature vector consistent with the language model input dimension, ensuring that information is not lost during the dimensionality conversion process. When generating the third feature vector corresponding to the second digital human as the listener, the language model comprehensively considers information from multiple modalities, including text, video, and speech, to generate a feature vector consistent with the input information. Through feature fusion and context construction, it captures long-range dependencies and feature information in the feature vector. The decoder performs prediction generation based on the constructed contextual representation, ensuring the coherence and consistency of the generated content with the input information. When generating audio and video of the second digital human as the speaker, the language model generates the spoken content and generates mouth feature vectors to accurately synchronize mouth movements with the speech content, enhancing the naturalness and vividness of the digital human's interaction. By concatenating the first vector codebooks corresponding to various action information to generate a video codebook, the initial language model is expanded based on the video codebook and the second vector codebook, enhancing the language model's processing capabilities for video and voice information. The vocabulary expansion and model expansion enable the language model to more accurately generate feature vectors related to the behavioral characteristics of the second digital human, further improving the quality of generated content. Through multimodal feature fusion and decoding, the video and voice of the current listener and the next speaker in the digital human dialogue interaction scenario are efficiently processed and generated, improving the naturalness, coherence, and vividness of the digital human interaction.

[0133] The following describes an exemplary application of the embodiments of the present application in a practical application scenario.

[0134] Conversational communication, in which speakers and listeners alternate roles, is the foundation of effective information exchange in social interactions. Speakers convey information through verbal and non-verbal means such as voice, facial expressions, and head gestures, while listeners provide real-time feedback through non-verbal behaviors such as nodding, smiling, and shaking their heads. With the increasing demand for human-computer interaction scenarios, both speaker and listener head generation technologies have made significant progress in this field. Human-computer interaction based on digital humans has been widely applied in multiple practical fields.

[0135] Related technologies combine the understanding and generation capabilities of large models to segment interactive videos, extract interaction information from the images, and transform the extracted interaction information into a holistic discrete behavior sequence. Model-based analysis and prediction of the transformed discrete behavior sequence determine the characteristic sequence of the digital human video, enabling the generation of digital human videos in conversational scenarios. Separately generating digital human videos for the speaker and listener fails to consider the alternating nature of the two roles in conversational scenarios. The speaker in one round may become the listener in the next round. Related digital human generation relies on mapping the correlation between the speaker's behavior sequence and the listener's behavior sequence, achieving diverse results through a certain degree of random sampling. However, the holistic transformation of digital human images or videos contains a large amount of redundant information. The extracted discrete behavior sequence contains redundant information that has little impact on the interaction. Important information is coupled with less important information, hindering the use of models and algorithms. This fails to prioritize the more important information inherent in the interaction scenario, resulting in poor performance and interactive capabilities in synthesized digital human videos. Changes in input or context can easily lead to unnatural or irrational movements in the generated digital human.

[0136] The embodiment of the present application generates videos of speakers and listeners in a unified manner, so that the same language model is modeled to have the ability to generate both audio and video of the speaker and video of the listener. By obtaining the conversation history data and the video and voice of the speaker in the current round, the speaker's video is discretized, the behavior sequence of the eyes, head and expression is determined and spliced ​​into a video behavior sequence, and the speaker's voice is discretized to obtain a voice behavior sequence. The video behavior sequence and the voice behavior sequence are predicted based on the multimodal language model to obtain the video behavior sequence when the digital human is the listener in the current round and the video behavior sequence and text behavior sequence when the digital human is the speaker in the next round. The video behavior sequence output by the model is rendered to obtain the video of the digital human, and the text behavior sequence is converted into voice to obtain the speaker's voice.

[0137] In some embodiments, see Figure 4A , Figure 4AThis is a schematic diagram of the digital human video generation process provided by an embodiment of the present application; the digital human generation model 401 includes an encoding module 4021, a language model 4022, and a decoding module 4023. The digital human generation model 402 processes the input first video 4011 of the speaker in the tth round, and the characteristic behavior sequence obtained by the encoding module 4021 is fused and decomposed by the decoding module 4023, and the listener video 403 of the tth round and the audio and video 404 of the speaker in the t+1th round are simultaneously output. Through the same language model 4022, the video tasks of the listener in the current round are mapped to the video tasks of the speaker in the next round, achieving a more natural and vivid multi-round human-computer interaction experience.

[0138] The following is a description with reference to the accompanying drawings. Figure 5 , Figure 5 This is a fourth flow chart of the data processing method provided in the embodiment of the present application, which will be combined with Figure 5 The steps shown are explained in detail.

[0139] In step 501, a first video feature sequence corresponding to a first video with a first digital person as the speaker is obtained.

[0140] In some embodiments, the first video contains motion information and background information about the first digital person acting as a speaker. In the first video, the motion information corresponding to the first digital person's eye movements, facial expressions, and head movements has a greater impact on interactive communication, while background information such as the person's clothing and scene background has a smaller impact on interactive communication. The motion information can also include hand and leg movements of the first digital person. A first encoding process is performed on the motion information in the first video that has a greater impact on interactive communication to obtain a first video feature sequence corresponding to the first video (corresponding to the first feature vector in step 301 above). The first encoding process is implemented using the PD-FGC algorithm and a vector quantized variational autoencoder of a specified dimension.

[0141] The action information in the first video is feature extracted using the PD-FGC algorithm to obtain a feature sequence corresponding to each action information, including: an eye movement feature sequence, a head movement feature sequence, and an expression movement feature sequence (corresponding to the fifth feature vector in step 3011 above). The feature sequence corresponding to each action is respectively input into a vector quantized variational autoencoder of a corresponding specified dimension. Based on the first vector codebook, the input feature sequence is mapped to a latent space by the vector quantized variational autoencoder of the specified dimension to obtain a discrete latent vector representation, i.e., a first vector sequence corresponding to the feature sequence in the first vector codebook, including a first vector sequence corresponding to eye movement, a first vector sequence corresponding to head movement, and a first vector sequence corresponding to expression change (corresponding to the first preset vector in step 3013 above).

[0142] In some embodiments, see Figure 6 , Figure 6 This is a schematic diagram of the PD-FGC algorithm provided in an embodiment of the present application. The PD-FGC algorithm is a single-image digital human algorithm that controls the digital human's mouth movements, eye movements, head movements, and facial expressions by decoupling latent representations. Based on a character image 601 to be generated, an identity encoder 603 obtains the task identity information in the character image 601. A speech encoder 604 encodes the digital human's voice information to generate the digital human's mouth movements based on the voice features, synchronizing the mouth movements with the voice information. The PD-FGC algorithm includes a head movement encoder 6051, an facial movement encoder 6061, and an eye movement encoder 6071. These three encoders learn feature representations of the head, facial expression, and eye features in a sample video 602. For example, head movement follows the leftmost video in the sample video 602, facial expression changes follow the video in the middle column, and eye movement follows the rightmost video. Head movement enables the digital human to perform various head movements, such as bowing, turning, nodding, and shaking its head, creating dynamic head gestures and enhancing interactivity. During the facial generation process, facial expression movement controls facial expression changes to express different emotional states, enabling the display of rich emotions and making the digital human more expressive and approachable. Eye movement controls the direction of eye gaze and blinking frequency, creating a natural gaze and blinking experience, making the digital human appear more vivid and realistic. The learned feature representations are used to obtain a sequence of changing features through the head feature model 6052, the expression feature model 6062, and the eye feature model 6072. The sequence of head, expression, and eye changing features obtained through the head feature model 6052, the expression feature model 6062, and the eye feature model 6072 is then rendered into a video using the generation model 608, generating a generated character video 609 that shares the same features learned in the sample video 602.

[0143] After separating the action information from the first video, the feature vectors carrying each type of information individually have low dimensionality, lower than the dimensionality of the feature vectors in the latent space of the language model. Directly upscaling the feature vectors of each action information from the low dimensionality to the latent space of the language model would result in a large dimensional expansion, which could lead to low information density and prevent the language model from accurately capturing the action information. By decoupling each type of action information, different feature sequences are obtained, which then need to be expanded and arranged. Longer feature sequences increase the computational complexity. Compared to predicting the entire first video, predicting the eyes, head, and expression separately can lead to erroneous results. For example, if the entire video has three possible outcomes: {white square, red triangle, white triangle}, then if the color and shape attributes are separated, the corresponding outcomes are {white, red} and {square, triangle}, respectively. If the color and shape are predicted separately, due to prediction errors, the color may be red and the shape may be square. However, the red square corresponding to this predicted combination does not actually exist. Even if the overall prediction still has errors, this is less likely to result in a non-existent red square. Therefore, after predicting the feature sequence corresponding to each action information, the reversible layer is used to ensure the integrity during the dimensional conversion process, avoid doubling the length of the feature sequence, and avoid the dimension of each attribute from being too high, thereby forming the first video feature sequence corresponding to the first video with the first digital person as the speaker (corresponding to the first feature vector in the above step 301), without having to predict each attribute separately, to avoid unreasonable combinations.

[0144] In some embodiments, see Figure 4B , Figure 4B4 is a schematic diagram of the encoding side structure of the digital human generation model provided by an embodiment of the present application; in the encoding module 4021, first, the action information in the first video 401 is feature extracted by the PD-FGC algorithm encoder 421 to obtain an eye movement feature sequence 422, a head movement feature sequence 423, and an expression change feature sequence 424 (corresponding to the fifth feature vector in the above step 3011). The eye movement feature sequence 422 is input into the eye discrete encoder 4221 for mapping processing, the head movement feature sequence 423 is input into the head discrete encoder 4231 for mapping processing, and the expression change feature sequence 424 is input into the expression discrete encoder 4241 for mapping processing. The eye discrete encoder 4221, the head discrete encoder 4231, and the expression discrete encoder 4241 are vector quantized variational autoencoders of specified dimensions. The eye discrete encoder 4221, the head discrete encoder 4231 and the expression discrete encoder 4241 respectively perform mapping processing based on the first vector codebook on the input eye movement feature sequence 422, the head movement feature sequence 423 and the expression change feature sequence 424 to obtain a discrete latent vector representation. The first vector sequence corresponding to the feature sequence in the first vector codebook includes: a first vector sequence for eye movement, a first vector sequence for head movement and a first vector sequence for expression change (corresponding to the first preset vector in the above step 3013).

[0145] As an example, a newly added reversible layer is used to ensure integrity during the dimensional conversion process, and the first vector sequence of eye movement, the first vector sequence of head movement, and the first vector sequence of expression change are aligned and spliced. The above-mentioned reversible layer can be implemented by a reversible neural network. In the latent space, the reversible neural network adjusts the dimension of the latent representation through a series of reversible transformations (corresponding to the above-mentioned first dimensional change processing in step 3012). These reversible transformations include reversible linear transformations and reversible nonlinear transformations. The latent representation after the reversible layer transformation is input into the vector quantization layer and mapped to the vector in the predefined discrete codebook to obtain the first feature sequence corresponding to eye movement, the first feature sequence corresponding to head movement, and the first feature sequence corresponding to expression change (corresponding to the above-mentioned sixth feature vector). The eye reversible layer 4222 performs a first dimensional transformation on the first vector sequence of eye movement, the head reversible layer 4232 performs a first dimensional transformation on the first vector sequence of head movement, and the expression reversible layer 4242 performs a first dimensional transformation on the first vector sequence of expression changes to obtain a first feature sequence corresponding to eye movement, a first feature sequence corresponding to head movement, and a first feature sequence corresponding to expression changes, i.e., the first video feature sequence (the above-mentioned first feature vector matching the sixth feature vector) ensures that the encoded information is completely decoded while not causing information to be lost when the vector dimension changes.

[0146] Continue to see Figure 5 In step 502, a first speech feature sequence corresponding to a first speech of a first digital person as a speaker is obtained.

[0147] In some embodiments, a pre-trained emotional speech embedding model is used to extract features of the first speech to obtain a first speech feature sequence (corresponding to the second feature vector in step 301 above). The pre-trained emotional speech embedding model is an emotion2vec model. After receiving the first speech, the emotion2vec model first extracts low-level features of the speech (corresponding to the seventh feature vector above), including Mel-frequency cepstral coefficients, which reflect the spectral characteristics of the speech, intonation, speaking rate and energy. These low-level features can capture the basic acoustic characteristics of the speech. The input layer of the pre-trained emotional speech embedding model receives the pre-processed speech signal, performs a nonlinear transformation on the speech signal through the hidden layer, extracts the emotional features of the speech, and the output layer generates a vector of fixed dimension as the first speech feature sequence of the first speech, which represents the emotional state and semantic information of the speech. The speech feature sequence is then input into the encoder of the speech discrete coding model for mapping to the corresponding second vector codebook. Each vector in the second vector codebook represents a typical speech feature pattern. The encoder of the speech discrete coding model will map the speech feature sequence to the latent space to obtain a continuous latent representation (corresponding to the above-mentioned second preset vector). The dimension of the latent representation is changed through the reversible layer. The vector quantization layer maps the continuous latent representation to the vector in the predefined discrete codebook to obtain a discrete latent vector sequence, namely the first speech feature sequence (corresponding to the above-mentioned second eigenvector matching the eighth eigenvector).

[0148] In some embodiments, see Figure 4B In the encoding module 4021, the first speech 4311 is subjected to feature extraction processing by the speech coding model 4321 to obtain a speech feature sequence 433. The speech coding model 4321 is a pre-trained emotional speech embedding model. Then, the speech feature sequence 433 is mapped to the second vector codebook corresponding to the speech information by the speech discrete encoder 434. The second vector codebook is a codebook corresponding to the speech information, and a speech vector sequence corresponding to the speech feature sequence 433 (corresponding to the above-mentioned second preset vector) is obtained. The speech discrete encoder 434 is a pre-trained speech VQ-VAE model. The speech vector sequence is subjected to dimensionality transformation processing (corresponding to the above-mentioned second dimensionality transformation processing) by the speech reversible layer 435 to obtain a first speech feature sequence (corresponding to the eighth feature vector).

[0149] In some embodiments, see Figure 7 , Figure 7: This is a schematic diagram of the discrete encoder structure of a specified dimension provided by an embodiment of the present application; the structures of the above-mentioned eye, head, expression and voice discrete encoders are the same, all including a discrete encoder 701, a reversible layer 702, a reversible layer 703 and a discrete decoder 704. The discrete encoder 701 maps the input feature sequence to the latent space to obtain continuous potential representations, which are high-dimensional continuous vectors. The reversible layer 702 and the reversible layer 703 are corresponding inverse operations to each other. The continuous potential representation is converted in dimension by a reversible neural network. The reversible neural network adjusts the dimension of the potential representation through a series of reversible transformations. These reversible transformations include reversible linear transformations and reversible nonlinear transformations. The potential representation after the reversible layer transformation is input to the vector quantization layer. The vector quantization layer maps each continuous potential vector to the closest discrete vector in the codebook through nearest neighbor search. The discrete decoder 704 restores the discrete potential vector sequence back to the original data space, generates an output feature sequence similar to the input feature sequence (corresponding to the second feature vector matching the eighth feature vector above), and obtains discrete eye, head, expression and voice feature sequences.

[0150] Continue to see Figure 5 In step 503, based on the first video feature sequence, the first voice feature sequence and the first prompt word, a second video feature sequence of the second digital person as a listener and a third video feature sequence and a second voice feature sequence of the second digital person as a speaker are generated.

[0151] In some embodiments, the first prompt word is a text description of the current dialogue scene. The first prompt word is segmented to obtain text features (corresponding to the ninth feature vector mentioned above), and the text features are projected into the latent vector space of the language model. The first video feature sequence is projected into the latent vector space of the language model through the video adaptation layer, and the first voice feature sequence is projected into the latent vector space of the language model through the video adaptation layer, together with the first prompt word as context information. The language model is obtained by expanding the pre-trained text model.

[0152] Through fine-tuning of instructions, the ability to understand and generate video and speech is expanded. In the process of training the language model, the spliced ​​video feature sequence is expanded into the vocabulary of the text model, and the encoding module of the input text model is expanded with the feature sequence so that the text model can be mapped to the corresponding vector in the video feature sequence. The original capabilities of the text model and the parameters of the text output head model are kept unchanged, and a video output head model is added. The video output head model and the text output head model are independent of each other, and each predicts a feature sequence based on the latent space vector. The video output head model predicts the video feature sequence corresponding to the video generation task, and the text head model outputs the text feature sequence predicted by the speech generation task. The encoding layer of the language model is expanded with video and speech feature sequences so that the video feature sequence can be mapped to the corresponding vector in the video codebook, and the speech codebook can be mapped to the corresponding vector in the speech codebook. The training set for training the language model collects a large number of multi-round two-person dialogue samples, among which (A s ,V s ,V l ) t is the triplet of the tth round, which is the speaker voice, speaker video, and listener video. s refers to the speaker and l refers to the listener. The i-th multi-round two-person dialogue sample can be expressed as D i ={(A s ,V s ,V l )1,(A s ,V s ,V l )2,…,(A s ,V s ,V 1 ) t ,…}, based on these samples, three databases are constructed for training language models: paired sample set D P 、Video Collection D V and speech set D A .

[0153] Paired sample set D P Each multi-round two-person conversation sample is split according to the rounds, and a pair sample set is constructed and recorded as Among them, H i,<t is the text record of the conversation of the i-th multi-round two-person conversation sample up to the t-1th round, and are the speech signal and face video of the speaker in the tth round of the i-th multi-round two-person conversation sample, It is to listen to the video of people’s faces; and They are the text and face video of the speaker in the t+1th round of the i-th multi-round two-person conversation sample. and It is the same digital person, the listener in round t will be the speaker in round t+1, and the videos in the sample are uniformly processed into 25 FPs.

[0154] Video Collection D V The set consists of all videos from all samples, regardless of whether they are speaking or listening. The encoder, using the PD-FGC algorithm, performs feature extraction, extracting action information such as eye movement, head movement, and facial expression changes from the set of videos from all samples to form a feature sequence. The eye movement, head movement, and facial expression change feature sequences are used as input to the eye, head, and facial expression VQ-VAE models, respectively, to learn the representation of action features. During the training process of the eye, head, and facial expression VQ-VAE models, the encoders of the eye, head, and facial expression VQ-VAEs map the input eye, head, and facial expression feature sequences into a latent space, generating continuous latent vectors. The vector quantization layer quantizes these continuous latent vectors to the nearest codebook vector for discretization. Each vector in the codebook represents a typical action feature pattern. The codebook vectors are continuously optimized and updated during training. Model parameters are optimized by minimizing the reconstruction loss and vector quantization loss, resulting in the optimized eye, head, and facial expression VQ-VAE models and the final first-vector codebook. The first-vector codebook effectively captures the main changes in the video feature sequence.

[0155] Voice Set D A A set of speech consisting of all samples, through the pre-trained emotional speech embedding model, extracts low-level features (such as Mel-frequency cepstral coefficients, intonation, speaking rate, energy, etc.) from the speech to form a speech feature sequence, which is used as the input of the speech VQ-VAE model to learn the representation of speech features. During the training process of the speech VQ-VAE model, the speech VQ-VAE model encoder maps the input speech feature sequence to the latent space to obtain a continuous latent vector. The vector quantization layer quantizes these continuous latent vectors to the nearest codebook vector to achieve discretization representation. Each vector in the codebook represents a typical speech feature pattern. By minimizing the reconstruction loss and vector quantization loss to optimize the parameters of the speech VQ-VAE model, the final speech codebook (corresponding to the second preset vector mentioned above) can effectively capture the main changes in the speech feature sequence.

[0156] For speech set D A Each speech segment A j , using the pre-trained emotion2vec model to obtain the speech feature sequence Get Collection Use collection Train the speech discretization coding model. i,<t , The corresponding video and speech feature sequences are the input part, The corresponding video and speech feature sequences are used as the expected output part to train the language model. The trained language model can perform prediction processing on the video and speech features, so that the output prediction layer of the expanded language model can predict the number of types of feature sequences obtained, and has multimodal understanding and generation capabilities.

[0157] See also Figure 4B , the word segmenter 412 performs text encoding processing on the first prompt word 411, decomposes the first prompt word 411 into a series of text tokens, maps each text token into a unique numerical identifier through the text model projection layer 413, and further converts it into a vector of fixed dimension through the embedding matrix in the text model projection layer 413 to form a unified text latent space vector 414 (i.e. <text>< / text> ).

[0158] The eye movement feature sequence, the head movement feature sequence, and the expression change feature sequence are spliced ​​together to obtain a first splicing result. The splicing is performed based on the Cartesian product combination. The first splicing result is projected and aligned by the video adaptation layer 425 to obtain a video latent space vector 426 (corresponding to the first latent space feature vector in step 3023). The video latent space vector 426 may be inconsistent with the latent vector space dimension of the language model 4022. The video adaptation layer 425 adjusts the video latent space vector 426 to a dimension compatible with the language model 4022 through a learnable linear transformation or other adaptation mechanism to obtain a video latent space vector 426, that is, <video>< / video> The speech adaptation layer 436 projects the first speech coding feature (corresponding to the second feature vector) into the latent vector space of the language model 4022 to obtain the speech latent vector space vector 437, that is, <audio>< / audio> (corresponding to the second latent space eigenvector mentioned above).

[0159] The language model generates a second video feature sequence (corresponding to the fourth feature vector mentioned above) of the second digital person as a listener, a third video feature sequence, and a second speech feature sequence of the second digital person as a speaker based on the text, video, and speech latent vector space vectors. The language model identifies the identifiers of the listener and speaker in the input text, video, and speech latent vector space vectors, switches the generation type of the head model according to the identified sequence representation, makes predictions based on the video head model in the language model, and simultaneously outputs the second video feature sequence of the second digital person as a listener and the third video feature sequence of the second digital person as a speaker. The second speech feature sequence of the second digital person as a speaker is output based on the prediction of the text head model in the language model.

[0160] In some embodiments, see Figure 4C , Figure 4C This is a schematic diagram of the decoding side structure of the digital human generation model provided in the embodiment of the present application; based on the listening person coding sequence 441 (ie <listener>< / listener> , corresponding to the third feature vector) and the speaker code sequence 442 (i.e. <speaker>< / speaker> The video generation task is mapped and decomposed using a video head model 4401. The listener's encoded sequence 441 is decomposed into eye movement feature mapping results, head movement feature mapping results, and expression change feature mapping results through alignment. These are then input into the eye reversible layer 4223, the head reversible layer 4233, and the expression reversible layer 4243 for inverse transformation, restoring these sequence encodings to continuous latent vector sequences of the original dimensions. These sequences are then input into the eye discrete decoder 4224, the head discrete decoder 4234, and the expression discrete decoder 4244 for decoding. The restored continuous latent vector sequences are gradually reconstructed into the features of each action corresponding to the original first video. The decoder for each action gradually increases the spatial dimension of the features through inverse convolution, recovers detailed information through inverse residual connections, and finally undergoes normalization. The vectors in the latent space are gradually restored to the feature sequences of each action corresponding to the original first video, including: eye movement feature sequence 427, head movement feature sequence 428, and expression change feature sequence 429.

[0161] Continue to see Figure 5 In step 504, when the second digital human acts as a listener, a video of the second digital human acting as a listener is generated based on the second video feature sequence.

[0162] In some embodiments, when the second digital human acts as a listener, the decoder of the PD-FGC algorithm decodes the second video feature sequence (corresponding to the fourth feature vector mentioned above), fuses features at different levels, and restores a complete video feature sequence, including information such as eye movement, head movement, and expression changes, to obtain a video feature sequence of the second digital human as a listener. The decoder of the PD-FGC algorithm renders the video feature sequence into specific video frames, and fuses identity features (such as facial images) with the video feature sequence to ensure that the generated video has correct identity information. The decoder of the PD-FGC algorithm generates video content frame by frame, ensuring that the features such as head posture, eye movement, and expression changes of each frame are consistent with the input video feature sequence, and synthesizes the generated video frames into a complete video as the video of the second digital human as a listener. See Figure 4C The PD-FGC algorithm decoder 460 decodes and renders the second video feature sequence (corresponding to the fourth feature vector mentioned above) and outputs the listener video 403.

[0163] Continue to see Figure 5 In step 505, when the second digital person acts as a speaker, the speech content of the second digital person as a speaker is generated based on the second voice feature sequence.

[0164] In some embodiments, when the second digital person acts as a speaker, the text latent space vector is extracted from the second speech feature sequence, and the text latent vector space is subjected to an inverse tokenizer through a text head model to perform an inverse conversion of the text tokens to obtain the speech content of the second digital person as a speaker. Figure 4C The inverse tokenizer 4511 in the text header model 4501 converts the digitized ID sequence of the text latent vector space vector 414 back into readable text form by searching the vocabulary of the language model 4022. The inverse tokenizer 4511 is the inverse operation of the tokenizer, mapping the digitized ID sequence output by the model back to the original text tokens, and ultimately combining them into the complete speaker's utterance 4521.

[0165] In step 506, based on the third video feature sequence and the speech content, an audio and video of the second digital human as the speaker is generated.

[0166] In some embodiments, when the second digital person acts as a speaker, text-to-speech processing is performed through a speech synthesis module based on the speech content to obtain the speech of the second digital person as the speaker, and features of the speaker's speech are extracted to generate a mouth feature sequence. The mouth feature sequence and the third video feature sequence are subjected to video rendering processing through a PD-FGC algorithm decoder to generate a video of the second digital person as the speaker. The video and voice are combined to obtain audio and video of the second digital person as the speaker.

[0167] See also Figure 4CSpeech synthesis module 4531 performs speech synthesis processing on spoken text 4521, removing extra spaces and special characters from spoken text 4521 to ensure that spoken text 4521 is properly formatted. Numbers, dates, abbreviations, and the like in spoken text 4521 are converted into a readable format, and spoken text 4521 is segmented into words and sentences to provide structured input for subsequent speech synthesis. Speech synthesis module 4531 converts each word in spoken text 4521 into a corresponding phoneme sequence. Phonemes are the basic units of speech. Based on language rules, hyphens, accents, and other conditions are processed. A pre-trained acoustic model (e.g., a deep learning-based acoustic model) is used to convert the phoneme sequence into speech features. A vocoder then converts the speech feature sequence into an actual speech signal, namely, speaker speech 4541. The speaker's speech 4541 is feature extracted through the speech-driven model 4550, and the speech feature sequence is mapped to mouth movements using a pre-trained deep learning model to generate a mouth feature sequence 456 synchronized with the speech signal to drive the digital human's mouth movements.

[0168] The PD-FGC algorithm decoder decodes the third video feature sequence and the mouth feature sequence, fusing features at different levels to restore the complete video feature sequence and mouth feature sequence. The decoder renders the video feature sequence and mouth feature sequence into specific video frames, generating video content frame by frame, ensuring that each frame's head posture, eye movement, expression changes, and mouth changes are consistent with the input video feature sequence. The generated video frames are synthesized into a complete video, which is then combined with the speaker's voice to produce a video of the second digital human speaking.

[0169] The data processing method provided in the embodiment of the present application has the following beneficial effects:

[0170] The PD-FGC algorithm extracts and decouples motion information from the digital human's first video, such as eye movements, head movements, and facial expressions. These feature sequences are discretized and encoded using a vector quantized variational autoencoder (VQ-VAE) with a specified dimension, effectively reducing information redundancy. A reversible layer ensures information integrity during dimensionality conversion. During the decoding phase, video features are gradually restored through inverse convolutions and inverse residual connections, ensuring that the generated video matches the input features in terms of detail and temporal coherence. A speech-driven model maps speech signals to mouth movement feature sequences, achieving precise synchronization between mouth movements and speech content. By constructing a large database of multi-turn two-person conversation samples, the language model learns behavioral patterns in different conversational scenarios, generating more natural and coherent video and speech content across multiple rounds of interaction. Through command fine-tuning and vocabulary expansion, the language model can simultaneously process data from multiple modalities, including text, video, and speech. This preserves the original capabilities of the text model during training and avoids performance degradation. Through multimodal fusion, fine-grained behavior control, voice-driven mouth movement synchronization, and efficient multi-round interaction capabilities, the performance of digital humans in complex dialogue scenarios is improved. At the same time, the audio and video of the speaker in the next round and the video of the listener in the current round are generated, which enhances the naturalness and vividness of digital human interactions.

[0171] The following continues to describe the exemplary structure of the data processing device 455 provided in the embodiment of the present application implemented as a software module. In some embodiments, such as Figure 2 As shown, the software modules stored in the data processing device 455 of the memory 450 may include: an encoding module 4551, which is used to perform a first encoding on each action information for the first video of the first digital human to obtain a first feature vector for each action information, and to perform a second encoding on the first voice of the first digital human to obtain a second feature vector; based on the first feature vector, the second feature vector and the first prompt word, a third feature vector of the second digital human is generated; a video generation module 4552, which is used to perform a first decoding on the corresponding action information of the third feature vector for each action information to obtain a fourth feature vector corresponding to the action information; and based on the fourth feature vector, a second video of the second digital human is generated.

[0172] In some embodiments, the encoding module 4551 is further used to extract the fifth eigenvector of each type of action information in the first video; for each type of action information, perform a first dimension transformation on the fifth eigenvector of the action information to obtain a sixth eigenvector of the action information; for each type of action information, determine, from a plurality of first preset vectors included in the first vector codebook of the action information, a first preset vector matching the sixth eigenvector as the first eigenvector of the action information.

[0173] In some embodiments, the encoding module 4551 is further used to extract the seventh eigenvector from the first speech; perform a second dimensional transformation on the seventh eigenvector to obtain an eighth eigenvector; and determine, from a plurality of second preset vectors included in the second vector codebook of the speech information, a second preset vector matching the eighth eigenvector as the second eigenvector.

[0174] In some embodiments, the encoding module 4551 is also used to perform text encoding on the first prompt word to obtain a ninth feature vector; perform splicing processing on the first feature vectors of multiple action information to obtain a first splicing result; project the first splicing result onto the feature space of the language model to obtain a first latent space feature vector; project the second feature vector onto the feature space of the language model to obtain a second latent space feature vector; and perform prediction processing on the ninth feature vector, the first latent space feature vector, and the second latent space feature vector through the language model to obtain a third feature vector of the second digital person.

[0175] In some embodiments, the video generation module 4552 is also used to perform mapping processing on the third feature vector corresponding to the video generation task to obtain a feature mapping result, and decompose the feature mapping result to obtain a feature mapping result for each action information; for each action information, the feature mapping result of the action information is subjected to an inverse dimensional transformation processing corresponding to the first dimensional transformation processing to obtain a first dimensional transformation result of the action information; for each action information, the first dimensional transformation result of the action information is decoded to obtain a fourth feature vector of the action information.

[0176] In some embodiments, the third feature vector of the second digital person is generated by the language model based on the first feature vector, the second feature vector and the first prompt word corresponding to multiple action information. The video generation module 4552 is also used to generate the speech content corresponding to the second digital person through the language model when generating the third feature vector corresponding to the second digital person; and generate audio and video of the second digital person based on the speech content.

[0177] In some embodiments, the video training module 4553 is further used to perform splicing processing on the first vector codebook corresponding to the action information to obtain a video codebook; and to perform expansion processing on the initial language model based on the video codebook and the second vector codebook to obtain a language model.

[0178] An embodiment of the present application provides a computer program product, which includes a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the data processing method described in the embodiment of the present application.

[0179] The embodiment of the present application provides a computer-readable storage medium in which computer-executable instructions or computer programs are stored. When the computer-executable instructions or computer programs are executed by a processor, the processor will execute the data processing method provided in the embodiment of the present application, for example, Figure 3A The data processing method is shown.

[0180] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or may be various devices including one or any combination of the above memories.

[0181] In some embodiments, computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0182] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files storing one or more modules, subroutines, or code portions).

[0183] By way of example, computer-executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed across multiple sites and interconnected by a communication network.

[0184] In summary, through the embodiment of the present application, the motion information and voice information of the first digital human are discretized and encoded separately to avoid interference between different modal information, and to achieve the separation processing of the multimodal features of the first digital human. Based on the first feature vector, the second feature vector and the first prompt word corresponding to the multiple motion information, the multimodal information such as motion, voice and prompt word are jointly processed to generate the third feature vector corresponding to the second digital human as a listener and the speech content of the second digital human as a speaker, and to integrate the multiple features of motion, voice and prompt words to ensure the logical association between the reaction action of the second digital human as a listener and the scene. The generated feature vectors are decomposed and fused and decoded corresponding to the multiple motion information, and the scattered motion codes are recombined into a complete video output, maintaining the details of multiple motions, and generating a more natural and coherent video and voice of the second digital human.

[0185] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.

Claims

1. A data processing method, characterized in that: The method comprises: For the first video of the first digital human, perform a first encoding on each type of action information to obtain a first feature vector for each type of action information, and perform a second encoding on the first voice of the first digital human to obtain a second feature vector; generating a third feature vector of the second digital person based on the first feature vector, the second feature vector, and the first prompt word; For each type of action information, performing a first decoding corresponding to the action information on the third eigenvector to obtain a fourth eigenvector corresponding to the action information; A second video of the second digital human is generated based on the fourth feature vector.

2. The method according to claim 1, characterized in that The first video of the first digital human performs a first encoding on each type of action information to obtain a first feature vector for each type of action information, including: extracting a fifth eigenvector of each type of action information in the first video; For each type of action information, performing a first dimension transformation process on the fifth eigenvector of the action information to obtain a sixth eigenvector of the action information; For each type of action information, a first preset vector matching the sixth eigenvector is determined from a plurality of first preset vectors included in the first vector codebook of the action information as the first eigenvector of the action information.

3. The method according to claim 2, characterized in that The performing a first decoding on the third eigenvector corresponding to the action information to obtain a fourth eigenvector corresponding to the action information includes: Performing mapping processing corresponding to the video generation task on the third feature vector to obtain a feature mapping result, and performing decomposition processing on the feature mapping result to obtain a feature mapping result for each type of action information; For each type of action information, performing an inverse dimensional transformation process corresponding to the first dimensional transformation process on a feature mapping result of the action information to obtain a first dimensional transformation result of the action information; For each type of action information, the first dimensional transformation result of the action information is decoded to obtain a fourth eigenvector of the action information.

4. The method according to claim 1, wherein The performing a second encoding on the first voice of the first digital person to obtain a second feature vector includes: extracting a seventh feature vector from the first speech; Performing a second dimension transformation on the seventh eigenvector to obtain an eighth eigenvector; A second preset vector matching the eighth eigenvector is determined from a plurality of second preset vectors included in the second vector codebook of the speech information as the second eigenvector.

5. The method according to claim 1, wherein Generating a third feature vector of the second digital person based on the first feature vector, the second feature vector, and the first prompt word includes: Perform text encoding on the first prompt word to obtain the ninth feature vector; performing splicing processing on the first feature vectors of the plurality of action information to obtain a first splicing result; Projecting the first concatenation result onto the feature space of the language model to obtain a first latent space feature vector; Projecting and aligning the second feature vector to the feature space of the language model to obtain a second latent space feature vector; The ninth eigenvector, the first latent space eigenvector, and the second latent space eigenvector are predicted and processed by the language model to obtain a third eigenvector of the second digital person.

6. The method according to claim 1, wherein The third feature vector of the second digital human is generated by a language model based on the first feature vectors, the second feature vectors, and the first prompt word corresponding to the plurality of action information; the method further includes: When generating the third feature vector corresponding to the second digital person, generating the speech content corresponding to the second digital person through the language model; The audio and video of the second digital human are generated based on the speech content.

7. The method according to claim 1, characterized in that The third feature vector is generated by a language model, the sum of the dimensions of the first feature vectors of the plurality of action information is the same as the dimension requirement of the input data of the language model, and the dimension of the second feature vector is the same as the dimension requirement of the input data of the language model.

8. A data processing device, characterized in that: The device comprises: an encoding module configured to perform a first encoding on each type of action information in a first video of a first digital human to obtain a first feature vector for each type of action information, perform a second encoding on the first speech of the first digital human to obtain a second feature vector, and generate a third feature vector for the second digital human based on the first feature vector, the second feature vector, and a first prompt word; The video generation module is used to perform a first decoding of the third eigenvector corresponding to each type of action information to obtain a fourth eigenvector corresponding to the action information; and generate a second video of the second digital human based on the fourth eigenvector.

9. An electronic device, characterized in that: The electronic device comprises: a memory for storing computer-executable instructions or computer programs; The processor is configured to implement the data processing method according to any one of claims 1 to 7 when executing the computer-executable instructions or computer programs stored in the memory.

10. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that: When the computer executable instructions or computer program are executed by a processor, the data processing method according to any one of claims 1 to 7 is implemented.

11. A computer program product comprising computer executable instructions or a computer program, characterized in that When the computer executable instructions or computer program are executed by a processor, the data processing method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Portrait dialogue video generation method, multi-person dialogue video generation method, product, equipment and storage medium

    CN121000955A