3D virtual human real-time interaction system and implementation method

CN118552672BActive Publication Date: 2026-09-29ZHEJIANG UNIV +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410505796.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-25
Publication Date
2026-09-29
Estimated Expiration
2044-04-25

AI Technical Summary

Technical Problem

[0005]本发明的目的在于提供一种3D虚拟人实时交互系统,解决现有技术中虚拟人交互困难繁琐、虚拟人口型合成不准确延迟高的问题

Benefits of technology

[0044]1.提出单张图片生成基于MetaHuman的3D虚拟人模型的方法,减少虚拟人建模时间,可以在较短时间内生成具有高保真形象的虚拟人。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118552672B_ABST
    Figure CN118552672B_ABST
Patent Text Reader

Abstract

The application discloses a 3D virtual human real-time interaction system and an implementation method, comprising: a UE5 front-end interaction module for 3D virtual human modeling, animation rendering and scene construction, obtaining multi-modal input, and performing real-time two-way data communication with a back-end server; an LLM server module for deploying an offline large language model, fine-tuning the large model according to an application scenario, enabling the virtual human to intelligently answer user questions, and communicating and interacting with the user; a back-end server module based on FastAPI, for deploying a voice field model and a virtual human Audio2Face model on a local server, and performing data communication with the UE5 front-end interaction module and the LLM server module to process related data. The application provides a one-stop virtual human deployment application solution from modeling to interaction with low cost, low delay and high fluency based on a MetaHuman model generation method, an interaction delay optimization method and an audio-driven mouth shape algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence and virtual human technology, and in particular to a 3D virtual human real-time interaction system and its implementation method. Background Technology

[0002] With the growing popularity of the metaverse concept, more and more virtual humans are coming into the public eye, such as the VOCALOID virtual singer Luo Tianyi and Siren, a high-fidelity virtual human co-developed by Tencent and Epic Games. The virtual human economy is booming, and the overall market size of digital virtual humans continues to expand.

[0003] Currently, the application of virtual human technology is mainly divided into service-oriented virtual humans and identity-oriented virtual humans. Service-oriented virtual humans include virtual anchors, while identity-oriented virtual humans include applications such as virtual IPs. However, regardless of the type of virtual human, the interaction method is often based on produced videos or animations. If real-time interaction is desired, motion capture or facial capture equipment is required, which is costly, and human control makes the virtual human lack intelligence. How to interact with virtual humans is a key issue for the further application of virtual human technology.

[0004] In recent years, the rise of AIGC and large language models has made it possible to endow virtual humans with intelligence. With the support of AI computing power, large language models can serve as the "brain" of virtual humans. In addition, how to drive virtual humans is also an important issue. Traditional solutions, such as the lip-sync plugin provided by Oculus, synthesize virtual human lip-sync animations based on visual pixels generated from audio, which often has the problem of inaccurate lip-syncing. Although the MetaHuman SDK provided by Microsoft Azure has accurate lip-syncing, it often requests data from its cloud server, resulting in high latency. Summary of the Invention

[0005] The purpose of this invention is to provide a real-time interactive system for 3D virtual humans, which solves the problems of difficult and cumbersome virtual human interaction and inaccurate virtual human lip-syncing with high latency in the prior art.

[0006] The technical solution adopted in this invention is as follows: A 3D virtual human real-time interactive system, comprising:

[0007] The UE5 front-end interaction module is used for 3D virtual human modeling, animation rendering and scene construction, to obtain multimodal input, and to establish a WebSocket connection with the back-end server module based on FastAPI for real-time bidirectional data communication.

[0008] The LLM server module is used to deploy an offline large language model. The large language model is fine-tuned according to the application scenario, receives audio inference results sent by the FastAPI backend server module, and uses the large language model to generate response text.

[0009] The backend server module based on FastAPI is used to deploy the speech domain model and the virtual human Audio2Face model on the local server, and to communicate with the UE5 front-end interaction module and LLM server module. It generates audio inference results and animation frames from the received audio and text data through the speech domain model and the virtual human Audio2Face model.

[0010] Furthermore, the UE5 front-end interaction module specifically includes:

[0011] MetaHuman modeling unit is used for the generation and editing of virtual human images;

[0012] Scene building units are used to provide user interaction interfaces;

[0013] The Animation Blueprint unit is used to establish a WebSocket connection with the backend server for data communication, as well as to display text, play audio, and render virtual human animations.

[0014] Furthermore, the LLM server module specifically includes:

[0015] The WebDemo unit is used to test the question-answering capabilities of large language models on the web.

[0016] The API unit is used to provide a data interaction interface for backend server modules based on FastAPI.

[0017] Furthermore, the FastAPI-based backend server module specifically includes:

[0018] The ASR unit is used to convert user-input audio sent by the UE5 front-end interaction module into text.

[0019] The LLM communication unit is used to send text as a prompt to the LLM server module and receive the returned text data.

[0020] The TTS unit is used to synthesize the responses returned by the LLM server module into audio.

[0021] The Audio2Face unit is used to convert synthesized audio into lip-syncing data to drive a virtual human.

[0022] Furthermore, the animation blueprint unit specifically includes:

[0023] The AudioImporter component is used to capture user-input audio and resample it, converting the audio into binary data for data transmission.

[0024] WebSocket component: Used to establish a WebSocket connection with the backend server, define callback events for receiving data, and establish a caching mechanism to handle data returned from the backend.

[0025] Furthermore, the ASR unit specifically includes:

[0026] Speech recognition model component: Used to deploy offline speech recognition models to convert audio data sent by the UE5 front-end interaction module into readable text;

[0027] Text correction model component: used to deploy offline text correction models to perform semantic correction on text inferred by speech recognition models;

[0028] Symbol calibration model component: Used to deploy offline symbol calibration models to add punctuation marks to text inferred by text correction models.

[0029] Furthermore, the TTS unit specifically includes:

[0030] Speech synthesis model definition component: used to deploy offline speech synthesis models, converting text returned by the LLM server into audio data;

[0031] Speech Synthesis Model Optimization Component: This component encodes and compresses the audio data obtained from offline speech synthesis model inference, reducing data transmission volume and increasing data transmission rate.

[0032] Furthermore, the Audio2Face unit specifically includes:

[0033] Audio-driven lip-sync model definition component: used to deploy offline audio-driven lip-sync models, converting audio data synthesized by TTS units into lip-sync animation data corresponding to the virtual human;

[0034] Audio-driven lip-sync model optimization component: Used for post-processing of lip-sync animation data obtained from offline audio-driven lip-sync model inference.

[0035] On the other hand, this invention also provides a real-time interaction method for 3D virtual humans, including:

[0036] Step S1: Construct a virtual human and interactive scene based on MetaHuman, and add hybrid deformation weights to the animation table of the virtual human; capture interactive audio and resample it to obtain processable audio data, and establish a WebSocket connection with the backend server based on FastAPI to send data;

[0037] Step S2: Process the interactive audio data sent by the front end through the backend server based on FastAPI, perform speech recognition, text correction, and symbol calibration, and send the inference results to the LLM server;

[0038] Step S3: After receiving the data, the LLM server uses the fine-tuned large model for inference and returns the text results in a streaming manner to the backend server based on FastAPI.

[0039] Step S4: After receiving the text data based on FastAPI, the backend server performs speech synthesis and lip-reading inference by segmenting the sentences, and encapsulates the results into JSON data containing mixed transformation weights and returns it to the frontend.

[0040] Step S5: The front end receives and parses the JSON data, displays the text data on the screen using a virtual human model, renders the lip-sync data, and completes the interaction.

[0041] Furthermore, the construction of the virtual human and interactive scene based on MetaHuman is specifically as follows: Select the face image to be modeled and import it into UE; create a new material and set the image's RGB to the material's self-illuminating color; create a cube and set the created material as a cube material; create a MetaHuman body, add the created cube as a mesh component, boost the frame rate, then track the active frames, perform MetaHuman body calculation, convert the mesh to MetaHuman, and finally enter MetaHumanCreator to sculpt the face, completing the virtual human modeling; use MetaShoot to build a photography studio as the virtual human placement scene, drag the MetaHuman into the scene, and set the lighting direction; add a movie camera component, set the Pawn to automatically control the player as player 0, set the focus method to tracking focus, and set the Actor to be tracked as MetaHuman, so that after entering the program, the user's view is directly facing the MetaHuman, and the MetaHuman display clarity can be automatically adjusted;

[0042] Modify the MetaHuman animation blueprint and disconnect the original Livelink real-time link pose from the animation graph. Add a modified curve node with cached pose as input and a mapping from the original MetaHuman face specification to the ARkit face specification as output. Add 52 hybrid deformation weights to the node and add 52 corresponding floating-point type variables to the animation graph.

[0043] The beneficial effects of this invention are as follows:

[0044] 1. A method for generating 3D virtual human models based on MetaHuman from a single image is proposed, which reduces the virtual human modeling time and can generate virtual humans with high fidelity in a shorter time.

[0045] 2. An optimization method to reduce the latency of virtual human interaction is proposed. By deploying the large language model, speech domain model and Audio2Face model offline on the local server, the model inference speed is improved and network latency is avoided. At the same time, when processing the data returned by the large model, streaming transmission and sentence segmentation are performed to avoid the latency caused by the TTS and Audio2Face models processing too long sentences.

[0046] 3. A method for audio-driven MetaHuman lip-syncing is proposed. By modifying the MetaHuman facial binding skeleton to conform to the widely applicable industry-standard binding specification of 52 Blendshape facial weights based on ARKit, the Audio2Face model can be used to drive virtual lip-syncing.

[0047] 4. Propose a one-stop solution for 3D virtual human from modeling to interaction, which can quickly build a low-cost, low-latency, and highly smooth 3D virtual human real-time interactive system. Attached Figure Description

[0048] Figure 1 This is a schematic diagram of a 3D virtual human real-time interactive system according to the present invention;

[0049] Figure 2 This is a flowchart of a real-time interactive method for 3D virtual humans according to the present invention;

[0050] Figure 3 This is a block diagram illustrating the UE5 front-end interaction implementation provided in an embodiment of the present invention.

[0051] Figure 4 This is a block diagram of an LLM server implementation provided in an embodiment of the present invention;

[0052] Figure 5 This is a block diagram of a backend server implementation based on FastAPI provided in an embodiment of the present invention. Detailed Implementation

[0053] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit the invention or its application or use. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without inventive effort are within the scope of protection of this invention.

[0054] See Figure 1 A 3D virtual human real-time interactive system, comprising:

[0055] The UE5 front-end interaction module is used for 3D virtual human modeling, animation rendering and scene construction, to obtain multimodal input such as voice, and to establish a WebSocket connection with the back-end server for real-time bidirectional data communication.

[0056] The UE5 front-end interaction module specifically includes:

[0057] MetaHuman modeling unit is used for the generation and editing of virtual human images;

[0058] Scene building units are used to provide a user-friendly interface.

[0059] The Animation Blueprint unit is used to establish a WebSocket connection with the backend server for data communication, as well as to display text, play audio, and render virtual human animations.

[0060] The scene construction unit includes:

[0061] Studio components: used to provide suitable display scenes and lighting conditions for virtual humans, making the virtual human image clear and lifelike;

[0062] Camera component: Used to automatically focus on the virtual human, providing users with a better visual experience.

[0063] The animation blueprint unit includes:

[0064] The AudioImporter component is used to capture user-input audio and resample it, converting the audio into binary data that can be used for data transmission.

[0065] WebSocket component: Used to establish a WebSocket connection with the backend server, define callback events for receiving data, and establish a caching mechanism to handle data returned from the backend.

[0066] The LLM server module is used to deploy offline large language models. It allows for fine-tuning of the large model according to application scenarios, enabling virtual humans to intelligently answer user questions and interact with users. The LLM server module includes:

[0067] The WebDemo unit is used to test the question-answering capabilities of large language models on the web.

[0068] The API unit is used to provide a data interaction interface for backend server modules based on FastAPI.

[0069] The backend server module based on FastAPI is used to deploy offline models in the speech domain and the virtual human Audio2Face model. After obtaining user input from the UE5 front-end interaction module, it performs speech recognition, text correction, and symbol calibration. The inference results are sent as a prompt to the LLM server module. After receiving the return results from the LLM server module, semantic analysis and speech synthesis are performed. Audio2Face is used to obtain virtual human lip-sync data based on the synthesized audio. The text, synthesized audio, and lip-sync animation data returned by the LLM server module are sent to the UE5 front-end interaction module for rendering.

[0070] The FastAPI-based backend server module includes:

[0071] The ASR unit is used to convert user-input audio sent by the UE5 front-end interaction module into text.

[0072] The LLM communication unit is used to send text as a prompt to the LLM server module to obtain a response;

[0073] The TTS unit is used to synthesize the responses returned by the LLM server module into audio.

[0074] The Audio2Face unit is used to convert synthesized audio into lip-syncing data to drive a virtual human.

[0075] The ASR unit includes:

[0076] Speech recognition model component: used to deploy offline speech recognition models, converting binary audio data sent by the UE5 front-end interaction module into readable text;

[0077] Text correction model component: used to deploy offline text correction models to perform semantic correction on text inferred by speech recognition models;

[0078] Symbol calibration model component: Used to deploy offline symbol calibration models to add punctuation marks to text inferred by text correction models.

[0079] The LLM communication unit includes:

[0080] Send Request Component: Used to send prompt request responses to the LLM server module;

[0081] Data receiving component: Used to receive text data returned by the LLM server module.

[0082] The TTS unit includes:

[0083] Speech synthesis model definition component: used to deploy offline speech synthesis models, converting text returned by the LLM server into audio data;

[0084] Speech Synthesis Model Optimization Component: This component encodes and compresses the audio data obtained from offline speech synthesis model inference, reducing data transmission volume and increasing data transmission rate.

[0085] The Audio2Face unit includes:

[0086] Audio-driven lip-sync model definition component: used to deploy offline audio-driven lip-sync models, converting audio data synthesized by TTS units into lip-sync animation data corresponding to the virtual human;

[0087] Audio-driven lip-sync model optimization component: Used for post-processing of lip-sync animation data obtained from offline audio-driven lip-sync model inference, such as smoothing.

[0088] See Figure 2 A method for real-time interaction with a 3D virtual human includes the following steps:

[0089] Step S1: Construct a virtual human and interactive scene based on MetaHuman, and add hybrid deformation weights to the animation table of the virtual human; use the terminal's built-in microphone and UE5 front-end interaction module to capture audio and resample it at 16kHz to obtain processable audio data, and establish a WebSocket connection with the back-end server based on FastAPI to send data.

[0090] Step S2: Process the audio data sent from the front end through the backend server module based on FastAPI, perform speech recognition, text correction, and symbol calibration, and send the inference results to the LLM server;

[0091] Step S3: After receiving the data, the LLM server module uses the fine-tuned large model for inference and returns the text results in a streaming manner to the backend server based on FastAPI.

[0092] Step S4: After receiving the text data, the backend server module based on FastAPI performs speech synthesis and lip-reading inference by segmenting the sentences, and encapsulates the results into JSON data containing mixed deformation weights and returns it to the UE5 frontend interaction module.

[0093] Step S5: The UE5 front-end interaction module receives JSON data, parses it, displays the text data on the screen, renders the lip-sync data, and completes the interaction.

[0094] Example:

[0095] See Figure 3This embodiment includes a UE5 front-end interaction module for 3D virtual human modeling, animation rendering, and scene construction. It acquires multimodal inputs such as voice and establishes a WebSocket connection with the back-end server for real-time bidirectional data communication. The specific implementation includes the following steps:

[0096] S1: Complete the virtual human modeling based on MetaHuman.

[0097] As an example, select the face image to be modeled and import it into UE. Create a new material and set the image's RGB to the material's self-illuminating color. Create a cube and set the created material as a cube material. Create a MetaHuman ontology, add the created cube as a mesh component, boost the frame, then track the active frames, perform MetaHuman ontology solving, convert the mesh to MetaHuman, and finally enter MetaHuman Creator to sculpt the face and complete the virtual human modeling.

[0098] S2: Complete the scene construction required for the front-end interactive interface.

[0099] As an example, MetaShoot is used to build a studio as a virtual human placement scene. MetaHuman is dragged into the scene and the lighting direction is set. A movie camera component is added, and the Pawn is set to automatically control the player as player 0. The focus method is set to tracking focus, and the Actor to be tracked is set to MetaHuman. This ensures that the user's view is facing MetaHuman after entering the program, and the display clarity of MetaHuman can be automatically adjusted.

[0100] S3: Complete the MetaHuman settings on the UE side and modify the face binding specifications.

[0101] As an example, the MetaHuman detail level is forced to LOD1 to match any resolution; the MetaHuman animation blueprint Face_AnimBP is modified, the original Livelink real-time link pose is disconnected in the AnimGraph, a modified curve node is added, the node input Source Pose is the cached pose BodyPose, and the output is the mapping mh_arkit_mapping_pose from the original MetaHuman face specification to the ARkit face specification. In addition, 52 blendshapeweights are added to the node, and 52 related floating-point type variables are added to the AnimGraph.

[0102] S4: Write the animation blueprint and complete the front-end interaction and WebSocket connection related logic functions.

[0103] As an example, this code writes a level blueprint, initializes the MetaHuman animation instance, obtains its animation blueprint Face_AnimBP, creates a Runtime Audio Importer component for audio capture, and creates a WebSocket component for establishing a connection with the server. It writes the audio capture logic, setting it to single-channel, 16kHz sampling rate, compressing the audio to save data transmission size, using WAV format with a quality of 100, and converting the captured data to a byte array. It then writes the WebSocket connection logic, using ws as the communication protocol to request a WebSocket connection from the backend server and implementing a connection success binding event. Finally, it writes the data reception callback event logic, parsing JSON data, establishing a cache queue, storing data in different queues based on different fields, creating an event Tick node, processing data frame by frame, retrieving a set of data from the queue each frame, displaying text, playing audio, and rendering animations. This involves two custom events: playing audio and rendering animation. The audio playback event converts the received Base64 data into WAV format and then plays it using UE's built-in "play audio" function. The animation rendering event sets an event timer every 0.03 seconds, and sets 52 blendshape weights variables for each event. All of the above processing uses a synchronous mechanism.

[0104] See Figure 4 The LLM server module is used to deploy offline large language models. It allows for fine-tuning of the large model based on application scenarios, enabling virtual humans to intelligently answer user questions and interact with users. The specific implementation includes the following steps:

[0105] S1: Fine-tune the large model according to the actual application field.

[0106] As an example, considering computing power and cost, the PEFT method is used to fine-tune the ChatGLM2-6B model. Taking the Su Shi large model as an example, ChatGPT can be used to generate a dataset for training, and fine-tuning can be performed based on the LoRA method.

[0107] S2: Write web-based scripts to test large models.

[0108] As an example, we will write a web-based demo based on Grado to test the model performance with different Maximum Length, TOP_P, and Temperature.

[0109] S3: Develop a Uvicorn-based server and encapsulate the API interface.

[0110] As an example, an LLM streaming output interface is encapsulated based on FastAPI, and the LLM is deployed on a local server using Uvicorn to keep it running in the background.

[0111] See Figure 5 The backend server module, based on FastAPI, is used to deploy the speech domain model and the virtual human Audio2Face model on the local server, and to communicate with the UE5 frontend interaction module and LLM server module to process relevant data. The specific implementation includes the following steps:

[0112] S1: Deploy the offline speech model and write the basic backend architecture code.

[0113] As an example, a speech recognition model is deployed based on PaddlePaddle, using the conformer_talcs English-Chinese mixed recognition model with an audio sampling rate of 16kHz. A speech synthesis model is deployed based on ModelScope, using a Su Shi timbre model trained on Sambert, along with audio compression code to reduce the size of the synthesized audio file. A text correction model is deployed based on Pycorrector, using the Kenlm model to correct phonetic, visual, and grammatical errors in Chinese text. A symbol calibration model is deployed based on PaddlePaddle, using the ernie_linear_p7_wudao model to calibrate Chinese symbols. The backend server is based on Uvicorn. Before service startup, the above models are initialized using a test voice sample for speech recognition, text correction, symbol calibration, and speech synthesis to initialize model parameters and avoid delays caused by repeated model initialization during subsequent model inference.

[0114] S2: Deploy the Audio2Face model.

[0115] As an example, an Audio2Face model is deployed based on PyTorch, using EmoTalk as the model. First, the names and order of the 52 blendshape weights are set to ensure consistency with the model output. Then, the animation frame format is written, encapsulating the model output data by frame number. One audio track corresponds to one JSON data set, and each JSON data set contains several frames, with each frame containing the 52 blendshape weight values. Finally, post-processing code is written for smoothing and other related tasks. Since the model architecture is a wav2vec model, the relevant model needs to be downloaded from HuggingFace. Therefore, initialization is performed before starting the background service, using the TTS-synthesized audio to test and complete the model parameter initialization.

[0116] S3: Write an LLM communication module to provide a front-end interaction interface.

[0117] As an example, an audio interface is written to handle interactive requests from the UE5 frontend. It continuously receives audio data sent by the client. When the audio is not empty, it performs speech recognition, text correction, and symbol calibration, sends the recognition results to the frontend interface for display, and sends the inference results to the LLM server. It receives the streamed text, segments it into sentences according to punctuation, performs speech synthesis sentence by sentence, performs Audio2Face on the synthesized audio, performs preliminary encapsulation of the model inference results to obtain animation frames, and finally encapsulates the segmented text, synthesized audio, and animation frames into JSON data and sends it to the UE5 frontend for interactive rendering. If the received audio is empty, it returns fixed text, audio, and animation data.

[0118] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A 3D virtual human real-time interactive system, characterized in that, include: The UE5 front-end interaction module is used for 3D virtual human modeling, animation rendering, and scene construction. It acquires multimodal input and establishes a WebSocket connection with the back-end server module based on FastAPI for real-time bidirectional data communication. After receiving JSON data returned by the back-end server module based on FastAPI, the UE5 front-end interaction module parses the JSON data into text data, audio data, and lip-sync animation data, and stores them into text cache queues, audio cache queues, and lip-sync animation cache queues, respectively. The corresponding data is retrieved from each cache queue frame by frame through the Tick event to synchronously display text, play audio, and render virtual human animation based on the lip-sync animation data. The LLM server module is used to deploy an offline large language model. The LLM server module includes a WebDemo unit and an API unit. The WebDemo unit is used to test the question-answering capabilities of the large language model on the web, and the API unit provides a data interaction interface for the FastAPI-based backend server module. The large language model is fine-tuned according to the application scenario. It receives text data sent by the FastAPI-based backend server module after speech recognition, text correction, and symbol calibration. Using this text data as a prompt input, the fine-tuned offline large language model generates response text and returns the response text in a streaming manner to the FastAPI-based backend server module through an LLM streaming output interface encapsulated in FastAPI. The FastAPI-based backend server module is used to deploy the speech domain model and the virtual human Audio2Face model on the local server, and to communicate with the UE5 frontend interaction module and the LLM server module. The FastAPI-based backend server module receives audio data sent by the UE5 frontend interaction module in a loop through the same audio interface. When the audio data is not empty, it performs speech recognition, text correction and symbol calibration in sequence, sends the obtained text data to the LLM server module, and receives the response text returned by the LLM server module in a streaming manner. The FastAPI-based backend server module segments the returned response text into sentences according to punctuation, performs speech synthesis and lip-syncing inference on each sentence, obtains synthesized audio and lip-syncing animation frames containing 52 mixed deformation weights, and encapsulates the sentence-segmented text, synthesized audio and lip-syncing animation frames into JSON data and returns it to the UE5 frontend interaction module. The UE5 front-end interaction module, LLM server module, and FastAPI-based back-end server module form a low-latency closed-loop processing link for real-time interaction of 3D virtual humans through synchronous rendering of the aforementioned cache queue, offline large language model streaming return, pipeline orchestration of the same audio interface, sentence-by-sentence speech synthesis and lip-reading inference, and JSON data return containing 52 hybrid deformation weights.

2. The 3D virtual human real-time interactive system as described in claim 1, characterized in that, The UE5 front-end interaction module specifically includes: MetaHuman modeling unit is used for the generation and editing of virtual human images; Scene building units are used to provide user interaction interfaces; The Animation Blueprint unit is used to establish a WebSocket connection with the FastAPI-based backend server module for data communication, as well as to display text, play audio, and render virtual human animations. The animation blueprint unit includes an AudioImporter component and a WebSocket component. The AudioImporter component is used to capture user-input audio and resample it at 16kHz, converting the audio into binary data for data transmission. The WebSocket component is used to establish a WebSocket connection with the FastAPI-based backend server module, define data reception callback events, parse the JSON data returned by the FastAPI-based backend server module, and establish the text cache queue, audio cache queue, and lip-sync animation cache queue. When rendering the virtual human animation, the animation blueprint unit sets an event timer every 0.03 seconds and sets 52 hybrid deformation weight variables once in each event.

3. The 3D virtual human real-time interactive system as described in claim 1, characterized in that, The ASR unit in the FastAPI-based backend server module specifically includes: The speech recognition model component is used to deploy an offline speech recognition model to convert audio data sent by the UE5 front-end interaction module into readable text. The text correction model component is used to deploy offline text correction models to perform semantic correction on the text inferred by the speech recognition model. The symbol calibration model component is used to deploy an offline symbol calibration model to add punctuation marks to the text inferred by the text correction model; Before starting the backend server module based on FastAPI, test voice is used to initialize the speech recognition model component, text correction model component, and symbol calibration model component to reduce the model initialization waiting time in subsequent real-time interaction.

4. A 3D virtual human real-time interactive system as described in claim 1, characterized in that, The TTS unit in the FastAPI-based backend server module specifically includes: The speech synthesis model definition component is used to deploy offline speech synthesis models, converting the sentence-by-sentence text returned by the LLM server module and segmented by punctuation into audio data. The speech synthesis model optimization component is used to encode and compress the audio data obtained from the offline speech synthesis model inference, thereby reducing the amount of data transmission and improving the data transmission rate. Specifically, before starting the backend server module based on FastAPI, the parameters of the offline speech synthesis model are initialized to reduce the model initialization waiting time in subsequent real-time interactions.

5. A 3D virtual human real-time interactive system as described in claim 1, characterized in that, The Audio2Face unit in the FastAPI-based backend server module specifically includes: The audio-driven lip-sync model definition component is used to deploy offline audio-driven lip-sync models and convert audio data synthesized by TTS units into lip-sync animation data corresponding to virtual humans. The audio-driven lip-sync model optimization component is used to post-process the lip-sync animation data obtained from offline audio-driven lip-sync model inference. The Audio2Face unit encapsulates the model output data according to the names and order of 52 hybrid deformation weights. Each sentence synthesized audio corresponds to a JSON data packet, each JSON data packet contains several frames, and each frame contains 52 hybrid deformation weight values.

6. A method for creating and interacting with a 3D virtual human in real time, characterized in that, include: Step S1: Create a 3D virtual human model using the MetaHuman method based on a single image. Select the face image to be modeled and import it into UE. Create a new material and set the image's RGB to the material's self-illuminating color. Create a cube and set the created material as a cube material. Create a MetaHuman body, add the created cube as a mesh component, boost the frame rate, then track the active frames, perform MetaHuman body calculation, convert the mesh to a MetaHuman, and finally enter MetaHuman Creator to sculpt the face, completing the virtual human modeling. Use MetaShoot to build a photography studio as the virtual human placement scene, drag the MetaHuman into the scene, and set the lighting direction. Add a movie camera component, set the Pawn to automatically control the player as player 0, set the focus method to tracking focus, and set the Actor to be tracked to MetaHuman, so that the user's view is directly facing the MetaHuman after entering the program, and the MetaHuman's display clarity can be automatically adjusted. Step S2: Construct a virtual human and interactive scene based on MetaHuman, modify the MetaHuman animation blueprint and disconnect the original Livelink real-time link pose in the animation graph, add modified curve nodes, with the node input being the cached pose and the output being the mapping from the original MetaHuman face specification to the ARKit face specification, add 52 hybrid deformation weights to the node, and add corresponding 52 floating-point type variables to the animation graph; capture interactive audio and resample it at 16kHz to obtain processable audio data, and establish a WebSocket connection with the backend server based on FastAPI to send data; Step S3: Receive interactive audio data sent by the front end through the same audio interface of the backend server based on FastAPI in a loop. When the interactive audio data is not empty, perform speech recognition, text correction and symbol calibration on the interactive audio data in sequence, and send the recognition results to the front end interface for display. At the same time, send the symbol-calibrated text data to the LLM server. Step S4: After receiving the text data through the API unit, the LLM server uses the text data as prompt input to perform inference on the fine-tuned offline large language model, and returns the response text stream to the FastAPI-based backend server through the LLM streaming output interface based on FastAPI. Step S5: After receiving the streaming response text based on FastAPI, the backend server segments the text into sentences according to punctuation, performs speech synthesis and lip-syncing inference sentence by sentence, obtains the segmented text, synthesized audio, and lip-syncing animation frames containing 52 mixed deformation weights, and encapsulates the segmented text, synthesized audio, and lip-syncing animation frames into JSON data and returns it to the frontend. Step S6: The front end receives and parses the JSON data, storing the text data, audio data, and lip-sync animation data into their respective cache queues. The front end retrieves the corresponding data from each cache queue frame by frame using the Tick event. The text data is displayed on the screen through the virtual human model, the audio data is transcoded into WAV format and played, and an event timer is set every 0.03 seconds to set 52 hybrid deformation weight variables to synchronously render the lip-sync animation data, completing the real-time interaction of the 3D virtual human.

Citation Information

Patent Citations

  • Realistic virtual human multi-modal interaction implementation method based on UE4

    CN111724457A

  • Historical ancient city digital protection and business travel application system based on artificial intelligence

    CN116894896A

  • Method and system for enhancing MetaHuman digital human dialogue effect by adopting ChatGPT and unreal engine

    CN117539425A