Virtual digital human interaction device, system and method thereof
Through the interaction method of virtual digital humans, multimodal sensors and large language models, the natural language communication and expression action output between virtual humans and users is realized, solving the problem of limited people and applications in the existing technology, and improving the fusion effect between virtual humans and real scenes.
Patent Information
- Application Number
- CN202410370711.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-29
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2044-03-29
AI Technical Summary
In the prior art, 3D virtual human IP relies on the personal capabilities of the person in the process, resulting in limited application in the on-site navigation environment. The large language model can only perform text and voice output, and cannot realize dialogue-based communication between people.
Through the interaction method of virtual digital humans, multi-modal sensors are used to obtain environmental information and user information, combined with the prompt word engineering of the large language model, the natural language communication between virtual digital humans and users is realized, and expressions and actions are output through virtual human animation control program.
Real-time natural language communication between virtual people and users in a specific environment is realized, the integration effect between virtual digital people and real scenes is improved, and the application restriction of virtual people in the on-site navigation environment is solved.
Smart Images

Figure CN118535005B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence applications, and particularly relates to an interaction method, device and storage medium for virtual digital humans. Background Art
[0002] The IP of 3D virtual humans with live motion capture has an increasingly good-looking image and more vivid performance. However, this 3D virtual human IP is based on the personal abilities of the "person behind the scenes". In the prior art, large language models can only input and output text and cannot conduct conversational communication between people, resulting in limited application of large language models in on-site navigation environments. Summary of the Invention
[0003] The present invention provides an interaction method, device and storage medium for virtual digital humans, aiming to solve at least one of the technical problems existing in the prior art.
[0004] The technical solution of the present invention relates to an interaction method, device and storage medium for virtual digital humans, characterized in that the method includes:
[0005] S100. Determine to receive a wake-up signal, cause the virtual digital human to enter the wake-up process, and switch from the standby state to the working state. The wake-up process includes a voice wake-up process and a visual wake-up process;
[0006] S200. Based on the multi-modal sensors of the virtual digital human, obtain environmental information and user information;
[0007] S300. Based on large language prompt engineering, the virtual digital human interacts with the user;
[0008] S400. Based on the prompt engineering of the large language model, adjust the character role setting of the virtual digital human.
[0009] Further, in the step S100, the voice wake-up process includes the following steps:
[0010] S110. In the standby state, collect environmental sounds and detect whether there is a human voice wake-up signal. The human voice wake-up signal includes one or a combination of human voice, wake-up keywords and voiceprint features;
[0011] S120. If the detected human voice wake-up signal reaches a preset threshold, wake up the virtual digital human and cause the virtual digital human to switch from the standby state to the working state.
[0012] Further, in the step S110, detecting whether there is a human voice wake-up signal includes the following steps:
[0013] S111. Perform noise reduction preprocessing on the received environmental sound. The noise reduction preprocessing includes frame splitting, windowing, pre-emphasis, and adaptive voice activity detection (VAD) on the sound signal;
[0014] S112. Extract speech features based on MDCC and input the speech features into the voiceprint recognition model;
[0015] S113. Detect whether there is a human voice wake-up signal and perform voiceprint matching between the speech features and the existing voiceprints in the voiceprint matching library;
[0016] S114. If the voiceprint matching is successful, retrieve the user information matching the voiceprint. If the voiceprint matching is unsuccessful, create a new user storage space.
[0017] Further, in step S112,
[0018] The speech features include linear prediction cepstral coefficients (LPCC) and Mel frequency cepstral coefficients (MFCC).
[0019] Further, in step S113, the human voice wake-up signal includes a wake-up word, and the wake-up word is stored in the common area of the voiceprint recognition model.
[0020] Further, in step S100, the visual wake-up process includes the following steps:
[0021] S130. In the standby state, control the visual sensor of the virtual digital human to collect images;
[0022] S140. Detect human skeleton recognition and / or face recognition, and detect whether someone is approaching;
[0023] S150. If it is detected that someone is approaching, wake up the virtual digital human and make the virtual digital human switch from the standby state to the working state.
[0024] Further, in step S100, after starting the virtual digital human, it is optional to skip the standby state and directly enter the working state.
[0025] Further, in step S200,
[0026] The multi-modal sensors of the virtual digital human include at least a depth sensor, an RGB camera, and a microphone. The multi-modal sensors are at least used for face recognition, skeleton recognition, posture, and gesture recognition of people.
[0027] Further, in step S300, it includes the following steps;
[0028] S310. After entering the working state, receive the user's image and voice through the multi-modal sensors, and input the voice and image into the voice and image extraction network to extract the voice and image;
[0029] S320. Convert the speech to text based on the speech-to-text service;
[0030] S330. Output the language text based on the large language model service;
[0031] S340. Convert the language text output by the large language model service to speech based on the text-to-speech service;
[0032] S350. Output sound, expressions and actions based on the virtual human animation control program;
[0033] S360. If the multi-modal sensor receives a response from the user, repeat steps S310 to S350 to provide feedback on the received content and output speech, expressions and actions.
[0034] Furthermore, in the step S310, the speech and image extraction network includes: a speech separation network and a visual separation network.
[0035] The speech separation network includes a mixed speech reception module, an STFT speech time-frequency transformation module, a speech feature extraction module, a speech upsampling module, and an ISTFT frequency-time transformation module connected in sequence.
[0036] The visual separation network includes a mixed visual reception module, a VGGFace feature extraction module, and a visual feature network module connected in sequence.
[0037] It further includes a fusion module. The input of the fusion module is respectively connected to the speech feature extraction module and the VGG feature extraction module, and the output of the fusion module is connected to the speech upsampling module.
[0038] Furthermore, in the step S330, the large language model includes a corpus system module, a pre-trained model, and a fine-tuning module connected in sequence.
[0039] The corpus system module includes pre-training corpus and fine-tuning corpus. The pre-training corpus includes text data collected from books, magazines, and encyclopedias, and the fine-tuning corpus includes annotated text data crawled from open-source code libraries, annotated by experts, and processed in the form of user conversations.
[0040] The pre-trained model is first trained in a large amount of general training using an unsupervised learning method on a large scale of training data to obtain a general and strong generalization ability language model. The training method of the pre-trained model includes at least word vector embedding, contrastive pre-training, and context learning.
[0041] The fine-tuning module includes freezing the bottom layer levels of the pre-trained model and adjusting the weights of the upper layer levels. The bottom layer levels include word vectors, and the upper layer levels include classifiers.
[0042] Further, the training steps of the pre-trained model include: collecting a dataset and performing supervised fine-tuning; collecting comparison data and training a scoring model, and using reinforcement learning for the reward model to optimize the cloud training model.
[0043] Further, the collecting a dataset and performing supervised fine-tuning includes sampling a prompt from the dataset; marking the preferred output answer, and performing supervised fine-tuning on the pre-trained model according to the collected data.
[0044] Further, the collecting comparison data and training a scoring model includes sampling a prompt and several corresponding model outputs; scoring and ranking the outputs, and training a scoring model on the scored and ranked data.
[0045] Further, the using reinforcement learning for the reward model to optimize the cloud training model includes resampling a prompt; the reward model scoring the output, and optimizing the model parameters.
[0046] Further, in step S300, it further includes receiving a task instruction, decomposing the task instruction into simple steps and executing them in sequence.
[0047] Further, in step S400, controlling the virtual digital human to change clothes, based on the character settings preset in the 3D clothing asset library, the 3D clothing asset library at least includes the clothes of the virtual digital human for hospital intelligent guidance, the clothes of the virtual digital human for museum explanation, and the clothes of the virtual digital human for restaurant greeting.
[0048] Further, in step S350, the virtual human animation control program includes a voice-to-expression service Audio2FaceService and a voice-to-action service Audio2ActionService. The voice-to-expression service Audio2FaceService is used to convert the virtual human voice information into the facial expressions of the virtual human, and the voice-to-action service Audio2ActionService is used to convert the virtual human voice information into the body movements of the virtual human.
[0049] Further, the voice-to-expression service Audio2FaceService is a voice animation synthesis model based on the BlendShapes method. The voice animation synthesis model based on the BlendShapes method includes a three-dimensional face control parameter prediction module based on different voice emotions, an expression basis construction module based on example expressions, and a three-dimensional face animation synthesis module.
[0050] The three-dimensional face control parameter prediction module based on different voice emotions includes a voice preprocessing module, a voice feature extraction module, an audio-visual mapping module, and a three-dimensional face control parameter module connected in sequence;
[0051] The example-expression-based expression basis construction module includes a sample expression model library, an expression feature extraction module, an expression migration module, and an expression basis construction module that are connected in sequence;
[0052] The three-dimensional human face animation synthesis module includes a three-dimensional human face control parameter smoothing module and a three-dimensional human face animation synthesis module that are connected in sequence. The input of the three-dimensional human face control parameter smoothing module is connected to the output of the three-dimensional human face control parameter module, and the input of the three-dimensional human face animation synthesis module is connected to the output of the expression basis construction module.
[0053] Furthermore, the present invention also proposes an interaction device for a virtual digital human, which is characterized by including:
[0054] A display device, which is a vertical large screen;
[0055] A multimodal input device, which includes a 4K high-definition RGB camera, a Tof depth camera, and a microphone array. The multimodal input device is arranged above the display device;
[0056] A real-time rendering device for performing real-time rendering output on the actions and expressions of the virtual digital human;
[0057] A heat dissipation device, which is used to dissipate heat for the real-time rendering device. The heat dissipation device is installed on the real-time rendering device;
[0058] A storage device, which includes a long-term storage device and a short-term storage device. The long-term storage device is a database, and the database includes a relational database and a vector database. The short-term storage device is a memory;
[0059] A processing device for processing the interaction data between the virtual digital human and the user. The processing device is installed on the back of the display device, and the display device, the multimodal input device, the real-time rendering device, and the storage device are respectively connected to the processing device.
[0060] Furthermore, the database includes an action library and an expression library. Each action record in the action library includes an action tag, and each expression record in the expression library includes an expression tag. The action tag and the expression tag are respectively called by the virtual digital human when replying to the user.
[0061] Furthermore, it further includes a tool library interface, and the tool library interface at least includes a printer interface, a search interface, an encyclopedia interface, a date and time interface, and an execution code interface.
[0062] Further, the processing device includes an AI algorithm service program, an interactive backend program, a background web program, and a virtual human rendering program.
[0063] Further, the AI algorithm service program is a Python program, and the AI algorithm service program includes an Audio2Face algorithm service, a Text2Action algorithm service, a large language model distribution service, and an Al Agent service.
[0064] Further, the interactive backend program is a Java program, and the interactive backend program includes an account management system, a memory configuration system, a 3D asset management system, and a large model interaction management system. The interactive backend program is connected to the database.
[0065] Further, the background web program is an H5 program, and the background web program includes an account management system page, a content configuration system page, a 3D asset management system page, and a large model interaction management page.
[0066] Further, the virtual human rendering program includes a UE5 program, and the virtual human rendering program includes virtual human rendering, action animation logic, and an external interaction program.
[0067] Further, the display device is an LED screen, a naked-eye 3D screen, or a holographic screen.
[0068] Further, the present invention also provides a computer-readable storage medium, on which program instructions are stored. When the program instructions are executed by a processor, the interactive method of the virtual digital human is implemented.
[0069] According to some embodiments of the present invention, the beneficial effects of the present invention are as follows:
[0070] The interactive method of the virtual digital human of the present invention can realize real-time natural language communication between the virtual human and the user in a specific environment, and cooperate with the facial expressions, actions, and professional costumes of the virtual digital human to achieve the technical effect of enhancing the integration of the virtual digital human and the real scene. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] Figure 1 It is a flowchart of the interactive method of the virtual digital human in the embodiment of the present invention.
[0072] Figure 2 It is a flowchart of the voice wake-up process and the visual wake-up process in the interactive method of the virtual digital human in the embodiment of the present invention.
[0073] Figure 3 It is a flowchart of the interaction between the virtual digital human and the user based on large language prompt engineering in the interactive method of the virtual digital human in the embodiment of the present invention.
[0074] Figure 4 This is a schematic diagram of the virtual human AI brain (AI agent) architecture in the embodiments of the present invention.
[0075] Figure 5 This is a schematic diagram of the overall software architecture in the embodiments of the present invention.
[0076] Figure 6 This is a schematic diagram of the visual wake-up process in the embodiments of the present invention.
[0077] Figure 7 This is a schematic diagram of the voice wake-up process in the embodiments of the present invention.
[0078] Figure 8 This is a schematic diagram of the AI multi-modal interaction 3D virtual human large screen all-in-one machine in the embodiments of the present invention.
[0079] Figure 9 This is a schematic diagram of the AI multi-modal interaction 3D virtual human transparent screen all-in-one machine in the embodiments of the present invention.
[0080] Figure 10 This is a flowchart of voiceprint recognition in the embodiments of the present invention.
[0081] Figure 11 This is a flowchart of data modeling for sound in the embodiments of the present invention.
[0082] Figure 12 This is a basic schematic diagram of the voice image extraction network in the embodiments of the present invention.
[0083] Figure 13 This is a basic schematic diagram of the large language model network structure in the embodiments of the present invention.
[0084] Figure 14 This is a flowchart of the training process of the pre-trained model in the embodiments of the present invention.
[0085] Figure 15 This is a basic structural schematic diagram of the voice animation synthesis model based on the BlendShapes method in the embodiments of the present invention.
[0086] Figure 16 This is a schematic diagram of the structure of the lip movement detection module in the embodiments of the present invention. Detailed implementation manners
[0087] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0088] The concept, specific structure, and technical effects of the present invention will be clearly and completely described below in combination with embodiments and the accompanying drawings to fully understand the objectives, solutions, and effects of the present invention.
[0089] It should be noted that unless otherwise specified, when a feature is referred to as "fixed" or "connected" to another feature, it can be directly fixed or connected to the other feature, or indirectly fixed or connected to the other feature. The singular forms "a", "the", and "said" used herein are also intended to include the plural forms unless the context clearly indicates otherwise. In addition, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this technology belongs. The terms used in the description of this specification are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0090] It should be understood that although the terms first, second, third, etc. may be used in this disclosure to describe various elements, these elements should not be limited to these terms. These terms are only used to distinguish elements of the same type from each other. For example, without departing from the scope of this disclosure, the first element may also be referred to as the second element, and similarly, the second element may also be referred to as the first element. The use of any and all example or exemplary language ("for example", "such as", etc.) provided herein is only intended to better illustrate the embodiments of the present invention and will not impose a limitation on the scope of the present invention unless otherwise required. In addition, the industry term "pose" used herein refers to the position and orientation of a certain element relative to a spatial coordinate system.
[0091] Refer to Figures 1 to 16 , embodiments of the present invention provide an interaction method, device, and storage medium for a virtual digital human, characterized in that, refer to Figure 1 , the method includes:
[0092] S100. Determine to receive a wake-up signal, cause the virtual digital human to enter a wake-up process, and switch from a standby state to a working state. The wake-up process includes a voice wake-up process and a visual wake-up process;
[0093] S200. Obtain environmental information and user information based on a multi-modal sensor of a virtual digital human;
[0094] S300. The virtual digital human interacts with the user based on large language prompt engineering;
[0095] S400. Adjust the character setting of the virtual digital human based on the prompt engineering of a large language model (LLM).
[0096] According to some embodiments of the present invention, the beneficial effects of the present invention are as follows:
[0097] The interaction method of the virtual digital human of the present invention can achieve real-time natural language communication between the virtual human and the user in a specific environment, and cooperate with the facial expressions, movements and professional costumes of the virtual digital human to achieve the technical effect of enhancing the integration of the virtual digital human and the real scene.
[0098] In the prior art, the IP of 3D virtual humans with live-action capture has become more and more good-looking and vivid in performance. However, this 3D virtual human IP is based on the personal ability of the "person behind the scenes". There have been incidents such as the departure of the "person behind the scenes" and the exposure of the "person behind the scenes" in A-soul on Bilibili. Therefore, the over-reliance of 3D virtual human IP on the ability of the "person behind the scenes" has affected the industrialization and productization of 3D virtual humans. If AI can replace the "person behind the scenes" as the skin of 3D virtual humans, the industrialization and productization problems of 3D virtual humans can surely be solved. Large language models (such as ERNIE Bot, ChatGPT, Bard, etc.) as question-and-answer robots can only output content in text and voice forms. For users, they cannot feel the service temperature brought by artificial intelligence. We believe that the next-generation interaction form will necessarily be an interaction mainly based on natural language. Therefore, it is an irresistible trend to give a virtual image to a large language model.
[0099] To implement the above ideas, the following three aspects of problems need to be solved. They are problems at the product level, including the problem that 3D virtual humans rely on the "zhongzhiren" (the person behind the virtual character), how to make the interaction between people and large language models more empathetic, which entity scenarios the product combining 3D virtual humans and large language models should serve, what problems the product of 3D virtual humans and large language models can actually solve, and how 3D virtual humans can enhance their expressiveness to make us feel that the virtual human is empathetic; it also includes problems at the software technology level: in the absence of a "zhongzhiren", the problems of virtual human motion and expression driving, the character setting and personality of 3D virtual humans, how 3D virtual humans obtain external information, i.e., "ears", "eyes", "space perception", etc., the problem that 3D virtual humans need to have memory capabilities, 3D virtual humans should be able to independently use tools, where the tools generally refer to software tools, and 3D virtual humans should be able to break down complex tasks into subtasks and execute them step by step; it also includes problems at the hardware level: on what medium to carry the image of the virtual human, and this image looks the same size as a real person, there needs to be one or more sensors for the virtual human to obtain external information, the real-time rendering problem of 3D virtual humans, the geometric surface of a sufficiently refined 3D virtual human has a high number of faces and a large texture size, which requires a high 3D rendering ability for the hardware, the high-performance rendering ability will inevitably bring high heat generation and heat dissipation problems, the fans added to solve the heat dissipation problem will cause acoustic structure problems, the problem of sound input in a noisy environment, and the problem of distinguishing voices and speakers in a multi-person environment.
[0100] Therefore, we propose the following solutions.
[0101] (1) Solutions to problems at the product level
[0102] In the past, 3D virtual humans relied on the "zhongzhiren" as the soul. The "zhongzhiren" is the soul, and a virtual human without a "zhongzhiren" can do nothing. Our 3D virtual humans use large language models to replace the "zhongzhiren" as the soul. At the same time, it also solves the problem that virtual humans can serve multiple scenarios. The Magic Realism All-in-One Machine (AI Multimodal Interaction 3D Virtual Human Large Screen) solves the traditional interaction solution where large language models can only output in voice and text, making the large language model the soul of the virtual human, and the virtual human becomes the beautiful shell of the large language model. The Magic Realism All-in-One Machine (AI Multimodal Interaction 3D Virtual Human Large Screen) adds an image to the output method of the large language model, making the service more empathetic. Most importantly, the Magic Realism All-in-One Machine (AI Multimodal Interaction 3D Virtual Human Large Screen) can replace human labor to independently complete part of the work, and has an elegant image, which can bring real foot traffic to merchants.
[0103] (2) Solutions to problems at the software technology level
[0104] Refer to Figure 4 and Figure 5, to enable the large language model to act as the soul of the virtual human, it is necessary to solve the driving problem of the virtual human. An Audio2Animation algorithm has been developed in-house. This algorithm consists of two parts. One part is Audio2Face, which converts speech into expressions, and Audio2Action, which converts speech into actions. The Magic Realization All-in-One Machine, through the Prompt Engineering of the large language model, has preset several character settings and assisted 3D clothing assets that match the character's occupation, giving the virtual human a professional attribute and making it appear more professional. Through multi-modal sensors, the virtual human can obtain various information by understanding the external environment. The depth sensor reconstructs the three-dimensional environment, enabling the virtual human to understand the environment. The RGB camera performs visual segmentation on the environment, object recognition, face recognition, human skeleton recognition, and posture and gesture recognition on the images segmented by vision. Through a localized vector database and a relational database, the virtual human has long-term memory. A sufficient tool library (including at least a printer interface, a search interface, an encyclopedia interface, a time and date interface, and an interface for executing code) is provided to the virtual human, allowing the virtual human to select suitable tools for use through the large language model. The Magic Realization All-in-One Machine, through the Prompt Engineering of the large language model, enables the virtual human to have the ability to solve complex problems. That is, the AI virtual human can break down complex task commands into relatively simple steps by itself and execute them sequentially. During the execution process, it can select tools or plugins and can also make autonomous selections.
[0105] (3) Solutions to Hardware-Level Problems
[0106] Select an 86-inch vertical large screen and an 86-inch transparent large screen, so that the rendered virtual human can be the same size as a real person, and on the transparent large screen, the virtual human has a naked-eye 3D feeling. Endow the virtual human with abilities such as "eyes", "ears", and "space perception" to obtain external information. Select a ring-shaped 7-microphone microphone array and a 4K high-definition RGB camera with HDR. And it is equipped with a Tof depth camera sensor. In order to integrate these three sensors, a depth vision sensor Azure Kinect DK used in industrial robots is selected. For the high-quality real-time rendering of 3D virtual humans, a high-performance image rendering card RTX3080 with ray tracing from NVIDIA is selected. Due to the heat dissipation problem caused by the selection of the high-performance image rendering card, as well as the chassis noise and vibration problems caused by the fans added to solve the heat dissipation problem, we choose a split modular design in the hardware design of the all-in-one machine, which not only solves the heat dissipation and sound collection problems, but also reduces the later maintenance cost and transportation difficulty. Regarding the problem that the vibration of the fan affects the built-in microphone array, we integrate the microphone with various sensors into a box and place it externally on the top of the chassis, perfectly solving the vibration problem of the chassis. Self-developed acoustic noise reduction algorithm and voice enhancement algorithm to solve the problem of voice input in a noisy environment. And through the narrow beam algorithm combined with face recognition for directional sound collection, the sound collection effect is better. In the case of multiple people speaking, through the combination of face recognition algorithm and lip movement recognition module, accurately direct the sound collection from multiple speakers, and combine the stored voiceprint library to separate the voiceprint of the target person from the voices of multiple people speaking.
[0107] In some specific embodiments, the lip movement recognition module includes a target detection module, a lip movement detection module, and an interaction module connected in sequence.
[0108] The target detection module needs to accurately find the lip position of the face and cut the image according to the predicted coordinates to ensure that the lips are located in the middle of the image. The project implementation algorithm is the Yolov5 algorithm, and the smallest pre-trained model is used for training. In the production of the target detection data set, it is necessary to ensure the integrity of the data, and the marked lips should be located in the middle of the image to achieve the effect of lip positioning, which can greatly accelerate the fitting speed of the subsequent classification network.
[0109] Lip movement detection module. The lip movement detection module adopts a composite network of 3DResNet and GRU. The data processed by the Yolov5 algorithm is input into this network. The residual structure extracts features, and GRU ensures the transmission and preservation of temporal information. Then, the predicted result is obtained through softmax. In the lip movement detection module, the quantity of training data is particularly crucial. Therefore, in preprocessing, data augmentation is used to enable the network to have sufficient data for training. Secondly, the structure of the network is also very important. The residual network ResNet is stacked by multiple layers of networks, which solves the problem of gradient disappearance caused by the network depth, makes the network deeper, and the more effective the image information features extracted are. The variant GRU of the recurrent neural network RNN, through the gate control mechanism, well preserves the temporal information, enabling the neural network to pay more attention to the dynamic change information of the lips in the time series.
[0110] Specifically, in some specific embodiments, referring to Figure 16 , the lip movement detection module is a GRU embedded in one layer of ResNet, also forming a composite network of CNN and RNN. However, the network does not simply use the structure of 3D convolution and pooling because, in the process of feature extraction, it is more inclined to use CNN to extract pixel features and GRU to extract features of temporal dynamic changes.
[0111] Interaction module. The predicted result is finally input into the interaction module, and the front-end and back-end interaction is completed relying on html and js to implement the recognition function of the system. In the interaction module, the framework used is the Flask framework. The advantage of this framework is that it is lightweight and flexible. In the login function, html directly submits the form to the back-end, and the back-end can complete the login function as long as it verifies the form through the database. In the recognition function, js cooperates with html to read the local video and then submit it to the background for processing. The processing process first involves video frame cutting, then the Yolov5 model predicts the lip coordinates of the frame images, then cuts and saves the images, and finally obtains the predicted result through the classification network. The recognized words are displayed on the front-end page to complete the entire functional process.
[0112] Specifically, in step S400, the prompt engineering gives the character setting of the virtual human, which is not only used for changing clothes, but also enables the virtual human to play a professional role, such as a doctor, a lawyer, a primary school English teacher, etc.
[0113] Furthermore, referring to Figure 2 and Figure 7 , in step S100, the voice wake-up process includes the following steps:
[0114] S110. In the standby state, collect the ambient sound and detect whether there is a voice wake-up signal, where the voice wake-up signal includes one or a combination of voice, wake-up keywords, and voiceprint features;
[0115] S120. If the detected voice wake-up signal reaches a preset threshold, wake up the virtual digital human, causing the virtual digital human to switch from the standby state to the working state.
[0116] Further, in step S110, detecting whether there is a voice wake-up signal includes the following steps:
[0117] S111. Perform noise reduction preprocessing on the received ambient sound, where the noise reduction preprocessing includes frame splitting, windowing, pre-emphasis, and voice activity detection (VAD) on the sound signal;
[0118] S112. Extract voice features based on MDCC and input the voice features into the voiceprint recognition model;
[0119] S113. Detect whether there is a voice wake-up signal and perform voiceprint matching between the voice features and the existing voiceprints in the voiceprint matching library;
[0120] S114. If the voiceprint matching is successful, retrieve the user information matching the voiceprint. If the voiceprint matching is unsuccessful, create a new user storage space.
[0121] In some specific embodiments, in step S114, the user storage space is set in Amazon's private repository.
[0122] In some specific embodiments, the voiceprint recognition process is to first extract voice features, then input the features into the model for training, and finally find the result with the highest score or the closest match. Refer to Figure 10 , voiceprint recognition includes a training stage and a testing stage. The training stage includes inputting training speech, feature extraction, model training, and voiceprint library search or generating a new target user registration speech. The testing stage includes inputting test speech, feature extraction, voiceprint matching scoring, and distinguishing target users or non-target users. For a virtual digital human, a target user can be defined as a user who has interacted with the virtual digital human and is stored in the memory, and a non-target user can be defined as a user who interacts with the virtual digital human for the first time, and the virtual digital human needs to create a new storage space to store the user's information.
[0123] In some specific embodiments, refer to Figure 11 , the flowchart of the virtual digital human performing data modeling on the sound. The virtual digital human first receives the sound, performs preprocessing, MFCC, and voiceprint recognition on the sound, and then communicates with the user or creates a new user according to the recognition result.
[0124] Further, in step S112,
[0125] The speech features include linear prediction coefficients LPCC and Mel cepstral coefficients MFCC.
[0126] Specifically, speech contains many characteristics. Feature extraction is to convert the time-domain continuous signal of the speech signal into a frequency-domain discrete signal, analyze the energy of the speech signal, and extract the required speech features from the time-domain continuous speech signal to represent this speech. There are many types of speech features, including linear prediction coefficients (LPCC) and Mel cepstral coefficients (MFCC) (Girin-Laurent).
[0127] The cochlea, an organ of the human body, helps humans receive speech signals and distinguish sounds in noisy environments. It is thus found that the cochlea plays a role in filtering out noise when humans receive speech signals. Therefore, a filter is constructed using the method of human sound discrimination to replace the function of the human cochlea to receive signals and achieve the purpose of noise filtering. The human cochlea is similar to a combination of multiple filters. It is very sensitive to the energy of low-frequency signals. It has a linear relationship with frequencies below 1000 Hz and a logarithmic relationship with frequencies above 1000 Hz. This irregular frequency division is an important feature that distinguishes the Mel cepstrum of speech signals from the ordinary cepstrum. The Mel cepstrum is one of the most common speech features in speech signal processing. Through research on the biological functions of the human ear, it is found that the human ear has different auditory sensitivities to speech signals of different amplitudes and frequencies. The acceptance range of the human ear for the amplitude and frequency of speech signals is limited. For speech signal amplitude frequencies that exceed what the human ear can receive, the human ear cannot hear them. For example, ultrasonic frequencies are around 20000 Hz, and speech signals with frequencies less than 20 Hz are also inaudible to the human ear. The frequency range of speech signals that the human ear can hear is within 20 Hz to 20000 Hz. Noise in the frequency range of 200 Hz to 5000 Hz has the greatest impact on speech signals, making the text content in the speech signal unclear.
[0128] Further, in the step S113, the human voice wake-up signal includes a wake-up word, and the wake-up word is stored in the common area of the voiceprint recognition model.
[0129] In some embodiments, the virtual digital human speech monitoring software collects environmental sounds, detects whether there is a human voice wake-up signal. If the detected human voice wake-up signal reaches the threshold, the virtual digital human is awakened and enters the working state; after entering the working state, the virtual digital human outputs a response signal to the user and collects the user's speech signal through a microphone; the virtual digital human converts the collected user speech signal into text through a speech-to-text service, inputs it into a large language model service, the large language model service outputs text, and outputs speech through a text-to-speech service; the virtual digital human outputs expressions and actions based on an animation control program.
[0130] Further, referring to Figure 2 and Figure 6 , in the step S100, the visual wake-up process includes the following steps:
[0131] S130. In the standby state, control the visual sensor of the virtual digital human to collect images;
[0132] S140. Detect human skeleton recognition and / or face recognition to detect whether someone is approaching;
[0133] S150. If someone is detected approaching, wake up the virtual digital human to convert the virtual digital human from the standby state to the working state.
[0134] In some specific embodiments, the visual sensor of the virtual digital human collects images, and detects whether someone is approaching through human skeleton recognition and face recognition; if someone is detected approaching, the virtual digital human requests the AI brain, and outputs language text through the large language model service; based on the text-to-speech service, converts the language text output by the large language model service into speech; based on the virtual human animation control program, outputs sound, expressions and actions; if the microphone receives the user's response, feedback and output the received content; receives the voice information of the user's response and the question asked through the microphone; based on the speech-to-text service, converts the voice information into text; based on the large language model service, inputs the text information and outputs the corresponding text information; based on the text-to-speech service, converts the corresponding text information into virtual human voice information; based on the virtual human animation control program, outputs expressions and actions.
[0135] In some specific embodiments, as an alternative to the wake-up algorithm, capture the state of the face and lips of the speaking object through vision, specifically including the following steps:
[0136] Lock the current interactive person through face or skeleton;
[0137] Capture the lip state of the current interactive person through computer vision algorithms;
[0138] When the lip state is open and closed, turn on the microphone for directional sound collection, and when the lips are in the closed state, turn off the microphone for directional sound collection, thus solving the problem of mis-sound collection in a noisy environment or when multiple people are speaking;
[0139] Through the lip reading recognition algorithm, continuously record the lip movement video images, and convert the lip movement video images into lip reading to enhance the speech recognition of the microphone.
[0140] Further, in the step S100, after starting the virtual digital human, it is optional to skip the standby state and directly enter the working state.
[0141] Specifically, after the virtual digital human is powered on, it can directly play animations and voices on the display screen without detecting whether the user is approaching, attracting the user's attention.
[0142] Furthermore, in the step S200,
[0143] The multi-modal sensors of the virtual digital human include at least a depth sensor, an RGB camera, and a microphone, and the multi-modal sensors are at least used for face recognition, bone recognition, and posture and gesture recognition of the person.
[0144] Furthermore, referring to Figure 3 、 Figure 6 and Figure 7 , in the step S300, the following steps are included;
[0145] S310. After entering the working state, the virtual digital human receives the user's image and voice through the multi-modal sensors, and transmits the voice and image to the voice and image extraction network to extract the voice and image;
[0146] S320. Based on the speech-to-text service, convert the voice into text;
[0147] S330. Based on the large language model service, output the language text;
[0148] S340. Based on the text-to-speech service, convert the language text output by the large language model service into voice;
[0149] S350. Based on the virtual human animation control program, output sounds, expressions, and actions;
[0150] S360. If the multi-modal sensors receive a response from the user, repeat steps S310 to S350 to feedback the received content and output voice, expressions, and actions.
[0151] Furthermore, referring to Figure 12 , in the step S310, the voice and image extraction network includes: a voice separation network and a visual separation network,
[0152] The voice separation network includes a mixed voice receiving module, an STFT voice time-frequency transformation module, a voice feature extraction module, a voice upsampling module, and an ISTFT frequency-time transformation module connected in sequence,
[0153] The visual separation network includes a mixed visual receiving module, a VGGFace feature extraction module, and a visual feature network module connected in sequence,
[0154] It also includes a fusion module, the input of the fusion module is respectively connected to the voice feature extraction module and the VGG feature extraction module, and the output of the fusion module is connected to the voice upsampling module.
[0155] Specifically, referring to Figure 12 , the voice-image extraction network structure mainly consists of two parts: a voice separation network and a visual separation network. The voice separation network receives a mixed voice spectrogram and outputs a spectral mask. The network model structure adopts an encoder-decoder structure similar to U-Net, which are respectively called a feature extraction module and an upsampling module. And there is a skip connection between the feature extraction module and the upsampling module. Its main function is to use the shallow features of the feature extraction module to promote the upsampling module to predict a high-resolution time-frequency mask. However, there are significant differences in the performance of feature extraction modules with different depths on different tasks. To enable the network to adaptively learn network features of different depths, a residual connection is added to the feature extraction module. The visual separation network receives a set of faces that exactly corresponds to the voice spectrogram and outputs visual features that can be distributed identically to the voice separation network. A residual connection is added on the basis of the U-Net model and it includes an audio-visual feature fusion module.
[0156] Furthermore, referring to Figure 13 , in the step S330, the large language model includes a corpus system module, a pre-trained model, and a fine-tuning module connected in sequence,
[0157] The corpus system module includes pre-training corpus and fine-tuning corpus. The pre-training corpus includes text data collected from books, magazines, and encyclopedia channels. The fine-tuning corpus includes annotated text data crawled from open-source code libraries, annotated by experts, and processed in the way of user conversations;
[0158] The pre-trained model is first trained in a large amount on a large-scale training data using an unsupervised learning method to obtain a general and strongly generalized language model. The training method of the pre-trained model at least includes word vector embedding, contrastive pre-training, and context learning;
[0159] The fine-tuning module includes freezing the bottom layer levels of the pre-trained model and adjusting the weights of the upper layer levels. The bottom layer levels include word vectors, and the upper layer levels include classifiers.
[0160] Specifically, the corpus system includes pre-training corpus and fine-tuning corpus. The latter includes code and dialogue fine-tuning corpus. The pre-training corpus includes a large amount of text data collected from channels such as books, magazines, and encyclopedias, enabling the model to learn the logical relationship expression methods of language; the fine-tuning corpus includes high-quality annotated text data processed in ways such as crawling from open-source code libraries, expert annotation, and user conversations, further enhancing its dialogue ability.
[0161] In some embodiments, pre-training is the basis for building large-scale language models, which refers to first conducting a large amount of general training on large-scale training data and using unsupervised learning methods for training to obtain a general and strong generalization ability language model. Based on the large-scale data, through pre-training, the model initially has the ability to understand human language and perform context learning, can capture the semantic similarity features of text fragments and code fragments, and thus generate more accurate text and code vectors to support subsequent fine-tuning tasks.
[0162] In some embodiments, fine-tuning is the guarantee for realizing the actual application of the model, which refers to further training the pre-trained model on the dataset of a specific task, usually including freezing the bottom layers (such as word vectors) of the pre-trained model and adjusting the weights of the upper layers (such as classifiers). Fine-tuning the pre-trained model will greatly shorten the training time, save computing resources and accelerate the training convergence speed. Based on the pre-trained model with strong generalization ability, by integrating training based on code data and fine-tuning based on instructions, and using a specific dataset for fine-tuning, it makes the model have a stronger ability to generate question-and-answer dialogue texts.
[0163] Further, referring to Figure 14 , the training steps of the pre-trained model include: collecting a dataset and performing supervised fine-tuning; collecting comparison data and training a scoring model and using reinforcement learning for the reward model to optimize the cloud training model.
[0164] Further, referring to Figure 14 , the collecting a dataset and performing supervised fine-tuning includes sampling a prompt from the dataset; marking the preferred output answer and performing supervised fine-tuning on the pre-trained model according to the collected data.
[0165] Further, referring to Figure 14 , the collecting comparison data and training a scoring model includes sampling a prompt and several corresponding model outputs; scoring and ranking the outputs and training a scoring model on the scored and ranked data.
[0166] Further, referring to Figure 14 , the using reinforcement learning for the reward model to optimize the cloud training model includes resampling a prompt; the reward model scoring the output and optimizing the model parameters.
[0167] Further, in step S300, it further includes receiving a task instruction, disassembling the task instruction into simple steps and executing them in sequence.
[0168] Further, in step S400, controlling the virtual digital human to change clothes, based on the character settings preset in the 3D clothing asset library, and the 3D clothing asset library at least includes the clothes of the virtual digital human for hospital intelligent guidance, the clothes of the virtual digital human for museum explanation, and the clothes of the virtual digital human for restaurant greeting.
[0169] Further, in the step S350, the virtual human animation control program includes a voice-to-expression service Audio2FaceService and a voice-to-action service Audio2ActionService. The voice-to-expression service Audio2FaceService is used to convert virtual human voice information into the facial expressions of the virtual human, and the voice-to-action service Audio2ActionService is used to convert virtual human voice information into the body movements of the virtual human.
[0170] Further, referring to Figure 15 , the voice-to-expression service Audio2FaceService is a voice animation synthesis model based on the BlendShapes method. The voice animation synthesis model based on the BlendShapes method includes a three-dimensional face control parameter prediction module based on different voice emotions, an expression basis construction module based on example expressions, and a three-dimensional face animation synthesis module.
[0171] The three-dimensional face control parameter prediction module based on different voice emotions includes a voice preprocessing module, a voice feature extraction module, an audio-video mapping module, and a three-dimensional face control parameter module connected in sequence.
[0172] The expression basis construction module based on example expressions includes an example expression model library, an expression feature extraction module, an expression migration module, and an expression basis construction module connected in sequence.
[0173] The three-dimensional face animation synthesis module includes a three-dimensional face control parameter smoothing module and a three-dimensional face animation synthesis module connected in sequence. The input of the three-dimensional face control parameter smoothing module is connected to the output of the three-dimensional face control parameter module, and the input of the three-dimensional face animation synthesis module is connected to the output of the expression basis construction module.
[0174] Further, referring to Figure 8 and Figure 9 , the present invention also proposes an interaction device for a virtual digital human, which is characterized by including:
[0175] A display device, and the display device is a vertical large screen;
[0176] A multi-modal input device, and the multi-modal input device includes a 4K high-definition RGB camera, a Tof depth camera, and a microphone array. The multi-modal input device is arranged above the display device;
[0177] A real-time rendering device for performing real-time rendering and output of the actions and expressions of the virtual digital human;
[0178] A heat dissipation device, which is used to dissipate heat for a real-time rendering device and is installed on the real-time rendering device;
[0179] A storage device, which includes a long-term storage device and a short-term storage device. The long-term storage device is a database, and the database includes a relational database and a vector database. The short-term storage device is a memory;
[0180] A processing device, which is used to process the interaction data between the virtual digital human and the user. The processing device is installed on the back of the display device, and the display device, the multi-modal input device, the real-time rendering device, and the storage device are respectively connected to the processing device.
[0181] Furthermore, the database includes an action library and an expression library. Each action record in the action library includes an action tag, and each expression record in the expression library includes an expression tag. The action tag and the expression tag are respectively called by the virtual digital human when replying to the user.
[0182] Furthermore, it also includes a tool library interface, and the tool library interface includes at least a printer interface, a search interface, an encyclopedia interface, a date and time interface, and an execution code interface.
[0183] Furthermore, the processing device includes an AI algorithm service program, an interaction backend program, a background web program, and a virtual human rendering program.
[0184] Furthermore, the AI algorithm service program is a Python program, and the AI algorithm service program includes an Audio2Face algorithm service, a Text2Action algorithm service, a large language model distribution service, and an Al Agent service.
[0185] Furthermore, the interaction backend program is a Java program, and the interaction backend program includes an account management system, a memory configuration system, a 3D asset management system, and a large model interaction management system. The interaction backend program is connected to the database.
[0186] Furthermore, the background web program is an H5 program, and the background web program includes an account management system page, a content configuration system page, a 3D asset management system page, and a large model interaction management page.
[0187] Furthermore, the virtual human rendering program includes a UE5 program, and the virtual human rendering program includes virtual human rendering, action animation logic, and an external interaction program.
[0188] Furthermore, the display device is an LED screen or a naked-eye 3D screen or a holographic screen.
[0189] Further, the present invention also provides a computer-readable storage medium storing program instructions, which, when executed by a processor, implement the interactive method of the virtual digital human described above.
[0190] It should be recognized that the method steps in the embodiments of the present invention can be implemented or carried out by computer hardware, a combination of hardware and software, or computer instructions stored in a non-transitory computer-readable memory. The method can use standard programming techniques. Each program can be implemented in a high-level procedural or object-oriented programming language to communicate with the computer system. However, if desired, the program can be implemented in assembly or machine language. In any case, the language can be a compiled or interpreted language. In addition, for this purpose, the program can run on a dedicated integrated circuit programmed for this purpose.
[0191] In addition, the operations of the processes described herein can be performed in any suitable order, unless otherwise indicated herein or otherwise clearly contradicted by the context. The processes described herein (or variations and / or combinations thereof) can be executed under the control of one or more computer systems configured with executable instructions, and can be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) executed commonly on one or more processors, by hardware, or a combination thereof. The computer program includes multiple instructions executable by one or more processors.
[0192] Further, the method can be implemented in any type of computing platform operably connected, including but not limited to personal computers, minicomputers, mainframes, workstations, network or distributed computing environments, separate or integrated computer platforms, or communicating with charged particle tools or other imaging devices, etc. Aspects of the present invention can be implemented in machine-readable code stored on a non-transitory storage medium or device, whether removable or integrated into the computing platform, such as a hard disk, optical read and / or write storage medium, RAM, ROM, etc., such that it can be read by a programmable computer and, when the storage medium or device is read by the computer, can be used to configure and operate the computer to execute the processes described herein. In addition, the machine-readable code, or portions thereof, can be transmitted via a wired or wireless network. When such media include instructions or programs that implement the steps described above in combination with a microprocessor or other data processor, the inventions described herein include these and other different types of non-transitory computer-readable storage media. When programmed according to the methods and techniques of the present invention, the present invention can also include the computer itself.
[0193] A computer program can be applied to input data to perform the functions described herein, thereby converting the input data to generate output data stored in a non-volatile memory. The output information can also be applied to one or more output devices such as a display. In a preferred embodiment of the present invention, the converted data represents physical and tangible objects, including a specific visual depiction of the physical and tangible objects generated on the display.
[0194] As described above, only the preferred embodiments of the present invention are given, and the present invention is not limited to the above embodiments. As long as the same means are used to achieve the technical effects of the present invention, any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention. Within the scope of protection of the present invention, its technical solutions and / or implementation manners can have various different modifications and changes.
Claims
1. A virtual digital human interaction method, characterized in that: The method includes: S100, determining to receive a wake-up signal, causing the virtual digital human to enter a wake-up process, switching from a standby state to a working state, wherein the wake-up process includes a voice wake-up process; S200, based on the multimodal sensor of virtual digital human, obtains environmental information and user information; S300, based on the large language prompt word project, the virtual digital human interacts with the user; S400, based on the large language model prompt word engineering, adjust the character settings of the virtual digital human; Wherein, in step S100, the voice wake-up process includes the following steps: S110, in standby mode, collecting environmental sounds and detecting whether there is a human voice wake-up signal, where the human voice wake-up signal includes a human voice, a wake-up keyword, and a voiceprint feature; S120, if the detected human voice wake-up signal reaches a preset threshold, wake up the virtual digital human, so that the virtual digital human is switched from a standby state to a working state; Wherein, in step S110, detecting whether there is a human voice wake-up signal includes the following steps: S111, performing noise reduction preprocessing on the received ambient sound, wherein the noise reduction preprocessing includes framing, windowing, pre-emphasis and adaptive endpoint detection (VAD) on the sound signal; S112, extracting speech features based on MDCC, and inputting the speech features into the voiceprint recognition model. The voiceprint recognition process is to first extract speech features, then input the features into the model for training, and finally find the result with the highest score or the closest score. Voiceprint recognition includes a training phase and a testing phase, wherein the training phase includes inputting training speech, feature extraction, model training, and voiceprint library search or generating a new target user registration voice; the testing phase includes inputting test speech, feature extraction, voiceprint matching scoring, and distinguishing target users or non-target users. For a virtual digital person, a target user can be defined as a user who has interacted with the virtual digital person and is stored in the memory, and a non-target user can be defined as a user who interacts with the virtual digital person for the first time. The virtual digital person needs to create a new storage space to store the user's information; the virtual digital person first receives the sound, pre-processes the sound, performs MFCC, and voiceprint recognition, and then communicates with the user or creates a new user based on the recognition result; S113, detecting whether there is a human voice wake-up signal, and matching the voice features with the existing voiceprints in the voiceprint matching library, wherein the human voice wake-up signal includes a wake-up word, and the wake-up word is stored in the public area of the voiceprint recognition model; the virtual digital human voice monitoring software collects environmental sounds, detects whether there is a human voice wake-up signal, and if the detected human voice wake-up signal reaches a threshold, wakes up the virtual digital human and enters the working state. After entering the working state, the virtual digital human outputs a response signal to the user, and collects the user's voice signal through a microphone. The virtual digital human converts the collected user voice signal into text through the voice-to-text service, and inputs it into the large language model service. The large language model service outputs text, and outputs voice through the text-to-speech service. The virtual digital human outputs expressions and actions based on the animation control program; S114, if the voiceprint matching is successful, retrieve the user information matching the voiceprint; if the voiceprint matching is unsuccessful, create a new user storage space; Wherein, the step S300 includes the following steps: S310, after entering the working state, receiving the user's image and voice through the multimodal sensor, and transmitting the voice and image to the voice and image extraction network to extract the voice and image; S320, converting speech into text based on the speech-to-text service; S330, outputting language text based on the large language model service; S340, based on the text-to-speech service, converting the language text output by the large language model service into speech; S350, outputting sounds, expressions and movements based on the virtual human animation control program; S360: If a response from the user is received, repeat steps S310 to S350 to provide feedback on the received content, outputting voice, expression and action; Wherein, in step S310, the speech image extraction network includes: a speech separation network and a visual separation network, The speech separation network includes a mixed speech receiving module, a STFT speech time-frequency conversion module, a speech feature extraction module, a speech upsampling module and an ISTFT frequency-time conversion module connected in sequence. The visual separation network includes a hybrid visual receiving module, a VGGFace feature extraction module, and a visual feature network module connected in sequence. It also includes a fusion module, the input of which is connected to the speech feature extraction module and the VGG feature extraction module respectively, and the output of which is connected to the speech upsampling module.
2. The method for interacting with a virtual digital human according to claim 1, characterized in that: In the step S112, The speech features include linear prediction coefficients LPCC and Mel cepstral coefficients MFCC.
3. The method for interacting with a virtual digital human according to claim 1, characterized in that: In the step S113, the human voice wake-up signal includes a wake-up word, and the wake-up word is stored in the common area of the voiceprint recognition model.
4. The method for interacting with a virtual digital human according to claim 1, characterized in that: In the step S200, The multimodal sensor of the virtual digital human includes at least a depth sensor, an RGB camera and a microphone, and the multimodal sensor is used at least for performing face recognition, skeleton recognition, posture and gesture recognition on the person.
5. The method for interacting with a virtual digital human according to claim 1, characterized in that: In step S330, the large language model includes a corpus system module, a pre-training model and a fine-tuning module connected in sequence. The corpus system module includes pre-trained corpus and fine-tuned corpus. The pre-trained corpus includes text data collected from books, magazines and encyclopedias, and the fine-tuned corpus includes annotated text data crawled from open source code libraries, annotated by experts, and processed by user dialogue. The pre-trained model is first trained in large quantities using an unsupervised learning method on large-scale training data to obtain a general language model with strong generalization ability. The training method of the pre-trained model at least includes word vector embedding, contrast pre-training and context learning; The fine-tuning module includes freezing the bottom layer of the pre-trained model and adjusting the weights of the upper layer, the bottom layer includes word vectors, and the upper layer includes classifiers.
6. The method for interacting with a virtual digital human according to claim 5, characterized in that: The training steps of the pre-trained model include: collecting data sets and performing supervised fine-tuning; collecting comparison data and training a scoring model and using reinforcement learning for a reward model to optimize the cloud training model.
7. The method for interacting with a virtual digital human according to claim 6, characterized in that: The collecting of a dataset and performing supervised fine-tuning includes sampling a prompt from the dataset; marking a preferred output answer and performing supervised fine-tuning on a pre-trained model based on the collected data.
8. The method for interacting with a virtual digital human according to claim 6, characterized in that: The collecting of comparison data and training of the scoring model includes sampling a prompt and a number of corresponding model outputs; scoring and sorting the outputs and training the scoring model on the scored and sorted data.
9. The method for interacting with a virtual digital human according to claim 6, characterized in that: The method of using reinforcement learning for a reward model to optimize a cloud-trained model includes resampling a prompt; the reward model scores the output and optimizes model parameters.
10. The method for interacting with a virtual digital human according to claim 1, characterized in that: The step S300 also includes receiving a task instruction, breaking the task instruction into simple steps and executing them in sequence.
11. The method for interacting with a virtual digital human according to claim 1, characterized in that: In the step S400, the virtual digital human is controlled to change clothes based on the character settings preset in the 3D clothing asset library, and the 3D clothing asset library at least includes hospital intelligent guidance virtual digital human clothing, museum explanation virtual digital human clothing and restaurant reception virtual digital human clothing.
12. The method for interacting with a virtual digital human according to claim 1, characterized in that: In step S350, the virtual human animation control program includes a voice-to-expression service Audio2FaceService and a voice-to-action service Audio2ActionService. The voice-to-expression service Audio2FaceService is used to convert the virtual human voice information into the virtual human's facial expressions, and the voice-to-action service Audio2ActionService is used to convert the virtual human voice information into the virtual human's body movements.
13. The method for interacting with a virtual digital human according to claim 12, characterized in that: The speech-to-expression service Audio2FaceService is a speech animation synthesis model based on the BlendShapes method, and the speech animation synthesis model based on the BlendShapes method includes a three-dimensional face control parameter prediction module based on different speech emotions, an expression base construction module based on sample expressions, and a three-dimensional face animation synthesis module. The three-dimensional face control parameter prediction module based on different voice emotions includes a voice preprocessing module, a voice feature extraction module, an audio and video mapping module and a three-dimensional face control parameter module connected in sequence; The expression base construction module based on sample expressions includes a sample expression model library, an expression feature extraction module, an expression migration module, and an expression base construction module connected in sequence; The 3D face animation synthesis module includes a 3D face control parameter smoothing module and a 3D face animation synthesis module which are connected in sequence. The input of the 3D face control parameter smoothing module is connected to the output of the 3D face control parameter module, and the input of the 3D face animation synthesis module is connected to the output of the expression base construction module.
14. An interactive device for a virtual digital human, characterized in that: The method for interacting with a virtual digital human according to any one of claims 1 to 13 is applied to the method; comprising: A display device, wherein the display device is a vertical large screen; A multimodal input device, the multimodal input device comprising a 4K high-definition RGB camera, a Tof depth camera and a microphone array, the multimodal input device being arranged above the display device; A real-time rendering device, used for real-time rendering and output of the movements and expressions of the virtual digital human; A heat dissipation device, the heat dissipation device is used to dissipate heat for the real-time rendering device, and the heat dissipation device is installed on the real-time rendering device; A storage device, wherein the storage device includes a long-term storage device and a short-term storage device, the long-term storage device is a database, the database includes a relational database and a vector database, and the short-term storage device is a memory; A processing device is used to process the interactive data between the virtual digital person and the user. The processing device is installed on the back of the display device. The display device, the multimodal input device, the real-time rendering device and the storage device are respectively connected to the processing device.
15. The interactive device for virtual digital human according to claim 14, characterized in that: The database includes an action library and an expression library, each action record in the action library includes an action tag, each expression record in the expression library includes an expression tag, and the action tag and the expression tag are respectively called by the virtual digital person when replying to the user.
16. The interactive device for virtual digital human according to claim 14, characterized in that: It also includes a tool library interface, which includes at least a printer interface, a search interface, an encyclopedia interface, a date and time interface, and an execution code interface. 17 . A computer-readable storage medium having program instructions stored thereon, wherein the program instructions implement the method according to claim 1 when executed by a processor.
Citation Information
Patent Citations
Method for personalized television voice wake-up by voiceprint and voice identification
CN104575504A
Virtual person-based multi-mode interactive processing method and system
CN107765852A
Virtual digital human-based interaction processing method, system, terminal, equipment and medium
CN117520498A