Virtual digital human interaction method and device based on large language model
Through the virtual digital human interaction method based on the large language model, the multimodal sensor and voice image extraction network are used to realize natural language communication between virtual digital humans and users, solving the problem that large language models cannot communicate with humans and 3D virtual humans and people depend on them, and improving the application and integration of virtual digital humans in real scenarios.
Patent Information
- Application Number
- CN202510883955.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-29
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2044-03-29
AI Technical Summary
The existing technology large language models cannot communicate with each other in conversation, which limits their application in on-site navigation environments, and the 3D virtual human IP is too dependent on people, affecting industrialization and productization.
The virtual digital human interaction method based on the large language model is adopted, and the environment and user information is obtained through multi-modal sensors, combined with voice image extraction network, voice to text and text to voice services, to realize natural language communication between virtual digital humans and users, and to output expressions and actions through the large language model service and virtual human animation control program, supporting multi-scene applications.
Real-time natural language communication between virtual digital people and users is realized, the integration of virtual digital people in real scenes and the diversity of application scenarios is improved, the problem of people dependence on 3D virtual people is solved, and the temperature and professionalism of the service is enhanced.
Smart Images

Figure CN120578296A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence applications, and in particular relates to a virtual digital human interaction method and device based on a large language model. Background Art
[0002] 3D virtual humans, using real-life motion capture, are becoming increasingly attractive and vivid. However, these 3D virtual human IPs rely on the individual abilities of the resident human. Existing large language models can only input and output text, not enable conversational communication between humans. This limits their application in on-site navigation environments. Summary of the Invention
[0003] The present invention provides a virtual digital human interaction method and device based on a large language model, aiming to solve at least one of the technical problems existing in the prior art.
[0004] The technical solution of the present invention relates to a virtual digital human interaction method based on a large language model, the method comprising:
[0005] S200, based on the multimodal sensor of the virtual digital human, obtains environmental information and user information;
[0006] S300, based on the large language prompt word project, the virtual digital human interacts with the user;
[0007] The step S300 includes the following steps:
[0008] S310: After entering the working state, the multimodal sensor receives the user's image and voice, and transmits the voice and image to the voice and image extraction network to extract the voice and image;
[0009] In step S310, the speech image extraction network includes: a speech separation network and a visual separation network.
[0010] The speech separation network includes a mixed speech receiving module, an STFT speech time-frequency conversion module, a speech feature extraction module, a speech upsampling module and an STFT frequency-time conversion module connected in sequence.
[0011] The visual separation network includes a hybrid visual receiving module, a VGGFace feature extraction module, and a visual feature network module connected in sequence.
[0012] It also includes a fusion module, wherein the input of the fusion module is connected to the speech feature extraction module and the VGG feature extraction module respectively, and the output of the fusion module is connected to the speech upsampling module;
[0013] S320, converting the speech into text based on the speech-to-text service;
[0014] S330: Outputting language text based on the large language model service;
[0015] S340: Based on the text-to-speech service, convert the language text output by the large language model service into speech;
[0016] S350, outputting sounds, expressions and movements based on the virtual human animation control program;
[0017] S360: If a response is received from the user, repeat steps S310 to S350 to provide feedback on the received content, outputting voice, expression, and action;
[0018] S400: Adjust the character settings of the virtual digital human based on the prompt word project of the large language model.
[0019] Furthermore, in step S200,
[0020] The multimodal sensor of the virtual digital human includes at least a depth sensor, an RGB camera and a microphone. The multimodal sensor is used at least to perform face recognition, skeleton recognition, posture and gesture recognition on the character.
[0021] Furthermore, in step S330, the large language model includes a corpus system module, a pre-training model and a fine-tuning module connected in sequence.
[0022] The corpus system module includes pre-trained corpus and fine-tuned corpus. The pre-trained corpus includes text data collected from books, magazines, and encyclopedias. The fine-tuned corpus includes annotated text data crawled from open source code libraries, annotated by experts, and processed through user dialogue.
[0023] The pre-trained model is first trained extensively on large-scale training data using an unsupervised learning method to obtain a general language model with strong generalization capabilities. The training method of the pre-trained model includes at least word vector embedding, contrastive pre-training, and contextual learning.
[0024] The fine-tuning module includes freezing the bottom layer of the pre-trained model and adjusting the weights of the upper layers, the bottom layer includes word vectors, and the upper layers include classifiers;
[0025] The training steps of the pre-trained model include: collecting a data set and performing supervised fine-tuning; collecting comparison data and training a scoring model and using reinforcement learning for the reward model to optimize the cloud training model;
[0026] The collecting of a dataset and performing supervised fine-tuning includes sampling a prompt from the dataset; marking a preferred output answer and performing supervised fine-tuning on a pre-trained model based on the collected data;
[0027] The collecting comparison data and training the scoring model includes sampling a prompt and a number of corresponding model outputs; scoring and sorting the outputs and training the scoring model on the scored and sorted data;
[0028] The method of using reinforcement learning for a reward model to optimize a cloud-trained model includes resampling a prompt; the reward model scores the output and optimizes model parameters.
[0029] Furthermore, the step S300 further includes receiving a task instruction, breaking the task instruction into simple steps and executing them in sequence.
[0030] Furthermore, in step S400, the virtual digital human is controlled to change clothes based on the character settings preset in the 3D clothing asset library, and the 3D clothing asset library at least includes hospital intelligent guidance virtual digital human clothing, museum explanation virtual digital human clothing, and restaurant reception virtual digital human clothing.
[0031] Furthermore, in step S350, the virtual human animation control program includes a voice-to-expression service Audio2FaceService and a voice-to-action service Audio2ActionService, wherein the voice-to-expression service Audio2FaceService is used to convert the virtual human's voice information into the virtual human's facial expression, and the voice-to-action service Audio2ActionService is used to convert the virtual human's voice information into the virtual human's body movements;
[0032] The audio-to-expression service Audio2FaceService is a voice animation synthesis model based on the BlendShapes method, which includes a three-dimensional face control parameter prediction module based on different voice emotions, an expression base construction module based on sample expressions, and a three-dimensional face animation synthesis module.
[0033] The three-dimensional face control parameter prediction module based on different voice emotions includes a voice preprocessing module, a voice feature extraction module, an audio and video mapping module and a three-dimensional face control parameter module connected in sequence;
[0034] The expression base construction module based on sample expressions includes a sample expression model library, an expression feature extraction module, an expression migration module, and an expression base construction module connected in sequence;
[0035] The three-dimensional facial animation synthesis module includes a three-dimensional facial control parameter smoothing module and a three-dimensional facial animation synthesis module connected in sequence. The input of the three-dimensional facial control parameter smoothing module is connected to the output of the three-dimensional facial control parameter module, and the input of the three-dimensional facial animation synthesis module is connected to the output of the expression base construction module.
[0036] Furthermore, the present invention also proposes a method for interacting with a virtual digital human, including the method for waking up the virtual digital human, and the method for interacting with the virtual digital human further includes:
[0037] S100: Determine whether a wake-up signal is received, causing the virtual digital human to enter a wake-up process, switching from a standby state to a working state, wherein the wake-up process includes a voice wake-up process and a visual wake-up process;
[0038] The voice wake-up process includes the following steps:
[0039] S110. In the standby state, collect ambient sound and detect whether there is a human voice wake-up signal, where the human voice wake-up signal includes one or more combinations of human voice, wake-up keywords, and voiceprint features;
[0040] S120: If the detected human voice wake-up signal reaches a preset threshold, wake up the virtual digital human and switch the virtual digital human from a standby state to a working state;
[0041] The visual awakening process includes the following steps:
[0042] S130, in standby mode, controlling the visual sensor of the virtual digital human to collect images;
[0043] S140, detecting human skeleton recognition and / or face recognition to detect whether someone is approaching;
[0044] S150: If it is detected that someone is approaching, wake up the virtual digital human and switch the virtual digital human from the standby state to the working state.
[0045] Furthermore, the present invention also proposes an interactive device for a virtual digital human, comprising:
[0046] A display device, wherein the display device is a vertical large screen;
[0047] A multimodal input device, comprising a 4K high-definition RGB camera, a Tof depth camera, and a microphone array, and disposed above the display device;
[0048] A real-time rendering device, used to render and output the movements and expressions of the virtual digital human in real time;
[0049] A heat dissipation device, the heat dissipation device is used to dissipate heat for the real-time rendering device, and the heat dissipation device is installed on the real-time rendering device;
[0050] A storage device, wherein the storage device includes a long-term storage device and a short-term storage device, the long-term storage device is a database, the database includes a relational database and a vector database, and the short-term storage device is a memory;
[0051] A processing device is used to process the interactive data between the virtual digital person and the user. The processing device is installed on the back of the display device. The display device, the multimodal input device, the real-time rendering device and the storage device are respectively connected to the processing device.
[0052] Furthermore, the database includes an action library and an expression library, each action record in the action library includes an action tag, and each expression record in the expression library includes an expression tag, and the action tag and the expression tag are respectively called by the virtual digital person when replying to the user;
[0053] It also includes a tool library interface, which includes at least a printer interface, a search interface, an encyclopedia interface, a date and time interface, and an execution code interface.
[0054] Furthermore, the present invention also proposes a computer-readable storage medium having program instructions stored thereon, and when the program instructions are executed by a processor, the interaction method of the virtual digital human is implemented.
[0055] According to some embodiments of the present invention, the beneficial effects of the present invention are as follows:
[0056] The virtual digital human interaction method of the present invention can realize real-time natural language communication between the virtual human and the user in a specific environment, and achieve the technical effect of enhancing the integration of the virtual digital human with the real scene by coordinating the virtual digital human's facial expressions, movements and professional attire. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 Flowchart of the virtual digital human interaction method in an embodiment of the present invention.
[0058] Figure 2 This is a flow chart of the voice wake-up process and the visual wake-up process in the interaction method of the virtual digital human in an embodiment of the present invention.
[0059] Figure 3 This is a flow chart of the interaction between a virtual digital human and a user based on a large language prompt word engineering in the interaction method of a virtual digital human in an embodiment of the present invention.
[0060] Figure 4 Schematic diagram of the architecture of the virtual human AI brain (AI agent) in an embodiment of the present invention.
[0061] Figure 5 Schematic diagram of the overall software architecture in an embodiment of the present invention.
[0062] Figure 6 Schematic diagram of the visual awakening process in an embodiment of the present invention.
[0063] Figure 7 Schematic diagram of the voice wake-up process in an embodiment of the present invention.
[0064] Figure 8 This is a schematic diagram of an AI multimodal interactive 3D virtual human large-screen all-in-one machine in an embodiment of the present invention.
[0065] Figure 9 This is a schematic diagram of an AI multimodal interactive 3D virtual human transparent screen all-in-one machine in an embodiment of the present invention.
[0066] Figure 10 This is a flow chart of voiceprint recognition in an embodiment of the present invention.
[0067] Figure 11 This is a flow chart of performing data modeling on sound in an embodiment of the present invention.
[0068] Figure 12 Schematic diagram of a speech image extraction network in an embodiment of the present invention.
[0069] Figure 13 Schematic diagram of the large language model network structure in an embodiment of the present invention.
[0070] Figure 14 Flowchart of the training process of the pre-training model in an embodiment of the present invention.
[0071] Figure 15 Schematic diagram of the basic structure of the speech animation synthesis model based on the BlendShapes method in an embodiment of the present invention.
[0072] Figure 16 Schematic diagram of the structure of the lip movement detection module in an embodiment of the present invention. DETAILED DESCRIPTION
[0073] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0074] The following will provide a clear and complete description of the concept, specific structure and technical effects of the present invention in conjunction with the embodiments and drawings to fully understand the purpose, scheme and effects of the present invention.
[0075] It should be noted that, unless otherwise specified, when a feature is referred to as being "fixed" or "connected" to another feature, it may be directly fixed or connected to the other feature, or it may be indirectly fixed or connected to the other feature. The singular forms "a", "said" and "the" used herein are also intended to include the plural forms, unless the context clearly indicates otherwise. In addition, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art. The terms used in this specification are only for describing specific embodiments and are not intended to limit the invention. The term "and / or" used herein includes any combination of one or more related listed items.
[0076] It should be understood that although the terms first, second, third, etc. may be used to describe various elements in the present disclosure, these elements should not be limited to these terms. These terms are only used to distinguish elements of the same type from each other. For example, without departing from the scope of the present disclosure, the first element may also be referred to as the second element, and similarly, the second element may also be referred to as the first element. The use of any and all examples or exemplary language ("for example", "such as", etc.) provided herein is only intended to better illustrate embodiments of the present invention, and unless otherwise required, will not impose limitations on the scope of the present invention. In addition, the industry term "posture" used herein refers to the position and attitude of a certain element relative to a spatial coordinate system.
[0077] Reference Figures 1 to 16 The embodiment of the present invention provides a method, device and storage medium for interacting with a virtual digital human, characterized in that, referring to Figure 1 , the method comprising:
[0078] S100: Determine whether a wake-up signal is received, causing the virtual digital human to enter a wake-up process, switching from a standby state to a working state, wherein the wake-up process includes a voice wake-up process and a visual wake-up process;
[0079] S200, based on the multimodal sensor of the virtual digital human, obtains environmental information and user information;
[0080] S300, based on the large language prompt word project, the virtual digital human interacts with the user;
[0081] S400: Adjust the character settings of the virtual digital human based on the prompt word engineering of the large language model (LLM).
[0082] According to some embodiments of the present invention, the beneficial effects of the present invention are as follows:
[0083] The virtual digital human interaction method of the present invention can realize real-time natural language communication between the virtual human and the user in a specific environment, and achieve the technical effect of enhancing the integration of the virtual digital human with the real scene by coordinating the virtual digital human's facial expressions, movements and professional attire.
[0084] In existing technology, the IPization of 3D virtual humans using real-life motion capture has become increasingly attractive and vivid. However, this type of 3D virtual human IP relies on the personal abilities of the embodying character. Bilibili's A-soul has experienced incidents such as the embodying character's escape and exposure. Therefore, the over-reliance of 3D virtual human IP on the embodying character's abilities has hampered the industrialization and commercialization of 3D virtual humans. If AI could replace the embodying character as the embodying character, this would undoubtedly solve the industrialization and commercialization issues of 3D virtual humans. Large language models (such as Wenxin Yiyan, ChatGPT, and Bard) as question-and-answer bots can only output content in text and voice. This makes it difficult for users to experience the warmth of AI-powered services. We believe that the next generation of interactive forms will inevitably be based primarily on natural language. Therefore, giving large language models an embodying character is the inevitable trend.
[0085] To realize the above ideas, we still need to solve the following three problems. These are product-level issues, including the 3D virtual human's reliance on the human within, how to make people's interactions with the large language model feel warm, which physical scenarios the product combining the 3D virtual human and the large language model should serve, what problems the 3D virtual human and the large language model product can solve, and how to improve the 3D virtual human's expressiveness so that we feel that the virtual human has warmth; they also include software technology issues: in the absence of a human within, the virtual human's movement and expression driving issues, the 3D virtual human's role setting, personality issues, how the 3D virtual human obtains external information, namely "ears," "eyes," "spatial perception," etc., the 3D virtual human's need for memory, the 3D virtual human's ability to use tools independently, where tools generally refer to software tools, the 3D virtual human's ability to break down complex tasks into subtasks and execute them step by step; they also include hardware-level issues: on what kind of medium is the virtual human's image carried, and how this image looks life-size, the need for one or more sensors to allow the virtual human to obtain external information, the real-time rendering of the 3D virtual human, and a sufficiently sophisticated 3D virtual human with a high geometric face count and large texture size. The hardware's 3D rendering capabilities are demanding, and high-performance rendering capabilities will inevitably lead to high heat generation and heat dissipation issues. The fans added to solve the heat dissipation problem will lead to acoustic deconstruction problems, sound input problems in noisy environments, and problems distinguishing between voices and speakers in multi-person environments.
[0086] To this end, we propose the following solutions:
[0087] (1) Solutions to product-level problems
[0088] In the past, 3D virtual humans relied on the human inside them as their soul. The human inside them was the soul, and without the human inside, the virtual human could do nothing. Our 3D virtual humans use a large language model to replace the human inside them as the soul. This also solves the problem of virtual humans being able to serve multiple scenarios. The all-in-one virtual machine (AI multimodal interactive 3D virtual human large screen) solves the traditional interaction solution where the large language model can only output sound and text. It makes the large language model the soul of the virtual human, and the virtual human becomes the beautiful skin of the large language model. The all-in-one virtual machine (AI multimodal interactive 3D virtual human large screen) adds an image to the output of the large language model, making the service more heartwarming. Most importantly, the all-in-one virtual machine (AI multimodal interactive 3D virtual human large screen) can replace human labor to independently complete some tasks, and its elegant appearance can bring real traffic to merchants.
[0089] (2) Solutions to software technology issues
[0090] Reference Figure 4 and Figure 5 To make the large language model the soul of the virtual human, the problem of driving the virtual human needs to be solved. We developed our own Audio2Animation algorithm, which consists of two parts: Audio2Face, which converts speech into expressions, and Audio2Action, which converts speech into actions. The virtual machine uses the large language model's Prompt Engineering to pre-set several character settings and supplements them with 3D clothing assets that complement their professions, giving the virtual human professional attributes and a more professional appearance. Multimodal sensors allow the virtual human to understand the external environment and obtain a variety of information. Depth sensors reconstruct the environment in three dimensions, enabling the virtual human to understand the environment. RGB cameras perform visual segmentation of the environment and perform object recognition, facial recognition, human skeleton recognition, and posture and gesture recognition on the segmented images. A localized vector database and relational database provide the virtual human with long-term memory. A comprehensive tool library (including at least a printer interface, search interface, encyclopedia interface, time and date interface, and code execution interface) is provided to the virtual human, allowing it to select appropriate tools using the large language model. The Huanzhen all-in-one machine uses prompt engineering, a large language model, to empower virtual humans with the ability to solve complex problems. This means the AI virtual human can automatically break down complex task commands into simpler steps and execute them sequentially. It can also choose tools or plug-ins during execution.
[0091] (3) Hardware-level problem solutions
[0092] An 86-inch vertical screen and an 86-inch transparent screen were chosen to render the virtual human life-size, creating a 3D effect with naked-eye glasses. This virtual human possesses "eyes," "ears," and spatial perception, allowing it to acquire external information. A circular 7-microphone array and a 4K HD RGB camera with HDR were selected, along with a ToF depth sensor. To integrate these three sensors, the Azure Kinect DK, a depth vision sensor used in industrial robots, was chosen. For high-quality real-time rendering of the 3D human, NVIDIA's high-performance RTX3080 graphics card with ray tracing was selected. Due to the heat dissipation issues associated with the high-performance graphics card and the chassis noise and vibration caused by the added fan, we opted for a split, modular design for the all-in-one hardware. This not only addresses heat dissipation and sound pickup issues, but also reduces maintenance costs and transportation. To address the issue of vibration caused by the fan affecting the microphone array, we integrated the microphones and various sensors into a single box, externally mounted on the top of the chassis, effectively addressing chassis vibration. Our proprietary acoustic noise reduction and voice enhancement algorithms address voice input issues in noisy environments. We also combine a narrow-beam algorithm with facial recognition for directional sound reception, enhancing audio quality. In situations where multiple speakers are speaking, we combine the facial recognition algorithm with a lip movement recognition module to accurately and directionally pick up sound from multiple speakers. We then use our stored voiceprint library to isolate the target speaker's voice from the multi-speaker conversation.
[0093] In some specific embodiments, the lip movement recognition module includes a target detection module, a lip movement detection module and an interaction module connected in sequence.
[0094] The target detection module needs to accurately locate the lips of a person's face and segment the image using the predicted coordinates to ensure that the lips are located in the center of the image. The project implementation algorithm is the Yolov5 algorithm, trained with a minimal pre-trained model. When preparing the target detection dataset, it is necessary to ensure data integrity, and the annotated lips should be located in the center of the image to achieve the effect of lip positioning, which can greatly speed up the fitting speed of the subsequent classification network.
[0095] The lip movement detection module uses a composite network of 3DResNet and GRU. Data processed by the Yolov5 algorithm is passed into the network. The residual structure extracts features, and the GRU ensures the transmission and preservation of temporal information. The prediction result is then obtained through softmax. In the lip movement detection module, the amount of training data is particularly critical, so data augmentation is used in preprocessing to provide the network with sufficient data training. Secondly, the structure of the network is also very important. The residual network ResNet is composed of a stack of multiple layers of networks to solve the gradient vanishing problem caused by the depth of the network. The deeper the network, the more image information features extracted, the more effective it is. The variant of the recurrent neural network RNN, GRU, uses a gate control mechanism to well preserve the temporal information, allowing the neural network to pay more attention to the dynamic change information of the lips in the time series.
[0096] Specifically, in some specific embodiments, referring to Figure 16 The lip movement detection module is a GRU layer embedded in ResNet, which also forms a composite network of CNN and RNN. However, the network does not simply use the 3D convolution and pooling structure because in the feature extraction process, CNN tends to extract pixel features and GRU to extract features that change dynamically over time.
[0097] The prediction results are finally passed to the interaction module, which relies on HTML and JS to complete the interaction between the front-end and back-end, thus realizing the recognition function of the system. In the interaction module, the framework used is the Flask framework. The advantage of this framework is that it is lightweight and flexible. In the login function, HTML directly submits the form to the back-end. As long as the back-end verifies the form through the database, the login function can be completed. In the recognition function, JS is used in conjunction with HTML to read the local video, and then it is submitted to the back-end for processing. The processing process first cuts the video frame, then the Yolov5 model predicts the lip coordinates of the frame image, and then cuts the image and saves it. Finally, the prediction result is obtained through the classification network. The recognized words are displayed on the front-end page to complete the entire function process.
[0098] Specifically, in step S400, the prompt word engineering gives the virtual person a character role setting, which is not only used for changing clothes, but also allows the virtual person to play a professional role, such as a doctor, lawyer, elementary school English teacher, etc.
[0099] Further, refer to Figure 2 and Figure 7 In step S100, the voice wake-up process includes the following steps:
[0100] S110. In the standby state, collect ambient sound and detect whether there is a human voice wake-up signal, where the human voice wake-up signal includes one or more combinations of human voice, wake-up keywords, and voiceprint features;
[0101] S120: If the detected human voice wake-up signal reaches a preset threshold, the virtual digital human is woken up, and the virtual digital human is switched from a standby state to a working state.
[0102] Furthermore, in step S110, detecting whether there is a human voice wake-up signal includes the following steps:
[0103] S111, performing noise reduction preprocessing on the received ambient sound, wherein the noise reduction preprocessing includes framing, windowing, pre-emphasis and adaptive endpoint detection (VAD) on the sound signal;
[0104] S112. Extract voice features based on MDCC and input the voice features into the voiceprint recognition model;
[0105] S113: Detect whether there is a human voice wake-up signal, and perform voiceprint matching between the voice features and the existing voiceprints in the voiceprint matching library;
[0106] S114: If the voiceprint is matched successfully, retrieve the user information that matches the voiceprint; if the voiceprint is matched unsuccessfully, create a new user storage space.
[0107] In some specific embodiments, in step S114, the user storage space is set in Amazon's private storage repository.
[0108] In some specific embodiments, the voiceprint recognition process is to first extract voice features, then put the features into the model for training, and finally find the result with the highest or closest score. Figure 10 Voiceprint recognition includes a training phase and a testing phase. The training phase includes inputting training voice, feature extraction, model training, and voiceprint library search or generating a new target user registration voice; the testing phase includes inputting test voice, feature extraction, voiceprint matching scoring, and distinguishing target users from non-target users. For virtual digital humans, target users can be defined as users who have interacted with the virtual digital humans and are stored in the memory, and non-target users can be defined as users who are interacting with the virtual digital humans for the first time. The virtual digital humans need to create a new storage space to store user information.
[0109] In some specific embodiments, referring to Figure 11 ,Flowchart of the virtual digital human data modeling of sound, the virtual digital human first receives the sound, pre-processes the sound, performs MFCC and voiceprint recognition, and then communicates with the user or creates a new user based on the recognition results.
[0110] Furthermore, in step S112,
[0111] The speech features include linear prediction coefficients LPCC and Mel cepstral coefficients MFCC.
[0112] Specifically, speech contains many characteristics. Feature extraction involves converting the continuous time-domain speech signal into a discrete frequency-domain signal, analyzing the energy of the speech signal, and extracting the necessary speech features from the continuous time-domain signal to represent the speech. There are many types of speech features, including Linear Prediction Coefficients (LPCCs) and Mel-Frequency Cepstral Coefficients (MFCCs) (Girin-Laurent).
[0113] The cochlea, a human organ, helps humans detect speech signals amidst noise. It has been discovered that the cochlea plays a role in filtering out noise when humans receive speech signals. Therefore, the idea is to leverage the human method of sound discrimination to construct a filter that replaces the cochlear function and achieves this goal. The human cochlea is similar to a combination of multiple filters. It is very sensitive to low-frequency signal energy. Its sensitivity is linear for frequencies below 1000 Hz and logarithmic for frequencies above 1000 Hz. This irregular frequency segmentation is a key feature that distinguishes the Mel cepstrum of speech signals from the standard cepstrum. The Mel cepstrum is one of the most common speech features in speech signal processing. Research on the biological function of the human ear has revealed that the human ear has varying sensitivity to speech signals of varying amplitudes and frequencies. The human ear has a limited range of amplitude and frequency, and cannot hear speech signals with amplitudes and frequencies exceeding its perceptibility. For example, ultrasound has a frequency of approximately 20,000 Hz, and speech signals below 20 Hz are also inaudible to the human ear. The frequency range of speech signals that can be heard by the human ear is within 20Hz to 20,000Hz. Noise with a frequency between 200Hz and 5,000Hz has the greatest impact on speech signals, making the text content within the speech signal unclear.
[0114] Furthermore, in step S113, the human voice wake-up signal includes a wake-up word, and the wake-up word is stored in the common area of the voiceprint recognition model.
[0115] In some embodiments, the virtual digital human voice monitoring software collects environmental sounds and detects whether there is a human voice wake-up signal. If the detected human voice wake-up signal reaches a threshold, the virtual digital human is awakened and enters the working state; after entering the working state, the virtual digital human outputs a response signal to the user and collects the user's voice signal through a microphone; the virtual digital human converts the collected user voice signal into text through the voice-to-text service, inputs it into the large language model service, the large language model service outputs text, and outputs voice through the text-to-speech service; the virtual familiar human outputs expressions and actions based on the animation control program.
[0116] Further, refer to Figure 2 and Figure 6 In step S100, the visual awakening process includes the following steps:
[0117] S130, in standby mode, controlling the visual sensor of the virtual digital human to collect images;
[0118] S140, detecting human skeleton recognition and / or face recognition to detect whether someone is approaching;
[0119] S150: If it is detected that someone is approaching, wake up the virtual digital human and switch the virtual digital human from the standby state to the working state.
[0120] In some specific embodiments, the virtual digital human's visual sensor collects images and detects whether someone is approaching through human skeleton recognition and face recognition; if someone is detected approaching, the virtual digital human requests the AI brain to output language text through the large language model service; based on the text-to-speech service, the language text output by the large language model service is converted into speech; based on the virtual human animation control program, sounds, expressions and actions are output; if the microphone receives the user's response, feedback is output for the received content; voice information of the user's response and questions is received through the microphone; based on the speech-to-text service, the voice information is converted into text; based on the large language model service, text information is input and the response text information is output; based on the text-to-speech service, the response text information is converted into virtual human voice information; based on the virtual human animation control program, expressions and actions are output.
[0121] In some specific embodiments, as an alternative to the wake-up algorithm, the state of the face and lips of the speaking object is captured visually, which specifically includes the following steps:
[0122] Lock the current interacting person by face or skeleton;
[0123] Capture the lip status of the current interacting person through computer vision algorithms;
[0124] When the lips are open or closed, the microphone is turned on for directional sound reception. When the lips are closed, the microphone is turned off for directional sound reception. This solves the problem of misreceiving sounds in noisy environments or when multiple people are talking.
[0125] Through the lip reading recognition algorithm, the continuously recorded lip movement video is converted into lip reading to enhance the microphone's voice recognition.
[0126] Furthermore, in step S100, after starting the virtual digital human, the standby state can be skipped and the working state can be directly entered.
[0127] Specifically, after the virtual digital human is turned on, it can directly play animations and voices on the display screen without detecting whether the user is approaching to attract the user's attention.
[0128] Furthermore, in step S200,
[0129] The multimodal sensor of the virtual digital human includes at least a depth sensor, an RGB camera and a microphone. The multimodal sensor is used at least to perform face recognition, skeleton recognition, posture and gesture recognition on the character.
[0130] Further, refer to Figure 3 、 Figure 6 and Figure 7 , said step S300 includes the following steps:
[0131] S310: After entering the working state, the virtual digital human receives the user's image and voice through the multimodal sensor, and transmits the voice and image to the voice and image extraction network to extract the voice and image;
[0132] S320, converting the speech into text based on the speech-to-text service;
[0133] S330: Outputting language text based on the large language model service;
[0134] S340: Based on the text-to-speech service, convert the language text output by the large language model service into speech;
[0135] S350, outputting sounds, expressions and movements based on the virtual human animation control program;
[0136] S360: If the multimodal sensor receives a response from the user, repeat steps S310 to S350 to provide feedback on the received content by outputting voice, expression, and action.
[0137] Further, refer to Figure 12 In step S310, the speech image extraction network includes: a speech separation network and a visual separation network.
[0138] The speech separation network includes a mixed speech receiving module, an STFT speech time-frequency conversion module, a speech feature extraction module, a speech upsampling module and an STFT frequency-time conversion module connected in sequence.
[0139] The visual separation network includes a hybrid visual receiving module, a VGGFace feature extraction module, and a visual feature network module connected in sequence.
[0140] It also includes a fusion module, the input of the fusion module is connected to the speech feature extraction module and the VGG feature extraction module respectively, and the output of the fusion module is connected to the speech upsampling module.
[0141] Specifically, refer to Figure 12 The network structure of speech image extraction mainly consists of two parts: speech separation network and visual separation network. The speech separation network receives the mixed speech spectrogram and outputs the spectrum mask. The network model structure adopts a codec structure similar to U-Net, which are respectively called feature extraction module and upsampling module. There is a jump connection between the feature extraction module and the upsampling module. Its main function is to use the shallow features of the feature extraction module to promote the upsampling module to predict the high-resolution time-frequency mask. However, the performance of feature extraction modules of different depths on different tasks is very different. In order to enable the network to adaptively learn network features of different depths, residual connections are added to the feature extraction module. The visual separation network receives a set of faces that completely corresponds to the speech spectrogram and outputs visual features that can be distributed in the same way as the speech separation network. Residual connections are added to the U-Net model and an audio and video feature fusion module is included.
[0142] Further, refer to Figure 13 In step S330, the large language model includes a corpus system module, a pre-training model and a fine-tuning module connected in sequence.
[0143] The corpus system module includes pre-trained corpus and fine-tuned corpus. The pre-trained corpus includes text data collected from books, magazines, and encyclopedias. The fine-tuned corpus includes annotated text data crawled from open source code libraries, annotated by experts, and processed through user dialogue.
[0144] The pre-trained model is first trained extensively on large-scale training data using an unsupervised learning method to obtain a general language model with strong generalization capabilities. The training method of the pre-trained model includes at least word vector embedding, contrastive pre-training, and contextual learning.
[0145] The fine-tuning module includes freezing the bottom layer of the pre-trained model and adjusting the weights of the upper layer, the bottom layer includes word vectors, and the upper layer includes classifiers.
[0146] Specifically, the corpus system includes pre-training data and fine-tuning data, the latter of which includes code and dialogue fine-tuning data. The pre-training data includes massive amounts of text data collected from books, magazines, encyclopedias, and other sources, enabling the model to learn the logical relationships of language. The fine-tuning data includes high-quality annotated text data processed through crawling open source code libraries, expert annotation, and user dialogue, further enhancing its conversational capabilities.
[0147] In some embodiments, pre-training is the foundation for building large-scale language models. This involves conducting extensive, generalized training on large-scale training data, using unsupervised learning methods, to develop a generalizable and highly generalizable language model. Through pre-training on this large-scale data, the model initially acquires the capabilities of human language understanding and contextual learning, capturing the semantic similarities between text and code snippets, thereby generating more accurate text and code vectors and supporting subsequent fine-tuning tasks.
[0148] In some embodiments, fine-tuning is the guarantee for the practical application of the model. It refers to further training the pre-trained model on a dataset for a specific task, which usually includes freezing the underlying layers of the pre-trained model (such as word vectors) and adjusting the weights of the upper layers (such as classifiers). Fine-tuning the pre-trained model will greatly shorten the training time, save computing resources and speed up the training convergence. Based on the pre-trained model with strong generalization ability, by integrating code data-based training and instruction-based fine-tuning, fine-tuning is performed using a specific dataset to make it have a stronger question-and-answer dialogue text generation capability.
[0149] Further, refer to Figure 14 The training steps of the pre-trained model include: collecting data sets and performing supervised fine-tuning; collecting comparison data and training the scoring model and using reinforcement learning for the reward model to optimize the cloud training model.
[0150] Further, refer to Figure 14 , the collecting a dataset and performing supervised fine-tuning includes sampling a prompt from the dataset; marking a preferred output answer and performing supervised fine-tuning on the pre-trained model based on the collected data.
[0151] Further, refer to Figure 14 The collecting of comparison data and training of the scoring model includes sampling a prompt and a number of corresponding model outputs; scoring and sorting the outputs and training the scoring model on the scored and sorted data.
[0152] Further, refer to Figure 14 , the use of reinforcement learning for a reward model to optimize a cloud training model includes resampling a prompt; the reward model scores the output and optimizes model parameters.
[0153] Furthermore, the step S300 further includes receiving a task instruction, breaking the task instruction into simple steps and executing them in sequence.
[0154] Furthermore, in step S400, the virtual digital human is controlled to change clothes based on the character settings preset in the 3D clothing asset library, and the 3D clothing asset library at least includes hospital intelligent guidance virtual digital human clothing, museum explanation virtual digital human clothing, and restaurant reception virtual digital human clothing.
[0155] Furthermore, in step S350, the virtual human animation control program includes a voice-to-expression service Audio2FaceService and a voice-to-action service Audio2ActionService. The voice-to-expression service Audio2FaceService is used to convert the virtual human's voice information into the virtual human's facial expressions, and the voice-to-action service Audio2ActionService is used to convert the virtual human's voice information into the virtual human's body movements.
[0156] Further, refer to Figure 15 The audio-to-expression service Audio2FaceService is a voice animation synthesis model based on the BlendShapes method. The voice animation synthesis model based on the BlendShapes method includes a 3D face control parameter prediction module based on different voice emotions, an expression base construction module based on sample expressions, and a 3D face animation synthesis module.
[0157] The three-dimensional face control parameter prediction module based on different voice emotions includes a voice preprocessing module, a voice feature extraction module, an audio and video mapping module and a three-dimensional face control parameter module connected in sequence;
[0158] The expression base construction module based on sample expressions includes a sample expression model library, an expression feature extraction module, an expression migration module, and an expression base construction module connected in sequence;
[0159] The three-dimensional facial animation synthesis module includes a three-dimensional facial control parameter smoothing module and a three-dimensional facial animation synthesis module connected in sequence. The input of the three-dimensional facial control parameter smoothing module is connected to the output of the three-dimensional facial control parameter module, and the input of the three-dimensional facial animation synthesis module is connected to the output of the expression base construction module.
[0160] Further, refer to Figure 8 and Figure 9 The present invention also proposes an interactive device for a virtual digital human, characterized by comprising:
[0161] A display device, wherein the display device is a vertical large screen;
[0162] A multimodal input device, comprising a 4K high-definition RGB camera, a Tof depth camera, and a microphone array, and disposed above the display device;
[0163] A real-time rendering device, used to render and output the movements and expressions of the virtual digital human in real time;
[0164] A heat dissipation device, the heat dissipation device is used to dissipate heat for the real-time rendering device, and the heat dissipation device is installed on the real-time rendering device;
[0165] A storage device, wherein the storage device includes a long-term storage device and a short-term storage device, the long-term storage device is a database, the database includes a relational database and a vector database, and the short-term storage device is a memory;
[0166] A processing device is used to process the interactive data between the virtual digital person and the user. The processing device is installed on the back of the display device. The display device, the multimodal input device, the real-time rendering device and the storage device are respectively connected to the processing device.
[0167] Furthermore, the database includes an action library and an expression library, each action record in the action library includes an action tag, and each expression record in the expression library includes an expression tag. The action tag and the expression tag are respectively called by the virtual digital person when replying to the user.
[0168] Furthermore, it also includes a tool library interface, which includes at least a printer interface, a search interface, an encyclopedia interface, a date and time interface, and an execution code interface.
[0169] Furthermore, the processing device includes an AI algorithm service program, an interactive backend program, a background web program and a virtual human rendering program.
[0170] Furthermore, the AI algorithm service program is a Python program, and the AI algorithm service program includes Audio2Face algorithm service, Text2Action algorithm service, large language model distribution service and Al Agent service.
[0171] Furthermore, the interactive backend program is a Java program, and the interactive backend program includes an account management system, a memory configuration system, a 3D asset management system and a large model interactive management system, and the interactive backend program is connected to the database.
[0172] Furthermore, the backend web program is an H5 program, and the backend web program includes an account management system page, a content configuration system page, a 3D asset management system page and a large model interaction management page.
[0173] Furthermore, the virtual human rendering program includes a UE5 program, and the virtual human rendering program includes virtual human rendering, action animation logic and external interaction program.
[0174] Furthermore, the display device is an LED screen, a naked-eye 3D screen, or a holographic screen.
[0175] Furthermore, the present invention also proposes a computer-readable storage medium having program instructions stored thereon, and when the program instructions are executed by a processor, the interaction method of the virtual digital human is implemented.
[0176] It should be appreciated that the method steps in the embodiments of the present invention can be implemented or executed by computer hardware, a combination of hardware and software, or by computer instructions stored in a non-transitory computer-readable memory. The method can use standard programming techniques. Each program can be implemented in a high-level procedural or object-oriented programming language to communicate with the computer system. However, if desired, the program can be implemented in assembly or machine language. In any case, the language can be a compiled or interpreted language. In addition, for this purpose, the program can be run on a programmed application-specific integrated circuit.
[0177] Furthermore, the operations of the processes described herein may be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by the context. The processes described herein (or variations and / or combinations thereof) may be performed under the control of one or more computer systems configured with executable instructions and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that is executed collectively on one or more processors, by hardware, or a combination thereof. The computer program includes a plurality of instructions that can be executed by one or more processors.
[0178] Further, the methods can be implemented in any type of computing platform that is operably connected to a suitable computer, including but not limited to a personal computer, a minicomputer, a mainframe, a workstation, a network or distributed computing environment, a separate or integrated computer platform, or in communication with a charged particle tool or other imaging device, etc. Various aspects of the present invention can be implemented as machine-readable code stored on a non-transitory storage medium or device, whether removable or integrated into a computing platform, such as a hard disk, an optical read and / or write storage medium, RAM, ROM, etc., so that it can be read by a programmable computer, and when the storage medium or device is read by the computer, it can be used to configure and operate the computer to perform the processes described herein. In addition, the machine-readable code, or portions thereof, can be transmitted over a wired or wireless network. When such media includes instructions or programs that implement the steps described above in conjunction with a microprocessor or other data processor, the invention described herein includes these and other different types of non-transitory computer-readable storage media. When programmed according to the methods and techniques of the present invention, the present invention can also include the computer itself.
[0179] The computer program can be applied to input data to perform the functions described herein, thereby converting the input data to generate output data that is stored in a non-volatile memory. The output information can also be applied to one or more output devices such as a display. In a preferred embodiment of the present invention, the converted data represents a physical and tangible object, including a specific visual depiction of the physical and tangible object produced on the display.
[0180] The above description is merely a preferred embodiment of the present invention. The present invention is not limited to the aforementioned embodiments. As long as the technical effects of the present invention are achieved by the same means, any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention. Within the scope of protection of the present invention, various modifications and variations of the technical solutions and / or implementation methods are possible.
Claims
1. A virtual digital human interaction method based on a large language model, characterized in that: The method includes: S200, based on the multimodal sensor of the virtual digital human, obtains environmental information and user information; S300, based on the large language prompt word project, the virtual digital human interacts with the user; The step S300 includes the following steps: S310: After entering the working state, the multimodal sensor receives the user's image and voice, and transmits the voice and image to the voice and image extraction network to extract the voice and image; In step S310, the speech image extraction network includes: a speech separation network and a visual separation network. The speech separation network includes a mixed speech receiving module, an STFT speech time-frequency conversion module, a speech feature extraction module, a speech upsampling module and an STFT frequency-time conversion module connected in sequence. The visual separation network includes a hybrid visual receiving module, a VGGFace feature extraction module, and a visual feature network module connected in sequence. It also includes a fusion module, wherein the input of the fusion module is connected to the speech feature extraction module and the VGG feature extraction module respectively, and the output of the fusion module is connected to the speech upsampling module; S320, converting the speech into text based on the speech-to-text service; S330: Outputting language text based on the large language model service; S340: Based on the text-to-speech service, convert the language text output by the large language model service into speech; S350, outputting sounds, expressions and movements based on the virtual human animation control program; S360: If a response is received from the user, repeat steps S310 to S350 to provide feedback on the received content, outputting voice, expression, and action; S400: Adjust the character settings of the virtual digital human based on the prompt word project of the large language model.
2. The virtual digital human interaction method based on a large language model according to claim 1 is characterized in that: In the step S200, The multimodal sensor of the virtual digital human includes at least a depth sensor, an RGB camera and a microphone. The multimodal sensor is used at least to perform face recognition, skeleton recognition, posture and gesture recognition on the character.
3. The virtual digital human interaction method based on a large language model according to claim 1, characterized in that: In step S330, the large language model includes a corpus system module, a pre-training model and a fine-tuning module connected in sequence. The corpus system module includes pre-trained corpus and fine-tuned corpus. The pre-trained corpus includes text data collected from books, magazines, and encyclopedias. The fine-tuned corpus includes annotated text data crawled from open source code libraries, annotated by experts, and processed through user dialogue. The pre-trained model is first trained extensively on large-scale training data using an unsupervised learning method to obtain a general language model with strong generalization capabilities. The training method of the pre-trained model includes at least word vector embedding, contrastive pre-training, and contextual learning. The fine-tuning module includes freezing the bottom layer of the pre-trained model and adjusting the weights of the upper layers, the bottom layer includes word vectors, and the upper layers include classifiers; The training steps of the pre-trained model include: collecting a data set and performing supervised fine-tuning; Collect comparative data and train the scoring model and use reinforcement learning for the reward model to optimize the cloud training model; The collecting of a dataset and performing supervised fine-tuning includes sampling a prompt from the dataset; marking a preferred output answer and performing supervised fine-tuning on a pre-trained model based on the collected data; The collecting comparison data and training the scoring model includes sampling a prompt and a number of corresponding model outputs; scoring and sorting the outputs and training the scoring model on the scored and sorted data; The method of using reinforcement learning for a reward model to optimize a cloud-trained model includes resampling a prompt; the reward model scores the output and optimizes model parameters.
4. The virtual digital human interaction method based on a large language model according to claim 1, characterized in that: The step S300 further includes receiving a task instruction, breaking the task instruction into simple steps and executing them in sequence.
5. The virtual digital human interaction method based on a large language model according to claim 1, characterized in that: In step S400, the virtual digital human is controlled to change clothes based on the character settings preset in the 3D clothing asset library, and the 3D clothing asset library at least includes hospital intelligent guide virtual digital human clothing, museum explanation virtual digital human clothing, and restaurant reception virtual digital human clothing.
6. The virtual digital human interaction method based on a large language model according to claim 1, characterized in that: In step S350, the virtual human animation control program includes a voice-to-expression service Audio2FaceService and a voice-to-action service Audio2ActionService. The voice-to-expression service Audio2FaceService is used to convert the virtual human's voice information into the virtual human's facial expression, and the voice-to-action service Audio2ActionService is used to convert the virtual human's voice information into the virtual human's body movements. The audio-to-expression service Audio2FaceService is a voice animation synthesis model based on the BlendShapes method, which includes a three-dimensional face control parameter prediction module based on different voice emotions, an expression base construction module based on sample expressions, and a three-dimensional face animation synthesis module. The three-dimensional face control parameter prediction module based on different voice emotions includes a voice preprocessing module, a voice feature extraction module, an audio and video mapping module and a three-dimensional face control parameter module connected in sequence; The expression base construction module based on sample expressions includes a sample expression model library, an expression feature extraction module, an expression migration module, and an expression base construction module connected in sequence; The three-dimensional facial animation synthesis module includes a three-dimensional facial control parameter smoothing module and a three-dimensional facial animation synthesis module connected in sequence. The input of the three-dimensional facial control parameter smoothing module is connected to the output of the three-dimensional facial control parameter module, and the input of the three-dimensional facial animation synthesis module is connected to the output of the expression base construction module.
7. A method for interacting with a virtual digital human, characterized in that: The method for waking up a virtual digital human according to any one of claims 1 to 5, wherein the method for interacting with the virtual digital human further comprises: S100, determining to receive a wake-up signal, causing the virtual digital human to enter a wake-up process, switching from a standby state to a working state, wherein the wake-up process includes a voice wake-up process; The voice wake-up process includes the following steps: S110. In standby mode, collect ambient sound and detect whether there is a human voice wake-up signal, where the human voice wake-up signal includes a human voice, a wake-up keyword, and a voiceprint feature; S120: If the detected human voice wake-up signal reaches a preset threshold, the virtual digital human is woken up, and the virtual digital human is switched from a standby state to a working state.
8. An interactive device for a virtual digital human, characterized in that: include: A display device, wherein the display device is a vertical large screen; A multimodal input device, comprising a 4K high-definition RGB camera, a Tof depth camera, and a microphone array, and disposed above the display device; A real-time rendering device, used to render and output the movements and expressions of the virtual digital human in real time; A heat dissipation device, the heat dissipation device is used to dissipate heat for the real-time rendering device, and the heat dissipation device is installed on the real-time rendering device; A storage device, wherein the storage device includes a long-term storage device and a short-term storage device, the long-term storage device is a database, the database includes a relational database and a vector database, and the short-term storage device is a memory; A processing device is used to process the interactive data between the virtual digital person and the user. The processing device is installed on the back of the display device. The display device, the multimodal input device, the real-time rendering device and the storage device are respectively connected to the processing device.
9. The interactive device for virtual digital human according to claim 8, characterized in that: The database includes an action library and an expression library, each action record in the action library includes an action tag, and each expression record in the expression library includes an expression tag, and the action tag and the expression tag are respectively called by the virtual digital person when replying to the user; It also includes a tool library interface, which includes at least a printer interface, a search interface, an encyclopedia interface, a date and time interface, and an execution code interface. 10 . A computer-readable storage medium having program instructions stored thereon, wherein the program instructions are executed by a processor to implement the method according to claim 1 .
Citation Information
Patent Citations
Virtual person-based multi-mode interactive processing method and system
CN107765852A
An interactive method and system based on a virtual human
CN109032328A
System and method for second-order cue word of large language model for digital human intellectualization
CN116881415A
Virtual digital human interaction system based on LLM language large model
CN117075732A
Digital interaction method and system based on artificial intelligence, and medium
CN117348736A
Cited By
Interaction scheme generation method and device, electronic equipment and storage medium
CN121834716A
Interaction scheme generation method and apparatus, electronic device, and storage medium
CN121834716B