Digital human real-time generation method and device based on interaction scene

Through the combination of large language models and multimodal sensors, natural language communication between virtual people and users is realized, solving the problem of people and communication restriction in 3D virtual people dependence, and improving the application effect of virtual people in real scenarios.

CN120491832APending Publication Date: 2025-08-15SHENZHEN MIRAGE FUTURE INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510883956.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-29
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In the prior art, 3D virtual human IP depends on people, resulting in instability of the system, and the large language model cannot realize natural language communication between people, limiting its application in on-site navigation environment.

Method used

The large language model is used as the soul of the virtual human, combined with multimodal sensors to obtain the environment and user information, switch to the working state through the voice and visual wake-up process, and use the storage device to compare user information to realize natural language communication between the virtual human and the user.

Benefits of technology

Real-time natural language communication between virtual people and users is realized, the integration effect between virtual digital people and real scenes is improved, the problems of industrialization and productization of 3D virtual people are solved, and the service temperature of artificial intelligence is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120491832A_ABST
    Figure CN120491832A_ABST
Patent Text Reader

Abstract

The invention relates to a digital human real-time generation method and device based on an interaction scene, and the method comprises the steps: obtaining environment sound and video, and detecting whether there is a wake-up signal; determining to receive a wake-up signal, enabling the virtual digital human to enter a wake-up process, and switching from a standby state to a working state; acquiring environment information and user information based on a multi-modal sensor; comparing the user information with the existing information of the storage device, if the user information is recorded in the storage device, calling the existing user information of the storage device, and if the user information is not recorded in the storage device, creating new user information in the storage device; and based on the large language model, the virtual human interaction all-in-one machine carries out conversation with the user. The virtual human interaction all-in-one machine can realize real-time natural language communication between a virtual human and a user in a specific environment, and achieves the technical effect of improving fusion of the virtual digital human and a real scene in cooperation with facial expressions, actions and career dress-up of the virtual digital human.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence applications, and in particular relates to a method and device for real-time generation of digital humans based on interactive scenes. Background Art

[0002] 3D virtual humans, using real-life motion capture, are becoming increasingly attractive and vivid. However, these 3D virtual human IPs rely on the individual abilities of the emulated human, making their roles unstable. In the event of a resident escaping, the system will inevitably struggle to recover quickly. Furthermore, existing large language models can only input and output text and cannot facilitate conversational communication, limiting their application in on-site navigation environments. Summary of the Invention

[0003] The present invention provides a method and device for real-time generation of digital humans based on an interactive scene, aiming to solve at least one of the technical problems existing in the prior art.

[0004] The technical solution of the present invention relates to a method for real-time generation of digital humans based on interactive scenes.

[0005] The method comprises the following steps:

[0006] S100: Acquire ambient sound and video to detect whether there is a wake-up signal;

[0007] S200: Determine whether a wake-up signal is received, causing the virtual digital human to enter a wake-up process, switching from a standby state to a working state, wherein the wake-up process includes a voice wake-up process and a visual wake-up process;

[0008] S300, acquiring environmental information and user information based on a multimodal sensor, wherein the user information includes user image information and user voice information;

[0009] S400, comparing the user information with existing information in the storage device. If the user information is already recorded in the storage device, retrieving the existing user information from the storage device. If the user information is not recorded in the storage device, creating new user information in the storage device.

[0010] S500, based on a large language model, the virtual human interactive all-in-one machine conducts conversations with users;

[0011] The step S500 includes:

[0012] S31. After entering the working state, the virtual digital human receives the user's image and voice through the multimodal sensor, and transmits the voice and image to the voice and image extraction network to extract the voice and image;

[0013] S32. Converting speech into text based on a speech-to-text service;

[0014] S33. Output language text based on the large language model service;

[0015] S34. Based on the text-to-speech service, convert the language text output by the large language model service into speech;

[0016] S35, outputting sounds, expressions and movements based on the virtual human animation control program;

[0017] S36: If the multimodal sensor 200 receives a response from the user, steps S31 to S35 are repeated to provide feedback on the received content in the form of voice, expression, and action.

[0018] Furthermore, in step S200,

[0019] The voice wake-up process includes the following steps:

[0020] S11. In standby mode, collect ambient sound and detect whether there is a human voice wake-up signal, where the human voice wake-up signal includes one or more combinations of human voice, wake-up keywords, and voiceprint features;

[0021] S12: If the detected human voice wake-up signal reaches a preset threshold, the virtual digital human is awakened, and the virtual digital human is switched from a standby state to a working state;

[0022] The visual awakening process includes the following steps:

[0023] S13, in standby mode, controlling the visual sensor of the virtual digital human to collect images;

[0024] S14, detecting human skeleton recognition and / or face recognition to detect whether someone is approaching;

[0025] S15. If it is detected that someone is approaching, the virtual digital human is awakened and switched from the standby state to the working state.

[0026] Furthermore, the present invention also proposes a virtual human interactive all-in-one machine, comprising:

[0027] A display device, wherein the display device is a vertical large screen;

[0028] a multimodal sensor, configured to detect the proximity of a user and acquire audio and video of the user, the multimodal sensor being disposed on the display device;

[0029] A storage device, wherein the storage device includes a long-term storage device and a short-term storage device, the long-term storage device is a database, the database includes a relational database and a vector database, and the short-term storage device is a memory;

[0030] A processing device is used to process interaction data between the virtual person and the user. The processing device is installed on the back of the display device. The display device, the multimodal sensor and the storage device are respectively connected to the processing device.

[0031] Furthermore, the display device is an LED screen, a naked-eye 3D screen, or a holographic screen.

[0032] Furthermore, the multimodal sensor includes one or more of a 4K high-definition RGB camera, a Tof depth camera, and a microphone array.

[0033] Furthermore, the database includes an action library and an expression library, each action record in the action library includes an action tag, and each expression record in the expression library includes an expression tag, and the action tag and the expression tag are respectively called by the virtual person when interacting with the user;

[0034] The database further includes a tool library interface, which includes at least a printer interface, a search interface, an encyclopedia interface, a date and time interface, and an execution code interface.

[0035] Furthermore, the processing device includes at least a real-time rendering device for rendering and outputting the virtual person's actions and expressions in real time, and the real-time processing device is connected to the processing device. The real-time processing device is installed on the display device or is independently installed with the processing device through an electrical connection.

[0036] Furthermore, it also includes a heat dissipation device, which is used to dissipate heat for the real-time rendering device and is installed on the real-time rendering device.

[0037] Furthermore, the processing device includes an AI algorithm service program, an interactive backend program, a backend web program, and a virtual human rendering program;

[0038] The AI algorithm service program is a Python program, and the AI algorithm service program includes Audio2Face algorithm service, Text2Action algorithm service, large language model distribution service and AI Agent service;

[0039] The interactive backend program is a Java program, and includes an account management system, a memory configuration system, a 3D asset management system, and a large model interactive management system. The interactive backend program is connected to the database;

[0040] The backend web program is an H5 program, which includes an account management system page, a content configuration system page, a 3D asset management system page, and a large model interaction management page;

[0041] The virtual human rendering program includes a UE5 program, and the virtual human rendering program includes virtual human rendering, action animation logic and external interaction program.

[0042] Furthermore, it also includes a shell, and the display device, the multimodal sensor, the storage device and the processing device are installed in the shell.

[0043] According to some embodiments of the present invention, the beneficial effects of the present invention are as follows:

[0044] The virtual human interactive all-in-one machine and its interactive method of the present invention can realize real-time natural language communication between the virtual human and the user in a specific environment, and achieve the technical effect of enhancing the integration of the virtual digital human with the real scene by coordinating the facial expressions, movements and professional attire of the virtual digital human. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 Schematic diagram of a virtual human interaction all-in-one machine in an embodiment of the present invention.

[0046] Figure 2 This is a schematic diagram of an AI multimodal interactive 3D virtual human large-screen all-in-one machine in an embodiment of the present invention.

[0047] Figure 3 This is a schematic diagram of an AI multimodal interactive 3D virtual human transparent screen all-in-one machine in an embodiment of the present invention.

[0048] Figure 4 Flowchart of the interactive method of the virtual human interactive all-in-one machine in an embodiment of the present invention.

[0049] Figure 5 Schematic diagram of the architecture of the virtual human AI brain (AI agent) in an embodiment of the present invention.

[0050] Figure 6 Schematic diagram of the overall software architecture in an embodiment of the present invention.

[0051] Figure 7 Schematic diagram of the visual awakening process in an embodiment of the present invention.

[0052] Figure 8 Schematic diagram of the voice wake-up process in an embodiment of the present invention.

[0053] Figure 9 This is a flow chart of voiceprint recognition in an embodiment of the present invention.

[0054] Figure 10 This is a flow chart of performing data modeling on sound in an embodiment of the present invention.

[0055] Figure 11 Schematic diagram of a speech image extraction network in an embodiment of the present invention.

[0056] Figure 12 Schematic diagram of the large language model network structure in an embodiment of the present invention.

[0057] Figure 13 Flowchart of the training process of the pre-training model in an embodiment of the present invention.

[0058] Figure 14 Schematic diagram of the basic structure of the speech animation synthesis model based on the BlendShapes method in an embodiment of the present invention.

[0059] Figure 15 Schematic diagram of the structure of the lip movement detection module in an embodiment of the present invention.

[0060] Reference numerals

[0061] 100. Display device; 200. Multimodal sensor; 300. Storage device; 400. Processing device; 410. Real-time rendering device. DETAILED DESCRIPTION

[0062] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0063] The following will provide a clear and complete description of the concept, specific structure and technical effects of the present invention in conjunction with the embodiments and drawings to fully understand the purpose, scheme and effects of the present invention.

[0064] It should be noted that, unless otherwise specified, when a feature is referred to as being "fixed" or "connected" to another feature, it may be directly fixed or connected to the other feature, or it may be indirectly fixed or connected to the other feature. The singular forms "a", "said" and "the" used herein are also intended to include the plural forms, unless the context clearly indicates otherwise. In addition, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art. The terms used in this specification are only for describing specific embodiments and are not intended to limit the invention. The term "and / or" used herein includes any combination of one or more related listed items.

[0065] It should be understood that although the terms first, second, third, etc. may be used to describe various elements in the present disclosure, these elements should not be limited to these terms. These terms are only used to distinguish elements of the same type from each other. For example, without departing from the scope of the present disclosure, the first element may also be referred to as the second element, and similarly, the second element may also be referred to as the first element. The use of any and all examples or exemplary language ("for example", "such as", etc.) provided herein is only intended to better illustrate embodiments of the present invention, and unless otherwise required, will not impose limitations on the scope of the present invention. In addition, the industry term "posture" used herein refers to the position and attitude of a certain element relative to a spatial coordinate system.

[0066] Reference Figures 1 to 15 The embodiment of the present invention provides a virtual human interactive all-in-one machine, referring to Figures 1 to 3 ,include:

[0067] Display device 100, the display device 100 is a vertical large screen;

[0068] a multimodal sensor 200 for detecting the proximity of a user and acquiring the user's audio and video, wherein the multimodal sensor 200 is disposed on the display device 100;

[0069] A storage device 300, wherein the storage device 300 includes a long-term storage device 300 and a short-term storage device 300, wherein the long-term storage device 300 is a database including a relational database and a vector database, and the short-term storage device 300 is a memory;

[0070] The processing device 400 is used to process the interaction data between the virtual person and the user. The processing device 400 is installed on the back of the display device 100. The display device 100, the multimodal sensor 200 and the storage device 300 are respectively connected to the processing device 400.

[0071] According to some embodiments of the present invention, the beneficial effects of the present invention are as follows:

[0072] The virtual human interactive all-in-one machine and its interactive method of the present invention can realize real-time natural language communication between the virtual human and the user in a specific environment, and achieve the technical effect of enhancing the integration of the virtual digital human with the real scene by coordinating the facial expressions, movements and professional attire of the virtual digital human.

[0073] Further, refer to Figures 1 to 3 , the display device 100 is an LED screen, a naked-eye 3D screen, or a holographic screen.

[0074] Further, refer to Figures 1 to 3The multimodal sensor 200 includes one or more of a 4K high-definition RGB camera, a Tof depth camera, and a microphone array.

[0075] Further, refer to Figures 1 to 3 The database includes an action library and an expression library. Each action record in the action library includes an action tag, and each expression record in the expression library includes an expression tag. The action tag and the expression tag are respectively called by the virtual person when interacting with the user.

[0076] Further, refer to Figures 1 to 3 The database also includes a tool library interface, which includes at least a printer interface, a search interface, an encyclopedia interface, a date and time interface, and an execution code interface.

[0077] Further, refer to Figures 1 to 3 The processing device 400 at least includes a real-time rendering device 410 for real-time rendering and outputting the actions and expressions of the virtual person. The real-time processing device 410 is connected to the processing device 400. The real-time processing device 410 is installed on the display device 100 or is independently installed through an electrical connection with the processing device 400.

[0078] Further, refer to Figures 1 to 3 , further comprising a heat dissipation device, which is used to dissipate heat for the real-time rendering device 410 , and the heat dissipation device is installed on the real-time rendering device 410 .

[0079] Further, refer to Figure 6 , the processing device 400 includes an AI algorithm service program, an interactive backend program, a backend web program and a virtual human rendering program;

[0080] The AI algorithm service program is a Python program, and the AI algorithm service program includes Audio2Face algorithm service, Text2Action algorithm service, large language model distribution service and AI Agent service;

[0081] The interactive backend program is a Java program, and includes an account management system, a memory configuration system, a 3D asset management system, and a large model interactive management system. The interactive backend program is connected to the database;

[0082] The backend web program is an H5 program, which includes an account management system page, a content configuration system page, a 3D asset management system page, and a large model interaction management page;

[0083] The virtual human rendering program includes a UE5 program, and the virtual human rendering program includes virtual human rendering, action animation logic and external interaction program.

[0084] Further, refer to Figures 1 to 3 , further comprising a shell, in which the display device 100, the multimodal sensor 200, the storage device 300 and the processing device 400 are installed.

[0085] Further, refer to Figure 4 The present invention also discloses an interactive method of a virtual human interactive all-in-one machine, comprising the following steps:

[0086] S100: Acquire ambient sound and video to detect whether there is a wake-up signal;

[0087] S200: Determine whether a wake-up signal is received, causing the virtual digital human to enter a wake-up process, switching from a standby state to a working state, wherein the wake-up process includes a voice wake-up process and a visual wake-up process;

[0088] S300, acquiring environmental information and user information based on the multimodal sensor 200, wherein the user information includes user image information and user voice information;

[0089] S400, comparing the user information with the existing information in the storage device 300. If the user information is already recorded in the storage device 300, retrieving the existing user information in the storage device 300; if the user information is not recorded in the storage device 300, creating new user information in the storage device 300;

[0090] S500, based on the large language model, the virtual human interactive all-in-one machine communicates with the user.

[0091] In existing technology, the IPization of 3D virtual humans using real-life motion capture has become increasingly attractive and vivid. However, this type of 3D virtual human IP relies on the personal abilities of the embodying figure. Bilibili's A-soul has experienced incidents such as the embodying figure's escape and exposure. Therefore, the over-reliance of 3D virtual human IP on the embodying figure's abilities has hampered the industrialization and commercialization of 3D virtual humans. If AI could replace the embodying figure as the embodying figure, this would undoubtedly solve the industrialization and commercialization issues of 3D virtual humans. II. As a question-and-answer robot, the large language model can only output content in text and voice. This hinders the user experience of the warmth of AI-powered services. We believe that the next generation of interactive forms will inevitably be based primarily on natural language. Therefore, giving the large language model an embodying figure is the inevitable trend.

[0092] To realize the above ideas, we still need to solve the following three problems. These are product-level issues, including the reliance of 3D virtual humans on the human within, how to make people's interactions with the large language model feel warm, which physical scenarios products combining 3D virtual humans and large language models should serve, what problems can 3D virtual humans and large language models solve, and how to improve the expressiveness of 3D virtual humans so that we feel that this virtual human has warmth. They also include software technology issues: in the absence of a human within, the movement and expression driving issues of virtual humans, the role setting and personality issues of 3D virtual humans, how 3D virtual humans obtain external information, namely "ears", "glasses", "spatial perception", etc., the need for 3D virtual humans to have memory, the ability of 3D virtual humans to use tools independently, where tools generally refer to software tools, the ability to break down complex tasks into subtasks and execute them step by step. They also include hardware-level issues: on what kind of medium should the image of the virtual human be carried, and should this image appear to be the same size as a real person, the need for one or more sensors for the virtual human to obtain external information, the real-time rendering of 3D virtual humans, and the high geometric count and large texture size of sufficiently refined 3D virtual humans. The hardware's 3D rendering capabilities are demanding, and high-performance rendering capabilities will inevitably lead to high heat generation and heat dissipation issues. The fans added to solve the heat dissipation problem will lead to acoustic deconstruction problems, sound input problems in noisy environments, and problems distinguishing between voices and speakers in multi-person environments.

[0093] To this end, we propose the following solutions:

[0094] (1) Solutions to product-level problems

[0095] In the past, 3D virtual humans relied on the human inside them as their soul. The human inside them was the soul, and without the human inside, the virtual human could do nothing. Our 3D virtual humans use a large language model to replace the human inside them as the soul. This also solves the problem of virtual humans being able to serve multiple scenarios. The all-in-one virtual machine (AI multimodal interactive 3D virtual human large screen) solves the traditional interaction solution where the large language model can only output sound and text. It makes the large language model the soul of the virtual human, and the virtual human becomes the beautiful skin of the large language model. The all-in-one virtual machine (AI multimodal interactive 3D virtual human large screen) adds an image to the output of the large language model, making the service more heartwarming. Most importantly, the all-in-one virtual machine (AI multimodal interactive 3D virtual human large screen) can replace human labor to independently complete some tasks, and its elegant appearance can bring real traffic to merchants.

[0096] (2) Solutions to software technology issues

[0097] Reference Figure 4 and Figure 5To make the large language model the soul of the virtual human, the problem of driving the virtual human needs to be solved. We developed our own Audio2Animation algorithm, which consists of two parts: Audio2Face, which converts speech into expressions, and Audio2Action, which converts speech into actions. The virtual machine uses the large language model's prompt engineering (PromptEngineering) to pre-set several character settings and supplements them with 3D clothing assets that complement their professions, giving the virtual human professional attributes and a more professional appearance. The multimodal sensor 200 allows the virtual human to understand the external environment and obtain a variety of information. The depth sensor reconstructs the environment in three dimensions, enabling the virtual human to understand the environment. The RGB camera performs visual segmentation of the environment and performs object recognition, facial recognition, human skeleton recognition, and posture and gesture recognition on the segmented images. A localized vector database and relational database provide the virtual human with long-term memory. The virtual human is provided with a comprehensive tool library (including at least a printer interface, a search interface, an encyclopedia interface, a time and date interface, and a code execution interface), allowing the virtual human to select appropriate tools using the large language model. The virtual reality machine uses prompt engineering, a large language model, to empower virtual humans with the ability to solve complex problems. This means the AI virtual human can automatically break down complex task commands into simpler steps and execute them sequentially. It can also choose tools or plug-ins during execution.

[0098] (3) Hardware-level problem solutions

[0099] An 86-inch vertical screen and an 86-inch transparent screen were chosen to render the virtual human life-size. The transparent screen gave the virtual human a naked-eye 3D feel. This gave the virtual human the ability to simultaneously "see," "hear," and "spatial perception" to acquire external information. A circular 7-microphone array and a 4K HD RGB camera with HDR were selected, along with a ToF depth sensor. To integrate these three sensors, the Azure Kinect DK, a depth vision sensor used in industrial robots, was chosen. For high-quality real-time rendering of the 3D human, NVIDIA's high-performance RTX3080 graphics card with ray tracing was selected. Due to the heat dissipation issues associated with the high-performance graphics card and the chassis noise and vibration caused by the added fan, we opted for a split, modular design for the all-in-one hardware. This not only addressed the heat dissipation and sound pickup issues, but also reduced maintenance costs and transportation. To address the issue of vibration caused by the fan affecting the microphone array, we integrated the microphones and various sensors into a single box, externally mounted on the top of the chassis, effectively addressing the chassis vibration. Our proprietary acoustic noise reduction and voice enhancement algorithms address voice input issues in noisy environments. We also combine a narrow beam algorithm with facial recognition for directional sound reception, enhancing sound reception. When multiple speakers are speaking, we combine facial recognition with lip movement recognition to accurately and directionally pick up sound from multiple speakers. By integrating this with a stored voiceprint database, we can isolate the target speaker's voiceprint from the multi-speaker conversation.

[0100] In some specific embodiments, the lip movement recognition module includes a target detection module, a lip movement detection module and an interaction module connected in sequence.

[0101] The target detection module needs to accurately locate the lips of a person's face and segment the image using the predicted coordinates to ensure that the lips are located in the center of the image. The project implementation algorithm is the Yolov5 algorithm, trained with a minimal pre-trained model. When preparing the target detection dataset, it is necessary to ensure data integrity, and the annotated lips should be located in the center of the image to achieve the effect of lip positioning, which can greatly speed up the fitting speed of the subsequent classification network.

[0102] The lip movement detection module uses a composite network of 3DResNet and GRU. Data processed by the Yolov5 algorithm is passed into the network. The residual structure extracts features, and the GRU ensures the transmission and preservation of temporal information. The prediction result is then obtained through softmax. In the lip movement detection module, the amount of training data is particularly critical, so data augmentation is used in preprocessing to provide the network with sufficient data training. Secondly, the structure of the network is also very important. The residual network ResNet is composed of a stack of multiple layers of networks to solve the gradient vanishing problem caused by the depth of the network. The deeper the network, the more image information features extracted, the more effective it is. The variant of the recurrent neural network RNN, GRU, uses a gate control mechanism to well preserve the temporal information, allowing the neural network to pay more attention to the dynamic change information of the lips in the time series.

[0103] Specifically, in some specific embodiments, referring to Figure 15 The lip movement detection module is a GRU layer embedded in ResNet, which also forms a composite network of CNN and RNN. However, the network does not simply use the 3D convolution and pooling structure because in the feature extraction process, CNN tends to extract pixel features and GRU to extract features that change dynamically over time.

[0104] The prediction results are finally passed to the interaction module, which relies on HTML and JS to complete the interaction between the front-end and back-end, thus realizing the recognition function of the system. In the interaction module, the framework used is the Flask framework. The advantage of this framework is that it is lightweight and flexible. In the login function, HTML directly submits the form to the back-end. As long as the back-end verifies the form through the database, the login function can be completed. In the recognition function, JS is used in conjunction with HTML to read the local video, and then it is submitted to the back-end for processing. The processing process first cuts the video frame, then the Yolov5 model predicts the lip coordinates of the frame image, and then cuts the image and saves it. Finally, the prediction result is obtained through the classification network. The recognized words are displayed on the front-end page to complete the entire function process.

[0105] In some specific embodiments, the interactive method of the virtual human interactive all-in-one machine includes the following steps:

[0106] S1. Determine that a wake-up signal is received, causing the virtual digital human to enter a wake-up process, switching from a standby state to a working state. The wake-up process includes a voice wake-up process and a visual wake-up process;

[0107] S2, acquiring environmental information and user information based on the multimodal sensor 200 of the virtual digital human;

[0108] S3. Based on the large language prompt word project, the virtual digital human interacts with the user;

[0109] S4. Based on the prompt word engineering of the large language model, adjust the character settings of the virtual digital human.

[0110] Specifically, in step S4, the prompt word engineering is used to set the character role of the virtual person, which is not only used to change clothes, but also to let the virtual person play a professional role, such as a doctor, a lawyer, an elementary school English teacher, etc.

[0111] Furthermore, in step S1, the voice wake-up process includes the following steps:

[0112] S11. In standby mode, collect ambient sound and detect whether there is a human voice wake-up signal, where the human voice wake-up signal includes one or more combinations of human voice, wake-up keywords, and voiceprint features;

[0113] S12: If the detected human voice wake-up signal reaches a preset threshold, the virtual digital human is woken up, and the virtual digital human is switched from a standby state to a working state.

[0114] In some specific embodiments, as an alternative to the wake-up algorithm, the state of the face and lips of the speaking object is captured visually, which specifically includes the following steps:

[0115] Lock the current interacting person by face or skeleton;

[0116] Capture the lip status of the current interacting person through computer vision algorithms;

[0117] When the lips are open or closed, the microphone is turned on for directional sound reception. When the lips are closed, the microphone is turned off for directional sound reception. This solves the problem of misreceiving sounds in noisy environments or when multiple people are talking.

[0118] Through the lip reading recognition algorithm, the continuously recorded lip movement video is converted into lip reading to enhance the microphone's voice recognition.

[0119] Further, in step S11, referring to Figure 10 , detecting whether there is a human voice wake-up signal includes the following steps:

[0120] S111, performing noise reduction preprocessing on the received ambient sound, wherein the noise reduction preprocessing includes framing, windowing, pre-emphasis and adaptive endpoint detection (VAD) on the sound signal;

[0121] S112, extracting speech features based on MDCC, and inputting the speech features into the voiceprint recognition model;

[0122] S113: Detect whether there is a human voice wake-up signal, and perform voiceprint matching between the voice features and the existing voiceprints in the voiceprint matching library;

[0123] S114: If the voiceprint is matched successfully, retrieve the user information that matches the voiceprint; if the voiceprint is matched unsuccessfully, create a new user storage space.

[0124] In some specific embodiments, in step S114, the user storage space is set in Amazon's private storage repository.

[0125] In some specific embodiments, the voiceprint recognition process first extracts voice features, then trains the features in a model, and finally finds the result with the highest or closest score. Voiceprint recognition includes a training phase and a testing phase. The training phase includes inputting training voice, feature extraction, model training, and searching the voiceprint library or generating a new target user registration voice. The testing phase includes inputting test voice, feature extraction, voiceprint matching scoring, and distinguishing between target and non-target users. For virtual digital humans, target users can be defined as users who have already interacted with the virtual digital human and are stored in the memory. Non-target users are defined as users who are interacting with the virtual digital human for the first time. The virtual digital human requires new storage space to store user information.

[0126] In some specific embodiments, a flowchart of data modeling of a virtual digital human for sound is provided. The virtual digital human first receives the sound, pre-processes the sound, performs MFCC and voiceprint recognition on the sound, and then communicates with the user or creates a new user based on the recognition result.

[0127] Furthermore, in step S112,

[0128] The speech features include linear prediction coefficients LPCC and Mel cepstral coefficients MFCC.

[0129] Specifically, speech contains many characteristics. Feature extraction involves converting the continuous time-domain speech signal into a discrete frequency-domain signal, analyzing the energy of the speech signal, and extracting the necessary speech features from the continuous time-domain signal to represent the speech. There are many types of speech features, including Linear Prediction Coefficients (LPCCs) and Mel-Frequency Cepstral Coefficients (MFCCs) (Girin-Laurent).

[0130] The cochlea, a human organ, helps humans detect speech signals amidst noise. It has been discovered that the cochlea plays a role in filtering out noise when humans receive speech signals. Therefore, the idea is to leverage the human method of sound discrimination to construct a filter that replaces the cochlear function and achieves this goal. The human cochlea is similar to a combination of multiple filters. It is very sensitive to low-frequency signal energy. Its sensitivity is linear for frequencies below 1000 Hz and logarithmic for frequencies above 1000 Hz. This irregular frequency segmentation is a key feature that distinguishes the Mel cepstrum of speech signals from the standard cepstrum. The Mel cepstrum is one of the most common speech features in speech signal processing. Research on the biological function of the human ear has revealed that the human ear has varying sensitivity to speech signals of varying amplitudes and frequencies. The human ear has a limited range of amplitude and frequency, and cannot hear speech signals with amplitudes and frequencies exceeding its perceptibility. For example, ultrasound has a frequency of approximately 20,000 Hz, and speech signals below 20 Hz are also inaudible to the human ear. The frequency range of speech signals that can be heard by the human ear is within 20Hz to 20,000Hz. Noise with a frequency between 200Hz and 5,000Hz has the greatest impact on speech signals, making the text content within the speech signal unclear.

[0131] Furthermore, in step S113, the human voice wake-up signal includes a wake-up word, and the wake-up word is stored in the common area of the voiceprint recognition model.

[0132] In some embodiments, the virtual digital human voice monitoring software collects environmental sounds and detects whether there is a human voice wake-up signal. If the detected human voice wake-up signal reaches a threshold, the virtual digital human is awakened and enters the working state; after entering the working state, the virtual digital human outputs a response signal to the user and collects the user's voice signal through a microphone; the virtual digital human converts the collected user voice signal into text through the voice-to-text service, inputs it into the large language model service, the large language model service outputs text, and outputs voice through the text-to-speech service; the virtual familiar human outputs expressions and actions based on the animation control program.

[0133] Further, in step S1, referring to Figure 7 , the visual awakening process includes the following steps:

[0134] S13, in standby mode, controlling the visual sensor of the virtual digital human to collect images;

[0135] S14, detecting human skeleton recognition and / or face recognition to detect whether someone is approaching;

[0136] S15. If it is detected that someone is approaching, the virtual digital human is awakened and switched from the standby state to the working state.

[0137] In some specific embodiments, the virtual digital human's visual sensor collects images and detects whether someone is approaching through human skeleton recognition and face recognition; if someone is detected approaching, the virtual digital human requests the AI brain to output language text through the large language model service; based on the text-to-speech service, the language text output by the large language model service is converted into speech; based on the virtual human animation control program, sounds, expressions and actions are output; if the microphone receives the user's response, feedback is output for the received content; voice information of the user's response and questions is received through the microphone; based on the speech-to-text service, the voice information is converted into text; based on the large language model service, text information is input and the response text information is output; based on the text-to-speech service, the response text information is converted into virtual human voice information; based on the virtual human animation control program, expressions and actions are output.

[0138] Furthermore, in step S1, after starting the virtual digital human, the standby state can be skipped and the working state can be directly entered.

[0139] Specifically, after the virtual digital human is turned on, it can directly play animations and voices on the display screen without detecting whether the user is approaching to attract the user's attention.

[0140] Further, in step S2,

[0141] The multimodal sensor 200 of the virtual digital human includes at least a depth sensor, an RGB camera and a microphone. The multimodal sensor 200 is used to perform at least face recognition, skeleton recognition, posture and gesture recognition on the human.

[0142] Further, in step S3, referring to Figure 7 and Figure 8 , comprising the following steps;

[0143] S31, after entering the working state, the virtual digital human receives the user's image and voice through the multimodal sensor 200, and transmits the voice and image to the voice and image extraction network to extract the voice and image;

[0144] S32. Converting speech into text based on a speech-to-text service;

[0145] S33. Output language text based on the large language model service;

[0146] S34. Based on the text-to-speech service, convert the language text output by the large language model service into speech;

[0147] S35, outputting sounds, expressions and movements based on the virtual human animation control program;

[0148] S36: If the multimodal sensor 200 receives a response from the user, steps S31 to S35 are repeated to provide feedback on the received content in the form of voice, expression, and action.

[0149] Further, in step S31, referring to Figure 11 , the speech image extraction network includes: speech separation network and visual separation network,

[0150] The speech separation network includes a mixed speech receiving module, an STFT speech time-frequency conversion module, a speech feature extraction module, a speech upsampling module and an STFT frequency-time conversion module connected in sequence.

[0151] The visual separation network includes a hybrid visual receiving module, a VGGFace feature extraction module, and a visual feature network module connected in sequence.

[0152] It also includes a fusion module, the input of the fusion module is connected to the speech feature extraction module and the VGG feature extraction module respectively, and the output of the fusion module is connected to the speech upsampling module.

[0153] Specifically, the speech image extraction network structure mainly consists of two parts: a speech separation network and a visual separation network. The speech separation network receives a mixed speech spectrogram and outputs a spectrum mask. The network model structure adopts a codec structure similar to U-Net, which are respectively called a feature extraction module and an upsampling module. There is also a jump connection between the feature extraction module and the upsampling module. Its main function is to use the shallow features of the feature extraction module to promote the upsampling module to predict a high-resolution time-frequency mask. However, the performance of feature extraction modules of different depths on different tasks varies greatly. In order to enable the network to adaptively learn network features of different depths, a residual connection is added to the feature extraction module. The visual separation network receives a set of faces that completely corresponds to the speech spectrogram and outputs visual features that can be distributed in the same way as the speech separation network. Residual connections are added to the U-Net model and an audio and video feature fusion module is included.

[0154] Further, in step S33, referring to Figure 12 The large language model includes a corpus system module, a pre-training model and a fine-tuning module connected in sequence.

[0155] The corpus system module includes pre-trained corpus and fine-tuned corpus. The pre-trained corpus includes text data collected from books, magazines, and encyclopedias. The fine-tuned corpus includes annotated text data crawled from open source code libraries, annotated by experts, and processed through user dialogue.

[0156] The pre-trained model is first trained extensively on large-scale training data using an unsupervised learning method to obtain a general language model with strong generalization capabilities. The training method of the pre-trained model includes at least word vector embedding, contrastive pre-training, and contextual learning.

[0157] The fine-tuning module includes freezing the bottom layer of the pre-trained model and adjusting the weights of the upper layer, the bottom layer includes word vectors, and the upper layer includes classifiers.

[0158] Specifically, the corpus system includes pre-training data and fine-tuning data, the latter of which includes code and dialogue fine-tuning data. The pre-training data includes massive amounts of text data collected from books, magazines, encyclopedias, and other sources, enabling the model to learn the logical relationships of language. The fine-tuning data includes high-quality annotated text data processed through crawling open source code libraries, expert annotation, and user dialogue, further enhancing its conversational capabilities.

[0159] In some embodiments, pre-training is the foundation for building large-scale language models. This involves conducting extensive, generalized training on large-scale training data, using unsupervised learning methods, to develop a generalizable and highly generalizable language model. Through pre-training on this large-scale data, the model initially acquires the capabilities of human language understanding and contextual learning, capturing the semantic similarities between text and code snippets, thereby generating more accurate text and code vectors and supporting subsequent fine-tuning tasks.

[0160] In some embodiments, fine-tuning is the guarantee for the practical application of the model. It refers to further training the pre-trained model on a dataset for a specific task, which usually includes freezing the underlying layers of the pre-trained model (such as word vectors) and adjusting the weights of the upper layers (such as classifiers). Fine-tuning the pre-trained model will greatly shorten the training time, save computing resources and speed up the training convergence. Based on the pre-trained model with strong generalization ability, by integrating code data-based training and instruction-based fine-tuning, fine-tuning is performed using a specific dataset to make it have a stronger question-and-answer dialogue text generation capability.

[0161] Further, refer to Figure 13 The training steps of the pre-trained model include: collecting data sets and performing supervised fine-tuning; collecting comparison data and training the scoring model and using reinforcement learning for the reward model to optimize the cloud training model.

[0162] Furthermore, the collecting a dataset and performing supervised fine-tuning includes sampling a prompt from the dataset; marking a preferred output answer and performing supervised fine-tuning on the pre-trained model based on the collected data.

[0163] Further, refer to Figure 13The collecting of comparison data and training of the scoring model includes sampling a prompt and a number of corresponding model outputs; scoring and sorting the outputs and training the scoring model on the scored and sorted data.

[0164] Further, the using reinforcement learning for the reward model to optimize the cloud training model includes resampling a prompt; the reward model scores the output and optimizes model parameters.

[0165] Furthermore, the step S3 also includes receiving a task instruction, breaking the task instruction into simple steps and executing them in sequence.

[0166] Furthermore, in step S4, the virtual digital human is controlled to change clothes based on the character settings preset in the 3D clothing asset library, and the 3D clothing asset library at least includes hospital intelligent guidance virtual digital human clothing, museum explanation virtual digital human clothing, and restaurant reception virtual digital human clothing.

[0167] Further, in step S35, the virtual human animation control program includes a voice-to-expression service Audio2FaceService and a voice-to-action service Audio2ActionService. The voice-to-expression service Audio2FaceService is used to convert the virtual human's voice information into the virtual human's facial expressions, and the voice-to-action service Audio2ActionService is used to convert the virtual human's voice information into the virtual human's body movements.

[0168] Furthermore, the speech-to-expression service Audio2FaceService is a speech animation synthesis model based on the BlendShapes method, which includes a three-dimensional face control parameter prediction module based on different speech emotions, an expression base construction module based on sample expressions, and a three-dimensional face animation synthesis module.

[0169] The three-dimensional face control parameter prediction module based on different voice emotions includes a voice preprocessing module, a voice feature extraction module, an audio and video mapping module and a three-dimensional face control parameter module connected in sequence;

[0170] The expression base construction module based on sample expressions includes a sample expression model library, an expression feature extraction module, an expression migration module, and an expression base construction module connected in sequence;

[0171] The three-dimensional facial animation synthesis module includes a three-dimensional facial control parameter smoothing module and a three-dimensional facial animation synthesis module connected in sequence. The input of the three-dimensional facial control parameter smoothing module is connected to the output of the three-dimensional facial control parameter module, and the input of the three-dimensional facial animation synthesis module is connected to the output of the expression base construction module.

[0172] It should be appreciated that the method steps in the embodiments of the present invention can be implemented or executed by computer hardware, a combination of hardware and software, or by computer instructions stored in a non-transitory computer-readable memory. The method can use standard programming techniques. Each program can be implemented in a high-level procedural or object-oriented programming language to communicate with the computer system. However, if desired, the program can be implemented in assembly or machine language. In any case, the language can be a compiled or interpreted language. In addition, for this purpose, the program can be run on a programmed application-specific integrated circuit.

[0173] Furthermore, the operations of the processes described herein may be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by the context. The processes described herein (or variations and / or combinations thereof) may be performed under the control of one or more computer systems configured with executable instructions and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that is executed collectively on one or more processors, by hardware, or a combination thereof. The computer program includes a plurality of instructions that can be executed by one or more processors.

[0174] Further, the methods can be implemented in any type of computing platform that is operably connected to a suitable computer, including but not limited to a personal computer, a minicomputer, a mainframe, a workstation, a network or distributed computing environment, a separate or integrated computer platform, or in communication with a charged particle tool or other imaging device, etc. Various aspects of the present invention can be implemented as machine-readable code stored on a non-transitory storage medium or device, whether removable or integrated into a computing platform, such as a hard disk, an optical read and / or write storage medium, RAM, ROM, etc., so that it can be read by a programmable computer, and when the storage medium or device is read by the computer, it can be used to configure and operate the computer to perform the processes described herein. In addition, the machine-readable code, or portions thereof, can be transmitted over a wired or wireless network. When such media includes instructions or programs that implement the steps described above in conjunction with a microprocessor or other data processor, the invention described herein includes these and other different types of non-transitory computer-readable storage media. When programmed according to the methods and techniques of the present invention, the present invention can also include the computer itself.

[0175] The computer program can be applied to input data to perform the functions described herein, thereby converting the input data to generate output data that is stored in a non-volatile memory. The output information can also be applied to one or more output devices such as a display. In a preferred embodiment of the present invention, the converted data represents a physical and tangible object, including a specific visual depiction of the physical and tangible object produced on the display.

[0176] The above description is merely a preferred embodiment of the present invention. The present invention is not limited to the aforementioned embodiments. As long as the technical effects of the present invention are achieved by the same means, any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention. Within the scope of protection of the present invention, various modifications and variations of the technical solutions and / or implementation methods are possible.

Claims

1. A real-time digital human generation method based on interactive scenes, characterized in that: The method comprises the following steps: S100: Acquire ambient sound and video to detect whether there is a wake-up signal; S200: Determine whether to receive a wake-up signal, and make the virtual digital human enter a wake-up process, switching from a standby state to a working state. The wake-up process includes a voice wake-up process. The voice wake-up process includes the following steps: In standby mode, the system collects ambient sound and detects whether there is a human voice wake-up signal, which includes a combination of one or more of human voice, wake-up keywords, and voiceprint features. If the detected human voice wake-up signal reaches the preset threshold, the virtual digital human will be awakened and switched from standby mode to working mode. Detecting whether a human voice wake-up signal exists comprises the following steps: The received ambient sound is pre-processed for noise reduction, wherein the noise reduction pre-processing includes framing, windowing, pre-emphasis and adaptive endpoint detection (VAD) of the sound signal. Extract speech features based on MDCC and input the speech features into the voiceprint recognition model. Detect whether there is a human voice wake-up signal, and match the voice features with the existing voiceprints in the voiceprint matching library. If the voiceprint matches successfully, retrieve the user information that matches the voiceprint; if the voiceprint matches unsuccessfully, create a new user storage space; S300, acquiring environmental information and user information based on a multimodal sensor, wherein the user information includes user image information and user voice information; S400, comparing the user information with existing information in the storage device. If the user information is already recorded in the storage device, retrieving the existing user information from the storage device. If the user information is not recorded in the storage device, creating new user information in the storage device. S500, based on a large language model, the virtual human interactive all-in-one machine conducts conversations with users; The step S500 includes: S31, after entering the working state, the virtual digital human receives the user's image and voice through the multimodal sensor 200, and transmits the voice and image to the voice and image extraction network to extract the voice and image; S32. Converting speech into text based on a speech-to-text service; S33. Output language text based on the large language model service; S34. Based on the text-to-speech service, convert the language text output by the large language model service into speech; S35, outputting sounds, expressions and movements based on the virtual human animation control program; S36. If the multimodal sensor 200 receives a response from the user, steps S310 to S350 are repeated to provide feedback on the received content by outputting voice, expression, and action.

2. A virtual human interactive all-in-one machine, characterized in that: include: A display device, wherein the display device is a vertical large screen; a multimodal sensor, configured to detect the proximity of a user and acquire audio and video of the user, the multimodal sensor being disposed on the display device; A storage device, wherein the storage device includes a long-term storage device and a short-term storage device, the long-term storage device is a database, the database includes a relational database and a vector database, and the short-term storage device is a memory; A processing device is used to process interaction data between the virtual person and the user. The processing device is installed on the back of the display device. The display device, the multimodal sensor and the storage device are respectively connected to the processing device.

3. The virtual human interactive all-in-one machine according to claim 2, characterized in that: The display device is an LED screen, a naked-eye 3D screen, or a holographic screen.

4. The virtual human interactive all-in-one machine according to claim 2, characterized in that: The multimodal sensor includes one or more of a 4K high-definition RGB camera, a Tof depth camera, and a microphone array.

5. The virtual human interactive all-in-one machine according to claim 2, characterized in that: The database includes an action library and an expression library, each action record in the action library includes an action tag, and each expression record in the expression library includes an expression tag, and the action tag and the expression tag are respectively called by the virtual person when interacting with the user; The database further includes a tool library interface, which includes at least a printer interface, a search interface, an encyclopedia interface, a date and time interface, and an execution code interface.

6. The virtual human interactive all-in-one machine according to claim 2, characterized in that: The processing device at least includes a real-time rendering device for real-time rendering and output of the virtual person's actions and expressions. The real-time processing device is connected to the processing device. The real-time processing device is installed on the display device or is independently installed with the processing device through an electrical connection.

7. The virtual human interactive all-in-one machine according to claim 6, characterized in that: It also includes a heat dissipation device, which is used to dissipate heat for the real-time rendering device and is installed on the real-time rendering device.

8. The virtual human interactive all-in-one machine according to claim 2, characterized in that: The processing device includes an AI algorithm service program, an interactive backend program, a backend web program, and a virtual human rendering program; The AI algorithm service program is a Python program, and the AI algorithm service program includes Audio2Face algorithm service, Text2Action algorithm service, large language model distribution service and AI Agent service; The interactive backend program is a Java program, and includes an account management system, a memory configuration system, a 3D asset management system, and a large model interactive management system. The interactive backend program is connected to the database; The backend web program is an H5 program, which includes an account management system page, a content configuration system page, a 3D asset management system page, and a large model interaction management page; The virtual human rendering program includes a UE5 program, and the virtual human rendering program includes virtual human rendering, action animation logic and external interaction program.

9. The virtual human interactive all-in-one machine according to claim 2, characterized in that: The device further comprises a housing in which the display device, the multimodal sensor, the storage device and the processing device are installed.