Virtual human control method and device based on multi-modal large model

By deploying virtual human control methods and devices of multimodal large models on the intranet, integrating Q&A database, RAG knowledge base and LLM large language model, the problem of difficulty in meeting security and flexibility in the existing technology is solved, and efficient and secure user Q&A processing and system scalability are achieved.

CN120144716APending Publication Date: 2025-06-13CHINA PETROLEUM & CHEMICAL CORP +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510283265.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The artificial intelligence dialogue systems in the prior art are difficult to meet the needs of security and flexibility at the same time, especially when processing user-input Q&A audio, there are problems of security risks and insufficient system scalability.

Method used

By deploying virtual human control methods and devices for multimodal large models on the intranet, integrating with the LLM large language model using the Q&A database and the RAG knowledge base, ensuring that the data is completely isolated from the external network in the processing process, enhancing data security, and allowing the voice recognition module to be connected to the external network to improve resource utilization efficiency.

Benefits of technology

It realizes the scalability and flexibility of the system while ensuring data security, and can more effectively process user-entered Q&A audio, provide high-quality responses and improve user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144716A_ABST
    Figure CN120144716A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a virtual human control method and device based on a multi-modal large model. The problem that in the prior art, an artificial intelligence dialogue system is difficult to meet safety and flexibility at the same time is solved. According to the virtual human control method based on the multi-modal large model, the question and answer database, the RAG knowledge base and the LLM large language model are all connected into the intranet, and the model and knowledge base data are completely isolated from an external network in the processing flow through intranet deployment, so that the processing efficiency is improved; parts of the modules needing to access the external internet are accessed through a unified internet outlet under the intranet, so that the security of the data in the transmission and storage process is ensured, and the risks of external attacks and data leakage of the data are avoided. Meanwhile, the voice recognition module can be connected with an external network, the resource utilization efficiency can be improved, and the data security is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of robots, and particularly to a virtual human control method and device based on a multimodal large model. Background Art

[0002] A robot is a virtual character created using 3D modeling, artificial intelligence, and interaction technologies, capable of simulating human behaviors and expressions with a realistic appearance, movements, and interaction methods, providing users with functions such as rich and diverse information display, entertainment interaction, and auxiliary services.

[0003] The large model dialogue technology is a large neural network model based on deep learning, aiming to simulate the ability of human natural language processing and interaction. This technology uses a large amount of text data for training, can understand and generate complex natural language texts, and achieve smooth and intelligent human-machine conversations. It can not only recognize the user's intentions, understand context information, but also generate appropriate and coherent responses according to the conversation content, greatly improving the naturalness and effectiveness of human-machine interaction.

[0004] Applying the large model in robots, with its powerful natural language processing ability and extensive knowledge reserve, provides an intelligent core for the robots, enabling them to more accurately understand user intentions and generate high-quality responses. The robots, through their anthropomorphic appearance, movements, and voices, bring a more intuitive and vivid interaction experience to users. The two together promote the wide application and innovative development of digital humans in fields such as customer service, content creation, and education and training.

[0005] In the prior art, most similar artificial intelligence dialogue systems adopt a single network architecture and completely rely on the external network for data processing and model inference. This single network architecture has many deficiencies and may face security risks, and information compliance cannot be guaranteed. Completely relying on the internal network for data processing and model inference can ensure the security of data at the physical level, but at the same time limits the scalability and flexibility of the system. For some modules sensitive to user experience, some require higher computing speeds of cloud resources to meet the performance requirements of instant messaging, which cannot be achieved in a full internal network environment. Summary of the Invention

[0006] In view of this, embodiments of the present invention provide a virtual human control method and device based on a multimodal large model, which solve the problem that it is difficult for artificial intelligence dialogue systems in the prior art to meet both security and flexibility.

[0007] As a first aspect of the present invention, the present invention provides a virtual human control method based on a multimodal large model, including:

[0008] Obtain the Q&A audio input by the user, where the Q&A audio includes the target Q&A, and convert the Q&A audio into text data;

[0009] Perform data processing on the text data to determine the target interface in the LLM large language model corresponding to the target Q&A;

[0010] When the target interface is the Q&A data interface, determine the target answer text data corresponding to the target Q&A based on the Q&A database;

[0011] When the target interface is the knowledge base Q&A interface, determine the target answer text data corresponding to the target Q&A based on the RAG knowledge base.

[0012] In an embodiment of the present invention, the performing data processing on the text data to determine the target interface in the LLM large language model corresponding to the target Q&A includes:

[0013] Based on the Q&A database, the database processing module queries in the Q&A database whether there is a target preset Q&A that matches the target Q&A;

[0014] When a target preset Q&A that matches the target Q&A is found in the Q&A database, determine that the Q&A database interface in the LLM large language model is called as the target interface;

[0015] Among them, the determining the target answer text data corresponding to the target Q&A based on the Q&A database includes:

[0016] According to the target browser address of the target preset Q&A, obtain the initial answer text data corresponding to the target preset Q&A;

[0017] The LLM large language model processes the initial answer text data based on the target Q&A to determine the target answer text data corresponding to the target Q&A.

[0018] In an embodiment of the present invention, the virtual human control method further includes:

[0019] Display the target answer text data and the target browser address.

[0020] In an embodiment of the present invention, the performing data processing on the text data to determine the target interface in the LLM large language model corresponding to the target Q&A further includes:

[0021] When no target preset Q&A that matches the target Q&A is found in the Q&A database, determine whether the target Q&A is a real-time data query according to the text data;

[0022] When it is determined according to the text data that the target Q&A is not a real-time data query, determine the knowledge base Q&A interface as the target interface, and perform word segmentation on the text data to generate multiple question vectors;

[0023] Among them, determining the target answer text data corresponding to the text data based on the RAG knowledge base includes:

[0024] The RAG knowledge base queries the target vector matching the question vector in the RAG knowledge base according to the question vector, and converts the target vector into target text;

[0025] Construct a context according to the target text and the text data corresponding to the Q&A audio;

[0026] The LLM large language model generates target Q&A text data based on the context and multiple target texts.

[0027] In an embodiment of the present invention, the virtual human control method further includes:

[0028] When it is determined according to the text data that the target Q&A is a real-time data query, process the text data through Prompt engineering to obtain query variables;

[0029] Determine the target data interface according to the query variables to call the target database corresponding to the target data interface;

[0030] Query the target answer text data matching the text data in the target database.

[0031] In an embodiment of the present invention, the virtual human control method further includes:

[0032] Convert the target answer text data into a target answer audio, and play the target answer audio.

[0033] As a second aspect of the present invention, the present invention also provides a virtual human control device based on a multi-modal large model, including:

[0034] An external network switch and an internal network switch;

[0035] A voice recognition module, which is used to convert the Q&A audio input by the user into text data, the Q&A audio includes the target Q&A input by the user, and the voice recognition module is connected to the external network switch;

[0036] An LLM large language model, which has a Q&A data interface, a knowledge base Q&A interface, and a data interface;

[0037] A data processing module, which is used to process the text data to determine a target interface in the LLM large language model corresponding to the target question and answer, wherein the data processing module is connected to the intranet switch;

[0038] A question and answer database, which stores multiple preset questions and answers and browser addresses corresponding to the preset questions and answers. The question and answer database interface is connected to the question and answer database, and the question and answer data module is connected to the intranet switch. The question and answer database is used to determine the target answer text data corresponding to the target question and answer based on the question and answer database when the target interface is a question and answer data interface;

[0039] A RAG knowledge base, which stores preset knowledge and preset vectors corresponding to the preset knowledge. The knowledge base question and answer interface is connected to the RAG knowledge base, and the RAG knowledge base is connected to the intranet switch. The RAG knowledge base is used to determine the target answer text data corresponding to the target question and answer based on the RAG knowledge base when the target interface is a knowledge base question and answer interface.

[0040] In an embodiment of the present invention, the virtual human control device further includes:

[0041] A playback module, which is communicatively connected to the LLM large language model. The playback module is used to convert the target answer text data into a target answer audio and play the target answer audio.

[0042] In an embodiment of the present invention, the virtual human control device further includes:

[0043] An action instruction module, which is connected to the LLM large language model and the intranet switch; and

[0044] A Unity client, which is connected to the action instruction module and the intranet switch.

[0045] In an embodiment of the present invention, the virtual human control device further includes:

[0046] A display screen, which is connected to the intranet switch. The display screen is used to display the target answer text data and the corresponding target browser address.

[0047] A virtual human control method based on a multimodal large model provided by the present invention, a Q&A database, a RAG knowledge base, and an LLM large language model are all connected to the internal network. Through the internal network deployment, the model and the knowledge base data are completely isolated from the external network in the processing flow. Some modules that need to access the external Internet access through a unified Internet exit under the internal network to ensure the security of data during transmission and storage and avoid the risks of external data attacks and data leakage. At the same time, the speech recognition module can be connected to the external network, which can improve resource utilization efficiency and strengthen data security. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 The figure shows a working block diagram of a virtual human control device based on a multimodal large model provided by an embodiment of the present invention.

[0049] Figure 2 The figure shows a working block diagram of a virtual human control device based on a multimodal large model provided by another embodiment of the present invention.

[0050] Figure 3 The figure shows a schematic flowchart of a virtual human control method based on a multimodal large model provided by an embodiment of the present invention.

[0051] Figure 4 The figure shows a schematic flowchart of a virtual human control method based on a multimodal large model provided by an embodiment of the present invention.

[0052] Figure 5 The figure shows a schematic flowchart of a virtual human control method based on a multimodal large model provided by another embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0053] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0054] As the first aspect of the present invention, the present invention provides a virtual human control device based on a multimodal large model. Figure 1 The figure shows a working block diagram of a virtual human control device based on a multimodal large model provided by an embodiment of the present invention. As Figure 1 shown, a virtual human control device based on a multimodal large model includes:

[0055] External network switch 11 and internal network switch 10; among them, the external network switch 11 is used to refer to the switch for connecting to the external network. The internal network switch 10 is mainly used to build an internal network, realize data transmission, device interconnection, and internal resource management. The internal network switch 10 can connect multiple terminal devices (such as computers, servers, network printers, etc.) and control the data flow direction to ensure stable and efficient data transmission in the internal network.

[0056] Speech recognition module 20, which is used to convert the Q&A audio input by the user into text data. The Q&A audio includes the target Q&A input by the user. The speech recognition module 20 is connected to the external network switch 10; specifically, the speech recognition module 20 can be ASR (Automatic Speech Recognition) speech recognition, which can convert the speech input by the user into computer-readable text data to achieve human-computer interaction and natural language processing.

[0057] Q&A database 30, which stores multiple preset Q&As and the corresponding browser addresses. The Q&A data module 30 is connected to the internal network switch 10;

[0058] Specifically, the Q&A database 30 is a pre-constructed database that stores multiple preset Q&As and the corresponding browser addresses. The browser address is the address of the preset answer that can answer the preset Q&A, that is, by opening the browser address, the preset answer corresponding to the preset Q&A can be obtained. It should be noted that one preset Q&A can correspond to one browser address or multiple browser addresses. Similarly, multiple preset Q&As can also correspond to one browser address.

[0059] RAG knowledge base 40, which stores preset knowledge and the corresponding preset vectors. The RAG knowledge base is connected to the internal network switch 10.

[0060] Specifically, the RAG knowledge base 40 is a knowledge storage and retrieval system built based on RAG (Retrieval-Augmented Generation) technology. The RAG knowledge base 40 is pre-constructed. The professional knowledge is stored in the local database in the form of vectors. In the internal network environment, enterprise-related business data is collected, including knowledge documents, real-time data, etc., the data is cleaned and preprocessed, and the real-time data is encapsulated through an interface to ensure the quality and accuracy of the data.

[0061] The LLM large language model 50 has a question-answer data interface, a knowledge base question-answer interface, and a data interface. The question-answer database interface is connected to the question-answer database, the knowledge base question-answer interface is connected to the RAG knowledge base, and the data interface is connected to multiple databases, such as a data dictionary.

[0062] Specifically, the LLM large language model 50 is a deep learning model trained based on massive text data, which can generate natural language text, deeply understand the meaning of the text, and handle various natural language tasks, such as text summarization, question answering, translation, etc. The LLM large language model in the present invention has three interfaces, the question answering data interface and the knowledge base question answering interface correspond to the question answering database and the RAG knowledge base respectively. The real-time data interface directly calls the LLM large language model.

[0063] The data processing module 60 is used to process the text data to determine the target interface corresponding to the relationship in the LLM large language model, wherein the data processing module 60 is connected to the intranet switch.

[0064] Specifically, the data processing module 60 is used to perform data processing on the text data to determine the target interface corresponding to the target question and answer in the LLM large language model, and call the target interface in the LLM large language model to call different data, so as to generate the target answer text data according to the called data. For example, when it is determined that the target interface is the question and answer data interface, the question and answer data interface of the LLM large language model is called, and at the same time, the question and answer database obtains the initial answer text data corresponding to the target preset question and answer according to the target browser address of the target preset question and answer; and the initial answer text data is transferred to the LLM large language model through the question and answer data interface, and the LLM large language model performs data processing according to the corresponding initial answer text data to determine the target answer text data.

[0065] The present invention provides a virtual human control device based on a multimodal large model. The question-answer database, RAG knowledge base, and LLM large language model are all connected to the intranet. Through the intranet deployment, the model and knowledge base data are completely isolated from the external network in the processing flow. Some modules that need to access the external Internet are accessed through the unified Internet export under the intranet, ensuring the security of data during transmission and storage, and avoiding the risk of external attacks and data leakage. At the same time, the voice recognition module can be connected to the external network, which can improve resource utilization efficiency and strengthen data security.

[0066] In one embodiment of the present invention, the virtual human control device based on the multimodal large model further includes: a playback module, which is in communication with the LLM large language model. The playback module can convert the target answer text data generated by the LLM large language model into a target answer audio, and play the target answer audio.

[0067] In an embodiment of the present invention, as Figure 2 shown, the virtual human control device based on the multimodal large model further includes: an action instruction module 70, the action instruction module 70 is connected to the LLM large language model 50, and the action instruction module 70 is connected to the intranet switch 10; and a Unity client 80, the Unity client 80 is connected to the action instruction module 70, and the Unity client 80 is connected to the intranet switch 10.

[0068] Specifically, the action instruction module 70 can convert the target answer text data into action instructions, and the action instructions include but are not limited to: expression action instructions, limb action instructions, etc.

[0069] The Unity client 80 uses a computer as the carrier of the program, uses the Unity engine, and through the designed digital image, cooperates with language and limb behaviors to perform data docking with the backend program. At the same time, combined with the three-dimensional rendering engine, it can convert the instructions into the dynamic performance of the digital human in real time, providing an immersive human-computer interaction experience. Thus, the virtual human control device based on the multimodal large model can combine the large language model and three-dimensional graphics rendering technology to achieve high-precision semantic understanding and dynamic generation.

[0070] Optionally, the virtual human control device based on the multimodal large model further includes a display screen, the display screen is connected to the intranet switch and the LLM large language model 50, and the display screen can display the target answer text data generated by the LLM large language model 50 and / or the target preset website corresponding to the target Q&A.

[0071] Optionally, the virtual human control device based on the multimodal large model may further include an interaction module: the interaction module can be in a manual interaction mode. For example, the interaction module can be a remote control recording pen, and the recording pen is equipped with buttons mapped to the computer. When the user presses different buttons, it will be mapped to the simulated keyboard input keys such as B, PageUP, PageDOWN, ESC, etc. on the computer. If these keys are detected to be pressed, the corresponding actions will be executed. Pressing and holding the B key is to start recording, and releasing the B key is to end recording. This operation can realize the function of recording the user's voice. Pressing PageUP is to answer a preset question, and this operation can enable the user to actively answer the preset question and enrich the interaction method. Pressing PageDOWN is for the virtual human to make a greeting gesture, and this operation can realize the interaction that the virtual human actively greets the visitors. Pressing the ESC key is to terminate the current operation of the virtual human and switch to the default state, which is convenient for the operator to stop in time when the virtual human tells wrong content or other unexpected situations. In addition, the remote control pen is also equipped with other buttons, and different functions can be customized according to user needs.

[0072] As a second aspect of the present invention, the present invention also provides a method for controlling a virtual human based on a multimodal large model. Based on the above-mentioned device for controlling a virtual human based on a multimodal large model, as Figure 3 shown, a device for controlling a virtual human based on a multimodal large model provided by an embodiment of the present invention includes the following steps:

[0073] S1: The speech recognition module obtains the Q&A audio input by the user. The Q&A audio includes the target Q&A, and converts the Q&A audio into text data;

[0074] Specifically, the speech recognition module can be ASR (Automatic Speech Recognition) speech recognition, which can convert the speech input by the user into computer-readable text data, realizing human-computer interaction and natural language processing.

[0075] The speech recognition module converts the Q&A audio input by the user into text data, and transmits the text data to the data processing module through the internal network. The Q&A audio includes the target Q&A of the user.

[0076] S2: The data processing module processes the text data to determine the target interface in the LLM large language model corresponding to the target Q&A;

[0077] After receiving the text data sent by the speech recognition module, the data processing module processes the text data to determine the target interface in the LLM large language model corresponding to the target Q&A. That is, the target database to be queried is determined. For example, when the target interface is the Q&A data interface, the target database is the Q&A database; when the target interface is the knowledge base Q&A interface, the target database is the RAG knowledge base. Specifically, after determining the target interface, that is, after determining the target database, the target database can be made to query in the target database according to the text data to query the initial data corresponding to the target Q&A.

[0078] S3: When the target interface is the Q&A data interface, determine the target answer text data corresponding to the target Q&A based on the Q&A database;

[0079] After determining that the Q&A database is the target database, the Q&A database can be called, and the text data is transmitted to the Q&A database. The Q&A database then queries the initial data corresponding to the target Q&A in the Q&A database, and then transmits the initial data to the LLM large language model for processing to generate the target answer text data.

[0080] S4: When the target interface is the knowledge base Q&A interface, determine the target answer text data corresponding to the target Q&A based on the RAG knowledge base.

[0081] Similarly, after determining that the RAG knowledge base is the target database, the RAG knowledge base can be called to transmit the text data to the RAG knowledge base. The RAG knowledge base then queries the initial data corresponding to the target Q&A in the RAG knowledge base based on the text data, and then transmits the initial data to the LLM large language model for processing to generate the target answer text data.

[0082] A virtual human control method based on a multimodal large model provided by the present invention obtains the user's Q&A audio, converts the Q&A audio into text data, determines the target database (such as a Q&A database or a RAG knowledge base) required for querying the answer based on the text data, calls the corresponding target interface according to the target database, and inputs the initial data queried in the target database into the LLM large language model for processing according to the target interface to generate the corresponding target answer text data. Since the Q&A database, the RAG knowledge base, and the LLM large language model are all connected to the internal network and deployed through the internal network, the model and knowledge base data are completely isolated from the external network in the processing flow. Some modules that need to access the external Internet access through a unified Internet exit under the internal network to ensure the security of data during transmission and storage and avoid the risks of external data attacks and data leaks. At the same time, the speech recognition module can be connected to the external network, which can improve resource utilization efficiency and strengthen data security. In addition, according to the different services to which the user's Q&A belongs, different databases can be selected to obtain the corresponding initial data first, and then the LLM large language model is used to process the initial data to generate the target answer text data, that is, before the LLM large language model answers, the content outside the LLM large language model itself is determined through the Q&A database or the RAG knowledge base, which improves the information acquisition and processing capabilities, optimizes the parameter order of the model, and reduces the computing power cost.

[0083] Optionally, the Q&A database stores not only the preset Q&A and the corresponding preset browser addresses, but also the preset answer text data corresponding to the preset Q&A. When the target interface is the knowledge base interface, after determining the target answer text data corresponding to the text data based on the RAG knowledge base, the target answer text data and the corresponding target Q&A can be stored in the Q&A database to continuously enrich the Q&A database.

[0084] In an embodiment of the present invention, as Figure 4 shown, S2 (the data processing module processes the text data to determine the target interface corresponding to the target Q&A in the LLM large language model) specifically includes the following steps:

[0085] S21: Based on the Q&A database, the database processing module queries the Q&A database for the target preset Q&A that matches the target Q&A;

[0086] Specifically, the Q&A database stores multiple preset Q&As and the corresponding preset browsers for the preset Q&As.

[0087] The specific matching method for whether there is a target preset Q&A that matches the target Q&A can be:

[0088] First, semantically parse the target Q&A and the preset Q&As to obtain the target semantics corresponding to the target Q&A and the reference semantics corresponding to each preset Q&A. The preset Q&A corresponding to the reference semantics whose similarity to the target semantics is greater than the preset similarity is the one that matches the target Q&A. When the number of preset Q&As determined to match the target Q&A based on the reference semantics is multiple, the target preset Q&A that matches the target Q&A can be further determined according to the overlap degree of the words included in the target Q&A and the preset Q&A. Conversely, when the similarity between the reference semantics of all preset Q&As and the target semantics of the target Q&A is less than the preset similarity, it means that there is no target preset Q&A in the Q&A database that matches the target Q&A.

[0089] First, determine whether there are key questions in the target Q&A input by the user through the Q&A database. When there are no key questions, then check whether it is a real-time data query.

[0090] Optionally, the number of target preset Q&As can be 1 or multiple.

[0091] S22: When a target preset Q&A that matches the target Q&A is found in the Q&A database, determine to call the Q&A database interface in the LLM large language model as the target interface;

[0092] When it is determined that a matching target preset Q&A is found, determine to call the Q&A database interface in the LLM large language model as the target interface.

[0093] S23: When no target preset Q&A that matches the target Q&A is found in the Q&A database, judge whether the target Q&A is a real-time data query according to the text data;

[0094] When it is determined that no matching target preset Q&A is found, it is necessary to further judge according to the text data whether the target Q&A is a real-time data query, that is, whether it is necessary to query real-time data.

[0095] S24: When it is judged according to the text data that the target Q&A is not a real-time data query, determine the knowledge base Q&A interface as the target interface and perform word segmentation on the text data to generate multiple question vectors;

[0096] When it is judged according to the text data that the target Q&A is not a real-time data query, determine the knowledge base Q&A interface as the target interface, that is, the target Q&A is a query for internal data. Then the knowledge base Q&A interface is the target interface.

[0097] Meanwhile, tokenize the text data to generate multiple question vectors.

[0098] Specifically, when the judgment result in S21 is yes, that is, when the target preset Q&A matching the target Q&A is found in the Q&A database, determine to call the Q&A database interface in the LLM as the target interface. In this case, S3 (determine the target answer corresponding to the target Q&A based on the Q&A database) specifically includes the following steps:

[0099] S31: Obtain the initial answer text data corresponding to the target preset Q&A according to the target browser address of the target preset Q&A;

[0100] When the target interface is the Q&A database interface, call the target browser address of the target preset Q&A from the Q&A database and obtain the initial answer text data corresponding to the target preset Q&A.

[0101] Specifically, when calling the target browser address, the intranet is also used for the call.

[0102] S32: The LLM processes the initial answer text data based on the target Q&A to determine the target answer text data corresponding to the target Q&A.

[0103] After obtaining the corresponding initial answer text data from the target browser address, the LLM can be used to process the initial answer text data to determine the target answer text data.

[0104] It should be noted that one target Q&A may correspond to multiple target preset Q&As, and one target preset Q&A may also correspond to multiple target browser addresses. Therefore, the number of initial answer text data corresponding to each target preset Q&A may also be multiple. Therefore, when the number of initial answer text data is multiple, when processing the initial answer text data, multiple initial text data corresponding to the same target preset Q&A can be merged, and then refined according to the merged text data to determine the target answer text data.

[0105] When it is determined according to the text data that the target Q&A is not a real-time data query, determine the knowledge base Q&A interface as the target interface, that is, the target Q&A is to query internal data. Then the knowledge base Q&A interface is the target interface. In this case, S4 (determine the target answer corresponding to the target Q&A based on the RAG knowledge base) specifically includes the following steps:

[0106] S41: The RAG knowledge base queries the target vector matching the question vector in the RAG knowledge base according to the question vector and converts the target vector into the target text;

[0107] When the target interface is the knowledge base interface, the RAG knowledge base is called. Since the RAG knowledge base stores preset knowledge and corresponding preset vectors for the preset knowledge. Therefore, after determining the question vector, the target vector matching the question vector can be queried in the RAG knowledge base according to the question vector. And the target vector is converted into target text.

[0108] S42: Construct a context based on the target text and the text data corresponding to the Q&A audio;

[0109] Specifically, the input prompt words (such as text data and target text) can be used to guide the pre-trained language model to generate the required response or complete a specific task (such as expanding the context) to obtain the context.

[0110] S43: The LLM large language model generates target Q&A text data based on the context and multiple target texts.

[0111] The LLM large language model mentions multiple target texts in the context and generates target Q&A text data.

[0112] The virtual human control method based on the multi-modal large model of the present invention can, by integrating the intelligent analysis module, enable the system to parse the user's query requirements in real time, quickly obtain relevant data using the data interface, and generate a natural and fluent conversation reply in combination with the prompt words, realizing the real-time data query function in the conversation and improving the user experience and business efficiency.

[0113] All of the above-mentioned text data includes key questions (calling the Q&A database). When it does not include key questions and is not for real-time data query, the RAG knowledge base is called. When it is for real-time data query, as Figure 4 shown, the virtual human control method based on the multi-modal large model of the present invention further includes the following steps:

[0114] S50: When it is determined according to the text data that the target Q&A is a real-time data query, the text data is processed through Prompt engineering to obtain query variables;

[0115] S51: Determine the target data interface according to the query variables to call the target database corresponding to the target data interface;

[0116] S52: Query the target answer text data matching the text data in the target database.

[0117] When it is determined according to the text data that the target Q&A is not a key question query but a real-time data query, the corresponding data interface is directly called to obtain the corresponding database, and the corresponding target answer text is searched in the database, and then the target answer text is organized by the LLM large language model. In another embodiment of the present invention, as Figure 5As shown, the virtual human control method based on the multimodal large model further includes the following steps:

[0118] S6: Convert the target answer text data into target answer audio and play the target answer audio.

[0119] Convert the target answer text data into target answer audio and play the target answer audio so that the user can hear the target answer corresponding to the target Q&A.

[0120] S7: Display the target answer text data and the target browser address.

[0121] It should be noted that the target answer text data is the initial data related to the target Q&A recorded in the target browser address. The target answer text data is obtained by processing the initial data, and the number of target browser addresses corresponding to a target Q&A may be multiple. Therefore, after the target browser address corresponding to the target answer text data is displayed, the user can open the target browser to obtain more knowledge related to the target Q&A.

[0122] Computer device

[0123] Based on the above embodiments, this embodiment provides a computer device, including a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of the method described in the above embodiments.

[0124] In some embodiments of this embodiment, a computer-readable storage medium is provided, on which a computer program is stored. It is characterized in that when the computer program is executed by a processor, the steps of the method described in the above embodiments are implemented.

[0125] In some embodiments of this embodiment, a computer program product is provided, including a computer program / instructions. It is characterized in that when the computer program is executed by a processor, the steps of the method described in the above embodiments are implemented.

[0126] The processor may include, but is not limited to, for example, one or more processors or microprocessors, etc. Each processor may be implemented by an Application Specific Integrated Circuit (ASIC), a Digital Signal Processor (DSP), a Digital Signal Processing Device (DSPD), a Programmable Logic Device (PLD), a Field Programmable Gate Array (FPGA), a controller, a microcontroller, a microprocessor, or other electronic components, and is used to execute the virtual human control method based on the multi-modal large model in the above embodiments.

[0127] The computer-readable storage medium may be implemented by any type of volatile or non-volatile storage device or a combination thereof. The computer-readable storage medium may include, but is not limited to, for example, Random Access Memory (RAM), Read-Only Memory (ROM), flash memory, EPROM memory, EEPROM memory, registers, computer storage media (such as hard disks, floppy disks, solid-state drives, removable disks, CD-ROMs, DVD-ROMs, Blu-ray discs, etc.).

[0128] The computer-readable storage medium may also store at least one computer-executable program / instructions, such as computer-readable instructions. The computer-readable storage medium includes, but is not limited to, for example, volatile memory and / or non-volatile memory. Volatile memory may include, for example, Random Access Memory (RAM) and / or cache memory, etc. The computer-readable storage medium may include, for example, Read-Only Memory (ROM), hard disks, flash memory, etc. For example, a non-transitory computer-readable storage medium may be connected to a computing device such as a computer. Then, when the computing device runs the computer-readable instructions stored on the computer-readable storage medium, the various virtual human control methods based on the multi-modal large model described above may be performed.

[0129] In addition, the computer device may also include (but is not limited to) a data bus, an Input / Output (I / O) bus, a display, and input / output devices (such as a keyboard, a mouse, a speaker, etc.).

[0130] The processor may communicate with external devices via the I / O bus through a wired or wireless network.

[0131] In one embodiment, the at least one computer-executable instruction may also be compiled into or comprise a software product / computer program product, wherein when one or more computer-executable instructions are run by a processor, the various functions and / or method steps in the embodiments described in this technology are executed.

[0132] In the embodiments provided in this disclosure, it should be understood that the disclosed devices and methods may also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of this disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code includes one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or actions, or may be implemented by a combination of dedicated hardware and computer instructions.

[0133] It should be noted that in this disclosure, the terms "include", "comprise", or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article, or device. Without further limitation, an element limited by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device including the element.

[0134] Although the embodiments disclosed in this disclosure are as above, the above content is only an embodiment adopted for the convenience of understanding this disclosure and is not intended to limit this disclosure. Any person skilled in the art within the technical field to which this disclosure pertains may make any modifications and changes in the form of implementation and details without departing from the spirit and scope disclosed in this disclosure. However, the scope of patent protection of this disclosure shall still be subject to the scope defined by the appended claims.

Claims

1. A virtual human control method based on a multimodal large model, characterized in that: include: Acquire a question and answer audio input by a user, the question and answer audio including a target question and answer, and convert the question and answer audio into text data; Performing data processing on the text data to determine a target interface corresponding to the target question and answer in the LLM large language model; When the target interface is a question and answer data interface, determining target answer text data corresponding to the target question and answer based on a question and answer database; When the target interface is a knowledge base question and answer interface, target answer text data corresponding to the target question and answer is determined based on the RAG knowledge base.

2. The virtual human control method according to claim 1, characterized in that: The processing of the text data to determine the target interface corresponding to the target question and answer in the LLM large language model includes: Based on the question and answer database, the database processing module queries the question and answer database for a target preset question and answer that matches the target question and answer; When a target preset question and answer matching the target question and answer is found in the question and answer database, determining to call the question and answer database interface in the LLM large language model as the target interface; Wherein, determining the target answer text data corresponding to the target question and answer based on the question and answer database includes: According to the target browser address of the target preset question and answer, initial answer text data corresponding to the target preset question and answer is obtained; The LLM large language model processes the initial answer text data based on the target question and answer, and determines the target answer text data corresponding to the target question and answer.

3. The virtual human control method according to claim 2, characterized in that: The virtual human control method further includes: The target answer text data and the target browser address are displayed.

4. The virtual human control method according to claim 2, characterized in that: The processing of the text data to determine the target interface corresponding to the target question and answer in the LLM large language model further includes: When no target preset question and answer matching the target question and answer is found in the question and answer database, determining whether the target question and answer is a real-time data query based on the text data; When it is determined according to the text data that the target question and answer is not a real-time data query, the knowledge base question and answer interface is determined as the target interface, and the text data is segmented to generate a plurality of question vectors; Wherein, determining the target answer text data corresponding to the text data based on the RAG knowledge base includes: The RAG knowledge base searches for a target vector matching the question vector in the RAG knowledge base according to the question vector, and converts the target vector into a target text; Constructing a context based on the target text and the text data corresponding to the question and answer audio; The LLM large language model generates target question-answer text data based on the context and multiple target texts.

5. The virtual human control method according to claim 4, characterized in that: The virtual human control method further includes: When it is determined according to the text data that the target question and answer is a real-time data query, the text data is processed by a prompt project to obtain a query variable; Determine a target data interface according to the query variable to call a target database corresponding to the target data interface; The target database is searched for target answer text data matching the text data.

6. The virtual human control method according to any one of claims 1 to 5, characterized in that: The virtual human control method further includes: The target answer text data is converted into a target answer audio, and the target answer audio is played.

7. A virtual human control device based on a multimodal large model, characterized in that: External network switches and internal network switches; A speech recognition module, the speech recognition module is used to convert the question and answer audio input by the user into text data, the question and answer audio includes the target question and answer input by the user, and the speech recognition module is connected to the external network switch; LLM large language model, the LLM large language model has a question-answering data interface, a knowledge base question-answering interface, and a data interface; A data processing module, the data processing module is used to perform data processing on the text data to determine to call the target interface corresponding to the target question and answer in the LLM large language model, wherein the data processing module is connected to the intranet switch; A question and answer database, wherein the question and answer database stores a plurality of preset questions and answers and browser addresses corresponding to the preset questions and answers, the question and answer database interface is connected to the question and answer database, the question and answer data module is connected to the intranet switch, and the question and answer database is used to determine the target answer text data corresponding to the target question and answer based on the question and answer database when the target interface is the question and answer data interface; A RAG knowledge base, wherein the RAG knowledge base stores preset knowledge and preset vectors corresponding to the preset knowledge, the knowledge base question and answer interface is connected to the RAG knowledge base, the RAG knowledge base is connected to the intranet switch, and the RAG knowledge base is used to determine the target answer text data corresponding to the target question and answer based on the RAG knowledge base when the target interface is the knowledge base question and answer interface.

8. The virtual human control device according to claim 7, characterized in that: Also includes: A playing module is communicatively connected to the LLM large language model, and is used for converting the target answer text data into a target answer audio and playing the target answer audio.

9. The virtual human control device according to claim 7, characterized in that: Also includes: An action instruction module, wherein the action instruction module is connected to the LLM large language model, and the action instruction module is connected to the intranet switch; as well as A Unity client, wherein the Unity client is connected to the action instruction module, and the Unity client is connected to the intranet switch.

10. The virtual human control device according to claim 7, characterized in that: Also includes: A display screen is connected to the intranet switch, and is used to display the target answer text data and the corresponding target browser address.