Digital human interaction method, server and storage medium

By generating image information and content expressions in digital human interaction, and combining knowledge base, question-and-answer library and large language model for multiple rounds of dialogue reasoning, the problem of incoherent interaction methods of digital humans is solved and the user experience is improved.

CN120353369APending Publication Date: 2025-07-22CHONGQING ZHONGKE YUNCONG TECH CO LTD +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510442640.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing digital human interaction methods lack fluency and coherence, and cannot effectively utilize user historical dialogue content for personalized interaction, resulting in poor user experience.

Method used

By responding to the user's selection operations on the interactive interface, the image information and content expression of digital people are generated, combined with the knowledge base, question-and-answer library and large language model, multiple rounds of dialogue reasoning are conducted, and historical answers are used to improve the logic and coherence of interaction.

Benefits of technology

The fluency and coherence of digital human interaction is realized, and the user experience of interaction with digital humans is improved. The generated digital humans can understand the logical relationship between the current question and historical answers, providing more logical answers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353369A_ABST
    Figure CN120353369A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, particularly provides a digital human interaction method, a server and a storage medium, and aims to solve the problem of how to improve the continuity of digital human interaction. In order to achieve the purpose, the method comprises the steps that in response to selection operation of a user on digital human image elements on an interactive interface, image information of a digital human is generated according to the selected image elements; in response to a selection operation of a user on the content elements of the digital person on the interactive interface, generating a content expression of the digital person according to the selected content elements; associating the image information with the content expression to generate a digital person; and interaction is carried out based on the digital human and an interactor, the interaction comprises calling a knowledge base or a question and answer library in content expression, and controlling a large language model to reasone the answer of the current question based on the called knowledge base or question and answer library and historical answers to obtain the answer of the current question. Based on the method, the answer logicality can be improved, and the interaction process is more coherent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly to a digital human interaction method, a server, and a storage medium. Background Art

[0002] A digital human is a digital human image, which is widely used in fields such as customer service, e-commerce, culture and tourism for user interaction. Currently, a major problem in using digital humans for interaction is that the interaction method is rigid, often in the form of question and answer, and it is impossible to achieve the fluency and naturalness similar to the conversation between people.

[0003] Specifically, the current digital human interaction method mainly pre-sets fixed questions and answers. When a user interacts with a digital human, it does not consider the historical conversation content between the user and the digital human, but only matches the user's current question with these fixed questions, and then obtains the successfully matched question and outputs the answer to that question. For example, if a user asks a question (hereinafter described as the first question), and then asks a related question (hereinafter described as the second question), when the digital human answers the second question, it does not consider the first question and its answer, that is, ignores the relevance between the first and second questions, but answers the first and second questions separately. In this case, the user may feel that the interaction ability of the digital human lags behind that of humans.

[0004] Correspondingly, a new technical solution is needed in this field to solve the above problems. Summary of the Invention

[0005] In order to overcome the above defects, this application is proposed to solve or at least partially solve the following technical problems: how to improve the fluency or coherence of digital human interaction, so as to improve the experience of users when interacting with digital humans.

[0006] In a first aspect, a digital human interaction method is provided, and the method includes:

[0007] Responding to a user's selection operation on a digital human image element on an interaction interface, generating image information of the digital human according to the selected image element;

[0008] Responding to a user's selection operation on a digital human content element on the interaction interface, generating a content expression of the digital human according to the selected content element, where the selected content element includes a knowledge base, a Q&A library, and a large language model, the knowledge base is used to store domain knowledge of a preset domain, and the Q&A library is used to store multiple preset Q&A pairs;

[0009] Associating the image information with the content expression to generate a digital human;

[0010] The digital human interacts with the interactant, and the interaction includes at least one round of conversation;

[0011] Among them, the interaction between the digital human and the interactant includes: invoking the knowledge base or Q&A library in the content expression, and controlling the large language model to reason about the answer to the current question based on the invoked knowledge base or Q&A library and the historical answers, so as to obtain the answer to the current question; the current question is the question raised by the interactant in the current round of conversation, and the historical answers are the answers to the questions raised by the interactant in multiple consecutive historical conversations before the current round of conversation.

[0012] In a technical solution of the above digital human interaction method, the large language model is a multimodal large language model, and the data processed by the multimodal large language model at least includes images and texts;

[0013] And / or, the invoking the knowledge base or Q&A library in the content expression includes:

[0014] Preferentially invoking the Q&A library, and determining whether to invoke the knowledge base according to the reasoning result of the large language model based on the Q&A library;

[0015] If the large language model reasons out the answer to the current question based on the Q&A library and the historical answers, then the knowledge base is no longer invoked; otherwise, the knowledge base is invoked.

[0016] In a technical solution of the above digital human interaction method, the large language model reasons about the answer to the current question based on the Q&A library and the historical answers in the following manner:

[0017] Determine the keywords of the current question according to the historical answers and the current question;

[0018] Match the keywords with the questions in each Q&A pair in the Q&A library;

[0019] According to the matching result, obtain the question that matches the keywords, and use the answer in the Q&A pair to which the question belongs as the answer to the current question.

[0020] In a technical solution of the above digital human interaction method, the Q&A library stores the Q&A pairs in vector form, and the large language model reasons about the answer to the current question based on the Q&A library and the historical answers in the following manner:

[0021] Determine the semantics of the current question according to the historical answers and the current question;

[0022] Perform vectorization processing on the semantics to obtain a semantic vector;

[0023] Semantically retrieve the questions of each Q&A pair in the Q&A library according to the semantic vector, obtain the question with the highest semantic similarity to the current question, and use the answer in the Q&A pair to which the question belongs as the answer to the current question.

[0024] In a technical solution of the above digital human interaction method, the method further includes updating the Q&A library in the content expression in the following manner:

[0025] Regularly obtain the interaction records generated after the digital human interacts with the interactant;

[0026] Use the large language model to extract Q&A pairs not stored in the Q&A library from the interaction records, and store the Q&A pairs in the Q&A library.

[0027] In a technical solution of the above digital human interaction method, the method further includes updating the Q&A library in the content expression in the following manner:

[0028] Obtain the answer evaluation information fed back by the interactant during the interaction process, where the answer evaluation information is used to indicate whether the answer provided by the digital human in the conversation is accurate;

[0029] Obtain the target answer according to the answer evaluation information, where the target answer is an inaccurate answer;

[0030] If the target answer is inferred by the large language model based on the Q&A library, update the target answer stored in the Q&A library.

[0031] In a technical solution of the above digital human interaction method, the selected content element further includes an interaction library, where the interaction library is used to store multiple questions and the interaction information corresponding to each question, and the interaction information includes at least one of the digital human's actions, images, displayed images, and displayed videos;

[0032] The interaction based on the digital human and the interactant further includes: obtaining the interaction information corresponding to the current question according to the interaction library in the content expression, and controlling the digital human to output the answer to the current question according to the interaction information.

[0033] In a technical solution of the above digital human interaction method, the method further includes:

[0034] In response to the user's configuration operation on the knowledge base in the interaction interface, obtain the first configuration information of the knowledge base, and create a knowledge base according to the first configuration information. The first configuration information includes a domain file of a preset domain, and the domain file is used to describe the domain knowledge of the preset domain;

[0035] In response to a user's configuration operation on the Q&A library in the interaction interface, obtain the second configuration information of the Q&A library, and create a Q&A library according to the second configuration information. The second configuration information includes a Q&A pair file, and the Q&A pair file is used to describe multiple preset Q&A pairs.

[0036] In response to a user's configuration operation on the interaction library in the interaction interface, obtain the third configuration information of the interaction library, and create an interaction library according to the third configuration information. The third configuration information includes multiple questions and corresponding interaction information for each question. The interaction information includes at least one of the digital human's actions, image, displayed image, and displayed video.

[0037] In a second aspect, a server is provided. The server includes at least one processor; and a memory communicatively connected to the at least one processor. Wherein, a computer program is stored in the memory, and when the computer program is executed by the at least one processor, the method described in any one of the technical solutions of the above digital human interaction method technical solution is implemented.

[0038] In a third aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores multiple program codes, and the program codes are adapted to be loaded and run by a processor to execute the method described in any one of the technical solutions of the above digital human interaction method technical solution.

[0039] One or more of the above technical solutions of the present application have at least one or more of the following Beneficial effects:

[0040] In a technical solution of implementing the digital human interaction method provided by the present application, it is possible to respond to a user's selection operation on the digital human image element in the interaction interface, generate the image information of the digital human according to the selected image element; respond to a user's selection operation on the digital human content element in the interaction interface, generate the content expression of the digital human according to the selected content element, and the selected content elements include a knowledge base, a Q&A library, and a large language model. The knowledge base is used to store domain knowledge in a preset domain, and the Q&A library is used to store multiple preset Q&A pairs; associate the image information with the content expression to generate a digital human; interact with the digital human and the interactant, and the interaction includes at least one round of dialogue; wherein, interacting with the digital human and the interactant includes: calling the knowledge base or Q&A library in the content expression, and controlling the large language model to infer the answer to the current question based on the called knowledge base or Q&A library and the historical answers, to obtain the answer to the current question; the current question is the question raised by the interactant in the current round of dialogue, and the historical answers are the answers to the questions raised by the interactant in multiple consecutive historical dialogues before the current round of dialogue.

[0041] Based on the above embodiments, users can customize the image information and content expression of the digital human, and the generated digital human can be understood as a digital human exclusive to the current user. In addition, when the digital human interacts with the interactant, the large language model simultaneously combines the question of the current round of conversation (i.e., the current question) and the answers of multiple consecutive rounds of historical conversations (i.e., historical answers) for reasoning. The combination of the current round of conversation and multiple consecutive rounds of historical conversations forms a coherent conversation process. In this way, when the large language model performs reasoning, it can not only understand the current question but also the logical relationship between the current question and the historical answers, improving the accuracy of reasoning, making the answer to the current question more logical, and thus enhancing the fluency or coherence of the digital human interaction. Brief Description of the Drawings

[0042] Referring to the accompanying drawings, the disclosure of the present application will become more understandable. It is easy for those skilled in the art to understand that these drawings are only for illustrative purposes and are not intended to limit the protection scope of the present application. Among them:

[0043] Figure 1 is a schematic flowchart of the main steps of a digital human interaction method according to an embodiment of the present application;

[0044] Figure 2 is a schematic diagram of an interaction interface for configuring a digital human image according to an embodiment of the present application;

[0045] Figure 3 is a schematic diagram of an interaction interface for digital human content expression according to an embodiment of the present application Figure 1 ;

[0046] Figure 4 is a schematic flowchart of the main steps of a large language model for reasoning the answer to the current question based on a Q&A library and historical answers according to an embodiment of the present application;

[0047] Figure 5 is a schematic flowchart of the main steps of a large language model for reasoning the answer to the current question based on a Q&A library and historical answers according to another embodiment of the present application;

[0048] Figure 6 is a schematic diagram of a conversation record according to an embodiment of the present application;

[0049] Figure 7 is a schematic diagram of an interaction interface for digital human content expression according to an embodiment of the present application Figure 2 ;

[0050] Figure 8 is a schematic diagram of a management interface for a knowledge base according to an embodiment of the present application;

[0051] Figure 9 It is a schematic diagram of the overall process of a digital human interaction method according to an embodiment of the present application;

[0052] Figure 10 It is a schematic diagram of the main structure of a server according to an embodiment of the present application;

[0053] Figure 11 It is a schematic diagram of the main structure of a digital human interaction system according to an embodiment of the present application;

[0054] Figure 12 It is a schematic diagram of the main structure of a digital human interaction system according to another embodiment of the present application.

[0055] Reference numerals:

[0056] 11: Memory; 12: Processor; 21: Digital human image module; 22: Digital human IP module; 23: Knowledge base management module; 24: Q&A library management module; 25: Interaction library management module; 26: Large model configuration module; 27: Pipeline module. Detailed implementation manners

[0057] The following describes some implementation manners of the present application with reference to the accompanying drawings. Those skilled in the art should understand that these implementation manners are only used to explain the technical principle of the present application and are not intended to limit the protection scope of the present application.

[0058] In the description of the present application, "module" and "processor" may include hardware, software, or a combination of both. A module may include a hardware circuit, various suitable sensors, communication ports, memory, and may also include a software part, such as program code, or a combination of software and hardware. The processor may be a central processing unit, a microprocessor, an image processor, a digital signal processor, or any other suitable processor. The processor has data and / or signal processing functions. The processor may be implemented in software, in hardware, or in a combination of both. The computer-readable storage medium includes any suitable medium for storing program code, such as magnetic disks, hard disks, optical disks, flash memories, read-only memories, random access memories, and so on. The term "A and / or B" represents all possible combinations of A and B, such as only A, only B, or A and B.

[0059] The following describes the embodiments of the digital human interaction method provided by the present application.

[0060] Refer to the attached Figure 1 , Figure 1 It is a schematic diagram of the main step process of a digital human interaction method according to an embodiment of the present application. As Figure 1 shown, the digital human interaction method in the embodiments of the present application mainly includes the following steps S101 to step S104.

[0061] Step S101: In response to a user's selection operation on the digital human image element on the interaction interface, generate the image information of the digital human according to the selected image element.

[0062] The digital human interaction method provided by this application can be applied to a server. On the display interface of the server, the above-mentioned interaction interface can be displayed. The interaction interface can display various image elements of the digital human, and each image element has multiple element information. The user can perform a selection operation on the interaction interface, select the image element to be configured, and select an element information for the image element. There can be controls for selection operations on the interaction interface, and the user can perform operations such as clicking and checking on the control to select the image element and the element information of the image element.

[0063] The image element can be understood as the type of the element, and the image information can be understood as the specific content of the element. For example, the image elements can include clothing, hairstyle, ear ornaments, facial ornaments (such as glasses), necklaces, socks, shoes, etc.; taking clothing and hairstyle as examples, the multiple element information of clothing is multiple different styles of clothing (such as clothing with different colors and styles), and the multiple element information of hairstyle is multiple different styles of hairstyles (such as hairstyles with different colors and styles).

[0064] In some embodiments, the user can also select the image type of the digital human on the interaction interface. The image types include 2D (two-dimensional) images and 3D (three-dimensional) images. Based on this, the digital human interaction method can also respond to the user's selection operation on the image type, and generate the image information of this image type according to the selected image type and the selected image element.

[0065] Refer to the appendix Figure 2 , Figure 2 exemplarily shows an interaction interface for configuring the digital human image. The interaction interface displays the image types (including 2D images and 3D images) and multiple image elements (including clothing, hairstyle, etc.). The user can select the image type and image elements on the interaction interface, and after selecting the image element, select the element information of the image element ( Figure 2 the element information is not shown), and the digital human interaction method can respond to these selection operations of the user, obtain the selected image type, image element and its element information, and generate the image information of the digital human according to these contents.

[0066] Step S102: In response to a user's selection operation on the digital human content element on the interaction interface, generate the content expression of the digital human according to the selected content element.

[0067] The types of content elements include knowledge bases, Q&A bases, and large language models (LLMs). There can be one or more knowledge bases, one or more Q&A bases, and one or more versions of the large language model. Users can select a knowledge base, a Q&A base, and a version of the large language model on the interaction interface. In this way, according to the selection operation, it can be determined that the selected content elements include a knowledge base, a Q&A base, and a large language model. Controls for selection operations can be set on the interaction interface, and users can perform operations such as clicking and checking on the controls to select each content element. In some embodiments, the version of the large language model can be represented by the engine address of the large language model.

[0068] The Q&A base is used to store multiple preset question-and-answer pairs (Question&Answer).

[0069] The knowledge base is used to store domain knowledge of a preset domain, and different knowledge bases are respectively used to store domain knowledge of different domains. The domain can also be understood as an industry, and the domain can include customer service, e-commerce, culture and tourism, etc. Domain knowledge is the knowledge information of the domain, and the knowledge information is pre-stored in the knowledge base. The knowledge base can store domain knowledge in the form of semantic vectors to facilitate the large language model in understanding, memorizing, and retrieving the domain knowledge in the knowledge base in subsequent step S103. For example, if the preset domain is the culture and tourism domain of Conghua District, Guangzhou City, the domain knowledge includes the cultural information and tourism information of Conghua District, Guangzhou City. The carriers of domain knowledge can include documents, pictures, etc. The documents can include PDF, Word, Excel, TXT, etc. The digital human interaction method provided in this application can extract domain knowledge from the above carriers, and then convert the domain knowledge into semantic vectors and store them in the knowledge base. For example, by using OCR technology to parse scanned documents and extract readable text content, and adopting the way of knowledge embedding to convert the readable text content into semantic vectors.

[0070] Refer to the appendix Figure 3 , Figure 3 An interaction interface for content expression is exemplarily shown. This interaction interface displays multiple knowledge bases, multiple Q&A bases, and versions of the large language model (i.e., Figure 3 the large model in Figure 3 ), namely Figure 3 LLM-1 in

[0071] In some embodiments, the large language model is a multimodal large language model, and the processing data of the multimodal large language model includes at least images and text. When the processing data of the multimodal large language model is images and text, the multimodal large language model can also be understood as a graphic and text understanding large model, that is, it can process (or understand) images and text simultaneously.

[0072] Step S103: Associate the image information with the content expression to generate a digital human.

[0073] The image information can be understood as the external image of the digital human, and the content expression can be understood as the internal image of the digital human. Associating the internal and external images forms a complete digital human.

[0074] Step S104: Interact with the digital human based on the digital human and the interactant. The interaction includes at least one round of conversation. Among them, interacting with the digital human based on the digital human and the interactant includes the following steps:

[0075] Call the knowledge base or Q&A library in the content expression, and control the large language model to reason about the answer to the current question based on the called knowledge base or Q&A library and the historical answers, so as to obtain the answer to the current question. The current question is the question raised by the interactant in the current round of conversation, and the historical answers are the answers to the questions raised by the interactant in the consecutive multiple rounds of historical conversations before the current round of conversation. In this embodiment, the digital human can output the question in various ways such as text, voice, image, video, etc. according to the answer to the current question.

[0076] If the current round of conversation is the first round of conversation, the historical answer is set to be empty, that is, the answer to the current question is only reasoned based on the called knowledge base or Q&A library.

[0077] The more the number of rounds of consecutive multiple rounds of historical conversations, the more historical answers, that is, the more input data for the large language model. When the input data increases, the reasoning accuracy of the large language model will be improved, but the reasoning efficiency may decrease. Therefore, when determining the number of rounds, the accuracy and efficiency of the large language model during reasoning can be tested in advance, and a number of rounds that can balance accuracy and efficiency as much as possible can be determined according to the test results, which can not only ensure the reasoning accuracy but also not significantly reduce the reasoning efficiency. For example, in some embodiments, the number of rounds is 5 rounds.

[0078] For example, Table 1 below exemplarily shows 3 rounds of conversations, where the 1st and 2nd rounds of conversations are historical rounds of conversations, and the 3rd round of conversation is the current round of conversation. Table 1

[0079] Based on the method described in the above steps S101 to S104, the user can personalized set the image information and content expression of the digital human, and the generated digital human can be understood as a digital human exclusive to the current user. In addition, when the digital human interacts with the interactant, the large language model combines the question of the current round of conversation and the answers of consecutive multiple rounds of historical conversations for reasoning. The combination of the current round of conversation and consecutive multiple rounds of historical conversations is a coherent conversation process. In this way, when the large language model conducts reasoning, it can not only understand the current question, but also understand the logical relationship between the current question and the historical answers, improving the accuracy of reasoning, making the answer to the current question more logical, and thus improving the fluency or coherence of the digital human interaction.

[0080] For example, in some application scenarios, a digital human can be generated, and the name of this digital human is "Conghua Xiaohua", an intelligent tour guide in Conghua, Guangdong. Among them, the image information is an AI tour guide wearing traditional Lingnan clothing (which can also be switched to an ancient scholar, a modern guide, etc.). The knowledge base in the content expression contains knowledge information such as the history and culture, hot spring health preservation, ecological tourism, and food recommendations of Conghua. "Conghua Xiaohua" can answer the questions of the interactant in various ways such as text, voice, image, and video. For example, Table 2 below exemplarily shows the ways in which "Conghua Xiaohua" replies to questions in 3 conversations. Table 2

[0081] Next, the embodiments of the digital human interaction method provided by the present application will be further described, specifically for the above step S104.

[0082] In some implementation manners of the above step S104, when the digital human interacts with the interactant, the knowledge base or Q&A library in the content expression can be called in the following way:

[0083] First, call the Q&A library in the content expression, obtain the reasoning result obtained by the large language model based on the Q&A library for reasoning, and determine whether to call the knowledge base in the content expression according to the reasoning result; if the large language model reasons out the answer to the current question based on the Q&A library and the historical answers, then the knowledge base will not be called; otherwise, the knowledge base will be called. In addition, after calling the knowledge base, the large language model will continue to be controlled to reason based on the knowledge base and the historical answers to obtain the answer to the current question.

[0084] Multiple preset question-and-answer pairs stored in the question-and-answer library, that is, some fixed questions and their answers are stored. The large language model performing reasoning based on the question-and-answer library can be understood as matching the question raised by the interactant with the preset questions in the question-and-answer library to determine whether the question raised by the interactant is a preset question in the question-and-answer library. If it is a preset question, the answer to the preset question is directly output, so that the answer can be quickly obtained through question matching; if it is not a preset question, the knowledge library is called to control the large language model to perform reasoning based on the domain knowledge in the knowledge library to obtain the answer.

[0085] The question-and-answer pairs stored in the question-and-answer library can be preset by the user in advance. In this embodiment, the content of the question-and-answer pairs is not specifically limited. For example, when setting the question-and-answer pairs, the user can determine some questions that the interactant may often ask, and then form question-and-answer pairs based on the questions and their answers.

[0086] In some embodiments of the above step S104, when interacting with the interactant based on the digital human, the large language model can Figure 4 perform reasoning on the answer to the current question based on the question-and-answer library and the historical answer through the following steps S1041 to S1043 shown.

[0087] Step S1041: Determine the keywords of the current question according to the historical answer and the current question. There are multiple keywords. Some keywords are words from the current question itself, and some keywords are words from the historical answer. At the same time, obtaining keywords according to the historical answer and the current question can more comprehensively and completely represent the meaning of the current question.

[0088] Taking Table 1 in the previous embodiment as an example, if the current question is "What is the ticket price?", and the historical answer is "The business hours of Scenic Spot A are XXX", the keyword determined according to the current question is "ticket price", and the keywords determined according to the historical answer can include "Scenic Spot A", so the final keywords of the current question are "Scenic Spot A, ticket price".

[0089] Step S1042: Match the keywords with the questions of each question-and-answer pair in the question-and-answer library.

[0090] Specifically, a conventional keyword matching method can be adopted to match the keywords of the current question with the questions in the Q&A pairs and obtain a matching degree. The greater the matching degree, the more similar the meaning of the current question is to the questions in the Q&A pairs. For this, a matching threshold can be set. If the matching degree between the keyword and the question is greater than this matching threshold, it is determined that the question is successfully matched with the keyword. When the matching degrees of multiple questions are all greater than the matching threshold at the same time, the question with the largest matching degree is used as the question successfully matched with the keyword. In some embodiments, when setting the value of the matching threshold, those skilled in the art can match the keywords of a large number of different test questions with the questions in the Q&A pairs and obtain the matching degree. The test questions have the same meaning as the questions in the Q&A pairs, and then obtain the minimum matching degree, and set the value of the matching threshold according to this minimum matching degree.

[0091] Step S1043: According to the matching result, obtain the question successfully matched with the keyword, and use the answer in the Q&A pair to which the question belongs as the answer to the current question.

[0092] Based on the method described in the above steps S1041 to S1043, the current question keywords can be obtained more accurately by using the current question and historical answers, which is beneficial to matching the accurate question answer from the Q&A library according to the keywords.

[0093] In some embodiments of the above step S104, the Q&A library stores the Q&A pairs in vector form. Specifically, an embedding model can be used to convert the questions and answers in the Q&A pairs into semantic vectors respectively. The Q&A library can use a vector database (such as databases like FAISS, Milvus, Chroma, etc.) to store the Q&A pairs converted into semantic vectors. Based on this, in some embodiments, when interacting between the digital human and the interactant, the large language model can Figure 5 perform reasoning on the answer to the current question based on the Q&A library and historical answers through the following steps S1044 to S1046 as shown.

[0094] Step S1044: Determine the semantics of the current question according to the historical answer and the current question. Specifically, the keywords of the current question can be determined according to the historical answer and the current question, and then the semantics of the current question can be determined according to the keywords. Among them, the method for obtaining keywords is similar to the method in step S1041 in the foregoing embodiments. For example, if the final keywords of the current question are "Scenic Spot A, ticket price", then the semantics of the current question is "What is the ticket price of Scenic Spot A?".

[0095] Step S1045: Perform vectorization processing on the semantics to obtain a semantic vector.

[0096] Specifically, conventional vectorization processing methods can be adopted for processing. For example, a vectorization model is used to convert semantics into semantic vectors.

[0097] Step S1046: Semantically retrieve the questions of each Q&A pair in the Q&A library based on the semantic vector, obtain the question with the highest semantic similarity to the current question, and use the answer in the Q&A pair to which the question with the highest similarity belongs as the answer to the current question.

[0098] Specifically, conventional semantic retrieval methods can be adopted to obtain the semantic similarity between the semantic vector and the questions of each Q&A pair.

[0099] Based on the method described in the above steps S1044 to S1046, the large language model can quickly match the accurate question answer from the Q&A library through semantic retrieval.

[0100] Next, the embodiments of the digital human interaction method provided in this application will be further described.

[0101] In some embodiments of this application, the Q&A library in the digital human content expression can be updated through the following steps 11 to 12:

[0102] Step 11: Regularly obtain the interaction records generated after the digital human interacts with the interactors. The interaction records can record the conversation records of each round in the interaction process.

[0103] Refer to the appendix Figure 6 , Figure 6 Exemplarily shows the conversation record of one round of conversation. As Figure 6 shown, the conversation record includes conversation content, conversation information, and user information.

[0104] The conversation content includes the questions raised by the interactors, the answers replied by the digital human, and the conversation start time, and at the same time supports voice broadcast of the conversation content.

[0105] The conversation information includes conversation serial number, conversation name, application name, model name, and model version, etc. The conversation serial number is the conversation round number, the conversation name is the question raised by the interactor, the application name can be an alternative name for the digital human interaction method provided in this application (such as intelligent knowledge Q&A), and the model name and model version are the name and version of the large language model in the content expression respectively.

[0106] The user information includes account name, user name, user unit, user occupation, mobile phone number, etc. When the server is used to execute the digital human interaction method provided in this application, the digital human interaction method is an application program of the server. When the user creates a digital human through the server, an account needs to be established to log in to this application program, and the account name is the name of this account.

[0107] Step 12: Use a large language model to extract question-and-answer pairs from the interaction record that are not stored in the question-and-answer library, and store the question-and-answer pairs in the question-and-answer library.

[0108] Specifically, the large language model can extract multiple question-and-answer pairs from the interaction record, then query the question-and-answer library based on the questions of the question-and-answer pairs to determine which questions are not stored in the question-and-answer library, and then store the question-and-answer pairs corresponding to the questions in the question-and-answer library. If the question-and-answer library stores question-and-answer pairs in vector form, the question-and-answer pairs need to be converted into vectors first and then stored in the question-and-answer library.

[0109] Based on the method described in the above Step 11 to Step 12, new question-and-answer pairs can be added to the question-and-answer library, gradually enriching the question-and-answer library, thereby improving the interaction ability of the digital human.

[0110] In some embodiments of the present application, the question-and-answer library in the digital human content expression can also be updated through the following Step 21 to Step 22:

[0111] Step 21: Obtain the answer evaluation information fed back by the interactant during the interaction, and the answer evaluation information is used to indicate whether the answer provided by the digital human in the conversation is accurate.

[0112] For example, when the digital human replies with an answer, an operation area for answer evaluation can be displayed on the interaction interface. The operation area can include a control indicating that the answer is accurate. When the interactant clicks on this control or performs other operations, it will feedback the answer evaluation information indicating that the answer is accurate; the operation area can also include a control indicating that the answer is inaccurate. When the interactant clicks on this control or performs other operations, it will feedback the answer evaluation information indicating that the answer is inaccurate; the operation area can also include an answer input field, where the interactant can input the answer that the interactant himself / herself thinks is correct. Based on this, the answer evaluation information will also include the answer input by the interactant.

[0113] Step 22: Obtain the target answer according to the answer evaluation information, and the target answer is an inaccurate answer; if the target answer is inferred by the large language model based on the question-and-answer library, update the target answer stored in the question-and-answer library.

[0114] Specifically, the update prompt information of the target answer can be output according to the answer evaluation information, so that the user can update the target answer according to the update prompt information. The digital human interaction method provided in the present application can respond to the user's update operation and replace the target answer with the answer indicated by the update operation. If the answer evaluation information includes the answer input by the interactant, the update prompt information can also include this answer, which is convenient for the user to quickly confirm whether the answer is correct and complete the update.

[0115] Based on the method described in the above steps 21 to 22, it is possible to confirm whether there are incorrect answers in the Q&A library according to the feedback of the interactors and correct them in a timely manner.

[0116] Next, the embodiments of the digital human interaction method provided in this application will be further described.

[0117] In some embodiments of this application, the types of content elements further include interaction libraries, and there can be one or more interaction libraries. The interaction libraries are used to store multiple questions and the corresponding interaction information for each question. The interaction information includes at least one of the digital human's actions, image, the images shown by the digital human, and the videos shown by the digital human. In different interaction libraries, the included questions may be different, and the interaction information corresponding to the questions may also be different. In addition to selecting a knowledge library, a Q&A library, and the version of the large language model on the interaction interface, the user can also select an interaction library. In this way, according to the selection operation, it can be determined that the selected content elements include a knowledge library, a Q&A library, an interaction library, and a large language model. As Figure 7 shown, the user selects the knowledge library 5, the Q&A library 5, the interaction library 5, and the large language model with the version of LLM-1 on the interaction interface.

[0118] Based on this, in some implementation manners of the foregoing step S104, when interacting with the interactor based on the digital human, the interaction information corresponding to the current question can also be obtained according to the interaction library in the content expression, and the digital human can be controlled to output the answer to the current question according to the interaction information. For example, the interaction information includes the actions of the digital human corresponding to the answer to the current question. While controlling the digital human to output the answer to the current question, the digital human is controlled to execute this action.

[0119] Based on the above implementation manners, the personalization degree of the digital human can be further improved, and at the same time, the fun or experience of the interactor interacting with the digital human can also be improved.

[0120] Next, the embodiments of the digital human interaction method provided in this application will be further described, specifically, the creation methods of the knowledge library, the Q&A library, and the interaction library will be described.

[0121] 1. Describe the creation method of the knowledge library.

[0122] Specifically, in response to a user's configuration operation on the knowledge base in the interaction interface, the first configuration information of the knowledge base can be obtained, and a knowledge base can be created according to the first configuration information. The first configuration information includes at least a domain file of a preset domain, and the domain file is used to describe the domain knowledge of the preset domain. The format of the domain file can be PDF, Word, Excel, TXT, etc. In addition, when creating the knowledge base, the domain knowledge can be extracted from the domain file, and then the domain knowledge can be converted into semantic vectors and stored in the knowledge base. For example, parse the PDF through OCR technology, extract the readable text content, and then convert the readable text content into semantic vectors.

[0123] See the appendix Figure 8 , Figure 8 Exemplarily shows a management interface of a knowledge base, and the management interface includes knowledge base information and knowledge base details.

[0124] The knowledge base information includes the knowledge base name, knowledge base label, number of knowledge items, start / stop status, knowledge base description, creation time, etc. The knowledge base label is a label used to mark or annotate the knowledge base, the number of knowledge items is the number of domain files, and the start / stop status includes enabled and disabled.

[0125] The knowledge base details include a control for inputting the file name, a control for adding new files, a control for batch operations (such as deletion) on domain files, and also include a list of all domain files. The list includes file number, file name, file format, creation time, and operations (including showing the details of the domain file, editing the domain file, deleting the domain file).

[0126] 2. Explain the creation method of the Q&A library.

[0127] Specifically, in response to a user's configuration operation on the Q&A library in the interaction interface, the second configuration information of the Q&A library can be obtained, and a Q&A library can be created according to the second configuration information. The second configuration information includes at least Q&A pair files, and the Q&A pair files are used to describe multiple preset Q&A pairs. The format of the Q&A pair files can also be PDF, Word, Excel, TXT, etc. When creating the Q&A library, the Q&A pairs can be extracted from the Q&A pair files, and then the Q&A pairs can be converted into semantic vectors and stored in the Q&A library. For example, parse the PDF through OCR technology, extract the readable text content, and then convert the readable text content into semantic vectors.

[0128] 3. Explain the creation method of the interaction library.

[0129] Specifically, in response to the user's configuration operation on the interaction library in the interaction interface, the third configuration information of the interaction library can be obtained, and an interaction library can be created according to the third configuration information. The third configuration information includes multiple questions and the corresponding interaction information for each question. The interaction information includes at least one of the digital human's actions, image, displayed image, and displayed video.

[0130] The interaction interface can display a variety of different actions, and the user can perform a selection operation on the interaction interface to select an action; the method for setting the image is the same as the method described in step S101 of the foregoing embodiment, and will not be elaborated here; the user can perform a data upload operation on the interaction interface, upload the images and videos to be displayed to the server (the server that executes the digital human interaction method), and add them to the third configuration information.

[0131] Based on the above methods for creating the knowledge base, Q&A library, and interaction library, users can flexibly create different knowledge bases, Q&A libraries, and interaction libraries according to their actual needs.

[0132] Next, the embodiments of the digital human interaction method provided by the present application will be further described.

[0133] Refer to the attached Figure 9 , Figure 9 which exemplarily shows the overall process of the digital human interaction method according to an embodiment of the present application. As Figure 9 shown, the digital human interaction method includes the following steps S201 to S204. Step S201: Generate a digital human image. Specifically, the method described in step S101 of the foregoing embodiment can be used to generate the image information of the digital human. Step S202: Configure the content expression of the digital human. Specifically, the relevant methods in the foregoing method embodiments can be used to configure the knowledge base, Q&A library, interaction library, and large language model in the digital human content expression. Step S203: Generate a digital human. Specifically, the image information and the content expression are associated to generate a digital human. Step S204: Perform an interaction based on the digital human. Specifically, the method described in step S104 of the foregoing method embodiments can be used to control the digital human to interact with the interactors. Similar to the foregoing method embodiments, based on the methods described in the above steps S201 to S204, the personalized setting of the digital human can also be achieved, and the fluency or coherence of the digital human interaction can be improved.

[0134] It should be noted that although the above embodiments describe the various steps in a specific order, those skilled in the art can understand that in order to achieve the effects of the present application, it is not necessary for different steps to be executed in such an order. They can be executed simultaneously (in parallel) or in other orders, and these adjusted solutions are equivalent technical solutions to the technical solutions described in the present application, and thus will also fall within the protection scope of the present application.

[0135] Those skilled in the art can understand that all or part of the processes in the methods of the above-mentioned embodiments of the present application can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable storage medium can include: any entity or device, medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory, random access memory, electrical carrier signal, telecommunication signal, and software distribution medium, etc., that can carry the computer program code.

[0136] On the other hand, the present application also provides a computer-readable storage medium.

[0137] In an embodiment of a computer-readable storage medium according to the present application, the computer-readable storage medium can be configured to store a program for executing the digital human interaction method of the above-mentioned method embodiment. The program can be loaded and run by a processor to implement the above-mentioned digital human interaction method. For the sake of convenience of description, only the parts related to the embodiments of the present application are shown. For the specific technical details not disclosed, please refer to the method part of the embodiments of the present application. The computer-readable storage medium can be a storage device formed by various electronic devices. Optionally, the computer-readable storage medium in the embodiments of the present application is a non-transitory computer-readable storage medium.

[0138] On the other hand, the present application also provides a server.

[0139] In an embodiment of a server according to the present application, the server can include at least one processor; and a memory communicatively connected to at least one processor; wherein, a computer program is stored in the memory, and when the computer program is executed by at least one processor, the method described in any of the above-mentioned embodiments is implemented. Refer to the attached Figure 10 , Figure 10 It is exemplarily shown in the figure that the memory 11 and the processor 12 are communicatively connected through a bus.

[0140] On the other hand, the present application also provides a digital human interaction system.

[0141] Refer to the attached Figure 11 , Figure 11 It is a schematic diagram of the main structure of a digital human interaction system according to an embodiment of the present application. As Figure 11 shown, the digital human interaction system can include a digital human image module 21 and a digital human IP module 22.

[0142] The digital human image module 21 can be configured to: in response to a user's selection operation on a digital human image element on the interaction interface, generate the image information of the digital human according to the selected image element. The digital human IP module 22 can be configured to: in response to a user's selection operation on a digital human content element on the interaction interface, generate the content expression of the digital human according to the selected content element, where the selected content elements include a knowledge base, a Q&A library, and a large language model, the knowledge base is used to store domain knowledge in a preset domain, and the Q&A library is used to store multiple preset Q&A pairs; associate the image information with the content expression to generate a digital human.

[0143] The digital human interaction system can interact with an interactant based on the digital human, and the interaction includes at least one round of conversation. Among them, interacting with the interactant based on the digital human includes: invoking the knowledge base or Q&A library in the content expression, and controlling the large language model to reason about the answer to the current question based on the invoked knowledge base or Q&A library and historical answers, to obtain the answer to the current question; the current question is the question raised by the interactant in the current round of conversation, and the historical answers are the answers to the questions raised by the interactant in multiple consecutive historical conversations before the current round of conversation.

[0144] For the description of the specific implementation functions of the above digital human image module 21, digital human IP module 22, and digital human interaction system, reference can be made to steps S101 to S104 in the foregoing method embodiment.

[0145] Refer to the appendix Figure 12 In some embodiments, the digital human interaction system may further include a knowledge base management module 23, a Q&A library management module 24, an interaction library management module 25, a large model configuration module 26, and a pipeline module 27.

[0146] The knowledge base management module 23 can be configured to: in response to a user's configuration operation on the knowledge base in the interaction interface, obtain first configuration information of the knowledge base, and create a knowledge base according to the first configuration information. The first configuration information at least includes a domain file of a preset domain, and the domain file is used to describe the domain knowledge of the preset domain. The Q&A library management module 24 can be configured to: in response to a user's configuration operation on the Q&A library in the interaction interface, obtain second configuration information of the Q&A library, and create a Q&A library according to the second configuration information. The second configuration information at least includes a Q&A pair file, and the Q&A pair file is used to describe multiple preset Q&A pairs. The interaction library management module 25 can be configured to: in response to a user's configuration operation on the interaction library in the interaction interface, obtain third configuration information of the interaction library, and create an interaction library according to the third configuration information. The third configuration information includes multiple questions and corresponding interaction information for each question, and the interaction information includes at least one of the actions, images, displayed images, and displayed videos of the digital human. For the description of the specific functions implemented by the above modules, reference can be made to the creation methods of the knowledge base, Q&A library, and interaction library in the foregoing method embodiments.

[0147] The large model configuration module 26 can be configured to manage different versions of large language models. The pipeline module 27 can be configured to manage the process time consumption, input, and output during the digital human interaction process, facilitating subsequent viewing and troubleshooting by the user.

[0148] The above digital human interaction system is used to execute Figure 1-9 the digital human interaction method embodiments shown. The technical principles, technical problems solved, and technical effects produced by the two are similar. Those skilled in the art of this technology can clearly understand that for the convenience and brevity of description, the specific working process and related descriptions of the digital human interaction system can refer to the content described in the embodiments of the digital human interaction method, which will not be elaborated here.

[0149] So far, the technical solution of the present application has been described in conjunction with an embodiment shown in the drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present application is obviously not limited to these specific embodiments. Without departing from the principle of the present application, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the protection scope of the present application.

Claims

1. A digital human interaction method, characterized in that, The method includes: In response to a user's selection operation on a digital human image element on the interaction interface, generating the image information of the digital human according to the selected image element; In response to a user's selection operation on a digital human content element on the interaction interface, generating the content expression of the digital human according to the selected content element, where the selected content elements include a knowledge base, a Q&A library, and a large language model, the knowledge base is used to store domain knowledge in a preset domain, and the Q&A library is used to store multiple preset Q&A pairs; Associating the image information with the content expression to generate a digital human; Based on the digital human, interacting with an interactant, where the interaction includes at least one round of conversation; Wherein, The interacting based on the digital human and the interactant includes: calling the knowledge base or the Q&A library in the content expression, and controlling the large language model to reason about the answer to the current question based on the called knowledge base or Q&A library and the historical answers to obtain the answer to the current question; The current question is the question raised by the interactant in the current round of conversation, and the historical answers are the answers to the questions raised by the interactant in multiple consecutive historical conversations before the current round of conversation.

2. The method according to claim 1, wherein The large language model is a multimodal large language model, and the data processed by the multimodal large language model at least includes images and texts; And / or, The calling the knowledge base or the Q&A library in the content expression includes: Preferentially calling the Q&A library, and determining whether to call the knowledge base according to the reasoning result of the large language model based on the Q&A library; If the large language model reasons about the answer to the current question based on the Q&A library and the historical answers, the knowledge base is no longer called; otherwise, the knowledge base is called.

3. The method according to claim 1 or 2, characterized in that, The large language model reasons about the answer to the current question based on the Q&A library and the historical answers in the following manner: Determining the keywords of the current question according to the historical answers and the current question; Matching the keywords with the questions in each Q&A pair in the Q&A library; According to the matching result, obtaining the question that matches the keywords successfully, and using the answer in the Q&A pair to which the question belongs as the answer to the current question.

4. The method according to claim 1 or 2, characterized in that, The Q&A library stores the Q&A pairs in vector form, and the large language model reasons about the answer to the current question based on the Q&A library and the historical answers in the following manner: Determining the semantics of the current question according to the historical answers and the current question; Performing vectorization processing on the semantics to obtain a semantic vector; Performing semantic retrieval on the questions in each Q&A pair in the Q&A library according to the semantic vector, obtaining the question with the highest semantic similarity to the current question, and using the answer in the Q&A pair to which the question belongs as the answer to the current question.

5. The method according to claim 1, wherein The method further includes updating the Q&A library in the content expression in the following manner: Regularly obtaining the interaction records generated after the digital human interacts with the interactant; Using the large language model to extract the Q&A pairs not stored in the Q&A library from the interaction records, and storing the Q&A pairs in the Q&A library.

6. The method according to claim 1 or 5, characterized in that The method further includes updating the Q&A library in the content expression in the following manner: Obtain answer evaluation information fed back by the interactant during the interaction, where the answer evaluation information is used to indicate whether the answer provided by the digital human in the conversation is accurate; Obtain a target answer according to the answer evaluation information, where the target answer is an inaccurate answer; If the target answer is inferred by the large language model based on the Q&A library, update the target answer stored in the Q&A library.

7. The method according to claim 1, wherein The selected content element further includes an interaction library, which is used to store multiple questions and the corresponding interaction information for each question, and the interaction information includes at least one of the actions, images, displayed images, and displayed videos of the digital human; Based on the interaction between the digital human and the interactant, it further includes: obtaining the interaction information corresponding to the current question according to the interaction library in the content expression, and controlling the digital human to output the answer to the current question according to the interaction information.

8. The method according to claim 1 or 7, characterized in that, The method further includes: In response to the user's configuration operation on the knowledge base on the interaction interface, obtain the first configuration information of the knowledge base, and create a knowledge base according to the first configuration information. The first configuration information includes a domain file of a preset domain, and the domain file is used to describe the domain knowledge of the preset domain; In response to the user's configuration operation on the Q&A library on the interaction interface, obtain the second configuration information of the Q&A library, and create a Q&A library according to the second configuration information. The second configuration information includes a Q&A pair file, and the Q&A pair file is used to describe multiple preset Q&A pairs; In response to the user's configuration operation on the interaction library on the interaction interface, obtain the third configuration information of the interaction library, and create an interaction library according to the third configuration information. The third configuration information includes multiple questions and the corresponding interaction information for each question, and the interaction information includes at least one of the actions, images, displayed images, and displayed videos of the digital human.

9. A server, characterized in that, Includes: At least one processor; And a memory communicatively connected to the at least one processor; Wherein, a computer program is stored in the memory, and when the computer program is executed by the at least one processor, it implements the digital human interaction method according to any one of claims 1 to 8.

10. A computer-readable storage medium storing multiple program codes, characterized in that, The program code is suitable for being loaded and run by a processor to execute the digital human interaction method according to any one of claims 1 to 8.

Citation Information

Cited By

  • Question and answer method and related equipment

    CN121561026A

  • Vehicle-mounted digital human voice interaction system and method based on unreal engine

    CN121617389A