Interaction method and apparatus, and computer-readable storage medium

By receiving and processing user queries, using machine learning models and Large Language Models (LLM) to judge and generate responses, and combining knowledge base retrieval, the system solves the problems of incomplete information display and poor interactivity in enterprise service interaction systems, achieving efficient and intelligent human-computer interaction.

WO2026091011A1PCT designated stage Publication Date: 2026-05-07BOE TECHNOLOGY GROUP CO LTD
View PDF 11 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
BOE TECHNOLOGY GROUP CO LTD
Filing Date
2024-10-31
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Existing enterprise service interaction systems suffer from incomplete information display, poor interactivity, and inability to update in real time, especially in terms of human-computer interaction.

Method used

By receiving input question information, a machine learning model is used to determine the type and matching degree of the question information. Combining the first information set and the second information set, corresponding response information is generated, including voice response and multimedia content. The Large Language Model (LLM) is used to classify, rewrite and enhance the retrieval of the question information. Finally, the response is generated by combining the retrieval results of the knowledge base.

Benefits of technology

It improves the accuracy and efficiency of human-computer interaction, enhances information security, provides natural language interaction and intelligent question-and-answer functions, and improves the intelligence and real-time update capabilities of the interaction system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024129037_07052026_PF_FP_ABST
    Figure CN2024129037_07052026_PF_FP_ABST
Patent Text Reader

Abstract

Provided are an interaction method and apparatus, and a computer-readable storage medium. The interaction method comprises: receiving input first question information; in response to the first question information belonging to a specified information type and matching a first information set, outputting reply information corresponding to the first question information; and in response to the first question information not matching the first information set but matching a second information set, outputting answer rejection information.
Need to check novelty before this filing date? Find Prior Art

Description

Interaction methods, devices and computer-readable storage media Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to an interaction method, an interaction device, a computer-readable storage medium, and a computer program product. Background Technology

[0002] Enterprise services, such as brand promotion, business consulting, knowledge management, and employee training, often rely on manual services or simple interactive systems. This approach suffers from problems such as incomplete information display, poor interactivity, and inability to update in real time.

[0003] In recent years, artificial intelligence (AI) technology has made significant progress, particularly in areas such as large-scale models, speech and semantics, and digital humans. Among these technologies, the introduction of advanced AI techniques to build interactive systems enables functions such as natural language interaction, intelligent question answering, and virtual demonstrations.

[0004] Summary of the Invention

[0005] According to some embodiments of this disclosure, an interaction method is provided, including: receiving input first question information; responding to the first question information belonging to a specified information type and matching a first information set, outputting response information corresponding to the first question information; and responding to the first question information not matching the first information set but matching a second information set, outputting a rejection message.

[0006] In some embodiments, the first query information belonging to a specified information type includes a specified person entity and / or a specified organization entity.

[0007] In some embodiments, the corresponding response information is generated using a machine learning model based on the retrieval results of the first question information in the knowledge base.

[0008] In some embodiments, the search results are determined based on a first candidate search result and a second candidate search result in the knowledge base. The first candidate search result is determined based on the similarity between the feature vector of the first query information and the feature vector of the information in the knowledge base. The second candidate search result is determined based on the term frequency and inverse document frequency of the first query information in the knowledge base.

[0009] In some embodiments, the search results are determined by filtering the first and second candidate search results by keywords and then reordering them.

[0010] In some embodiments, the interaction method further includes: in response to a mismatch between the first question information and both the first information set and the second information set, outputting corresponding response information, wherein the corresponding response information is generated using a machine learning model.

[0011] In some embodiments, outputting a rejection message includes: in response to the first question information belonging to a specified information type, and the first question information not matching the first information set but matching the second information set, outputting a rejection message; and / or, outputting a corresponding response message includes: in response to the first question information belonging to a specified information type, and the first question information not matching either the first information set or the second information set, outputting a corresponding response message.

[0012] In some embodiments, the interaction method further includes: in response to a first question information matching preset topic information, outputting multimedia content corresponding to the preset topic information; and / or, in response to a first question information being related to a designated organization, outputting corresponding response information.

[0013] In some embodiments, the corresponding multimedia content is determined based on topic tags, which are determined using a machine learning model based on the first question information.

[0014] In some embodiments, the information type to which the first question information belongs is determined based on the second question information, and the second question information is generated by using a machine learning model to process the referential resolution and / or semantic missingness in the first question information based on the first prompt information.

[0015] In some embodiments, the first prompt information is used to instruct the machine learning model to process the information based on context information in response to the presence of pronouns and / or missing sentence components in the first prompt information.

[0016] In some embodiments, the information type to which the first question information belongs is determined using a machine learning model based on the second prompt information, and the second prompt information includes relevant information of multiple information types.

[0017] In some embodiments, outputting the response information corresponding to the first question information includes: outputting the voice response information corresponding to the first question information using a machine learning model based on the first question information, wherein the first question information includes question vocabulary in a first language and question vocabulary in a second language, and the voice response information includes response vocabulary in the first language and response vocabulary in the second language.

[0018] In some embodiments, the first question information includes a first question voice, and outputting the voice response information corresponding to the first question information using a machine learning model includes: using a machine learning model to output the voice response information corresponding to the first question information based on the voice recognition result of the first question voice.

[0019] In some embodiments, outputting voice response information corresponding to the first question information using a machine learning model includes: outputting voice response information using a digital human model based on the lip shape information corresponding to the voice response information.

[0020] According to some other embodiments of this disclosure, an interactive device is provided, including: a receiving unit for receiving input first question information; and an output unit for outputting response information corresponding to the first question information in response to the first question information belonging to a specified information type and matching a first information set, and outputting a rejection message in response to the first question information not matching the first information set but matching a second information set.

[0021] In some embodiments, the first query information belonging to a specified information type includes a specified person entity and / or a specified organization entity.

[0022] In some embodiments, the corresponding response information is generated using a machine learning model based on the retrieval results of the first question information in the knowledge base.

[0023] In some embodiments, the search results are determined based on a first candidate search result and a second candidate search result in the knowledge base. The first candidate search result is determined based on the similarity between the feature vector of the first query information and the feature vector of the information in the knowledge base. The second candidate search result is determined based on the term frequency and inverse document frequency of the first query information in the knowledge base.

[0024] In some embodiments, the search results are determined by filtering the first and second candidate search results by keywords and then reordering them.

[0025] In some embodiments, the output unit responds to the fact that the first question information does not match either the first information set or the second information set by outputting the corresponding response information, which is generated using a machine learning model.

[0026] In some embodiments, the output unit responds to the first question information belonging to a specified information type, and the first question information does not match the first information set but matches the second information set, by outputting a rejection message, and / or, the output of the corresponding response information includes: the output unit responds to the first question information belonging to a specified information type, and the first question information does not match either the first information set or the second information set, by outputting the corresponding response information.

[0027] In some embodiments, the output unit outputs multimedia content corresponding to the preset topic information in response to a first question information matching preset topic information, and / or outputs corresponding response information in response to a first question information being related to a designated organization.

[0028] In some embodiments, the corresponding multimedia content is determined based on topic tags, which are determined using a machine learning model based on the first question information.

[0029] In some embodiments, the information type to which the first question information belongs is determined based on the second question information, and the second question information is generated by using a machine learning model to process the referential resolution and / or semantic missingness in the first question information based on the first prompt information.

[0030] In some embodiments, the first prompt information is used to instruct the machine learning model to process the information based on context information in response to the presence of pronouns and / or missing sentence components in the first prompt information.

[0031] In some embodiments, the information type to which the first question information belongs is determined using a machine learning model based on the second prompt information, and the second prompt information includes relevant information of multiple information types.

[0032] In some embodiments, the output unit outputs voice response information corresponding to the first question information using a machine learning model based on the first question information. The first question information includes question vocabulary in a first language and question vocabulary in a second language, and the voice response information includes response vocabulary in the first language and response vocabulary in the second language.

[0033] In some embodiments, the first question information includes a first question voice, and the output unit uses a machine learning model to output voice response information corresponding to the first question information based on the voice recognition result of the first question voice.

[0034] In some embodiments, the output unit outputs voice response information based on the lip shape information corresponding to the voice response information using a digital human model.

[0035] According to further embodiments of this disclosure, an interactive device is provided, comprising: a memory; and a processor coupled to the memory, the processor being configured to execute the interactive method of any of the above embodiments based on instructions stored in the memory device.

[0036] According to further embodiments of the present disclosure, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the interactive methods of any of the above embodiments.

[0037] According to some embodiments of this disclosure, a computer program product is also provided, including instructions that, when executed by a processor, cause the processor to perform an interactive method according to any of the foregoing embodiments.

[0038] Other features and advantages of this disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description

[0039] The accompanying drawings, which are included to provide a further understanding of this disclosure and form part of this application, illustrate exemplary embodiments of the present disclosure and are used to explain the disclosure, but do not constitute an undue limitation thereof. In the drawings:

[0040] Figure 1 shows a flowchart of an interaction method according to some embodiments of the present disclosure;

[0041] Figures 2a-2d illustrate schematic diagrams of interactive systems according to some embodiments of the present disclosure;

[0042] Figures 3a-3c show schematic diagrams of speech recognition methods according to some embodiments of the present disclosure;

[0043] Figure 4 shows a block diagram of an interactive device according to some embodiments of the present disclosure;

[0044] Figure 5 shows a block diagram of some other embodiments of the interactive device of this disclosure;

[0045] Figure 6 shows a block diagram of some further embodiments of the interactive device of this disclosure. Detailed Implementation

[0046] The technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit this disclosure or its application or use. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without creative effort are within the scope of protection of this disclosure.

[0047] Unless otherwise specifically stated, the relative arrangement, numerical expressions, and values ​​of the components and steps set forth in these embodiments do not limit the scope of this disclosure. It should also be understood that, for ease of description, the dimensions of the various parts shown in the drawings are not drawn to actual scale. Techniques, methods, and devices known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and devices should be considered part of the specification. In all examples shown and discussed herein, any specific values ​​should be interpreted as merely exemplary and not as limitations. Therefore, other examples of exemplary embodiments may have different values. It should be noted that similar reference numerals and letters in the following drawings denote similar items; therefore, once an item is defined in one drawing, it need not be further discussed in subsequent drawings.

[0048] The inventors of this disclosure have discovered the following problem in the aforementioned related technologies: the inability to effectively determine the response method to a question leads to poor human-computer interaction. Therefore, this disclosure proposes an interaction technology solution that can effectively determine the response method to a question based on the information content and type of the question, combined with a first information set and a second information set, thereby improving the human-computer interaction effect of the computer.

[0049] For example, the technical solution of this disclosure can be implemented through the following embodiments.

[0050] Figure 1 shows a flowchart of an interaction method according to some embodiments of the present disclosure.

[0051] As shown in Figure 1, in step 110, the first question information is received. For example, the first question information may be voice, text, or other question information input by the user through an interactive platform, interactive application, or other interactive system.

[0052] In step 120, in response to the first question information belonging to the specified information type and matching the first information set, the answer information corresponding to the first question information is output.

[0053] In some embodiments, the specified information type includes a first information type, and the first question information belonging to the first information type includes a specified person entity and / or a specified organization entity. For example, the first question information is "What are the main responsibilities of XX person in the company?", which is related to the person entity "XX person", and it can be determined that the first question information belongs to the specified information type; in this case, it can be further determined whether the first question information matches the question sample data in the first information set.

[0054] In some embodiments, the first information set contains multiple specified first question sample data. Whether the first question information matches the first information set can be determined by judging whether there is first question sample data in the first information set that matches the first question information. For example, the first question information is "What are the main responsibilities of XX person in the company?", and the first information set includes one first question sample data "The responsibilities of XX person". The semantic similarity between the first question information and the first question sample data can be determined by methods such as the similarity of their feature vectors, thus determining that the first question information matches the first question sample data, and consequently, that the first question information matches the first information set. In this case, the corresponding answer information can be further generated and output.

[0055] In step 130, in response to the first query information not matching the first information set but matching the second information set, a rejection message is output. Steps 120 and 130 have no execution order.

[0056] In some embodiments, in response to the first question information belonging to a specified information type and not matching a first information set but matching a second information set, a rejection message is output. For example, if the first question information is "What is the office location of XX person in the company?", which is related to the person entity "XX person", it can be determined that the first question information belongs to a specified information type; in response to the absence of first question sample data matching the first question information in the first information set, it can be further determined whether the first question information matches the question sample data in the second information set.

[0057] In some embodiments, the second information set contains multiple specified second question sample data. Whether the first question information matches the second information set can be determined by judging whether there is second question sample data in the second information set that matches the first question information. For example, the first question information is "What is the office location of XX person in the company?", and the second information set includes one second question sample data "XX person's office location". The semantic similarity between the first question information and the second question sample data can be determined by the similarity of their feature vectors, thus determining that the first question information matches the second question sample data, and consequently, that the first question information matches the second information set. In this case, a rejection message can be output.

[0058] In the above embodiments, a second information set including multiple second question sample data that refuse to answer can be preset, such as question samples involving personal privacy, corporate secrets, etc.; by determining whether the current question information matches these question samples that refuse to answer, information security can be improved.

[0059] In some embodiments, in response to a first question not matching either the first information set or the second information set, a corresponding response is output, which is generated using a machine learning model. For example, in response to a first question belonging to a specified information type and not matching either the first or second information set, a corresponding response is output. For example, if the first question is "What are XX celebrity's recent works?", which is related to the person entity "XX celebrity", it can be determined that the first question belongs to a specified information type. In this case, it can be further determined whether the first question matches the question sample data in the first information set. In response to no matching question sample data in the first information set, it can be further determined whether the first question matches the question sample data in the second information set. In response to no matching question sample data in the second information set, a response can be directly generated using the model. In the above embodiments, based on the information type of the question, the first information set is used first to determine the response method; based on the determination result, the second information set is combined to further determine the response method, thereby improving the human-computer interaction effect of the computer.

[0060] The following examples illustrate how the query information is processed in step 110.

[0061] In some embodiments, the information type to which the first question information belongs is determined using a machine learning model based on the second prompt information, and the second prompt information includes relevant information of multiple information types.

[0062] For example, the information type to which the first question belongs can include multiple information types, such as the first information type, the second information type, the third information type, and the fourth information type. The first information type can be question information that includes a specified person entity and / or a specified organization entity; the second information type can be question information that matches preset topic information; the third information type can be question information related to a specified organization; and the fourth information type is other question information, such as general questions related to life, entertainment, or consultation. For example, the information type specified in step 120 can be the first information type; in response to the first question information belonging to the first information type, it is necessary to further determine whether the first question information matches the first information set.

[0063] For example, a user inputs the first question, "What are the core values ​​of XX company?", through the interactive system. This first question can be input into the LLM (Large Language Model), along with a prompt message (the second prompt): {"instruction": "You are a task classifier. Given a question, which category (A, B, C, D) does it belong to? A: Related to XX person; B: Matches XX topic; C: Related to XX company; D: General question information. Classify the user's question according to these four categories and simply output ABCD."}

[0064] "input": "What are the core values ​​of XX company?", "output": "B"}.

[0065] The "input" field contains the user's question, while the "instruction" field provides guidance to the LLM (Local Management Module) on categorizing the question. This instruction includes the four information types mentioned above. The LLM categorizes the question based on the meaning of these four information types for subsequent processing.

[0066] For example, if the second prompt information indicates that the question information belongs to information type A using LLM, the question information can be determined to belong to the specified information type, and subsequent processing can be performed in any of the above embodiments; if the second prompt information indicates that the question information belongs to information type B using LLM, the third prompt information can be used to determine the topic tag corresponding to the question information to identify the topic corresponding to the question information, and then output multimedia content corresponding to the topic.

[0067] In this way, the machine learning model can be further enhanced to adapt to the task of classifying question information by fine-tuning instructions (such as a second prompt message), thereby improving the accuracy of information type judgment and improving the effect of human-computer interaction.

[0068] In some embodiments, the information type to which the first question information belongs is determined based on the second question information, and the second question information is generated by using a machine learning model to process the referential resolution and / or semantic missingness in the first question information based on the first prompt information.

[0069] For example, the dialogue understanding and dialogue generation capabilities of machine learning models (such as large language models) can be used to rewrite the first question information to address issues such as the resolution of references and the lack of semantics, in order to generate the second question information.

[0070] For example, a user interacts with a computer through an interactive system: the first round of questions is "Introduce the vision of XX Company"; the first round of responses is "XX Company's vision is to become a respected company"; the second round of questions is "What are the core values?". After the first round of questions and responses, the user's second round of questions, "What are the core values?", lacks a subject. A machine learning model can be used to determine, based on the context, that the user actually wants to know "What are the core values ​​of XX Company?" and complete the subject "XX Company".

[0071] In this way, by resolving pronoun substitution and / or handling semantic gaps, the semantics of the question information become clearer, thereby improving the accuracy of information type judgment and enhancing the effect of human-computer interaction.

[0072] In some embodiments, the first prompt information is used to instruct the machine learning model to process the information based on contextual information in response to the presence of pronouns and / or missing sentence components in the first question information. For example, the first prompt information can instruct the machine learning model to replace pronouns in the first question information with corresponding nouns based on contextual information, or to complete missing subjects, objects, or other components in the first question information.

[0073] For example, prompts such as the first, second, third, and fourth prompts are visible to the development side (e.g., the system's maintenance personnel and developers), but not to the user side (e.g., users asking questions through the interactive system).

[0074] For example, you can input the above dialogue into an LLM program and provide the prompt message: "Your task is to rewrite the following sentence: 1. If the sentence contains pronouns, missing subjects, missing predicates, or missing objects, you must appropriately replace or fill them in based on historical information to eliminate referential ambiguity; 2. If there is no historical information or the historical information is not very relevant, you can directly retain the original sentence; 3. Please do not provide answers, additional information, or explanations, just ensure that the sentence is grammatically correct, structurally clear, and complete. The sentence to be rewritten is:\nIntroduce the vision of XX Company. The vision of XX Company is to be a respected company. What are its core values?" Based on the above prompt message, the LLM program outputs the rewritten question: "What are the core values ​​of XX Company?"

[0075] In this way, the machine learning model can be further enhanced to adapt to tasks involving the resolution of pronouns and / or semantic missing information by fine-tuning instructions (such as the first prompt message), thereby improving the effect of human-computer interaction.

[0076] The following examples illustrate how the response information is generated in step 120.

[0077] In some embodiments, the corresponding response information is generated using a machine learning model based on the retrieval results of the first question information in the knowledge base. For example, for an enterprise interaction system, a knowledge base can be established based on the knowledge and experience information accumulated within the enterprise; content related to the first question information can be retrieved from the knowledge base as the retrieval results; and LLM can be used to generate response information based on the retrieval results.

[0078] In this way, by effectively managing, deeply mining, and integrating the knowledge and experience information accumulated within the enterprise, and utilizing machine learning models, an intelligent and efficient interactive system is provided, thereby improving the effectiveness of human-computer interaction.

[0079] In some embodiments, in response to the first question information belonging to a first information type and the first question information matching a first information set, or the first question information belonging to a third information type, a response information is generated using a machine learning model based on the retrieval results of the first question information in the knowledge base.

[0080] In this way, different response processing methods are adopted according to different types of question information, which solves the illusion problem of machine learning models and improves the effect of human-computer interaction. Moreover, in response to user questions involving institutions and people, the response information is generated by searching a pre-established knowledge base, which can improve the accuracy of the response and thus improve the effect of human-computer interaction.

[0081] In some embodiments, the search results are determined based on a first candidate search result and a second candidate search result in the knowledge base. The first candidate search result is determined based on the similarity between the feature vector of the first query information and the feature vector of the information in the knowledge base. The second candidate search result is determined based on the TF (Term Frequency) and IDF (Inverse Document Frequency) of the first query information in the knowledge base. For example, the search results are determined by filtering the first and second candidate search results by keywords and then re-ranking them.

[0082] For example, firstly, the knowledge content information (such as document information) and the first question information in the knowledge base can be vectorized using a vector model (such as Sentence-BERT or BGE-M3). The vectorized knowledge content information can also be pre-stored in a vector database to improve processing speed. Then, based on the cosine similarity between the vectors of the knowledge content information and the vectors of the first question information, the top n most similar knowledge content information (Top-n) can be selected as the first candidate retrieval results.

[0083] For example, similar document algorithms (such as BM25) can be used to retrieve the top n most relevant knowledge content information (Top-n) from the knowledge base as second candidate search results. The relevance between the knowledge content information D and the first query information Q can be calculated based on TF, IDF, and information length (such as document length).

[0084] q i It is the i-th word in the first question information, f(q) i D) is the word q i The word frequency in knowledge content information D, IDF(q) i ) is the word q i In the knowledge base, IDF, |D| is the document length of knowledge content information D, and avgdl is the average document length of all knowledge content information in the knowledge base. k1 and b are adjustable parameters; k1 can range from 1.2 to 2, and b is less than 1 and greater than 0 (e.g., 0.75). N is the total number of knowledge content information in the knowledge base, n(q i ) is a word containing q i The amount of knowledge content information.

[0085] For example, the first candidate search results based on vector similarity and the second candidate search results based on similar document algorithms are filtered to select knowledge content information containing keyword information through keyword filtering and similarity threshold screening. The top k similar knowledge content information (Top-k) is determined from the filtered results based on relevance through a secondary sorting algorithm as the final search results. The first question information and the final search results are input into the LLM so that the LLM can generate the response information.

[0086] In this way, by employing multi-path recall retrieval methods (such as vector similarity, similar document algorithms, keyword filtering, etc.) and secondary ranking strategies, the accuracy of retrieval can be effectively improved, thereby enhancing the human-computer interaction effect.

[0087] In some embodiments, in response to a match between a first query and preset topic information, multimedia content (such as images, audio, video, etc.) corresponding to the preset topic information is output. For example, the corresponding multimedia content is determined based on topic tags, which are determined using a machine learning model based on the first query.

[0088] For example, preset topic information can be product-related information such as the production process of a specified product (e.g., production line) or information about a specified product (e.g., platform); it can output PPT content corresponding to these topics for explanation, thereby improving the effect of human-computer interaction.

[0089] For example, the first question and the third prompt are input into a machine learning model to determine the topic tag corresponding to the first question. The third prompt is used to instruct the machine learning model to determine the topic tag corresponding to the first question among the specified topic tags.

[0090] For example, semantic tags can be used for topic identification, such as tag categories: {"1": "Reasons for building a central control platform", "2": "Definition of a central control platform", "3": "Origin and development history of a central control platform", "4": "Artificial intelligence in a central control platform", "5": "Internet of Things in a central control platform", "6": "User value of a central control platform", ...}. Compared with numerical tags, semantic tags contain semantic information, which can improve the effect of human-computer interaction.

[0091] For example, the third prompt message can be configured as follows, with the final output topic tag being "Reasons for building a central control platform":

[0092] For example, the third prompt message can be configured as follows, with the final output topic tag being "User Value of the Central Control Platform":

[0093] For example, if a user's first question is "What problem is the central control platform created to solve?", the machine learning model can be used to determine the topic tag corresponding to the first question as "the reason for creating the central control platform" based on the third prompt information mentioned above. The video content that introduces the central control platform and is bound to the topic tag "the reason for creating the central control platform" can be called from the database and displayed to the user on the display device.

[0094] In the above embodiments, different prompts are assigned to the machine learning model in response to different tasks (such as question information classification, referential resolution and / or semantic missing tasks, topic recognition, etc.). Thus, through this prompting process, the adaptability of the machine learning model to different tasks is further enhanced, thereby improving the effectiveness of human-computer interaction.

[0095] In some embodiments, in response to the first question information belonging to the fourth information type, there is no need to search in the knowledge base; instead, the answer information is directly generated using a machine learning model to improve the efficiency of human-computer interaction.

[0096] The following examples illustrate how the response information is output in step 120.

[0097] In some embodiments, based on the first question information, a machine learning model is used to output the voice response information corresponding to the first question information. The first question information includes question vocabulary in a first language and question vocabulary in a second language, and the voice response information includes response vocabulary in the first language and response vocabulary in the second language.

[0098] For example, a speech synthesis algorithm based on mixed Chinese and English can be used with an end-to-end speech synthesis model to output voice responses. The speech synthesis model can be trained using mixed Chinese and English speech data, enabling the same model to generate speech that includes both Chinese and English. This minimizes information loss during speech generation and reduces the awkwardness caused by switching between Chinese and English speech models, thereby improving the quality of voice responses.

[0099] In some embodiments, a machine learning model is used to output voice response information corresponding to the first question based on the speech recognition result of the first question.

[0100] For example, based on the framework of speech recognition models, the WSFT (Weighted Finite State Transducer) hot word algorithm can be used to weight and enhance the vocabulary in a specified domain, thereby effectively improving the recognition accuracy of the vocabulary in the specified domain.

[0101] For example, corpora can be constructed based on vocabulary from a specific domain, and datasets can be recorded and collected to fine-tune ASR (Automatic Speech Recognition) models, thereby increasing the number of speech samples containing vocabulary from that domain during the training phase. This allows the machine learning model to become familiar with the vocabulary of that domain, thus improving the accuracy of its speech responses.

[0102] In some embodiments, a digital human model is used to output voice response information based on the lip shape information corresponding to the voice response information.

[0103] For example, lip-sync algorithms can be used to align the lip shape of a digital human with its speech. This allows audio synthesized using TTS (Text-to-Speech) algorithms to be input into a pre-trained deep neural network, outputting a 3D vertex mesh for the digital human. This enables real-time creation of facial animations corresponding to the digital human's speech responses, thereby improving the effectiveness of human-computer interaction.

[0104] In the above embodiments, the human-computer interaction system was optimized at the interaction level based on technologies such as machine learning models and 3D digital humans. For example, based on the semantic understanding and generation capabilities of large language models, the computer's human-computer interaction capabilities were enhanced through the classification and processing of question information, the rewriting of question information, and the generation of response information based on retrieval enhancement technology. The anthropomorphism of the digital human can also be improved through technologies such as 3D rendering and lip-syncing. Furthermore, the effect of human-computer interaction can be improved through voice interaction technologies such as speech recognition and speech synthesis, as well as visualization technologies.

[0105] In this way, 3D digital human technology can be combined with interactive systems. Based on technologies such as lip-syncing, body movement choreography, and 3D rendering, the interactive effect of digital humans becomes more realistic, thereby improving the effect of human-computer interaction.

[0106] The interactive system of this disclosure is illustrated below by way of examples from the embodiments shown in Figures 2a-2c.

[0107] Figures 2a-2c show schematic diagrams of interactive systems according to some embodiments of the present disclosure.

[0108] As shown in Figure 2a, the interactive application 22a can be optimized at the interaction level based on technologies such as the backend large model 21a and digital human 23a. For example, based on the semantic understanding and generation capabilities of the backend large model 21a (such as a large language model), the human-computer interaction capabilities of the interactive application 22a can be enhanced through classification and processing of question information, rewriting of question information, and generation of response information based on retrieval enhancement technology. For example, the anthropomorphism of the digital human 23a can be improved through technologies such as 3D rendering and lip-syncing; and the human-computer interaction effect of the interactive application 22a can be improved through voice interaction technologies such as speech recognition and speech synthesis, as well as visualization presentation technologies.

[0109] As shown in Figure 2b, in step 210b, the interactive application receives a question input by the user via voice.

[0110] In step 220b, speech recognition is performed on the question information based on LLM.

[0111] In step 230b, the information type of the question is determined based on the speech recognition result; and the processing method for responding to the question is determined according to the information type.

[0112] For example, if the question information belongs to the first information type, the response information can be generated through retrieval enhancement. For instance, a machine learning model can be used to generate the corresponding response information based on the retrieval results of the first question information in the knowledge base.

[0113] For example, if the question information belongs to the second information type, meaning the question information matches preset topic information, multimedia content corresponding to the preset topic information can be output. For instance, a machine learning model can be used to determine the topic tags that match the first question information; the corresponding multimedia content can then be determined based on the topic tags.

[0114] For example, if the question information belongs to the fourth type of information, there is no need to search the knowledge base; the answer information can be directly generated using a machine learning model.

[0115] In step 240b, the lip shape of the digital human can be aligned with the response speech based on the lip-driven algorithm to generate the facial information of the digital human.

[0116] In step 250b, a digital human can be displayed on the screen based on facial information.

[0117] In step 260b, response information can be output using digital human voice based on facial information and response voice.

[0118] As shown in Figure 2c, the interactive system disclosed herein may include a content review module, a question information rewriting module, a question information classification module, a topic recognition module, a search enhancement module, and a response generation module.

[0119] In step 210c, the interactive application receives a question input by the user via voice and uses a content moderation module to review the question. For example, based on the prompt project, an LLM can be used to review whether the question contains specified keywords; if the specified keywords are included, the answer can be directly rejected; the prompt project's prompt information can include the specified keywords and information about the review logic.

[0120] By adding a content moderation module to the input end of the interactive system, security issues introduced at the input end can be avoided; by building a specified keyword library, a polite refusal message can be provided for questions that fail the moderation, thereby improving the effectiveness of human-computer interaction.

[0121] In step 220c, the prompt process uses the first instruction to instruct the machine learning model to replace pronouns in the input with corresponding nouns based on the context information (history), and can also complete missing subjects, objects, and other components in the input (output).

[0122] For example, the first prompt message can be configured as follows:

[0123] In step 230c, it is determined whether the question information belongs to a first information type, a second information type, a third information type, or a fourth information type. For example, the first information type can be question information that includes a specified person entity and / or a specified organization entity; the second information type can be question information that matches preset topic information; the third information type can be question information related to a specified organization; and the fourth information type is other question information, such as general questions related to life, entertainment, or consultation. For example, a machine learning model can be used to determine the information type to which the question information belongs based on the second prompt information.

[0124] For example, in response to a match between the question information and preset topic information, it can be determined that the question information belongs to the second information type; a machine learning model is used to generate topic tags that match the first question information, and the corresponding multimedia content is determined based on the topic tags; the multimedia content corresponding to the preset topic information is then output. For example, a machine learning model can be used to determine the topic tags that match the question information based on the third prompt information.

[0125] In this way, by introducing preset topic information into the human-computer interaction process, the digital human can recognize the intent of the question and explain it based on the multimedia content (PPT, pictures, etc.) corresponding to the topic, thereby improving the effect of human-computer interaction.

[0126] For example, if the first question information belongs to the fourth information type, there is no need to search the knowledge base. Instead, the machine learning model can be used to directly generate the response information to improve the efficiency of human-computer interaction.

[0127] For example, if the first query information belongs to the third information type, a response can be generated through retrieval enhancement. Retrieval enhancement can be achieved through embedding, using a vector model to search the knowledge base, and generating a response based on the search results. Alternatively, a prompt process can be used, with a fourth prompt instructing the LLM to generate a response based on the search results. For instance, a knowledge base can be built based on the company's accumulated knowledge and experience; multi-path recall (vector similarity, BM25, keyword filtering, etc.) and secondary ranking strategies can be employed to effectively improve retrieval accuracy.

[0128] This can solve the illusion problem caused by the lack of knowledge in specific domains in LLM, thereby improving the effectiveness of human-computer interaction.

[0129] For example, in response to the first question information belonging to a first information type (which can be resolved using named entity recognition technology), the response method can be determined through steps 240c and 250c. In step 240c, it is determined whether the question information matches the first information set. If they match, a response is generated using retrieval enhancement; otherwise, step 250c is executed. In step 250c, it is determined whether the question information matches the second information set. If they match, the response is rejected; otherwise, there is no need to search the knowledge base, and a response is directly generated using a machine learning model.

[0130] In the above embodiments, the semantic understanding and response generation capabilities of LLM are used to construct an interactive system, forming the brain of the digital human; speech recognition and speech synthesis technologies are used to form the digital human's hearing and mouth, enabling the digital human to hear and speak. In this way, this digital human can interact with the user as the persona of the entire interactive system, thereby improving the effectiveness of human-computer interaction.

[0131] The retrieval enhancement method of this disclosure is illustrated below by way of an embodiment shown in Figure 2d.

[0132] Figure 2d shows a schematic diagram of an interactive system according to some embodiments of the present disclosure.

[0133] As shown in Figure 2d, firstly, the document information and the first question information in the knowledge base can be vectorized using a vector model. The vectorized knowledge content information can also be pre-stored in a vector database to improve processing speed. Then, based on the cosine similarity between the vectors of the document information and the vectors of the first question information, the top n most similar documents (Top-n) can be selected as the first candidate retrieval results. Alternatively, algorithms such as BM25 can be used to retrieve the top n most relevant documents (Top-n) from the knowledge base as the second candidate retrieval results.

[0134] The system can retrieve first candidate search results based on vector similarity and second candidate search results based on similar document algorithms. By filtering keywords and similarity thresholds, document information containing keyword information can be selected. Using a ranking model and a secondary ranking algorithm, the top k similar knowledge content information (Top-k) is determined from the selected results based on relevance as the final search result. The first question information and the final search results are input into the LLM so that the LLM can generate response information.

[0135] In this way, by employing multi-path recall retrieval methods (such as vector similarity, similar document algorithms, keyword filtering, etc.) and secondary ranking strategies, the accuracy of retrieval can be effectively improved, thereby enhancing the human-computer interaction effect.

[0136] The speech recognition method of this disclosure is illustrated below by way of examples from the embodiments shown in Figures 3a-3c.

[0137] Figures 3a-3c show schematic diagrams of speech recognition methods according to some embodiments of the present disclosure.

[0138] As shown in Figure 3a, during the training process of the speech recognition model, speech data from general domains and speech data from specific domains (such as specific technical fields, knowledge domains of specific enterprises, etc.) can be used as training samples; the features of the extracted training samples are input into the encoder and processed in conjunction with a language model (such as an N-gram model); the encoding results are input into the decoder, and the decoder is trained using the text annotations corresponding to the speech.

[0139] As shown in Figure 3b, during the inference process of the speech recognition model, the features of the extracted user-input speech data can be input into the encoder and processed in conjunction with a language model (such as an N-gram model); the encoded result is then input into the decoder for further processing to output the recognition result. For example, based on the framework of the speech recognition model, hot word algorithms such as WSFT can be used to weight and enhance vocabulary in a specific domain, thereby effectively improving the recognition accuracy of vocabulary in that domain.

[0140] As shown in Figure 3c, the decoder described above can be built based on the transformer model; the speech recognition model can include multiple decoders, and the model of each decoder can adopt the structure shown in Figure 3c. For example, each decoder can mainly contain three parts: MLP (Multilayer Perceptron), RMSNorm (Root Mean Square Normalization) layer, and Attention Layer.

[0141] For example, the Attention Layer can add a RoPE (Rotation Position Encoding) module to the Q and K tensors. This rotation position encoding relies on absolute position encoding in form, but can be transformed into relative position encoding during computation to improve the ability to learn the contextual semantics of long texts.

[0142] For example, Feed Forward (i.e., MLP) needs to learn the up-linear and down-linear transformations, as well as the ReLU activation function FFN(x,W1,W2,b1,b2) = max(0,xW1+b1)W2+b2 used for these two linear transformations; MLP also includes Gate units, which can be calculated in the following ways: Based on the above formula, the activation function SwiGLU can be obtained as follows: β is a specified constant; MLP can also include an SLU (spoken language understanding) module.

[0143] For example, during the data feedback process, LoRA (Low-Rank Adaptation) technology can be used to create a branch next to the LLM, and then fine-tuning can be performed using the feedback data. This reduces the loss of pre-learned knowledge in the LLM while allowing the addition of new knowledge, eliminating the need for full fine-tuning and saving computational costs.

[0144] For example, in the generative model's prediction process, information from the first K-1 words is used as prior information, while information from the Kth word onwards is ignored, and the prediction is made for the Kth word. In this self-supervised learning approach, the Kth token is used as the prediction target for the (K-1)th token, and the loss value is calculated as follows:

[0145] The loss represents the sum of the losses of the n words in the response to be predicted, x1, x2, ..., xn. m This indicates a prompt message.

[0146] In the above embodiments, the interactive system employs techniques such as prompt engineering, model fine-tuning, and retrieval enhancement. Based on prompt engineering, fine-tuning is performed by combining prompt information and execution domain data, enhancing the machine learning model's ability to perform relevant tasks given a prompt. This involves optimizing the task-based prompt into a system-integrated prompt. The question information classification module responds to different question information inputs into the interactive system, determines the information type of the current task, and then completes the corresponding processing through the execution of sub-task modules. Through referential resolution and / or semantic missing processing, multi-turn dialogues can be completed without contextual information during the machine learning model's response generation process. This mitigates the semantic deviation of the classification module and downstream sub-task modules caused by adding context, thereby improving the accuracy of the entire interactive system's understanding and response generation.

[0147] Figure 4 shows a block diagram of an interactive device according to some embodiments of the present disclosure.

[0148] As shown in Figure 4, the interactive device 4 includes: a receiving unit 41 for receiving input first question information; and an output unit 42 for outputting response information corresponding to the first question information in response to the first question information belonging to a specified information type and matching a first information set, and outputting a rejection response information in response to the first question information not matching the first information set but matching a second information set.

[0149] In some embodiments, the first query information belonging to a specified information type includes a specified person entity and / or a specified organization entity.

[0150] In some embodiments, the corresponding response information is generated using a machine learning model based on the retrieval results of the first question information in the knowledge base.

[0151] In some embodiments, the search results are determined based on a first candidate search result and a second candidate search result in the knowledge base. The first candidate search result is determined based on the similarity between the feature vector of the first query information and the feature vector of the information in the knowledge base. The second candidate search result is determined based on the term frequency and inverse document frequency of the first query information in the knowledge base.

[0152] In some embodiments, the search results are determined by filtering the first and second candidate search results by keywords and then reordering them.

[0153] In some embodiments, the output unit 42 responds to the fact that the first question information does not match either the first information set or the second information set, and the corresponding response information is generated using a machine learning model.

[0154] In some embodiments, the output unit 42 responds to the first question information belonging to a specified information type, and the first question information does not match the first information set but matches the second information set, by outputting a rejection message, and / or, outputting the corresponding response information includes: the output unit 42 responds to the first question information belonging to a specified information type, and the first question information does not match either the first information set or the second information set, by outputting the corresponding response information.

[0155] In some embodiments, the output unit 42 outputs multimedia content corresponding to the preset topic information in response to the first question information matching the preset topic information, and / or outputs corresponding response information in response to the first question information being related to a designated organization.

[0156] In some embodiments, the corresponding multimedia content is determined based on topic tags, which are determined using a machine learning model based on the first question information.

[0157] In some embodiments, the information type to which the first question information belongs is determined based on the second question information, and the second question information is generated by using a machine learning model to process the referential resolution and / or semantic missingness in the first question information based on the first prompt information.

[0158] In some embodiments, the first prompt information is used to instruct the machine learning model to process the information based on context information in response to the presence of pronouns and / or missing sentence components in the first prompt information.

[0159] In some embodiments, the information type to which the first question information belongs is determined using a machine learning model based on the second prompt information, and the second prompt information includes relevant information of multiple information types.

[0160] In some embodiments, the output unit 42 outputs voice response information corresponding to the first question information using a machine learning model based on the first question information. The first question information includes question vocabulary in a first language and question vocabulary in a second language, and the voice response information includes response vocabulary in the first language and response vocabulary in the second language.

[0161] In some embodiments, the first question information includes a first question voice, and the output unit 42 uses a machine learning model to output the voice response information corresponding to the first question information based on the voice recognition result of the first question voice.

[0162] In some embodiments, the output unit 42 outputs the voice response information using a digital human model based on the lip shape information corresponding to the voice response information.

[0163] Figure 5 shows a block diagram of some other embodiments of the interactive device of this disclosure.

[0164] As shown in FIG5, the interactive device 5 of this embodiment includes a memory 51 and a processor 52 coupled to the memory 51. The processor 52 is configured to execute the interactive method in any embodiment of this disclosure based on instructions stored in the memory 51.

[0165] The memory 51 may include, for example, system memory, fixed non-volatile storage media, etc. The system memory stores, for example, the operating system, application programs, boot loader, database, and other programs.

[0166] Figure 6 shows a block diagram of some further embodiments of the interactive device of this disclosure.

[0167] As shown in FIG6, the interactive device 6 of this embodiment includes a memory 610 and a processor 620 coupled to the memory 610. The processor 620 is configured to execute the interactive method in any of the foregoing embodiments based on instructions stored in the memory 610.

[0168] The memory 610 may include, for example, system memory, fixed non-volatile storage media, etc. The system memory may store, for example, the operating system, application programs, boot loader, and other programs.

[0169] Device 6 may also include an input / output interface 630, a network interface 640, and a storage interface 650. These interfaces 630, 640, and 650, as well as the memory 610 and processor 620, can be connected, for example, via a bus 860. Specifically, the input / output interface 630 provides a connection interface for input / output devices such as a monitor, mouse, keyboard, touchscreen, microphone, and speakers. The network interface 640 provides a connection interface for various networked devices. The storage interface 650 provides a connection interface for external storage devices such as SD cards and USB flash drives.

[0170] Those skilled in the art will understand that embodiments of this disclosure can be provided as methods, systems, or computer program products. Therefore, this disclosure can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this disclosure can take the form of a computer program product embodied on one or more computer-usable non-transitory storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0171] The present disclosure has now been described in detail. To avoid obscuring the concept of this disclosure, some details known in the art have not been described. Those skilled in the art will fully understand how to implement the technical solutions disclosed herein based on the above description.

[0172] The methods and systems of this disclosure may be implemented in many ways. For example, they may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above-described order of steps for the methods is for illustrative purposes only, and the steps of the methods of this disclosure are not limited to the specific order described above unless otherwise specifically stated. Furthermore, in some embodiments, this disclosure may also be implemented as a program recorded on a recording medium, the program including machine-readable instructions for implementing the methods according to this disclosure. Thus, this disclosure also covers recording media storing programs for performing the methods according to this disclosure.

[0173] While specific embodiments of this disclosure have been described in detail by way of example, those skilled in the art should understand that the examples are for illustrative purposes only and not intended to limit the scope of this disclosure. Those skilled in the art should understand that modifications can be made to the above embodiments without departing from the scope and spirit of this disclosure. The scope of this disclosure is defined by the appended claims.

Claims

1. An interaction method, comprising: Receive the first question input; In response to the fact that the first question information belongs to the specified information type and matches the first information set, the answer information corresponding to the first question information is output; In response to the first question not matching the first information set but matching the second information set, a rejection message is output.

2. The interaction method according to claim 1, wherein, The first query information belonging to the specified information type includes a specified person entity and / or a specified organization entity.

3. The interaction method according to claim 1, wherein, The corresponding response information is generated using a machine learning model based on the retrieval results of the first question information in the knowledge base.

4. The interaction method according to claim 3, wherein, The retrieval results are determined based on the first candidate retrieval results and the second candidate retrieval results in the knowledge base. The first candidate retrieval results are determined based on the similarity between the feature vector of the first question information and the feature vector of the information in the knowledge base. The second candidate retrieval results are determined based on the word frequency and inverse document frequency of the first question information in the knowledge base.

5. The interaction method according to claim 4, wherein, The search results are determined by filtering the first and second candidate search results by keywords and then re-sorting them.

6. The interaction method according to claim 1, further comprising: In response to the first question information not matching either the first information set or the second information set, the corresponding response information is output, which is generated using a machine learning model.

7. The interaction method according to claim 6, wherein, The output of the rejection message includes: in response to the first question information belonging to the specified information type, and the first question information not matching the first information set but matching the second information set, the output is... The statement indicates a refusal to answer information, and / or The output of the corresponding response information includes: in response to the first question information belonging to the specified information type, and the first question information not matching either the first information set or the second information set, outputting the corresponding response information.

8. The interaction method according to any one of claims 1-7, further comprising: In response to the first question information matching the preset topic information, the multimedia content corresponding to the preset topic information is output; and / or In response to the first question being related to a designated organization, the corresponding response information is output.

9. The interaction method according to claim 8, wherein, The corresponding multimedia content is determined based on topic tags, which are determined using a machine learning model based on the first question information.

10. The interaction method according to any one of claims 1-7, wherein, The information type of the first question information is determined based on the second question information. The second question information is generated by using a machine learning model to process the referential resolution and / or semantic missingness in the first question information based on the first prompt information.

11. The interaction method according to claim 10, wherein, The first prompt information is used to instruct the machine learning model to process the information based on the context in response to the presence of pronouns and / or missing sentence components in the first prompt information.

12. The interaction method according to claim 10, wherein, The information type of the first question is determined using a machine learning model based on the second prompt information, which includes relevant information of multiple information types.

13. The interaction method according to any one of claims 1-7, wherein, The response information corresponding to the first question includes: Based on the first question information, a machine learning model is used to output the voice response information corresponding to the first question information. The first question information includes question vocabulary in a first language and question vocabulary in a second language. The voice response information includes response vocabulary in the first language and response vocabulary in the second language.

14. The interaction method according to claim 13, wherein, The first question information includes the first question voice. The step of using a machine learning model to output the voice response information corresponding to the first question includes: The machine learning model is used to output the voice response information corresponding to the first question based on the speech recognition result of the first question.

15. The interaction method according to claim 13, wherein, The step of using a machine learning model to output the voice response information corresponding to the first question includes: Based on the lip shape information corresponding to the voice response information, the voice response information is output using a digital human model.

16. An interactive device, comprising: The receiving unit is used to receive the first input query information; The output unit is configured to output the response information corresponding to the first question information in response to the first question information belonging to a specified information type and matching a first information set, and to output the rejection information in response to the first question information not matching the first information set but matching a second information set.

17. An interactive device, comprising: Memory; and A processor coupled to the memory, the processor being configured to execute the interaction method of any one of claims 1-15 based on instructions stored in the memory device.

18. A computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the interactive method as described in any one of claims 1-15.

19. A computer program product comprising instructions that, when executed by a processor, cause the processor to perform the interactive method according to any one of claims 1-15.

Citation Information

Patent Citations

  • Fuzzy recognition system for legal consultation

    CN111324719A

  • Information acquisition method and device, electronic equipment and computer readable storage medium

    CN111368093A

  • Method, device, equipment and medium for human-computer interaction

    CN112286366A

  • Method, device and equipment for man-machine interaction and medium

    CN114578969A

  • Law and regulation question-answering system based on generative language model and construction method

    CN117271716A