Conversation processing method and device, program product and electronic equipment

By matching the user's conversation information with multiple tools and calling the target tool to generate replies, the problem that existing Q&A system cannot handle dynamic data is solved, improving the accuracy and user experience of the answers.

CN120067268APending Publication Date: 2025-05-30HANGZHOU NETEASE CLOUD MUSIC TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510213860.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing Q&A system is based on static text data and cannot effectively handle questions related to dynamic data, resulting in incorrect answers or inability to answer, reducing user experience.

Method used

By determining the first dialogue information and matching it with the information of the associated multiple tools, if the target tool is matched, the tool function is generated and the target tool execution tool function is called to obtain the reply information.

Benefits of technology

It improves the accuracy and efficiency of dialogue processing, can effectively handle issues related to dynamic data and operation and maintenance operations, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067268A_ABST
    Figure CN120067268A_ABST
Patent Text Reader

Abstract

The invention provides a dialogue processing method and device, a program product and electronic equipment, and relates to the technical field of computers. The method comprises the following steps: determining first dialogue information; matching the first dialogue information with information of a plurality of associated tools; and if the first dialogue information is matched with the information of a target tool in the plurality of tools, generating a tool function according to the target tool and the first dialogue information, and calling the target tool to execute the tool function so as to obtain reply information corresponding to the first dialogue information. According to the embodiment of the invention, the association with the plurality of tools is established, so that after the dialogue information is obtained, the dialogue information can be matched with the plurality of tools, the target tool which is most matched with the dialogue information is selected to process the dialogue information, the dialogue information can be timely and effectively responded, and the use experience of a user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of computer technology. More specifically, embodiments of the present disclosure relate to a dialogue processing method, a dialogue processing device, a computer program product, and an electronic device. Background Art

[0002] This section aims to provide background or context for the embodiments of the present disclosure stated in the claims. The description herein is not admitted to be prior art merely because it is included in this section.

[0003] Currently, most existing Q&A systems adopt the Retrieval-augmented generation (RAG) technology, that is, by retrieving relevant knowledge to enable a Large Language Model (LLM) to obtain higher-quality outputs. Summary of the Invention

[0004] However, since the Q&A system using the RAG technology answers based on static text data, it cannot answer or answers incorrectly for questions related to dynamic data, reducing the user experience.

[0005] In view of this, the present disclosure provides a dialogue processing method, a dialogue processing device, a computer program product, and an electronic device to improve the accuracy of dialogue processing to a certain extent.

[0006] According to a first aspect of the present disclosure, there is provided a dialogue processing method, the method comprising:

[0007] Determine first dialogue information;

[0008] Match the first dialogue information with information of a plurality of associated tools;

[0009] If the first dialogue information matches the information of a target tool among the plurality of tools, generate a tool function according to the target tool and the first dialogue information, and call the target tool to execute the tool function to obtain a reply information corresponding to the first dialogue information.

[0010] In a possible implementation manner, the method further comprises:

[0011] If the first dialogue information does not match the information of any tool among the plurality of tools, analyze the first dialogue information based on a large language model to obtain a reply information corresponding to the first dialogue information.

[0012] In a possible implementation manner, determining first dialogue information includes:

[0013] Receive current question information;

[0014] Historical conversation information before determining the current problem information; the historical conversation information includes the problem information and the corresponding answer information before the current problem information;

[0015] Determine the first conversation information according to the current problem information and the conversation history information before the current problem information.

[0016] In a possible implementation manner, matching the first conversation information with the information of multiple associated tools includes:

[0017] Determine the type information of the first conversation information;

[0018] Filter out candidate tools from the multiple tools according to the type information;

[0019] Match the first conversation information with the information of the candidate tools.

[0020] In a possible implementation manner, filtering out candidate tools from the multiple tools according to the type information includes:

[0021] Filter out candidate tools from the multiple tools according to the type information and the prompt template;

[0022] Wherein, the prompt template is configured for the large language model and includes the tool names, tool input parameters, call formats, output formats, and tool introductions of all tools.

[0023] In a possible implementation manner, filtering out candidate tools from the multiple tools according to the type information and the prompt template includes:

[0024] If the type information is an operation and maintenance operation type, filter out candidate tools of the atomic capability api type according to the operation and maintenance operation type and the tool introduction in the prompt template;

[0025] If the type information is a static text type, filter out candidate tools of the knowledge base type according to the static text type and the tool introduction in the prompt template.

[0026] In a possible implementation manner, calling the target tool to execute the tool function to obtain the reply information corresponding to the first conversation information includes:

[0027] Call the target tool to execute the tool function to obtain a tool result;

[0028] When it is determined to execute the direct return working mode, use the tool result as the reply information;

[0029] If the current execution cache working mode is determined, the tool result is used as the historical first conversation information, and the processing of the second conversation information is continued to obtain the response information corresponding to the first conversation information; the second conversation information is obtained based on the association with the first conversation information.

[0030] According to a second aspect of the present disclosure, there is provided a dialogue processing device, the device comprising:

[0031] A determination unit, configured to determine first conversation information;

[0032] A matching unit, configured to match the first conversation information with the information of a plurality of associated tools;

[0033] A processing unit, configured to, if the first conversation information matches the information of a target tool among the plurality of tools, generate a tool function according to the target tool and the first conversation information, and call the target tool to execute the tool function to obtain the response information corresponding to the first conversation information.

[0034] In a possible implementation manner, the processing unit is further configured to:

[0035] If the first conversation information does not match the information of any tool among the plurality of tools, analyze the first conversation information based on a large language model to obtain the response information corresponding to the first conversation information.

[0036] In a possible implementation manner, the determination unit is specifically configured to:

[0037] Receive the current problem information;

[0038] Determine the historical conversation information before the current problem information; the historical conversation information includes the problem information and the corresponding answer information before the current problem information;

[0039] Determine the first conversation information according to the current problem information and the conversation history information before the current problem information.

[0040] In a possible implementation manner, the matching unit is specifically configured to:

[0041] Determine the type information of the first conversation information;

[0042] Filter out candidate tools from the plurality of tools according to the type information;

[0043] Match the first conversation information with the information of the candidate tools.

[0044] In a possible implementation manner, the matching unit is specifically configured to:

[0045] Screen out candidate tools from the multiple tools according to the type information and the prompt template;

[0046] Among them, the prompt template is configured for the large language model and includes the tool names, tool input parameters, call formats, output formats, and tool introductions of all tools.

[0047] In a possible implementation manner, the matching unit is specifically configured to:

[0048] If the type information is an operation and maintenance operation type, screen out candidate tools of the atomic capability api type according to the operation and maintenance operation type and the tool introductions in the prompt template;

[0049] If the type information is a static text type, screen out candidate tools of the knowledge base type according to the static text type and the tool introductions in the prompt template.

[0050] In a possible implementation manner, the processing unit is specifically configured to:

[0051] Call the target tool to execute the tool function and obtain the tool result;

[0052] When it is determined to execute the direct return working mode, use the tool result as the reply information;

[0053] When it is currently determined to execute the cache working mode, use the tool result as the historical first dialogue information, and continue to process the second dialogue information to obtain the reply information corresponding to the first dialogue information; the second dialogue information is obtained based on the first dialogue information.

[0054] According to a third aspect of the present disclosure, there is provided a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the method of the first aspect and its possible implementation manners.

[0055] According to a fourth aspect of the present disclosure, there is provided an electronic device, including: a processor; and a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the executable instructions to execute the method of the first aspect and its possible implementation manners.

[0056] The technical solution of the present disclosure has the following beneficial effects:

[0057] In the embodiments of the present disclosure, first dialogue information can be determined; then the first dialogue information is matched with the information of a plurality of associated tools; further, if the first dialogue information matches the information of a target tool among the plurality of tools, a tool function is generated based on the target tool and the first dialogue information, and the target tool is called to execute the tool function to obtain a reply information corresponding to the first dialogue information. That is to say, since the embodiments of the present disclosure are associated with a plurality of tools, when dialogue information is obtained, the dialogue information can be matched with the plurality of tools, so as to select the target tool that best matches the dialogue information to process the dialogue information, and thus the dialogue information can be responded to in a timely and effective manner, improving the user experience.

[0058] Other features and advantages of the present disclosure will be described in the following specification, and, in part, will be obvious from the specification, or will be understood by implementing the present disclosure. The objectives and other advantages of the present disclosure can be realized and obtained by the structures specifically pointed out in the written specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] To more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following briefly introduces the drawings required to be used in the embodiments of the present disclosure. Obviously, the following introduced drawings are only some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0060] Figure 1 Schematic diagram showing the question answering process of a question answering system in the related art;

[0061] Figure 2 Schematic diagram showing an application scenario in this exemplary embodiment;

[0062] Figure 3 Flowchart showing a dialogue processing method in this exemplary embodiment;

[0063] Figure 4 Schematic diagram showing a dialogue information in this exemplary embodiment;

[0064] Figure 5 Schematic diagram showing a prompt word template in this exemplary embodiment;

[0065] Figure 6 Schematic diagram showing the question answering of an intelligent operation and maintenance question answering system in this exemplary embodiment;

[0066] Figure 7 Schematic diagram showing the structure of a dialogue processing device in this exemplary embodiment;

[0067] Figure 8 Schematic structural diagram of an electronic device in this exemplary embodiment is shown. Specific embodiments

[0068] To make the objectives, technical solutions, and advantages of the present disclosure clearer and more understandable, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part, rather than all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without creative efforts fall within the scope of protection of the present disclosure. Without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other arbitrarily. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0069] The terms "including" and any variations thereof in the specification and claims of the present disclosure are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.

[0070] One or more in the embodiments of the present disclosure, "multiple" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one (item)" or a similar expression thereof refers to any combination of these items, including any combination of single item (item) or plural items (items). For example, at least one (item) of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.

[0071] It should be noted that the terms "first", "second", etc. in the specification, claims and the above-mentioned drawings of the present disclosure are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order, sequence, size and priority. For example, the first conversation information and the second conversation information in the embodiments of the present disclosure are only used to distinguish different conversation information. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described here can be implemented in an order other than those illustrated or described here. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0072] The following makes an explanation of the exemplary embodiments of the present disclosure with reference to the accompanying drawings. The drawings are schematic diagrams of the present disclosure and are not necessarily drawn to scale. Some of the block diagrams shown in the drawings may be functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or in hardware modules or integrated circuits, or in networks, processors or microcontrollers. The embodiments can be implemented in multiple forms and should not be construed as limited to the examples set forth herein. The features, structures or characteristics described in the present disclosure can be combined in one or more embodiments in any suitable manner. In the following description, many specific details are provided to give a full explanation of the embodiments of the present disclosure. However, those skilled in the art should be aware that one or more specific details can be omitted when implementing the technical solutions of the present disclosure, or other methods, components, devices, steps, etc. can be used to replace one or more specific details. It should be noted that in the embodiments of the present disclosure, the dissemination, use, etc. of data all comply with relevant national laws and regulations. Summary of the Invention

[0074] Currently, with the continuous development of science and technology, more and more technical products are being generated. When a mature technical product is presented to users, it generally requires supporting documents, and the more functions the technical product has, the more cumbersome the documents become, thus bringing certain inconveniences and barriers to users in obtaining useful information related to the technical product and reducing the user experience. Based on this, a question-answering system is provided in the related art, through which professional answers can be given to the questions raised by users.

[0075] Most of the existing Q&A systems provided in the related art adopt the Retrieval-augmented generation (RAG) technology. Through this technology, the Q&A system can retrieve relevant information from an external knowledge base before generating an answer. Since the relevant information usually exists in the form of multiple document fragments, several fragments most relevant to the question raised by the user can be screened out through semantic similarity calculation. Subsequently, the several fragments determined by the foregoing screening and most relevant to the question raised by the user, together with the original question, are sent to a Large Language Model (LLM), so that the large language model generates a more accurate and well-founded answer to the question raised by the user. That is to say, through the RAG technology, not only can the hallucination problem of the LLM be solved, making the generated answer more credible, but also the operation and maintenance manpower can be saved.

[0076] For example, please refer to Figure 1 as shown in Figure 1 FIG. is a schematic diagram of the Q&A process of the Q&A system in the related art. Specifically, the documents corresponding to the technical products can be preprocessed, and the documents are preprocessed into multiple fragments. Each fragment will be converted into a vector form as an index and stored in a vector database. In this way, when a question raised by the user is received, the question can be converted into a vector form, and the most relevant several document fragments can be retrieved from the vector database. Further, the LLM can use the question received from the user and the retrieved document fragments as inputs to generate a final answer.

[0077] It can be seen that although the Q&A system using the RAG technology can, to a certain extent, rely on an external database to make the LLM better answer the questions raised by the user, there are the following problems:

[0078] (1) The Q&A system can only answer questions related to the data in the knowledge base. When the question raised by the user is not related to the database, the Q&A system will still retrieve the corresponding vector database once, resulting in irrelevant information being introduced into the reply to the question raised by the user. And although threshold truncation can be performed by calculating semantic similarity, it is impossible to ensure that all irrelevant texts are filtered, and the increase in overall time consumption is inevitable, affecting the user experience.

[0079] (2) Due to the randomness of the LLM, the same question may have different answer results. When the user inputs the same question, even if the retrieval process is the same and the questions and documents received by the LLM are the same, the answers of the LLM will be inconsistent, and even semantic contradictions may occur. That is, the controllability of the answer corresponding to the question in the Q&A system in the related art is poor.

[0080] (3) The Q&A system can only query and answer the processed static text data, making it difficult to process dynamic data and perform some operation and maintenance operations (such as permission application, database writing, etc.). For example, when a user fails to start a container instance on the xx machine learning platform of a music software or music application (APP), it may be because the number of available GPUs is insufficient. When the user consults the Q&A system, the real-time dynamic information should be informed to the user, but this kind of information cannot be obtained from the preprocessed static text data, resulting in the Q&A system being unable to answer the questions raised by the user or answering incorrectly, etc.

[0081] In view of this, the embodiments of the present disclosure provide a dialogue processing method. Through this method, the first dialogue information can be determined; then the first dialogue information is matched with the information of multiple associated tools; further, if the first dialogue information matches the information of the target tool among the multiple tools, a tool function is generated according to the target tool and the first dialogue information, and the target tool is called to execute the tool function to obtain the reply information corresponding to the first dialogue information. It can be seen that since there are multiple tools provided in the embodiments of the present disclosure, the multiple tools include different types of knowledge bases and different types of atomic capability apis. In this way, when the question raised by the user is related to a certain knowledge base, the data of the most relevant knowledge base tool can be automatically retrieved to assist in obtaining high-quality reply information; when the question raised by the user is not related to the knowledge base but involves some operation and maintenance operations (such as related issues such as platform usage permission application, database query, etc.), the relevant atomic capability api tool can be automatically called to perform corresponding operations; when the question raised by the user is not related to the knowledge base or operation and maintenance operations, the answer can be refused or just chat with the user, thereby improving the accuracy and effectiveness of the reply to the question raised by the user and enhancing the user experience.

[0082] After introducing the basic principle of the present disclosure, the various non-limiting embodiments of the present disclosure will be specifically introduced below.

[0083] Overview of Application Scenarios

[0084] To better understand the technical solutions provided by the embodiments of the present disclosure, the application scenarios applicable to the technical solutions provided by the embodiments of the present disclosure will be briefly introduced below. It should be noted that the application scenarios introduced below are only used to illustrate the embodiments of the present disclosure rather than limiting them. In specific implementation, the technical solutions provided by the embodiments of the present disclosure can be flexibly applied according to actual needs.

[0085] In the embodiments of the present disclosure, the dialogue processing technology can be applied to various business scenarios of dialogue processing. For example, in the xx machine learning platform of a music software or music APP, when replying to questions input by users, or in other question-and-answer systems or reply systems. The embodiments of the present disclosure do not limit this.

[0086] Please refer to Figure 2 as shown in Figure 2 This is an application scenario to which the technical solution of the embodiments of the present disclosure can be applied. In this scenario schematic diagram, it includes device 101 deployed with an intelligent operation and maintenance Q&A system, device 102 deployed with tools such as atomic capability APIs, device 103 deployed with a knowledge base, and terminal device 104. Among them, device 101 deployed with an intelligent operation and maintenance Q&A system can be connected to multiple device 102, multiple device 103, and multiple terminal devices 104. Multiple terminal devices 104 are the devices corresponding to users who submit dialogue requests or questions. Only one is shown in Figure 2 this figure.

[0087] Among them, device 101 deployed with an intelligent operation and maintenance Q&A system, device 102 deployed with tools such as atomic capability APIs, device 103 deployed with a knowledge base, and terminal device 104, and between each device can be directly or indirectly communicatively connected through one or more networks.

[0088] In the embodiments of the present disclosure, the intelligent operation and maintenance Q&A system integrates the RAG technology and related atomic capability APIs based on the LLM. That is to say, the intelligent operation and maintenance Q&A system provided in the embodiments of the present disclosure can not only answer questions about multiple knowledge bases, but also perform intelligent operation and maintenance operations, and even can control the answer template to a certain extent.

[0089] In the embodiments of the present disclosure, according to the actual operation requirements or Q&A requirements in implementation, document segmentation, vectorization, vector database entry operations, etc. can be performed according to multiple category knowledge bases, and each knowledge base is provided with a corresponding detailed introduction. Among them, the detailed introduction can include content description, retrieval timing, etc. of the knowledge base. Of course, the detailed introduction can also supplement other useful information, such as the disabled time, update time, etc. of the knowledge base. The embodiments of the present disclosure do not limit this.

[0090] For example, assume that the intelligent operation and maintenance system is associated with three knowledge bases, namely the first knowledge base, the second knowledge base, and the third knowledge base, and the corresponding introductions of the first knowledge base, the second knowledge base, and the third knowledge base are: "Introduction to the Use of the Platform Container Application Module", "Introduction to the Use of the Platform Operator Management Module", "Introduction to the Persons in Charge of Each Module of the Platform". In this way, when the question asked by the user is: "Who is the person in charge of Module A", the LLM can automatically retrieve the third knowledge base according to the semantic matching degree.

[0091] In the embodiments of the present disclosure, the atomic capability API can be integrated in the form of Python functions, and the required input parameters (including the input parameter names, meanings, and types) and the introduction of the functions can be defined. The LLM can extract the input parameters from the questions raised by the user and execute the relevant APIs to obtain data, and finally return the answers to the questions raised by the user according to a fixed template. For example, a function container_fail is predefined in advance. When the user consults questions related to the failure of container startup, the LLM will call this function and extract the container instance name from the questions consulted by the user (optionally, if not provided, the user will be reminded to provide it in the conversation) to query the specific reason. Based on this function, the two atomic capability APIs, check_status and check_resource, can be called successively. These two atomic capability APIs respectively perform real-time queries on the container status and the required resources, obtain the latest information, and feedback the obtained information as the answer to the user according to a fixed template.

[0092] It can be seen that in the embodiments of the present disclosure, the LLM will regard the above knowledge base and Python functions as available tools. When the user raises a question, it automatically determines whether to call a tool and which tool to call. Optionally, when it is determined that a tool needs to be called, the corresponding tool name can be output and the corresponding input parameters can be extracted, and the answer output by the tool can be obtained after execution. In addition, the LLM can think about the answer to determine whether it can answer the user's question. If it meets the requirements, it will be directly returned; if it does not meet the requirements, it will try to use other tools to obtain more information and answer again until the final answer is obtained or the maximum number of tool calls is reached and then the answer is refused. Optionally, when the LLM determines that no tool needs to be called, the LLM will use the world knowledge contained in itself to communicate with the user to complete the reply to the user.

[0093] In the embodiments of the present disclosure, the user can log in to the corresponding intelligent operation and maintenance Q&A platform based on the terminal device 104, and then trigger the question information. Thus, the device 101 deployed with the intelligent operation and maintenance Q&A system can receive the question information and determine the conversation information, and can determine the first conversation information; then match the first conversation information with the information of multiple associated tools; further, if the first conversation information matches the information of the target tool among the multiple tools, a tool function is generated according to the target tool and the first conversation information, and the device 102 or the device 103 is linked to call the target tool to execute the tool function to obtain the reply information corresponding to the first conversation information.

[0094] Among them, Figure 2The devices 101 with an intelligent operation and maintenance Q&A system deployed therein, the devices 102 with tools such as atomic capability APIs deployed therein, and the devices 103 with a knowledge base deployed therein can also be independent physical servers, or server clusters or distributed systems composed of multiple physical servers, or cloud servers or cloud server clusters that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, but are not limited thereto.

[0095] Figure 2 The terminal device 104 therein can be a mobile phone, a tablet computer (PAD), a personal computer (PC), a smart TV, a smart watch, a smart speaker, a smart vehicle-mounted device, a wearable device, etc., but is not limited thereto.

[0096] Of course, the method provided by the embodiments of the present disclosure is not limited to Figure 2 the application scenarios shown, and can also be used in other possible application scenarios, and the embodiments of the present disclosure do not impose any restrictions.

[0097] Exemplary Method

[0098] To further illustrate the technical solutions provided by the embodiments of the present disclosure, the following will be described in detail in conjunction with the accompanying drawings and specific implementation manners. Although the embodiments of the present disclosure provide method operation steps as shown in the following embodiments or drawings, more or fewer operation steps may be included in the method based on routine or non-creative labor. In steps where there is no necessary causal relationship logically, the execution order of these steps is not limited to the execution order provided by the embodiments of the present disclosure. When the method is actually processed or executed by a device, it can be executed in the method order shown in the embodiments or drawings or executed in parallel.

[0099] Please refer to Figure 3 , Figure 3 which is a schematic flowchart of a dialogue processing method in the embodiments of the present disclosure. The process of the method can be executed by an electronic device, such as Figure 2 the device 101 with an intelligent operation and maintenance Q&A system deployed therein as shown therein. The specific implementation process of the method is as follows:

[0100] Step 301: Determine the first dialogue information.

[0101] In the embodiments of the present disclosure, an electronic device may receive current problem information and then determine historical conversation information before the current problem information; wherein the historical conversation information includes problem information before the current problem information and corresponding answer information; thus, first conversation information may be determined according to the current problem information and the conversation history information before the current problem information.

[0102] For example, refer to Figure 4 as shown in Figure 4 which is a schematic diagram of conversation information provided by the embodiments of the present disclosure. In Figure 4 "Llm-lch" is the current problem information received by the electronic device, then Figure 4 in it, "Please correctly enter the name of your container instance. You can find the instance name in the basic information of the xxxxxxxxxxxx instance list", "My container cannot be started", and "Hello, I am the xx machine learning platform Q&A assistant s, currently in the testing phase, and can help you query the person in charge of a specific field and the problem of container startup failure. You can ask: 1: The person in charge of the relevant field. For example, who is the person in charge of datahub? 2: The reason for container startup failure. For example, why does the container instance lch fail to start?" are historical conversation information, and thus the first conversation information can be determined.

[0103] Optionally, if the electronic device determines that there is no historical conversation information before the current problem information, the current problem information may be used as the first conversation information.

[0104] It can be seen that in the embodiments of the present disclosure, the electronic device not only considers the problem information currently proposed by the user, but also considers the historical conversation information between the user and the intelligent Q&A system, and thus can comprehensively determine the problem information that the user really wants to propose by integrating the historical conversation information, avoiding the situation of incorrect answers caused by inaccurate or incomplete problem descriptions, and improving the user experience.

[0105] Step 302: Match the first conversation information with the information of multiple associated tools.

[0106] In the embodiments of the present disclosure, the intelligent operation and maintenance Q&A system may determine the type information of the first conversation information, and then screen out candidate tools from multiple tools according to the type information, and match the first conversation information with the information of the candidate tools.

[0107] In the embodiments of the present disclosure, the intelligent operation and maintenance Q&A system may screen out candidate tools from multiple tools according to the type information and a prompt word template; wherein the prompt word template is configured for the large language model and includes the tool names, tool input parameters, call formats, output formats, and tool introductions of all tools.

[0108] For example, please refer toFigure 5 , Figure 5 This is a schematic diagram of the prompt template in the embodiments of the present disclosure. Among them, the prompt template includes an introduction to the assistant of the intelligent operation and maintenance Q&A system, such as Figure 5 in "{##Role; You are an assistant for answering questions on the machine learning platform. Your purpose is to answer and solve users' questions. When the user's question is not related to the platform, please chat with the user based on prior knowledge.}".

[0109] Among them, the prompt template includes an introduction to all tools of the intelligent operation and maintenance Q&A system, such as Figure 5 in "{##Tools You can use a variety of tools. You need to plan the use of these tools by yourself and complete the task at hand in the order you think is appropriate. This may require decomposing the task into subtasks and using different tools to complete each subtask}".

[0110] Among them, the tool names included in the prompt template are, for example Figure 5 in "{tool desc}".

[0111] Among them, the tool call format included in the prompt template is, for example Figure 5 in "When answering questions, if you need to use tools, please use the following format, Thought: Tool name (one of {tool names}) Action: Tool name (one of {tool_names}) Action Input: The input of the tool, represented in JSON format as kwargs (e.g., {{"input": "hello world", "num_beams": 5}}). Please ***always*** start your answer with Thought. Please use valid JSON format as the action input. Do not write it like {{"input": "hello world", 'num_beams': 5}}. If you use this format, the user will respond in the following format: Observation: Tool response".

[0112] Among them, the tool output format included in the prompt template is, for example Figure 5 in "Respond in one of the following two formats: Thought: I can answer the user's question now. Answer: [Your answer]; Thought: I cannot answer the user's question. Answer: Sorry, I can't solve your problem temporarily. I can provide you with the person in charge of the corresponding field. Which field are you consulting?".

[0113] In the embodiments of the present disclosure, if the type information is the operation and maintenance operation type, candidate tools of the atomic capability api type are filtered according to the operation and maintenance operation type and the tool introduction in the prompt template;

[0114] In an embodiment of the present disclosure, if the type information is of the static text type, candidate tools of the knowledge base type are filtered out according to the static text type and the tool introduction in the prompt template.

[0115] Step 303: If the first conversation information matches the information of the target tool among multiple tools, a tool function is generated according to the target tool and the first conversation information, and the target tool is called to execute the tool function to obtain a reply message corresponding to the first conversation information.

[0116] In an embodiment of the present disclosure, if the first conversation information does not match the information of any tool among multiple tools, the first conversation information is analyzed based on a large language model to obtain a reply message corresponding to the first conversation information.

[0117] In an embodiment of the present disclosure, the target tool can be called to execute the tool function to obtain a tool result; when it is determined to execute the direct return working mode, the tool result is used as the reply message; when it is currently determined to execute the cache working mode, the tool result is used as the historical first conversation information, and the processing of the second conversation information is continued to obtain a reply message corresponding to the first conversation information, where the second conversation information is obtained based on the association of the first conversation information.

[0118] For example, please refer to Figure 6 , Figure 6 which is a schematic diagram of answering questions of an intelligent operation and maintenance question answering system provided by an embodiment of the present disclosure.

[0119] In an embodiment of the present disclosure, multiple constructed knowledge bases and related atomic capability apis can be integrated into the form of tools (python functions), and information such as tool names, tool introductions, and tool input parameters can be pre-configured into the intelligent operation and maintenance question answering system. At the same time, the working modes of each tool can also be configured, and the working modes are, for example, the direct return working mode or the cache working mode. For example, the working mode of tool A is configured as the direct return working mode, that is to say, the result fed back by tool A can be directly used as the answer to be returned to the user. Obviously, based on this intelligent operation and maintenance question answering system, the template of the answer can be controlled, increasing the controllability of the answer. For example, the direct return working mode can be executed for tools similar to "the knowledge base corresponding to the introduction of the person in charge of each module of the platform".

[0120] In other words, in the embodiments of the present disclosure, for non-complex scenarios, the intelligent operation and maintenance Q&A system can configure the working mode of the corresponding tool as the direct return working mode. The direct return working mode is, for example, a working mode in which the reply result determined by the tool is directly returned to the user. In this way, the overall time consumption can be significantly reduced and the response efficiency can be improved. For example, since at least one LLM call is reduced, the overall time consumption can be reduced by about 33% - 50%, thereby improving the response efficiency to the questions raised by users and further enhancing the user experience.

[0121] In addition, the LLM central prompt word template can be configured in advance, and the information of all tools can be injected into the prompt word template, so that the intelligent operation and maintenance Q&A system can select the corresponding tool based on the prompt word template for the questions raised by users in the later stage.

[0122] In the embodiments of the present disclosure, when the intelligent operation and maintenance Q&A system receives the question information of the user, it can determine the first conversation information corresponding to the user, that is, Figure 6 the conversation history therein, and then based on the conversation history and the LLM central prompt word template, determine whether to use the tool. When the tool needs to be used, the LLM will perform semantic judgment and output the tool information. The intelligent operation and maintenance Q&A system will extract the tool name and tool input parameters through regular expressions, and then execute the corresponding tool function to obtain the tool result. If the tool is configured with direct return, it will be directly returned to the user as an answer, otherwise it will continue to execute as a part of the conversation history until the final answer is determined and feedback to the user. When the tool does not need to be used, the intelligent operation and maintenance Q&A system extracts the final answer after Answer through regular expressions and returns it to the user.

[0123] In the embodiments of the present disclosure, since the time consumption of calling the LLM is almost proportional to the number of output characters, in order to reduce the overall time consumption, the intelligent operation and maintenance Q&A system can limit the number of output characters of the LLM to be as small as possible while the key information is as much as possible, so as to reduce the overall time consumption. For example, considering that the actually required LLM output is the tool name and tool input parameters, it can be controlled that the part of Figure 6 Thought in is as small as possible, and even can be directly removed.

[0124] Exemplary Device

[0125] The exemplary embodiments of the present disclosure also provide a dialogue processing device. Referring to Figure 7 shown, the dialogue processing device 700 includes the following program units:

[0126] A determination unit 701, configured to determine the first conversation information;

[0127] A matching unit 702, configured to match the first conversation information with the information of a plurality of associated tools;

[0128] A processing unit 703, configured to, if the first conversation information matches the information of a target tool among the multiple tools, generate a tool function according to the target tool and the first conversation information, and call the target tool to execute the tool function to obtain a reply message corresponding to the first conversation information.

[0129] In a possible implementation manner, the processing unit 703 is further configured to:

[0130] If the first conversation information does not match the information of any tool among the multiple tools, analyze the first conversation information based on a large language model to obtain a reply message corresponding to the first conversation information.

[0131] In a possible implementation manner, the determining unit 701 is specifically configured to:

[0132] Receive the current problem information;

[0133] Determine the historical conversation information before the current problem information; the historical conversation information includes the problem information and the corresponding answer information before the current problem information;

[0134] Determine the first conversation information according to the current problem information and the conversation history information before the current problem information.

[0135] In a possible implementation manner, the matching unit 702 is specifically configured to:

[0136] Determine the type information of the first conversation information;

[0137] Screen out candidate tools from the multiple tools according to the type information;

[0138] Match the first conversation information with the information of the candidate tools.

[0139] In a possible implementation manner, the matching unit 702 is specifically configured to:

[0140] Screen out candidate tools from the multiple tools according to the type information and a prompt template;

[0141] Wherein, the prompt template is configured for the large language model and includes the tool names, tool input parameters, call formats, output formats, and tool introductions of all tools.

[0142] In a possible implementation manner, the matching unit 702 is specifically configured to:

[0143] If the type information is an operation and maintenance operation type, candidate tools of the atomic capability API type are filtered according to the operation and maintenance operation type and the tool introduction in the prompt word template;

[0144] If the type information is a static text type, candidate tools of the knowledge base type are filtered according to the static text type and the tool introduction in the prompt word template.

[0145] In a possible implementation manner, the processing unit 703 is specifically configured to:

[0146] Call the target tool to execute the tool function to obtain a tool result;

[0147] When it is determined to execute the direct return working mode, the tool result is used as the reply information;

[0148] When it is currently determined to execute the cache working mode, the tool result is used as the historical first conversation information, and the processing of the second conversation information is continued to obtain the reply information corresponding to the first conversation information; the second conversation information is obtained based on the first conversation information.

[0149] The specific details of each part in the above device have been described in detail in the implementation manner of the method part. The details not disclosed can be seen in the implementation manner content of the method part, and thus will not be repeated.

[0150] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the exemplary embodiments of the present disclosure, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0151] Exemplary Program Product

[0152] The exemplary embodiments of the present disclosure further provide a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the above-mentioned conversation processing method is implemented.

[0153] In one embodiment, a computer program product may be a tangible product containing a computer program, such as a computer-readable storage medium storing the computer program. The readable storage medium may be a storage medium based on signals such as electricity, magnetism, light, electromagnetic, infrared, etc., including but not limited to: random access memory (RAM), read-only memory (ROM), magnetic tape, floppy disk, flash memory, mechanical hard disk drive (HDD), solid state drive (SSD), and so on. Exemplarily, the computer program product may be implemented as a non-volatile storage medium storing the computer program, such as read-only memory, NAND flash memory, etc.

[0154] In one embodiment, a computer program product may be an intangible product containing a computer program. Exemplarily, the computer program product may be implemented as a virtual digital product, such as an executable file storing the computer program, digital files such as installation packages.

[0155] The code of the computer program can be written in one or more programming languages. Programming languages such as C, Java, C++, etc. The program code can be executed entirely on the user's computing device, or partially on the user's computing device, or executed as an independent software package, or partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device through any type of network, such as a local area network (LAN), wide area network (WAN), etc., or can be connected to an external computing device (e.g., through an Internet connection provided by an operator).

[0156] The computer program can be carried or transmitted by signals such as electricity, magnetism, light, electromagnetic, infrared, etc. The electronic device can convert the signal carrying the computer program into a digital signal and then run the computer program. When the computer program runs on the electronic device, its code is used to cause the electronic device to execute (more specifically, can cause the processor of the electronic device to execute) the method steps of various exemplary embodiments of the present disclosure. For example, it can execute the above-mentioned dialogue processing method, which includes the following steps: Step 301: Determine the first dialogue information. Step 302: Match the first dialogue information with the information of multiple associated tools. Step 303: If the first dialogue information matches the information of the target tool among the multiple tools, generate a tool function based on the target tool and the first dialogue information, and call the target tool to execute the tool function to obtain the reply information corresponding to the first dialogue information.

[0157] By implementing the above method steps through a computer program, the first conversation information can be determined; then, the first conversation information is matched with the information of multiple associated tools; further, if the first conversation information matches the information of the target tool among the multiple tools, a tool function is generated based on the target tool and the first conversation information, and the target tool is called to execute the tool function to obtain the reply information corresponding to the first conversation information. It can be seen that since multiple tools are provided in the embodiments of the present disclosure, the multiple tools include different types of knowledge bases and different types of atomic capability apis. In this way, when the question raised by the user is related to a certain knowledge base, the data of the most relevant knowledge base tool can be automatically retrieved to assist in obtaining high-quality reply information; when the question raised by the user is not related to the knowledge base but involves some operation and maintenance operations (such as related issues such as platform usage permission application and database query), the relevant atomic capability api tool can be automatically called to perform corresponding operations; when the question raised by the user is not related to the knowledge base or operation and maintenance operations, the answer can be refused or just chat with the user, thereby improving the accuracy and effectiveness of the reply to the question raised by the user and enhancing the user experience.

[0158] Exemplary Electronic Device

[0159] An exemplary embodiment of the present disclosure also provides an electronic device. The electronic device may include a processor and a memory. The memory stores executable instructions of the processor, such as a computer program. The processor executes the executable instructions to execute the method steps of various exemplary embodiments of the present disclosure.

[0160] The following refers to Figure 8 , and the electronic device is exemplarily described in the form of a general computing device. It should be understood that Figure 8 The electronic device 800 shown is only an example and should not impose limitations on the functions and usage scope of the embodiments of the present disclosure.

[0161] As Figure 8 shown, the electronic device 800 may include: a processor 810, a memory 820, a bus 830, an I / O (input / output) interface 840, and a network adapter 850.

[0162] The memory 820 may include volatile memory, such as RAM 821 and a cache unit 822, and may also include non-volatile memory, such as ROM 823. The memory 820 may also include one or more program modules 824. Such program modules 828 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment. For example, the program module 824 may include each unit in the above device.

[0163] The processor 810 may include one or more processing units. For example, the processor 810 may include an AP (Application Processor), a modem processor, a GPU (Graphics Processing Unit), an ISP (Image Signal Processor), a controller, an encoder, a decoder, a DSP (Digital Signal Processor), a baseband processor, and / or an NPU (Neural-Network Processing Unit), etc.

[0164] The processor 810 can be used to execute the executable instructions stored in the memory 820. For example, it can execute the above-mentioned dialogue processing method, which includes the following steps: Step 301: Determine the first dialogue information. Step 302: Match the first dialogue information with the information of multiple associated tools. Step 303: If the first dialogue information matches the information of the target tool among the multiple tools, generate a tool function according to the target tool and the first dialogue information, and call the target tool to execute the tool function to obtain the reply information corresponding to the first dialogue information.

[0165] By executing the above method steps through the processor 810, the first dialogue information can be determined; then the first dialogue information is matched with the information of multiple associated tools; further, if the first dialogue information matches the information of the target tool among the multiple tools, a tool function is generated according to the target tool and the first dialogue information, and the target tool is called to execute the tool function to obtain the reply information corresponding to the first dialogue information. It can be seen that since there are multiple tools provided in the embodiments of the present disclosure, and these multiple tools include different types of knowledge bases and different types of atomic capability apis, in this way, when the question raised by the user is related to a certain knowledge base, the data of the most relevant knowledge base tool can be automatically retrieved to assist in obtaining high-quality reply information; when the question raised by the user is not related to the knowledge base but involves some operation and maintenance operations (such as related issues like platform usage permission application, database query, etc.), the relevant atomic capability api tool can be automatically called to perform corresponding operations; when the question raised by the user is neither related to the knowledge base nor related to operation and maintenance operations, the answer can be refused or just chat with the user, thereby improving the accuracy and effectiveness of the reply to the question raised by the user and enhancing the user experience.

[0166] The bus 830 is used to implement the connection between different components of the electronic device 800 and may include a data bus, an address bus, and a control bus.

[0167] The electronic device 800 can communicate with one or more external devices 900 (such as a keyboard, a mouse, an external controller, etc.) through the I / O interface 840.

[0168] The electronic device 800 can communicate with one or more networks through the network adapter 850. For example, the network adapter 850 can provide mobile communication solutions such as 3G / 4G / 5G, or wireless communication solutions such as wireless local area network, Bluetooth, near field communication, etc. The network adapter 850 can communicate with other modules of the electronic device 800 through the bus 830.

[0169] Although Figure 8 not shown in the figure, other hardware and / or software modules can also be provided in the electronic device 800, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0170] As can be seen from the above, the technical solution of the present disclosure can be implemented as a method, a device, a system, a computer program product, a storage medium, an electronic device, etc. Those skilled in the art can understand that various aspects of the present disclosure can be specifically implemented in the following forms, namely: a complete hardware implementation manner, a complete software implementation manner (including firmware, microcode, etc.), or an implementation manner combining hardware and software aspects, such as can be respectively referred to as "circuit", "module" or "system".

[0171] It should be understood that the present disclosure is not limited to the specific method steps or structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. Based on the specific implementation manners provided by the present disclosure, those skilled in the art will easily think of other implementation manners. Therefore, the specific implementation manners provided by the present disclosure are only exemplary, and the scope and spirit of the present disclosure are pointed out by the claims, and should cover any variations, uses or adaptive changes of the present disclosure, and these variations, uses or adaptive changes follow the general principles of the present disclosure and include the well-known common knowledge or conventional technical means in the technical field not disclosed by the present disclosure.

Claims

1. A method for processing a conversation, characterized in that: The method comprises: Determining first conversation information; matching the first conversation information with information of the associated multiple tools; If the first dialogue information matches the information of a target tool among the multiple tools, a tool function is generated according to the target tool and the first dialogue information, and the target tool is called to execute the tool function to obtain reply information corresponding to the first dialogue information.

2. The method according to claim 1, characterized in that: The method further comprises: If the first dialogue information does not match the information of any tool among the multiple tools, the first dialogue information is analyzed based on the large language model to obtain reply information corresponding to the first dialogue information.

3. The method according to claim 1, characterized in that Determine the first conversation information, including: Receive information about current issues; Determine historical dialogue information before the current question information; the historical dialogue information includes question information before the current question information and corresponding answer information; The first dialogue information is determined according to the current question information and dialogue history information before the current question information.

4. The method according to claim 1, characterized in that Matching the first conversation information with information of multiple associated tools includes: Determining type information of the first dialogue information; Filtering candidate tools from the multiple tools according to the type information; The first conversation information is matched with the information of the candidate tool.

5. The method according to claim 4, characterized in that Screening out candidate tools from the multiple tools according to the type information includes: Filtering candidate tools from the multiple tools according to the type information and the prompt word template; The prompt word template is configured for a large language model and includes tool names, tool input parameters, calling formats, output formats, and tool introductions of all tools.

6. The method according to claim 5, characterized in that Screening out candidate tools from the multiple tools according to the type information and the prompt word template includes: If the type information is an operation type, screening out candidate tools of the atomic capability api type according to the operation type and the tool introduction in the prompt word template; If the type information is of a static text type, candidate tools of the knowledge base type are screened out according to the static text type and the tool introduction in the prompt word template.

7. The method according to any one of claims 1 to 6, characterized in that: Calling the target tool to execute a tool function to obtain reply information corresponding to the first dialogue information includes: Calling the target tool to execute the tool function and obtain the tool result; When it is determined to execute the direct return to the working mode, the tool result is used as the reply information; At present, it is determined to execute the cache working mode, then the tool result is used as the historical first dialogue information, and the processing of the second dialogue information is continued to obtain the reply information corresponding to the first dialogue information; the second dialogue information is obtained based on the association of the first dialogue information.

8. A dialogue processing device, characterized in that: The device comprises: A determining unit, configured to determine first dialogue information; A matching unit, configured to match the first dialogue information with information of a plurality of associated tools; A processing unit is used to generate a tool function according to the target tool and the first dialogue information if the first dialogue information matches the information of the target tool among the multiple tools, and call the target tool to execute the tool function to obtain reply information corresponding to the first dialogue information.

9. An electronic device, characterized in that: include: processor; A memory, configured to store executable instructions of the processor; The processor is configured to perform the method of any one of claims 1 to 7 by executing the executable instructions.

10. A computer program product having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.