Virtual dialogue system performance evaluation and enrichment

CN116806339BActive Publication Date: 2026-09-11INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202280010890.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-01-29
Filing Date
2022-01-04
Publication Date
2026-09-11
Estimated Expiration
2042-01-04

AI Technical Summary

Technical Problem

例如,缺少同义词或概念关系可限制AI平台确定由客户或客户端输入的问题等效于或相关于数据集或数据库内可得到答案的已知问题的能力

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116806339B_ABST
    Figure CN116806339B_ABST
Patent Text Reader

Abstract

Embodiments are provided relating to computer systems, computer program products, and computer-implemented methods for improving performance of a virtual dialog agent system employing an automated virtual dialog agent. Embodiments relate to generating ground truth (GT) from a user's knowledge base, and utilizing the GT to evaluate performance of the virtual dialog agent using the GT. The evaluation measures quality of multi-turn virtual dialog, and generates a remediation plan for algorithmic improvements to the virtual dialog agent.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] One or more embodiments of the present invention relate to a virtual dialogue system employing an automated virtual dialogue agent (such as a chatbot), and related computer program products and computer-implemented methods. In a particular exemplary embodiment, quality criteria for corresponding automated virtual dialogue agent interactions are evaluated and selectively parsed, the parsing being aimed at selectively applying one or more remedial actions to the automated virtual dialogue agent, for example, to improve performance regarding natural language (NL) dialogue events.

[0002] Automated virtual conversational agents use artificial intelligence (AI) as a platform to perform NL interactions between automated virtual conversational agents and users, typically such as consumers or clients, or even another conversational agent. Interactions can involve product sales, customer service, information retrieval, or other types of interactions or transactions. Chatbots interact with users through dialogue, which is typically text-based (e.g., online or via text) or auditory (e.g., via telephone). Chatbots are known in the art to act as a question-and-answer component between the user and the AI ​​platform. The quality of the question (or query) and answer (or response) is derived from the quality of question understanding, question transformation, and answer parsing. Common reasons for failing to meet quality standards are often found in inappropriate or inefficient question generation that requests a response. This could be due to a lack of knowledge to effectively transform the question into an equivalent knowledge representation mapped to the answer, or it could be due to inefficiencies within the AI ​​platform or chatbot. For example, a lack of synonyms or conceptual relationships can limit the AI ​​platform's ability to determine that a question entered by a customer or client is equivalent to or related to a known question within a dataset or database that has an answer.

[0003] Enterprises can specify particular requirements for virtual assistance that they expect to be met before the commercial deployment of the virtual system, such as accuracy or interaction quality. For example, a virtual system might have a minimum performance requirement of 50% accuracy for supporting a proxy user base, or a minimum performance requirement of 90% accuracy for an end-user base. Therefore, it is desirable to benchmark or quality-test the dialogue system before deployment. Summary of the Invention

[0004] Examples include systems, computer program products, and methods for improving the performance of dialog systems. This "Summary" is provided to introduce, in a simplified form, a selection of representative concepts further described in the following "Detailed Description." This "Summary" is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter in any way.

[0005] In one aspect, a computer system is provided having a processor operatively coupled to memory and an artificial intelligence (AI) platform operatively coupled to the processor. The AI ​​platform includes one or more tools for improving the performance of a virtual dialogue agent. The tools include a ground truth (GT) manager, a simulator, an evaluation manager, and a remediation manager. The GT manager is configured to automatically generate GTs from a knowledge source. The simulator is configured to use the virtual dialogue agent to simulate NL dialogue interactions. More specifically, the simulator is configured to utilize the GTs to drive the output generated by the simulated NL dialogue and create a corresponding simulation log. The evaluation manager is configured to evaluate the performance of the virtual dialogue agent with respect to the simulation logs, taking the GTs into account. The remediation manager is configured to identify and selectively implement one or more remedial actions for the dialogue system, taking performance thresholds into account.

[0006] In another aspect, a computer program product is provided, having a computer-readable storage medium and program code stored on the computer-readable storage medium. The program code is executable by a computer processor to improve the performance of a virtual dialogue agent. Program code is provided to automatically generate ground truth (GT) from a knowledge source. Program code is also provided to simulate NL dialogue interactions using the virtual dialogue agent. The simulation utilizes the GT to drive the output generated by the simulated NL dialogue and to create a corresponding simulation log. Program code is provided to evaluate the performance of the virtual dialogue agent with respect to the simulation log, taking the GT into account, and to identify and selectively implement one or more remedial actions for the dialogue system, taking performance thresholds into account.

[0007] In another aspect, a computer-implemented method for improving the performance of a virtual dialogue agent is provided. The method is configured to automatically generate ground truth (GT) from a knowledge source. A non-linear dialogue interaction undergoes simulation using the virtual dialogue agent. The simulation utilizes the GT to drive the output generated by the simulated NL dialogue interaction and to create a corresponding simulation log. The performance of the virtual dialogue agent with respect to the simulation log is evaluated, taking the GT into account. One or more remedial actions for the dialogue system are identified and selectively implemented, taking performance thresholds into account.

[0008] These and other features and advantages will become apparent from the following detailed description of the present exemplary embodiments taken in conjunction with the accompanying drawings. Attached Figure Description

[0009] The accompanying drawings, which are referenced herein, form part of the specification and are incorporated herein by reference. Unless otherwise indicated, the features shown in the drawings are intended to illustrate some embodiments only, and not all embodiments.

[0010] Figure 1 A system diagram illustrating an artificial intelligence platform computing system in a network environment is presented.

[0011] Figure 2 Depicting as shown Figure 1 A block diagram showing and describing the artificial intelligence platform tools and their associated application programming interfaces.

[0012] Figure 3 A flowchart illustrating an embodiment of a method for automatically generating ground reality from corresponding knowledge sources is provided.

[0013] Figure 4 A flowchart illustrating an embodiment of the method for generating GT based on the used GT is depicted.

[0014] Figure 5 A flowchart illustrating an embodiment of a method for generating a curation-based ground truth (GT) is provided.

[0015] Figure 6 A flowchart illustrating an embodiment of a method for simulating and interacting with a dialogue system is provided.

[0016] Figure 7 A flowchart illustrating an embodiment of a method for evaluating the performance of a virtual dialogue system is provided.

[0017] Figure 8 A block diagram illustrating an example of a cloud-based support computer system / server is provided, which is used to implement the above-mentioned... Figure 1-7 The system and process described.

[0018] Figure 9 A block diagram illustrating a cloud computing environment is depicted.

[0019] Figure 10 A block diagram illustrating a set of functional abstraction model layers provided by a cloud computing environment is depicted. Detailed Implementation

[0020] It will be readily understood that the components of embodiments of the invention generally described and illustrated in the accompanying drawings can be arranged and designed in various different configurations. Therefore, the following detailed description of exemplary embodiments of the apparatuses, systems, methods, and computer program products presented in the drawings is not intended to limit the scope of the claimed embodiments, but merely to illustrate selected embodiments.

[0021] Throughout this specification, references to "selected embodiment," "an embodiment," "exemplary embodiment," or "embodiment" refer to a specific feature, structure, or characteristic described in connection with that embodiment being included in at least one embodiment. Therefore, the phrases "selected embodiment," "in an embodiment," "in an exemplary embodiment," or "in an embodiment" appearing in various places throughout this specification do not necessarily refer to the same embodiment. The embodiments described herein can be combined with each other and modified to include features of each other. Furthermore, the features, structures, or characteristics described in the various embodiments can be combined and modified in any suitable manner.

[0022] The illustrated embodiments will be best understood by referring to the accompanying drawings, in which like parts are designated by like reference numerals throughout. The following description is intended only as an example and simply illustrates specific selections of devices, systems, products, and processes consistent with the embodiments claimed herein.

[0023] In the field of artificial intelligence computer systems, natural language systems (such as IBM...) Artificial intelligence computer systems or other natural language systems process natural language based on knowledge acquired by the system. To process natural language, the system can be trained with data derived from databases or knowledge corpora; however, the results may be incorrect or inaccurate for various reasons.

[0024] Machine learning (ML) (a subset of artificial intelligence (AI)) uses algorithms to learn from data and create visions based on that data. AI refers to the intelligence of a machine when it can make decisions based on information, maximizing the chances of success on a given subject. More specifically, AI can learn from datasets to solve problems and provide relevant recommendations. Cognitive computing is a hybrid of computer science and cognitive science. Cognitive computing utilizes self-learning algorithms that use data minimization, visual recognition, and natural language processing to solve problems and optimize human processes.

[0025] At the heart of AI and related reasoning lies the concept of similarity. Understanding natural language and objects requires reasoning from the perspective of potentially challenging relationships. Structures (including static and dynamic structures) prescribe a definite output or action for a given, definite input. More specifically, the determined output or action is based on explicit or inherent relationships within the structure. This arrangement can be satisfactory for the chosen situations and conditions. However, it should be understood that dynamic structures inherently undergo change, and the output or action can change accordingly. Existing solutions for effectively identifying objects and understanding natural language, as well as handling responses to the identified and understood content and changes in structure, are extremely difficult at the practical level.

[0026] An automated virtual agent (referred to herein as a chatbot) is an artificial intelligence (AI) program that simulates interactive human conversation using pre-computed phrases and auditory or text-based signals. Chatbots are increasingly used in electronic platforms for customer service support. In one embodiment, a chatbot can function as an intelligent virtual agent. Each chatbot experience comprises a set of communications, including user actions and dialogue system actions, where the experience exhibits distinctive behavioral patterns. It should be understood in the art that chatbot conversations can be evaluated and diagnosed to identify elements of the chatbot that can be guaranteed to be changed to improve future chatbot experiences.

[0027] A system, computer program product, and method are disclosed that evaluate the performance of an automated virtual dialogue agent by automatically generating benchmark data (also referred to herein as ground truth (GT)) from a knowledge base, and, in an exemplary embodiment, evaluate the performance of a multi-turn dialogue system to evaluate the automated virtual dialogue agent. In the exemplary embodiment, the GT is automatically generated from a user's knowledge base, rather than from a standard or general dataset. Benchmark data generation serves as a venue for extracting the GT within the knowledge base. Simulated dialogue interactions are performed with the automated virtual dialogue agent supported by the GT. As shown and described herein, the automated virtual dialogue agent undergoes a performance evaluation involving comparing corresponding simulated logs while taking the GT into account. Metrics used to measure the evaluation of the automated virtual dialogue agent include, for example, determining the correctness of the automated virtual dialogue agent's responses, the relevance of the disambiguation or follow-up questions posed by the automated virtual dialogue agent, the number of disambiguation or follow-up questions posed by the automated virtual dialogue agent, and / or the order in which the disambiguation or follow-up questions posed by the automated virtual dialogue agent are performed.

[0028] An automated virtual dialogue agent (also referred to herein as a chatbot platform or chatbot) serves as the AI ​​interaction interface. As shown and described herein, the chatbot platform undergoes evaluation based on comparisons with simulated interactions with ground truth (GT). Ground truth (also referred to herein as GT) is a term used in machine learning that refers to information provided through direct observation (e.g., empirical evidence), as opposed to information provided by inference. As explained in more detail below, GT data includes, for example, content-based information (such as information from knowledge graphs or knowledge bases), usage logs (especially with feedback information about these logs), subject matter expert (SME) records, or any combination thereof.

[0029] refer to Figure 1A schematic diagram of an artificial intelligence (AI) platform and a corresponding system (100) is depicted. As shown, a server (110) is provided to communicate with multiple computing devices (180), (182), (184), (186), (188), and (190) across a network connection (e.g., a computer network (105)). The server (110) is configured with a processing unit (e.g., a processor) that communicates across a bus and memory. The server (110) is shown to have an AI platform (150) that is operatively coupled to a dialogue system (160), a corresponding virtual agent (162) (e.g., a chatbot), and an associated knowledge base (170) (e.g., a data source). The computing devices (180), (182), (184), (186), (188), and (190) may have a visual display, an audio interface, an audio-video interface, or other types of interfaces configured to allow users to interface with a representation of the virtual agent (e.g., a chatbot) (162).

[0030] The AI ​​platform (150) is operatively coupled to the network (105) to support interaction from one or more of the computing devices (180), (182), (184), (186), (188), and (190) with the virtual dialogue agent (162). More specifically, the computing devices (180), (182), (184), (186), (188), and (190) communicate with each other and with other devices or components via one or more wired and / or wireless data communication links, wherein each communication link may include one or more of wires, routers, switches, transmitters, receivers, etc. In this networking arrangement, the server (110) and the network connection (105) enable communication detection, identification, and resolution. Other embodiments of the server (110) may be used with components, systems, subsystems, and / or devices other than those described herein.

[0031] The AI ​​platform (150) is also shown here as being operationally coupled to a knowledge base (170) (also referred to herein as an information corpus). As shown in the figure, the knowledge base (170) is configured with multiple libraries, which are shown here as library A (172) as an example. A ) and Library B (172 B Although in Figure 1 Two libraries are shown, but it should be understood that a knowledge base (170) may include fewer or more libraries. Furthermore, multiple libraries (e.g., library A (172)) may also be present. A ) and Library B (172 B Multiple libraries (library A (172)) can be combined together. A ) and Library B (172 BA knowledge base (170) can exist across multiple knowledge domains (including knowledge bases (170) and other knowledge domains (not shown)). Each knowledge base is populated with data in structured or unstructured form. For example, in an exemplary embodiment, the structured data can take the form of a knowledge graph. By way of example, knowledge base A (172) A The knowledge graphs (KGs) are filled with structured knowledge domains represented as knowledge graphs (KGs). Figure 1 The middle is shown as KG0(172) A,0 ), KG1 (172) A,1 ) and KG2 (172 A,2 ).

[0032] The AI ​​platform (150) is shown herein as having multiple tools to support the evaluation, benchmarking, and improvement of the performance of the dialogue system (160) and the corresponding automated virtual agent (e.g., chatbot) (162). These tools include the GT Manager (152), the simulator (154), the evaluation manager (156), and the remediation manager (also referred to herein as the guide) (158).

[0033] GT Manager (152) is configured to draw from one or more knowledge sources (e.g., knowledge domains, which are shown here by way of example as library A (172)). A The generated ground truth (GT) is automatically generated. The generated GT can be content-based, usage-based, and / or regulatory-based. Content-based GTs are automatically generated by leveraging the corresponding structured dataset to generate issues based on symptoms, issue variations, and graph traversal, and to obtain relevant entities for the symptoms. In one exemplary embodiment, a symptom is a phrase describing some problems or issues concerning the system or any component of the system. Figure 3 The details of content-based ground truth (GT) generation are shown and described. Usage-based GT generation targets GT generation from usage logs, which take the form of records of collected data from workflows as query text. Figure 4 The details of the GT generation based on usage are shown and described. The regulatory-based GT targets data manually generated by subject matter experts (SMEs). In this embodiment, the SME provides assistance with selection options from a knowledge base to generate test data. Figure 5 The details of regulation-based GT generation are shown and described. Accordingly, content-based, usage-based, and regulation-based GTs each utilize corresponding knowledge domains in structured or unstructured formats to support and enable the automatic generation of GTs.

[0034] Each library is populated with one or more knowledge domains represented as structured knowledge, such as a knowledge graph processed by the GT manager (152). As illustrated by example, the first knowledge domain is represented as a first knowledge graph (KG) (in Figure 1The middle is shown as KG0(172) A,0 And is shown to have a corresponding content-based GT (GT 0,0 (174 0,0 )), based on the GT used (GT) 0,1 (174 0,1 )), and regulatory-based GT 0,2 (174 0,2 Similarly, the second knowledge domain is represented as the second KG (in Figure 1 The middle is shown as KG1(172) A,1 ), and is shown to have content-based GT (GT 1,0 (174 1,0 )), based on the GT used (GT) 1,1 (174 1,1 )), and regulatory-based GT 1,2 (174 1,2 ), and the third knowledge domain is represented as the third KG (in Figure 1 The middle is shown as KG2(172) A,2 ), and is shown to have content-based GT (GT 2,0 (174 2,0 )), based on the GT used (GT) 2,1 (174 2,1 )), and regulatory-based GT 2,2 (174 2,2 The three categories of GT shown herein (e.g., content, use, and regulation) can serve different roles in the evaluation of virtual agents. In one exemplary embodiment, any combination of the three categories of GT can be used for virtual agent evaluation. The number of GT categories shown herein should not be considered limiting. In one exemplary embodiment, the number of categories may include subsets of categories, combinations of categories, or new categories.

[0035] Interactions with the chatbot (162) take the form of a sequence of queries and corresponding responses, as well as subsequent disambiguation questions and their responses. Such interactions, and the data specifically associated with the interactions, are recorded and populated in one or more libraries of a knowledge base (170). In one or more embodiments, an initial NL query and a result are generated, taking into account the corresponding structured knowledge (e.g., KG). In one or more other embodiments, an initial NL query is generated, and as part of a multi-turn or multi-step session or interaction, one or more subsequent NL queries are generated to obtain the NL result. For example, the generation of one or more subsequent queries is particularly useful when the initial response to the initial NL query does not provide a satisfactory response (whether due to ambiguity in the initial response or for another reason). In such a case, a first subsequent or disambiguation query is generated, taking into account the corresponding structured knowledge (e.g., KG). If the first subsequent query does not provide a satisfactory response, a second subsequent or disambiguation query is generated, taking into account the corresponding structured knowledge. This multi-turn session can continue until disambiguation is satisfactorily resolved. For illustrative purposes, only the first set of subsequent queries and the second set of subsequent queries have been described above. However, it should be understood that additional (e.g., third, fourth, etc.) follow-up queries may be generated as part of a multi-turn session or interaction. Therefore, a content-based GT is generated by the GT manager (152), which takes the form of the generated question (which is based on symptoms, question variants, and knowledge graph traversal to obtain relevant entries for the symptoms).

[0036] In addition to content-based GTs, the GT manager (152) generates usage-based GTs and regulatory-based GTs. Usage-based GTs are referred to herein as GTs. 0,1 (174 0,1 ), GT 1,1 (174 1,1 ) and GT 2,1 (174 2,1 The GT used includes log and feedback data. As illustrated by the example, the GT used (GT) 0,1 (174 0,1 )) is shown as having a log 0,1 (176 0,1 ) and feedback 0,1 (178 0,1 ), based on the GT (GT 1,1 (174 0,1 )) is shown as having a log 1,1 (176 1,1 ) and feedback 1,1 (178 1,1 ), and based on the use of GT (GT 2,1 (1742,1 )) is shown as having a log 2,1 (176 2,1 ) and feedback 2,1 (178 2,1 Similarly, regulatory-based GT (GT 0,2 (174 0,2 The data is populated with regulatory data, referred to here as c_data. 0,2 (178 0,2 ), based on regulation of GT (GT 1,2 (174 1,2 )) is shown as having c_ data 1,2 (178 1,2 ), and regulatory-based GT (GT 2,2 (174 2,2 )) is shown as having c_ data 2,2 (178 2,2 Accordingly, the GT manager (152) generates GTs for multiple categories, each of which is associated with a corresponding knowledge domain and stored in the knowledge base (170).

[0037] The dialogue system (160) is an interactive AI interface configured to support communication between virtual agents and non-virtual agents (such as users (e.g., end users), which can be human or software, and potentially AI virtual agents). The interactions that occur generate so-called conversational or dialogue interactions, which contain the content of such conversational or dialogue interactions between the user and the virtual agent.

[0038] The AI ​​platform (150) is shown herein as an operatively coupled virtual dialogue agent (162) to a dialogue system (160), the dialogue system (160) being configured to receive input (102) from different sources across a network (105). For example, the dialogue system (160) may receive input across the network (105) and utilize one of the knowledge domains and the corresponding ground truth (GT) to create output or response content. The created output or response content may be returned as output (104) across the computer network (105) to the same source and / or one or more other sources.

[0039] Different computing devices (180), (182), (184), (186), (188), and (190) communicating with the network (105) may include access points to the dialogue system (160). In various embodiments, the network (105) may include local network connectivity and remote connectivity, enabling the AI ​​platform (150) to operate in environments of any size, including local and global, such as the Internet. Furthermore, the AI ​​platform (150) acts as a backend system that makes various kinds of knowledge extracted or represented from documents, web-accessible sources, and / or structured data sources available. In this way, some processing populates the AI ​​platform (150), which also includes input interfaces for receiving requests and responding accordingly.

[0040] As shown in the figure, users can access the AI ​​platform (150) and the operationally coupled dialogue system (160) via a network connection or Internet connection to the network (105), and can submit natural language (NL) input to the dialogue system (160). The AI ​​platform (150) can effectively determine the output response related to the input from the natural language (NL) input by utilizing the operationally coupled knowledge base (170) and tools including the AI ​​platform (150).

[0041] The simulator (154) interfaces with the dialogue system (160) to simulate one or more NL dialogue interactions using the dialogue system's (160) automated virtual dialogue agent (162). In one exemplary embodiment, the simulator (154) utilizes an operationally coupled simulator application (154... A ) to simulate interaction with the chatbot (162). Figure 6 The details of the simulation are shown and described herein. The simulation defines a set of test queries with corresponding answers existing in a corresponding knowledge domain, which is represented as a knowledge graph in one exemplary embodiment. The output from the simulation, referred to herein as simulation data, includes a log of all queries and their corresponding responses, which in one exemplary embodiment includes solutions or one or more disambiguation options. Second Library (Library) B (172 B The data is populated in the knowledge base (170) and further populated with simulated data (referred to as s_data in this paper). As illustrated by example in this paper, s_data0 (1540) represents the data used to utilize the knowledge domain (172). A,0 The simulated data of the chatbot (162) is s_data1 (1541), which represents the data used to utilize the knowledge domain (172). A,1 The simulated data of the chatbot (162) and s_data2 (1542) represent the data used to utilize the knowledge domain (172). A,2The simulation data is used for simulating the chatbot (162). Although only one set of simulation data is shown as being associated with each knowledge domain, it should be understood that any knowledge domain can be used for simulating the chatbot (162), where each simulation generates separate or additional simulation data. Similarly, although each knowledge domain is shown as having simulation data, it should be understood that in one exemplary embodiment, not all knowledge domains have been used for simulating the chatbot (162) and therefore will not have corresponding simulation data. Therefore, for each knowledge domain utilized by the interactive simulator (154), an output in the form of simulation data is generated and the output is associated with the corresponding knowledge domain.

[0042] As illustrated herein, an evaluation manager (156), operationally coupled to the simulator (154), is configured to evaluate the performance of the automated virtual dialogue agent (162). The evaluation manager (156) compares simulated interactions, represented as simulated data, with ground truth (GT) used for the corresponding knowledge domain. The GT used in the comparison can include one or more GT types, including content-based, usage-based, and regulatory-based GTs. Figure 7 The details of the simulated interactive evaluation are shown and described. The output from the evaluation manager is multidimensional, including the number of disambiguation questions posed and the differences with respect to the test data, whether the questions were asked in a specific order, and whether the presented solution matches the expected solution. As illustrated in the example, an output of 0 (1560) indicates the use of the knowledge domain (172). A,0 The multidimensional output of the evaluation of the simulation data (1540) is output 1 (1561), which represents the use of the knowledge domain (172). A,1 The multidimensional output of the evaluation of the simulation data (1541), and output 2 (1562) representing the use of the knowledge domain (172) A,2 The evaluation manager (156) evaluates the chatbot (162) based on the simulated data (1542) and records the evaluation as corresponding output data.

[0043] The output data includes insights and recommendations based on the various metrics collected. Business objectives can be predefined based on metrics and corresponding metric measurements (such as the expected accuracy of a chatbot) and an acceptable range of error. Examples of such metrics include, but are not limited to, accuracy, interaction overhead, interaction length, quality of follow-up questions, and response time. In one exemplary embodiment, metrics can be prioritized, for example, prioritizing accuracy over response time. Recommendations corresponding to the collected metrics are automatic comparisons of the defined or predefined business objectives with actual metrics reflecting the performance and identification of one or more corresponding remedial actions. As shown herein, a remedial manager (158), operatively coupled to an evaluation manager (156), is used to identify one or more remedial actions to be applied to the dialogue system (160) based on the corresponding output. For example, in one embodiment, one or more remedial actions can be identified when the performance evaluation of a virtual dialogue agent (162) fails to meet a performance threshold. In one exemplary embodiment, the recommended actions(s), also referred to herein as a recommendation plan, are designed to improve interaction overhead, which can be achieved by collecting additional real-time data and reducing interaction length. In one embodiment, other recommendations can be implemented, and therefore the examples provided herein should not be considered limiting. Accordingly, the remediation manager (158) is configured to perform one or more remediation actions to improve the performance of the automatic virtual dialogue agent (162).

[0044] Dialogue events created or enabled by the dialogue system (160) can be handled by IBM. The server (110) and the corresponding AI platform (150) handle the processing. The GT manager (152) generates GTs from the user's knowledge base and facilitates and enables the evaluation of the dialogue system (160) supported by the generated GTs. In some illustrative embodiments, the server (110) may be an IBM product available from International Business Machines Corporation in Armonk, New York. The system is enhanced using the mechanisms of the illustrative embodiments described below.

[0045] The GT Manager (152), simulator (154), evaluation manager (156), and remediation manager (158) (collectively referred to below as AI tools) are shown as being contained in or integrated within the AI ​​platform (150) of the server (110). The AI ​​tools may be implemented in a separate computing system (e.g., 190) connected to the server (110) across a network (105). Regardless of where they are implemented, the AI ​​tools are used to evaluate dialogue events, extract behavioral features from requests and responses, and selectively identify and apply one or more corresponding remedial actions to improve the performance of the dialogue system (160).

[0046] The range of information processing systems that can utilize the artificial intelligence platform (150) is from small handheld devices such as handheld computers / mobile phones (180) to mainframe systems such as mainframe computers (182). Examples of handheld computers (180) include personal digital assistants (PDAs), personal entertainment devices such as MP4 players, portable televisions, and CD players. Other examples of information processing systems include pen or tablet computers (184), laptop or notebook computers (186), personal computer systems (188), and servers (190). As shown, different information processing systems can be networked together using a computer network (105). Types of computer networks (105) that can be used to interconnect different information processing systems include local area networks (LANs), wireless local area networks (WLANs), the Internet, the public switched telephone network (PSTN), other wireless networks, and any other network topologies that can be used to interconnect information processing systems. Many information processing systems include non-volatile data storage, such as hard disk drives and / or non-volatile memory. Some information processing systems may use a separate non-volatile data storage; for example, a server (190) utilizes non-volatile data storage (190). A ), mainframe computers (182) utilize non-volatile data memory (182 A Non-volatile data memory (182) A It can be a component outside of a different information processing system or it can be inside one of the information processing systems.

[0047] Information processing systems used to support AI platforms (150) can take many forms, some of which are... Figure 1 As shown in the diagram. For example, an information processing system can take the form of a desktop computer, server, portable, laptop, notebook, or other form factor computer or data processing system. Furthermore, an information processing system can take other form factors, such as a personal digital assistant (PDA), gaming device, ATM machine, portable telephone device, communication device, or other device including a processor and memory.

[0048] Application Programming Interface (API) is understood in this art as a software intermediary between two or more applications. About Figure 1 The artificial intelligence platform (150) shown and described herein includes one or more APIs that can be used to support one or more of tools (152), (154), (156), and (158) and their associated functions. References Figure 2A block diagram (200) is provided illustrating the tools (152), (154), (156), and (158) and their associated APIs. As shown, multiple tools are embedded within the AI ​​platform (205), including a GT manager (252) associated with API0 (212), a simulator (254) associated with API1 (222), an evaluation manager (256) associated with API2 (232), and a remediation manager (258) associated with API3 (242). Each API can be implemented using one or more languages ​​and interface specifications. API0 (212) provides functionality to automatically generate GTs from knowledge sources; API1 (222) provides functionality to simulate NL dialogues by utilizing GTs through automated virtual agents; API2 (232) provides functionality to evaluate the performance of automated virtual dialogue agents based on simulations; and API3 (242) provides functionality to selectively identify and implement one or more remedial actions aimed at improving the performance of the dialogue system. As shown in the figure, each of APIs (212), (222), (232), and (242) is operatively coupled to an API orchestrator (260), which is also referred to as an orchestration layer and is understood in the art to act as an abstraction layer to transparently chain the individual APIs together. In one embodiment, the functionality of the individual APIs can be joined or combined. Therefore, the configuration of the APIs shown herein should not be considered limiting. Accordingly, as shown herein, the functionality of these tools can be embodied or supported by their respective APIs.

[0049] refer to Figure 3A flowchart (300) illustrating the process for automatically generating ground truth (GT) from a corresponding knowledge source is provided. As shown and described, the knowledge source may be in a structured form (such as a knowledge graph) or an unstructured form. For descriptive purposes, the GT generation process is described relative to a structured knowledge source, although such a structured format should not be considered limiting. A relevant knowledge source is identified and a set of symptoms is obtained from that source (302). In one exemplary embodiment, one or more selection criteria are used to identify a subset of symptoms. For each symptom, a natural language query is generated (304), which in one exemplary embodiment employs variance generation and adds or removes entities. In one embodiment, variance generation is a natural language equivalent of a phrase and is used here to broaden the scope of the query by identifying comparable or equivalent terms. A knowledge graph (e.g., a structured representation of a knowledge domain) is searched to find query text that matches the symptom (306). In one exemplary embodiment, text matching techniques, such as general sentence encoding, are utilized in step (306). The search from step (306) generates an output (308) in the form of matching symptoms (also referred to herein as matching). Each matching symptom has a corresponding score or weight. In one exemplary embodiment, approximate matching of two or more phrases or sentences is a common operation in natural language processing. A threshold evaluation is performed on the set of matching symptoms from step (308), which, in one exemplary embodiment, evaluates the quality of the matching symptoms. Each matching symptom has constraints. For each matching symptom, a constraint node and the nodes connected to that constraint node are extracted (310). A constraint node is not connected to another constraint node. The extraction at step (310) involves identifying both the constraint node(s) and all other nodes connected to the constraint node(s). For example, the solution for the symptom “battery problem during charging” and the specific solution for this symptom can be limited by a specific hardware model and series. Accordingly, the constraint node(s) are connected to all relevant nodes in the graph as indicators of the constraints.

[0050] VARIABLE X Total The number of constraints assigned (312) and the corresponding constraint count variable X are initialized (314). Based on the constraints (e.g., constraints... XA set of disambiguation questions with answer options is generated (316). In one exemplary embodiment, multiple constraints represent a multi-step session, and disambiguation questions and answer options are generated at each step, and this process is repeated until disambiguation is no longer required. After step (316), the constraint count variable X is incremented (318), and it is determined whether each constraint has been processed (320). A negative response is followed by a return to step (316), and a positive response ends question and answer generation. Thus, one or more disambiguation questions and one or more corresponding answer options for each constraint and each disambiguation selection path are recorded and saved as GT.

[0051] like Figure 3 The methods shown and described automatically generate content-based ground truth (GT) by leveraging knowledge graphs to generate questions and perform graph traversal to identify one or more related entities corresponding to one or more answers. Figure 1 As shown, two other forms of ground truth (GT) are generated, including GTs based on usage and regulation. (Reference) Figure 4 A flowchart (400) is provided to illustrate the process used to generate a usage-based ground truth (GT). A usage log is provided, which records or has been recorded the original query text presented to the user and all subsequent questions, as well as the choices made by the user using the final solution or action plan (402). The usage log includes feedback on whether the query represented in the query text was answered satisfactorily. Variable X Total The number of queries assigned to the log with feedback indicating at least a satisfactory resolution (404). Initialize the corresponding query count variable X (406). For queries... X The query text (408) was obtained from the usage log, and the query was also obtained from the usage log. XThe subsequent questions (410) and the user selection (412) are obtained. In the interaction between the user and the chatbot, the chatbot may provide the user with subsequent questions and optional options as answers. This interaction (including the selection of answers to the options) is collected or obtained in steps (408)-(412). It is then determined whether the solution to the query has been reached (414). The system knows that the information it is sending to the user is a subsequent question, and also knows what information is the solution to the user's question. When the system has sent a solution instead of the next subsequent question, it is determined that the solution to the query has been reached. A negative response to the determination at step (414) is followed by an increment of the query count variable (416), and a return to step (408). Conversely, a positive response to the determination at step (414) is followed by the user selection being identified as the solution (418). The query text, subsequent questions, and user selection obtained in steps (408)-(412) and (418), respectively, are recorded as a workflow and are referred to herein as the usage-based GT. Accordingly, the evaluation uses logs to identify one or more questions and user choices corresponding to the query text and records one or more questions and user choices corresponding to the query text as usage-based GT.

[0052] refer to Figure 5 A flowchart (500) is provided to illustrate the process for generating a regulatory-based GT. In one exemplary embodiment, the regulatory-based GT is manually generated by one or more subject matter experts (SMEs). As shown, a list of symptoms from a knowledge base is provided to the SME for reference (502). The SME writes one or more text queries (504), and for reference, one or more potential answers to the queries and related entities (e.g., constraints) for each question are provided to the SME (506). The SME selects the preferred answer to the query and optionally selects a preferred sequence of subsequent questions for disambiguation (508). In the case of disambiguation questions, a consistency check is performed to verify that the disambiguation options are consistent with the corresponding knowledge base, and the knowledge base may optionally be updated to make the representation consistent (510). The flow of text queries and corresponding answers, as well as one or more subsequent questions for disambiguation in one exemplary embodiment, are recorded and saved as GT. Accordingly, a record of the regulatory-based GT is provided with the assistance of the SME.

[0053] like Figure 1 As shown and described, a simulator (154) is provided to support the simulation of NL dialogue interactions. Reference Figure 6 A flowchart (600) is provided to illustrate the process used to simulate interaction with the dialogue system (160). The disambiguation selection path count variable N (602) is initialized, GT data is used as a source to drive interaction with the virtual dialogue agent, and queries (e.g., queries) are generated using an operationally coupled simulator application. NThe query is sent to the automated virtual agent (604). The query is logged in the corresponding simulation log (606). The automated virtual agent responds to the query with a solution or a set of disambiguation options (608). If the response is a solution, the solution is logged in the corresponding log (610), and if the response is a set of disambiguation questions, the disambiguation selection path count variable N is incremented (612), and the GT is consulted to find the disambiguation question for the query in the GT (614). In an exemplary embodiment, if no disambiguation question for the query is found in the GT, the process may stop, select "Any" (if offered as an option), or randomly select one of the offered options, depending on the configuration. If a question is selected, the process returns to step (606). After step (610), the number of disambiguation selection paths is assigned to the variable N. total (616). Subsequently, a log of questions and answers obtained or identified from the simulation is recorded in the simulation log (618). Accordingly, the simulator application creates a simulation interaction log that records the inputs received from the corresponding GT and the outputs generated.

[0054] The dialogue system (160) and the corresponding automated virtual agent (162) undergo performance evaluation by utilizing ground truth (GT) and corresponding simulated logs. (See reference) Figure 7 A flowchart (700) is provided to illustrate the process used for performance evaluation of a virtual dialogue system. As shown in the figure, variable N... total The number of query-response pairs (702) is assigned to be logged in the simulation log, and the corresponding counter variable N is initialized (704). For each query-response pair... N Find the corresponding entry in the GT (706). Compare the query-response in the simulation log with the query-response in the GT (708). Generate a multidimensional output based on the comparison at step (708), including: 1. the difference between the number of disambiguation questions actually queried and the number in the GT, 2. whether the questions were queried in the same order, and 3. whether the solutions presented in the simulation log match the solutions in the GT. Insights and recommendations are generated based on the outputs generated for each of the multiple dimensions (710). In an exemplary embodiment, one or more additional dimensions may be added to the evaluation, or conversely, a reduced number of dimensions may be used for the evaluation. Accordingly, as shown herein, the comparison of the simulation log with the GT provides insights into the performance of the dialogue system (160).

[0055] The output data at step (710) includes insights and recommendations based on the various metrics collected. Business objectives can be predefined based on metrics and corresponding metric measurements (such as the expected accuracy of a chatbot) and an acceptable range of error. Examples of such metrics include, but are not limited to, accuracy, interaction overhead, interaction length, quality of follow-up questions, and response time. In one exemplary embodiment, metrics can be prioritized, for example, prioritizing accuracy instead of response time. Recommendations corresponding to the collected metrics are an automatic comparison of the defined or predefined business objectives with actual metrics reflecting the performance and identification of one or more corresponding remedial actions. As shown herein, one or more remedial actions are identified (712) and selectively implemented (714) for application to the dialogue system (160) based on the corresponding output. For example, in one embodiment, one or more remedial actions can be identified when the performance evaluation of the virtual dialogue agent (162) fails to meet a performance threshold. In one exemplary embodiment, the recommendations(s), also referred to herein as recommendation plans, aim to improve interaction overhead, which can be achieved by collecting additional real-time data and reducing interaction length. In one embodiment, other recommendations can be implemented, and therefore the examples provided herein should not be considered limiting. Accordingly, the remedial action involves improving the performance of the dialogue system (160) and the corresponding automatic virtual dialogue agent (162).

[0056] As in Figure 1-7 The computer systems, program products, and methods shown and described herein are provided for evaluating the performance of a multi-turn automated virtual agent using an automatically generated ground truth (GT) from an operationally coupled knowledge source. The automated virtual agent is used to simulate a multi-turn (NL) dialogue by driving the corresponding dialogue using the GT. A log is created to record the simulation. The performance of the automated virtual agent is evaluated by comparing the simulation log with the corresponding GT. Based on the performance evaluation of the simulation and simulation log, one or more remedial actions aimed at improving the performance of the automated virtual agent are identified and selectively implemented.

[0057] The embodiments shown and described herein can take the form of a computer system used in conjunction with an intelligent computing platform to improve the performance of the dialogue system and the corresponding automated virtual agent. Aspects of tools (152), (154), (156), and (158) and their associated functionality can be embodied in a single-location computer system / server, or, in one embodiment, can be configured in a cloud-based system with shared computing resources. References Figure 8 A block diagram (800) is provided illustrating an example of a computer system / server (802), hereinafter referred to as a host (802) communicating with a cloud-based support system (810) to achieve the above. Figures 1 to 7The systems, tools, and processes described herein. In one embodiment, the host (802) is a node in a cloud computing environment. The host (802) may operate with many other general-purpose or special-purpose computing system environments or configurations. Examples of well-known computing systems, environments, and / or configurations suitable for use with the host (802) include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and file systems (e.g., distributed storage environments and distributed cloud computing environments) that include any of the above systems, devices, and their equivalents.

[0058] The host (802) can be described within the general context of computer system executable instructions (such as program modules) executed by the computer system. Generally, program modules may include routines, programs, objects, components, logic, data structures, etc., that perform specific tasks or implement specific abstract data types. The host (802) can operate in a distributed cloud computing environment, where tasks are performed by remote processing devices linked via a communication network. In a distributed cloud computing environment, program modules can reside in local and remote computer system storage media, including memory storage devices.

[0059] like Figure 8 As shown, the host (802) is illustrated as a general-purpose computing device. Components of the host (802) may include, but are not limited to, one or more processors or processing units (804) (e.g., hardware processors), system memory (806), and buses (808) that couple the various system components, including the system memory (806), to the processors (804). The bus (808) represents any one or more of several types of bus architectures, including memory buses or memory controllers, peripheral buses, accelerated graphics ports, and processor or local buses using any of a variety of bus architectures. By way of example and not limitation, such architectures include the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MCA) bus, the Enhanced ISA (EISA) bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnect (PCI) bus. The host (802) typically includes a variety of computer system readable media. Such media can be any available media accessible by the host (802), and includes volatile and non-volatile media, removable and non-removable media.

[0060] System memory (806) may include computer system-readable media in the form of volatile memory, such as random access memory (RAM) (830) and / or cache memory (832). By way of example only, storage system (834) may be provided for reading from and writing to non-removable, non-volatile magnetic media (not shown, and generally referred to as "hard disk drives"). Although not shown, disk drives may be provided for reading from or writing to removable non-volatile disks (e.g., "floppy disks"), and optical disk drives may be provided for reading from or writing to removable non-volatile optical disks (such as CD-ROMs, DVD-ROMs, or other optical media). In such a case, each may be connected to a bus (808) via one or more data media interfaces.

[0061] A program / utility (840) having a set (at least one) of program modules (842), along with (by way of example and not limitation) an operating system, one or more applications, other program modules, and program data, may be stored in system memory (806). Each or some combination of the operating system, one or more applications, other program modules, and program data may include implementations of a networked environment. Program modules (842) generally perform the functions and / or methods of an embodiment to dynamically interpret and understand request and action descriptions and effectively augment corresponding domain knowledge. For example, the set of program modules (842) may include, for example, Figure 1 The tools shown are (152), (154), (156) and (158).

[0062] The host (802) can also communicate with one or more external devices (814) (e.g., keyboard, pointing device, etc.), a display (824), one or more devices that enable a user to interact with the host (802), and / or any device that enables the host (802) to communicate with one or more other computing devices (e.g., network interface card, modem, etc.). Such communication may occur via an input / output (I / O) interface (822). Furthermore, the host (802) can communicate with one or more networks such as a local area network (LAN), a general wide area network (WAN), and / or a public network (e.g., the Internet) via a network adapter (820). As depicted, the network adapter (820) communicates with other components of the host (802) via a bus (808). In one embodiment, multiple nodes of a distributed file system (not shown) communicate with the host (802) via the I / O interface (822) or via the network adapter (820). It should be understood that, although not shown, other hardware and / or software components may be used in conjunction with the host (802). Examples include, but are not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archiving storage systems.

[0063] In this document, the terms “computer program media,” “computer-usable media,” and “computer-readable media” are used to refer in general to media such as main memory (806) (including RAM (830)), cache (832), and storage systems (834) (such as removable storage drives and hard disks installed in hard disk drives).

[0064] A computer program (also known as computer control logic) is stored in memory (806). The computer program can also be received via a communication interface (such as a network adapter (820)). Such a computer program, when run, enables the computer system to perform the features of the embodiments of the invention discussed herein. Specifically, the computer program, when run, enables the processing unit (804) to perform the features of the computer system. Therefore, such a computer program represents the controller of the computer system.

[0065] Computer-readable storage media can be tangible devices capable of retaining and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example, but not limited to, electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes: portable computer disks, hard disks, dynamic or static random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), magnetic storage devices, portable compact disc read-only memory (CD-ROM), digital universal disc (DVD), memory sticks, floppy disks, mechanical encoding devices such as punch cards or recessed structures with instructions recorded thereon, and any suitable combination of the foregoing. As used herein, computer-readable storage media should not be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0066] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a suitable computing / processing device, or downloaded via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network) to an external computer or external storage device. The network may include copper cables, optical fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to a computer-readable storage medium within the suitable computing / processing device.

[0067] Computer-readable program instructions used to perform operations according to embodiments of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​(such as Java, Smalltalk, C++, etc.) and conventional procedural programming languages ​​(such as the "C" programming language or similar programming languages). The computer-readable program instructions may be executed entirely on a user's computer, partially on a user's computer, as a standalone software package, partially on a user's computer and partially on a remote computer, or entirely on a remote computer, server, or server cluster. In the latter case, the remote computer may be connected to the user's computer via any type of network (including a local area network (LAN) or wide area network (WAN)) or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs) may execute computer-readable program instructions by utilizing state information from the computer-readable program instructions to personalize the electronic circuitry in order to perform aspects of embodiments of the present invention.

[0068] The functional tools described in this specification are labeled as managers. Managers can be implemented in programmable hardware devices such as field-programmable gate arrays, programmable array logic, programmable logic devices, etc. Managers can also be implemented in software for processing by different types of processors. The manager of the identified executable code can, for example, comprise one or more physical or logical blocks of computer instructions, which can be organized, for example, into objects, procedures, functions, or other constructs. However, the executable files of the identified managers do not need to be physically located together, but can include different instructions stored in different locations that, when logically combined, include the manager and achieve the manager's stated purpose.

[0069] In practice, the manager of executable code can be a single instruction or many instructions, and can even be distributed across several different code segments, different applications, and multiple storage devices. Similarly, operational data can be identified and displayed within the manager, and can be represented in any suitable form and organized within any suitable type of data structure. Operational data can be collected as a single dataset, or it can be distributed across different locations (including different storage devices), and can exist at least partially as electronic signals on the system or network.

[0070] See now Figure 9An illustrative cloud computing network (900) is shown. As illustrated, the cloud computing network (900) includes a cloud computing environment (950) with one or more cloud computing nodes (910), whose local computing devices used by cloud consumers can communicate with the cloud computing nodes. Examples of these local computing devices include, but are not limited to, personal digital assistants (PDAs) or cellular phones (954A), desktop computers (954B), laptop computers (954C), and / or automotive computer systems (954N). Individual nodes within the nodes (910) can further communicate with each other. They can be physically or virtually grouped (not shown) in one or more networks, such as private clouds, community clouds, public clouds, or hybrid clouds, or combinations thereof, as described above. This allows the cloud computing environment (900) to provide Infrastructure as a Service, Platform as a Service, and / or Software as a Service, without requiring cloud consumers to maintain resources on their local computing devices. It should be understood that... Figure 9 The types of computing devices (954A-N) shown are intended to be illustrative only, and the cloud computing environment (950) can communicate with any type of computerized device via any type of network and / or network-addressable connection (e.g., using a web browser).

[0071] See now Figure 10 This shows the result of Figure 9 The cloud computing network provides a set of functional abstraction layers (1000). It should be understood in advance that... Figure 10 The components, layers, and functions shown are intended to be illustrative only, and the embodiments are not limited thereto. As depicted, the following layers and corresponding functions are provided: hardware and software layer (1010), virtualization layer (1020), management layer (1030), and workload layer (1040).

[0072] The hardware and software layer (1010) includes hardware and software components. Examples of hardware components include mainframes, which in one example are... System; a server based on a RISC (Reduced Instruction Set Computer) architecture, in one example being an IBM... System; IBM System; IBM Systems; storage devices; networking and interconnection components. Examples of software components include network application server software, one example being IBM. Application server software; and database software, in one example, IBM. Database software. (IBM, zSeries, pSeries, xSeries, BladeCenter, WebSphere, and DB2 are trademarks of IBM registered in many jurisdictions worldwide.)

[0073] The virtualization layer (1020) provides an abstraction layer from which the following examples of virtual entities can be provided: virtual servers; virtual storage devices; virtual networks, including virtual private networks; virtual applications and operating systems; and virtual clients.

[0074] In one example, the management layer (1030) can provide the following functions: resource provisioning, metering and pricing, user portal, service layer management, and SLA planning and enforcement. Resource provisioning provides dynamic procurement of computing resources and other resources used to perform tasks within the cloud computing environment. Metering and pricing provides cost tracking as resources are utilized within the cloud computing environment and bills or invoices for the consumption of these resources. In one example, these resources may include application software licenses. Security provides authentication for cloud consumers and tasks, as well as protection for data and other resources. The user portal provides consumers and system administrators with access to the cloud computing environment. Service layer management provides the allocation and management of cloud computing resources to meet required service layer requirements. Service layer agreement (SLA) planning and enforcement provides the pre-scheduling and procurement of cloud computing resources, anticipating future requirements for cloud computing resources based on the SLA.

[0075] The workload layer (1040) provides examples of functionalities that can be leveraged in a cloud computing environment. Examples of workloads and functionalities that can be provided from this layer include, but are not limited to: mapping and navigation; software development and lifecycle management; virtual classroom education delivery; data analytics and processing; transaction processing; and evaluation and enrichment of virtual dialogue systems.

[0076] While specific embodiments of the invention have been shown and described, it will be apparent to those skilled in the art that changes and modifications can be made based on the teachings herein without departing from the embodiments and their broader aspects. Therefore, the appended claims include, within their scope, all such changes and modifications within the true spirit and scope of the embodiments. Furthermore, it should be understood that the embodiments are defined solely by the appended claims. Those skilled in the art will understand that if a specific number of introduced claim elements were intended, such an intention would be explicitly stated in the claims, and without such a statement, there is no such limitation. As a non-limiting example, to aid understanding, the following appended claims contain the use of the introductory phrases “at least one” and “one or more” to introduce claim elements. However, the use of such phrases should not be construed as implying that introducing a claim element by the indefinite article “a(a)” or “an” limits any particular claim containing such an introduced claim element to an embodiment containing only one such element, even when the same claim includes the introductory phrases “one or more” or “at least one” and indefinite articles such as “a(a)” or “an”; the same applies to the use of definite articles in the claims. As used herein, the term “and / or” means either one or both (or one or any combination or all of the terms, or the expressions they refer to).

[0077] Embodiments of the present invention may be systems, methods, and / or computer program products. Furthermore, selected aspects of the embodiments of the present invention may take the form of entirely hardware embodiments, entirely software embodiments (including firmware, resident software, microcode, etc.), or embodiments combining software and / or hardware aspects, which may be collectively referred to herein as “circuit,” “module,” or “system.” Additionally, aspects of the embodiments of the present invention may take the form of computer program products embodied in a computer-readable storage medium having computer-readable program instructions thereon for causing a processor to execute aspects of the embodiments of the present invention. Therefore, the disclosed systems, methods, and / or computer program products are operable to support the evaluation and improvement of virtual dialogue systems.

[0078] This document describes aspects of embodiments of the invention with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0079] These computer-readable program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / actions specified in one or more blocks of a flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, programmable data processing apparatus, and / or other devices to operate in a particular manner, such that the computer-readable storage medium in which the instructions are stored includes an article of writing comprising instructions for implementing aspects of the functions / actions specified in one or more blocks of a flowchart and / or block diagram.

[0080] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer-implemented process, such that the instructions, which execute on the computer, other programmable apparatus or other device, perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0081] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of instructions comprising one or more executable instructions for implementing a specified logical function. In some alternative embodiments, the functions indicated in the blocks may occur in a non-consecutive order as shown in the figures. For example, two blocks shown consecutively may actually be executed substantially simultaneously, or these blocks may sometimes be executed in reverse order, depending on the functions involved. It will also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified function or action or executes a combination of dedicated hardware and computer instructions.

[0082] It should be understood that although specific embodiments have been described herein for illustrative purposes, various modifications may be made without departing from the spirit and scope of the embodiments. Therefore, the scope of protection of the embodiments is defined only by the appended claims and their equivalents.

Claims

1. A computer system, comprising: A processor that is operationally coupled to memory; and an artificial intelligence (AI) platform, which is operationally coupled to the processor, the AI ​​platform including one or more tools for interfacing with a virtual dialogue agent, the tools further including: Ground-based GT Manager, which is configured to automatically generate GT from knowledge sources; A simulator configured to use the virtual dialogue agent to simulate NL dialogue interactions, the simulator being configured to use the GT to drive the output generated by the simulated NL dialogue and create corresponding simulation logs; The evaluation manager is configured to evaluate the performance of the virtual dialogue agent with respect to the created simulation logs, taking into account the ground truth (GT); and The Rescue Manager is configured as follows: In response to the assessed performance failing to meet a performance threshold, one or more remedial actions are identified for the virtual dialogue agent; and Selectively implement one or more of the identified remedial actions.

2. The computer system of claim 1, wherein, The GT data includes usage logs and corresponding feedback, structured data, records generated by subject matter experts, or any combination thereof.

3. The computer system according to claim 1, wherein, The evaluation manager is configured to compare query-response pairs in the GT with corresponding query-response pairs in the simulation log.

4. The computer system according to claim 1, wherein, The GT Manager is further configured to edit the first disambiguation selection path, the editing including: Generate an NL query and at least one disambiguation NL query; In response to the at least one disambiguation NL query, generate NL results; and Record the first log used for the first disambiguation selection path.

5. The computer system according to claim 4, wherein, The simulator is further configured to edit a second disambiguation selection path, the editing including: Generate a test NL query and at least one test disambiguation NL query; Generate a test NL response to the at least one test disambiguation NL query; and Record the second log of the second disambiguation selection path.

6. The computer system according to claim 5, wherein, The evaluation manager is further configured to compare the recorded first log with the recorded second log.

7. A computer program product for improving the performance of a virtual dialogue agent, the computer program product comprising: Computer-readable storage medium; as well as Program code, which is stored on the computer-readable storage medium and is executable by a computer processor to perform the following operations: Automatically generate ground-based ground-based ground (GT) data from knowledge sources; Using the virtual dialogue agent to simulate NL dialogue interaction includes: using the GT to drive the output generated by the simulated NL dialogue, and creating a corresponding simulation log; Evaluate the performance of the virtual dialogue agent with respect to the created simulated logs, taking into account the GT; and In response to the assessed performance failing to meet a performance threshold, identify one or more remedial actions for the dialogue system; and One or more remedial actions may be selectively performed.

8. The computer program product according to claim 7, wherein, The GT data includes usage logs and corresponding feedback, knowledge graphs, records generated by subject matter experts, or any combination thereof.

9. The computer program product according to claim 7, wherein, The program code that can be executed by the computer processor to evaluate performance includes: computer code that can be executed by the computer processor to compare query-response pairs in the GT with corresponding query-response pairs in the simulation log.

10. The computer program product according to claim 7, wherein, The program code executable by the computer processor to utilize the GT data includes: program code executable by the computer processor to edit the first disambiguation selection path, the editing including: Generate an NL query and at least one disambiguation NL query; In response to the at least one disambiguation NL query, generate NL results; and Record the first log used for the first disambiguation selection path.

11. The computer program product according to claim 10, wherein, The program code executable by the computer processor to perform the simulation further includes: program code executable by the computer processor to edit the second disambiguation selection path, the editing including: Generate a test NL query and at least one test disambiguation NL query; Generate a test NL response to the at least one test disambiguation NL query; and Record the second log of the second disambiguation selection path.

12. The computer program product according to claim 11, wherein, The program code, which can be executed by the computer processor to evaluate the performance of the automated virtual dialogue agent, further includes: program code, which can be executed by the computer processor to compare the recorded first log with the recorded second log.

13. A computer-implemented method relating to improving the performance of a virtual dialogue agent system, the method comprising: Ground-based ground tracking (GT) is automatically generated from knowledge sources by a computer processor. The computer processor uses the virtual dialogue agent to simulate NL dialogue interaction, including: using the GT to drive the output generated by the simulated NL dialogue interaction, and creating a corresponding simulation log; The computer processor evaluates the performance of the virtual dialogue agent with respect to the created simulated logs, taking into account the ground truth (GT). In response to the assessed performance failing to meet a performance threshold, the computer processor identifies one or more remedial actions for the dialogue system; and The computer processor selectively performs one or more of the identified remedial actions.

14. The computer-implemented method according to claim 13, wherein, The GT data includes usage logs and corresponding feedback, structured data, records generated by subject matter experts, or any combination thereof.

15. The computer-implemented method according to claim 13, wherein, The evaluation includes comparing the query-response pairs in the GT with the corresponding query-response pairs in the simulation log.

16. The computer-implemented method according to claim 13, wherein, Using the GT data includes editing a first disambiguation selection path by the computer processor, the editing including: Generate an NL query and at least one disambiguation NL query; In response to the at least one disambiguation NL query, generate NL results; and Record the first log used for the first disambiguation selection path.

17. The computer-implemented method according to claim 16, wherein, The simulation further includes editing a second disambiguation selection path by the computer processor, the editing including: Generate a test NL query and at least one test disambiguation NL query; Generate a test NL response to the at least one test disambiguation NL query; and Record a second log for the second disambiguation path selection.

18. The computer-implemented method according to claim 17, wherein, Evaluating the performance of the automated virtual dialogue agent further includes comparing the recorded first log with the recorded second log.

Citation Information

Patent Citations

  • View synthesis using neural networks

    CN111698463A

  • Machine natural language processing for summarization and sentiment analysis

    US20190349321A1