Performance evaluation and enhancement of virtual dialogue systems

By introducing tools such as GT Manager, simulator, and evaluation manager into the automated dialogue system, benchmark data is automatically generated and dialogue interactions are simulated, which solves the problem of insufficient question-and-answer quality of chatbots and achieves optimization and performance improvement of the dialogue system.

JP7805074B2Active Publication Date: 2026-01-23INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2023544551
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-01-29
Filing Date
2022-01-04
Publication Date
2026-01-23
Estimated Expiration
2042-01-04

AI Technical Summary

Technical Problem

The chatbots in existing automated dialogue systems fail to meet quality standards in natural language dialogue events, mainly due to insufficient question-and-answer quality and the inability to effectively convert questions into equivalent knowledge representations, resulting in inappropriate responses or low efficiency.

Method used

A system, including a processor, memory, and an artificial intelligence platform, is employed to automatically generate benchmark data, simulate conversational interactions, evaluate the chatbot's performance, and implement remedial measures based on performance thresholds, through a GT manager, simulator, evaluation manager, and repair manager.

Benefits of technology

This improves the quality of natural language dialogue in chatbots, ensuring their evaluation and optimization before commercial deployment to meet accuracy and interaction quality requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007805074000001
    Figure 0007805074000001
  • Figure 0007805074000002
    Figure 0007805074000002
  • Figure 0007805074000003
    Figure 0007805074000003
Patent Text Reader

Abstract

Embodiments are provided for a computer system, computer program product, and computer-implemented method for improving the performance of a virtual dialogue agent system using an automated virtual dialogue agent. The embodiments include generating a ground truth (GT) from a knowledge base and utilizing the GT to evaluate how the virtual dialogue agent performs with the GT. The evaluation measures the quality of multi-turn virtual dialogues and generates a remediation plan directed to algorithmically improving the virtual dialogue agent.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] One or more of the present embodiments relate to a virtual dialogue system that uses automated virtual dialogue agents, such as "chatbots," and associated computer program products and computer-implemented methods. In certain exemplary embodiments, quality metrics of corresponding automated virtual dialogue agent interactions are evaluated and selectively resolved, where the resolution is directed, for example, to selectively applying one or more remediation actions to the automated virtual dialogue agent to improve performance with respect to natural language (NL) dialogue events. [Background technology]

[0002] Automated virtual conversational agents use artificial intelligence (AI) as a platform to conduct NL interactions between the automated virtual conversational agent and a user, typically a consumer or client, or another conversational agent. The interactions may involve product sales, customer service, information acquisition, or other types of interactions or transactions. Chatbots often interact with users through either textual (e.g., online or text) or auditory (e.g., telephone) dialogue. It is known in the art that chatbots function as a question-answering component between a user and an AI platform. The quality of questions (i.e., queries) and answers (i.e., responses) is derived from the quality of question understanding, question transformation, and answer resolution. Failure to achieve quality standards is commonly due to either inappropriate question formulation or inefficient question formulation to solicit a corresponding response. This may be due to a lack of knowledge to effectively convert questions into equivalent knowledge representations that map to answers, or due to inefficiencies within the AI ​​platform or chatbot. For example, a lack of synonyms or conceptual relationships may limit the ability of an AI platform to determine that a question entered by a customer or client is equivalent to or related to a known question for which an answer is available in a dataset or database.

[0003] Companies may impose certain requirements, such as accuracy or interaction quality in virtual assistance, that are expected to be met before commercial deployment of the virtual system. For example, the virtual system may have minimum performance requirements, such as 50 percent accuracy for a user base of support agents or 90 percent accuracy for a user base of end users. Therefore, it is desirable for dialogue systems to be subjected to benchmarking or quality testing before deployment. Summary of the Invention [Means for solving the problem]

[0004] The embodiments include systems, computer program products, and methods for improving the performance of dialogue systems. This Summary is provided to introduce in a simplified form a selection of representative concepts that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used in any way to limit the scope of the claimed subject matter.

[0005] In one aspect, a computer system is provided that includes a processor operatively connected to a memory and an artificial intelligence (AI) platform operatively connected to the processor. The AI ​​platform includes one or more tools for improving performance of a virtual dialogue agent. The tools include a ground truth (GT) manager, a simulator, an evaluation manager, and a remediation manager. The GT manager is configured to automatically generate a GT from a knowledge source. The simulator is configured to simulate a NL dialogue interaction with the virtual dialogue agent. More particularly, the simulator is configured to utilize the GT to drive simulated NL dialogue-generated outputs and create a corresponding simulation log. The evaluation manager is configured to evaluate performance of the virtual dialogue agent with respect to the simulation log in terms of the GT. The remediation manager is configured to identify and selectively implement one or more remediation actions for the dialogue system in terms of a performance threshold.

[0006] In another aspect, a computer program product is provided comprising one or more computer-readable storage media and program code stored on the one or more computer-readable storage media. The program code is executable by a computer processor to improve performance of a virtual dialogue agent. The program code is provided for automatically generating a ground truth (GT) from a knowledge source. The program code is further provided for simulating a NL dialogue interaction using the virtual dialogue agent. The simulation utilizes the GT to drive simulated NL dialogue-generated outputs and creates a corresponding simulation log. Program code is provided for evaluating performance of the virtual dialogue agent with respect to the simulation log in terms of the GT, and identifying and selectively executing one or more remediation actions for the dialogue system in terms of a performance threshold.

[0007] In yet another aspect, a computer-implemented method for improving performance of a virtual dialogue agent is provided. The method is configured to automatically generate a ground truth (GT) from a knowledge source. NL dialogue interactions are simulated using the virtual dialogue agent. The simulation utilizes the GT to drive simulated NL dialogue-generated outputs and creates a corresponding simulation log. Performance of the virtual dialogue agent is evaluated in terms of the GT with respect to the created simulation log. One or more remediation actions for the dialogue system are identified and selectively implemented in terms of a performance threshold.

[0008] These and other features and advantages will become apparent from the following detailed description of one or more presently exemplary embodiments, considered in conjunction with the accompanying drawings.

[0009] The drawings referenced herein form a part of this specification and are incorporated herein by reference. Features shown in the drawings are intended to illustrate only some embodiments, and not all embodiments, unless otherwise expressly stated. [Brief explanation of the drawings]

[0010] [Figure 1] FIG. 1 illustrates a system diagram illustrating an artificial intelligence platform computing system in a network environment. [Figure 2] FIG. 2 is a block diagram illustrating the artificial intelligence platform tools and their associated application program interfaces shown and described in FIG. [Figure 3] FIG. 3 illustrates a flow chart diagram showing one embodiment of a method for automatically generating ground truth from corresponding knowledge sources. [Figure 4] FIG. 4 illustrates a flow chart diagram showing one embodiment of a method for generating a usage-based GT. [Figure 5] FIG. 5 illustrates a flow chart diagram showing one embodiment of a method for generating a curation-based GT. [Figure 6] FIG. 6 illustrates a flow chart diagram showing one embodiment of a method for simulating an interaction with a dialogue system. [Figure 7] FIG. 7 illustrates a flow chart diagram showing one embodiment of a method for performance evaluation of a virtual dialogue system. [Figure 8] FIG. 8 illustrates a block diagram of an example computer system / server of a cloud-based support system for implementing the systems and processes described above with respect to FIGS. [Figure 9] FIG. 9 illustrates a block diagram illustrating a cloud computing environment. [Figure 10]FIG. 10 depicts a block diagram illustrating one set of functional abstraction model layers provided by a cloud computing environment. DETAILED DESCRIPTION OF THE INVENTION

[0011] It will be readily understood that the components of the present embodiments, as generally described and illustrated in the figures accompanying this specification, could be arranged and designed in a wide variety of different configurations. Thus, the following detailed description of exemplary embodiments of devices, systems, methods and computer program products, as set forth in the figures, is not intended to limit the scope of the embodiments claimed in the appended claims, but is merely representative of selected embodiments.

[0012] References throughout this specification to "selected embodiments," "one embodiment," "exemplary embodiment," or "an embodiment" mean that a particular feature, structure, or characteristic described in connection with an embodiment is included in at least one embodiment. Thus, the appearances of the phrases "selected embodiments," "one embodiment," "exemplary embodiment," or "in an embodiment" in various places throughout this specification do not necessarily refer to the same embodiment. The embodiments described herein can be combined with each other and modified to include each other's features. Moreover, the described features, structures, or characteristics of various embodiments can be combined and modified in any suitable manner.

[0013] The illustrated embodiments will be best understood by referring to the drawings, where like parts are designated by like numerals throughout. The following description is intended as an example only and merely illustrates certain selected embodiments of devices, systems, products and processes consistent with embodiments claimed herein.

[0014] In the field of artificial intelligence computer systems, natural language systems (e.g., IBM (登録商標) Watson (登録商標) Artificial intelligence computer systems (and other natural language systems) process natural language based on knowledge acquired by the system. To process natural language, the system may be trained with data obtained from a database or corpus of knowledge, but the resulting results may be inaccurate or incorrect for various reasons.

[0015] Machine learning (ML) is a subset of artificial intelligence (AI) that uses algorithms to learn from data and make predictions based on this data. AI refers to the intelligence of machines to make informed decisions and maximize the chances of success in a given subject. More specifically, AI can learn from data sets to solve problems and provide relevant recommendations. Cognitive computing is a combination of computer science and cognitive science. Cognitive computing utilizes self-learning algorithms that use data minimization, visual recognition, and natural language processing to solve problems and optimize human processes.

[0016] Similar concepts are at the core of AI and associated reasoning. The process of understanding natural language and objects requires reasoning in terms of relationships, which can be challenging. A structure, comprising both static and dynamic structures, dictates a deterministic output or action given a deterministic input. More specifically, the determined output or action is based on explicit or inherent relationships within the structure. This arrangement may be satisfactory for selected situations and conditions. However, it is understood that dynamic structures are inherently variable, and outputs or actions may change accordingly. Existing solutions for efficiently identifying objects and understanding natural language, and for processing content responses to changes to the structure as well as the identification and understanding, are extremely challenging at a practical level.

[0017] An automated virtual agent, referred to herein as a chatbot, is an artificial intelligence (AI) program that simulates interactive human conversation using pre-computed phrases and auditory or text-based signals. Chatbots are increasingly being used in electronic platforms for customer service support. In one embodiment, the chatbot can function as an intelligent virtual agent. Each chatbot experience consists of a set of communications made up of user actions and dialogue system actions, where the experience has an identifiable behavioral pattern. It is understood in the art that chatbot interactions can be evaluated and subjected to diagnostics to identify chatbot elements that may warrant modification to remediate future chatbot experiences.

[0018] The systems, computer program products, and methods evaluate automated virtual dialogue agents, and in exemplary embodiments, multi-turn dialogue systems, by automatically generating benchmark data, also referred to herein as ground truth (GT), from a knowledge base and evaluating the automated virtual dialogue agent. In exemplary embodiments, the GT is generated automatically from a user's knowledge base and not from a standard or generic dataset. The generation of the benchmark data serves as a platform for extracting a GT within the knowledge base. As supported by the GT, simulated dialogue interactions are conducted with the automated virtual dialogue agent. As shown and described herein, the automated virtual dialogue agent is subjected to a performance evaluation, including a comparison of corresponding simulation logs, from the perspective of the GT. Metrics for measuring the evaluation of the automated virtual dialogue agent include, for example, a determination of the correctness of the automated virtual dialogue agent's responses, the relevance of one or more disambiguation or follow-up questions asked by the automated virtual dialogue agent, the number of disambiguation or follow-up questions asked by the automated virtual dialogue agent, or the order of disambiguation or follow-up questions asked by the automated virtual dialogue agent, or a combination thereof.

[0019] The automated virtual conversational agent (also referred to herein as a "chatbot platform" or "chatbot") functions as an AI interaction interface. As shown and described herein, the chatbot platform is evaluated based on a comparison of simulated interactions with a GT. Ground truth (also referred to herein as "GT") is a term used in machine learning to refer to information provided by direct observation, e.g., rules of thumb, as opposed to information provided by inference. As explained in more detail below, GT data includes, for example, content-based information, such as a knowledge graph or knowledge base, usage logs (especially feedback information related to those logs), subject matter expert (SME) records, or any combination thereof.

[0020] Referring to Figure 1, a schematic diagram of an artificial intelligence (AI) platform and corresponding system (100) is shown. As shown in Figure 1, a server (110) is provided in communication with multiple computing devices (180), (182), (184), (186), (188), and (190) via a network connection, e.g., a computer network (105). The server (110) is configured with a processing unit, e.g., a processor, in communication with a memory via a bus. The server (110) is shown to include an AI platform (150) operatively connected to a dialogue system (160), a corresponding virtual agent (162), e.g., a chatbot, and an associated knowledge base (170), e.g., a data source. The computing devices 180, 182, 184, 186, 188, and 190 may include a visual display, an audio interface, an audio-video interface, or other type of interface configured to allow a user to interface with a representation of a virtual agent 162, e.g., a chatbot.

[0021] The AI ​​platform 150 is operatively connected to a computer network 105 to support interactions with the virtual conversational agent 162 from one or more of the computing devices 180, 182, 184, 186, 188, and 190. More particularly, the computing devices 180, 182, 184, 186, 188, and 190 communicate with each other and with other devices or components via one or more wired or wireless data communication links, or a combination thereof, where each communication link may comprise one or more wires, routers, switches, transmitters, receivers, etc. In this networked arrangement, the server 110 and network connections 105 enable communication detection, recognition, and resolution. Other implementations of the server 110 may be used with components, systems, subsystems, or devices, or combinations thereof, other than those illustrated herein.

[0022] The AI ​​platform (150) is also shown herein to be operatively connected to a knowledge base (170), also referred to herein as a corpus of information. As shown, the knowledge base (170) is illustratively referred to herein as a library A (172 A ) and libraries B (172 B ) in the knowledge base 170. Although two libraries are shown in FIG. 1, it should be understood that the knowledge base 170 may include fewer or more libraries. Furthermore, the libraries, e.g., A (172 A ) and libraries B (172 B ) may be connected together. A (172 A ) and librariesB (172 B ), may span multiple knowledge domains, e.g., knowledge domains comprising knowledge bases (170) and other knowledge domains (not shown). Each library captures data in either structured or unstructured form. For example, in one exemplary embodiment, the structured data may be in the form of a knowledge graph. For example, a library A (172 A ) is implemented as a structured knowledge domain represented as a knowledge graph (KG), and is shown in Figure 1 as KG0 (172 A,0 ), KG1(172 A,1 ), and KG2(172 A,2 )

[0023] The AI ​​platform (150) is presented herein with several tools to assist in evaluating, benchmarking, and improving the performance of dialogue systems (160) and corresponding automated virtual agents (e.g., chatbots) (162) experiences. The tools include a GT manager (152), a simulator (154), an evaluation manager (156), and a remediation manager (also referred to herein as the "director") (158).

[0024] The GT manager (152) can access one or more knowledge sources, such as libraries, as used herein as examples. A (172 AThe system is configured to automatically generate GTs from knowledge domains, such as those shown in Figure 1. The generated GTs can be content-based, usage-based, or curation-based, or a combination thereof. Content-based GTs are automatically generated by leveraging corresponding structured datasets to generate questions based on symptoms, question variants, and graph traversal, and then retrieving related entities for the symptoms. In one exemplary embodiment, a symptom is a phrase that describes some problem or issue with the system or any of its components. Details of content-based GT generation are shown and described in Figure 3. Usage-based GTs are directed to generating GTs from usage logs in the form of records of collected data, as a workflow for query text. Details of usage-based GT generation are shown and described in Figure 4. Curation-based GTs are directed to data manually generated by subject matter experts (SMEs). In this embodiment, the SMEs provide assistance with selecting options from the knowledge base to generate test data. Details of the curation-based GT generation are shown and described in Figure 5. Thus, each of the content-based, usage-based, and curation-based GTs leverages a corresponding knowledge domain in either a structured or unstructured format to assist and enable the automatic generation of GTs.

[0025] Each library is populated with one or more knowledge domains, represented as structured knowledge, e.g., knowledge graphs, that are processed by the GT Manager (152). As an example, the first knowledge domain is shown as KG0 (172) in FIG. A,0 ) and the corresponding content-based GT, i.e., GT 0,0(174 0,0 ), usage-based GT, i.e., GT 0,1 (174 0,1 ), and curation-based GT 0,2 (174 0,2 ), and the second knowledge domain is represented as a second KG, denoted as KG1 (172 A,1 ) as well as content-based GT, i.e., GT 1,0 (174 1,0 ), usage-based GT, i.e., GT 1,1 (174 1,1 ), and curation-based GT 1,2 (174 1,2 ), and the third knowledge domain is represented as a third KG, KG2 (172 A,2 ) as well as content-based GT, i.e., GT 2,0 (174 2,0 ), usage-based GT, i.e., GT 2,1 (174 2,1 ), and curation-based GT 2,2 (174 2,2 ), are shown. The three categories of GT presented herein, e.g., content, usage, and curation, may play different roles for the evaluation of a virtual agent. In one exemplary embodiment, any combination of the three categories of GT may be utilized for the evaluation of a virtual agent. The amount of GT categories presented herein should not be considered limiting. In one exemplary embodiment, the amount of categories may include a subset of the categories, a combination of the categories, or new categories.

[0026] Interactions with the chatbot (162) are in the form of sequences of queries and corresponding responses, as well as follow-up disambiguation questions and their responses. Such interactions, and specifically data associated with the interactions, are recorded and captured in one or more libraries of the knowledge base (170). In one or more embodiments, an initial NL query and a result are generated in terms of corresponding structured knowledge, e.g., a KG. In one or more other embodiments, an initial NL query is generated, and then one or more follow-up NL queries are generated to arrive at an NL result as part of a multi-turn or multi-step conversation or interaction. Generating a follow-up query or multiple follow-up queries is particularly useful, for example, when an initial response to an initial NL query does not provide a satisfactory answer, whether due to ambiguity in the initial response or for other reasons. In such cases, an initial follow-up query or disambiguation query is generated in terms of corresponding structured knowledge, e.g., a KG. If the first follow-up query does not provide a satisfactory answer, a second follow-up or disambiguation query is generated in light of the corresponding structured knowledge. This multi-turn conversation may continue until the disambiguation is satisfactorily resolved. For purposes of explanation, only the first and second sets of follow-up queries are described above. However, it should be understood that additional (e.g., third, fourth, etc.) follow-up queries may be generated as part of the multi-turn conversation or interaction. Thus, content-based GTs in the form of symptom-based generated questions, question variations, and knowledge graph traversals to obtain related entries for the symptoms are generated by the GT manager (152).

[0027] In addition to content-based GTs, the GT manager 152 generates usage-based GTs and curation-based GTs. The usage-based GTs are herein referred to as GTs. 0,1 (174 0,1 ), GT 1,1 (174 1,1 ), and GT 2,1 (174 2,1 ) The usage-based GT is composed of both log and feedback data. As an example, the usage-based GT, i.e., GT 0,1 (174 0,1 ), is a log 0,1 (176 0,1 ) and feedback 0,1 (178 0,1 ) and the usage-based GT, i.e., GT 1,1 (174 1,1 ), is a log 1,1 (176 1,1 ) and feedback 1,1 (178 1,1 ) and the usage-based GT, i.e., GT 2,1 (174 2,1 ), is a log 2,1 (176 2,1 ) and feedback 2,1 (178 2,1 ) Similarly, curation-based GT, i.e., GT 0,2 (174 0,2 ), is referred to herein as c_data 0,2 (178 0,2 ) is incorporated into the curation data, and it is a curation-based GT, that is, GT 1,2 (174 1,2 ), is c_data 1,2 (178 1,2 ) and curation-based GT, i.e., GT 2,2 (174 2,2 ), is c_data 2,2 (178 2,2) Thus, the GT manager (152) generates a plurality of categories of GTs, each of which is associated with a corresponding knowledge domain and stored in the knowledge base (170).

[0028] The dialogue system (160) is a conversational AI interface configured to support communication between a virtual agent and a non-virtual agent, which may be a human or software, such as a user (e.g., an end user), potentially an AI virtual agent. The interactions that occur generate what is referred to as a conversation or dialogue interaction, with the content of such conversation or dialogue interaction between the user and the virtual agent.

[0029] The AI ​​platform (150) is shown herein operatively connected to a dialogue system (160) and its virtual dialogue agent (162), which are configured to receive input (102) from various sources via a computer network (105). For example, the dialogue system (160) may receive input via the computer network (105) and utilize one of the knowledge domains and corresponding GTs to generate output or response content. The generated output or response content may be returned via the computer network (105) as output (104) to the same source, one or more other sources, or a combination thereof.

[0030] Various computing devices 180, 182, 184, 186, 188, and 190 communicating with the computer network 105 may provide access points to the dialogue system 160. The computer network 105 may include local network connections and remote connections in various embodiments, allowing the AI ​​platform 150 to operate in environments of any size, including local and global environments, such as the Internet. Additionally, the AI ​​platform 150 functions as a backend system that can make available various knowledge extracted from or represented in documents, network-access sources, or structured data sources, or a combination thereof. In this manner, several processes are incorporated into the AI ​​platform 150, which also includes an input interface for receiving requests and responding accordingly.

[0031] As shown, a user may access the AI ​​platform (150) and an operably connected dialogue system (160) via a network connection or an internet connection to the computer network (105) and may submit natural language (NL) input to the dialogue system (160), from which the AI ​​platform (150) may effectively determine an output response for the input by utilizing an operably connected knowledge base (170) and tools provided with the AI ​​platform (150).

[0032] The simulator (154) interfaces with the dialogue system (160) to simulate one or more NL dialogue interactions with the dialogue system's (160) automated virtual dialogue agent (162). In an exemplary embodiment, the simulator (154) includes an operably connected simulator application (154) for conducting simulated interactions with the chatbot (162). A), details of the simulation are shown and described in FIG. 6. The simulation defines a set of test queries with each answer existing in a corresponding knowledge domain, represented in an exemplary embodiment as a knowledge graph. The output from the simulation, referred to herein as simulation data, includes a log of all queries and corresponding responses, which in an exemplary embodiment includes solutions or one or more disambiguation options. A second library, i.e., a library B (172 B ) is captured in the knowledge base (170), and further captured is simulation data, referred to herein as data. As shown as an example herein, s_data0 (1540) is captured in the knowledge domain (172). A,0 ) and s_data1 (1541) represents the simulation data for the simulation of the chatbot (162) that utilizes the knowledge domain (172 A,1 ) and s_data2 (1542) represents the simulation data for the simulation of the chatbot (162) that utilizes the knowledge domain (172 A,2 ) represents simulation data for a simulation of a chatbot (162) utilizing the knowledge domains. While only one set of simulation data associated with each knowledge domain is shown, it is understood that any one of the knowledge domains may be utilized for the simulation of the chatbot (162), where each simulation may generate separate or additional simulation data. Similarly, while each knowledge domain is shown with simulation data, it is understood that in an exemplary embodiment, not all of the knowledge domains may be utilized for the simulation of the chatbot (162) and would not have corresponding simulation data. Thus, for each knowledge domain utilized by the dialogue simulator (154), output in the form of simulation data is created and associated with the corresponding knowledge domain.

[0033] As shown herein, the evaluation manager (156) is operatively connected to the simulator (154) and configured to evaluate the performance of the automated virtual dialogue agent (162). The evaluation manager (156) compares simulated interactions, represented as simulation data, with the GT for the corresponding knowledge domain. The GT used in the comparison may include one or more of the following GT types: content-based GT, usage-based GT, and curation-based GT. Details of the simulated interaction evaluation are shown and described in FIG. 7. The output from the evaluation manager is multidimensional, including the amount of disambiguation questions asked and their difference relative to the test data, whether the questions were asked in a particular order, and whether the proposed solution matches the intended solution. As shown by way of example herein, the output (1560) is a metric representing the performance of the knowledge domain (172). A,0 ) represents the multidimensional output of the evaluation of the simulation data (1540) utilizing the knowledge domain (172 A,1 ) and Output 2 (1562) represents the multidimensional output of the evaluation of the simulation data (1541) using the knowledge domain (172 A,2 ) represents the multidimensional output of the evaluation of the simulation data (1542) utilizing the data. Accordingly, the evaluation manager (156) performs an assessment of the chatbot (162) and documents the assessment in the form of corresponding output data.

[0034] The output data includes insights and recommendations based on the various metrics collected. Business objectives may be predefined in terms of metrics and corresponding metric measurements, such as the chatbot's expected accuracy, along with acceptable error ranges. Examples of such metrics include, but are not limited to, accuracy, interaction overhead, interaction length, quality of follow-up questions, and response time. In one exemplary embodiment, the metrics may be prioritized, e.g., accuracy may be prioritized instead of response time. Recommendations corresponding to the collected metrics are directed toward an automated comparison of defined or predefined business objectives against actual metrics reflecting performance and the identification of one or more corresponding remediation actions. As shown herein, a remediation manager (158) is operatively connected to the assessment manager (156) and functions to identify one or more remediation actions to apply to the dialogue system (160) based on the corresponding output. For example, in one embodiment, one or more remediation actions may be identified when a performance evaluation of the virtual dialogue agent (162) fails to meet a performance threshold. In one exemplary embodiment, the one or more recommended actions, also referred to herein as a recommendation plan, are directed to improving interaction overhead and shortening the length of interactions, which may be implemented by collecting additional real-time data. In one embodiment, other recommendations may be implemented, and therefore the examples provided herein should not be considered limiting. Accordingly, the remediation manager (158) is configured to implement one or more remediation actions to improve the performance of the automated virtual dialogue agent (162).

[0035] The interaction events generated or enabled by the interaction system (160) are (登録商標) Watson (登録商標)The server 110 may be processed by a corresponding AI platform 150. The GT manager 152 facilitates and enables the generation of GTs from the user's knowledge base and the evaluation of the dialogue system 160 as supported by the generated GT. In some exemplary embodiments, the server 110 may be a GT manager from International Business Machines Corporation of Armonk, NY, augmented with mechanisms of the exemplary embodiments described below. (登録商標) Available from IBM (登録商標) Watson (登録商標) It can be a system.

[0036] The GT manager 152, simulator 154, evaluation manager 156, and remediation manager 158, hereinafter collectively referred to as AI tools, are shown as embodied within or integrated into the AI ​​platform 150 of the server 110. The AI ​​tools may also be implemented within another computing system (e.g., 190) connected to the server 110 via the computer network 105. Wherever embodied, the AI ​​tools function to evaluate dialogue events, extract behavioral characteristics from requests and responses, and selectively identify and apply one or more corresponding remediation actions to improve the performance of the dialogue system 160.

[0037] The types of information handling systems that can utilize the artificial intelligence platform (150) range from small handheld devices, e.g., handheld computers / cell phones (180), to large mainframe systems, e.g., mainframe computers (182). Examples of handheld computers (180) include personal digital assistants (PDAs), personal entertainment devices (e.g., MP4 players), portable televisions, and compact disc players. Other examples of information handling systems include pen or tablet computers (184), laptop or notebook computers (186), personal computer systems (188), and servers (190). As shown, various information handling systems can be networked together using a computer network (105). Types of computer networks (105) that can be used to interconnect various information handling systems include local area networks (LANs), wireless local area networks (WLANs), the Internet, public switched telephone networks (PSTNs), other wireless networks, and any other network topology that can be used to interconnect information handling systems. Many information handling systems include a non-volatile data store, such as a hard drive or non-volatile memory, or a combination thereof. Some information handling systems may use a separate non-volatile data store (e.g., the server (190) may use a non-volatile data store (190)). A ), and the mainframe computer (182) uses a non-volatile data store (182 A )). Non-volatile data store (182 A ) can be a component external to the various information handling systems, or can be internal to one of the information handling systems.

[0038] The information handling system used to support the AI ​​platform 150 may take many forms, some of which are shown in Figure 1. For example, the information handling system may take the form of a desktop, server, portable, laptop, notebook, or other form factor computer or data processing system. In addition, the information handling system may take the form of other form factors, such as a personal digital assistant (PDA), gaming device, ATM machine, mobile phone device, communications device, or other device with a processor and memory.

[0039] An application program interface (API) is understood in the art as a software intermediary between two or more applications. With respect to the artificial intelligence platform 150 shown and described in FIG. 1, one or more APIs may be utilized to support one or more of the tools 152, 154, 156, and 158 and their associated functionality. Referring to FIG. 2, a block diagram 200 is provided illustrating the tools 152, 154, 156, and 158 and their associated APIs. As shown, multiple tools are embedded within the AI ​​platform 205, including a GT manager 252 associated with API 0 212, a simulator 254 associated with API 1 222, an evaluation manager 256 associated with API 2 232, and a remediation manager 258 associated with API 3 242. Each API may be implemented in one or more languages ​​and interface specifications. API 0 (212) provides functional support for automatically generating a GT from a knowledge source; API 1 (222) provides functional support for simulating a NL interaction with the automated virtual agent utilizing the GT; API 2 (232) provides functional support for evaluating the performance of the automated virtual dialogue agent based on the simulation; and API 3 (242) provides functional support for selectively identifying and implementing one or more remediation actions aimed at improving the performance of the dialogue system. As shown, each of APIs 212, 222, 232, and 242 is operably connected to an API orchestrator 270, also known as an orchestration layer, which is understood in the art to act as an abstraction layer for transparently threading separate APIs together. In one embodiment, the functionality of separate APIs may be combined or merged.As such, the configuration of APIs shown herein should not be considered limiting, and the functionality of the multiple tools may be embodied or supported by their respective APIs as shown herein.

[0040] Referring to FIG. 3 , a flowchart (300) illustrating a process for automatically generating a ground truth (GT) from corresponding knowledge sources is provided. As shown and described, the knowledge sources can be in a structured format, such as a knowledge graph, or in an unstructured format. For purposes of explanation, the GT generation process is described with respect to a structured knowledge source, but such a structured format should not be considered limiting. Relevant knowledge sources are identified, and a set of symptoms from the knowledge sources is obtained (302). In one exemplary embodiment, a subset of symptoms is identified using one or more selection criteria. For each symptom, a natural language query is generated (304), which, in an exemplary embodiment, uses variance generation and the addition or deletion of entities. In one embodiment, variance generation is the natural language equivalent of a phrase and is utilized herein to expand the scope of a query through the identification of comparable or equivalent terms. The knowledge graph, e.g., a structured representation of a knowledge domain, is searched for query text matching the symptoms (306). In one exemplary embodiment, a text matching technique, e.g., universal sentence encoding, is utilized in step (306). Output is generated from the search in step (306) in the form of matching symptoms, also referred to herein as matches (308). Each matching symptom has a corresponding score or weight. In one exemplary embodiment, approximate matching of two or more phrases or sentences is a common operation in natural language processing. The set of matching symptoms from step (308), in an exemplary embodiment, is subjected to a threshold evaluation that is directed to the quality of the matching symptom. Each matching symptom has a constraint. A constraint node and nodes connected to the constraint node are fetched for each matching symptom (310). A constraint node is not connected to other constraint nodes.The fetching in step 310 is directed to identifying both one or more constraint nodes and all other nodes connected to the one or more constraint nodes. For example, a solution to the symptom "battery problem while charging," and a specific solution to this symptom, may be constrained by a specific hardware model number and series. Thus, the one or more constraint nodes are connected to all relevant nodes in the graph as indices of the constraint.

[0041] Variable X 合計 is assigned to the quantity of the constraint (312), and the corresponding constraint count variable X is initialized (314). X , a set of disambiguation questions with answer options is generated (316). In one exemplary embodiment, multiple constraints represent a multiple-step conversation, and at each step, a disambiguation question and answer option are generated, and the process is repeated until no more disambiguation is required. Following step (316), a constraint count variable X is incremented (318), and it is determined (320) whether each of the constraints has been processed. A negative response to the determination returns to step (316), while an affirmative response terminates the generation of the question and one or more answers. Thus, one or more disambiguation questions and one or more corresponding answer options for each constraint and each disambiguation selection path are recorded and saved as a GT.

[0042] As shown and described in FIG. 3, a content-based GT is automatically generated by utilizing a knowledge graph to generate a question and performing graph traversal to identify one or more relevant entities for one or more corresponding answers. Two other forms of GTs are generated, including usage-based GTs and curation-based GTs, as shown and described in FIG. 1. Referring to FIG. 4, a flowchart diagram (400) is provided to explain the process for generating a usage-based GT. A usage log is provided (402) that records or records the original query text, any follow-up questions posed to the user, and the selections made by the user along with a final solution or action plan. The usage log includes feedback on whether the query as expressed in the query text was satisfactorily answered. The variable X 合計 is assigned 404 to the amount of queries in the usage log that have feedback indicating at least a satisfactory solution. A corresponding query count variable X is initialized 406. X If so, the query text is obtained from the usage log (408), and the query XA follow-up question for the query is also retrieved from the usage log (410), and a user selection is retrieved (412). In an interaction between the user and the chatbot, the chatbot may provide the user with a follow-up question and selectable options as answers. This interaction, including the selected answers for the options, is collected or retrieved in steps (408) through (412). Next, it is determined (414) whether a solution to the query has been reached. The system knows that the information it is sending to the user is a follow-up question and also knows which information is a solution to the user's problem. A solution to the query is determined to have been reached when the system sends the solution and does not send another follow-up question. A negative response to the determination in step (414) is followed by incrementing the query count variable (416) and returning to step (408). Conversely, a positive response to the determination in step (414) is followed by identifying the user selection as a solution (418). The query text, one or more follow-up questions, and user selections obtained in steps 408-412 and 418, respectively, are recorded as a workflow, referred to herein as a usage-based GT. Accordingly, the usage log is evaluated to identify and record one or more questions and user selections corresponding to the query text as a usage-based GT.

[0043] Referring to FIG. 5, a flowchart diagram (500) is provided illustrating a process for generating a curation-based GT. In one exemplary embodiment, the curation-based GT is manually generated by one or more subject matter experts (SMEs). As shown, the SME is provided with a list of symptoms in a knowledge base for reference (502). One or more text queries are submitted to the SME (504), and the SME is provided with one or more potential answers to the text queries and associated entities, e.g., constraints, for each question for reference (506). The SME selects a preferred answer to the query and, optionally, a preferred sequence of follow-up questions for disambiguation (508). In the case of a disambiguation question, a consistency check is performed to verify whether one or more disambiguation options match a corresponding knowledge base, which may optionally be updated for consistency of expression (510). The text query and corresponding answer, and in an exemplary embodiment, a sequence of one or more follow-up questions for disambiguation, are recorded and stored as a GT. Thus, a record of the curated GT is provided with the assistance of an SME.

[0044] As shown and described in Figure 1, a simulator (154) is provided to support the simulation of NL dialogue interactions. Referring to Figure 6, a flow chart diagram (600) is provided to illustrate a process for simulating an interaction with a dialogue system (160). A disambiguation selection path count variable N is initialized (602), GT data is utilized as a source to drive the interaction with the virtual dialogue agent, and a query, e.g., query N, is generated and sent to an automated virtual agent using an operationally connected simulator application (604). The query is recorded in a corresponding simulation log (606). The automated virtual agent responds to the query with a solution or a set of disambiguation options (608). If the response is the solution, the solution is recorded in the corresponding log (610). If the response is a set of disambiguation questions, a disambiguation selection path count variable N is incremented (612), and then the GT is consulted to find a disambiguation question in the GT for this query (614). In one exemplary embodiment, if the disambiguation question is not found in the GT for this query, depending on the configuration, the process may stop, select "any" if options are provided, or randomly select one of multiple options provided. If a question is selected, the process returns to step (606). Following step (610), the number of disambiguation selection paths is incremented to variable N. 合計 is assigned 616. A log of the questions and answers obtained or identified from the simulation is then recorded in a simulation log 618. Thus, the simulator application creates a simulation interaction log that records the inputs received and outputs generated from the corresponding GT.

[0045] The dialogue system (160) and the corresponding automated virtual agent (162) are subjected to a performance evaluation by utilizing the GT and the corresponding simulation log. Referring to Figure 7, a flow chart (700) is provided to illustrate a process for conducting a performance evaluation of the virtual dialogue system. As shown, the variable N 合計 is assigned the number of query-responses recorded in the simulation log (702), and a corresponding count variable N is initialized (704). NFor each of the multiple dimensions, a corresponding entry in the GT is found (704). The query-response in the simulation log is compared to the query-response in the GT (706). From the comparison in step (706), a multidimensional output is generated, including: 1. the difference between the amount of disambiguation questions actually asked and the amount in the GT; 2. whether the questions were asked in the same order; and 3. whether the solution presented in the simulation log matches the GT solution. Based on the output generated for each of the multiple dimensions, insights and recommendations are generated (708). In one exemplary embodiment, one or more additional dimensions may be added to the evaluation, or conversely, a reduced amount of dimensionality may be used for the evaluation. Thus, as illustrated herein, the comparison of the simulation log and the GT provides insight into the performance of the dialogue system (160).

[0046] The output data from step 708 includes insights and recommendations based on the various metrics collected. Business objectives may be predefined in terms of metrics and corresponding metric measurements, such as the chatbot's expected accuracy, along with acceptable error ranges. Examples of such metrics include, but are not limited to, accuracy, interaction overhead, interaction length, quality of follow-up questions, and response time. In one exemplary embodiment, the metrics may be prioritized, e.g., accuracy may be prioritized instead of response time. Recommendations corresponding to the collected metrics are directed toward automated comparison of defined or predefined business objectives against actual metrics reflecting performance and identification of one or more corresponding remediation actions. As illustrated herein, one or more remediation actions are identified 710 and selectively implemented 712 for application to the dialogue system 160 based on the corresponding output. For example, in one embodiment, one or more remediation actions may be identified when a performance evaluation of the virtual dialogue agent (162) fails to meet a performance threshold. In one exemplary embodiment, the one or more recommended actions, also referred to herein as a recommendation plan, are directed to improving interaction overhead and shortening the length of the interaction, which may be implemented by collecting additional real-time data. In one embodiment, other recommendations may be implemented, and therefore the examples provided herein should not be considered limiting. Thus, the remediation actions are directed to remediating the performance of the dialogue system (160) and the corresponding automated virtual dialogue agent (162).

[0047] As shown and described in Figures 1-7, computer systems, program products, and methods are provided for evaluating the performance of a multi-turn automated virtual agent using a GT automatically generated from an operably connected knowledge source. A simulation of a NL interaction is performed using the automated virtual agent and utilizing the GT to drive the corresponding interaction. A log is created to record the simulation. Performance of the automated virtual agent is evaluated by comparing the simulation log with the corresponding GT. One or more remediation actions directed at improving the performance of the automated virtual agent are identified and selectively implemented based on the simulation and performance evaluation of the simulation log.

[0048] The embodiments shown and described herein may be in the form of a computer system for use with an intelligent computing platform for enhancing the performance of a dialogue system and the corresponding automated virtual agent. Aspects of tools 152, 154, 156, and 158 and their associated functionality may be embodied in a computer system / server at a single location or, in one embodiment, configured in a cloud-based system sharing computing resources. Referring to FIG. 8, a block diagram 800 is provided illustrating an example of a computer system / server 802, referred to as a host 802, in communication with a cloud-based support system 810 for executing the systems, tools, and processes described above in FIGS. 1-7. In one embodiment, the host 802 is a node in a cloud computing environment. The host 802 is operable in numerous other general-purpose or special-purpose computing system environments or configurations. Examples of well-known computing systems, environments, or configurations, or combinations thereof, suitable for use with host (802) include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and file systems (e.g., distributed storage environments and distributed cloud computing environments) comprising any of the above systems, devices, and their equivalents.

[0049] The host 802 may be described in the general context of computer system-executable instructions, such as program modules, being executed by a computer system. Generally, program modules may include routines, programs, objects, components, logic, data structures, etc. that perform particular tasks or implement particular abstract data types. The host 802 may be implemented in a distributed cloud computing environment where tasks are performed by remote processing devices that are linked through a communications network. In a distributed cloud computing environment, program modules may be located in both local and remote computer system storage media, such as memory storage devices.

[0050] As shown in Figure 8, the host (802) is depicted in the form of a general-purpose computing device. Components of the host (802) may include, but are not limited to, one or more processors or processing units (804), such as a hardware processor, a system memory (806), and a bus (808) connecting various system components, including the system memory (806), to the one or more processors or processing units (804). The bus (808) represents any one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, and a processor or local bus using any of a variety of bus architectures. By way of example and not limitation, such architectures include the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MCA) bus, the Enhanced ISA (EISA) bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnect (PCI) bus. The host 802 typically includes a variety of computer system-readable media. Such media may be any available media accessible by the host 802 and include both volatile and nonvolatile media, removable and non-removable media.

[0051] The system memory 806 may include computer system-readable media in the form of volatile memory, such as random access memory (RAM) 830 or cache memory 832, or a combination thereof. By way of example only, a storage system 834 may be provided for reading from and writing to non-removable, non-volatile magnetic media (not shown, typically referred to as a "hard drive"). Although not shown, a magnetic disk drive may be provided for reading from and writing to removable, non-volatile magnetic disks (e.g., floppy disks), and an optical disk drive may be provided for reading from and writing to removable, non-volatile optical disks, such as CD-ROMs, DVD-ROMs, or other optical media. In such cases, each may be connected to the bus 808 by one or more data media interfaces.

[0052] The programs / utilities 840 may include at least one set of program modules 842, which may be stored in the system memory 806, including, but not limited to, an operating system, one or more application programs, other program modules, and program data. Each of the operating system, one or more application programs, other program modules, and program data, or some combination thereof, may comprise an implementation of a networking environment. The program modules 842 generally perform an embodiment function or method, or a combination thereof, for dynamically interpreting and understanding requests and action descriptions and effectively augmenting corresponding domain knowledge. For example, the set of program modules 842 may include tools 152, 154, 156, and 158, as shown in FIG. 1.

[0053] The host 802 may also communicate with one or more external devices 814, such as a keyboard, pointing device, etc.; a display 824; one or more devices that allow a user to interact with the host 802; or any device that allows the host 802 to communicate with one or more other computing devices (e.g., a network card, modem, etc.); or any combination thereof. Such communication may occur via one or more input / output (I / O) interfaces 822. The host 802 may also communicate with one or more networks, such as a local area network (LAN), a general wide area network (WAN), or a public network (e.g., the Internet), or combinations thereof, via a network adapter 820. As shown, the network adapter 820 communicates with the other components of the host 802 via a bus 808. In one embodiment, multiple nodes (not shown) of the distributed file system communicate with the host (802) via an I / O interface (822) or via a network adapter (820). Although not shown, it should be understood that other hardware or software components, or combinations thereof, may be used with the host (802). Examples include, but are not limited to, microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, data archival storage systems, and the like.

[0054] As used herein, the terms "computer program medium," "computer usable medium," and "computer readable medium" are used to generally refer to media such as main memory (806), including RAM (830) and cache (832), storage systems (834), such as removable storage drives and hard disks installed in hard disk drives.

[0055] Computer programs (also referred to as computer control logic) are stored in memory 806. Computer programs may also be received via a communications interface, such as a network adapter 820. When executed, such computer programs enable the computer system to perform the features of the present embodiments as discussed herein. In particular, when executed, such computer programs enable the processing device 804 to perform the functions of the computer system. Thus, such computer programs represent the controller of the computer system.

[0056] The computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction-execution device. The computer-readable storage medium can be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. A non-exhaustive list of more specific examples of the computer-readable storage medium includes the following: a portable computer diskette (登録商標), hard disk, random access memory (RAM), read only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanically encoded device such as punch cards or raised structures in grooves on which instructions are recorded, or any suitable combination thereof. As used herein, the computer-readable storage medium should not be construed to be a transitory signal per se, such as an electric wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., light pulses passing through a fiber optic cable), or an electrical signal transmitted over an electrical wire.

[0057] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing device / processing device, or to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof. The network may be comprised of copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. A network adapter card or network interface in each computing device / processing device receives the computer-readable program instructions from the network and transmits the computer-readable program instructions for storage in a computer-readable storage medium within the respective computing device / processing device.

[0058] The computer-readable program instructions for carrying out the operations of the present invention may be either assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, configuration data for an integrated circuit, or source or object code written in any combination of one or more programming languages, such as object-oriented programming languages, e.g., Smalltalk, C++, etc., or conventional procedural programming languages ​​(e.g., the "C" programming language or similar programming languages). The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, partially on the user's computer as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any kind of network, such as a local area network (LAN) or a wide area network (WAN), or the connection may be to an external computer (e.g., over the Internet using an Internet Service Provider). In some embodiments, electronic circuits, such as programmable logic circuits, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), may execute computer-readable program instructions by utilizing state information of the computer-readable program instructions to personalize the electronic circuitry to perform aspects of the invention.

[0059] The functional tools described herein are labeled as managers. Managers may be implemented in programmable hardware devices, such as field programmable gate arrays, programmable array logic, or programmable logic devices. The managers may also be implemented in software for processing by various types of processors. An identified manager of executable code may comprise, for example, one or more physical or logical blocks of computer instructions, which may be organized as objects, procedures, functions, or other constructs. Nevertheless, the executable code of an identified manager need not be physically located together and may include heterogeneous instructions stored in different locations that, when logically connected, constitute the manager and achieve the manager's stated purpose.

[0060] In fact, a manager of executable code may be a single instruction or many instructions, and may be distributed among several different code segments, different applications, and across several memory devices. Similarly, operational data may be identified in the manager and illustrated herein, may be embodied in any suitable form, and may be organized within any suitable type of data structure. The operational data may be collected as a single data set or may be distributed in different locations, including on different storage devices, and may exist at least in part as electronic signals on a system or network.

[0061] Referring now to FIG. 9, an exemplary cloud computing network (900) is illustrated. As shown, the cloud computing network (900) comprises a cloud computing environment (950) having one or more cloud computing nodes (910) with which local computing devices used by cloud consumers may communicate. Examples of these local computing devices include, but are not limited to, a personal digital assistant (PDA) or mobile phone (954A), a desktop computer (954B), a laptop computer (954C), or an automobile computer system (954N), or combinations thereof. The nodes within the nodes (910) may also communicate with each other. They may be physically or virtually grouped together in one or more networks (not shown), such as a private cloud, community cloud, public cloud, or hybrid cloud, or combinations thereof, as described herein. This allows the cloud computing environment (900) to provide infrastructure, platform, or software, or combinations thereof, as a service without the cloud consumer having to maintain resources on their local computing devices. It is understood that the types of computing devices (954A-954N) shown in FIG. 9 are intended to be illustrative only, and that the cloud computing environment (950) can communicate with any type of computerized device (e.g., using a web browser) via any type of network or network-addressable connection or combination thereof.

[0062] Referring now to Figure 10, there is shown a set of functional abstraction layers (1000) provided by the cloud computing network of Figure 9. It should be understood that the components, layers, and functions shown in Figure 10 are intended to be merely exemplary, and that the present embodiment is not limited thereto. As shown, the following layers and corresponding functions are provided: a hardware and software layer (1010), a virtualization layer (1020), a management layer (1030), and a workload layer (1040).

[0063] The hardware and software layer (1010) includes hardware components and software components. An example of a hardware component is a mainframe, for example, an IBM (登録商標) zSeries (登録商標) Systems: RISC (Reduced Instruction Set Computer) architecture-based servers, such as IBM (登録商標) pSeries (登録商標) Systems; IBM (登録商標) xSeries (登録商標) Systems; IBM (登録商標) BladeCenter (登録商標) Examples of software components include network application server software, such as IBM (登録商標) WebSphere (登録商標) Application server software; and database software, such as IBM (登録商標) DB2 (登録商標) Includes database software (IBM (登録商標) , zSeries, pSeries, xSeries, BladeCenter, WebSphere, and DB2 are trademarks of International Business Machines Corporation, registered in many jurisdictions worldwide. (登録商標) are trademarks of

[0064] The virtualization layer (1020) provides an abstraction layer from which the following examples of virtual entities can be provided: virtual servers; virtual storage; virtual networks, such as those described above, including virtual private networks; virtual applications and operating systems; and virtual clients.

[0065] In one example, the management layer (1030) may provide the following functions: resource provisioning, metering and pricing, a user portal, service level management, and service level agreement (SLA) planning and fulfillment. Resource provisioning provides dynamic procurement of computing and other resources used to perform tasks within the cloud computing environment. Metering and pricing provides cost tracking as resources are utilized within the cloud computing environment and billing or invoicing for the consumption of these resources. In one example, these resources may include application software licenses. Security provides identity verification for cloud consumers and tasks and protection for data and other resources. The user portal provides access to the cloud computing environment for consumers and system administrators. Service level management provides allocation and management of cloud computing resources so that required service levels are met. Service level agreement (SLA) planning and fulfillment provides pre-allocation and procurement of cloud computing resources when future requirements are anticipated according to SLAs.

[0066] The workload layer (1040) provides examples of functions for which the cloud computing environment may be utilized. Examples of workloads and functions that may be provided from this layer include, but are not limited to: mapping and navigation; software development and lifecycle management; virtual classroom instruction delivery; data analytics processing; transaction processing; and virtual interactive system evaluation and enhancement.

[0067] While particular embodiments of the present invention have been shown and described, it will be apparent to those skilled in the art that, based on the teachings herein, changes and modifications can be made without departing from the present invention and its broader aspects. Therefore, the appended claims are intended to encompass within their scope all such changes and modifications as are within the true spirit and scope of the present invention. Moreover, it is to be understood that the present invention is defined solely by the appended claims. Where a specific number of introduced claim elements is intended, such intent will be expressly set forth in the claims; in the absence of such recitation, those skilled in the art will understand that no such limitation exists. As a non-limiting example, and to aid in understanding, the appended claims include the use of the introductory phrases "at least one" and "one or more" to introduce claim elements. However, the use of such phrases should not be construed as limiting the introduction of a claim element by the indefinite article "a" or "an" to embodiments in which any particular claim containing that introduced claim element includes only one of that element, even if the same claim contains the introductory phrases "one or more" or "at least one" and an indefinite article such as "a" or "an," and the same is true even if a definite article is used in the claim. As used herein, the word "and / or" means either one or both (or either or any combination or all of the referenced words or phrases).

[0068] The present embodiments may be a system, a method, or a computer program product, or a combination thereof. In addition, selected aspects of the present embodiments may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, all of which may be generally referred to herein as a "circuit," "module," or "system." Moreover, aspects of the present embodiments may take the form of a computer program product embodied in one or more computer-readable storage media having computer-readable program instructions thereon for causing a processor to execute aspects of the present embodiments. The disclosed system, method, or computer program product, or a combination thereof, so embodied, may be operable to assist in the evaluation and enhancement of virtual interaction systems.

[0069] Aspects of the present embodiments are described herein with reference to flowchart illustrations or block diagrams, or combinations thereof, of methods, apparatus (systems), and computer program products according to the embodiments. It will be understood that each block of the flowchart illustrations or block diagrams, or combinations thereof, and combinations of blocks in the flowchart illustrations or block diagrams, or combinations thereof, can be implemented by computer-readable program instructions.

[0070] These computer-readable program instructions may be provided to a processor of a computer or other programmable data processing apparatus, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts identified in one or more blocks of the flowchart diagrams or block diagrams, or a combination thereof, to produce a machine. These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer-programmable data processing apparatus or other device, or a combination thereof, to function in a particular manner, such that a computer-readable storage medium having stored instructions includes an article of manufacture including instructions that implement aspects of the functions / acts identified in one or more blocks of the flowchart diagrams or block diagrams, or a combination thereof.

[0071] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device such that the instructions, which execute on the computer, other programmable data processing apparatus, or other device, implement the functions / acts identified in one or more blocks of the flowchart diagrams or block diagrams, or a combination thereof, to cause the computer, other programmable apparatus, or other device to perform a series of operational steps to generate a computer-implemented process.

[0072] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products or computer programs according to various embodiments of the present invention. In this regard, each block in the flowcharts or block diagrams may represent a module, segment, or portion of instructions, which includes one or more executable instructions for implementing one or more specified logical functions. In some alternative implementations, the functions shown in the blocks may occur out of the order shown in the figures. For example, two blocks shown in succession may actually be accomplished as a single step performed simultaneously, substantially simultaneously, partially, or fully in a time-overlapping manner, depending on the functionality involved, or the blocks may be performed in the reverse order. It should be noted that each block of the block diagrams or flowchart diagrams or combinations thereof, and combinations of multiple blocks in the block diagrams or flowchart diagrams or combinations thereof, may be implemented by a special-purpose hardware-based system that performs the specified functions or operations, or may execute a combination of special-purpose hardware and computer instructions.

[0073] Although particular embodiments have been described herein for purposes of illustration, it will be understood that various modifications can be made without departing from the spirit and scope of the embodiments. Accordingly, the scope of protection for the embodiments is limited only by the appended claims and their equivalents.

Claims

1. 1. A computer system, comprising: a processor operably connected to a memory; and an artificial intelligence platform operably connected to the processor; the artificial intelligence platform comprises one or more tools for interfacing with a virtual conversational agent; The tool may include: a ground truth manager configured to automatically generate ground truth from the knowledge source; a simulator configured to simulate a natural language dialogue interaction with the virtual dialogue agent, wherein the simulator is configured to utilize the ground truth to drive simulated natural language dialogue generated output and to create a corresponding simulation log; an evaluation manager configured to evaluate the performance of the virtual conversational agent with respect to the created simulation log in terms of the ground truth; and a remediation manager configured to identify one or more remediation actions for the virtual dialogue agent in response to the evaluated performance not meeting a performance threshold, and to selectively implement one or more of the identified remediation actions. It is equipped with the identified remediation action is improving interaction overhead, which may be implemented by collecting additional real-time data, or shortening the length of the interaction; The computer system.

2. 1. A computer system, comprising: a processor operably connected to a memory; and an artificial intelligence platform operably connected to the processor; the artificial intelligence platform comprises one or more tools for interfacing with a virtual conversational agent; The tool may include: a ground truth manager configured to automatically generate ground truth from the knowledge source; a simulator configured to simulate a natural language dialogue interaction with the virtual dialogue agent, wherein the simulator is configured to utilize the ground truth to drive simulated natural language dialogue generated output and to create a corresponding simulation log; an evaluation manager configured to evaluate the performance of the virtual conversational agent with respect to the created simulation log in terms of the ground truth; and a remediation manager configured to identify one or more remediation actions for the virtual dialogue agent in response to the evaluated performance not meeting a performance threshold, and to selectively implement one or more of the identified remediation actions. It is equipped with The ground truth manager is further configured to compile a first disambiguation selection path, the compiling comprising: generating a natural language query and at least one disambiguating natural language query; generating natural language results in response to the at least one disambiguated natural language query; and recording a first log about the first disambiguation selection path; The computer system.

3. A computer system as described in claim 1 or 2, wherein the ground truth data includes usage logs and feedback corresponding to the usage logs, structured data, records generated by subject matter experts, or any combination thereof.

4. The computer system of claim 1 or 2, wherein the evaluation manager is configured to compare pairs of query responses in the ground truth with corresponding pairs of query responses in the simulation log.

5. The simulator is further configured to compile a second disambiguation selection path, the compiling comprising: generating a test natural language query and at least one test disambiguating natural language query; generating a test natural language response to the at least one test disambiguated natural language query; and recording a second log about the second disambiguation selection path; 3. The computer system of claim 2, comprising:

6. The computer system of claim 5 , wherein the assessment manager is further configured to compare the first recorded log with the second recorded log.

7. 1. A computer program for improving the performance of a virtual conversational agent, comprising: Automatically generating ground truth from knowledge sources; simulating a natural language dialogue interaction with the virtual dialogue agent, wherein the simulating includes utilizing the ground truth to drive simulated natural language dialogue generated output and creating a corresponding simulation log; evaluating the performance of the virtual conversational agent with respect to the created simulation log in terms of the ground truth; identifying one or more remediation actions for the virtual dialogue agent in response to the evaluated performance not meeting a performance threshold; and Selectively implementing one or more of the remediation actions. causing a computer processor to execute each step of the method, including the identified remediation action is improving interaction overhead, which may be implemented by collecting additional real-time data, or shortening the length of the interaction; The computer program.

8. 1. A computer program for improving the performance of a virtual conversational agent, comprising: Automatically generating ground truth from knowledge sources; simulating a natural language dialogue interaction with the virtual dialogue agent, wherein the simulating includes utilizing the ground truth to drive simulated natural language dialogue generated output and creating a corresponding simulation log; evaluating the performance of the virtual conversational agent with respect to the created simulation log in terms of the ground truth; identifying one or more remediation actions for the virtual dialogue agent in response to the evaluated performance not meeting a performance threshold; and Selectively implementing one or more of the remediation actions. causing a computer processor to execute each step of the method, including Leveraging the ground truth data includes compiling a first disambiguation selection path, the compiling comprising: generating a natural language query and at least one disambiguating natural language query; generating natural language results in response to the at least one disambiguated natural language query; and recording a first log about the first disambiguation selection path; Including, The computer program.

9. A computer program as described in claim 7 or 8, wherein the ground truth data includes usage logs and feedback corresponding to the usage logs, a knowledge graph, records generated by subject matter experts, or any combination thereof.

10. The computer program product of claim 7 or 8, wherein the evaluating comprises comparing pairs of query responses in the ground truth with corresponding pairs of query responses in the simulation log.

11. The simulating further includes compiling a second disambiguation selection path, the compiling comprising: generating a test natural language query and at least one test disambiguating natural language query; generating a test natural language response to the at least one test disambiguated natural language query; and recording a second log about the second disambiguation selection path; 9. The computer program of claim 8, comprising:

12. The computer program of claim 11, wherein evaluating the performance of the virtual conversational agent further comprises comparing the recorded first log with the recorded second log.

13. 1. A computer-implemented method directed to improving the performance of a virtual conversational agent, comprising: automatically generating, by a computer processor, a ground truth from a knowledge source; simulating, by the computer processor, a natural language dialogue interaction with the virtual dialogue agent, wherein the simulating includes utilizing the ground truth to drive simulated natural language dialogue generated output and creating a corresponding simulation log; evaluating, by the computer processor, the performance of the virtual conversational agent with respect to the created simulation log in terms of the ground truth; identifying, by the computer processor, one or more remediation actions for the virtual conversational agent in response to the evaluated performance not meeting a performance threshold; and selectively implementing, by said computer processor, said identified one or more remediation actions. Including, the identified remediation action is improving interaction overhead, which may be implemented by collecting additional real-time data, or shortening the length of the interaction; The method.

14. 1. A computer-implemented method directed to improving the performance of a virtual conversational agent, comprising: automatically generating, by a computer processor, a ground truth from a knowledge source; simulating, by the computer processor, a natural language dialogue interaction with the virtual dialogue agent, wherein the simulating includes utilizing the ground truth to drive simulated natural language dialogue generated output and creating a corresponding simulation log; evaluating, by the computer processor, the performance of the virtual conversational agent with respect to the created simulation log in terms of the ground truth; identifying, by the computer processor, one or more remediation actions for the virtual conversational agent in response to the evaluated performance not meeting a performance threshold; and selectively implementing, by said computer processor, said identified one or more remediation actions. Including, Leveraging ground truth data includes compiling, by the computer processor, a first disambiguation selection path, the compiling comprising: generating a natural language query and at least one disambiguating natural language query; generating natural language results in response to the at least one disambiguated natural language query; and recording a first log about the first disambiguation selection path; Including, Computer-implemented methods.

15. A computer-implemented method as described in claim 13 or 14, wherein the ground truth data includes usage logs and feedback corresponding to the usage logs, structured data, records generated by subject matter experts, or any combination thereof.

16. 15. The computer-implemented method of claim 13 or 14, wherein the evaluating comprises comparing pairs of query responses in the ground truth with corresponding pairs of query responses in the simulation log.

17. The simulating includes compiling, by the computer processor, a second disambiguation selection path, the compiling comprising: generating a test natural language query and at least one test disambiguating natural language query; generating a test natural language response to the at least one test disambiguated natural language query; and recording a second log about the second disambiguation selection path; 15. The computer-implemented method of claim 14, comprising:

18. The computer-implemented method of claim 17, wherein evaluating the performance of the virtual conversational agent further comprises comparing the recorded first log with the recorded second log.

Citation Information

Patent Citations

  • Cognitive Moderator for Cognitive Instances

    JP2020532804A

  • Collaborative-filtering based user simulation for dialog systems

    US10706086B1

  • Systems and methods for conducting multi-task oriented dialogues

    US20200110915A1