Solution-guided dialog system response generation
By identifying topics in user speech through a processor, generating dialogue system responses using text classification and sequence-to-sequence models, and combining solution segments provided by subject matter experts, this approach solves the problems of high cost and difficulty in learning in existing technologies, and achieves intelligent and automated dialogue generation for a more intelligent dialogue system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- INTERNATIONAL BUSINESS MACHINE CORPORATION
- Filing Date
- 2022-06-09
- Publication Date
- 2026-06-19
AI Technical Summary
Existing solutions generation methods for dialogue systems are costly, labor-intensive, and struggle to effectively learn and recognize external information, leading to modeling difficulties.
The system uses a processor to identify topics in user speech, generates responses using text classification and sequence-to-sequence machine learning models, and combines solution segments provided by topic experts with artificial intelligence models for natural language processing and text generation to generate responses for the dialogue system.
It reduces the modeling cost of dialogue systems, improves the efficiency and accuracy of dialogue generation, and enables intelligent and automated response generation of dialogue systems.
Smart Images

Figure CN115565530B_ABST
Abstract
Description
Background Technology
[0001] This disclosure generally relates to the field of dialogue systems, and more particularly to response generation in solution-guided dialogue systems.
[0002] A conversational system is an intelligent machine capable of understanding language and engaging in written or spoken conversation with a user. Two common approaches to creating conversational systems are: subject matter experts (“SMEs”) manually creating conversational flows using domain knowledge and data-driven modeling. Data-driven modeling involves learning from chat logs, where problem-solving is implicitly learned from both chat logs and external knowledge, providing a further foundation for generating responses.
[0003] SME-based modeling is time-consuming, expensive, and manual. Furthermore, learning from chat logs is difficult because the model needs to learn both language and business logic. In either case, identifying and representing the necessary external information is challenging. Therefore, a solution for data-driven dialogue systems that is less costly, less labor-intensive, and easier to model is needed. Summary of the Invention
[0004] Embodiments of this disclosure include methods, computer program products, and systems for generating response-guided solutions for dialogue systems.
[0005] In some embodiments, the processor may receive first speech data associated with a first user utterance in a conversation within a guided dialogue system. The processor may identify a first topic from a set of topics associated with the first user utterance from the first speech data. The processor may identify a first solution associated with the first topic, the first solution having one or more solution segments for performing a topic-related task. The processor may generate a first response from a second user based on the first solution segment of the first solution and the first speech data.
[0006] In some embodiments, the processor may receive first speech data associated with a first user utterance in a conversation within a guided dialogue system. The processor may identify a first topic from a set of topics associated with the first user utterance from the first speech data. The processor may identify a first solution associated with the first topic, the first solution having one or more solution segments for completing a topic-related task. The processor may generate the first solution using a corpus of documents from entities associated with the first topic. The processor may generate a first response from a second user based on the first solution segments of the first solution and the first speech data.
[0007] In some embodiments, the processor may receive first speech data associated with a first user utterance in a conversation in a guided dialogue system. The processor may identify a first topic from a set of topics associated with the first user utterance from the first speech data. The processor may identify a first solution associated with the first topic, the first solution having one or more solution segments for completing a topic-related task. The processor may generate the first solution using a text generation AI model that generates solutions from sample conversations. The processor may generate a first response from a second user based on the first solution segments of the first solution and the first speech data.
[0008] In some embodiments, a sequence-to-sequence machine learning model can be used to generate the first response.
[0009] In some embodiments, a text classification model may be used to identify the first topic.
[0010] In some embodiments, the processor may receive second voice data associated with a second user utterance in a conversation within a guided dialogue system. The processor may confirm that the second user utterance is not associated with another topic in a topic set. The processor may generate a second response from the second user based on the first voice data, the second user's first response, the second voice data, and a second solution segment of the first solution.
[0011] In some embodiments, the processor may receive third voice data associated with a third user utterance in a conversation within a guided dialogue system. The processor may identify a second topic associated with the third user utterance. The processor may identify a second solution associated with the second topic. The processor may generate a third response from the second user based on first voice data, a first response from the second user, second voice data, a second response from the second user, third voice data, and a solution segment of the second solution.
[0012] The above description is not intended to depict every illustrated embodiment or every implementation of this disclosure. Attached Figure Description
[0013] The accompanying drawings included in this disclosure are incorporated in and form a part of this specification. They illustrate embodiments of the present disclosure and, together with the description, serve to explain the principles of the disclosure. The drawings are merely illustrative of certain embodiments and are not intended to limit the scope of the disclosure.
[0014] Figure 1 This is a block diagram of an exemplary system generated from responses for solution guidance based on various aspects of this disclosure.
[0015] Figure 2 This is a flowchart of an exemplary method for generating a response for solution guidance based on various aspects of this disclosure.
[0016] Figure 3A A cloud computing environment based on various aspects of this disclosure is shown.
[0017] Figure 3B An abstract model layer is shown according to various aspects of this disclosure.
[0018] Figure 4 A high-level block diagram of an example computer system is shown that can be used to implement one or more of the methods, tools, and modules described herein, and any related functions, according to various aspects of this disclosure.
[0019] While the embodiments described herein may have various modifications and alternatives, their details have been shown by way of example in the accompanying drawings and will be described in detail. However, it should be understood that the specific embodiments described should not be construed as limiting. Rather, the invention will cover all modifications, equivalents, and alternatives falling within the spirit and scope of this disclosure. Detailed Implementation
[0020] The aspects of this disclosure generally relate to the field of dialogue systems, and more particularly to solution-guided response generation in dialogue systems. While this disclosure is not necessarily limited to such applications, its various aspects can be understood through discussion of various examples within this context.
[0021] In some embodiments, the processor may receive first speech data associated with a first user utterance in a conversation within a guided dialogue system. In some embodiments, the processor may identify a first topic from a set of topics associated with the first user utterance from the first speech data. In some embodiments, a text classification model may be used to identify the first topic.
[0022] In some embodiments, user speech may relate to requests, questions, concerns, or problems addressed, addressed, resolved, or completed by guiding conversation in a dialogue system. In some embodiments, the conversation may be communication between a caller (e.g., a first user) and an agent (e.g., a second user), wherein the caller and the agent take turns speaking to or responding to each other. In some embodiments, the first voice data may include a text transcription of at least a portion of the speech spoken by the first user during their turn to speak in the conversation. In some embodiments, the guided dialogue system may be used by a virtual assistant to assist a user in completing various tasks, and may include conversations between the agent and the user to complete those tasks.
[0023] In some embodiments, the text classification model may analyze the text spoken by a first user and assign a set of predefined labels or categories (e.g., a first topic) to the text based on its context. In some embodiments, the text classification model may utilize natural language processing for sentiment analysis, topic detection, intent detection, entity recognition, and language detection. In some embodiments, the identified topics may include, but are not limited to, one or a combination of symptoms, themes, actions, intents, requests, questions, concerns, or difficulties, or any other identifier associated with a task, request, question, concern, or difficulty that the user wishes to assist with in relation to the service or system in which the dialogue system operates.
[0024] For example, a guided dialogue system can be initiated by a caller making a verbal request to extend the due date of an outstanding bill. The caller can make an initial request by stating, "I want to postpone paying my electricity bill." The entire user utterance can be transcribed and fed to an AI model with natural language processing capabilities, which identifies the topic of the user utterance from the speech data. The topic of the user utterance can relate to the subject the user is querying, the question the user wants help with, the problem or concern the caller wants to resolve, the action the user wants to perform, etc. From the user utterance "I want to postpone paying my electricity bill," the topic "payment postponement" can be identified.
[0025] In some embodiments, the processor can identify a first solution associated with a first topic. In some embodiments, the first solution may have one or more solution segments for performing a task related to the topic. In some embodiments, the processor can identify the first solution based on an identifier of the first topic. In some embodiments, one or more solution segments may be a series of steps, actions, or communications through which a task related to the topic can be performed. For example, for the topic "payment deferral," the processor can identify a solution that includes a series of steps through which an agent can perform the task of obtaining a payment deferral from a caller. To obtain a payment deferral from a caller, the solution segment may include the following steps: verifying the caller's phone number; sending a PIN number to the caller and requesting the caller to provide the PIN number to the agent; inquiring about the date he / she wants to settle the payment (e.g., obtaining the deferral period until); reminding the caller that it is important to comply with the new payment arrangement and that failure to comply may result in late fees; recording the payment deferral in a database; notifying the caller that the payment deferral has been granted; and providing the caller with a reference number.
[0026] In some embodiments, one or more solution segments may include subtasks to be performed, including: obtaining various types of information from the user, notifying the caller of various results of the actions performed, verifying relevant background information from the caller (e.g., verifying the user's identity or account information), and obtaining information related to the task to be completed (e.g., when the user wants to extend the due date). In some embodiments, subtasks or solution segments may include: communication exchanges, queries, instructions to be provided, questions to be asked, answers to be provided, information to be obtained, user responses to be received, etc. In some embodiments, the solution and solution segments may be a roadmap or instructions on how to complete a task through communication with the user via a guided dialogue system.
[0027] In some embodiments, the solution may be selected by an artificial intelligence (“AI”) model that links the selected solution to a topic identified from the speech data. In some embodiments, the AI model may be a text classification model that utilizes natural language processing. In some embodiments, the solution selection model may be trained using a dataset that links solutions and topics to text from a conversation between two users (e.g., a caller and an agent).
[0028] In some embodiments, the processor may generate a first response from a second user based on a first solution segment of a first solution and first voice data. For example, based on the first voice data “I want to defer payment of my electricity bill” and the first solution segment of the first solution “Confirm caller’s phone number”, the agent’s (e.g., the second user’s) first response could be: “I can help you, but I need to verify your identity first. Can you provide the phone number associated with this account?”. In some embodiments, the agent (e.g., the second user) is an automated agent that relays the response to the user / caller.
[0029] In some embodiments, a sequence-to-sequence machine learning model may be used to generate the first response. In some embodiments, the sequence-to-sequence model may be a deep learning model that generates text. In some embodiments, the sequence-to-sequence model may be a deep learning model that generates text using a recurrent neural network (RNN), long short-term memory (LSTM), or grating recurrent unit (GRU) architecture. In some embodiments, the context of each item is the output from a previous step. In some embodiments, the main components of the sequence-to-sequence model are encoder and decoder networks. In some embodiments, the encoder transforms each item into a corresponding hidden vector containing the item and its context. In some embodiments, the decoder reverses this process, using the previous output as input context and transforming the vector into an output item. In some embodiments, the sequence-to-sequence model may include BART, Generative Pre-trained Transformer 2 (“GPT2”), Generative Pre-trained Transformer 3 (“GPT3”), etc. In some embodiments, the sequence-to-sequence model may be trained to take both the conversation context (e.g., user utterances and agent responses) and the identified solution (or solution segment) as input to produce a generated response for a second user.
[0030] In some embodiments, the processor may receive second voice data associated with a second user utterance in a conversation within a guided dialogue system. In some embodiments, the processor may confirm that the second user utterance is not associated with another topic in a set of topics. In some embodiments, the processor may generate a second response from the second user based on first voice data, a first response from the second user, the second voice data, and a second solution segment of the first solution.
[0031] Continuing the previous example, the second voice data could be "My phone number is 123 345 6443". The processor can analyze the text provided by the user to determine if this information is relevant to the second topic. The processor can then generate a second response based on the conversation so far and the second solution segment of the first solution. The processor can then feed the first and second voice data, along with the agent's first response, into a machine learning model:
[0032] First user: "I want to postpone paying my electricity bill."
[0033] Agent: "I can help you, but I need to verify your identity first. Could you provide the phone number associated with this account?"
[0034] First user: "My phone number is 123 345 6443."
[0035] The processor can also input a second solution segment of the first solution into the machine learning model, which reads, "Send the PIN number to the caller and request the caller to provide the PIN number to the agent," to generate a second response: "Our system will send the PIN number to your phone via text message. Please provide me with the four digits sent to you."
[0036] In some embodiments, the processor may receive third voice data associated with a third user utterance in a conversation within a guided dialogue system. In some embodiments, the processor may identify a second topic associated with the third user utterance. In some embodiments, the processor may identify a second solution associated with the second topic. In some embodiments, the processor may generate a third response from a second user based on first voice data, a first response from a second user, second voice data, a second response from a second user, third voice data, and a solution segment of the second solution.
[0037] Continuing the previous example, the third user utterance could be, "I received PIN number I is 3476, but I actually want to update the address associated with this account." The processor can identify the second topic "update account information" within the third user utterance. The processor can identify a second solution that provides a series of steps to update the account information. The processor can then generate a third response for the agent based on the entire conversation history (e.g., the first, second, and third user utterances and the first and second responses of the second user / agent) and the solution segment for the solution to update the account information. The solution segment for the solution to update the account information could be, "Verify the type of account information to be updated." The third response generated for the agent could be, "You want to update the address on your account profile, is that right?". The third response can be generated based on the conversation history and the solution segment associated with updating account information.
[0038] In some embodiments, each of the first, second, and third responses can be generated predictively based on patterns detected by a machine learning model.
[0039] In some implementations, the solution (e.g., a first solution or a second solution) may be identified and prepared by a subject matter expert. In some embodiments, one or more steps (e.g., solution segments) included in the solution may be manually created, generated, exported, or prepared by the subject matter expert. In some embodiments, the subject matter expert is able to identify a series of solution segments required to complete a subject-related task based on the subject matter expert's knowledge of the processes, steps, or actions taken to complete the task.
[0040] In some embodiments, the subject matter expert can examine sample exchanges between the user and the agent and identify the steps or processes required to complete a subject-related task based on the communication between the user and the agent. In some embodiments, the subject matter expert can review any other relevant reference materials (e.g., manuals, how-to websites, process announcements) that provide information on how to complete the task (e.g., information that helps the subject matter expert determine, identify, or break down what steps need to be taken). Regardless of the subject matter expert embodiment used, the information provided by the subject matter expert is annotated and presented to the disclosed system for storage and subsequent use during conversations. Furthermore, the information provided by the subject matter expert can be analyzed by a natural language processing system and stored / tagged as a topic / subtopic.
[0041] In some embodiments, a text generation AI model that generates solutions from sample conversations can be used to generate solutions (e.g., a first solution or a second solution). In some embodiments, the text generation AI model can generate a desired solution token based on a token of available conversation context. In some embodiments, the text generation AI model may include BART, GPT2, and GPT3.
[0042] In some embodiments, a text generation model can be trained using transcripts of a conversation between a caller and an agent that have already been annotated by subject matter experts. In some embodiments, sample transcripts can be annotated to identify parts of the conversation that are relevant to solution components. In some embodiments, the text generation AI model can be trained to associate parts of the conversation (e.g., the language spoken by the caller or agent) with solution components.
[0043] In some embodiments, once the text generation model is trained, it can be used to generate solutions by applying the text generation model to a corpus of conversations. In some embodiments, based on sample conversations provided as input, the text generation model is able to output additional solutions (e.g., solutions consisting of a series of solution segments) from sample conversations in the conversation corpus.
[0044] In some embodiments, a corpus of documents from entities related to a topic (e.g., a first topic or a second topic, respectively) may be used to generate a solution (e.g., a first solution or a second solution). In some embodiments, the entities related to the topic may be individuals, groups of individuals, organizations, databases, libraries, etc., which have or can provide access to information related to the topic, the solutions associated with the topic, and / or one or more solution segments for performing tasks related to the topic.
[0045] In some embodiments, a rule-based approach can be used to generate solutions using a document corpus. In some embodiments, the rule-based approach can identify portions of documents in the document corpus that are relevant to a topic, solution, or solution segment. In some embodiments, the rules can describe how to identify topics, solutions, or solution segments and the associated conversational text (e.g., from user speech).
[0046] For example, a solution can be generated using rules associated with Document Object Model (“DOM”) elements on a specified set of web pages (e.g., web pages of an organization that provide step-by-step instructions to resolve common problems encountered with products purchased from the organization). In some implementations, the solutions generated using rules (e.g., associated with DOM elements on the web pages) may be reviewed by a subject matter expert (e.g., the author of the web pages). In some embodiments, feedback from the subject matter expert regarding the solution may be used to update the rules upon which the solution is generated.
[0047] In some embodiments, a subject matter expert may review certain documents from a document corpus (e.g., a user manual for a company product) and generate or draft solutions from paragraphs. For example, the subject matter expert may provide annotations identifying which parts of a paragraph lead to which sub-part of a solution. In some embodiments, the annotations may be used to train a text generation model to generate other solutions from similar documents. In some embodiments, the text generation model may learn how to link paragraphs from the user manual to solution paragraphs. In some embodiments, the text generation model can then generate new solutions from the user manual it receives as input.
[0048] Now for reference Figure 1 A block diagram of a system 100 for response generation in solution guidance is shown. System 100 includes user device 102 and system device 104. System device 104 includes a conversation context database 106, a solution selector 108, a solution 110, a response generator 112, and a response provider 114. User device 102 and system device 104 are configured to communicate with each other. User device 102 and system device 104 can be any device containing a processor configured to perform one or more of the functions or steps described in this disclosure.
[0049] In some embodiments, system device 104 receives first voice data associated with a first user utterance in a conversation from user device 102. The first voice data is stored in a conversation context database 106. A solution selector 108 of system device 104 identifies a first topic from a set of topics associated with the first user utterance from the first voice data, and identifies a first solution 110 associated with the first topic. The first solution 110 has one or more solution segments for performing a task related to the topic. A response generator 112 of system device 104 generates a first response for a second user (e.g., an agent in a guided dialogue system) based on the first solution segments of the first solution and the first voice data. In some embodiments, the response generator generates the response using a sequence-to-sequence machine learning model. The first response is transmitted to the first user (e.g., to user device 102) via a response provider 114.
[0050] In some embodiments, a first response is stored in a conversation context database 106 and used to generate a second response for the agent. In some embodiments, system device 104 receives second voice data associated with a second user utterance in the conversation. In some embodiments, solution selector 108 confirms that the second user utterance is not associated with another topic in the topic set. In some embodiments, the solution selector may use text classification model 116 to confirm that the second user utterance is not associated with another topic in the topic set. In some embodiments, response generator 112 generates a second response for the agent based on the first voice data, the agent's first response, the second voice data, and a second solution segment of the first solution.
[0051] In some embodiments, first and second voice data, as well as first and second responses, are stored in a conversation context database 106 and used to generate a third response for the agent. In some embodiments, system device 104 receives third voice data associated with a third user utterance in a conversation. In some embodiments, solution selector 108 identifies a second topic associated with the third user utterance. In some embodiments, solution selector 108 identifies a second solution associated with the second topic. In some embodiments, response generator 112 generates a third response for the agent based on the first voice data, the agent's first response, the second voice data, the agent's second response, the third voice data, and the solution segment of the second solution.
[0052] Now for reference Figure 2A flowchart of an exemplary method 200 for generating a response for solution guidance according to embodiments of the present disclosure is shown. In some embodiments, a processor of the system may perform the operations of method 200. In some embodiments, method 200 begins at operation 202. At operation 202, the processor receives first voice data associated with a first user utterance in a conversation in a guided dialogue system. In some embodiments, method 200 proceeds to operation 204, wherein the processor identifies a first topic from a set of topics associated with the first user utterance from the first voice data.
[0053] In some embodiments, method 200 proceeds to operation 206. In operation 206, the processor identifies a first solution associated with a first topic, the first solution having one or more solution segments for performing a task related to the topic. In some embodiments, method 200 proceeds to operation 208. In operation 208, the processor generates a first response for a second user based on the first solution segment of the first solution and first voice data.
[0054] As discussed in more detail herein, it is conceivable that some or all of the operations of method 200 may be performed in an alternative order or may not be performed at all; furthermore, multiple operations may occur simultaneously or as part of a larger process.
[0055] It should be understood that although this disclosure includes a detailed description of cloud computing, the implementation of the teachings set forth herein is not limited to a cloud computing environment. Rather, embodiments of this disclosure can be implemented in conjunction with any other type of computing environment now known or developed hereafter.
[0056] Cloud computing is a service delivery model for enabling convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be rapidly provisioned and released with minimal management effort or interaction with service providers. This cloud model may include at least five features, at least three service models, and at least four deployment models.
[0057] The characteristics are as follows:
[0058] On-demand self-service: Cloud consumers can unilaterally and automatically provide computing power, such as server time and network storage, as needed, without requiring manual interaction with the service provider.
[0059] Wide Area Network (WAN) Access: Capabilities are available on the network and accessed through standard mechanisms that facilitate the use of heterogeneous thin or thick client platforms (e.g., mobile phones, laptops, and PDAs).
[0060] Resource pooling: A provider's computing resources are pooled to serve multiple consumers using a multi-tenant model, where different physical and virtual resources are dynamically allocated and reallocated based on demand. Partial independence exists because consumers typically cannot control or know the exact portion of the resources provided, but may be able to specify portions at a higher level of abstraction (e.g., country, state, or data center).
[0061] Rapid Flexibility: In some cases, the ability to scale outwards and inwards quickly and flexibly can be provided. For consumers, the available capacity often appears unlimited and can be purchased in any quantity at any time.
[0062] Measurement services: Cloud systems automatically control and optimize resource usage by leveraging metering capabilities at a level of abstraction appropriate to the service type (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, and reported, providing transparency to both the providers and consumers of the services being utilized.
[0063] The service model is as follows:
[0064] Software as a Service (SaaS): The capability offered to consumers is the ability to use the provider's applications running on cloud infrastructure. Applications can be accessed from various client devices through thin client interfaces such as web browsers (e.g., web-based email). Consumers do not manage or control the underlying cloud infrastructure, including the network, servers, operating system, storage, or even individual application capabilities, with possible exceptions such as limited user-specific application configuration settings.
[0065] Platform as a Service (PaaS): This provides consumers with the ability to deploy consumer-created or acquired applications onto cloud infrastructure using programming languages and tools supported by the provider. Consumers do not manage or control the underlying cloud infrastructure, including networks, servers, operating systems, or storage, but they have control over the deployed applications and the configuration of any application hosting environments.
[0066] Infrastructure as a Service (IaaS): This provides consumers with the capability to deliver processing, storage, networking, and other basic computing resources that enable them to deploy and run arbitrary software, which may include operating systems and applications. Consumers do not manage or control the underlying cloud infrastructure, but they do have control over the operating system, storage, deployed applications, and possibly limited control over selected networking components (e.g., host firewalls).
[0067] The deployment model is as follows:
[0068] Private cloud: Cloud infrastructure operated solely by an organization. It can be managed by the organization or a third party and can exist on-site or off-site.
[0069] Community cloud: Cloud infrastructure shared by several organizations and supporting a specific community with shared concerns (e.g., tasks, security requirements, policies, and compliance considerations). It can be managed by an organization or a third party and can exist on-site or off-site.
[0070] Public cloud: Cloud infrastructure available to the general public or large industrial groups and owned by organizations that sell cloud services.
[0071] Hybrid cloud: A cloud infrastructure is a combination of two or more clouds (private, community, or public) that remain a single entity but are bound together by standardized or proprietary technologies that enable data and applications to be ported together (e.g., cloud bursting for load balancing between clouds).
[0072] Cloud computing environments are service-oriented, focusing on statelessness, loose coupling, modularity, and semantic interoperability. At the heart of cloud computing is the infrastructure of a network of interconnected nodes.
[0073] Figure 3A A cloud computing environment 310 is illustrated. As shown, the cloud computing environment 310 includes one or more cloud computing nodes 300 to which local computing devices used by cloud consumers can communicate, such as personal digital assistants (PDAs) or cellular phones 300A, desktop computers 300B, laptop computers 300C, and / or automotive computer systems 300N. The nodes 300 can communicate with each other. They can be physically or virtually grouped (not shown) in one or more networks, such as private clouds, community clouds, public clouds, or hybrid clouds, or combinations thereof, as described above.
[0074] This allows cloud computing environments 310 to provide infrastructure, platforms, and / or software as services, without requiring cloud consumers to maintain resources on their local computing devices. It should be understood that... Figure 3A The types of computing devices 300A-N shown are for illustrative purposes only, and computing node 300 and cloud computing environment 310 can communicate with any type of computing device via any type of network and / or network-addressable connection (e.g., using a web browser).
[0075] Figure 3B This demonstrates the cloud computing environment 310 ( Figure 3A This provides a set of functional abstractions. It should be understood beforehand that... Figure 3B The components, layers, and functions shown are for illustrative purposes only, and embodiments of this disclosure are not limited thereto. The following layers and corresponding functions are provided.
[0076] The hardware and software layer 315 includes hardware and software components. Examples of hardware components include: a host 302; a server 304 based on a RISC (Reduced Instruction Set Computer) architecture; a server 306; a blade server 308; a storage device 311; and a network and networking component 312. In some embodiments, software components include network application server software 314 and database software 316.
[0077] The virtualization layer 320 provides an abstraction layer from which the following examples of virtual entities can be provided: virtual server 322; virtual storage 324; virtual network 326, including virtual private network; virtual application and operating system 328; and virtual client 330.
[0078] In one example, management layer 340 may provide the functionality described below. Resource provisioning 342 provides dynamic procurement of computing resources and other resources used to perform tasks within the cloud computing environment. Metering and pricing 344 provides cost tracking when utilizing resources in the cloud computing environment, as well as billing or invoicing for the consumption of these resources. In one example, these resources may include application software licenses. Security provides authentication for cloud consumers and tasks, as well as protection for data and other resources. User portal 346 provides access to the cloud computing environment for consumers and system administrators. Service level management 348 provides cloud resource allocation and management to ensure that required service levels are met. Service level agreement (SLA) planning and fulfillment 350 provides pre-scheduling and procurement of cloud resources, where future needs are anticipated according to the SLA.
[0079] The workload layer 360 provides examples of functionalities that can be leveraged in a cloud computing environment. Examples of workloads and functionalities that can be provided from this layer include: mapping and navigation 362; software development and lifecycle management 364; virtual classroom education delivery 366; data analytics and processing 368; transaction processing 370; and solution-guided dialogue system response generation 372.
[0080] Figure 4A high-level block diagram of an example computer system 401 according to an embodiment of the present disclosure is shown. This computer system can be used to implement one or more of the methods, tools, and modules described herein, as well as any associated functions (e.g., using one or more processor circuits or a computer processor). In some embodiments, the main components of the computer system 401 may include one or more CPUs 402, a memory subsystem 404, a terminal interface 412, a storage interface 416, an I / O (input / output) device interface 414, and a network interface 418. All these components may be directly or indirectly communicatively coupled to enable inter-component communication via a memory bus 403, an I / O bus 408, and an I / O bus interface unit 410.
[0081] Computer system 401 may include one or more general-purpose programmable central processing units (CPUs) 402A, 402B, 402C, and 402D, collectively referred to herein as CPU 402. In some embodiments, computer system 401 may include a typical multiple processors of a relatively large system; however, in other embodiments, computer system 401 may alternatively be a single-CPU system. Each CPU 402 may execute instructions stored in memory subsystem 404 and may include one or more levels of onboard cache.
[0082] System memory 404 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 422 or cache memory 424. Computer system 401 may also include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 426 may be provided for reading from and writing to non-removable, non-volatile magnetic media such as a "hard disk drive". Although not shown, a disk drive may be provided for reading from and writing to a removable, non-volatile disk (e.g., a "floppy disk"), or an optical disc drive or other optical media may be provided for reading from or writing to a removable, non-volatile optical disc such as a CD-ROM or DVD-ROM. Additionally, memory 404 may include flash memory, such as a flash stick drive or flash drive. Memory devices may be connected to memory bus 403 via one or more data media interfaces. Memory 404 may include at least one program product having a set (e.g., at least one) of program modules configured to perform functions of various embodiments.
[0083] One or more programs / utilities 428 may be stored in memory 404, each program / utility having at least one set of program modules 430. Programs / utilities 428 may include a hypervisor (also known as a virtual machine monitor), one or more operating systems, one or more applications, other program modules, and program data. Each of the operating system, one or more applications, other program modules, and program data, or some combination thereof, may include an implementation of a networking environment. Programs 428 and / or program modules 430 typically perform the functions or methods of various embodiments.
[0084] Although memory bus 403 is Figure 4 While shown as a single bus structure providing a direct communication path between CPU 402, memory subsystem 404, and I / O bus interface 410, in some embodiments, memory bus 403 may include multiple different buses or communication paths, which may be arranged in any of a variety of forms, such as hierarchical point-to-point links, star or mesh configurations, multi-layer buses, parallel and redundant paths, or any other suitable type of configuration. Furthermore, although I / O bus interface 410 and I / O bus 408 are shown as a single corresponding unit, in some embodiments, computer system 401 may include multiple I / O bus interface units 410, multiple I / O buses 408, or both. Additionally, although multiple I / O interface units are shown separating I / O bus 408 from various communication paths to various I / O devices, in other embodiments, some or all I / O devices may be directly connected to one or more system I / O buses.
[0085] In some embodiments, computer system 401 may be a multi-user mainframe computer system, a single-user system, a server computer, or a similar device that has little or no direct user interface but receives requests from other computer systems (clients). Furthermore, in some embodiments, computer system 401 may be implemented as a desktop computer, portable computer, laptop or notebook computer, tablet computer, pocket computer, telephone, smartphone, network switch or router, or any other suitable type of electronic device.
[0086] Notice, Figure 4 The description aims to depict representative main components of an exemplary computer system 401. However, in some embodiments, the various components may have more... Figure 4 The greater or lesser complexity represented therein can exist differently from... Figure 4 The components shown, or other components, and the number, type, and configuration of these components may vary.
[0087] As discussed in more detail herein, it is conceivable that some or all of the operations of some embodiments of the methods described herein may be performed in an alternative order or may not be performed at all; furthermore, multiple operations may occur simultaneously or as internal parts of a larger process.
[0088] This disclosure can be a system, method, and / or computer program product at any possible level of technical detail integration. A computer program product may include a computer-readable storage medium (or media) having computer-readable program instructions thereon for causing a processor to execute aspects of this disclosure.
[0089] Computer-readable storage media can be tangible devices capable of retaining and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example, but not limited to, electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable optical disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices such as punch cards or recessed structures with instructions recorded thereon, and any suitable combination of the foregoing. As used herein, computer-readable storage media should not be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.
[0090] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a suitable computing / processing device, or via a network, such as the Internet, a local area network (LAN), a wide area network (WAN), and / or a wireless network, to an external computer or external storage device. The network may include copper cables, optical fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to a computer-readable storage medium within the respective computing / processing device.
[0091] Computer-readable program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, integrated circuit configuration data, or source code or object code written in any combination of one or more programming languages (including object-oriented programming languages such as Smalltalk, C++, etc.) and procedural programming languages (such as the "C" programming language or similar programming languages). The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs) may execute computer-readable program instructions to personalize the electronic circuitry in order to perform aspects of this disclosure by utilizing the status information of the computer-readable program instructions.
[0092] This document describes aspects of the disclosure with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the disclosure. It will be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0093] These computer-readable program instructions may be provided to a processor of a computer or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / actions specified in one or more blocks of a flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, programmable data processing apparatus, and / or other devices to operate in a particular manner, such that the computer-readable storage medium in which the instructions are stored includes an article of writing comprising instructions for implementing aspects of the functions / actions specified in one or more blocks of a flowchart and / or block diagram.
[0094] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer-implemented process, such that the instructions, which execute on the computer, other programmable apparatus or other device, perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0095] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of instructions comprising one or more executable instructions for implementing a specified logical function. In some alternative embodiments, the functions indicated in the blocks may occur in a non-linear order as shown in the figures. For example, two blocks shown consecutively may actually be implemented as a single step, executed simultaneously, substantially simultaneously, with partial or complete time overlap, or these blocks may sometimes be executed in reverse order, depending on the functions involved. It will also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified function or action or executes a combination of dedicated hardware and computer instructions.
[0096] Various embodiments of this disclosure have been described for illustrative purposes, but are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein has been chosen to best explain the principles of the embodiments, their practical application, or improvements to existing technologies in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
[0097] Although this disclosure has been described with reference to specific embodiments, changes and modifications thereto will be apparent to those skilled in the art. Therefore, the appended claims are intended to be construed as covering all such alterations and modifications that fall within the true spirit and scope of this disclosure.
Claims
1. A computer-implemented method for generating responses in a dialogue system, the method comprising: The processor receives first voice data associated with first user utterances in a conversation of a second user in a guided dialogue system, wherein the second user is an automated agent of the guided dialogue system; Identify a first topic from the set of topics associated with the first user's speech from the first speech data; Identify a first solution associated with the first topic, the first solution having a series of solution segments for completing a task related to the topic; Based on a first solution segment from a series of solution segments of the first solution and the first speech data, a first response for a second user is generated, wherein the first response is generated using a sequence-to-sequence machine learning model, the machine learning model including generating a pre-trained transformer trained to take the first speech data and the first solution segment as input and output the first response; and The first response is relayed to the first user via the second user.
2. The method of claim 1, wherein, The first topic was identified using a text classification model.
3. The method according to claim 1, further comprising: Receive second voice data associated with a second user utterance in the conversation in the guided dialogue system; Confirm that the second user utterance is not associated with another topic in the topic set; as well as Based on the first voice data, the second user's first response, the second voice data, and the second solution segment from a series of solution segments of the first solution, the second user's second response is generated.
4. The method according to claim 3, further comprising: Receive third voice data associated with a third user utterance in the conversation in the guided dialogue system; Identify a second topic associated with the third user's utterance; Identify a second solution associated with the second topic; and Based on the first voice data, the second user's first response, the second voice data, the second user's second response, the third voice data, and the solution segment of the second solution, the second user's third response is generated.
5. The method of claim 1, wherein the first solution is generated using a text generation artificial intelligence model that generates solutions from sample conversations.
6. The method of claim 5, wherein, The first topic was identified using a text classification model.
7. The method according to claim 5, further comprising: Receive second voice data associated with a second user utterance in the conversation in the guided dialogue system; Confirm that the second user utterance is not associated with another topic in the topic set; as well as Based on the first voice data, the second user's first response, the second voice data, and the second solution segment from a series of solution segments of the first solution, the second user's second response is generated.
8. The method according to claim 7, further comprising: Receive third voice data associated with a third user utterance in the conversation in the guided dialogue system; Identify a second topic associated with the third user's utterance; Identify a second solution associated with the second topic; and Based on the first voice data, the second user's first response, the second voice data, the second user's second response, the third voice data, and the solution segment of the second solution, the second user's third response is generated.
9. The method of claim 1, wherein the first solution is generated using a corpus of documents from entities related to the first topic.
10. The method of claim 9, wherein, The first topic was identified using a text classification model.
11. The method of claim 9, further comprising: Receive second voice data associated with a second user utterance in the conversation in the guided dialogue system; Confirm that the second user utterance is not associated with another topic in the topic set; as well as Based on the first voice data, the second user's first response, the second voice data, and the second solution segment from a series of solution segments of the first solution, the second user's second response is generated.
12. The method of claim 11, further comprising: Receive third voice data associated with a third user utterance in the conversation in the guided dialogue system; Identify a second topic associated with the third user's utterance; Identify a second solution associated with the second topic; and Based on the first voice data, the second user's first response, the second voice data, the second user's second response, the third voice data, and the solution segment of the second solution, the second user's third response is generated.
13. A system for generating responses in a dialogue system, comprising: Memory; as well as A processor communicating with the memory, the processor being configured to perform operations including the following: Receive first voice data associated with first user utterances in a conversation of a second user in a guided dialogue system, wherein the second user is an automated agent of the guided dialogue system; Identify a first topic from the set of topics associated with the first user's speech from the first speech data; Identify a first solution associated with the first topic, the first solution having a series of solution segments for completing a task related to the topic; Based on a first solution segment from a series of solution segments of the first solution and the first speech data, a first response for a second user is generated, wherein the first response is generated using a sequence-to-sequence machine learning model, the machine learning model including generating a pre-trained transformer trained to take the first speech data and the first solution segment as input and output the first response; and The first response is relayed to the first user via the second user.
14. The system of claim 13, wherein, The first topic was identified using a text classification model.
15. The system of claim 13, wherein the processor is further configured to perform operations including: Receive second voice data associated with a second user utterance in the conversation in the guided dialogue system; Confirm that the second user utterance is not associated with another topic in the topic set; and Based on the first voice data, the second user's first response, the second voice data, and the second solution segment from a series of solution segments of the first solution, the second user's second response is generated.
16. The system of claim 15, wherein the processor is further configured to perform operations including: Receive third voice data associated with a third user utterance in the conversation in the guided dialogue system; Identify a second topic associated with the third user's utterance; Identify a second solution associated with the second topic; and Based on the first voice data, the second user's first response, the second voice data, the second user's second response, the third voice data, and the solution segment of the second solution, the second user's third response is generated.
17. A computer program product comprising program instructions executable by a processor to cause the processor to perform operations, the operations including: Receive first voice data associated with first user utterances in a conversation of a second user in a guided dialogue system, wherein the second user is an automated agent of the guided dialogue system; Identify a first topic from the set of topics associated with the first user's speech from the first speech data; Identify a first solution associated with the first topic, the first solution having a series of solution segments for completing a task related to the topic; Based on a first solution segment from a series of solution segments of the first solution and the first speech data, a first response for a second user is generated, wherein the first response is generated using a sequence-to-sequence machine learning model, the machine learning model including generating a pre-trained transformer trained to take the first speech data and the first solution segment as input and output the first response; and The first response is relayed to the first user via the second user.
18. The computer program product of claim 17, wherein, The first topic was identified using a text classification model.
19. The computer program product of claim 17, wherein the processor is further configured to perform operations including: Receive second voice data associated with a second user utterance in the conversation in the guided dialogue system; Confirm that the second user utterance is not associated with another topic in the topic set; and Based on the first voice data, the second user's first response, the second voice data, and the second solution segment from a series of solution segments of the first solution, the second user's second response is generated.
20. The computer program product of claim 19, wherein the processor is further configured to perform operations including: Receive third voice data associated with a third user utterance in the conversation in the guided dialogue system; Identify a second topic associated with the third user's utterance; Identify a second solution associated with the second topic; and Based on the first voice data, the second user's first response, the second voice data, the second user's second response, the third voice data, and the solution segment of the second solution, the second user's third response is generated.