Computer implementation methods, systems, and computer programs (solutions for interactive systems and guided response generation)

The processor-based solution for dialogue systems uses text classification and machine learning to identify topics and generate responses, addressing the inefficiencies of SME-based modeling and data-driven learning, enhancing response generation accuracy and efficiency.

JP7868934B2Active Publication Date: 2026-06-02INTERNATIONAL BUSINESS MACHINE CORPORATION

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
INTERNATIONAL BUSINESS MACHINE CORPORATION
Filing Date
2022-06-29
Publication Date
2026-06-02

Smart Images

  • Figure 0007868934000001
    Figure 0007868934000001
  • Figure 0007868934000002
    Figure 0007868934000002
  • Figure 0007868934000003
    Figure 0007868934000003
Patent Text Reader

Abstract

To provide solutions for data driven dialog systems that are less costly, less labor-intensive, and easier to model.SOLUTION: A processor may receive first voice data associated with a first user utterance in conversation in a guided dialog system. The processor may identify from the first voice data a first topic of a set of topics associated with the first user utterance. The processor may identify a first solution associated with the first topic, the first solution having one or more solution segments for accomplishing a task related to the topic. The processor may generate a first response for a second user based on a first solution segment of the first solution and the first voice data.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to the field of dialogue systems, and more particularly to Solution elicitation Response of the expression generation for dialogue systems.

Background Art

[0002] A dialogue system is an intelligent machine that can understand language and communicate with a user in writing or orally. Two common ways to create a dialogue system are for a content area expert (an "SME") to This involves manually creating a dialog flow using [a specific method / tool]. - data-driven modeling Toga There is. Data-driven modeling includes learning from chat logs where problem solving is implicitly learned from both chat logs and external knowledge, thereby providing more basis for generating responses.

Summary of the Invention

Problems to be Solved by the Invention

[0003] SME-based modeling requires a great deal of time, cost, and manpower. Furthermore, since the model needs to learn not only business logic but also language, learning from chat logs is difficult. In either case, it is difficult to identify and represent the necessary external information. Therefore, there is a need for a solution for a data-driven dialogue system that is cost-effective, does not require a large labor force, and is easier to model. resolution is needed.

Means for Solving the Problems

[0004] Embodiments of the present disclosure include Solution elicitation Response of the expression methods, computer program products, and systems for generation for dialogue systems.

[0005] In some embodiments, the processor may receive first speech data associated with a first user utterance in a conversation in an assisted dialogue system. From the first speech data, the processor may identify a first topic of a set of topics associated with the first user utterance. The processor may identify a first topic associated with the first topic Solution It is possible to identify the first Solution This is one or more tasks related to the topic. Solution It has a segment. The processor is the first Solution The first Solution Based on the segment and the first voice data, a first response for a second user can be generated.

[0006] In some embodiments, the processor may receive first speech data associated with a first user utterance in a conversation in an assisted dialogue system. From the first speech data, the processor may identify a first topic of a set of topics associated with the first user utterance. The processor may identify a first topic associated with the first topic Solution It is possible to identify the first Solution This is one or more tasks related to the topic. Solution It has segments. The processor is the entity associated with the first topic. of Using a document corpus, the first Solution The processor can generate the first Solution The first Solution Based on the segment and the first voice data, a first response for a second user can be generated.

[0007] In some embodiments, the processor may receive first speech data associated with a first user utterance in a conversation in an assisted dialogue system. From the first speech data, the processor may identify a first topic of a set of topics associated with the first user utterance. The processor may identify a first topic associated with the first topic Solution It is possible to identify the first SolutionThis is one or more tasks related to the topic. Solution It has segments. The processor uses sample conversations to process them. Solution Using an artificial intelligence model that generates text, the first Solution The processor can generate the first Solution The first Solution Based on the segment and the first voice data, a first response for a second user can be generated.

[0008] In some embodiments, the first response may be generated using a sequence-to-sequence machine learning model.

[0009] In some embodiments, the first topic may be identified using a text classification model.

[0010] In some embodiments, the processor may receive second speech data associated with a second user utterance in a conversation in a guided dialogue system. The processor may verify that the second user utterance is not associated with another topic in a set of topics. The processor may process the first speech data, the first response of the second user, the second speech data, and the first Solution The second Solution Based on the segment, a second response can be generated for a second user.

[0011] In some embodiments, the processor may receive third speech data associated with a third user utterance in a conversation in a guided dialogue system. The processor may identify a second topic associated with the third user utterance. The processor may identify a second topic associated with the second topic. Solution The processor can identify the first voice data, the first response of the second user, the second voice data, the second response of the second user, the third voice data, and the second Solution of Solution Based on the segment, a third response for a second user can be generated.

[0012] The above summary is not intended to describe every illustrated embodiment or every implementation of the present disclosure.

Brief Description of the Drawings

[0013] The drawings included in the present disclosure are incorporated herein and form a part of this specification. They illustrate embodiments of the present disclosure and, together with the description, serve to explain the principles of the present disclosure. The drawings merely illustrate specific embodiments and do not limit the present disclosure.

[0014] [Figure 1] FIG. 1 is a block diagram of an exemplary system for solution-guided response generation according to an aspect of the present disclosure.

[0015] [Figure 2] FIG. 2 is a flowchart of an exemplary method for solution-guided response generation according to an aspect of the present disclosure.

[0016] [Figure 3A] FIG. 3 shows a cloud computing environment according to an aspect of the present disclosure.

[0017] [Figure 3B] FIG. 4 shows an abstraction model layer according to an aspect of the present disclosure.

[0018] [Figure 4] FIG. 5 shows a high-level block diagram of an exemplary computer system that can be used in implementing one or more of the methods, tools, and modules described herein, and any related functionality, according to an aspect of the present disclosure.

[0019] The embodiments described herein are subject to various modifications and alternative forms, specific examples of which are shown in the drawings and described in detail. However, it should be understood that the specific embodiments described herein should not be construed as restrictive. Rather, the intention is to cover all modifications, equivalents, and alternatives that fall within the spirit and scope of this disclosure. [Modes for carrying out the invention]

[0020] The aspects of this disclosure generally relate to the field of dialogue systems, and more specifically to dialogue systems. Solution Guidance Response of the expression Regarding generation. While this disclosure is not necessarily limited to such uses, various aspects of this disclosure can be understood through discussion of various examples of its use in this context.

[0021] In some embodiments, the processor may receive first speech data associated with a first user utterance in a conversation in a guided dialogue system. In some embodiments, the processor may identify a first topic from the first speech data, which is part of a set of topics associated with the first user utterance. In some embodiments, the first topic may be identified using a text classification model.

[0022] In some embodiments, user utterances are handled through conversation in a guided dialogue system. ResolvedThe conversation may relate to a request, issue, concern, or problem to resolve or perform. In some embodiments, the conversation may be communication between a speaker (e.g., a first user) and an agent (e.g., a second user), where the speaker and agent take turns speaking and responding to each other. In some embodiments, the first audio data may include a text transcript of at least a portion of the conversation spoken by the first user during the first user's turn to speak in the conversation. In some embodiments, the guided dialogue system may be utilized by a virtual assistant to help the user perform a wide variety of tasks, and may include conversations between the agent and the user to perform those tasks.

[0023] In some embodiments, a text classification model may analyze text spoken by a first user and, based on the context of the text, assign a set of predefined tags or categories (e.g., a first topic) to the text. In some embodiments, the text classification model may utilize natural language processing for sentiment analysis, topic detection, intent detection, entity identification, and language detection. In some embodiments, the identified topics may be, but are not limited to, signs, topics, actions, intentions, requests, issues, concerns, or one or a combination of issues, or the dialogue system may be identified. process This may include any other identifiers associated with a task, request, issue, concern, or problem related to the service or system for which the user wishes to receive assistance.

[0024] For example, a guided dialogue system may be initiated by a speaker who verbally requests an extension of the due date for paying a utility bill. The speaker would say, "I would like an extension on the payment of my electricity bill." Guided dialogueThe system can make an initial request. The entire user utterance can be transcribed and provided to an artificial intelligence model with natural language processing capabilities that can identify the topic of the user utterance from the audio data. The topic of the user utterance may relate to the topic the user is asking about, the problem the user wants help with, the issue or concern the speaker wants to address, or the action the user wants to take. From the user utterance, "I would like an extension on my electricity bill payment," the topic "payment extension" can be identified.

[0025] In some embodiments, the processor is associated with a first topic Solution It is possible to identify the first Solution This is one or more tasks related to the topic. Solution It may have segments. In some embodiments, the processor identifies a first topic based on the identification of the first topic Solution It is possible to identify one or more. In some embodiments, one or more Solution A segment can be a set of steps, actions, or communications that can perform a task related to a topic. For example, in the case of the topic "payment extension," the processor would have the agent communicate with the speaker. for This includes a series of steps that may be taken to perform the task of obtaining a payment extension. Solution It is possible to identify this. In order to obtain an extension of payment for the speaker, Solution The segment includes checking the speaker's phone number. Step Then, it sends a PIN number to the speaker and requests the speaker to provide the PIN number to the agent. Step Then, ask the speaker what date they would like the payment to be settled. Step (For example, obtain the extension period) and inform the speaker that it is important to comply with the new payment agreement, and that failure to do so may result in late fees. Step The payment extension is then recorded in the database. Step This notifies the speaker that the payment extension has been implemented. Step This provides the speaker with a reference number. Step This may include.

[0026] In some embodiments, one or more Solution The segment may include subtasks that need to be performed, which include obtaining various types of information from the user, notifying the speaker of various outcomes when an action is performed, confirming relevant background information from the speaker (e.g., confirming the user's ID or account information), and obtaining information related to the task that needs to be performed (e.g., until when the user wants to extend the deadline). In some embodiments, subtasks or Solution Segments may include communication exchanges, inquiries, instructions provided, questions asked, answers provided, information obtained, and user responses received. In some embodiments, Solution and Solution A segment can be a roadmap or instructions on how to perform a task through communication and exchange with the user via a guided dialogue system.

[0027] In some embodiments, Solution is selected Solution The audio data may be selected by an artificial intelligence ("AI") model that associates it with a topic identified from the audio data. In some embodiments, the AI ​​model may be a text classification model that utilizes natural language processing. In some embodiments, Solution The selected model is, Solution The system can be trained using datasets that associate topics with text from conversations between two users (e.g., a speaker and an agent).

[0028] In some embodiments, the processor is a first Solution The first Solution Based on the segment and the first voice data, a first response for a second user may be generated. For example, the first voice data "I would like to extend the payment period for my electricity bill" and the first Solution The first SolutionBased on the segment "Verify speaker's phone number," the first response for the agent (e.g., a second user) could be: "We can extend your contract, but we first need to verify your identity. Could you please provide the phone number associated with your account?" In some embodiments, the agent (e.g., a second user) is an automated agent that relays the response to the user / speaker.

[0029] In some embodiments, the first response may be generated using a sequence-to-sequence machine learning model. In some embodiments, the sequence-to-sequence model may be a deep learning model that generates text. In some embodiments, the sequence-to-sequence model may be a deep learning model that generates text by using a recurrent neural network (RNN), long short-term memory (LSTM), or gated recurrent unit (GRU) architecture. In some embodiments, the context of each item is the output from the previous step. In some embodiments, the main components of the sequence-to-sequence model are encoder and decoder networks. In some embodiments, the encoder transforms each item into a corresponding hidden vector containing the item and its context. In some embodiments, the decoder reverses the process by transforming the vector into an output item, using the previous output as the input context. In some embodiments, the sequence-to-sequence model may include BART, generative pre-trained transformer 2 ("GPT2"), generative pre-trained transformer 3 ("GPT3"), etc. In some embodiments, the sequence-to-sequence model includes the conversational context (e.g., user utterances and responses by the agent) and identified Solution (or Solution Both segments are inputs. 、 Second user Therefore to raw Growth done response Output Bringing It is good to be trained to do so. .

[0030] In some embodiments, the processor may receive second speech data associated with a second user utterance in a conversation in a guided dialogue system. In some embodiments, the processor may verify that the second user utterance is not associated with another topic in a set of topics. In some embodiments, the processor receives first speech data, a first response from the second user, second speech data, and first Solution The second Solution Based on the segment, a second response can be generated for a second user.

[0031] Continuing from the previous example, the second voice data could be "My phone number is 123 345 6443." The processor, The To determine that the information is not related to the second topic, the information provided by the user The The text can be analyzed. The processor then analyzes the previous conversation and the first Solution The second Solution A second response can be generated based on the segment. The processor then provides the machine learning model with The Audio data 1 and 2 and Agent A first response may be entered.

[0032] First user: "I would like to request an extension on the payment period for my electricity bill."

[0033] Agent: "We can extend your contract, but we first need to verify your ID. Could you please tell us the phone number associated with your account?"

[0034] First user: "My phone number is 123 345 6443."

[0035] The processor also provides the machine learning model with a first Solution The second Solution Segment: "Send a PIN number to the speaker and request the speaker to provide the PIN number to the agent." too input As a result,The second response was, "Our system will send a PIN number to your mobile phone via text message. Please let us know the 4-digit number you received." but Generate So obtain.

[0036] In some embodiments, the processor may receive third speech data associated with a third user utterance in a conversation in an assisted dialogue system. In some embodiments, the processor may identify a second topic associated with the third user utterance. In some embodiments, the processor may identify a second topic associated with the second topic. Solution It is possible to identify the first voice data, the first response of the second user, the second voice data, the second response of the second user, the third voice data, and the second Solution of Solution Based on the segment, a third response for a second user can be generated.

[0037] Continuing from the previous example, the third user utterance could be, "The PIN number you provided is 3476. We will update the address associated with this account." The processor may identify the second topic, "updating account information," in the third user utterance. The processor may then provide a set of steps for updating the account information. Solution The processor can then identify the entire conversation history (e.g., the first, second, and third user utterances, and the first and second responses of the second user / agent) and update the account information. Solution of Solution Based on the segment, a third response for the agent can be generated. Account information of Update Solution of Solution The segment could be defined as "Confirm the type of account information to update." forThe generated third response could be, "We will update the address in your account profile. Is this the correct account?" The third response is associated with updating the conversation history and account information. Solution of Solution It can be generated based on the segment.

[0038] In some embodiments, each of the first, second, and third responses may be generated predictively based on patterns detected by a machine learning model.

[0039] In some embodiments, Solution (For example, the first Solution or second Solution ) can be identified and prepared by a subject area expert. In some embodiments, Solution One or more steps included (for example, Solution Segments can be manually created, generated, derived, or prepared by a content area expert. In some embodiments, the content area expert will perform the task Take Based on the subject matter expert knowledge regarding the process, steps, or actions to be obtained. tree , to A series of tasks required to perform tasks related to picking Solution Identify the segment Good .

[0040] In some embodiments, a content area expert provides samples between the user and the agent. exchange Based on this, and the communication between the user and the agent, the necessary steps or processes to perform a task related to the topic can be identified. In some embodiments, the subject area expert provides information on how to perform the task (e.g., information that helps determine, identify, or categorize the steps that the subject area expert needs to perform). any You can consult other relevant reference materials (e.g., manuals, instructions on how to access the website, procedure reports). Inside Implementation of a Specialist in a Scoped Field Use Regardless, information provided by subject matter experts is tagged with annotations. Note To remember, and the subsequent During the conversation For use, the information will be presented to the disclosed system. Furthermore, information provided by content domain experts may be analyzed by a natural language processing system and stored / tagged by topic / subtopic.

[0041] In some embodiments, Solution (For example, the first Solution or second Solution ) is from a sample conversation Solution It can be generated using a text generation artificial intelligence model that generates the necessary tokens based on the available conversational context. In some embodiments, the text generation artificial intelligence model generates the necessary tokens based on the available conversational context. Solution Tokens can be generated. In some embodiments, the text generation artificial intelligence model may include BART, GPT2, and GPT3.

[0042] In some embodiments, the text generation model may be trained using a transcript of a conversation between an utterancer and an agent, annotated by a content domain expert. In some embodiments, the sample transcript is: Solutions in conversation Related to components Department Minutes can be annotated to identify them. In some embodiments, a text-generating artificial intelligence model can be used in conversations. (For example, the language spoken by the speaker or agent) The part Solution It can be trained to associate with constituent elements.

[0043] In some embodiments, Once Once the text generation model is trained, the text generation model teeth By applying a text generation model to a conversational corpus Solution generate It may be used for that purpose. In some embodiments, based on a sample conversation provided as input, the text generation model generates additional data from the sample conversation in the conversation corpus. Solution (For example, a series of Solution Created in segments Solution ) outputs It is good to be able to do so. .

[0044] In some embodiments, Solution (For example, the first Solution or second Solution ) are entities related to topics (for example, the first topic or the second topic, respectively). Origin It can be generated using a document corpus. In some embodiments, entities related to a topic are the topic, associated with the topic. Solution , or one or more tasks related to the topic Solution Information related to segments or combinations thereof Having or such information Access proposal This could include individuals, groups of individuals, organizations, databases, libraries, etc.

[0045] In some embodiments, Solution This can be generated using a document corpus with a rule-based method. In some embodiments, the rule-based method uses topics, Solution ,or Solution The portion of a document within a document corpus related to the segment can be identified. In some embodiments, the rule is a topic, Solution ,or Solution This could describe how to identify segments and the conversational text to which they are associated (e.g., from user utterances).

[0046] for example, Solution This refers to a specified set of web pages (for example, the organization's web pages that provide step-by-step instructions for resolving common problems that occur with products purchased from the organization). The above It can be generated using rules related to Document Object Model ("DOM") elements. In some embodiments, it can be generated using rules (for example, related to the DOM elements of a web page). SolutionThis can be verified by a content domain expert (e.g., the creator of the web page). In some embodiments, Solution Using feedback from experts in the relevant subject area, Solution generate underlying The rules may be updated.

[0047] In some embodiments, a content domain expert reviews a specific document (e.g., a user manual for a company's product) from a document corpus. text part from Solution It can be generated or drafted. For example, a subject area expert can sentence The throat part Solution Annotations can be provided to identify which sub-component it is associated with. In some embodiments, annotations are used to train a text generation model from similar documents to other Solution It can generate the following. In some embodiments, the text generation model generates the user manual. sentence of Solution It can learn how to associate segments. In some embodiments, the text generation model then generates new data from the user manual received as input. Solution It can generate.

[0048] Now, referring to Figure 1, Solution Guidance Response of the expression A block diagram of system 100 for generation is shown. System 100 includes user device 102 and system device 104. System device 104 is a conversation context database 106. Solution Selector 108, Solution This includes 110, a response generator 112, and a response provider 114. The user device 102 and the system device 104 are configured to communicate with each other. The user device 102 and the system device 104 may be any devices including a processor configured to perform one or more of the functions or steps described herein.

[0049] In some embodiments, the system device 104 receives first voice data associated with a first user utterance in a conversation from the user device 102. The first voice data is stored in the conversation context database 106. Solution The selector 108 identifies a first topic from a set of topics associated with the first user utterance from the first audio data, and the first topic associated with the first topic Solution Identify 110. First Solution 110 is one or more tasks related to the topic Solution It has a segment. The response generator 112 of the system device 104 is first Solution The first Solution Based on the segment and first voice data, a first response is generated for a second user (e.g., an agent in an induction dialogue system). In some embodiments, the response generator generates the response using a sequence-to-sequence machine learning model. The first response is communicated to the first user (e.g., a user device 102) via the response provider 114.

[0050] In some embodiments, the first response is stored in the conversation context database 106 and used to generate a second response for the agent. In some embodiments, the system device 104 receives second voice data associated with a second user utterance in the conversation. In some embodiments, Solution The selector 108 verifies that the second user utterance is not associated with another topic in the set of topics. In some embodiments, Solution The selector may use the text classification model 116 to verify that the second user utterance is not associated with another topic in a set of topics. In some embodiments, the response generator 112 processes the first speech data, the agent's first response, the second speech data, and the first Solution The second Solution Based on the segment, a second response is generated for the agent.

[0051] In some embodiments, the first and second voice data and the first and second responses are stored in the conversation context database 106 and used to generate a third response for the agent. In some embodiments, the system device 104 receives third voice data associated with a third user utterance in the conversation. In some embodiments, Solution Selector 108 identifies a second topic associated with a third user utterance. In some embodiments, Solution Selector 108 is the second topic associated with the second topic Solution It is possible to identify the following. In some embodiments, the response generator 112 generates first voice data, the agent's first response, second voice data, the agent's second response, third voice data, and second Solution of Solution Based on the segment, a third response is generated for the agent.

[0052] Referring now to Figure 2, according to the embodiment of the present disclosure, Solution Guidance Response of the expression A flowchart of an exemplary method 200 for generation is shown. In some embodiments, the system's processor may perform the operations of method 200. In some embodiments, method 200 begins with operation 202. In operation 202, the processor receives first speech data associated with a first user utterance in a conversation in an assisted dialogue system. In some embodiments, method 200 proceeds to operation 204, in which the processor identifies a first topic of a set of topics associated with the first user utterance from the first speech data. In some embodiments, method 200 proceeds to operation 206. In operation 206, the processor identifies a first topic associated with the first topic Solution Identify the first Solution This is one or more tasks related to the topic. Solution It has a segment. In some embodiments, method 200 proceeds to operation 208. In operation 208, the processor has a first SolutionThe first Solution Based on the segment and the first voice data, a first response is generated for the second user.

[0053] As will be discussed in more detail herein, some or all of the operations of Method 200 may be performed in an alternative order or may not be performed at all, and furthermore, multiple operations may occur simultaneously or as internal parts of a larger process.

[0054] While this disclosure includes a detailed description of cloud computing, it should be understood that the implementation of the teachings described herein is not limited to cloud computing environments. Rather, embodiments of this disclosure can be implemented in combination with any other type of computing environment that is currently known or may be developed in the future.

[0055] Cloud computing is a service delivery model that enables convenient on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be rapidly provisioned and deployed with minimal administrative effort or interaction with service providers. This cloud model may include at least five features, at least three service models, and at least four deployment models.

[0056] The features are as follows:

[0057] On-demand self-service: Cloud consumers can automatically provision computing power, such as server time and network storage, unilaterally as needed, without requiring human interaction with a service provider.

[0058] Broad network access: Capabilities are available over the network and accessed through standard mechanisms that facilitate use by heterogeneous thin client platforms or thick client platforms (e.g., mobile phones, laptops, PDAs).

[0059] Resource pooling: A provider's computing resources are pooled to serve multiple consumers using a multi-tenant model, with various physical and virtual resources dynamically allocated and reallocated as needed. Consumers generally have a sense of partial independence, meaning they have no control or knowledge of the exact portion of the resources provided, but can identify portions at a higher level of abstraction (e.g., country, state, data center).

[0060] Rapid resilience: Capabilities can be provisioned quickly and resiliently, and in some cases, automatically and rapidly scale out, rapidly release, and rapidly scale in. To consumers, the capacity available for provisioning often appears unlimited and can be purchased in any quantity at any time.

[0061] Measured Services: Cloud systems automatically control and optimize resource usage by leveraging metric capabilities at a level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, active user accounts). They can monitor, control, and report on resource usage, providing transparency to both providers and consumers of the services they utilize.

[0062] The service model is as follows:

[0063] Software as a Service (SaaS): The ability offered to consumers is the use of a provider's applications running on cloud infrastructure. These applications can be accessed from various client devices via thin client interfaces, such as web browsers (e.g., web-based email). Consumers do not manage or control the underlying cloud infrastructure, including networks, servers, operating systems, storage, or individual application functionalities, except for limited, user-specific application configuration settings.

[0064] Platform as a Service (PaaS): The ability offered to consumers is the deployment of applications they have created or acquired, written using programming languages ​​and tools supported by the provider, onto a cloud infrastructure. Consumers do not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, and storage, but they do control the configuration of the deployed applications and, in some cases, the application hosting environment.

[0065] Infrastructure as a Service (IaaS): The ability offered to consumers is the provisioning of processing, storage, networking, and other fundamental computing resources that consumers can deploy and run any software they choose, including operating systems and applications. Consumers do not manage or control the underlying cloud infrastructure, but they do control the operating system, storage, and deployed applications, and, in some cases, limit their control over selected network components (e.g., host firewalls).

[0066] The deployment model is as follows:

[0067] Private Cloud: Cloud infrastructure is operated solely for an organization. It may be managed by the organization or a third party, and may reside on-premises or off-premises.

[0068] Community Cloud: The cloud infrastructure is shared by multiple organizations, supporting a specific community with shared concerns (e.g., mission, security requirements, policies, compliance considerations). It can be managed by an organization or a third party and can reside on-premises or off-premises.

[0069] Public cloud: Cloud infrastructure is made available to the general public or large industry groups, and is owned by the organization that sells the cloud services.

[0070] Hybrid Cloud: Cloud infrastructure is a configuration of two or more clouds (private, community, or public) that remain unique entities but are coupled together by standardized or proprietary technologies (e.g., cloudburst for load balancing between clouds) that enable data and application portability.

[0071] Cloud computing environments are service-oriented, emphasizing statelessness, low coupling, modularity, and semantic interoperability. At the heart of cloud computing is the infrastructure, including a network of interconnected nodes.

[0072] Figure 3A illustrates a cloud computing environment 310. As shown, the cloud computing environment 310 includes one or more cloud computing nodes 300 that can communicate with local computing devices used by cloud consumers, such as a personal digital assistant (PDA) or mobile phone 300A, a desktop computer 300B, a laptop computer 300C, or an automotive computer system 300N, or a combination thereof. The nodes 300 can communicate with each other. They may be grouped (not shown) physically or virtually into one or more networks, such as the aforementioned private cloud, community cloud, public cloud, or hybrid cloud, or a combination thereof.

[0073] This allows the cloud computing environment 310 to provide infrastructure, a platform, or software, or a combination thereof, as a service that does not require cloud consumers to maintain resources on their local computing devices. The types of computing devices 300A-N shown in Figure 3A are for illustrative purposes only, and it should be understood that the computing node 300 and the cloud computing environment 310 can communicate with any type of computerized device via any type of network or network-addressable connection (e.g., using a web browser), or a combination thereof.

[0074] Figure 3B shows a set of functional abstraction layers provided by the cloud computing environment 310 (Figure 3A). The components, layers, and functionalities shown in Figure 3B are for illustrative purposes only, and embodiments of this disclosure are not limited thereto. The following layers and corresponding functionalities are provided, as shown below:

[0075] The hardware and software layer 315 includes hardware and software components. Examples of hardware components include a mainframe 302, a RISC (Reduced Instruction Set Computer) architecture-based server 304, server 306, blade server 308, storage device 311, and network and network components 312. In some embodiments, the software components include network application server software 314 and database software 316.

[0076] The virtualization layer 320 provides an abstraction layer that may provide the following examples of virtual entities: a virtual server 322, virtual storage 324, a virtual network 326 including a virtual private network, a virtual application and operating system 328, and a virtual client 330.

[0077] For example, the management layer 340 may provide the following functions: Resource provisioning 342 provides dynamic procurement of computing and other resources used to perform tasks within the cloud computing environment. Metering and pricing 344 provides cost tracking as resources are used within the cloud computing environment and billing or invoicing for the consumption of these resources. For example, these resources may include application software licenses. Security provides identification and verification for cloud consumers and tasks, as well as protection for data and other resources. The user portal 346 provides consumers and system administrators with access to the cloud computing environment. Service level management 348 provides allocation and management of cloud computing resources to ensure that the required service levels are met. Service level agreement (SLA) planning and execution 350 provides proactive preparation and procurement of cloud computing resources for which future requirements are anticipated in accordance with the SLA.

[0078] The workload layer 360 provides examples of functions that can be utilized in a cloud computing environment. Examples of workloads and functions that can be provided from this layer include mapping and navigation 362, software development and lifecycle management 364, virtual classroom education delivery 366, data analysis processing 368, transaction processing 370, and responses for conversational systems 372. Solution Includes induced generation.

[0079] Figure 4 shows a high-level block diagram of an exemplary computer system 401 that may be used to implement one or more of the methods, tools, and modules described herein (for example, using one or more processor circuits or computer processors of a computer) and any related functions according to embodiments of the present disclosure. In some embodiments, the main components of the computer system 401 may include one or more CPUs 402, a memory subsystem 404, a terminal interface 412, a storage interface 416, an I / O (input / output) device interface 414, and a network interface 418, all of which may be directly or indirectly coupled to communicate with each other for intercomponent communication via a memory bus 403, an I / O bus 408, and an I / O bus interface unit 410.

[0080] The computer system 401 may include one or more general-purpose programmable central processing units (CPUs) 402A, 402B, 402C, and 402D, which are generally referred to herein as CPUs 402. In some embodiments, the computer system 401 may have multiple processors, which is typical for relatively large systems, but in other embodiments, the computer system 401 may alternatively be a single CPU system. Each CPU 402 may include one or more levels of onboard cache and may execute instructions stored in the memory subsystem 404.

[0081] The system memory 404 may include computer system-readable media in the form of volatile memory, such as random access memory (RAM) 422 or cache memory 424. The computer system 401 may further include other removable / non-removable, volatile / non-volatile computer system storage media. As just one example, a storage system 426 may be provided for reading to and writing to a non-removable non-volatile magnetic medium, such as a “hard drive”. Not shown, a magnetic disk drive for reading to and writing to a removable non-volatile magnetic disk (e.g., a “floppy disk”), or an optical disk drive for reading to or writing to a removable non-volatile optical disk, such as a CD-ROM, DVD-ROM, or other optical medium, may be provided. Furthermore, the memory 404 may include flash memory, such as a flash memory stick drive or flash drive. Memory devices can be connected to the memory bus 403 by one or more data media interfaces. The memory 404 may include at least one program product having a set of program modules (e.g., at least one) configured to perform the functions of various embodiments.

[0082] One or more programs / utilities 428, each having at least one set of program modules 430, may be stored in memory 404. A program / utility 428 may include a hypervisor (also called a virtual machine monitor), one or more operating systems, one or more application programs, other program modules, and program data. Each of the operating systems, one or more application programs, other program modules, and program data, or some combination thereof, may include an implementation of a networking environment. A program 428 or program module 430, or both, generally performs functions or methodologies in various embodiments.

[0083] Although the memory bus 403 is shown in Figure 4 as a single bus structure providing a direct communication path between the CPU 402, the memory subsystem 404, and the I / O bus interface 410, the memory bus 403 may include multiple different buses or communication paths, which can be arranged in any of various forms, such as point-to-point links in a hierarchical, star, or web configuration, multiple hierarchical buses, parallel and redundant paths, or any other suitable type of configuration. Furthermore, although the I / O bus interface 410 and the I / O bus 408 are shown as single units, the computer system 401 may include multiple I / O bus interface units 410, multiple I / O buses 408, or both, in some embodiments. Furthermore, although multiple I / O interface units are shown to isolate the I / O bus 408 from various communication paths to various I / O devices, in other embodiments, some or all of the I / O devices may be directly connected to one or more system I / O buses.

[0084] In some embodiments, the computer system 401 may be a multi-user mainframe computer system, a single-user system, or a server computer or similar device that has little or no direct user interface but receives requests from other computer systems (clients). Furthermore, in some embodiments, the computer system 401 may be implemented as a desktop computer, a portable computer, a laptop or notebook computer, a tablet computer, a pocket computer, a telephone, a smartphone, a network switch or router, or any other suitable type of electronic device.

[0085] It should be noted that Figure 4 is intended to show typical main components of an exemplary computer system 401. However, in some embodiments, individual components may be more or less complex than those shown in Figure 4, or there may be components other than those shown in Figure 4, or components in addition to those components, and the number, type, and configuration of such components may vary.

[0086] As will be discussed in more detail herein, some or all of the operations of some embodiments of the methods described herein may be performed in an alternative order or may not be performed at all, and furthermore, multiple operations may occur simultaneously or as internal parts of a larger process.

[0087] This disclosure may be a system, method, or computer program product, or combination thereof, in an integration of any possible level of technical detail. A computer program product may include a computer-readable storage medium (or more) having computer-readable program instructions thereon for causing a processor to perform aspects of this disclosure.

[0088] A computer-readable storage medium can be a tangible device capable of holding and storing instructions for use by an instruction-executing device. A computer-readable storage medium may, for example, be an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. A non-exhaustive list of more specific examples of computer-readable storage media includes portable computer diskettes, hard disks, random-access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random-access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital multipurpose disks (DVDs), memory sticks, floppy disks, mechanically encoded devices such as punched cards or grooved structures on which instructions are recorded, and any suitable combination thereof. The computer-readable storage media used herein should not be interpreted as transient signals in themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., optical pulses through fiber optic cables), or electrical signals transmitted through wires.

[0089] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof. The network may include copper transmission cables, optical transmission fibers, wireless transmissions, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. A network adapter card or network interface within each computing / processing device receives computer-readable program instructions from the network and transfers the computer-readable program instructions for storage in a computer-readable storage medium within each computing / processing device.

[0090] The computer-readable program instructions for performing the operations of the Disclosure may be assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, integrated circuit configuration data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk®, C++, and procedural programming languages ​​such as the C programming language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or on a server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or wide area network (WAN), or it may be connected to an external computer (for example, via the Internet using an Internet service provider). In some embodiments, for example, an electronic circuit including a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA) may execute computer-readable program instructions by utilizing state information of computer-readable program instructions for personalizing the electronic circuit in order to perform an aspect of the present disclosure.

[0091] Aspects of this disclosure are described herein with reference to flowcharts, block diagrams, or both, of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It will be understood that each block in a flowchart or block diagram, or both, and any combination of blocks in a flowchart or block diagram, or both, can be implemented by computer-readable program instructions.

[0092] These computer-readable program instructions are provided to a computer processor or other programmable data processing device, which can generate a machine, and as a result, instructions executed via the processor of the computer or other programmable data processing device create means for performing functions / operations specified in one or more blocks of a flowchart or block diagram or both. These computer-readable program instructions can also be stored in a computer-readable storage medium that can instruct a computer, a programmable data processing device or other device or both to function in a particular way, and as a result, a computer-readable storage medium having instructions stored therein comprises a product containing instructions that perform modes of functions / operations specified in one or more blocks of a flowchart or block diagram or both.

[0093] Computer-readable program instructions can also be loaded into a computer, other programmable data processing device, or other device, causing a series of operational steps on the computer, other programmable device, or other device to generate a computer implementation process, such that the instructions executed on the computer, other programmable device, or other device implement a function / operation specified in one or more blocks of a flowchart or block diagram or both.

[0094] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of the systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or part of an instruction having one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions shown in a block may occur out of the order shown in the figure. For example, two consecutively shown blocks may actually be executed as a single step, simultaneously or substantially simultaneously, in a way that partially or entirely overlaps in time, or the blocks may sometimes be executed in reverse order depending on the related functions. It should also be noted that each block in a block diagram or flowchart, or both, and any combination of blocks in a block diagram or flowchart, or both, may be implemented by a special-purpose hardware-based system that performs a specified function or operation, or a combination of special-purpose hardware and computer instructions.

[0095] The descriptions of the various embodiments of this disclosure are presented for illustrative purposes only and are not intended to be exhaustive or limit to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the embodiments described. The terms used herein have been selected to best describe the principles, practical applications, or technological improvements beyond the technology available on the market, or to enable those skilled in the art to understand the embodiments disclosed herein.

[0096] While this disclosure has described specific embodiments, any changes and modifications thereto are expected to be apparent to those skilled in the art. Therefore, the following claims are intended to encompass all such changes and modifications that fall within the true spirit and scope of this disclosure.

Claims

1. A computer implementation method, The processor receives the first audio data associated with the first utterance of the first user in a conversation in the guided dialogue system, A step of identifying a first topic from a series of topics associated with the first utterance of the first user from the first audio data, A step of identifying a first solution associated with the first topic, wherein the first solution has a series of solution segments for performing tasks related to the first topic, A step of generating a first response for a second user based on the first solution segment of the series of solution segments of the first solution and the first voice data. A method that includes [a certain feature].

2. The method according to claim 1, wherein the step of generating a first response for a second user includes the step of generating using a sequence-to-sequence machine learning model which includes a generative pre-trained transformer that takes the first audio data and the first solution segment as input and is trained to output the first response.

3. The method according to claim 1 or 2, wherein the first topic is identified using a text classification model.

4. The steps include receiving second audio data associated with the second utterance of the first user in the conversation in the guided dialogue system, The step of confirming that the second utterance is not associated with another topic in the series of topics, A step of generating a second response for the second user based on the first audio data, the first response of the second user, the second audio data, and the second solution segment of the series of solution segments of the first solution. The method according to claim 3, further comprising:

5. The steps include receiving third audio data associated with the third utterance of the first user in the conversation in the guided dialogue system, The step of identifying a second topic associated with the third utterance, The steps include identifying a second solution associated with the second topic, A step of generating a third response for the second user based on the first audio data, the first response of the second user, the second audio data, the second response of the second user, the third audio data, and the solution segment of the second solution. The method according to claim 4, further comprising:

6. A computer implementation method, The processor receives the first audio data associated with the first utterance of the first user in a conversation in the guided dialogue system, A step of identifying a first topic from a series of topics associated with the first utterance of the first user from the first audio data, A step of identifying a first solution associated with the first topic, wherein the first solution has a series of solution segments for performing tasks related to the first topic, and the first solution is generated using a text generation artificial intelligence model that generates solutions from sample conversations. A step of generating a first response for a second user based on the first solution segment of the series of solution segments of the first solution and the first voice data. A method that includes [a certain feature].

7. The method according to claim 6, wherein the step of generating a first response for a second user includes the step of generating using a sequence-to-sequence machine learning model which includes a generative pre-trained transformer that takes the first audio data and the first solution segment as input and is trained to output the first response.

8. The method according to claim 6 or 7, wherein the first topic is identified using a text classification model.

9. The steps include receiving second audio data associated with the second utterance of the first user in the conversation in the guided dialogue system, The step of confirming that the second utterance is not associated with another topic in the series of topics, A step of generating a second response for the second user based on the first audio data, the first response of the second user, the second audio data, and the second solution segment of the series of solution segments of the first solution. The method according to claim 8, further comprising:

10. The steps include receiving third audio data associated with the third utterance of the first user in the conversation in the guided dialogue system, The step of identifying a second topic associated with the third utterance, The steps include identifying a second solution associated with the second topic, A step of generating a third response for the second user based on the first audio data, the first response of the second user, the second audio data, the second response of the second user, the third audio data, and the solution segment of the second solution. The method according to claim 9, further comprising:

11. A computer implementation method, The processor receives the first audio data associated with the first utterance of the first user in a conversation in the guided dialogue system, A step of identifying a first topic from a series of topics associated with the first utterance of the first user from the first audio data, A step of identifying a first solution associated with the first topic, wherein the first solution has a set of solution segments for performing tasks related to the first topic, and the first solution is generated using a document corpus derived from entities related to the first topic. A step of generating a first response for a second user based on the first solution segment of the series of solution segments of the first solution and the first voice data. A method that includes [a certain feature].

12. The method according to claim 11, wherein the step of generating a first response for a second user includes the step of generating using a sequence-to-sequence machine learning model which includes a generative pre-trained transformer that takes the first audio data and the first solution segment as input and is trained to output the first response.

13. The method according to claim 11 or 12, wherein the first topic is identified using a text classification model.

14. The steps include receiving second audio data associated with the second utterance of the first user in the conversation in the guided dialogue system, The step of confirming that the second utterance is not associated with another topic in the series of topics, A step of generating a second response for the second user based on the first audio data, the first response of the second user, the second audio data, and the second solution segment of the series of solution segments of the first solution. The method according to claim 13, further comprising:

15. The steps include receiving third audio data associated with the third utterance of the first user in the conversation in the guided dialogue system, The step of identifying a second topic associated with the third utterance, The steps include identifying a second solution associated with the second topic, A step of generating a third response for the second user based on the first audio data, the first response of the second user, the second audio data, the second response of the second user, the third audio data, and the solution segment of the second solution. The method according to claim 14, further comprising:

16. It is a system, Memory and The system comprises a processor that communicates with the memory, and the processor A procedure for receiving the first audio data associated with the first utterance of the first user in a conversation in an assisted dialogue system, A procedure for identifying a first topic from a series of topics associated with the first utterance of the first user, from the first audio data, A procedure for identifying a first solution associated with the first topic, wherein the first solution has a set of solution segments for performing tasks related to the first topic, A system configured to perform an operation comprising a procedure for generating a first response for a second user based on the first solution segment of the series of solution segments of the first solution and the first voice data.

17. The system according to claim 16, wherein the step of generating a first response for a second user includes a step of generating using a sequence-to-sequence machine learning model which includes a generative pre-trained transformer that takes the first audio data and the first solution segment as input and is trained to output the first response.

18. The system according to claim 16 or 17, wherein the first topic is identified using a text classification model.

19. The aforementioned processor, A procedure for receiving second audio data associated with the second utterance of the first user in the conversation in the guided dialogue system, A procedure to confirm that the second utterance is not associated with another topic in the series of topics, The system according to claim 18, further configured to perform an operation comprising a step of generating a second response for the second user based on the first audio data, the first response of the second user, the second audio data, and the second solution segment of the series of solution segments of the first solution.

20. The aforementioned processor, A procedure for receiving third audio data associated with the third utterance of the first user in the conversation in the guided dialogue system, A procedure for identifying a second topic associated with the third utterance, A procedure for identifying a second solution associated with the second topic described above, The system according to claim 19, further configured to perform an operation comprising a procedure for generating a third response for the second user based on the first audio data, the first response of the second user, the second audio data, the second response of the second user, the third audio data, and a solution segment of the second solution.

21. In the processor, A procedure for receiving the first audio data associated with the first utterance of the first user in a conversation in an assisted dialogue system, A procedure for identifying a first topic from a series of topics associated with the first utterance of the first user, from the first audio data, A procedure for identifying a first solution associated with the first topic, wherein the first solution has a set of solution segments for performing tasks related to the first topic, A procedure for generating a first response for a second user based on the first solution segment of the series of solution segments of the first solution and the first voice data, and A computer program designed to execute something.

22. The computer program according to claim 21, wherein the step of generating a first response for a second user includes a step of generating using a sequence-to-sequence machine learning model which includes a generative pre-trained transformer that takes the first audio data and the first solution segment as input and is trained to output the first response.

23. The computer program according to claim 21 or 22, wherein the first topic is identified using a text classification model.

24. The aforementioned processor, A procedure for receiving second audio data associated with the second utterance of the first user in the conversation in the guided dialogue system, A procedure to confirm that the second utterance is not associated with another topic in the series of topics, The computer program according to claim 23, which causes a procedure to be performed to generate a second response for the second user based on the first audio data, the first response of the second user, the second audio data, and the second solution segment of the series of solution segments of the first solution.

25. The aforementioned processor, A procedure for receiving third audio data associated with the third utterance of the first user in the conversation in the guided dialogue system, A procedure for identifying a second topic associated with the third utterance, A procedure for identifying a second solution associated with the second topic described above, The computer program according to claim 24, further comprising the steps of generating a third response for the second user based on the first audio data, the first response of the second user, the second audio data, the second response of the second user, the third audio data, and the solution segment of the second solution.