Systems and methods for bot authoring by mining intent from conversation data via intent seeding

JP2023545947A5Active Publication Date: 2025-11-19GENESIS CLOUD SERVICES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023519241
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-09-25
Filing Date
2021-09-27
Publication Date
2025-11-19
Estimated Expiration
2041-09-27

AI Technical Summary

Technical Problem

Existing methods for automating bot authoring workflows and mining intentions from natural language conversational data are time-consuming and expensive, requiring manual inspection of large volumes of chat transcripts to identify intents and utterances.

Method used

An intent mining process that automatically mines intents and associated utterances from conversations using an algorithm, allowing for efficient bot authoring by importing conversation data, identifying utterances with intent, generating candidate intents, and grouping them semantically to create intent labels and alternatives, which can be uploaded to conversational bots for automated interactions.

Benefits of technology

This approach significantly reduces the time and cost of bot authoring by automating the intent identification process, enabling faster development and deployment of conversational bots with improved efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

1. A method for authoring a conversational bot, the method including: receiving conversation data; receiving seed intent data including seed intents having seed intent labels and utterances having sample intents; mining the conversation data using an intent mining algorithm to determine new utterances to associate with the seed intents; expanding the seed intent data to include the mined new utterances associated with the seed intents; and uploading the expanded seed intent data to the conversational bot. The intent mining algorithm may include identifying utterances having intents, identifying candidate intents, and, for each seed intent, identifying seed intent alternatives from utterances having sample intents, and associating the utterances having intents from the conversation data with the seed intents via determining a degree of semantic similarity between the candidate intents of the utterances having intent and the seed intent alternatives.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] (Cross - reference to related applications) This application claims priority to U.S. Provisional Patent Application No. 63 / 083,561, filed on September 25, 2020, which was converted to U.S. Patent Application No. 17 / 218,456, entitled "SYSTEMS AND METHODS RELATING TO BOT AUTHORING BY MINING INTENTS FROM NATURAL LANGUAGE CONVERSATIONS", filed on March 31, 2021.

Background Art

[0002] The present invention generally relates to a telecommunications system in the field of customer relationship management, such as customer support via Internet - based service options. More particularly, but not by way of limitation, the present invention relates to systems and methods for implementing an intent mining process for automating bot authoring workflows and / or mining intents and associated utterances from natural language conversation data using an intent seeding process.

Summary of the Invention

[0003] The present invention includes a computer implementation method for authoring a conversational bot, which provides intent mining using intent seeding. The method may include receiving conversational data, wherein the conversational data includes text derived from conversations, and each conversation is between a customer and a customer service representative; receiving seed intent data, which may include seed intents, wherein each seed intent includes a seed intent label and utterances having sample intents associated with the seed intent; using an intent mining algorithm to automatically mine the conversational data to determine new utterances to associate with seed intents; expanding the seed intent data to include the newly mined utterances associated with seed intents; and uploading the expanded seed intent data to a conversational bot, which may use the conversational bot to conduct automated conversations with other customers. The intent mining algorithm may include analyzing utterances occurring within conversations in the conversational data to identify intent-containing utterances. Each utterance may include a turn in a conversation, thereby communicating in the form of customer utterances or customer service representative utterances. An intentional utterance can be defined as one of the utterances determined to be highly likely to express an intention. An intention mining algorithm may further include analyzing the identified intentional utterances to identify candidate intentions. Each candidate intention is identified as a text phrase occurring within one of the intentional utterances, having two parts: an action which may contain a word or phrase describing a purpose or task, and an object which may contain a word or phrase describing an object or thing on which the action operates. For each seed intention, the intention mining algorithm may further include identifying seed intention surrogates from sample intentional utterances associated with the seed intention.A seed intent substitute is identified as a text phrase occurring within one of the sample intent utterances, which may consist of two parts: an action that may contain a word or phrase describing a purpose or task, and an object that may contain a word or phrase describing the object or thing on which the action operates. The intent mining algorithm may further include associating intent utterances from conversational data with seed intents by determining the degree of semantic similarity between candidate intents present within the intent utterances and seed intent substitutes belonging to each of the seed intent labels.

[0004] These and other features of the present application will become more apparent by considering the following detailed description of exemplary embodiments in conjunction with the drawings and the attached claims. [Brief explanation of the drawing]

[0005] A more complete understanding of the present invention will become more readily apparent, as it is better understood by referring to the following detailed description when the invention is considered in conjunction with the accompanying drawings in which similar reference numerals indicate similar components. [Figure 1] A schematic block diagram of a computing device that may be enabled or implemented by exemplary embodiments of the present invention is shown. [Figure 2] A schematic block diagram of a communications infrastructure or contact center that may be enabled or implemented by exemplary embodiments of the present invention is shown. [Figure 3] This is a schematic block diagram showing further details of a chat server operating as part of a chat system according to an embodiment of the present invention. [Figure 4] This is a schematic block diagram of a chat module according to an embodiment of the present invention. [Figure 5] This is an exemplary customer chat interface according to an embodiment of the present invention. [Figure 6] This is a block diagram of a customer automation system according to an embodiment of the present invention. [Figure 7] This is a flowchart of a method for automating customer interactions according to an embodiment of the present invention. [Figure 8] This is a workflow for authoring conversational bots. [Figure 9] This is an exemplary flowchart for intent mining according to the present invention. [Figure 10] This is an exemplary flowchart for intent mining by seeding intent according to the present invention. [Modes for carrying out the invention]

[0006] For the purpose of facilitating an understanding of the principles of the present invention, exemplary embodiments illustrated in the drawings are described herein using specific language. However, it will be apparent to those skilled in the art that detailed materials provided in the examples may not be necessary for carrying out the present invention. In other examples, well-known materials or methods are not described in detail to avoid obscuring the present invention. In addition, further modifications to the examples provided or to the application of the principles of the present invention, as presented herein, are intended to be as commonly conceivable to those skilled in the art.

[0007] As used herein, the language specifying non-limiting embodiments and examples includes "e.g.," "i.e.," "for example," and "for instance." Furthermore, throughout this specification, the terms "embodiment," "one embodiment," and "this embodiment" are used. References to "embodiments," "exemplary embodiments," and "specific embodiments" mean that certain features, structures, or characteristics described in relation to a given embodiment may be included in at least one embodiment of the present invention. Therefore, the appearance of phrases such as "embodiment," "one embodiment," "this embodiment," "exemplary embodiment," and "specific embodiment" does not necessarily refer to the same embodiment or example. Furthermore, certain features, structures, or characteristics may be combined in any preferred combination and / or partial combination in one or more embodiments or examples.

[0008] Those skilled in the art will recognize from this disclosure that various embodiments can be computer-implemented using many different types of data processing equipment, and that embodiments can be implemented as devices, methods, or computer program products. Accordingly, exemplary embodiments may take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Exemplary embodiments may further take the form of computer program products embodied by computer-usable program code in any tangible medium of expression. In any case, exemplary embodiments may generally be referred to as “modules,” “systems,” or “methods.”

[0009] The flowcharts and block diagrams provided in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to exemplary embodiments of the present invention. In this regard, it will be understood that each block or combination thereof in the flowcharts and / or block diagrams may represent a module, segment, or portion of program code having one or more executable instructions for implementing a specified logical function. Similarly, it will be understood that each block or combination thereof in the flowcharts and / or block diagrams may be implemented by a dedicated hardware-based system or a combination of dedicated hardware and computer instructions that performs a particular operation or function. Such computer program instructions may also be stored in computer-readable media that can instruct a computer or other programmable data processing device to function in a particular way so as to produce a product containing instructions that implement the functions or operations specified in each block or combination thereof in the flowcharts and / or block diagrams.

[0010] Computing devices It will be understood that the systems and methods of the present invention can be computer-implemented using many different forms of data processing equipment, such as a digital microprocessor and associated memory, which execute appropriate software programs. For background, Figure 1 illustrates a schematic block diagram of an exemplary computing device 100 according to embodiments of the present invention, and / or which such embodiments may enable or implement. It should be understood that Figure 1 is provided as a non-limiting embodiment.

[0011] Computing device 100 may be implemented, for example, through firmware (e.g., application-specific integrated circuits), hardware, or a combination of software, firmware, and hardware. It will be understood that each of the servers, controllers, switches, gateways, engines, and / or modules (which may be collectively referred to as servers or modules) in the following diagrams may be implemented through one or more computing devices 100. In embodiment, various servers may be processes operating on one or more processors of one or more computing devices 100 that execute computer program instructions to perform various functions described herein and interact with other systems or modules. Unless particularly limited, functions described in relation to multiple computing devices may be integrated into a single computing device, or various functions described in relation to a single computing device may be distributed across several computing devices. Furthermore, relating to the computing systems shown in the following diagrams, such as the contact center system 200 in Figure 2, its various servers and computer devices may be located on local computing devices 100 (i.e., on-site or in the same physical location as the contact center agents), remote computing devices 100 (i.e., away from the site or in a cloud computing environment, for example, in a remote data center connected to the contact center via a network), or some combination thereof. The functions provided by servers located on remote computing devices may be provided as if such servers were on-site, on a virtual private network (VPN). The functions may be accessed and provided via the Internet, or provided using software as a service (SaaS) accessed over the Internet using various protocols, such as exchanging data via extensible markup language (XML), JSON, etc.

[0012] As illustrated in the illustrated examples, the computing device 100 includes a central processing unit (CPU) or processor 105 and a main memory The computing device 100 may also include a memory device 115, a removable media interface 120, a network interface 125, an I / O controller 130, and one or more input / output (I / O) devices 135, which may include a display device 135A, a keyboard 135B, and a pointing device 135C, as shown. The computing device 100 may further include additional elements such as a memory port 140, a bridge 145, I / O ports, one or more additional input / output devices 135D, 135E, 135F, and a cache memory 150 that communicates with the processor 105.

[0013] The processor 105 may be any logic circuit that responds to and processes instructions fetched from the main memory 110. For example, the processor 105 may be implemented by an integrated circuit, such as a microprocessor, microcontroller, or graphics processing unit, or in a field-programmable gate array or application-specific integrated circuit. As shown, the processor 105 may communicate directly with the cache memory 150 via a secondary bus or backside bus. The cache memory 150 typically has a faster response time than the main memory 110. The main memory 110 may be one or more memory chips capable of storing data, allowing the stored data to be directly accessed by the central processing unit 105. The storage device 115 may provide storage for an operating system and other software that controls scheduling tasks and access to system resources. Unless otherwise specified, the computing device 100 may include an operating system and software capable of performing the functionalities described herein.

[0014] As illustrated in the illustrated embodiments, the computing device 100 may include a wide variety of I / O devices 135, one or more of which may be connected via an I / O controller 130. Input devices may include, for example, a keyboard 135B and a pointing device 135C, such as a mouse or optical pen. Output devices may include, for example, a video display device, a speaker, and a printer. The I / O devices 135 and / or the I / O controller 130 may include suitable hardware and / or software to enable the use of multiple display devices. The computing device 100 may also support one or more removable media interfaces 120, such as a disk drive, a USB port, or any other suitable device for reading data from or writing data to computer-readable media. More generally, the I / O devices 135 may include any conventional devices for performing the functions described herein.

[0015] Computing device 100 may be, but is not limited to, any workstation, desktop computer, laptop or notebook computer, server machine, virtual machine, mobile or smartphone, portable telecommunications device, media playback device, game system, mobile computing device, or any other type of computing, telecommunications, or media device capable of performing the operations and functions described herein. Computing device 100 includes multiple devices connected by a network or connected to other systems and resources via a network. As used herein, a network includes one or more computing devices, machines, clients, client nodes, client machines, client computers, client devices, endpoints, or endpoint nodes that communicate with one or more other computing devices, machines, clients, client nodes, client machines, client computers, client devices, endpoints, or endpoint nodes. For example, a network may be a private or public switched telephone network (PSTN), a wireless carrier network, or a local area network. This may include a local area network (LAN), a private wide area network (WAN), or a public WAN such as the Internet, and appropriate communication is required. A connection is established using a communication protocol. More generally, unless otherwise specified, it should be understood that computing device 100 can communicate with other computing devices 100 over any type of network using any conventional communication protocol. Furthermore, the network may be a virtual network environment in which various network components are virtualized. For example, various machines may be virtual machines implemented as software-based computers running on a physical machine, or a "hypervisor" type of virtualization may be used in which multiple virtual machines run on the same host physical machine. Other types of virtualization are also conceivable.

[0016] Contact Center Referring here to Figure 2, a communications infrastructure or contact center system 200 according to an exemplary embodiment of the present invention and / or which an exemplary embodiment of the present invention may enable or implement is shown. It should be understood that in this specification, the term “contact center system” is used to refer to the system and / or its components shown in Figure 2, while the term “contact center” is used more generally to refer to the contact center system, the customer service provider operating these systems, and / or the organization or company associated therewith. Therefore, unless specifically limited, the term “contact center” generally refers to the contact center system (such as contact center system 200), the associated customer service provider (such as a specific customer service provider providing customer services through contact center system 200), and the organization or company on which customer services are provided.

[0017] As background, customer service providers generally provide many types of services through contact centers. Such contact centers can have employees or customer service agents (or simply "agents") located therein, and the agents function as an interface between a company, business, government agency, or organization (hereinafter interchangeably referred to as an "organization" or "business") and people such as users, individuals, or customers (hereinafter interchangeably referred to as "individuals" or "customers"). For example, contact center agents can assist customers in making purchase decisions, taking orders, or resolving issues related to products or services already received. Within the contact center, such interactions between contact center agents and external entities or customers can occur via various communication channels, such as, for example, voice (e.g., a telephone call or voice over IP, i.e., a VoIP call), video (e.g., a video conference), text (e.g., an email and a text chat), screen sharing, co-browsing, and the like.

[0018] In practice, contact centers generally strive to provide high-quality service to customers while minimizing costs. For example, one way a contact center operates is to handle all customer interactions with live agents. While this approach can be quite successful from a service quality standpoint, it is likely to be prohibitively expensive due to the high cost of agent labor. For this reason, most contact centers utilize a level of automated processes instead of live agents, such as interactive voice response (IVR) systems, interactive media response (IMR) systems, internet robots, or "bots," and automated chat modules, or "chatbots." In many cases, this has proven to be a successful strategy because automated processes can be highly efficient at handling certain types of interactions and can be effective in reducing the need for live agents. Such automation allows contact centers to target customer interactions that are more difficult for human agents to handle, while automated processes handle more repetitive or routine tasks. Furthermore, automated processes can be structured in a way that optimizes efficiency and promotes repeatability. Human agents, i.e., live agents, may forget to answer specific questions or thoroughly pursue certain details, but such errors are typically avoided through the use of automated processes. Customer service providers are increasingly relying on automated processes to interact with customers, while the use of such technologies by customers remains far less developed. Therefore, on the contact center side of the interaction, IVR systems, IMR systems, and / or bots are used to automate parts of the interaction, while customer-side actions remain performed manually by the customer.

[0019] Referring specifically to FIG. 2, contact center system 200 can be used by a customer service provider to provide various types of services to customers. For example, contact center system 200 can be involved in interactions where an automated process (or bot) or a human agent communicates with a customer and is used to manage the interactions. As understood, contact center system 200 can be an in-house facility of a business or enterprise for implementing sales and customer service functions related to products and services available through the enterprise. In another aspect, contact center system 200 can be operated by a third-party service provider that contracts to provide services on behalf of another organization. Further, contact center system 200 can be deployed on equipment dedicated to an enterprise or a third-party service provider and / or in a remote computing environment such as a private or public cloud environment with infrastructure for supporting multiple contact centers for multiple enterprises, for example. Contact center system 200 can include software applications or programs that can be executed on-premises or remotely, or some combination thereof. Further, it should be understood that various components of contact center system 200 can be distributed across various geographical locations and are not necessarily included in a single location or computing environment.

[0020] Furthermore, unless otherwise specified, it should be understood that any of the computing elements of the present invention may be implemented in a cloud-based or cloud computing environment. As used herein, “cloud computing” or simply “cloud” is defined as a model for enabling ubiquitous, convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, servers, storage devices, applications, and services) that can be rapidly provisioned via virtualization, released with minimal administrative effort or service provider interaction, and then scaled as appropriate. Cloud computing encompasses a variety of characteristics (e.g., on-demand self-service, wide-area network access, resource pooling, rapid resilience, measurable services, etc.), service models (e.g., Software as a Service ("SaaS"), Platform as a Service ("PaaS"), Infrastructure as a Service). Infrastructure as a Service (IaaS), and deployment models (e.g., platform Clouds can consist of (private clouds, community clouds, public clouds, etc.). Cloud execution models, often referred to as "serverless architectures," generally include service providers that dynamically manage the allocation and provisioning of remote servers to achieve desired functionality.

[0021] According to the illustrated embodiment in Figure 2, the components or modules of the contact center 200 include a plurality of customer devices 205A, 205B, 205C, a communication network (or simply "Network") 210, a switch / media gateway 212, a call controller 214, a bidirectional media response (IMR) server 216, a routing server 218, a storage device 220, a statistics (or "stat") server 226, a plurality of agent devices 230A, 230B, 230C, each including workbins 232A, 232B, 232C, a multimedia / social media server 234, a knowledge management server 236 coupled to a knowledge system 238, a chat server 240, a web server 242, an interaction (or "iXn") server 244, and a universal contact server (or "universal contact") This may include a server (UCS) 246, a reporting server 248, a media service server 249, and an analysis module 250. It should be understood that any computer implementation component, module, or server shown in relation to Figure 2, or any of the following figures, may be implemented via a computing device of a type such as the computing device 100 in Figure 1. As understood, the contact center system 200 generally manages resources (e.g., personnel, computers, telecommunications equipment, etc.) to enable the delivery of services via telephone, email, chat, or other communication mechanisms. Such services may vary depending on the type of contact center and may include, for example, customer service, help desk functions, emergency response, telemarketing, order taking, etc.

[0022] A customer wishing to receive services from the contact center system 200 may initiate inbound communication (e.g., telephone calls, emails, chats, etc.) to the contact center system 200 via a customer device 205. Figure 2 shows three such customer devices, namely customer devices 205A, 205B, and 205C, but it should be understood that any number may exist. The customer device 205 may be a communication device such as a telephone, smartphone, computer, tablet, or laptop. According to the functions described herein, a customer may generally use the customer device 205 to initiate, manage, and perform communications with the contact center system 200, such as telephone calls, emails, chats, text messages, web browsing sessions, and other multimedia transactions.

[0023] Inbound and outbound communications to customer devices 205 may traverse network 210, where the nature of the network depends on the type of customer device being used and the form of communication. Examples of network 210 include telephone, cellular, and / or data service communication networks. Network 210 may be a private or public switched telephone network (PSTN), a local area network (LAN), a private wide area network (WAN), and / or a public WAN such as the Internet. Furthermore, network 210 may also be a code division multiple access (CDMA) network, a mobile communication network. Global system for mobile communications (GSM) This may include wireless carrier networks, including, but not limited to, 3G, 4G, LTE, 5G, and any other wireless network / technology commonly used in the art.

[0024] With respect to the switch / media gateway 212, it may be coupled to the network 210 to receive and transmit telephone calls between customers and the contact center system 200. The switch / media gateway 212 may be a telephone switch or communications switch configured to function as a central switch for agent-level routing within the center. The switch may be a hardware switching system or implemented via software. For example, switch 215 may be an automated call distributor, a private branch exchange (PBX), or an IP network. The system may include a software switch and / or any other switch having dedicated hardware and software configured to receive internet-sourced and / or telephone network-sourced interactions from the customer and route these interactions to, for example, one of the agent devices 230. Thus, generally, the switch / media gateway 212 establishes a voice connection between the customer and the agent by establishing a connection between the customer device 205 and the agent device 230.

[0025] As further shown, the switch / media gateway 212 may be coupled to a call controller 214, which functions as an adapter or interface between the switch and other routing, monitoring, and communication processing components of the contact center system 200. The call controller 214 may be configured to handle PSTN calls, VoIP calls, etc. For example, the call controller 214 may include computer-telephone integration (CTI) software for interfaced with the switch / media gateway and other components. The call controller 214 handles session initiation protocol (SIP) calls. It may include a SIP server for this purpose. The call controller 214 can also extract data about incoming interactions, such as the customer's phone number, IP address, or email address, and then communicate these with other contact center components when processing the interaction.

[0026] With respect to the two-way media response (IMR) server 216, it may be configured to enable self-help or virtual assistant functionality. Specifically, the IMR server 216 may be similar to an interactive voice response (IVR) server, except that the IMR server 216 is not limited to voice and may also cover various media channels. In the example illustrating voice, the IMR server 216 may consist of IMR scripts to query the customer about their needs. For example, a bank contact center may, via an IMR script, tell the customer to "press 1" if they want to retrieve their account balance. Through continuous interaction with the IMR server 216, the customer can receive service without having to speak to an agent. The IMR server 216 may also be configured to determine why the customer is contacting the contact center so that the communication can be routed to the appropriate resources.

[0027] With respect to the routing server 218, it may function to route incoming interactions. For example, if it is determined that an inbound communication should be handled by a human agent, the functionality within the routing server 218 may select the most appropriate agent and route the communication to that agent. This agent selection may be based on which available agent is best suited to handle the communication. More specifically, the selection of the appropriate agent may be based on a routing strategy or algorithm implemented by the routing server 218. In doing so, the routing server 218 may query data related to the incoming interaction, such as data related to a specific customer, available agents, and the type of interaction, and this data may be stored in a specific database, as described further below. Once an agent is selected, the routing server 218 may interact with the call controller 214 to route (i.e., connect) the incoming interaction to the corresponding agent device 230. As part of this connection, information about the customer may be provided to the selected agent via that agent device 230. This information is intended to enhance the services that the agent can provide to the customer.

[0028] With respect to data storage, the contact center system 200 may include one or more mass storage devices, generally represented by storage devices 220, for storing data in one or more databases related to the functions of the contact center. For example, storage device 220 may store customer data maintained in customer database 222. Such customer data may include customer profiles, contact information, service level agreements (SLAs), and interaction history (e.g., previous interactions). This may include details of a previous interaction with a particular customer, including the nature of the interaction, disposal data, wait times, processing times, and actions taken by the contact center to resolve the customer's issue. In another embodiment, the storage device 220 may store agent data in the agent database 223. Agent data maintained by the contact center system 200 may include agent availability and agent profiles, schedules, skills, processing times, etc. In another embodiment, the storage device 220 may store interaction data in the interaction database 224. Interaction data may include data related to numerous past interactions between the customer and the contact center. More generally, unless otherwise specified, the storage device 220 may be configured to include databases and / or store data related to any of the types of information described herein, and it should be understood that these databases and / or data are accessible to other modules or servers of the contact center system 200 in a manner that facilitates the functions described herein. For example, a server or module of the contact center system 200 may query such databases to retrieve data stored therein or to transmit data therefor to be stored. The storage device 220 can take the form of any conventional storage medium, for example, and may be locally located or operated from a remote location. For example, the database may be a Cassandra database, a NoSQL database, or an SQL database, and may be managed by a database management system such as Oracle, IBM DB2, Microsoft SQL Server, Microsoft Access, or PostgreSQL.

[0029] With respect to the stat server 226, it may be configured to record and aggregate data related to the performance and operating characteristics of the contact center system 200. Such information may be compiled by the stat server 226 and made available to other servers and modules, such as the reporting server 248, which may then use the data to produce reports used to manage the operating characteristics of the contact center and to perform automated actions in accordance with the functions described herein. Such data may relate to the status of the contact center's resources, such as average wait time, discard rate, agent occupancy rate, and others that the functions described herein may require.

[0030] The agent device 230 of the contact center 200 may be a communication device configured to interact with various components and modules of the contact center system 200 in a manner that facilitates the functions described herein. For example, the agent device 230 may include a telephone adapted for regular telephone calls or VoIP calls. The agent device 230 may further include a computing device configured to communicate with the server of the contact center system 200, perform data processing associated with its operation, and interface with customers via voice, chat, email, and other multimedia communication mechanisms in accordance with the functions described herein. Figure 2 shows three such agent devices, namely agent devices 230A, 230B, and 230C, but it should be understood that any number may exist.

[0031] With respect to the multimedia / social media server 234, it may be configured to facilitate (non-voice) media interactions with customer devices 205 and / or server 242. Such media interactions may relate to, for example, email, voicemail, chat, video, text messaging, the web, social media, co-browsing, etc. The multimedia / social media server 234 may take the form of any IP router conventional in the art, having dedicated hardware and software for receiving, processing, and transmitting multimedia events and communications.

[0032] With respect to the knowledge management server 234, it may be configured to facilitate interaction between the customer and the knowledge system 238. Generally, the knowledge system 238 may be a computer system that can receive questions or queries and provide answers accordingly. The knowledge system 238 may be included as part of the contact center system 200 or operated remotely by a third party. The knowledge system 238 may include an artificial intelligence computer system that can answer questions presented in natural language by retrieving information from sources such as encyclopedias, dictionaries, newswire articles, literary works, or other documents submitted to the knowledge system 238 as reference material, as is known in the art. In embodiment, the knowledge system 238 may be embodied as IBM Watson or a similar system.

[0033] Regarding the chat server 240, it may be configured to conduct, orchestrate, and manage electronic chat communications with customers. In general, the chat server 240 The system is configured to implement and maintain chat conversations and generate chat transcripts. Such chat communications may be conducted by the chat server 240 in a manner such that a customer communicates with an automated chatbot, a human agent, or both. In an exemplary embodiment, the chat server 240 may function as a chat orchestration server that dispatches chat conversations between chatbots and available human agents. In such a case, the processing logic of the chat server 240 may be rules that are thus driven to leverage intelligent workload distribution among available chat resources. The chat server 240 may further implement, manage, and facilitate UIs related to chat functionality, including their user interfaces (also referred to as UIs) generated on either the customer device 205 or the agent device 230. The chat server 240 may be configured to forward chat between automated and human sources, for example, so that within a single chat session with a particular customer, the chat session is forwarded from a chatbot to a human agent, or from a human agent to a chatbot. The chat server 240 may also be coupled with the knowledge management server 234 and the knowledge system 238 to receive suggestions and answers to inquiries presented by the customer during the chat, for example, links to relevant articles may be provided.

[0034] Regarding web server 242, such servers allow customers to subscribe to various social media platforms such as Facebook, Twitter, and Instagram. It may be included to provide site hosting for the contact site. Although shown as part of the contact center system 200, it should be understood that the web server 242 may be provided by a third party and / or maintained remotely. The web server 242 may also provide web pages of a company or organization supported by the contact center system 200. For example, a customer may browse a web page to receive information about a particular company's products and services. Within such a company's web page, a mechanism may be provided for initiating interaction with the contact center system 200, for example, via web chat, voice, or email. An example of such a mechanism is a widget that may be deployed on a web page or website hosted on the web server 242. As used herein, a widget refers to a user interface component that performs a particular function. In some implementation examples, a widget may include a graphical user interface control that can be overlaid on a web page displayed to a customer over the Internet. A widget may include a button or other control that allows the user to access a particular function, such as displaying information in a window or text box, or sharing or opening a file, or initiating communication. In some implementations, widgets include user interface components with portable portions of code that can be installed and executed within a separate web page without compilation. Some widgets may include corresponding or additional user interfaces and may be configured to access various local resources (e.g., calendar or contact information on a customer's device) or remote resources over a network (e.g., instant messaging, email, or social networking updates).

[0035] With respect to the Interaction (iXn) server 244, it may be configured to manage deferred activities in a contact center and the routing of those activities to human agents for completion. As used herein, deferred activities include back-office work that can be performed offline, such as responding to emails, participating in training, and other activities that do not involve real-time communication with customers. In an embodiment, the Interaction (iXn) server 244 may be configured to interact with the routing server 218 to select an agent suitable for handling each of the deferred activities. Once assigned to a particular agent, the deferred activity is pushed to that agent so that it appears on the agent device 230 of the selected agent. The deferred activity may appear in the workbin 232 as a task for the selected agent to complete. The functionality of the workbin 232 may be implemented via any conventional data structure, such as a linked list or an array. Each of the agent devices 230 may include a workbin 232 with workbins 232A, 232B, and 232C maintained in agent devices 230A, 230B, and 230C, respectively. In one embodiment, the workbin 232 may be maintained in the buffer memory of the corresponding agent device 230.

[0036] With respect to the universal contact server (UCS) 246, it may be configured to retrieve information stored in the customer database 222 and / or transmit information therefor to be stored there. For example, UCS 246 may be used as part of a chat function to facilitate maintaining a history of how chats with a particular customer were handled, and this history may then be used as a reference for how future chat communications should be handled. More generally, UCS 246 may be configured to facilitate maintaining a history of customer preferences, such as preferred media channels and best times to contact. To do this, UCS 246 may be configured to identify data related to each customer's interaction history, such as data on comments from agents, customer communication history, etc. Each of these data types may then be stored in the customer database 222 or other modules and retrieved when required by the functions described herein.

[0037] With respect to the reporting server 248, it may be configured to generate reports from data compiled and aggregated by the statistics server 226 or other sources. Such reports may include quasi-real-time or historical reports and may relate to the state of contact center resources and performance characteristics, such as average latency, discard rate, and agent occupancy. Reports may be generated automatically or in response to specific requests from requesters (e.g., agents, administrators, contact center applications). The reports may then be used to manage the operation of the contact center in accordance with the functions described herein.

[0038] With respect to the media service server 249, it may be configured to provide audio and / or video services to support contact center functions. According to the functions described herein, such functions include prompting for IVR or IMR systems (e.g., playing audio files), hold music, voicemail / single-party recording, multi-party recording (e.g., audio and / or video calls), speech recognition, dual-tone multi-frequency (DTMF) recognition, and more. Audio and video transcoding, secure real-time transport protocol (SRTP), teleconferencing, video conferencing, This may include coaching (for example, support for coaches to overhear interactions between clients and agents, and for coaches to provide comments to agents without the client hearing them), call analysis, keyword spotting, etc.

[0039] With respect to the analysis module 250, it may be configured to provide a system and method for performing analysis on data received from multiple different data sources, where the functions described herein may be required. According to exemplary embodiments, the analysis module 250 may also generate, update, train, and modify a predictor or model 252 based on collected data, such as customer data, agent data, and interaction data. Model 252 may include a customer or agent behavior model. The behavior model may be used to predict customer or agent behavior in various situations, for example, thereby enabling embodiments of the invention to adjust interactions based on such predictions or allocate resources in preparation for predicted characteristics of future interactions, thereby improving overall contact center performance and customer experience. Although the analysis module 250 is shown as part of a contact center, it will be understood that such a behavior model may also be implemented in a customer system (or the "customer side" of an interaction, as used herein) and used for the benefit of the customer.

[0040] According to exemplary embodiments, the analysis module 250 may have access to data stored in the storage device 220, including a customer database 222 and an agent database 223. The analysis module 250 may also have access to an interaction database 224 that stores data related to interactions and interaction content (e.g., transcripts of interactions and events detected therein), interaction metadata (e.g., customer identifier, agent identifier, medium of interaction, length of interaction, start and end times of interaction, department, tagged category), and application settings (e.g., interaction paths through the contact center). Furthermore, as will be discussed further below, the analysis module 250 may be configured to retrieve data stored in the storage device 220 for use in developing and training algorithms and models 252, for example, by applying machine learning techniques.

[0041] One or more of the included Model 252 may be configured to predict customer or agent behavior and / or aspects related to the operation and performance of a contact center. Furthermore, one or more of the Model 252 may be used for natural language processing, including, for example, intent recognition. A Model 252 may be developed based on 1) known first-principles equations describing a system, 2) data resulting from an empirical model, or 3) a combination of known first-principles equations and data. When developing a model for use in this embodiment, it may generally be preferable to build an empirical model based on collected and stored data, since first-principles equations are often unavailable or not easily derived. It may be preferable for Model 252 to be nonlinear in order to appropriately capture the relationship between the instrumental / disturbance variables and the control variables of a complex system. This is because a nonlinear model may show a curvilinear relationship, rather than a linear relationship, between the instrumental / disturbance variables and the control variables, which is common in complex systems such as those discussed herein. Given the requirements described above, machine learning or neural network-based approaches are currently preferred embodiments for implementing Model 252. For example, neural networks can be developed based on empirical data using sophisticated regression algorithms.

[0042] The analysis module 250 may further include an optimizer 254. As can be understood, the optimizer can be used to minimize the “cost function” object to which a set of constraints is applied, where the cost function is a mathematical representation of the desired objective or system behavior. Since Model 252 may be nonlinear, the optimizer 254 may be a nonlinear programming optimizer. However, the present invention is intended to be implemented by using, individually or in combination, various different types of optimization approaches, including, but not limited to, linear programming, quadratic programming, mixed-integer nonlinear programming, stochastic programming, global nonlinear programming, genetic algorithms, particle / swarm techniques, and the like.

[0043] According to some exemplary embodiments, Model 252 and Optimizer 254 may be used together within an optimization system 255. For example, the analysis module 250 may utilize the optimization system 255 as part of an optimization process in which the performance and behavior of the contact center are optimized, or at least enhanced. This may include embodiments relating to customer experience, agent experience, interaction routing, natural language processing, intent recognition, or other functions related to automated processes, for example.

[0044] The various components, modules, and / or servers in Figure 2 (and other figures included herein) may each include one or more processors that execute computer program instructions and interact with other system components to perform various functions described herein. Such computer program instructions may be stored in memory implemented using standard memory devices such as random-access memory (RAM), or on other non-temporary computer-readable media such as CD-ROMs or flash drives. While each function of a server is described as being provided by a particular server, those skilled in the art should understand that the functions of various servers may be combined or integrated into a single server, or that the functions of a particular server may be distributed across one or more other servers without departing from the scope of the invention. Furthermore, the terms “interaction” and “communication” are used interchangeably and generally refer to any real-time and non-real-time interaction using any communication channel, including but not limited to telephone calls (PSTN or VoIP calls), email, V-mail, video, chat, screen sharing, text messages, social media messages, WebRTC calls, etc. Access to and control of the components of the contact system 200 may be affected through a user interface (UI) that may be generated on the customer device 205 and / or agent device 230. As already mentioned, the contact center system 200 may be operated as a hybrid system in which some or all of its components are hosted remotely in a cloud-based environment or a cloud computing environment.

[0045] Chat system Referring to Figures 3, 4, and 5, various embodiments of chat systems and chatbots are shown. As can be understood, these embodiments generally include, or may be enabled by, such chat functionality that enables the exchange of text messages between different parties. These parties may include living people such as customers and agents, as well as automated processes such as bots or chatbots.

[0046] As background, bots (also known as “Internet bots”) are software applications that perform automated tasks or scripts over the internet. Typically, bots perform simple, structurally repetitive tasks at a much higher rate than possible for humans. A chatbot is a specific type of bot, defined as part of software and / or hardware that engages in conversation via auditory or textual means. To be understood, chatbots are designed to convincingly simulate how a human would behave as a conversational partner. Chatbots are typically used in dialogue systems for a variety of practical purposes, such as customer service or information retrieval. Some chatbots use sophisticated natural language processing systems, while simpler chatbots scan keywords in the input and then select a response from a database based on matching keywords or phrasing patterns.

[0047] Before continuing the description of the present invention, note the reference system components described in any of the previous figures, such as modules, servers, and other components. Where referred to hereafter, whether or not the corresponding numerical identifiers used in the previous figures are included, the references are intended to encompass the examples described in the previous figures and, unless particularly specifically limited, may be implemented according to either those embodiments or other prior art capable of fulfilling the desired functionality as understood by those skilled in the art. Thus, for example, hereafter referred to as “Contact Center System” should be understood as referring to the exemplary “Contact Center System 200” in Figure 2 and / or other prior art for implementing a Contact Center System. As an additional example, hereafter referred to as “Customer Device,” “Agent Device,” “Chat Server,” or “Computing Device” should be understood as referring to the exemplary “Customer Device 205,” “Agent Device 230,” “Chat Server 240,” or “Computing Device 200” in Figures 1 and 2, respectively, and prior art for fulfilling the same functionality.

[0048] Next, chat functionality and chatbots will be described in more detail with reference to exemplary embodiments of the chat server, chatbot, and chat interface shown in Figures 3, 4, and 5, respectively. While these examples are provided for chat systems implemented on the contact center side, such chat systems may also be used on the customer side of the interaction. Therefore, it should be understood that the exemplary chat systems in Figures 3, 4, and 5 can be modified for similar customer-side implementations, including the use of customer-side chatbots configured to interact with contact center agents and chatbots on behalf of customers. It should also be understood that chat functionality may be utilized via voice communication through text-to-speech and / or speech-to-text conversion.

[0049] Referring in particular to Figure 3, a more detailed block diagram of a chat server 240 that may be used to implement a chat system and its functions is provided. The chat server 240 may be connected to (i.e., electronically communicate with) a customer device 205 operated by the customer via a data communication network 210. The chat server 240 may be operated by a company as part of a contact center for implementing and organizing chat conversations with customers, for example, including both automated chat and chat with human agents. With respect to automated chat, the chat server 240 may host chat automation modules or chatbots 260A-260C (collectively referred to as 260) configured using computer program instructions to engage in chat conversations. Thus, the chat server 240 generally implements chat functions, including the exchange of text-based communication or chat communication between the customer device 205 and the agent device 230 or chatbot 260. As will be discussed in more detail below, the chat server 240 may include a customer interface module 265 and an agent interface module 266 for generating specific UIs on the customer device 205 and agent device 230, respectively, to facilitate chat functionality.

[0050] Each of the chatbots 260 can operate as an executable program that is launched on demand. For example, the chat server 240 can operate as the execution engine for the chatbots 260, similar to loading a VoiceXML file into a media server for interactive voice response (IVR) functionality. Loading and unloading can be controlled by the chat server 240, similar to how a VoiceXML script may be controlled in the context of interactive voice response. The chat server 240 may further provide means for acquiring and collecting customer data in a unified manner, similar to customer data capture in the context of IVR. Such data can be stored, shared, and used in subsequent conversations, whether it is from the same chatbot, different chatbots, agent chat, or even different media types. In an exemplary embodiment, the chat server 240 is configured to coordinate the sharing of data among various chatbots 260 when an interaction is transferred or migrated from one chatbot to another, or from one chatbot to a human agent. Data acquired during interaction with a specific chatbot may be transmitted along with a call request to a second chatbot or human agent.

[0051] In exemplary embodiments, the number of chatbots 260 may vary depending on the design and functionality of the chat server 240 and is not limited to the number illustrated in Figure 3. Furthermore, different chatbots may be created to have different profiles, which may be selected to match a particular chat or a particular customer's subject matter. For example, a particular chatbot's profile may include expertise to assist customers with a particular subject or communication style tailored to a particular customer's preferences. More specifically, one chatbot may be designed to handle a first communication topic (e.g., opening a new account at a business), and another chatbot may be designed to handle a second communication topic (e.g., technical support regarding products or services offered by a business). Alternatively, chatbots may be configured to use various dialects or slang, or may have various personalities or characteristics. The involvement of chatbots with profiles that cater to specific types of customers may enable more effective communication and results. Chatbot profiles may be selected based on information known about the other party, such as demographic information, interaction history, or data available on social media. The chat server 240 may host a default chatbot that is invoked when there is insufficient information about the customer to call a more specialized chatbot. Optionally, different chatbots may be customer-selectable. In an exemplary embodiment, the chatbot 260's profile may be stored in a profile database hosted within the memory device 220. Such a profile may include the chatbot's personality, demographics, areas of expertise, etc.

[0052] The customer interface module 265 and the agent interface module 266 may be configured to generate a user interface (UI) on the customer device 205 to facilitate chat communication between the customer and the chatbot 260 or a human agent. Similarly, the agent interface module 266 may generate a specific UI on the agent device 230 to facilitate chat communication between the agent operating the agent device 230 and the customer. The agent interface module 266 may also generate a UI on the agent device 230 that allows the agent to monitor the status of an ongoing chat between the chatbot 260 and the customer. For example, during a chat session, the customer interface module 265 may transmit a signal to the customer device 205 configured to generate a specific UI on the customer device 205, which may include the display of text messages being sent from the chatbot 260 or a human agent, as well as other non-text graphics intended to accompany the text messages, such as emoticons or animations. Similarly, during a chat session, the agent interface module 266 may transmit a signal to the agent device 230 configured to generate a UI on the agent device 230. Such a UI could include an interface that facilitates agent selection of non-text graphics accompanying text messages sent to customers.

[0053] In an exemplary embodiment, the chat server 240 may be implemented in a layered architecture using a media layer, a media control layer, and a chatbot executed by the IMR server 216 (similar to the execution of VoiceXML on the IVR media server). As described above, the chat server 240 may be configured to interact with the knowledge management server 234 and query the server for knowledge information. For example, the query may be based on a question received from the customer during the chat. The response received from the knowledge management server 234 may then be provided to the customer as part of the chat response.

[0054] Referring specifically to Figure 4, a block diagram of an exemplary chat automation module or chatbot 260 is provided. As shown, the chatbot 260 may include several modules, including a text analysis module 270, a dialogue manager 272, and an output generator 274. In a more detailed description of the chatbot's functionality, it will be understood that other subsystems or modules may be described, for example, modules related to intent recognition, text-to-speech or speech-to-text conversion modules, and modules related to script storage, retrieval, and data field processing according to information stored in agent or customer profiles. However, such topics are covered more completely in other areas of this disclosure (for example, in relation to Figures 6 and 7) and are therefore not repeated here. Nevertheless, it should be understood that disclosures made in those areas may be used in a similar manner to the functionality of the chatbot according to the functions described herein.

[0055] The text analysis module 270 may be configured to analyze and understand natural language. In this regard, the text analysis module may consist of a language dictionary, a syntactic parser, a semantic parser, and grammatical rules for dividing phrases provided by the customer device 205 into syntactic and semantic internal representations. The configuration of the text analysis module depends on the specific profile associated with the chatbot. For example, certain words may be included in the dictionary of one chatbot but excluded from the dictionary of another chatbot.

[0056] The Dialogue Manager 272 receives syntactic and semantic expressions from the Text Analysis Module 270 and manages the general flow of the conversation based on a set of decision rules. In this regard, the Dialogue Manager 272 maintains the history and state of the conversation and generates outbound communications based on them. Communications may follow a script of a specific conversation path selected by the Dialogue Manager 272. As will be described in more detail below, conversation paths may be selected based on an understanding of the specific purpose or topic of the conversation. The script of a conversation path may be in the relevant technical field, such as Artificial Intelligence Markup Language (AIML), SCXML, etc. It can be generated using any of the various conventional languages ​​and frameworks.

[0057] During a chat conversation, the Dialogue Manager 272 selects a response deemed appropriate at a specific point in the conversation flow / script and outputs the response to the Output Generator 274. In an exemplary embodiment, the Dialogue Manager 272 may also be configured to computer-process the confidence level of the selected response and provide the confidence level to the Agent Device 230. Every segment, step, or input in the chat communication may have a corresponding list of possible responses. Responses may be categorized based on the topic (determined using a suitable text analysis and topic detection scheme) and assigned a proposed next action. Actions may include, for example, a response with an answer, an additional question, or transfer to a human agent to provide assistance. Confidence levels may be used to help the system determine whether the detection, analysis, and response to customer input are appropriate or whether a human agent should be involved. For example, a threshold confidence level may be assigned to prompt human agent intervention based on one or more business rules. In an exemplary embodiment, confidence levels may be determined based on customer feedback. As described, the responses selected by the Dialogue Manager 272 may include information provided by the Knowledge Management Server 234.

[0058] In an exemplary embodiment, the output generator 274 retrieves a semantic representation of the response provided by the dialogue manager 272, maps the response to a chatbot profile or personality (for example, by adjusting the language of the response according to the chatbot's dialect, vocabulary, or personality), and outputs output text to be displayed on the customer device 205. The output text may be presented intentionally so that the customer interacting with the chatbot does not realize that they are interacting with an automated process rather than a human agent. As can be understood, according to other embodiments, the output text may be linked to visual representations, such as emoticons or animations, that are incorporated into the customer's user interface.

[0059] Referring here to Figure 5, a webpage 280 having an exemplary implementation of the chat function 282 is presented. The webpage 280 may be related to, for example, a corporate website and may be intended to initiate interaction between prospective and current customers visiting the webpage and a contact center associated with the corporate. As can be understood, the chat function 282 may be generated on any type of customer device 205, including personal computing devices such as laptops, tablet devices, or smartphones. Furthermore, the chat function 282 may be generated as a window within the webpage or implemented as a full-screen interface. As in the illustrated example, the chat function 282 may be contained in a defined part of the webpage 280 and may be implemented as a widget, for example, through the systems and components described above and / or any other conventional means. In general, the chat function 282 may include exemplary methods for customers to enter text messages to deliver to the contact center.

[0060] As an example, the web page 280 may be accessed by a customer via a customer device, such as a customer device that provides a communication channel for chatting with a chatbot or live agent. In an exemplary embodiment, as shown, the chat function 282 includes generating a user interface on the display of the customer device, which is referred to herein as the customer chat interface 284. The customer chat interface 284 may be generated by a customer interface module of a chat server, such as a chat server already described. As described, the customer interface module 265 may send a signal to a customer device 205 configured to generate a desired customer chat interface 284 according to the content of a chat message originated by a chat source, which in this example is a chatbot or agent named "Kate". The customer chat interface 284 may be contained within a designated area or window, which occupies a designated portion of the web page 280. The customer chat interface 284 may also include a text display area 286, which is an area dedicated to the chronological display of received and sent text messages. The customer chat interface 284 further includes a text input area 288, which is a designated area where the customer enters the text for their next message. As you can see, other configurations are also possible.

[0061] Customer automation system Embodiments of the present invention include systems and methods for automating and enhancing customer actions during various stages of interaction with a customer service provider or contact center. As understood, these various stages of interaction may be classified as pre-contact, in-contact, and post-contact stages (or, respectively, pre-interaction, in-interaction, and post-interaction stages). Referring particularly to Figure 6, an exemplary customer automation system 300 that may be used with embodiments of the present invention is shown. To better illustrate how the customer automation system 300 works, also refer to Figure 7, which provides a flowchart 350 of an exemplary method for automating customer actions when a customer interacts with a contact center, for example. Further information regarding customer automation is provided in U.S. Patent Application No. 16 / 151,362, “System and Method for Customer Experience Automation,” filed on 4 October 2018.

[0062] The customer automation system 300 in Figure 6 represents a system that may be commonly used for customer-side automation, where used herein, referring to the automation of actions taken on behalf of a customer in an interaction with a customer service provider or contact center. Such interactions may also be referred to as “customer-contact center interactions” or simply “customer interactions.” Furthermore, when discussing such customer-contact center interactions, it should be understood that references to “contact center” or “customer service provider” are generally intended to refer to any customer service department or other service provider associated with an organization or enterprise (e.g., a business, government agency, non-profit organization, school, etc.) with which the user or customer has business, transaction, operations, or other vested interests.

[0063] In exemplary embodiments, the customer automation system 300 may be implemented as a software program or application running on a mobile device or other computing device, a cloud computing device (e.g., a computer server connected to the customer device 205 via a network), or a combination thereof (for example, some modules of the system may be implemented in a local application, while others may be implemented in the cloud. For convenience, embodiments will be described primarily in the context of implementation via an application running on the customer device 205. However, it should be understood that embodiments are not limited thereto.

[0064] The customer automation system 300 may include several components or modules. In the example shown in Figure 6, the customer automation system 300 includes a user interface 305, a natural language processing (NLP) module 310, and intent inference. The system includes module 315, script storage module 320, script processing module 325, customer profile database or module (or simply "customer profile") 330, communication manager module 335, text-to-speech conversion module 340, speech-to-text conversion module 342, and application programming interface (API) 345, each of which will be described in more detail with reference to the flowchart 350 in Figure 7. It will be understood that some of the components of the customer automation system 300 and some of the functions associated with them may overlap with the chatbot system described above in relation to Figures 3, 4, and 5. If the customer automation system 300 and such a chatbot system are adopted together as part of a customer-side implementation, such overlap may include the sharing of resources between the two systems.

[0065] In one example of operation, referring specifically to flowchart 350 in Figure 7, the customer automation system 300 may receive input in the first step or operation 355. Such input may come from several sources. For example, the primary source of input may be the customer, and such input may be received via the customer device. The input may also include data received from other parties, in particular from parties that interact with the customer through the customer device. For example, information or communications sent from a contact center to the customer may provide a form of input. In any case, the input may be provided in the form of free speech or text (e.g., unstructured natural language input). The input may also include other forms of data received or stored on the customer device.

[0066] Continuing with the flowchart 350, in operation 360, the customer automation system 300 parses the natural language of the input using the NLP module 310 and infers the intent from there using the intent inference module 315. For example, if the input is provided as speech from a customer, the speech may be rewritten into text by a speech-to-text system (such as a large-vocabulary continuous speech recognition, i.e., an LVCSR system) as part of the parsing by the NLP module 310. The rewriting may be performed locally on the customer device 205, or the speech may be transmitted over a network for conversion to text by a cloud-based server. In certain embodiments, for example, the intent inference module 315 may automatically infer the customer's intent from the text of the provided input using artificial intelligence or machine learning techniques. Such artificial intelligence techniques may include, for example, identifying one or more keywords from the customer input and searching a database of potential intents corresponding to given keywords. The database of potential intents and keywords corresponding to intents may be automatically mined from a collection of historical interaction records. If the customer automation system 300 cannot understand the intent from the input, several intent options may be provided to the customer in the user interface 305. The customer can then clarify their intent by selecting one of the alternatives, or request that other alternatives be provided.

[0067] After the customer intent is determined, the flowchart 350 proceeds to action 365, and the customer automation system 300 loads a script associated with the given intent. Such a script may be stored in, for example, a script storage module 320 and retrieved from there. Such a script may include a set of commands or actions, prewritten speech or text, and / or parameter or data fields (also called “data fields”) that represent data for which an action is requested to be automated on behalf of the customer. For example, a script may include commands, text, and data fields required to solve a problem specified by the customer intent. A script may be specific to a particular contact center and may be tailored to solve a particular problem. Scripts can be organized in several ways, for example, hierarchically, where all scripts related to a particular organization are derived from a common “parent” script that defines common characteristics. Scripts can be created by mining data, actions, and dialogues from previous customer interactions. Specifically, a series of statements made during a request to solve a particular problem can be automatically mined from a collection of historical interactions between the customer and the customer service provider. A system and method for automatically mining valid sequences of statements and comments, such as those written by a contact center agent, may be employed, as described in U.S. Patent Application No. 14 / 153,049, “Computing Suggested Actions in Caller Agent Phone Calls By Using Real-Time Speech Analytics and Real-Time Desktop Analytics,” filed with the U.S. Patent and Trademark Office on 12 January 2014.

[0068] Once a script is retrieved, flowchart 350 proceeds to action 370, in which the customer automation system 300 processes or "loads" the script. This action may also be performed by script processing module 325, which does so by populating the script's data fields with appropriate data about the customer. More specifically, script processing module 325 may extract customer data relevant to an expected interaction, the relevance of which is predetermined by the script selected as corresponding to the customer's intent. Data for many of the data fields in the script may be automatically loaded along with data retrieved from data stored in customer profile 330. As can be understood, customer profile 330 may store specific data related to the customer, such as the customer's name, date of birth, address, account number, credentials, and other types of information related to customer service interactions. The data selected to be stored in customer profile 330 may be based on data used by the customer in previous interactions and / or may include data values ​​directly obtained by the customer. If there is ambiguity regarding data fields or missing information in the script, the script processing module 325 may include a function to prompt and enable the customer to manually enter the required information.

[0069] Referring again to flowchart 350, in operation 375, the loaded script may be sent to a customer service provider or contact center. As will be discussed further below, the loaded script may include commands and customer data necessary to automate at least part of the interaction with the contact center on behalf of the customer. In an exemplary embodiment, API 345 is used to interact directly with the contact center. The contact center may define a protocol for making common requests to their system, and API 345 is configured to do this. Such an API uses Simple, which uses Extended Markup Language (XML). It can be implemented via various standard protocols, such as Object Access Protocol (SOAP), Representational State Transfer (REST) ​​APIs with messages formatted using XML or JavaScript Object Notation (JSON). Thus, the customer automation system 300 can automatically generate messages formatted according to a defined protocol for communication with the contact center, and the messages will include information specified by the script in the appropriate parts of the formatted message.

[0070] Bot authoring using intent mining automation In recent years, artificial intelligence (AI) and computing technologies have With several breakthroughs in this area, interest in applications, automation systems, chatbots, or bots capable of engaging in natural language conversations with humans is growing. In recent years, we have witnessed significant growth in the adoption of AI-powered chatbots and virtual assistants that can converse naturally with humans and perform a wide variety of tasks in a self-service manner. Such conversational bots work by first analyzing the user's input and then attempting to understand the meaning of that input. This is called Natural Language Understanding (or "NLU") and typically involves identifying the user's intent and specific keywords or entities within the user's input utterance. Once the intent and entities are determined, the bot can respond to the user with appropriate follow-up actions.

[0071] Various machine learning algorithms are used to train NLU models. Training typically involves teaching the system to recognize patterns present in natural language input and associate them with a predefined set of intentions. The quality of the training data is a crucial factor in determining model performance. A sufficiently large dataset with adequate diversity in the input utterances is essential for building a good NLU model.

[0072] As used herein, the term “bot authoring” refers to the process of creating a conversational bot or chatbot with NLU capabilities. This process generally involves defining intent, identifying entities, formulating utterances, training an NLU model, testing the bot, and finally publishing it. This is a largely manual process that can typically take weeks or even months to complete. Generally, identifying intent and formulating utterances takes up the majority of this time. Organizations may already possess a large volume of chat conversations between their customers and customer support staff, such as contact center agents, but the process of manually examining these raw chat transcripts to identify intent and utterances is both time-consuming and costly.

[0073] As used herein, an intent mining engine or process (which may generally be referred to as an “intent mining process”) is a system or method that makes a bot authoring workflow more efficient. As understood, the intent mining process of the present invention works by mining intent from tens of thousands of conversations, finding a robust and diverse set of utterances to which each belongs. Furthermore, the intent mining process helps to gain insights into conversations by providing conversation analysis. It also provides bot authors with the opportunity to analyze and modify intent. Finally, these intents and utterances may be exported to a variety of chatbot authoring platforms, such as those commercially available on Genesys Dialog Engine, Google Dialogflow, and Amazon Lex. As understood, this results in a flexible and efficient bot authoring workflow that significantly reduces overall development time.

[0074] Referring next to Figure 8, various stages or steps of a bot authoring workflow 400 using the intent mining process of the present invention (or simply “the intent mining process”) are shown. To start the workflow 400, a conversation or conversational data may be imported for mining. Such conversational data may consist of previously occurring interactions between an agent and a customer. Such conversational data may be a natural language conversation consisting of multiple round-trip messaging turns. The conversation may take place, for example, via a chat interface, through text, or via a voice call. In the latter case, the conversation may be transcribed into text via speech recognition before mining begins.

[0075] In the first step 405, the bot authoring workflow 400 may include importing conversational data (i.e., conversational text data) for use in the intent mining process. This can be done in several ways. For example, conversational data may be imported via a text file (in a supported format such as JSON) containing the conversations to be mined. Conversational data may also be imported from cloud storage.

[0076] In step 410, the bot authoring workflow 400 may include mining intent from conversational data. Intent may be mined according to an intent mining algorithm, as described in relation to Figures 9 and 10 below.

[0077] In step 415, the bot authoring workflow 400 may include testing the mined intent. This may include interacting with the output of the intent mining process. That is, at this stage of the workflow, the bot author interacts with the mined output to make edits, which may include fine-tuning and pruning the intent and associated utterances before exporting them to the bot for training. The bot author can perform various actions on the mined output, such as selecting intents and utterances belonging to those intents, merging two or more intents into a single intent to produce a merged set of those selected utterances, splitting an intent into multiple intents to produce a split set of corresponding utterances, and renaming intent labels. At the end of this business logic-driven process, a set of modified intents and associated utterances is created, which can then be used to train a chatbot.

[0078] In step 420, the bot authoring workflow 400 may include importing mined intent and utterances into the bot. For example, the mined intent may be uploaded to a conversational bot, which may be used to conduct automated conversations with customers. This intent mining process can provide multiple methods for adding the mined or modified intent and utterances to the bot. The data may be downloaded in CSV format for convenient review. The data may also be exported to multiple bot formats, thus providing support for a wider variety of conversational AI chatbot services, such as Genesys Dialog Engine, Google Dialogflow, or Amazon Lex.

[0079] The bot authoring process may also include additional steps. According to some embodiments, this intent mining process may be significantly involved in the steps already described above and less involved in the later development stages. These later steps may include optional editing steps, bot design steps, and finally, final testing and publication steps.

[0080] Referring next to Figure 9, an exemplary algorithm for implementing the intent mining engine or process 500 is described below. As can be understood, this algorithm may be roughly broken down into several steps, which are referred to herein as 1) identifying intentful utterances, 2) generating candidate intents, 3) identifying prominent intents, 4) semantic grouping of intents, 5) intent labeling, and 6) utterance-intent association. Other steps may include masking personally identifiable information within utterances. Another additional step may include computer processing of intent analysis. These steps are discussed below. As can be understood, the steps are described in relation to imported conversational data, for example, data containing natural language conversations between customers interacting with customer service representatives or agents, but it should be understood that the process may also be applicable to other contexts with other types of users and conversation types.

[0081] In accordance with the first step 505, the intent mining process processes conversational data to identify intentful turns or utterances. As used herein, an intentful utterance is an utterance that is determined to be likely to contain or describe a customer's intent. Thus, this first step in the intent mining process is to identify intentful utterances from a given conversation. For example, a conversation typically consists of multiple message turns or utterances from multiple parties, such as an agent (which may include an automated system or bot or a human agent) and a customer.

[0082] For example, a bot-generated message might look like this: "Hello, thank you for contacting us. All chats may be monitored or recorded for quality and training purposes. We will be with you." "Shortly to help you with your request." Such bot-generated messages tend to be generic and fail to shed light on the intent behind the conversation, so they can be safely discarded. An actual conversation begins with either the agent or the customer sending a substantial communication or message. For example, during an interaction, the customer may explain their reason or "intent" for contacting customer care. The subsequent agent-customer conversation turn is based on this intent expressed by the customer.

[0083] Based on the analysis of real-world customer-agent conversations, the present invention includes several heuristics or strategies for identifying intentful utterances. For example, intentful turns have been observed to typically occur towards the beginning of the customer's side of the conversation. Therefore, to identify intent, it is generally necessary to process only a few of the initial customer utterances, and the rest of the conversation can be discarded. This further helps to reduce system latency and memory footprint. Furthermore, word count constraints can be used to discard other utterances as less likely to contain customer intent.

[0084] For example, identifying intentful utterances may include: selecting a set of consecutive customer utterances in a conversation. This set may include customer utterances that occur at the beginning of the conversation. Furthermore, a word count constraint may be used to disqualify some of the customer utterances in this initial set. That is, for an utterance to be eligible, the number of words in each turn must be greater than a minimum threshold. Such word count or length constraints help discard some customer turns that are irrelevant to intent mining purposes, such as conventional greetings like "Hello," "Hi there," and "How are you?". For example, this minimum word count threshold may be set to 2-5.

[0085] This intent mining process can concatenate utterances from consecutive customer turns containing intent into a single combined utterance. Before this is done, each customer turn may be pruned based on a maximum length threshold, as longer sentences tend to produce inconsistent or noisy results. For example, the maximum number of words per utterance can be set to 50 words. Thus, at the end of this step, a combined utterance is obtained from each conversation that is likely to contain the intent expressed by the customer. If a conversation does not contain a message turn that meets the above criteria, that conversation may be discarded without obtaining a combined utterance from it. Since this intent mining process is used to obtain dominant intent from hundreds or thousands of conversations, it can be safely assumed that customer intent is repeated across multiple conversations. Therefore, conversations that do not meet the above heuristic criteria may be discarded without affecting the system's functionality for greater robustness in intent identification.

[0086] In accordance with the second step 510, candidate intentions are generated based on an analysis of the combined utterances. That is, once utterances from intention-bearing turns are taken from a conversation and combined, the next task includes identifying possible or likely intentions, which are referred to herein as “candidate intentions.” As used herein, a candidate intention is a text phrase consisting of two parts: 1) an action, which is a word or phrase representing a tangible purpose, task, or activity; and 2) an object, which is a word or phrase on which the action acts or operates.

[0087] There are various ways to obtain these action-object pairs from utterances. As you may understand, the choice may depend on the language models and resources available for a particular language. Typically, for example, a syntactic dependency parser. This method analyzes the grammatical structure of an utterance and retrieves the relationships between "head" words and "tokens" or words that modify these heads. These relationships between the tokens of an utterance and their heads, along with their part-of-speech (POS) tags, are used to identify potential or candidate intentions for a given utterance.

[0088] As an example, the process of obtaining such action-object pairs may include the following: First, all token-head pairs in the utterance can be obtained using a dependent parser. From these, pairs are selected that have the token's POS tag and its associated head, which is a noun and a verb, respectively. The use of universal POS tags helps to make the system language agnostic and therefore extensible to multiple language domains.

[0089] The "action" portion is typically a token that has the "verb" as its associated POS tag. If the token is a "basic verb" with a "particle" token, the token forms a "phrasal verb" of the utterance. The associated "particle" token is also included along with the verb token. Therefore, the entire phrasal verb becomes the action portion of the candidate intent. The "object" portion is typically a token that has the "noun" as its associated POS tag. If the token is part of a "compound" and all its constituent tokens have the "noun" POS tag, the entire compound is considered an object. Similarly, if the token is part of an adjectival phrase, the entire phrase is considered an object. If the token is associated with a juxtaposition modifier, all tokens constituting the latter are appended to the current token to form the object portion of the candidate intent. If only universal POS tags are available for the language and universal dependencies are not available, "verb" tokens and "noun" tokens are considered the action portion and object portion, respectively. As a next step, action-object ordered pairs may be lemmatized to translate the candidate intent into a more standard form. For further normalization, the number of cases in the lemmatized pair can be reduced.

[0090] Therefore, one or more normalized action-object pairs can be obtained from each utterance, which together form a candidate intent for the conversation. If no such pairs are obtained, the utterance is discarded. With this in mind, consider the following first example utterance: "I'm looking to contact the instructor for this course. Can you provide "His email please?" In this case, possible intentions could include "contact instructor" and "provide email." Consider the following second example utterance: "I just finished my bachelor's program yesterday on my account; it says you must complete a graduation application, but when I click it, it goes to a page that says messages and only shows potential scholarships. What should I do?" In this case, possible intentions could include "finish program," "complete graduation application," "say message," and "show potential scholarship."

[0091] In accordance with the third step 515, the prominent intent is identified. As used herein, the term “prominent intent” refers to a refined list of intents from the candidate intents identified in the previous step, the refinement being based, for example, on relevance, importance, certainty, and / or attention. Thus, from the set of candidate intents, the intent that describes the customer’s actual intent is identified as the prominent intent. As you may understand, this task is not always straightforward. In some cases, the customer’s intent may be implicit in nature. However, in other cases, especially in utterances containing multiple candidate intents, there may be differing opinions regarding the customer’s actual intent.

[0092] Consider the examples provided above. In the first example utterance, it can be argued that both "contact instructor" and "provide email" describe the customer's intent. In the second example utterance, the customer is facing a problem while completing their undergraduate program and finalizing their graduation application. While this intent is more implicit, the closest explicit approximation might be the candidate intent "complete graduation application." The decision of whether "contact instructor" or "provide email" should be selected as the intent for the first utterance, or even whether "finish program" or "complete graduation application" should be selected as the intent for the second utterance, may be better determined by business logic than by any algorithmic formulation. That is, a bot author can apply appropriate business logic to arrive at a final decision regarding such intents. A bot author may also choose to maintain multiple intents, or even describe a hierarchy of intents, in order to achieve appropriate business objectives or goals within a particular business domain.

[0093] Since the objective is to make the bot authoring process more efficient, this intent mining process can narrow down the list of candidate intents to the most prominent ones, which the bot author can then review for relevance. In such cases, singularity can be defined in multiple ways based on different criteria. For example, according to an exemplary embodiment, the frequency of candidate intents in the entire set of utterances can be an indicator of singularity, i.e., the more candidate intents there are, the more relevant they are. According to another embodiment of the present invention, singular intents can be found using criteria based on Latent Semantic Analysis (LSA). LSA is a topic modeling technique used in natural language understanding (NLU) tasks. To do this, each utterance described in terms of candidate intent action-object pairs is considered a document. LSA then analyzes the relationship between these documents and the terms they contain (i.e., action-object pairs) by creating a set of concepts associated with the documents and the terms they contain. Each concept is described in terms of candidate intents with associated weights. These weights provide insight into the relative singularity of candidate intents within each group of concepts.

[0094] As an example, according to the present invention, the process for identifying significant intent may include the following: First, the LSA is applied to an utterance describing candidate intent action-object pairs, with the number of LSA components set to a predetermined limit, e.g., 50. Next, the candidate intents of each conceptual group are sorted in descending order with respect to their weights, and the top candidate intents, e.g., the top five, are selected. Next, the selected candidate intents obtained from each conceptual group are matched and sorted in descending order with respect to their weights. Next, duplicate entries are discarded, and entries with higher weights are retained. Then, a predetermined number of these may be considered significant candidate intents or simply “significant intents”. The predetermined number may be based on the maximum number of intents that need to be mined. For example, this maximum number of intents may be determined by the intent mining process based on real-world contact center interaction patterns, or it may be selected by a bot author based on appropriate business logic and use cases.

[0095] In the fourth step, according to step 520, the prominent intentions are semantically grouped. As can be seen, since only the syntactic structure of the utterances is used to generate candidate intentions, many of the prominent intentions identified by the system may be semantically similar. Therefore, semantically similar prominent intentions can be grouped together for optimal downstream functionality. The output of this intention mining process may also be used to train a Natural Language Understanding (NLU) model, which in this case effectively forms the "brain" of the natural language chatbot. For these models to identify intentions associated with diverse utterances, the NLU models must be trained with utterances that are syntactically different but semantically similar. Therefore, the bot authoring process must enable the creation of intentions associated with utterances that have appropriate diversity. Grouping semantically similar prominent intentions helps create this diversity in the mined intentions.

[0096] This step generally involves calculating the semantic similarity between significant intentions, which can be completed as follows, for example: First, embeddings or word embeddings associated with the texts of significant intentions are processed by computer. As understood, such embeddings represent the texts in question, e.g., words, phrases, or sentences, such that semantically similar texts have similar embeddings. Such word embeddings generally involve converting the text data into a numerical format through an encoding process, and such word embeddings can be extracted from the text data using various conventional encoding techniques. The embeddings can then be efficiently compared to determine a measure of semantic similarity between the texts. For example, global vectors (or "Global Vector") GloVe is an algorithm that can be used to obtain a vector representation of a word. A GloVe model may have, for example, 300 dimensions. In an exemplary embodiment, word embeddings of significant intents may be computer-processed using an inverse document frequency (IDF) weighted average of GloVe embeddings of the constituent tokens. As is understood, IDF is a numerical statistic that reflects a measure of whether a term is common or rare in a given document corpus. Used in this way, the set of all candidate intents or significant intents can be considered, here, as the document corpus for the purpose of IDF computer processing.

[0097] Once word embeddings for the text of the prominent intentions are obtained, these word embeddings can be used to calculate the semantic similarity between pairs of prominent intentions. For example, cosine similarity can be used to provide a measure of semantic closeness between word embeddings in a higher-dimensional space. Once this is obtained, the prominent intentions can then be grouped according to pairs of embeddings whose cosine similarity is greater than a predetermined similarity threshold, which can be set in the range of 0 to 1. As can be understood, a higher threshold will group less prominent intentions together, thereby creating more homogeneous groups, while a lower threshold will group more semantically diverse intentions together, creating less homogeneous groups. This homogeneity value may be preset in the system selected by the bot author (e.g., to 0.8), as in the case of selecting the maximum intention described above. In the latter case, the bot author may look at multiple combinations of output intentions and utterances and select a value that is appropriate for the best bot result.

[0098] In accordance with step 525 of the fifth step, intent labels are identified. Each of the grouped significant intents (or “significant intent groups”) can ultimately become a mined intent (or “mined intent”). Thus, for each of these significant intent groups, an intent label is selected to serve as the label or identifier for the mined intent. According to an exemplary embodiment, this labeling may be performed by computer-processing the IDF of each significant intent within a given significant intent group. For this calculation, utterances describing candidate intents are considered documents, and action-object pairs considered as single units are considered tokens that constitute the elements. The significant intent in each group with the highest calculated IDF is then made the intent sample or “intent label” for the group, while the other significant intents within the group are referred to as “intent substitutes”.

[0099] In accordance with the sixth step 530, utterances are associated with mined intentions (each mined intention, now reflected by intention labels and their respective prominent intention groups). As understood, this next step determines which utterances are associated with each mined intention. As in the previous step, semantic similarity techniques using embeddings can also be used here. For example, semantic similarity is computer-processed between candidate intentions derived from each intention-having utterance and each prominent intention within a given prominent intention group. An utterance is then associated with a given prominent intention group (which may also be called mined intentions or simply intentions) if it is determined that the similarity of any of its constituent candidate intentions is the highest for the prominent intentions of that prominent intention group and exceeds a minimum threshold (e.g., 0.8). Furthermore, with respect to each prominent intention group, the candidate intentions of the intention-having utterances that produced the highest similarity with each particular prominent intention group may be brought into that particular prominent intention group as an "intention aid." Here again, a minimum threshold may be required. Therefore, within this step, an utterance with a specific intention is associated with one of the prominent intention groups, and the candidate intentions that are components of that particular intention utterance are associated with each prominent intention group as intention auxiliary. Thus, each mined intention can include an intention label, as well as one or more intention surnames and / or one or more intention auxiliary, as described above. As can be understood, such a formulation does not prevent an utterance with a single intention from being associated with multiple intention groups. This is because an utterance with a single intention may have multiple candidate intentions that are added as intention auxiliary to different ones across multiple mined intentions. This introduces greater flexibility and robustness in the downstream functionality. The bot author can choose to retain or discard such utterances from one or more groups. It has been observed that utterances repeated across multiple intentions help instruct the NLU model about the inherent confusion present in them, and thus help build a more realistic and robust model.

[0100] In another step (not shown in the diagram), personally identifiable information within the utterance is removed or masked. To ensure customer privacy, all personally identifiable information present in the associated utterance is masked. Of course, this step can be omitted if the input conversation is anonymized before being provided to the intent mining process. Such personally identifiable information may include customer names, phone numbers, email addresses, and social security information. In addition, entities related to geographical location, dates, and numbers may be masked as an additional precaution. For example, in the utterance "Hi, I need to book a flight from Washington" Please consider "DC to Miami on August 15 under the name of John Honai." After masking, the utterance should be, "Hi, I need to book a flight from <geo> <geo>to <geo>on <date> <date>under the name of <person> <person>This could result in ".". In addition to protecting privacy, such masking can allow bot authors to quickly identify different entities present within an intent utterance. This can help bot authors create utterances that are similar but with varying slot values ​​for these entities. This leads to greater diversity in utterances, which in turn helps in creating better NLU models.

[0101] According to another possible step (not shown), intent analysis can be computerized. That is, apart from mining intent and associated utterances, this intent mining process can also generate analyses and metrics related to conversational data that help companies identify customer interaction patterns. Two such metrics are as follows:

[0102] The first analysis is intent volume analysis, which analyzes the extent to which a conversation addresses a particular intent. This analysis can also be expressed as a percentage. Intention volume analysis can help understand the relative importance of intents based on the frequency of intent occurrences within conversational data. Since only a single utterance is taken from each conversation, this metric is essentially the number of utterances belonging to each intent.

[0103] The second type of analysis is intent duration analysis, which is an analysis of the duration of conversations dealing with a particular intent. This analysis can also be expressed as a percentage. As you can see, this metric helps to compare intents based on the total conversation time associated with that intent. The time spent in a conversation is processed by computer as the difference between the last customer / agent timestamp and the first customer / agent timestamp. The sum of the durations of the individual conversations belonging to an intent gives the duration of that intent. As you can see, this type of analysis can help bot authors and businesses better understand customers and contact center staff.

[0104] This section describes a method for authoring a conversational bot and intent mining. This method may include receiving conversational data, wherein the conversational data includes text derived from conversations between customers and customer service representatives; automatically mining intent from the conversational data using an intent mining algorithm, wherein each mined intent includes intent labels, intent surrogates, and associated utterances; and uploading the mined intent to a conversational bot and using the conversational bot to conduct automated conversations with other customers.

[0105] According to an exemplary embodiment, an intent mining algorithm may include analyzing utterances occurring within a conversation in conversational data to identify intent utterances. Each utterance may contain a turn within the conversation, thereby communicating whether a customer is communicating in the form of customer utterances or a customer service representative is communicating in the form of customer service representative utterances. An intent utterance is defined as one of the utterances determined to be highly likely to express an intent. The intent mining algorithm may further include analyzing the identified intent utterances to identify candidate intents. Each candidate intent may be identified as a text phrase occurring within one of the intent utterances, having two parts: an action which may contain a word or phrase describing a purpose or task, and an object which may contain a word or phrase describing an object or thing on which the action operates. The intent mining algorithm may further include selecting prominent intents from the candidate intents according to one or more criteria. The intent mining algorithm may further include grouping the selected prominent intents into prominent intent groups according to the degree of semantic similarity between the prominent intents. An intent mining algorithm may further include, for each group of prominent intents, selecting one prominent intent as the intent label and designating the other prominent intents as intent substitutes. The intent mining algorithm may further include associating intent-having utterances with prominent intent groups by determining the degree of semantic similarity between candidate intents present in intent-having utterances and intent substitutes within each of the prominent intent groups. Each mined intent may include a given one of the prominent intent groups, each defined by one of the prominent intent groups selected as the intent label and the other prominent intents designated as substitute intents, and an intent-having utterance associated with that given one of the prominent intent groups.

[0106] According to an exemplary embodiment, the step of identifying an intentional utterance may include selecting a first portion of a customer utterance as an intentional utterance and discarding a second portion of the customer utterance in the conversation data. The first portion of a customer utterance may be defined as a predetermined number of consecutive customer utterances occurring at the beginning of each conversation, and the second portion may be defined as the remainder of each conversation.

[0107] According to an exemplary embodiment, the step of identifying an intent utterance may further include discarding customer utterances in a first portion of a customer utterance that do not satisfy a word count constraint. The word count constraint may include a minimum word count constraint, in which customer utterances in a first portion of a customer utterance having fewer words than the minimum word count constraint are discarded, and / or a maximum word count constraint, in which customer utterances in a first portion of a customer utterance having more words than the maximum word count constraint are discarded. The minimum word count constraint may take values ​​of 2 to 5 words. The maximum word count constraint may take values ​​of 40 to 50 words.

[0108] According to an exemplary embodiment, the step of identifying an intentional utterance may include concatenating customer utterances occurring within each first part of a conversation into a combined customer utterance.

[0109] According to an exemplary embodiment, the step of identifying candidate intentions may include: identifying head-token pairs by analyzing the grammatical structure of an intent utterance using a syntactic-dependent parser, wherein each head-token pair includes a head word modified by a token word; and identifying head-token pairs as candidate intentions by using part-of-speech (POS) tagging to tag the part of speech of an intent utterance, wherein the POS tag of the head word may include a noun tag and the POS tag of the token word may include a verb tag.

[0110] According to an exemplary embodiment, the step of selecting a significant intention from candidate intentions may include selecting those candidate intentions that are determined to appear more frequently than others in intent-containing utterances. One or more criteria for selecting a significant intention from candidate intentions may include criteria based on latent semantic analysis (LSA). The step of selecting a significant intention from candidate intentions may include generating a set of documents, each having a document corresponding to each of the candidate intentions, such that each document covers an action-object pair defined by one of the corresponding candidate intentions; generating conceptual groups based on terms appearing in the action-object pairs contained in the set of documents; calculating weight values ​​for each of the candidate intentions for each conceptual group, such that the weight values ​​measure the degree of relevance between a given candidate intention in the documents and a given one of the conceptual groups; and selecting a predetermined number of candidate intentions in each conceptual group as significant intentions and creating weight values ​​therefore indicating a higher degree of relevance.

[0111] According to an exemplary embodiment, the step of grouping significant intentions according to their degree of semantic similarity may include: calculating an embedding for each significant intention, wherein the embedding may include encoded representations of texts whose semantically similar texts have similar encoded representations; comparing the calculated embeddings to determine the degree of semantic similarity between pairs of significant intentions; and grouping significant intentions whose degree of semantic similarity exceeds a predetermined threshold. The embedding may be calculated as the inverse document frequency (IDF) mean of global vector embeddings of head-token pairs that are components of the significant intention. Comparing the calculated embeddings may include cosine similarity.

[0112] According to an exemplary embodiment, the step of labeling each of the groups of significant intentions with an intention identifier may include selecting one representative significant intention within each of the groups of significant intentions.

[0113] According to an exemplary embodiment, the step of associating utterances from conversational data with significant intention groups may include repeatedly performing a first process to cover each utterance having an intention in relation to each of the significant intention groups. When described in relation to an exemplary first case involving first and second significant intention groups and an utterance having a first intention that includes first and second candidate intentions, the first process may include: computer processing the degree of semantic similarity between each of the first and second candidate intentions and each of the intention substitutes in the first significant intention group; computer processing the degree of semantic similarity between each of the first and second candidate intentions and each of the intention substitutes in the second significant intention group; determining which of the intention substitutes produced the highest computer-processed degree of semantic similarity; and associating the utterance having a first intention with the first of the first and second significant intention groups that includes the intention substitute that was determined to produce the highest computer-processed degree of semantic similarity. The step of associating utterances from conversational data with significant intent groups may further include associating intent substitutes that produce the highest computer-processed semantic similarity, only if it is found that the highest computer-processed semantic similarity also exceeds a predetermined similarity threshold.

[0114] Referring next to Figure 10, the various stages of the alternative bot authoring workflow are shown, and the intent mining method disclosed above in relation to Figure 9 is extended by seeding of the intent process. The intent mining process using the intent seeding process will be discussed below after a brief introduction. To facilitate distinction between references, the intent mining process using intent seeding will hereafter be referred to as the “intent mining process by seeding” or simply “intent mining by seeding,” while the aforementioned process of intent mining disclosed above in relation to Figure 9 (i.e., intent mining without seeding) will hereafter be referred to as the “general intent mining process” or simply “general intent mining.”

[0115] In its normal operating mode, general intent mining mines intent for both intent labels and the associated utterances from conversational data, such as a set of agent-customer conversations. As already discussed, this process is guided by the syntactic structure and semantic content of the conversation. For example, syntactic dependencies and POS tags can be used to find candidate intents from intent-bearing utterances in a conversation, and methods such as latent semantic analysis (LSA) can be used to narrow down the prominent intents from the utterances. Intent labels, intent surnames, and intent aids are then obtained by associating semantically similar prominent intents, which helps link utterances to specific intents. As already disclosed, a bot author can use the data mined via general intent mining to train an NLU model, which then powers the conversational bot.

[0116] It will be understood that this general framework of intent mining proceeds from the assumption that bot authors are unaware of the intent typically present in conversational data, and / or that NLU models have not yet been developed for conversations or sets of similar conversational domains. Thus, a general intent mining process, for example, a process using the intent mining engine disclosed above, essentially starts without prior domain knowledge and derives or mines intent based solely on the conversational content of the data.

[0117] However, in many cases, this assumption does not hold true. That is, a bot author may already possess knowledge about intentions within a particular domain. In such cases, the bot author can understand intentions that are typically present in a particular conversation or that are expected to be present in a particular conversational domain. This may be true, for example, in the banking or travel domain. Furthermore, in many cases, the NLU model may already be trained, and bots such as travel or banking bots have already been published by the bot author. In such scenarios, the existing domain knowledge can be used to guide the intention mining process to mine specific intentions by employing an intention seeding process. As can be understood, as part of intention mining by seeding, the existing domain knowledge is supplied to the mining process in the form of seed intention data. Such seed intention data may consist of intention labels, which may be called “seed intentions” or “seed intention labels,” and sample utterances associated with each. The intention mining process then uses this seed intention data to mine more utterances from the conversational data for each of the seed intentions, while also finding utterances for any other prominent intentions that can be found in the conversational data. As described above, this mining process is referred to herein as the “seeding-based intent mining process” or simply “seeding-based intent mining.”

[0118] As can be understood, intent mining by seeding can help bot authors quickly identify more utterances belonging to a seed intent, which can be used to train or improve NLU models. Since such a system can mine other prominent intents in addition to a given seed intent, this process can help bot authors identify customer intents that change over different time frames.

[0119] Similar to the general intent mining methods described above, the intent seeding intent mining process can be initiated by importing conversation data. Generally, the other steps of intent mining by seeding may be the same as or similar to the steps disclosed above in relation to general intent mining. Therefore, for the sake of brevity, we will focus primarily on the areas in which intent mining by seeding differs from the general intent mining process presented above in relation to Figure 9.

[0120] According to the present invention, intent mining by seeding uses seed intent data. When used herein, seed intent data includes one or more seed intents and, for each of the one or more seed intents, a set of associated sample utterances. Intent mining by seeding then processes the seed intent data together with conversational data to obtain intent substitutes and / or other utterances to associate with the seed intents. Such intent substitutes are obtained in substantially the same manner as those used to generate candidate intents given in the above section. In this case, the seed intents and associated sample utterances are considered to be utterances with intents provided by the customer within the conversational data. The normalized action-object pairs obtained from them constitute the intent substitutes for each seed intent.

[0121] Once an intent substitute is obtained for each seed intent, a seed intent aid is identified from the set of candidate intents derived from the conversation data, as described above in relation to Figure 9. With respect to the steps of finding the seed intent aid and associating utterances with seed intents, the intent mining process by seeding may be the same as or similar to those described above in general intent mining processes. As in the previous section, semantic similarity techniques using embeddings can also be used here. Similarity can be processed computer-wise between each candidate intent of an intent-containing utterance and the intent substitute of each seed intent. An intent-containing utterance is associated with a seed intent if it is determined that a) the semantic similarity of any of the candidate intents that are components of the intent-containing utterance is the highest with respect to the intent substitute of that seed intent, and b) the semantic similarity exceeds a minimum threshold (e.g., a score above 0.8). Furthermore, as previously stated, the candidate intent that produces the highest similarity score with respect to one of the seed intents is brought into the seed intent as an "intent aid," or more specifically, as a "seed intent aid."

[0122] Intention mining by seeding may also include deriving other significant intentions found in conversation data that are different from the intentions identified in the seed intention data. With respect to identifying such significant intentions, the intention mining process by seeding may be the same as or similar to that described above in the general intention mining process. That is, candidate intentions are identified, then sorted by weight within a conceptual group, and a predetermined number of higher-weighted candidate intentions are selected from the group. Of these selected candidate intentions, duplicate entries are discarded, and entries with higher weights become the entries to be retained. In completing this step, the intention mining process by seeding may include additional steps from those disclosed above with respect to general intention mining. Specifically, these additional steps may include discarding any of the identified candidate intentions that have already been identified as seed intention aids.

[0123] The next step is to find utterances associated with the mined intentions. As in the previous section, a semantic similarity technique using embeddings is used here. Similarity is processed computer-wise between the utterance's candidate intentions and the intentions of all groups. An utterance is associated with an intention group if the similarity of any of its constituent candidate intentions is the highest for that group's intention and exceeds a minimum threshold (e.g., 0.8). The candidate intention that produced the highest similarity is brought into the group and is called the "intention helper". Candidate intentions that have already been identified as seed intention helpers are discarded from this run.

[0124] Given the capabilities discussed above regarding intent mining without seeding (i.e., the general intent mining process described in relation to Figure 9) and intent mining with seeding (i.e., the intent mining process with seeding described in relation to Figure 10), it should be understood that several different use cases or applications are possible. In the first case, intent mining is performed without seeding. This can be used to mine significant intents and their associated utterances from a given conversational data. The second case involves a hybrid case in which general intent mining and intent mining with seeding are performed. As understood, this case can be used on a given conversational data to mine both significant intents and associated utterances, as well as more utterances to associate with a given set of seed intents. In the third case, intent mining with seeding is used to provide focused mining on a given set of intent seeds. This last case can be used to mine additional utterances to associate with each of the seed intents in a given set.

[0125] Referring particularly to Figure 10, a method 600 for intent mining using intent seeds is provided. In an exemplary embodiment, method 600 includes a first step 605 of receiving seed intents. Each seed intent includes an intent label and an utterance having a sample intent. In step 610, utterances having intents are identified from conversational data. In step 615, candidate intents are selected from the utterances having intents. In step 620, seed intent substitutes are identified from the utterances having sample intents. Then, in step 625, the new utterances are associated with seed intents. These steps will now be described in more detail in the following embodiments.

[0126] According to an exemplary embodiment, a computer implementation method for authoring a conversational bot and intent mining using intent seeding is provided. The method may include receiving conversational data, wherein the conversational data includes text derived from conversations, and each conversation is between a customer and a customer service representative; receiving seed intent data, which may include seed intents, wherein each seed intent includes a seed intent label and an utterance having a sample intent associated with the seed intent; automatically mining the conversational data using an intent mining algorithm to determine new utterances to be associated with seed intents; expanding the seed intent data to include the newly mined utterances associated with seed intents; and uploading the expanded seed intent data to a conversational bot and using the conversational bot to conduct automated conversations with other customers.

[0127] In the case of seed intent mining, intent mining algorithms may include analyzing utterances occurring within conversations in conversational data to identify intent-containing utterances. Each utterance may contain a turn within a conversation, thereby communicating whether a customer is speaking in the form of customer utterances or a customer service representative is speaking in the form of customer service representative utterances. An intent-containing utterance can be defined as one of the utterances determined to be highly likely to express an intent. Intent mining algorithms may further include analyzing the identified intent-containing utterances to identify candidate intents. Each candidate intent is identified as a text phrase occurring within one of the intent-containing utterances, having two parts: an action which may contain a word or phrase describing a purpose or task, and an object which may contain a word or phrase describing an object or thing on which the action operates. For each seed intent, intent mining algorithms may further include identifying seed intent surrogates from sample intent-containing utterances associated with the seed intent. A seed intent substitute is identified as a text phrase occurring within one of the sample intent utterances, which may consist of two parts: an action that may contain a word or phrase describing a purpose or task, and an object that may contain a word or phrase describing the object or thing on which the action operates. The intent mining algorithm may further include associating intent utterances from conversational data with seed intents by determining the degree of semantic similarity between candidate intents present within the intent utterances and seed intent substitutes belonging to each of the seed intent labels.

[0128] According to an exemplary embodiment, the step of identifying an intentional utterance may include selecting a first portion of a customer utterance as an intentional utterance and discarding a second portion of the customer utterance in the conversation data. The first portion of a customer utterance may be defined as a predetermined number of consecutive customer utterances occurring at the beginning of each conversation, and the second portion may be defined as the remainder of each conversation. The step of identifying an intentional utterance may further include discarding customer utterances in the first portion of a customer utterance that do not satisfy a word count constraint. The word count constraint may include a minimum word count constraint, where customer utterances in the first portion of a customer utterance having fewer words than the minimum word count constraint are discarded, and / or a maximum word count constraint, where customer utterances in the first portion of a customer utterance having more words than the maximum word count constraint are discarded.

[0129] According to an exemplary embodiment, the step of identifying candidate intentions may include: identifying head-token pairs by analyzing the grammatical structure of an intent utterance using a syntactic-dependent parser, wherein each head-token pair includes a head word modified by a token word; and identifying head-token pairs as candidate intentions by using part-of-speech (POS) tagging to tag the part of speech of an intent utterance, wherein the POS tag of the head word may include a noun tag and the POS tag of the token word may include a verb tag.

[0130] According to an exemplary embodiment, the step of identifying seed intent substitutes may include: identifying head-token pairs by analyzing the grammatical structure of utterances with sample intent using a syntactic-dependent parser, wherein each head-token pair includes a head word modified by a token word; and using part-of-speech (hereinafter, "POS") tagging to tag the parts of speech of utterances with sample intent, identifying head-token pairs as candidate intents where the POS tag of the head word may include a noun tag and the POS tag of the token word may include a verb tag.

[0131] According to an exemplary embodiment, the step of associating an intent-containing utterance from conversational data with a seed intent may include repeatedly performing a first process to cover each intent-containing utterance in relation to each seed intent, as described in relation to an exemplary first case involving first and second seed intents, and first and second candidate intents, with an intent-containing utterance. The first process may include: computer-processing the degree of semantic similarity between each of the first and second candidate intents and each of the intent substitutes in the first seed intent; computer-processing the degree of semantic similarity between each of the first and second candidate intents and each of the intent substitutes in the second seed intent; determining which of the intent substitutes produced the highest computer-processed degree of semantic similarity; and associating the first intent-containing utterance with the first seed intent, which of the first and second seed intents includes the intent substitute that was determined to produce the highest computer-processed degree of semantic similarity.

[0132] In an alternative use case, the method of the present invention includes automatically mining new intentions, along with mining new utterances for association with a given set of seed intentions using an intention mining algorithm. In such a case, the method may include expanding the seed intention data to include the newly mined intentions. In this case, the intention mining algorithm may further include selecting prominent intentions from candidate intentions present in utterances that have intentions not yet associated with one of the seed intentions (hereinafter, "utterances with unassociated intentions") according to one or more criteria; grouping the selected prominent intentions into prominent intention groups according to the degree of semantic similarity between the prominent intentions; for each prominent intention group, selecting one of the prominent intentions as the intention label and designating the other prominent intentions as intention substitutes; and associating utterances with unassociated intentions from conversational data into prominent intention groups by determining the degree of semantic similarity between the candidate intentions present in utterances with unassociated intentions and the intention substitutes within each of the prominent intention groups. Each newly mined intent may include a given one from a group of significant intents, each defined by one of the significant intents selected as an intent label and another of the significant intents designated as alternative intents, and an utterance having an unassociated intent that becomes associated with the given one from the group of significant intents.

[0133] According to an exemplary embodiment, the step of identifying candidate intentions may include: identifying head-token pairs by analyzing the grammatical structure of an intent utterance using a syntactic-dependent parser, wherein each head-token pair includes a head word modified by a token word; and identifying head-token pairs as candidate intentions by using part-of-speech (POS) tagging to tag the part of speech of an intent utterance, wherein the POS tag of the head word may include a noun tag and the POS tag of the token word may include a verb tag.

[0134] According to an exemplary embodiment, one or more criteria for selecting a significant intent from candidate intents may include criteria based on latent semantic analysis (LSA). The step of selecting a significant intent from candidate intents may include generating a set of documents, each having a document corresponding to each of the candidate intents, such that each document covers an action-object pair defined by one of the corresponding candidate intents; generating conceptual groups based on terms appearing in the action-object pairs contained in the set of documents; calculating weight values ​​for each of the candidate intents for each of the conceptual groups, such that the weight values ​​measure the degree of relevance between a given candidate intent in the documents and a given one of the conceptual groups; and selecting a predetermined number of candidate intents in each of the conceptual groups as significant intents, and creating weight values ​​thereon that indicate a higher degree of relevance.

[0135] According to an exemplary embodiment, the step of grouping significant intentions according to their degree of semantic similarity may include: calculating an embedding for each significant intention, wherein the embedding may include encoded representations of texts whose semantically similar texts have similar encoded representations; comparing the calculated embeddings to determine the degree of semantic similarity between pairs of significant intentions; and grouping significant intentions whose degree of semantic similarity exceeds a predetermined threshold. The embedding is calculated as the inverse document frequency (IDF) mean of global vector embeddings of head-token pairs that are components of the significant intention. Comparing the calculated embeddings may include cosine similarity.

[0136] According to an exemplary embodiment, the step of associating an utterance with an unassociated intent from conversational data with a significant intent group may include repeatedly performing the first process to cover each utterance with an unassociated intent in relation to each significant intent group. When described in relation to an exemplary first case involving first and second significant intent groups and a first utterance with an unassociated intent that includes first and second candidate intents, the first process may include: computer processing the degree of semantic similarity between each of the first and second candidate intents and each of the intent substitutes in the first significant intent group; computer processing the degree of semantic similarity between each of the first and second candidate intents and each of the intent substitutes in the second significant intent group; determining which of the intent substitutes produced the highest computer-processed degree of semantic similarity; and associating the first utterance with an unassociated intent with the first utterance with an unassociated intent with the first of the first and second significant intent groups that includes the intent substitute that was determined to produce the highest computer-processed degree of semantic similarity.

[0137] As those skilled in the art will understand, many of the various features and configurations described above in relation to some exemplary embodiments may be further selectively applied to form other possible embodiments of the present invention. For the sake of brevity and in consideration of the ability of those skilled in the art, each possible iteration is not provided or discussed in detail, but all combinations and possible embodiments or others encompassed by some of the following claims are intended to be part of this application. Furthermore, from the above description of various exemplary embodiments of the present invention, those skilled in the art will understand improvements, changes, and modifications. Such improvements, changes, and modifications within the scope of the skills of those skilled in the art are also intended to be covered by the appended claims. Furthermore, the above relates only to the described embodiments of this application, and it will be apparent that a number of changes and modifications can be made herein without departing from the spirit and scope of this application as defined by the following claims and their equivalents.< / person> < / person> < / date> < / date> < / geo> < / geo> < / geo>

Claims

1. 1. A computer-implemented method for authoring a conversational bot, comprising: receiving conversation data including text derived from the conversations, each of the conversations being between a customer and a customer service representative; receiving seed intent data including seed intents, each of the seed intents including a seed intent label and a sample utterance having an intent associated with the seed intent; automatically mining the conversation data using an intent mining algorithm to determine new utterances to associate with the seed intent; expanding the seed intent data to include the mined new utterances associated with the seed intent; uploading the augmented seed intent data to the conversational bot and using the conversational bot to conduct automated conversations with other customers; The intent mining algorithm analyzing utterances occurring within the conversation of the conversation data to identify utterances having intent, each of the utterances comprises a turn in the conversation, whereby the customer is communicating in the form of a customer utterance or the customer service representative is communicating in the form of a customer service representative utterance; identifying an utterance having an intention, wherein the utterance is defined as one of the utterances determined to be likely to express the intention; analyzing the identified utterances having the intent to identify candidate intents, each of the candidate intents being identified as a text phrase occurring within one of the utterances having the intent, the text phrase having two parts: an action, which includes a word or phrase describing a purpose or task, and an object, which includes a word or phrase describing an object or thing on which the action operates; for each of the seed intentions, identifying a seed intention alternative from sample utterances having the intent associated with the seed intention, wherein the seed intention alternative is identified as being a text phrase occurring within one of the sample utterances having the intent that includes two parts: an action including a word or phrase describing a purpose or task, and an object including a word or phrase describing an object or thing on which the action operates; and associating an utterance having the intention from the conversation data with the seed intention by determining a degree of semantic similarity between the candidate intention present in the utterance having the intention and the seed intention alternatives belonging to each of the seed intention labels.

2. wherein identifying the intended utterance includes selecting a first portion of the customer utterance as the intended utterance and discarding a second portion of the customer utterance in the conversation data; 2. The method of claim 1, wherein the first portion of customer utterances is defined as a predetermined number of consecutive customer utterances occurring at the beginning of each of the conversations, and the second portion is defined as the remainder of each of the conversations.

3. identifying the intended utterance further comprises discarding the customer utterance within the first portion of the customer utterance that does not satisfy a word count constraint; The word count constraint is: a minimum word count constraint, wherein customer utterances in the first portion of the customer utterance having fewer words than the minimum word count constraint are discarded; a maximum word count constraint, wherein the customer utterances in the first portion of the customer utterance having more words than the maximum word count constraint are discarded.

4. Identifying candidate intents includes: analyzing the grammatical structure of the intended utterance using a syntax-sensitive parser to identify head-token pairs, each head-token pair including a head word modified by a token word; using part-of-speech (hereinafter "POS") tagging to tag parts of speech of the utterance having the intention, and identifying as the candidate intention the head-token pairs in which the POS tag of the head word includes a noun tag and the POS tag of the token word includes a verb tag.

5. said identifying a seed intent alternative comprising: analyzing the grammatical structure of the sample utterance having the intent using a syntax-sensitive parser to identify head-token pairs, each head-token pair including a head word modified by a token word; using part-of-speech (hereinafter "POS") tagging to tag parts of speech of the sample utterances having the intent, and identifying as the candidate intent the head-token pairs in which the POS tag of the head word includes a noun tag and the POS tag of the token word includes a verb tag.

6. The associating of utterances having the intention from the conversation data with the seed intention includes iteratively performing a first process to cover each of the utterances having the intention in relation to each of the seed intentions, wherein when described in relation to an exemplary first case involving first and second seed intentions and utterances having a first intention including first and second candidate intentions, the first process: computing a degree of semantic similarity between each of the first and second candidate intents and each of the intent alternatives within the first seed intent; computing a degree of semantic similarity between each of the first and second candidate intents and each of the intent alternatives within the second seed intent; determining which of the intent alternatives produced the highest computationally processed degree of semantic similarity; and associating the utterance having the first intent with one of the first and second seed intents that includes the intent alternative determined to produce the highest computationally processed degree of semantic similarity.

7. automatically mining new intents using the intent mining algorithm, wherein each of the mined new intents includes an intent label, an intent alternative, and an associated utterance; and expanding the seed intent data to include the mined new intents; The intent mining algorithm: selecting a salient intent from the candidate intents present in an utterance having an intent that is not yet associated with one of the seed intents (hereinafter, "an utterance having an unassociated intent") according to one or more criteria; Grouping the selected salient intents into salient intent groups according to the degree of semantic similarity between the salient intents; For each of the salient intent groups, selecting one of the salient intents as the intent label and designating the other salient intents as intent alternatives; 2. The method of claim 1, further comprising: associating the utterances having the unassociated intentions from the conversation data with the salient intent groups via determining a degree of semantic similarity between the candidate intents present in the utterances having the unassociated intentions and the intent alternatives in each of the salient intent groups.

8. Each of the new mined intents: a given one of said salient intent groups, each of which comprises: the one of the salient intents selected as the intent label; and a given one defined by the other of the salient intents designated as the alternative intent; and utterances having the unassociated intent that become associated with the given one of the salient intent groups.

9. Identifying candidate intents includes: analyzing the grammatical structure of the intended utterance using a syntax-sensitive parser to identify head-token pairs, each head-token pair including a head word modified by a token word; using part-of-speech (hereinafter "POS") tagging to tag parts of speech of the utterance having the intention, and identifying as the candidate intention the head-token pairs in which the POS tag of the head word includes a noun tag and the POS tag of the token word includes a verb tag.

10. The selecting the salient intent from the candidate intents includes: generating a document set having a document corresponding to each one of the candidate intentions, each of the documents covering an action-object pair defined by the corresponding one of the candidate intentions; generating concept groups based on terms appearing in the action-object pairs contained in the set of documents; calculating a weight value for each of the candidate intents for each of the concept groups, the weight value measuring a degree of relevance between the candidate intent of a given one of the documents and the given one of the concept groups; and selecting a predetermined number of the candidate intents in each of the concept groups as the salient intents and generating a weight value indicating a higher degree of relevance based thereon.

11. The grouping of the salient intents according to the degree of semantic similarity includes: Computing an embedding for each of the salient intents, the embedding including coded representations of text where semantically similar text has similar coded representations; comparing the computed embeddings to determine the degree of semantic similarity between the pair of salient intents; and grouping the salient intents having a degree of semantic similarity above a predetermined threshold.

12. the embeddings are computed as the inverse document frequency average of the global vector embeddings of the head-token pairs that are components of the salient intent; The method of claim 11 , wherein the comparing the computed embeddings comprises cosine similarity.

13. The associating of utterances having the unassociated intentions from the conversation data to the salient intent groups includes iteratively performing a first process to cover each of the utterances having the unassociated intentions in association with each of the salient intent groups, wherein when described in relation to an exemplary first case involving first and second salient intent groups and an utterance having a first unassociated intention that includes first and second candidate intents, the first process: computing a degree of semantic similarity between each of the first and second candidate intentions and each of the intention alternatives in the first salient intention group; computing a degree of semantic similarity between each of the first and second candidate intentions and each of the intention alternatives in the second salient intention group; determining which of the intent alternatives produced the highest computationally processed degree of semantic similarity; and associating the utterance having the first unassociated intent with one of the first and second salient intent groups that includes the intent alternative determined to produce the highest computationally processed degree of semantic similarity.

14. 1. A system for automating aspects of authoring a conversational bot, the system comprising: a processor; a memory storing instructions that, when executed by the processor, cause the processor to: receiving conversation data including text derived from the conversations, each of the conversations being between a customer and a customer service representative; receiving seed intent data including seed intents, each of the seed intents including a seed intent label and a sample utterance having an intent associated with the seed intent; automatically mining the conversation data using an intent mining algorithm to determine new utterances to associate with the seed intent; expanding the seed intent data to include the mined new utterances associated with the seed intent; uploading the augmented seed intent data to the conversational bot and using the conversational bot to conduct automated conversations with other customers; The intent mining algorithm analyzing utterances occurring within the conversation of the conversation data to identify utterances having intent, each of the utterances comprises a turn in the conversation, whereby the customer is communicating in the form of a customer utterance or the customer service representative is communicating in the form of a customer service representative utterance; identifying an utterance having an intention, wherein the utterance is defined as one of the utterances determined to be likely to express the intention; analyzing the identified utterances having the intent to identify candidate intents, each of the candidate intents being identified as a text phrase occurring within one of the utterances having the intent, the text phrase having two parts: an action, which includes a word or phrase describing a purpose or task, and an object, which includes a word or phrase describing an object or thing on which the action operates; for each of the seed intentions, identifying a seed intention alternative from sample utterances having the intent associated with the seed intention, wherein the seed intention alternative is identified as being a text phrase occurring within one of the sample utterances having the intent that includes two parts: an action including a word or phrase describing a purpose or task, and an object including a word or phrase describing an object or thing on which the action operates; and associating an utterance having the intention from the conversation data with the seed intention by determining a degree of semantic similarity between the candidate intention present in the utterance having the intention and the seed intention alternatives belonging to each of the seed intention labels.

15. wherein identifying the intended utterance includes selecting a first portion of the customer utterance as the intended utterance and discarding a second portion of the customer utterance in the conversation data; 15. The system of claim 14, wherein the first portion of customer utterances is defined as a predetermined number of consecutive customer utterances occurring at the beginning of each of the conversations, and the second portion is defined as the remainder of each of the conversations.

16. identifying the intended utterance further comprises discarding the customer utterance within the first portion of the customer utterance that does not satisfy a word count constraint; The word count constraint is: a minimum word count constraint, wherein customer utterances in the first portion of the customer utterance having fewer words than the minimum word count constraint are discarded; a maximum word count constraint, wherein the customer utterances in the first portion of the customer utterance having more words than the maximum word count constraint are discarded.

17. Identifying candidate intents includes: analyzing the grammatical structure of the intended utterance using a syntax-sensitive parser to identify head-token pairs, each head-token pair including a head word modified by a token word; using part-of-speech (hereinafter "POS") tagging to tag parts of speech of the utterance having the intention, and identifying as the candidate intention the head-token pairs in which the POS tag of the head word includes a noun tag and the POS tag of the token word includes a verb tag.

18. said identifying a seed intent alternative comprising: analyzing the grammatical structure of the sample utterance having the intent using a syntax-sensitive parser to identify head-token pairs, each head-token pair including a head word modified by a token word; using part-of-speech (hereinafter "POS") tagging to tag parts of speech of the sample utterances having the intent, and identifying as the candidate intent the head-token pairs in which the POS tag of the head word includes a noun tag and the POS tag of the token word includes a verb tag.

19. The associating of utterances having the intention from the conversation data with the seed intention includes iteratively performing a first process to cover each utterance having the intention in relation to each of the seed intentions, wherein when described in relation to an exemplary first case involving first and second seed intentions and utterances having a first intention including first and second candidate intentions, the first process: computing a degree of semantic similarity between each of the first and second candidate intents and each of the intent alternatives within the first seed intent; computing a degree of semantic similarity between each of the first and second candidate intents and each of the intent alternatives within the second seed intent; determining which of the intent alternatives produced the highest computationally processed degree of semantic similarity; and associating the utterance having the first intent with one of the first and second seed intents that includes the intent alternative determined to produce the highest computationally processed degree of semantic similarity.

20. automatically mining new intents using the intent mining algorithm, wherein each of the mined new intents includes an intent label, an intent alternative, and an associated utterance; and expanding the seed intent data to include the mined new intents; The intent mining algorithm selecting a salient intent from the candidate intents present in an utterance having an intent that is not yet associated with one of the seed intents (hereinafter, "an utterance having an unassociated intent") according to one or more criteria; Grouping the selected salient intents into salient intent groups according to the degree of semantic similarity between the salient intents; For each of the salient intent groups, selecting one of the salient intents as the intent label and designating the other salient intents as intent alternatives; 15. The system of claim 14, further comprising: associating the utterances having the unassociated intentions from the conversation data with the salient intent groups via determining a degree of semantic similarity between the candidate intents present in the utterances having the unassociated intentions and the intent alternatives in each of the salient intent groups.

21. Each of the new mined intents: a given one of said salient intent groups, each of which comprises: the one of the salient intents selected as the intent label; and a given one defined by the other of the salient intents designated as the alternative intent; and utterances having the unassociated intent that become associated with the given one of the salient intent groups.

22. Identifying candidate intents includes: analyzing the grammatical structure of the intended utterance using a syntax-sensitive parser to identify head-token pairs, each head-token pair including a head word modified by a token word; using part-of-speech (hereinafter "POS") tagging to tag parts of speech of the utterance having the intention, and identifying as the candidate intention the head-token pairs in which the POS tag of the head word includes a noun tag and the POS tag of the token word includes a verb tag.

23. The selecting the salient intent from the candidate intents includes: generating a document set having a document corresponding to each one of the candidate intents, each of the documents being defined by the corresponding one of the candidate intents; Actions that cover object pairs, generating them, and generating concept groups based on terms appearing in the action-object pairs contained in the set of documents; calculating a weight value for each of the candidate intents for each of the concept groups, the weight value measuring a degree of relevance between the candidate intent of a given one of the documents and the given one of the concept groups; and selecting a predetermined number of the candidate intents in each of the concept groups as the salient intents and generating a weight value indicating a higher degree of relevance based thereon.

24. The grouping of the salient intents according to the degree of semantic similarity includes: Computing an embedding for each of the salient intents, the embedding including coded representations of text where semantically similar text has similar coded representations; comparing the computed embeddings to determine the degree of semantic similarity between the pair of salient intents; and grouping the salient intents having a degree of semantic similarity above a predetermined threshold.

25. the embeddings are computed as the inverse document frequency average of the global vector embeddings of the head-token pairs that are components of the salient intent; 25. The system of claim 24, wherein the comparing the computed embeddings comprises cosine similarity.

26. The associating of utterances having the unassociated intentions from the conversation data to the salient intent groups includes iteratively performing a first process to cover each of the utterances having the unassociated intentions in association with each of the salient intent groups, wherein when described in relation to an exemplary first case involving first and second salient intent groups and an utterance having a first unassociated intention that includes first and second candidate intents, the first process: computing a degree of semantic similarity between each of the first and second candidate intentions and each of the intention alternatives in the first salient intention group; computing a degree of semantic similarity between each of the first and second candidate intentions and each of the intention alternatives in the second salient intention group; determining which of the intent alternatives produced the highest computationally processed degree of semantic similarity; and associating the utterance having the first unassociated intent with one of the first and second salient intent groups that includes the intent alternative determined to produce the highest computationally processed degree of semantic similarity.