Computer-based systems and methods for creating conversational bots
Patent Information
- Application Number
- CN202180065821.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-09-25
- Filing Date
- 2021-09-27
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2041-09-27
Smart Images

Figure CN116235177B_ABST
Abstract
Description
[0001] Cross-referencing of related patent applications
[0002] This application relates to U.S. Provisional Patent Application No. 63 / 083,561, filed September 25, 2020, with the U.S. Patent and Trademark Office, entitled “SYSTEMS AND METHODS RELATING TO BOT AUTHORING AND / OR AUTOMATING THE MINING OF INTENTSFROM NATURAL LANGUAGE CONVERSATIONS”, which was converted into pending U.S. Patent Application No. 17 / 218456, filed March 31, 2021, entitled “SYSTEMS AND METHODS RELATING TO BOTAUTHORING BY MINING INTENTS FROM CONVERSATION DATA VIA INTENT SEEDING”. Background Technology
[0003] This invention relates generally to telecommunications systems in the field of customer relationship management, including customer assistance via Internet-based service options. More specifically, but not as a limitation, this invention relates to systems and methods for automating robot creation workflows and / or implementing intent mining processes for using intent seeding processes to mine intents and associated utterances from natural language conversational data. Summary of the Invention
[0004] This invention includes a computer-implemented method for creating a conversational bot and for intent mining using intent seeding. The method may include: receiving conversation data comprising text derived from conversations, wherein each of these conversations is between a customer and a customer service representative; receiving seed intent data that may include seed intents, each seed intent including a seed intent tag and a sample of intent-related utterances associated with the seed intent; using an intent mining algorithm to automatically mine the conversation data to determine new utterances to be associated with the seed intents; expanding the seed intent data to include the mined new utterances associated with the seed intents; and uploading the expanded seed intent data to a conversational bot and using the conversational bot to engage in automated conversations with other customers. The intent mining algorithm may include analyzing utterances appearing within the conversation data to identify intent-related utterances. These utterances may each include a turn within the conversation, whereby the customer is communicating in the form of customer utterances or the customer service representative is communicating in the form of customer service representative utterances. Intent-related utterances may be defined as utterances that are determined to have a greater probability of expressing an intent. The intent mining algorithm may also include analyzing the identified intent-related utterances to identify candidate intents. Candidate intents are each identified as text phrases appearing within a sentence of an intent-laden utterance. These text phrases have two parts: an action and an object. The action may include words or phrases describing a purpose or task, and the object may include words or phrases describing the object or thing to which the action is performed. The intent mining algorithm may further include: for each seed intent in the seed intents, identifying seed intent substitutes from sample intent-laden utterances associated with that seed intent. Seed intent substitutes are identified as text phrases appearing within a sentence of a sample intent-laden utterance. These text phrases may include two parts: an action and an object. The action may include words or phrases describing a purpose or task, and the object may include words or phrases describing the object or thing to which the action is performed. The intent mining algorithm may further include: associating intent-laden utterances from session data with seed intents by determining the semantic similarity between candidate intents present in the intent-laden utterances and seed intent substitutes belonging to each seed intent tag in the seed intent tags.
[0005] These and other features of this application will become more apparent when read in conjunction with the accompanying drawings and the appended claims in the following detailed description of exemplary embodiments. Attached Figure Description
[0006] A more complete understanding of the invention will become apparent when considered in conjunction with the accompanying drawings and with reference to the following detailed description, wherein similar reference numerals indicate similar parts in the drawings:
[0007] Figure 1A schematic block diagram of a computing device is shown that enables or practices an exemplary embodiment of the invention according to the present invention.
[0008] Figure 2 A schematic block diagram of a communication infrastructure or contact center that can be used to enable or practice the invention according to an exemplary embodiment of the invention is shown.
[0009] Figure 3 This is a schematic block diagram illustrating further details of a chat server operating as part of a chat system according to an embodiment of the present invention;
[0010] Figure 4 This is a schematic block diagram of a chat module according to an embodiment of the present invention;
[0011] Figure 5 This is an exemplary customer chat interface according to an embodiment of the present invention;
[0012] Figure 6 This is a block diagram of a customer automation system according to an embodiment of the present invention;
[0013] Figure 7 This is a flowchart of a method for automating interactions on behalf of a customer, according to an embodiment of the present invention;
[0014] Figure 8 This is a workflow for creating conversational bots;
[0015] Figure 9 This is an exemplary flowchart of the intent mining according to the present invention; and
[0016] Figure 10 This is an exemplary flowchart of intention mining via seeding intention according to the present invention. Detailed Implementation
[0017] To facilitate understanding of the principles of the invention, reference will now be made to exemplary embodiments illustrated in the accompanying drawings, and these embodiments will be described using specific language. However, it will be apparent to those skilled in the art that the detailed material provided in the examples may not be necessary for practicing the invention. In other instances, well-known materials or methods have not been described in detail to avoid obscuring the invention. Furthermore, as would normally be expected by those skilled in the art, as presented herein, further modifications to the provided examples or applications of the principles of the invention can be contemplated.
[0018] As used herein, language specifying non-limiting examples and descriptions includes "for example (e.g., for example / for instance)," "that is," etc. Furthermore, throughout this specification, terms such as "implementation," "an embodiment," "an embodiment of the invention," "exemplary embodiment," "certain embodiments," etc., refer to a particular feature, structure, or characteristic described in connection with a given example that may be included in at least one embodiment of the invention. Therefore, the appearance of the phrases "implementation," "an embodiment," "an embodiment of the invention," "exemplary embodiment," "certain embodiments," etc., does not necessarily refer to the same embodiment or example. Moreover, in one or more embodiments or examples, a particular feature, structure, or characteristic may be combined in any suitable combination and / or sub-combination.
[0019] Those skilled in the art will recognize from this disclosure that various embodiments can be computers implemented using many different types of data processing devices, wherein the embodiments are implemented as apparatus, methods, or computer program products. Therefore, exemplary embodiments may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Exemplary embodiments may also take the form of a computer program product embodied in computer-usable program code in any tangible medium. In each case, the exemplary embodiments may generally be referred to as a “module,” a “system,” or a “method.”
[0020] The flowcharts and block diagrams provided in the accompanying drawings illustrate the architecture, functionality, and operation of possible specific implementations of systems, methods, and computer program products according to exemplary embodiments of the present invention. In this regard, it should be understood that each block or combination of blocks in the flowcharts and / or block diagrams may represent a module, segment, or portion of program code having one or more executable instructions for implementing a specified logical function. Similarly, it should be understood that each block or combination of blocks in the flowcharts and / or block diagrams may be implemented by a system based on dedicated hardware or a combination of dedicated hardware and computer instructions that perform a specified action or function. Such computer program instructions may also be stored in a computer-readable medium that instructs a computer or other programmable data processing apparatus to operate in a particular manner, such that the program instructions in the computer-readable medium produce an article of art including instructions that implement the function or action specified in each block or combination of blocks in the flowcharts and / or block diagrams.
[0021] computing devices
[0022] It should be understood that the systems and methods of the present invention can be implemented using a computer employing many different forms of data processing devices (e.g., digital microprocessors and associated memory) that execute appropriate software programs. Considering the background factors, Figure 1A schematic block diagram of an exemplary computing device 100 according to embodiments of the present invention and / or utilizing it to enable or practice those embodiments is shown. It should be understood that... Figure 1 Provided as a non-restrictive example.
[0023] Computing device 100 may be implemented, for example, via firmware (e.g., an application-specific integrated circuit), hardware or software, or a combination of firmware and hardware. It should be understood that each of the servers, controllers, switches, gateways, engines, and / or modules (collectively referred to as servers or modules) in the figure below may be implemented via one or more of the computing devices 100. For example, various servers may be processes running on one or more processors of one or more computing devices 100, which may execute computer program instructions and interact with other systems or modules to perform the various functions described herein. Unless otherwise expressly limited, the functions described with respect to multiple computing devices may be integrated into a single computing device, or the various functions described with respect to a single computing device may be distributed across several computing devices. Furthermore, regarding the computing system described in the figure below, such as… Figure 2 The contact center system 200, various servers, and their computer equipment may be located on local computing devices 100 (i.e., on-site or at the same physical location as the contact center agents), on remote computing devices 100 (i.e., off-site or in a cloud computing environment, such as in a remote data center connected to the contact center via a network), or some combination thereof. Functionality provided by servers located on off-site computing devices may be accessed and provided via a Virtual Private Network (VPN) as if such servers were on-site, or functionality may be provided using Software as a Service (SaaS) (which uses various protocols to be accessed over the Internet), such as exchanging data via Extensible Markup Language (XML), JSON, etc.
[0024] As illustrated in the example, computing device 100 may include a central processing unit (CPU) or processor 105 and main memory 110. Computing device 100 may also include storage device 115, removable media interface 120, network interface 125, I / O controller 130, and one or more input / output (I / O) devices 135, which may include a display device 135A, a keyboard 135B, and a pointing device 135C, as shown. Computing device 100 may also include additional elements such as memory port 140, bridge 145, I / O ports, one or more additional input / output devices 135D, 135E, 135F, and cache memory 150 communicating with processor 105.
[0025] Processor 105 can be any logic circuit that responds to and processes instructions fetched from main memory 110. For example, process 105 can be implemented by an integrated circuit (e.g., a microprocessor, microcontroller, or graphics processing unit) or in a field-programmable gate array or application-specific integrated circuit. As shown, processor 105 can communicate directly with cache memory 150 via an auxiliary bus or back bus. Cache memory 150 typically has a faster response time than main memory 110. Main memory 110 can be one or more memory chips capable of storing data and allowing central processing unit 105 to directly access the stored data. Storage device 115 can provide storage for operating systems and other software that control scheduling tasks and access to system resources. Unless otherwise limited, computing device 100 can include operating systems and software capable of performing the functions described herein.
[0026] As depicted in the illustrated example, computing device 100 may include various I / O devices 135, one or more of which may be connected via I / O controller 130. Input devices may include, for example, a keyboard 135B and pointing devices 135C, such as a mouse or optical pen. Output devices may include, for example, a video display device, speakers, and a printer. I / O devices 135 and / or I / O controller 130 may include suitable hardware and / or software for enabling multiple display devices. Computing device 100 may also support one or more removable media interfaces 120, such as a disk drive, a USB port, or any other device suitable for reading data from or writing data to a computer-readable medium. More generally, I / O devices 135 may include any conventional devices for performing the functions described herein.
[0027] Computing device 100 can be any workstation, desktop computer, laptop or notebook computer, server machine, virtualization machine, mobile or smartphone, portable telecommunications equipment, media playback device, gaming system, mobile computing device, or any other type of computing, telecommunications, or media device capable of (but not limited to) performing the operations described herein. Computing device 100 includes multiple devices connected to or via a network to other systems and resources. As used herein, a network includes one or more computing devices, machines, clients, client nodes, client machines, client computers, client devices, endpoints, or endpoint nodes that communicate with one or more other computing devices, machines, clients, client nodes, client machines, client computers, client devices, endpoints, or endpoint nodes. For example, a network can be a private or public switched telephone network (PSTN), a wireless carrier network, a local area network (LAN), a private wide area network (WAN), or a public WAN such as the Internet, where appropriate communication protocols are used to establish connections. More generally, it should be understood that, unless otherwise limited, computing device 100 can communicate with other computing devices 100 via any type of network using any conventional communication protocols. Furthermore, the network can be a virtual network environment in which various network components are virtualized. For example, various machines can be virtual machines implemented as software-based computers running on physical machines, or multiple virtual machines can run on the same host physical machine using a "hypervisor" type of virtualization. Other types of virtualization are also conceivable.
[0028] Contact Center
[0029] Now for reference Figure 2 This illustrates a communication infrastructure or contact center system 200 according to exemplary embodiments of the present invention and / or exemplary embodiments thereof that enable or practice the present invention. It should be understood that the term "contact center system" is used herein to refer to... Figure 2 The term “contact center” refers to the system and / or its components, while it is more generally used to refer to a contact center system, the customer service providers operating those systems, and / or the organizations or enterprises associated with them. Therefore, unless otherwise expressly limited, the term “contact center” generally refers to a contact center system (such as contact center system 200), the associated customer service providers (such as a specific customer service provider that provides customer service through contact center system 200), and the organizations or enterprises that provide those customer services on its behalf.
[0030] In the back office, customer service providers typically offer a variety of services through contact centers. These contact centers may be staffed with employees or customer service agents (or simply "agents") who act as intermediaries between companies, businesses, government agencies, or organizations (hereinafter referred to as "organizations" or "enterprises") and individuals such as users, individuals, or customers (hereinafter referred to as "individuals" or "customers"). For example, agents at a contact center can assist customers in making purchasing decisions, receiving orders, or resolving issues with received products or services. Within a contact center, such interactions between contact center agents and external entities or customers can take place over various communication channels, such as via voice (e.g., telephone calls or VoIP calls), video (e.g., video conferencing), text (e.g., email and text chat), screen sharing, shared browsing, etc.
[0031] Operationally, contact centers generally strive to provide high-quality service to customers while minimizing costs. For example, one way contact centers operate is by handling each customer's interaction with a live agent. While this approach may score well in terms of service quality, it can also be very expensive due to the high cost of agent labor. Therefore, most contact centers utilize some degree of automation to replace live agents, such as interactive voice response (IVR) systems, interactive media response (IMR) systems, internet bots or "bots," automated chat modules or "chatbots," etc. In many cases, this has proven to be a successful strategy because automated processes can handle certain types of interactions very efficiently and effectively reduce the need for live agents. Such automation allows contact centers to use human agents for more challenging customer interactions, while automated processes handle more repetitive or routine tasks. Furthermore, automated processes can be built in ways that optimize efficiency and promote repeatability. While human or live agents may forget to ask certain questions or follow up on specific details, such errors can often be avoided by using automated processes. Although customer service providers are increasingly reliant on automated processes to interact with customers, customers still use such technologies far less. Therefore, although IVR systems, IMR systems, and / or robots are used to automate some interactions on the contact center side of the interaction, the actions on the customer side are still performed manually by the customer.
[0032] For details, please refer to the following: Figure 2Customer service providers can use contact center system 200 to provide various types of services to customers. For example, contact center system 200 can be used to participate in and manage interactions between automated processes (or robots) or human agents and customers. It should be understood that contact center system 200 can be an internal facility of a business or enterprise, used to perform sales and customer service functions relative to the products and services available to the enterprise. On the other hand, contact center system 200 can be operated by a third-party service provider contracted to provide services to another organization. Furthermore, contact center system 200 can be deployed on equipment dedicated to an enterprise or third-party service provider, and / or deployed in remote computing environments, such as, for example, private or public cloud environments with infrastructure for supporting multiple contact centers for multiple enterprises. Contact center system 200 may include software applications or programs that can execute on-site, remotely, or in some combination thereof. It should also be understood that the various components of contact center system 200 can be distributed across various geographical locations and are not necessarily contained in a single location or computing environment.
[0033] It should also be understood that, unless otherwise expressly limited, any of the computing elements in this invention may also be implemented in a cloud-based or cloud computing environment. As used herein, “cloud computing” (or simply “cloud”) is defined as a model for enabling ubiquitous, convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, servers, storage devices, applications, and services), which can be rapidly provisioned via virtualization and deployed with minimal management effort or service provider interaction, and then scaled accordingly. Cloud computing can be comprised of various features (e.g., on-demand self-service, extensive network access, resource pooling, rapid elasticity, metered services, etc.), service models (e.g., Software as a Service (“SaaS”), Platform as a Service (“PaaS”), Infrastructure as a Service (“IaaS”), and deployment models (e.g., private cloud, community cloud, public cloud, hybrid cloud, etc.). Cloud execution models are often referred to as “serverless architectures,” which typically involve service providers that dynamically manage the allocation and configuration of remote servers to achieve the required functionality.
[0034] according to Figure 2As shown in the example, components or modules of the contact center system 200 may include: multiple client devices 205A, 205B, 205C; a communication network (or simply "network") 210; a switch / media gateway 212; a call controller 214; an interactive media response (IMR) server 216; a routing server 218; a storage device 220; a statistics (or "stat") server 226; multiple agent devices 230A, 230B, 230C, each including work areas 232A, 232B, 232C; a multimedia / social media server 234; a knowledge management server 236 coupled to a knowledge system 238; a chat server 240; a web server 242; an interactive (or "iXn") server 244; a universal contact server (or "UCS") 246; a reporting server 248; a media service server 249; and an analytics module 250. It should be understood that, relative to... Figure 2 Or any of the computer-implemented components, modules, or servers described in any of the following figures can be accessed via various types of computing devices (such as, for example...) Figure 1 The contact center system 200 is implemented using a computing device 100. As can be seen, the contact center system 200 generally manages resources (e.g., personnel, computers, telecommunications equipment, etc.) to enable the delivery of services via telephone, email, chat, or other communication mechanisms. Such services may vary depending on the type of contact center and may include, for example, customer service, help desk functions, emergency response, remote marketing, order taking, etc.
[0035] Customers wishing to receive services from contact center system 200 can initiate inbound communications (e.g., telephone calls, emails, chats, etc.) to contact center system 200 via customer equipment 205. Although Figure 2 Three such client devices are shown, namely client devices 205A, 205B, and 205C, but it should be understood that any number of such client devices may exist. Client device 205 may be, for example, a communication device such as a telephone, smartphone, computer, tablet, or laptop. Based on the functions described herein, a customer can generally use client device 205 to initiate, manage, and conduct communications with contact center system 200, such as telephone calls, emails, chat, text messages, web browsing sessions, and other multimedia transactions.
[0036] Inbound and outbound communications to and from client equipment 205 can traverse network 210, the nature of which typically depends on the type of client equipment used and the form of communication. For example, network 210 may include communication networks for telephone, cellular, and / or data services. Network 210 may be a private or public switched telephone network (PSTN), a local area network (LAN), a private wide area network (WAN), and / or a public WAN such as the Internet. Furthermore, network 210 may include a wireless carrier network, including Code Division Multiple Access (CDMA) networks, Global System for Mobile Communications (GSM) networks, or any wireless network / technology conventional in the art, including but not limited to 3G, 4G, LTE, 5G, etc.
[0037] Regarding switch / media gateway 212, it is coupled to network 210 for receiving and transmitting telephone calls between the customer and contact center system 200. Switch / media gateway 212 may include a telephone exchange or communication exchange configured to act as a central exchange for agent-level routing within the center. The exchange may be a hardware switching system or implemented via software. For example, switch 215 may include an automatic call distributor, a private branch exchange (PBX), an IP-based software exchange, and / or any other exchange with dedicated hardware and software configured to receive interactions from the Internet and / or from the telephone network from the customer and route those interactions to, for example, one of the agent devices 230. Thus, generally speaking, switch / media gateway 212 establishes a voice connection between the customer and the agent by establishing a connection between customer device 205 and agent device 230.
[0038] As further shown, the switch / media gateway 212 may be coupled to a call controller 214, which may serve as, for example, an adapter or interface between the switch and other routing, monitoring, and communication processing components of the contact center system 200. The call controller 214 may be configured to handle PSTN calls, VoIP calls, etc. For example, the call controller 214 may include computer telephony integration (CTI) software for engaging with the switch / media gateway and other components. The call controller 214 may include a Session Initiation Protocol (SIP) server for handling SIP calls. The call controller 214 may also extract data about incoming interactions, such as a customer's phone number, IP address, or email address, and then communicate this data with other contact center components while processing the interaction.
[0039] Regarding the Interactive Media Response (IMR) server 216, it can be configured to enable self-service or virtual assistant functions. Specifically, the IMR server 216 can be similar to an Interactive Voice Response (IVR) server, except that the IMR server 216 is not limited to voice and can also cover various media channels. In the example illustrating voice, the IMR server 216 can be configured with an IMR script to inquire about the customer's needs. For example, a bank's contact center can inform a customer via an IMR script to "press 1" if they wish to retrieve their account balance. By continuing to interact with the IMR server 216, the customer can receive service without speaking to an agent. The IMR server 216 can also be configured to determine why the customer contacted the contact center, allowing communications to be routed to the appropriate resources.
[0040] Regarding routing server 218, it can be used to route incoming interactions. For example, once it is determined that inbound communication should be handled by a human agent, functionality within routing server 218 can select the most appropriate agent and route the communication to them. This agent selection can be based on which available agent is best suited to handle the communication. More specifically, the selection of the appropriate agent can be based on a routing strategy or algorithm implemented by routing server 218. In doing so, routing server 218 can query data related to the incoming interaction, such as data related to a specific customer, available agents, and interaction type, as described in more detail below, which can be stored in a specific database. Once an agent is selected, routing server 218 can interact with call controller 214 to route (i.e., connect) the incoming interaction to the corresponding agent device 230. As part of this connection, information about the customer can be provided to the selected agent via their agent device 230. This information is intended to enhance the service that the agent can provide to the customer.
[0041] Regarding data storage, contact center system 200 may include one or more mass storage devices (generally represented by storage device 220) for storing data in one or more databases related to the functions of the contact center. For example, storage device 220 may store customer data maintained in customer database 222. Such customer data may include customer profiles, contact information, service level agreements (SLAs), and interaction history (e.g., details of previous interactions with a particular customer, including the nature of the previous interaction, handling data, wait times, processing times, and actions taken by the contact center to resolve customer issues). As another example, storage device 220 may store agent data in agent database 223. Agent data maintained by contact center system 200 may include agent availability and agent profiles, schedules, skills, processing times, etc. As yet another example, storage device 220 may store interaction data in interaction database 224. Interaction data may include data related to numerous past interactions between customers and the contact center. More generally, it should be understood that, unless otherwise specified, storage device 220 may be configured to include databases and / or store data relating to any type of information described herein, wherein such databases and / or data can be accessed by other modules or servers of contact center system 200 in a manner that facilitates the functionality described herein. For example, servers or modules of contact center system 200 may query such databases to retrieve data stored therein or transfer data thereto for storage. For example, storage device 220 may take the form of any conventional storage medium and may be locally located or operated from a remote location. For example, the database may be a Cassandra database, a NoSQL database, or an SQL database, and is managed by a database management system such as Oracle, IBM DB2, Microsoft SQL Server, Microsoft Access, or PostgreSQL.
[0042] Regarding stat server 226, it can be configured to log and aggregate data related to the performance and operational aspects of contact center system 200. This information can be compiled by stat server 226 and made available to other servers and modules, such as reporting server 248, which can then use the data to generate reports for managing operational aspects of the contact center and performing automated actions according to the functions described herein. This data may relate to the status of contact center resources, such as average wait times, abandonment rates, agent occupancy rates, and other data required for the functions described herein.
[0043] The agent device 230 of contact center 200 may be a communication device configured to interact with various components and modules of contact center system 200 in a manner that facilitates the functions described herein. For example, agent device 230 may include a telephone suitable for regular telephone calls or VoIP calls. Agent device 230 may also include a computing device configured to communicate with the server of contact center system 200 according to the functions described herein, perform data processing associated with operations, and interact with customers via voice, chat, email, and other multimedia communication mechanisms. Although Figure 2 Three such agent devices are shown, namely agent devices 230A, 230B and 230C, but it should be understood that any number of agent devices may exist.
[0044] Regarding the multimedia / social media server 234, it can be configured to facilitate media interaction (excluding voice) with client device 205 and / or server 242. Such media interaction may be related to, for example, email, voicemail, chat, video, text messaging, networking, social media, shared browsing, etc. The multimedia / social media server 234 may take the form of any IP router conventional in the art with dedicated hardware and software for receiving, processing, and forwarding multimedia events and communications.
[0045] Regarding the knowledge management server 234, it can be configured to facilitate interaction between clients and the knowledge system 238. Generally, the knowledge system 238 can be a computer system capable of receiving questions or queries and providing answers as responses. The knowledge system 238 can be included as part of a contact center system 200 or remotely operated by a third party. The knowledge system 238 may include an artificial intelligence computer system capable of answering questions posed in natural language, as is known in the art, by retrieving information from sources such as encyclopedias, dictionaries, newsletter articles, literary works, or other documents submitted to the knowledge system 238 as reference material. For example, the knowledge system 238 can be embodied in IBM Watson or a similar system.
[0046] Regarding chat server 240, it can be configured to conduct, orchestrate, and manage electronic chat communications with clients. Generally, chat server 240 is configured to implement and maintain chat sessions and generate chat transcripts. Such chat communications can be conducted by chat server 240 in a manner where the client communicates with an automated chatbot, a human agent, or both. In an exemplary embodiment, chat server 240 can be used as a chat orchestration server that schedules chat sessions between chatbots and available human agents. In such cases, the processing logic of chat server 240 can be rule-driven to leverage intelligent workload distribution among available chat resources. Chat server 240 can also implement, manage, and facilitate user interfaces (also referred to as UIs) associated with chat features, including those UIs generated at client device 205 or agent device 230. Chat server 240 can be configured to transfer chat within a single chat session between automated and human resources, such as transferring a chat session from a chatbot to a human agent or vice versa. Chat server 240 can also be coupled to knowledge management server 234 and knowledge system 238 to receive suggestions and answers to queries made by customers during chat, such as providing links to relevant articles.
[0047] Regarding web server 242, such servers may be included to provide site hosting for various social interaction sites (such as Facebook, Twitter, Instgraph, etc.) subscribed to by customers. Although depicted as part of contact center system 200, it should be understood that web server 242 may be provided and / or remotely maintained by a third party. Web server 242 may also provide web pages for businesses or organizations supported by contact center system 200. For example, customers may browse web pages and receive information about the products and services of a particular business. Within such business web pages, mechanisms may be provided for initiating interactions with contact center system 200, for example, via web chat, voice, or email. Examples of such mechanisms are widgets that may be deployed on web pages or websites hosted on web server 242. As used herein, a widget refers to a user interface component that performs a specific function. In some implementations, a widget may include a graphical user interface control that may be overlaid on a web page displayed to a customer via the Internet. Widgets may display information, such as in windows or text boxes, or include buttons or other controls that allow customers to access certain functions, such as sharing or opening files or initiating communications. In some implementations, a widget includes a user interface component with a portable portion of code that can be installed and executed within a separate web page without compilation. Some widgets may include a corresponding or additional user interface and may be configured to access various local resources (e.g., calendar or contact information on a client device) or remote resources via a network (e.g., instant messaging, email, or social network updates).
[0048] Regarding the interactive (iXn) server 244, it can be configured to manage deferred activities in the contact center and their routing to human agents for completion. As used herein, deferred activities include background work that can be performed offline, such as replying to emails, attending training sessions, and other activities that do not require real-time communication with customers. For example, the interactive (iXn) server 244 can be configured to interact with the routing server 218 to select the appropriate agent to handle each deferred activity among the deferred activities. Once assigned to a specific agent, the deferred activity is pushed to that agent, making it appear on the selected agent's agent device 230. The deferred activity may appear in workspace 232 as a task to be completed by the selected agent. The functionality of workspace 232 can be implemented via any conventional data structure such as, for example, a linked list, an array, etc. Each agent device in agent device 230 may include workspace 232, wherein workspaces 232A, 232B, and 232C are maintained in agent devices 230A, 230B, and 230C, respectively. As an example, workspace 232 can be stored in the buffer memory of the corresponding agent device 230.
[0049] Regarding the Universal Contact Server (UCS) 246, it can be configured to retrieve information stored in the customer database 222 and / or transmit information to it for storage therein. For example, the UCS 246 can be used as part of chat features to maintain a history of how chats with specific customers were handled, which can then be used as a reference for how future chats should be handled. More generally, the UCS 246 can be configured to facilitate the maintenance of a history of customer preferences, such as preferred media channels and optimal contact times. To this end, the UCS 246 can be configured to identify data related to the interaction history of each customer, such as data related to comments from agents, customer communication history, etc. Each of these data types can then be stored in the customer database 222 or on other modules and retrieved as needed according to the functional requirements described herein.
[0050] Regarding report server 248, it can be configured to generate reports from data compiled and aggregated by statistics server 226 or other sources. Such reports may include near real-time or historical reports and relate to the status and performance characteristics of contact center resources, such as, for example, average wait time, abandonment rate, and agent occupancy rate. Reports may be generated automatically or in response to specific requests from requesters (e.g., agents, administrators, contact center applications, etc.). These reports can then be used to manage contact center operations according to the functionality described herein.
[0051] Regarding media service server 249, it can be configured to provide audio and / or video services to support contact center features. Based on the functionality described herein, such features may include prompts for IVR or IMR systems (e.g., playback of audio files), hold music, voicemail / one-way recording, multi-way recording (e.g., multi-way recording of audio and / or video calls), speech recognition, dual-tone multi-frequency (DTMF) recognition, fax, audio and video transcoding, Secure Real-Time Transport Protocol (SRTP), audio conferencing, video conferencing, tutorials (e.g., enabling coaches to listen to interactions between customers and agents and enabling coaches to provide comments to agents when customers have not heard the comments), call analytics, keyword targeting, etc.
[0052] Regarding the analysis module 250, it can be configured to provide systems and methods for performing analysis on data received from multiple different data sources, as the functionality described herein may require. According to exemplary embodiments, the analysis module 250 can also generate, update, train, and modify a predictor or model 252 based on collected data, such as, for example, customer data, agent data, and interaction data. Model 252 may include behavioral models of customers or agents. Behavioral models can be used to predict, for example, customer or agent behavior in various situations, thereby allowing embodiments of the invention to tailor interactions or allocate resources to prepare predictive characteristics for future interactions based on such predictions, thereby improving the overall performance of the contact center and the customer experience. It should be understood that while the analysis module 250 is depicted as part of a contact center, such behavioral models can also be implemented on the customer system (or, as used herein, on the "customer side" of the interaction) and used for customer benefits.
[0053] According to an exemplary embodiment, the analysis module 250 can access data stored in storage device 220, including a customer database 222 and an agent database 223. The analysis module 250 can also access an interaction database 224, which stores data related to interactions and interaction content (e.g., transcriptions of detected interactions and events), interaction metadata (e.g., customer identifiers, agent identifiers, interaction media, interaction duration, interaction start and end times, department, tagged category), and application settings (e.g., interaction paths via contact centers). Furthermore, as discussed in more detail below, the analysis module 250 can be configured to retrieve data stored in storage device 220 for use, for example, developing and training algorithms and models 252 by applying machine learning techniques.
[0054] One or more of the included models 252 can be configured to predict customer or agent behavior and / or aspects related to contact center operation and performance. Furthermore, one or more of the models 252 can be used for natural language processing and, for example, include intent recognition. Model 252 can be developed based on: 1) known first-principles formulas describing the system; 2) data, generating an empirical model; or 3) a combination of known first-principles formulas and data. When developing models for use with embodiments of the invention, since first-principles formulas are often unavailable or not easily derived, it is generally preferable to build empirical models based on collected and stored data. To accurately capture the relationship between the manipulated / interference variables and the controlled variables of a complex system, it is likely preferable that model 252 be nonlinear. This is because nonlinear models can represent a curvilinear relationship between the manipulated / interference variables and the controlled variables rather than a linear one, which is common for complex systems such as those discussed herein. In view of the foregoing requirements, machine learning or neural network-based methods are currently preferred embodiments for implementing model 252. For example, advanced regression algorithms can be used to develop neural networks based on empirical data.
[0055] Analysis module 250 may also include optimizer 254. It should be understood that an optimizer can be used to minimize a “cost function” subject to a set of constraints, where the cost function is a mathematical representation of the desired objective or system operation. Since model 252 may be nonlinear, optimizer 254 may be a nonlinear programming optimizer. However, it is conceivable that the present invention can be implemented by using a variety of different types of optimization methods, individually or in combination, including but not limited to linear programming, quadratic programming, mixed-integer nonlinear programming, random programming, global nonlinear programming, genetic algorithms, particle / swarm optimization techniques, etc.
[0056] According to an exemplary implementation, model 252 and optimizer 254 may be used together within optimization system 255. For example, analysis module 250 may utilize optimization system 255 as part of an optimization process to optimize or at least enhance various aspects of contact center performance and operation. This may include aspects related to customer experience, agent experience, interaction routing, natural language processing, intent recognition, or other functionalities related to automation processes.
[0057] Figure 2The various components, modules, and / or servers (and other figures included herein) may each include one or more processors that execute computer program instructions and interact with other system components to perform the various functions described herein. Such computer program instructions may be stored in memory implemented using standard storage devices such as, for example, random access memory (RAM), or in other non-transitory computer-readable media such as, for example, CD-ROMs, flash drives, etc. Although the functionality of each server is described as being provided by a specific server, those skilled in the art will recognize that the functionality of various servers may be combined or integrated into a single server, or the functionality of a specific server may be distributed across one or more other servers, without departing from the scope of the invention. Furthermore, the terms “interaction” and “communication” are used interchangeably and generally refer to any real-time or non-real-time interaction using any communication channel, including but not limited to telephone calls (PSTN or VoIP calls), email, voicemail, video, chat, screen sharing, text messages, social media messages, WebRTC calls, etc. Access to and control of components of the communication system 200 can be influenced by a user interface (UI) that may be generated on client equipment 205 and / or agent equipment 230. As already noted, the contact center system 200 can operate as a hybrid system in which some or all components are remotely hosted, such as in a cloud-based or cloud computing environment.
[0058] Chat system
[0059] Go to Figure 3 , Figure 4 and Figure 5 This illustrates various aspects of chat systems and chatbots. As will be seen, embodiments of the invention may include, or be enabled by, such chat features, which generally enable the exchange of text messages between different parties. These parties may include field personnel, such as customers and agents, and automated processes, such as bots or chatbots.
[0060] In the background, a bot (also known as an "internet bot") is a software application that runs automated tasks or scripts over the internet. Typically, bots perform simple, structured, and repetitive tasks much faster than a human. A chatbot is a specific type of bot and, as used herein, is defined as software and / or hardware that engages in conversation through auditory or textual methods. It should be understood that chatbots are generally designed to convincingly mimic human behavior as a conversational partner. Chatbots are commonly used in conversational systems for a variety of practical purposes, including customer service or information gathering. Some chatbots use sophisticated natural language processing systems, while simpler chatbots scan for keywords in the input and then select responses from a database based on matching keyword or phrasing patterns.
[0061] Before further describing the invention, an explanation of references to system components (e.g., modules, servers, and other components) already introduced in any prior figures will be provided. Regardless of whether subsequent references include corresponding numerical identifiers used in the preceding figures, it should be understood that such references incorporate the examples described in the preceding figures and, unless otherwise expressly limited, can be implemented based on those examples or other conventional techniques capable of achieving the desired functionality, as will be understood by one of ordinary skill in the art. Therefore, for example, subsequent references to “contact center system” should be understood to refer to… Figure 2 The exemplary “Contact Center System 200” and / or other conventional technologies used to implement a contact center system. As an additional example, subsequent references to “client equipment,” “agent equipment,” “chat server,” or “computing device” should be understood as referring to respectively Figures 1 to 2 Examples include “client device 205”, “agent device 230”, “chat server 240” or “computing device 200”, as well as conventional technologies for achieving the same functionality.
[0062] Now refer to Figure 3 , Figure 4 and Figure 5 The exemplary implementations of the chat server, chatbot, and chat interface described separately discuss chat features and chatbots in more detail. While these examples are provided relative to chat systems implemented on the contact center side, such chat systems can be used on interactive clients. Therefore, it should be understood that... Figure 3 , Figure 4 and Figure 5 The exemplary chat system can be modified for similar customer-side implementations, including the use of a customer-side chatbot configured to interact with contact center agents and chatbots on behalf of customers. It should also be understood that voice communication can leverage chat features by converting text to speech and / or speech to text.
[0063] Now for specific reference Figure 3 A more detailed block diagram of a chat server 240, which can be used to implement chat systems and features, is provided. The chat server 240 can be coupled to (i.e., communicate electronically with) a client device 205 operated by a client via a data communication network 210. For example, the chat server 240 can be operated by an enterprise as part of a contact center to implement and coordinate chat sessions with customers, including automated chat and chat with human agents. Regarding automated chat, the chat server 240 can host chat automation modules or chatbots 260A-260C (collectively referred to as 260), which are configured with computer program instructions for participating in chat sessions. Thus, generally speaking, the chat server 240 implements chat functionality, including the exchange of text-based or chat communication between the client device 205 and the agent device 230 or chatbot 260. As discussed in more detail below, the chat server 240 may include a client interface module 265 and an agent interface module 266, which are respectively used to generate specific UIs at the client device 205 and agent device 230 to facilitate chat functionality.
[0064] Regarding the chatbots 260, each can operate as an executable program launched on demand. For example, chat server 240 can operate as the execution engine of chatbot 260, similar to loading a VoiceXML file onto a media server for use in interactive voice response (IVR) functionality. Loading and unloading can be controlled by chat server 240, similar to how VoiceXML scripts are controlled in an interactive voice response scenario. Chat server 240 may also provide means for capturing and collecting customer data in a uniform manner, similar to customer data capture in an IVR scenario. Regardless of whether the same chatbot, different chatbots, agent chat, or even different media types are used, such data can be stored, shared, and used in subsequent sessions. In an exemplary embodiment, chat server 240 is configured to coordinate data sharing between various chatbots 260 when the interaction shifts or transitions from one chatbot to another or from one chatbot to a human agent. Data captured during interaction with a particular chatbot can be transmitted along with a request to invoke a second chatbot or human agent.
[0065] In an exemplary implementation, the number of chatbots 260 may vary depending on the design and functionality of the chat server 240, and is not limited to this. Figure 3The number shown. Furthermore, different chatbots can be created with different profiles, and selection can then be made between these profiles to match a specific chat or a specific customer's topic. For example, a particular chatbot's profile may include expertise in assisting the customer on a specific topic or a communication style tailored to a specific customer's preferences. More specifically, one chatbot may be designed to participate in a first communication topic (e.g., opening a new account for a business), while another chatbot may be designed to participate in a second communication topic (technical support for products or services offered by the business). Alternatively, chatbots may be configured to use different dialects or slang, or have different personality traits or characteristics. Using chatbots with profiles catering to specific types of customers enables more effective communication and results. Chatbot profiles can be selected based on known information about the other party, such as demographic information, interaction history, or data available on social media. Chat server 240 may host a default chatbot, which is invoked if insufficient information about the customer is available to invoke a more specialized chatbot. Optionally, different chatbots may be customer-selectable. In an exemplary embodiment, the profiles of chatbot 260 may be stored in a profile database hosted in storage device 220. Such profiles can include the chatbot's personality, demographics, area of expertise, and so on.
[0066] Client interface module 265 and agent interface module 266 can be configured to generate a user interface (UI) for display on client device 205, which facilitates chat communication between the client and chatbot 260 or a human agent. Similarly, agent interface module 266 can generate a specific UI on agent device 230 that facilitates chat communication between the agent operating agent device 230 and the client. Agent interface module 266 can also generate a UI on agent device 230 that allows the agent to monitor various aspects of the ongoing chat between chatbot 260 and the client. For example, client interface module 265 can transmit a signal to client device 205 during a chat session configured to generate a specific UI on client device 205, which may include displaying text messages sent from chatbot 260 or a human agent, as well as other non-text graphics intended to accompany the text messages, such as emoticons or animations. Similarly, agent interface module 266 can transmit a signal to agent device 230 during a chat session configured to generate a UI on agent device 230. Such UIs may include interfaces that allow agents to select non-text graphics to attach outgoing text messages to customers.
[0067] In an exemplary implementation, chat server 240 may be implemented in a layered architecture having a media layer, a media control layer, and a chatbot executed via IMR server 216 (similar to executing VoiceXML on an IVR media server). As described above, chat server 240 may be configured to interact with knowledge management server 234 to query knowledge information from the server. Queries may, for example, be based on questions received from a customer during a chat. The responses received from knowledge management server 234 may then be provided to the customer as part of the chat response.
[0068] Now for specific reference Figure 4 A block diagram of an exemplary chat automation module or chatbot 260 is provided. As shown, chatbot 260 may include several modules, including a text analysis module 270, a dialogue manager 272, and an output generator 274. It should be understood that other subsystems or modules may be described in a more detailed discussion of the operability of the chatbot, including, for example, modules related to intent recognition, text-to-speech or speech-to-text modules, and modules related to script storage, retrieval, and data field processing based on information stored in agent or customer profiles. However, in other areas of this disclosure (e.g., relative to...) Figure 6 and Figure 7 Such topics are covered more comprehensively and will not be repeated here. However, it should be understood that the public disclosures made in these areas can be used in a similar manner to implement the operability of chatbots based on the functions described herein.
[0069] The text analysis module 270 can be configured to analyze and understand natural language. In this regard, the text analysis module can be configured with a language dictionary, a syntax / semantic parser, and grammar rules for decomposing phrases provided by the client device 205 into internal syntactic and semantic representations. The configuration of the text analysis module depends on the specific profile associated with the chatbot. For example, certain words may be included in the dictionary of one chatbot but excluded from the dictionary of another.
[0070] The dialogue manager 272 receives syntactic and semantic representations from the text analysis module 270 and manages the general flow of the session based on a set of decision rules. In this respect, the dialogue manager 272 maintains the history and state of the session and generates outbound communications based on these. Communications may follow a script for a specific session path selected by the dialogue manager 272. As described in further detail below, the session path may be selected based on an understanding of the specific purpose or topic of the session. The script for the session path can be generated using any of a variety of languages and frameworks conventional in the art, such as, for example, Artificial Intelligence Markup Language (AIML), SCXML, etc.
[0071] During a chat session, the conversation manager 272 selects a response deemed appropriate at a specific point in the conversation flow / script and outputs the response to the output generator 274. In an exemplary embodiment, the conversation manager 272 may also be configured to calculate a confidence level for the selected response and provide that confidence level to the agent device 230. Each segment, step, or input in the chat communication may have a corresponding list of possible responses. Responses may be categorized based on a topic (determined using appropriate text analysis and topic detection schemes) and suggested next actions may be assigned. Actions may include, for example, a response with an answer, an additional question, transfer to a human agent for assistance, etc. The confidence level can be used to help the system determine whether the detection, analysis, and response to customer input are appropriate, or whether a human agent should be involved. For example, a threshold confidence level may be assigned based on one or more business rules to invoke human agent intervention. In an exemplary embodiment, the confidence level may be determined based on customer feedback. As mentioned above, the response selected by the conversation manager 272 may include information provided by the knowledge management server 234.
[0072] In an exemplary implementation, output generator 274 uses a semantic representation of the response provided by dialogue manager 272 to map the response to a chatbot profile or personality (e.g., by adjusting the language of the response according to the chatbot's dialect, vocabulary, or personality), and outputs the output text to be displayed at client device 205. The output text may be presented intentionally so that a client interacting with the chatbot is unaware that they are interacting with an automated process rather than a human agent. As will be seen, according to other implementations, the output text may be linked to visual representations (such as emojis or animations) integrated into the client's user interface.
[0073] Now refer to Figure 5 The example presents a webpage 280 with an exemplary implementation of chat feature 282. For example, webpage 280 may be associated with a business website and is designed to initiate interaction between a potential or current customer visiting the webpage and a contact center associated with the business. It should be understood that chat feature 282 can be generated on any type of customer device 205, including personal computing devices such as laptops, tablets, or smartphones. Furthermore, chat feature 282 can be generated as a window within a webpage or implemented as a full-screen interface. As in the example shown, chat feature 282 may be contained within a defined portion of webpage 280 and may be implemented as a widget, for example, via the aforementioned systems and components and / or any other conventional means. Generally, chat feature 282 may include an exemplary manner in which a customer enters a text message to be delivered to a contact center.
[0074] For example, webpage 280 may be accessed by a customer via a customer device, such as a customer device that provides a communication channel for chatting with a chatbot or live agent. In an exemplary embodiment, as shown, chat feature 282 includes generating a user interface on the display of the customer device, which is referred to herein as customer chat interface 284. For example, customer chat interface 284 may be generated by a customer interface module of a chat server, such as a chat server already described. As described, customer interface module 265 may send a signal to customer device 205 configured to generate the desired customer chat interface 284, for example, based on the content of chat messages published by a chat source, in this example, a chatbot or agent named "Kate". Customer chat interface 284 may be contained in a designated area or window that covers a designated portion of webpage 280. Customer chat interface 284 may also include a text display area 286, which is a dedicated area for displaying received and sent text messages over time. Customer chat interface 284 also includes a text input area 288, which is a designated area where the customer enters the text for their next message. It should be understood that other configurations are also possible.
[0075] Customer automation system
[0076] Embodiments of the present invention include systems and methods for automating and enhancing customer actions during various stages of interaction with a customer service provider or contact center. As will be seen, those various stages of the interaction can be categorized as a pre-contact stage, a contact stage, and a post-contact stage (or pre-interaction stage, contact stage, and post-interaction stage, respectively). Specific reference will now be made to... Figure 6 An exemplary customer automation system 300 that can be used with embodiments of the present invention is shown. Reference will also be made to better explain how the customer automation system 300 operates. Figure 7 , Figure 7 A flowchart 350 is provided for an exemplary method of automating customer actions, such as when a customer interacts with a contact center. Additional information relating to customer automation is provided in the following patent application: U.S. Application Serial No. 16 / 151,362, filed October 4, 2018, entitled “System and Method for Customer Experience Automation”.
[0077] Figure 6The term "customer automation system 300" generally refers to a system that can be used for customer-side automation. As used herein, customer-side automation refers to the automation of actions taken on behalf of a customer when interacting with a customer service provider or contact center. Such interactions may also be referred to as "customer-contact center interactions" or simply "customer interactions." Furthermore, in discussing such customer-contact center interactions, it should be understood that references to "contact center" or "customer service provider" generally refer to any customer service department or other service provider (such as, for example, a company, government agency, non-profit organization, school, etc.) associated with an organization or enterprise with whom users or customers have business, transaction, business, or other interests.
[0078] In an exemplary embodiment, the customer automation system 300 may be implemented as a software program or application running on a mobile device or other computing device, a cloud computing device (e.g., a computer server connected to the client device 205 via a network), or a combination thereof (e.g., some modules of the system are implemented in a local application, while others are implemented in the cloud). For convenience, the embodiments are described primarily in the context of the specific implementation via an application running on the client device 205. However, it should be understood that embodiments of the invention are not limited thereto.
[0079] The customer automation system 300 may include several components or modules. Figure 6 In the example shown, the customer automation system 300 includes a user interface 305, a natural language processing (NLP) module 310, an intent inference module 315, a script storage module 320, a script processing module 325, a customer profile database or module (or simply "customer profile") 330, a communication manager module 335, a text-to-speech module 340, a speech-to-text module 342, and an application programming interface (API) 345, each of which will also refer to Figure 7 The flowchart 350 describes this in more detail. It should be understood that some of the components and functions associated with the customer automation system 300 may be related to those described above. Figure 3 , Figure 4 and Figure 5 The chatbot systems overlap. When the customer automation system 300 and such chatbot systems are used together as part of a customer-side implementation, this overlap may include resource sharing between the two systems.
[0080] In the example of the operation, please refer to the specific details now. Figure 7In flowchart 350, the customer automation system 300 may receive input at the initial step or operation 355. Such input may come from several sources. For example, the primary source of input may be the customer, where such input is received via the customer's device. Input may also include data received from other parties, particularly those interacting with the customer via the customer's device. For example, information or communications sent to the customer from a contact center may provide various aspects of the input. In either case, input may be provided in free speech or text form (e.g., unstructured natural language input). Input may also include other forms of data received or stored on the customer's device.
[0081] Continuing with flowchart 350, at operation 360, the customer automation system 300 uses NLP module 310 to parse the natural language input and thereby uses intent inference module 315 to infer intent. For example, if the input is provided as speech from a customer, the speech can be transcribed into text by a speech-to-text system (such as large-vocabulary continuous speech recognition or LVCSR systems) as part of the analysis by NLP module 310. Transcription can be performed locally on customer device 205, or the speech can be transmitted over a network for conversion into text by a cloud-based server. In some implementations, for example, intent inference module 315 can automatically infer the customer's intent from the text of the provided input using artificial intelligence or machine learning techniques. Such artificial intelligence techniques may include, for example, identifying one or more keywords from the customer input and searching a database of potential intents corresponding to given keywords. A database of potential intents and keywords corresponding to those intents can be automatically mined from a collection of historical interaction records. If the customer automation system 300 cannot understand the intent from the input, the customer can be presented with a selection of several intents in user interface 305. The customer can then clarify their intent by selecting one of the alternatives or request other alternatives.
[0082] After determining the customer intent, flowchart 350 proceeds to operation 365, in which customer automation system 300 loads a script associated with the given intent. Such scripts can be stored and retrieved, for example, from script storage module 320. These scripts may include a set of commands or actions, pre-written voice or text and / or parameter or data fields (also referred to as “data fields”) representing the data required for the customer automation action. For example, a script may include commands, text, and data fields that would be needed to resolve the problem specified by the customer intent. Scripts may be specific to a particular contact center and tailored to resolve a specific problem. Scripts can be organized in various ways, such as hierarchically, where all scripts related to a particular organization originate from a common “parent” script that defines common characteristics. Scripts can be generated by mining data, actions, and dialogues from previous customer interactions. Specifically, sequences of statements made during requests for resolution of specific problems can be automatically mined from a collection of historical interactions between the customer and the customer service provider. As described from the contact center agent side, systems and methods for automatically mining valid sequences of statements and comments are described in the following U.S. patent application: U.S. Patent Application No. 14 / 153,049, filed January 12, 2014, with the U.S. Patent and Trademark Office, entitled “Computing Suggested Actions in CallerAgent Phone Calls By Using Real-Time Speech Analytics and Real-Time Desktop Analytics”.
[0083] Upon retrieving the script, flowchart 350 proceeds to operation 370, where the customer automation system 300 processes or "loads" the script. This action can be performed by script processing module 325, which populates the script's data fields with appropriate customer-related data. More specifically, script processing module 325 can extract customer data relevant to the anticipated interaction, the relevance pre-determined by a script selected to correspond to the customer's intent. Data in many data fields within the script can be automatically loaded from data retrieved from the customer profile 330. It should be understood that customer profile 330 may store customer-specific data, such as the customer's name, date of birth, address, account number, authentication information, and other types of information relevant to customer service interactions. The data selected for storage in customer profile 330 may be based on data used by the customer in previous interactions and / or include data values directly obtained by the customer. In the event of any ambiguity regarding missing information in data fields or within the script, script processing module 325 may include prompts and allow the customer to manually enter the required information.
[0084] Referring again to flowchart 350, at operation 375, the loaded script can be transferred to the customer service provider or contact center. As discussed in more detail below, the loaded script may include commands and customer data necessary to automate at least a portion of the interactions conducted on behalf of the customer with the contact center. In an exemplary implementation, API 345 is used to interact directly with the contact center. The contact center may define a protocol for making general requests to its systems, and API 345 is configured to execute that protocol. Such APIs can be implemented through various standard protocols, such as the Simple Object Access Protocol (SOAP) using Extensible Markup Language (XML), the Representational State Transfer (REST) API for formatting messages using XML or JavaScript Object Notation (JSON), etc. Therefore, the customer automation system 300 can automatically generate formatted messages according to the defined protocol used for communicating with the contact center, wherein the message includes information specified by the script in the appropriate portion of the formatted message.
[0085] Robots that use intention mining to automate creation
[0086] In recent years, with several breakthroughs in artificial intelligence (AI) and computing technologies, interest in applications, automated systems, chatbots, or robots capable of engaging in natural language conversations with humans has increased. There has been significant growth in the application of AI-enabled chatbots and virtual assistants that can converse naturally with humans and perform a variety of tasks autonomously. These conversational robots work by first analyzing the user's input and then attempting to understand the meaning of that input. This is known as natural language understanding (or "NLU") and typically involves recognizing the user's intentions or "intents" and certain keywords or "entities" in the user's input utterance. Once the intentions and entities are determined, the robot can respond to the user with appropriate follow-up actions.
[0087] Various machine learning algorithms are used to train NLU models. Training typically involves teaching the system to recognize patterns present in natural language input and associate these patterns with a set of predefined intentions. The quality of the training data is a key factor in determining model performance. A sufficiently large dataset with adequate diversity in the input utterances is essential for building a good NLU model.
[0088] As used in this article, the term "bot creation" refers to the process of creating a conversational bot or chatbot with NLU capabilities. This process typically involves defining intents, identifying entities, formulating utterances, training an NLU model, testing the bot, and finally deploying it. This is usually a largely manual process that can take weeks or months to complete. Typically, the majority of this time is spent identifying intents and formulating utterances. While organizations may already have a large number of chat sessions between their customers and customer support staff (such as contact center agents), manually traversing these raw chat logs to identify intents and utterances is both time-consuming and costly.
[0089] As used herein, an intent mining engine or process (which may generally be referred to as an "intent mining process") is a system or method that makes bot creation workflows more efficient. As will be seen, the intent mining process of this invention works by mining intents from thousands of conversations and finding a robust and diverse set of utterances belonging to each conversation. Furthermore, the intent mining process helps gain a deep understanding of the conversation by providing conversation analysis. The process also provides bot creators with the opportunity to analyze intents and make modifications. Finally, these intents and utterances can be exported to a variety of chatbot creation platforms, such as those commercially available in Genesys Conversation Engine, Google's Dialogflow, and Amazon Lex. As will be seen, this results in a flexible and efficient bot creation workflow with a significant reduction in total development time.
[0090] Now for reference Figure 8 The present intent mining process (or simply the "current intent mining process") illustrates the various stages or steps of the robot creation workflow 400. To initiate workflow 400, sessions or session data can be imported for mining. Such session data may consist of prior interactions between an agent and a customer. This session data can be a natural language session consisting of multiple rounds of back-and-forth messaging. For example, the session may have occurred via a chat interface, text, or voice call. In the latter case, the session can be transcribed into text via speech recognition before mining begins.
[0091] At initial step 405, the robot creation workflow 400 may include importing session data (i.e., session text data) for use during intent mining. This can be done in several ways. For example, session data can be imported via a text file containing the session to be mined (in a supported format, such as JSON). Session data can also be imported from cloud storage.
[0092] At step 410, the robot creation workflow 400 may include mining intents from session data. As shown below relative to... Figure 9 and Figure 10The intent discussed can be mined using intent mining algorithms.
[0093] At step 415, the chatbot creation workflow 400 may include testing the mined intents. This may include interacting with the output of the intent mining process. That is, at this stage of the workflow, the chatbot creator interacts with the mined output for editing, which may include fine-tuning and refining the intents and associated utterances before exporting them to the chatbot for training. The chatbot creator may perform various actions on the mined output, such as, for example: selecting an intent and utterances belonging to that intent; merging two or more intents into a single intent, which may result in the merging of the utterances they selected; splitting an intent into multiple intents, which may result in the splitting of the corresponding utterances; and renaming the intent tags. At the end of this business logic-driven process, a set of modified intents and associated utterances is generated that can then be used to train the chatbot.
[0094] At step 420, the bot creation workflow 400 may include importing the mined intents and utterances into the bot. For example, the mined intents may be uploaded to a conversational bot, which can then be used to engage in automated conversations with customers. This intent mining process provides several ways to add mined or modified intents and utterances to the bot. The data can be downloaded in CSV format for easy viewing. The data can also be exported to various bot formats, providing support for a wider range of conversational AI chatbot services such as Genesys Conversation Engine, Google's Dialogflow, or Amazon Lex.
[0095] The robot creation process may also include additional steps. According to some implementations, this intent-mining process may significantly involve the steps already described above, and less so the later development phases. These later steps may include optional editing steps, robot design steps, and finally, final testing and release steps.
[0096] Now for reference Figure 9 The exemplary algorithm used to implement this intent mining engine or process 500 will now be discussed. As will be seen, the algorithm can be roughly broken down into several steps, referred to herein as: 1) identifying utterances with intent; 2) generating candidate intents; 3) identifying salient intents; 4) semantically grouping intents; 5) intent tagging; and 6) utterance intent association. Other steps may include masking personally identifiable information in the utterances. Another additional step may include computation of intent analysis. These steps will now be discussed. As will be seen, these steps will be described relative to imported conversation data (e.g., data including natural language conversations between customers interacting with customer service representatives or agents), but it should be understood that the process can also be applied to other contexts involving other types of users and conversation types.
[0097] According to step 505, this intent mining process processes session data to identify intentional rounds or utterances. As used herein, intentional utterances are those determined to be likely to include or describe customer intent. Therefore, this initial step in the intent mining process is to identify intentional utterances from a given session. For example, a session typically consists of multiple message rounds or utterances from multiple parties such as agents (which may include automated systems or robotic or human agents) and customers.
[0098] For example, a bot-generated message might look like this: “Hello, thank you for contacting us. For quality and training purposes, all chats can be monitored or recorded. We will work with you soon to help you resolve your request.” Such bot-generated messages can be safely discarded because they tend to be generic and do not clarify the intent found in the conversation. An actual conversation begins when an agent or customer sends a substantive communication or message. For example, during the interaction, the customer may explain the reason or “intent” for contacting customer service. Subsequent agent-customer conversation rounds occur based on that intent expressed by the customer.
[0099] Based on analysis of real-world customer-agent conversations, this invention includes several heuristics or strategies for identifying intent-laden utterances. For example, it has been observed that intent-laden utterances typically occur at the beginning of the conversation on the customer's side. Therefore, usually only a few initial customer utterances need to be processed to identify intent, and the remainder of the conversation can be discarded. This further helps reduce system latency and memory footprint. Furthermore, word count constraints can be used to discard other utterances that are unlikely to contain customer intent.
[0100] For example, identifying intent-driven utterances can include the following: Selecting a set of consecutive customer utterances from a conversation. This set can include customer utterances that appear at the beginning of the conversation. Additionally, a word count constraint can be used to disqualify some customer utterances within this initial set. That is, to qualify, the word count in each round must exceed a minimum threshold. Such a word count or length constraint helps discard some customer rounds that are irrelevant to the intent-mining objective, such as habitual greetings like "Hello," "Hey, hello," "How are you?" etc. For example, the minimum word count threshold could be set between 2 and 5.
[0101] This intent mining process concatenates utterances from consecutive customer rounds within a round containing intent into a single composite utterance. Before doing so, each customer round can be trimmed based on a maximum length threshold, as longer sentences tend to be disjointed or produce distracting results. For example, the maximum number of characters per sentence can be set to 50. Therefore, at the end of this step, composite utterances are obtained from each session that is likely to contain an intent expressed by the customer. If a session does not contain message rounds that meet the above criteria, it can be discarded without obtaining composite utterances from it. Since this intent mining process is used to obtain primary intents from hundreds or even thousands of sessions, it can be reliably assumed that customer intents are repeated across multiple sessions. Therefore, for greater robustness in intent identification, sessions that fail to meet the above heuristic criteria can be discarded without affecting the functionality of the system.
[0102] According to step 510, candidate intentions are generated based on the analysis of the combined discourse. That is, once the discourse from the conversation is obtained and combined from the round with the intention, the next task involves identifying possible or likely intentions, which will be referred to herein as “candidate intentions”. As used herein, a candidate intention is a textual phrase consisting of two parts: 1) an action, which is a word or phrase that indicates a tangible purpose, task, or activity, and 2) an object, which indicates those words or phrases that the action will be performed or act upon.
[0103] There are different ways to obtain these action-object pairs from discourse. It should be understood that the choice can depend on the language model and the resources available for a particular language. Typically, for example, a syntactic dependency parser is used to analyze the grammatical structure of the discourse and obtain the relationships between “head” words and “tokens” or words that modify those heads. These relationships between the tokens of a discourse and their heads, along with their part-of-speech (POS) tags, are used to identify the latent or candidate intentions of a given discourse.
[0104] For example, the process of obtaining such action-object pairs may include the following. First, a dependency parser can be used to obtain all morphemes and head pairs in the discourse. Based on this, POS tags for selecting morphemes and their associated heads are noun-verb pairs. The use of general POS tags helps to make the system language agnostic and thus extend to multiple linguistic fields.
[0105] The "action" part is typically a verbal unit with an associated POS tag. If the unit is a basic verb with a "particle" unit, then that unit forms the "phrasal verb" of the discourse. The associated "particle" unit is also included in the verbal unit. Thus, the entire phrasal verb becomes the action part of the candidate intention. The "object" part is typically a noun with an associated POS tag. If the unit is part of a compound word whose constituent units all have a "noun" POS tag, then the entire compound word is considered the object. Similarly, if the unit is part of an adjective-modifying phrase, then the entire phrase is considered the object. If the unit is associated with an appositive modifier, then all the units constituting the latter are appended to the current unit to form the object part of the candidate intention. If only general POS tags are available for language rather than general dependencies, then the "verb" and "noun" units are considered the action and object parts, respectively. As a next step, the action-object ordered pair can be morphologically reduced to transform the candidate intention into a more standard form. To further normalize, the number of word form restoration pairs can be reduced.
[0106] Therefore, one or more normalized action-object pairs can be obtained from each utterance, which together form the candidate intents for the conversation. If no such pair is obtained, the utterance is discarded. With this in mind, consider the first exemplary utterance: “I want to contact the lecturer for this course. Can you provide his email address?” In this case, candidate intents could include “contact the lecturer” and “provide his email address.” Consider the second exemplary utterance: “I just completed my bachelor’s degree course on my account yesterday, and it says I have to complete my graduation application, but when I clicked on it, I was taken to a page with a message that only showed potential scholarships. What should I do?” In this case, candidate intents could include “complete the course,” “complete the graduation application,” “the page with the message,” and “show potential scholarships.”
[0107] According to step 3.515, salient intentions are identified. As used herein, the term "salient intention" refers to a narrowed list of intentions from the candidate intentions identified in the previous step, where the narrowing is based on, for example, relevance, salience, certainty, and / or obviousness. Therefore, among this set of candidate intentions, those intentions that describe the customer's actual intentions are identified as salient intentions. It should be understood that this task is not always straightforward. In some cases, the customer's intentions may be implicit in nature. However, in other cases, there may be differing opinions about the customer's actual intentions, especially in discourses that contain multiple candidate intentions.
[0108] Consider the examples provided above. In the first example, both "contact the lecturer" and "provide an email" can be considered to describe the customer's intention. In the second example, the customer has completed his / her bachelor's degree program and is experiencing problems completing their graduation application. While this intention is more implicit, the closest explicit approximation could be the candidate intention "complete the graduation application." The decision of whether "contact the lecturer" or "provide an email" should be chosen as the intention for the first statement, or even whether "complete the program" or "complete the graduation application" should be chosen as the intention for the second statement, is best determined by business logic rather than any algorithmic formula. That is, the bot creator can apply appropriate business logic to make the final decision regarding such intentions. The bot creator can also choose to retain multiple intentions or even a hierarchy of intentions to achieve appropriate business goals or objectives within a specific business domain.
[0109] Since the goal is to make the robot creation process more efficient, this intent mining process can narrow down the list of candidate intents to the most salient ones, which the robot creator can then examine to see if they are appropriate. In this case, salientity can be defined in various ways based on different criteria. For example, according to an exemplary embodiment, the frequency of candidate intents across the entire set of utterances can be an indicator of salientity; that is, the more candidate intents there are, the higher the relevance. According to other embodiments of the invention, salient intents can be found using a criterion based on Latent Semantic Analysis (LSA). LSA is a topic modeling technique used in Natural Language Understanding (NLU) tasks. For this purpose, each utterance described in terms of candidate intent action-object pairs is considered a document. LSA then analyzes the relationships between these documents and the terms they include (i.e., action-object pairs) by generating a set of concepts associated with these documents and the terms they contain. Each concept is described in terms of candidate intents with associated weights. These weights provide a deep understanding of the relative salientity of candidate intents within each concept set.
[0110] For example, according to the present invention, the process of identifying salient intentions may include the following: First, an LSA is applied to the utterance described in terms of candidate intention action-object pairs, wherein the number of LSA components is set to a predetermined limit, such as 50. Then, candidate intentions for each concept group are sorted in descending order relative to their weights, and the top candidate intentions, such as the first 5, are selected. The selected candidate intentions obtained from each concept group are then proofread and sorted in descending order relative to their weights. Duplicate entries are then discarded, and entries with higher weights are retained. These intentions, then a predetermined number, may be considered salient candidate intentions, or simply “salient intentions.” The predetermined number may be based on the maximum number of intentions to be mined. For example, this maximum number of intentions may be determined by this intention mining process based on real-world contact center interaction patterns, or selected by the bot creator based on appropriate business logic and use cases.
[0111] According to step 520, salient intentions are semantically grouped. As will be understood, since only the syntactic structure of the utterances is used to generate candidate intentions, many salient intentions that the system may identify are semantically similar. Therefore, semantically similar salient intentions can be grouped together to obtain optimal downstream functionality. The output of this intention mining process can be used to train a Natural Language Understanding (NLU) model, which then effectively forms the "brain" of the natural language chatbot. In order for these models to identify intentions associated with different utterances, the NLU models must be trained with utterances that are syntactically different but semantically similar. Therefore, the robot creation process must enable the creation of intentions associated with utterances of sufficient diversity. Grouping semantically similar salient intentions helps to generate this diversity of mined intentions.
[0112] This step typically involves calculating the semantic similarity between salient intents, which can be accomplished as follows: First, embeddings or word embeddings associated with the text of the salient intent are calculated. It should be understood that such embeddings represent topical text, such as words, phrases, or sentences, such that semantically similar texts have similar embeddings. Such word embeddings typically involve converting text data into a numerical format via an encoding process, and various conventional encoding techniques can be used to extract such word embeddings from the text data. Embeddings can then be efficiently compared to determine a measure of semantic similarity between texts. For example, global vectors (or “GloVe”) are algorithms that can be used to obtain vector representations of words. For example, a GloVe model may have 300 dimensions. In an exemplary implementation, the word embeddings of salient intents can be calculated using an inverse document frequency (IDF) weighted average of the GloVe embeddings of the constituent morphemes. It should be understood that IDF is a numerical statistic reflecting a measure of whether a term is common or rare in a given document corpus. Used in this way, the set of all candidate intents or salient intents can be considered here as the document corpus used for IDF calculation.
[0113] Once word embeddings of the text with salient intents are obtained, these embeddings can be used to calculate the semantic similarity between pairs of salient intents. For example, cosine similarity can be used to provide a measure of semantic proximity between word embeddings in a higher-dimensional space. With word embeddings available, salient intents can be grouped according to pairs of embeddings with cosine similarity greater than a predetermined similarity threshold, which can be set between 0 and 1. It should be understood that a higher threshold results in less salient intents being grouped together, producing more homogeneous groups, while a lower threshold leads to semantically more diverse intents being grouped together, producing less homogeneous groups. For example, in selecting the maximum intent mentioned above, this homogeneity value can be preset in the system chosen by the bot creator (e.g., 0.8). In the latter case, the bot creator will be able to view multiple combinations of output intents and utterances and select the value that best suits the bot's outcome.
[0114] According to step 525, intent markers are identified. Each salient intent in a group of salient intents (or “salient intent groups”) can ultimately be a mined intent (or “mined intent”). Therefore, for each of these salient intent groups, intent markers are selected to serve as markers or identifiers for the mined intent. According to an exemplary embodiment, this marking can be accomplished by calculating the IDF of each salient intent within a given salient intent group. For this calculation, the utterance described in relation to the candidate intent is considered a document, and action-object pairs considered as individual units are considered as constituent morphemes. The salient intent in each group with the highest calculated IDF is then used as the intent paradigm or “intent marker” for that group, while other salient intents within that group are referred to as “intent substitutes”.
[0115] According to step 6.530, the utterance is associated with the mined intent (the mined intent reflected by the intent marker and each mined intent in the corresponding salient intent group). It should be understood that this next step determines the utterance associated with each mined intent in the mined intent group. Similar to the previous step, embedded semantic similarity techniques can also be used here. For example, the semantic similarity between the candidate intents derived from each utterance with intent and each salient intent within a given salient intent group is calculated. If any of the constituent candidate intents of the utterance has the highest similarity to the salient intent of the given salient intent group (which may also be referred to as the mined intent or simply intent) and is also determined to be above a minimum threshold (e.g., 0.8), then the utterance is associated with that salient intent group. Furthermore, for each salient intent group, the candidate intent of the utterance with intent that produces the highest similarity to each specific salient intent group can be included in that specific salient intent group as an "intent aid." A minimum threshold can also be required. Therefore, within this step, a specific intentional utterance is associated with one of the salient intent groups, and the constituent candidate intents of that specific intentional utterance are associated with corresponding salient intent groups as intent aids. Thus, each mined intent may include intent tags as described above, as well as one or more intent substitutions and / or one or more intent aids. It should be understood that such a formulation does not preclude the possibility of a single intentional utterance being associated with multiple intent groups. This is because a single intentional utterance can have multiple candidate intents, which are added as intent aids to different intent groups across multiple mined intents. This introduces greater flexibility and robustness in downstream functions. Bot creators can choose to retain or discard such utterances from one or more groups. It has been observed that repeating utterances across multiple intents helps to teach NLU models about the inherent confusion present in these utterances, and thus helps to build more realistic and robust models.
[0116] In another step (not illustrated), personally identifiable information in the utterance is removed or masked. To ensure customer privacy, all personally identifiable information present in the associated utterance is masked. This step can, of course, be omitted if the input session is anonymized before being provided to the current intent mining process. Such personally identifiable information may include customer name, phone number, email address, social security, etc. In addition, as an additional precaution, entities associated with geographic location, date, and numbers can be masked. For example, consider the utterance: “Hi, I want to book a flight from Washington to Miami for John Honai on August 15th.” After masking, the utterance becomes: “Hi, I want to book a flight from <person> <person> <date> <date> from <geographical location> <geographical location> to <geographical location>.” Besides protecting privacy, such masking allows bot creators to quickly identify different entities present in the utterance of intent. This helps bot creators create similar utterances, but with different slot values for these entities. This results in greater utterance diversity, which further contributes to creating better NLU models.
[0117] According to another possible step (not illustrated), intent analysis can be calculated. That is, in addition to mining intent and associated utterances, this intent mining process can also generate analysis and metrics related to session data that help businesses identify customer interaction patterns. Two such metrics are as follows.
[0118] The first analysis is intent volume analysis, which is an analysis of the extent to which a conversation processes a specific intent. This analysis can also be expressed as a percentage. Intent volume analysis can help understand the relative importance of intents based on their frequency of occurrence in conversation data. Since only a single utterance is extracted from each conversation, this metric essentially becomes the number of utterances belonging to each intent.
[0119] The second analysis is intent duration analysis, which is an analysis of the duration of sessions that process a specific intent. This analysis can also be expressed as a percentage. It should be understood that this metric helps compare intents based on the total session time associated with them. The time spent in a session is calculated as the difference between the timestamps of the last and first customer / agent rounds. The sum of the durations of the individual sessions belonging to an intent gives the duration of that intent. It should be understood that this type of analysis can assist bot creators and businesses in better understanding customer and contact center staffing.
[0120] Examples of methods for creating chatbots and intent mining will now be discussed. The method may include: receiving conversation data, wherein the conversation data includes text derived from a conversation between a customer and a customer service representative; using an intent mining algorithm to automatically mine intents from the conversation data, each mined intent including intent tags, intent alternatives, and associated utterances; and uploading the mined intents to a chatbot and using the chatbot to engage in automated conversations with other customers.
[0121] According to an exemplary embodiment, the intent mining algorithm may include analyzing utterances occurring within a session of session data to identify utterances with intent. These utterances may each comprise a turn within the session, whereby a customer is communicating in the form of customer utterances or a customer service representative is communicating in the form of customer service representative utterances. Furthermore, an intent-based utterance is defined as one of the utterances determined to be more likely to express an intent. The intent mining algorithm may also include analyzing the identified intent-based utterances to identify candidate intents. Candidate intents may each be identified as a text phrase appearing within an intent-based utterance, the text phrase having two parts: an action and an object. The action may include words or phrases describing a purpose or task, and the object may include words or phrases describing the object or thing to which the action is performed. The intent mining algorithm may also include selecting salient intents from candidate intents based on one or more criteria. The intent mining algorithm may further include grouping the selected salient intents into salient intent groups based on semantic similarity between salient intents. The intent mining algorithm may further include, for each salient intent group, selecting one salient intent as an intent marker and designating other salient intents as intent substitutes. The intent mining algorithm may further include associating the intentd utterance with a salient intent group by determining the semantic similarity between candidate intents present in the intentd utterance and intent alternatives within each salient intent group in the salient intent group. The mined intents may each include a given salient intent group within the salient intent group, each of which is defined as: a salient intent selected as an intent marker and other salient intents designated as alternative intents; and the intentd utterance associated with a given salient intent group within the salient intent group.
[0122] According to an exemplary implementation, the step of identifying intentional utterances may include: selecting a first portion of a customer utterance as the intentional utterance, and discarding a second portion of the customer utterances within the session data. The first portion of the customer utterances may be defined as a predetermined number of consecutive customer utterances that appear at the beginning of each session in the session, and the second portion may be defined as the remaining portion of each session in the session.
[0123] According to an exemplary implementation, the step of identifying intent-laden utterances may further include discarding customer utterances in the first part of the customer utterance that fail to meet word count constraints. Word count constraints may include: a minimum word count constraint, wherein customer utterances in the first part of the customer utterance with fewer words than the minimum word count constraint are discarded; and / or a maximum word count constraint, wherein customer utterances in the first part of the customer utterance with more words than the maximum word count constraint are discarded. The minimum word count constraint may include a value between 2 and 5 words. The maximum word count constraint may include a value between 40 and 50 words.
[0124] According to an exemplary implementation, the step of identifying intentional utterances may include concatenating customer utterances that appear in the first part of each session in a conversation into a combined customer utterance.
[0125] According to an exemplary implementation, the step of identifying candidate intentions may include: using a syntactic dependency parser to analyze the grammatical structure of the utterance with intention to identify head-to-word pairs, each head-to-word pair including a head word modified by a word; and using part-of-speech (hereinafter referred to as "POS") tags to tag the part of speech of the utterance with intention, and identifying the head-to-word pairs as candidate intentions, wherein the POS tag of the head word may include a noun tag, and the POS tag of the word may include a verb tag.
[0126] According to an exemplary implementation, the step of selecting salient intents from candidate intents may include selecting candidate intents that are determined to appear more frequently in the utterance containing the intent than other candidate intents in the candidate intents. One or more criteria for selecting salient intents from candidate intents may include criteria based on Latent Semantic Analysis (LSA). The step of selecting salient intents from candidate intents may include: generating a set of documents having documents corresponding to the corresponding candidate intents in the candidate intents, wherein each document in the documents covers action-object pairs defined by the corresponding candidate intent in the candidate intents; generating concept groups based on items appearing in the action-object pairs contained in the set of documents; calculating a weight value for each candidate intent in the candidate intents for each concept group in the concept groups, the weight value measuring the degree of relevance between the candidate intent of a given document in the documents and a given concept group in the concept groups; and selecting a predetermined number of candidate intents as salient intents in each concept group in the concept groups based on weight values indicating a higher degree of relevance generated from a predetermined number of candidate intents.
[0127] According to an exemplary embodiment, the step of grouping salient intents based on semantic similarity may include: calculating an embedding for each salient intent, wherein the embedding may include an encoded representation of text, wherein semantically similar texts have similar encoded representations; comparing the calculated embeddings to determine semantic similarity between pairs of salient intents; and grouping salient intents with semantic similarity above a predetermined threshold. The embedding may be calculated as the average inverse document frequency (IDF) of the global vector embeddings of the constituent center-terminal pairs of the salient intents. Comparing the calculated embeddings may include cosine similarity.
[0128] According to an exemplary implementation, the step of labeling each salient intent group in a salient intent group with an intent identifier may include selecting a representative salient intent from the salient intents within each salient intent group in the salient intent group.
[0129] According to an exemplary implementation, the step of associating utterances from session data with salient intent groups may include repeatedly performing a first process to cover each intentional utterance in the intentional utterances associated with each salient intent group in the salient intent groups. If described relative to an exemplary first case involving a first salient intent group and a second salient intent group, and a first intentional utterance containing a first candidate intent and a second candidate intent, the first process may include: calculating a semantic similarity between each of the first and second candidate intents and each of the intent substitutions in the first salient intent group; calculating a semantic similarity between each of the first and second candidate intents and each of the intent substitutions in the second salient intent group; determining which intent substitutions produce the highest calculated semantic similarity; and associating the first intentional utterance with one of the intent substitutions in the first and second salient intent groups that is determined to produce the highest calculated semantic similarity. The step of associating utterances from session data with salient intent groups may further include associating the intent substitutions that produce the highest calculated semantic similarity only if the highest calculated semantic similarity is also found to exceed a predetermined similarity threshold.
[0130] Now for reference Figure 10 This illustrates the various stages of an alternative robotic creation workflow, where the above is relative to... Figure 9 The disclosed intent mining method extends the intent seeding process. The intent mining process using the intent seeding process will be discussed below after a brief introduction. For ease of distinction, the intent mining process using intent seeding will be referred to below as "intent mining by seeding" or simply "intent mining by seeding," whereas the preceding discussion above... Figure 9The disclosed intent mining process (i.e., intent mining without seeding) will be referred to below as the "general intent mining process" or simply "general intent mining".
[0131] In normal operation, general intent mining extracts intents from conversational data (such as collections of agent-customer conversations), specifically both intent tags and utterances associated with the intent. As already discussed, this process is guided by the grammatical structure and semantic content of the conversation. For example, syntactic dependencies and POS tags can be used to find candidate intents from intent-laden utterances within a conversation, while methods such as Latent Semantic Analysis (LSA) can be used to narrow down salient intents within the utterance. Intent tags, intent substitutes, and intent aids are then obtained by associating semantically similar salient intents, which in turn helps link utterances to specific intents. As already disclosed, bot creators can use the data mined via general intent mining to train NLU models, which then drive conversational bots.
[0132] It should be understood that this general framework for intent mining is based on the assumption that the bot creator is unaware of the intents typically present in the session data, and / or has not yet developed an NLU model relative to the set of sessions or similar session domains. Therefore, the general intent mining process (e.g., the process using the intent mining engine disclosed above) essentially begins without prior domain knowledge and derives or mines intents solely based on the session content of the data.
[0133] However, this assumption does not always apply. That is, the bot creator may already know the intents within a specific domain. In this case, the bot creator can understand intents that are typically present in certain conversations or are expected to exist in a specific conversation domain. For example, this might be true in the banking or travel domain. Furthermore, in many cases, the NLU model may have already been trained, and bots such as travel or banking bots have already been released by the bot creator. In such scenarios, existing domain knowledge can be used to guide the intent mining process to mine specific intents by using an intent seeding process. As will be seen, as part of intent mining by seeding, existing domain knowledge is fed into the mining process in the form of seed intent data. This seed intent data may consist of intent tags (which may be referred to as “seed intents” or “seed intent tags”) and sample utterances associated with each. This intent mining process then uses the seed intent data to mine more utterances from the conversation data for each seed intent in the seed intents, while also finding utterances for any other salient intents that can be found in the conversation data. As mentioned above, this mining process is referred to herein as the “intention mining process by seeding” or simply “intent mining by seeding”.
[0134] As will be seen, intent mining through seeding can help bot creators quickly identify more utterances belonging to the seed intent, which can be used to train or improve NLU models. Because such systems can mine other salient intents besides a given seed intent, the process can help bot creators identify customer intents that vary across different timeframes.
[0135] Similar to the general intent mining methods discussed above, the process of intent mining using seeding can be initiated by importing session data. Generally, the other steps of this seeding-based intent mining can be the same as or similar to those disclosed above relative to general intent mining. Therefore, for the sake of brevity, the focus will be primarily on the differences between seeding-based intent mining and the methods described above relative to general intent mining. Figure 9 The places where the general intention of the process of excavation is presented.
[0136] According to the present invention, intent mining via seeding utilizes seed intent data. As used herein, seed intent data comprises one or more seed intents and a set of associated sample utterances for each of the one or more seed intents. Intent mining via seeding then processes the seed intent data with session data to obtain intent substitutions and / or additional utterances associated with the seed intents. Such intent substitutions are obtained in a manner substantially similar to that given in the sections above for generating candidate intents. In this case, the seed intent and associated sample utterances are considered to be intent-laden utterances provided by the client within the session data. Normalized action-object pairs obtained from these utterances constitute the intent substitution for each seed intent.
[0137] Once an intent substitution is obtained for each seed intent, seed intent aids are identified from the set of candidate intents derived from the session data, as described above. Figure 9 The process of intention mining via seeding, which involves finding seed intention aids and associating utterances with seed intentions, can be the same as or similar to the general intention mining process described above. As in the preceding section, embedded semantic similarity techniques can also be employed here. The similarity between the candidate intentions of each utterance with intention and the intention alternatives of each seed intention can be calculated. An intentional utterance is associated with a seed intention if: a) any of the constituent candidate intentions of the intentional utterance has the highest semantic similarity to the intention alternative of that seed intention; and b) the semantic similarity is determined to be above a minimum threshold (e.g., a score above 0.8). Furthermore, as previously stated, the candidate intention that produces the highest similarity score relative to one of the seed intentions is included in the seed intention as an "intention aid" or more specifically as a "seed intention aid."
[0138] Intent mining via seeding may also include deriving additional salient intents found within the session data that differ from those identified in the seed intent data. The process for identifying such salient intents via seeding may be the same as or similar to the general intent mining process described above. That is, candidate intents are identified, and then those candidate intents within a concept group are ranked relative to their weights, with a predetermined number of higher-weighted candidate intents from that group being selected. Among these selected candidate intents, duplicate entries are discarded, and those with higher weights are retained. In completing this step, the process for intent mining via seeding may include additional procedures from the process disclosed above relative to general intent mining. Specifically, this additional procedure includes discarding any identified candidate intents that have already been identified as seed intent aids.
[0139] The next step is to identify the utterances associated with the mined intent. Similar to the previous section, this employs an embedded semantic similarity technique. The similarity between the candidate intents of an utterance and the intents of all groups is calculated. If any of the constituent candidate intents of an utterance has the highest similarity to the intent of an intent group and is above a minimum threshold (e.g., 0.8), then the utterance is associated with that intent group. The candidate intents that produce the highest similarity are included in that group and are referred to as “intent aids.” Candidate intents that have been identified as seed intent aids are discarded from this exercise.
[0140] It should be understood that, given the above discussion of intentional excavation without sowing (i.e., relative to...), Figure 9 The general intent mining process discussed) and seeded intent mining (i.e., relative to...) Figure 10 The process of intent mining via seeding (discussed here) and the functions discussed therein are possible in several different use cases or applications. In the first case, intent mining is performed without seeding. This can be used to mine salient intents and utterances associated with those salient intents from given session data. The second case involves a hybrid approach of performing general intent mining and intent mining via seeding. It should be understood that this case can be used with given session data to mine both salient intents and associated utterances, as well as additional utterances for association with a given set of seed intents. In the third case, intent mining with seeding is used to provide focused mining of a pre-determined set of seed intents. This last case can be used to mine additional utterances for association with each seed intent within a pre-determined set of seed intents.
[0141] For details, please refer to the following: Figure 10A method 600 for intent mining using intent seeds is provided. In an exemplary embodiment, method 600 includes an initial step 605 of receiving seed intents. Each seed intent includes an intent tag and a sample utterance with the intent. At step 610, utterances with the intent are identified from session data. At step 615, candidate intents are selected from the utterances with the intent. At step 620, seed intent alternatives are identified from the sample utterances with the intent. Then, at step 625, new utterances are associated with seed intents. These steps will now be discussed in more detail in the following examples.
[0142] According to an exemplary embodiment, a computer-implemented method is provided for creating a chatbot and performing intent mining using intent seeding. The method may include: receiving session data comprising text derived from sessions, wherein each of these sessions is between a customer and a customer service representative; receiving seed intent data that may include seed intents, each seed intent including a seed intent tag and a sample of intent-related utterances associated with the seed intent; using an intent mining algorithm to automatically mine the session data to determine new utterances to be associated with the seed intents; expanding the seed intent data to include the mined new utterances associated with the seed intents; and uploading the expanded seed intent data to a chatbot and using the chatbot to automatically engage in conversations with other customers.
[0143] In terms of mining seed intentions, intention mining algorithms may include analyzing utterances occurring within a session of session data to identify intentional utterances. These utterances may each comprise a turn within the session, whereby a customer is communicating in the form of customer utterances or a customer service representative is communicating in the form of customer service representative utterances. An intentional utterance may be defined as an utterance that is determined to be more likely to express an intention. The intention mining algorithm may also include analyzing the identified intentional utterances to identify candidate intentions. Each candidate intention is identified as a text phrase appearing within an intentional utterance, the text phrase having two parts: an action and an object. The action may include words or phrases describing a purpose or task, and the object may include words or phrases describing the object or thing to which the action is performed. The intention mining algorithm may also include, for each seed intention in the seed intentions, identifying seed intention alternatives from sample intentional utterances associated with that seed intention. Seed intent substitutions are identified as text phrases appearing within a sample intentional utterance. These text phrases can include two parts: an action and an object. The action can include words or phrases describing a purpose or task, and the object can include words or phrases describing the object or thing to which the action is performed. The intent mining algorithm may further include associating intentional utterances from session data with seed intents by determining the semantic similarity between candidate intents present in the intentional utterance and seed intent substitutions belonging to each seed intent tag in the seed intent tagging.
[0144] According to an exemplary implementation, the step of identifying intentional utterances may include: selecting a first portion of a customer utterance as the intentional utterance, and discarding a second portion of the customer utterances within the session data. The first portion of the customer utterance may be defined as a predetermined number of consecutive customer utterances appearing at the beginning of each session in the session, and the second portion may be defined as the remaining portion of each session in the session. The step of identifying intentional utterances may further include discarding customer utterances in the first portion of the customer utterances that fail to meet a word count constraint. The word count constraint may include: a minimum word count constraint, wherein customer utterances in the first portion of the customer utterances with fewer words than the minimum word count constraint are discarded; and / or a maximum word count constraint, wherein customer utterances in the first portion of the customer utterances with more words than the maximum word count constraint are discarded.
[0145] According to an exemplary implementation, the step of identifying candidate intentions may include: using a syntactic dependency parser to analyze the grammatical structure of the utterance with intention to identify head-to-word pairs, each head-to-word pair including a head word modified by a word; and using part-of-speech (hereinafter referred to as "POS") tags to tag the part of speech of the utterance with intention, and identifying the head-to-word pairs as candidate intentions, wherein the POS tag of the head word may include a noun tag, and the POS tag of the word may include a verb tag.
[0146] According to an exemplary implementation, the step of identifying seed intent substitution may include: using a syntactic dependency parser to analyze the grammatical structure of a sample intent-laden utterance to identify head-to-word pairs, each head-to-word pair including a head word modified by a word; and using part-of-speech (hereinafter “POS”) tags to tag the part of speech of the sample intent-laden utterance and to identify head-to-word pairs as candidate intents, wherein the POS tag of the head word may include a noun tag and the POS tag of the word may include a verb tag.
[0147] According to an exemplary implementation, the step of associating intentional utterances from session data with seed intentions may include repeatedly performing a first process to cover each intentional utterance in the intentional utterances associated with each seed intention in the seed intentions, wherein, if described relative to an exemplary first case involving a first salient intention group and a second salient intention group and a first intentional utterance containing a first candidate intention and a second candidate intention, the first process may include: calculating a semantic similarity between each of the first candidate intention and the second candidate intention and each of the intentional substitutions in the first seed intention; calculating a semantic similarity between each of the first candidate intention and the second candidate intention and each of the intentional substitutions in the second seed intention; determining which intentional substitutions produce the highest calculated semantic similarity; and associating the first intentional utterance with one of the intentional substitutions in the first seed intention and the second seed intention that is determined to produce the highest calculated semantic similarity.
[0148] In an alternative use case, the method of the present invention includes using an intent mining algorithm to automatically mine new intents and to mine new utterances for association with a given set of seed intents. In such cases, the method may include expanding the seed intent data to include the mined new intents. In this case, the intent mining algorithm may further include: selecting salient intents from candidate intents (hereinafter “unassociated utterances with intent”) existing in utterances with intents that are not yet associated with one of the seed intents, based on one or more criteria; grouping the selected salient intents into salient intent groups based on semantic similarity between salient intents; for each salient intent group, selecting one salient intent as an intent marker and designating other salient intents as intent substitutes; and associating unassociated utterances with intents from session data with salient intent groups by determining semantic similarity between candidate intents existing in unassociated utterances with intents and intent substitutes within each salient intent group. The newly discovered intentions may each include a given set of salient intentions within a set of salient intentions, each of which is defined as: a salient intention selected as an intention marker from among the salient intentions and other salient intentions designated as alternative intentions from among the salient intentions; and an unassociated utterance with intention that becomes associated with a given set of salient intentions within the set of salient intentions.
[0149] According to an exemplary implementation, the step of identifying candidate intentions may include: using a syntactic dependency parser to analyze the grammatical structure of the utterance with intention to identify head-to-word pairs, each head-to-word pair including a head word modified by a word; and using part-of-speech (hereinafter referred to as "POS") tags to tag the part of speech of the utterance with intention, and identifying the head-to-word pairs as candidate intentions, wherein the POS tag of the head word may include a noun tag, and the POS tag of the word may include a verb tag.
[0150] According to an exemplary implementation, one or more criteria for selecting salient intents from candidate intents may include criteria based on Latent Semantic Analysis (LSA). The steps of selecting salient intents from candidate intents may include: generating a set of documents having documents corresponding to corresponding candidate intents in the candidate intents, wherein each document in the documents covers action-object pairs defined by the corresponding candidate intent in the candidate intents; generating concept groups based on items appearing in the action-object pairs contained in the set of documents; calculating a weight value for each candidate intent in the candidate intents for each concept group in the concept groups, the weight value measuring the degree of relevance between a candidate intent of a given document in the documents and a given concept group in the concept groups; and selecting a predetermined number of candidate intents as salient intents in each concept group in the concept groups based on generating weight values indicating a higher degree of relevance.
[0151] According to an exemplary embodiment, the step of grouping salient intents based on semantic similarity may include: calculating an embedding for each salient intent, wherein the embedding may include an encoded representation of text, and semantically similar texts have similar encoded representations; comparing the calculated embeddings to determine semantic similarity between pairs of salient intents; and grouping salient intents with semantic similarity above a predetermined threshold. The embeddings are calculated as the average inverse document frequency (IDF) of the global vector embeddings of the constituent center-terminal pairs of the salient intents. Comparing the calculated embeddings may include cosine similarity.
[0152] According to an exemplary implementation, the step of associating unrelated intentional utterances from session data with salient intent groups may include repeatedly performing a first process to cover each unrelated intentional utterance in the unrelated intentional utterances associated with each salient intent group in the salient intent groups. If described relative to an exemplary first case involving a first salient intent group and a second salient intent group, and a first unrelated intentional utterance containing a first candidate intent and a second candidate intent, the first process may include: calculating a semantic similarity between each of the first candidate intent and the second candidate intent and each of the intent substitutions in the first salient intent group; calculating a semantic similarity between each of the first candidate intent and the second candidate intent and each of the intent substitutions in the second salient intent group; determining which intent substitutions produce the highest calculated semantic similarity; and associating the first unrelated intentional utterance with one of the intent substitutions in the first salient intent group and the second salient intent group that is determined to produce the highest calculated semantic similarity.
[0153] Those skilled in the art will understand that many of the different features and configurations described above in conjunction with several exemplary embodiments can be selectively applied to form other possible embodiments of the invention. For the sake of brevity and considering the capabilities of those skilled in the art, not every possible iteration of the possible iterations is provided or discussed in detail, but all combinations and possible embodiments contained in the following claims or otherwise are intended to be part of this application. Furthermore, improvements, changes, and modifications will occur to those skilled in the art from the foregoing description of several exemplary embodiments of the invention. Such improvements, changes, and modifications within the scope of the art are also intended to be covered by the appended claims. Moreover, it should be apparent that the foregoing relates only to the embodiments described in this application, and many changes and modifications may be made herein without departing from the spirit and scope of this application as defined by the following claims and their equivalents.
Claims
1. A computer-implemented method for creating a conversational robot, the computer-implemented method comprising: Receive session data, which includes text derived from sessions, wherein each session is between a customer and a customer service representative; Receive seed intent data including seed intents, each seed intent including a seed intent tag and a sample of intent-laden utterances associated with the seed intent; The intent mining algorithm is used to automatically mine the session data to determine new utterances to be associated with the seed intent; Expand the seed intent data to include newly mined discourses associated with the seed intent; as well as The expanded seed intent data is uploaded to the chatbot, and the chatbot is used to conduct automated conversations with other customers; The intent mining algorithm includes: Analyze the utterances occurring within the session in the session data to identify utterances with intent, wherein: Each of the statements comprises a turn within the session, thereby the customer communicating in the form of customer statements or the customer service representative communicating in the form of customer service representative statements; and Intentional utterances are defined as those utterances that are determined to be more likely to express an intention; The identified intentional utterances are analyzed to identify candidate intentions, wherein each candidate intention is identified as a text phrase appearing within an intentional utterance in the intentional utterance, the text phrase having two parts: an action and an object, the action comprising words or phrases describing a purpose or task, and the object comprising words or phrases describing the object or thing to which the action is performed. For each of the seed intentions, a seed intention substitution is identified from the sample intentional utterances associated with the seed intention, wherein the seed intention substitution is identified as a text phrase appearing within a sample intentional utterance in the sample intentional utterance, the text phrase having two parts: action and object; The intentional utterance from the session data is associated with the seed intention by determining the semantic similarity between: the candidate intention present in the intentional utterance; and the seed intention replacement belonging to each seed intention tag in the seed intention tags.
2. The method of claim 1, wherein identifying the intentional utterance comprises: Select the first part of the customer's utterance as the utterance with intent, and discard the second part of the customer's utterance within the session data; and The first portion of the customer utterance is defined as a predetermined number of consecutive customer utterances that appear at the beginning of each session in the session, and the second portion is defined as the remainder of each session in the session.
3. The method of claim 2, wherein identifying the intentional utterance further comprises: Discard any portion of the customer statement that fails to meet the word count constraint in the first part of the customer statement; The word count constraints mentioned above include: Minimum word count constraint, wherein customer utterances in the first part of the customer utterance that have fewer words than the minimum word count constraint are discarded; and Maximum word count constraint, wherein customer utterances in the first part of the customer utterance that have more words than the maximum word count constraint are discarded.
4. The method according to claim 1, wherein identifying candidate intent comprises: A syntactic dependency parser is used to analyze the grammatical structure of the intentional utterance to identify head-morpheme pairs, each head-morpheme pair including a head word modified by a morpheme word. The POS tags of the intentional utterance are used to tag the part of speech, and the center-word pairs are identified as the candidate intentions, wherein the POS tags of the center words include noun tags, and the POS tags of the word pairs include verb tags.
5. The method of claim 1, wherein the identification of seed intent substitution comprises: A syntactic dependency parser is used to analyze the grammatical structure of the sample intentional utterance to identify head-to-word pairs, each head-to-word pair including a head word modified by a morpheme. The part-of-speech (POS) tags are used to label the part of speech of the sample with intent, and the center-word pair is identified as the candidate intent, wherein the POS tags of the center word include noun tags and the POS tags of the word pair include verb tags.
6. The method of claim 5, wherein associating the intentioned utterance from the session data with the seed intention comprises repeatedly performing a first process to cover each intentioned utterance of the intentioned utterances associated with each seed intention in the seed intentions, wherein, If described relative to an exemplary first case involving a first seed intention and a second seed intention, and a first intentional utterance containing a first candidate intention and a second candidate intention, then the first process includes: Calculate the semantic similarity between each of the first candidate intent and the second candidate intent and each of the intent substitutions in the first seed intent; Calculate the semantic similarity between each of the first candidate intent and the second candidate intent and each of the intent substitutions in the second seed intent; Determine which of the stated intention substitutions produce the highest computed semantic similarity; and The first intentional utterance is associated with one of the intentional substitutes included in the first seed intention and the second seed intention, which is determined to produce the highest calculated semantic similarity.
7. The method according to claim 1, further comprising: The intent mining algorithm is used to automatically mine new intents, and each of the new intents mined includes intent tags, intent substitutions, and associated utterances. as well as Expand the seed intent data to include newly mined intents; The intent mining algorithm further includes: According to one or more criteria, significant intentions are selected from unrelated intentional utterances, wherein the unrelated intentional utterances are the candidate intentions that exist in the intentional utterances but have not yet been associated with one of the seed intentions; Based on the semantic similarity between the salient intentions, the selected salient intentions are grouped into salient intention groups; For each of the salient intent groups, select one salient intent as the intent marker and designate other salient intents as intent substitutes; and The unrelated intentional utterances from the session data are associated with the salient intent groups by determining the semantic similarity between: the candidate intent present in the unrelated intentional utterances; and the intent substitutions within each salient intent group in the salient intent groups.
8. The method of claim 7, wherein each of the newly discovered intentions comprises: Given a salient intent group within the salient intent group, each salient intent group is defined as follows: The salient intent selected as the intent marker from the salient intents; and The other significant intents designated as replacements for the aforementioned intent; and The unrelated intentional utterance becomes associated with the given salient intention group within the salient intention group.
9. The method of claim 8, wherein identifying candidate intent comprises: A syntactic dependency parser is used to analyze the grammatical structure of the intentional utterance to identify head-morpheme pairs, each head-morpheme pair including a head word modified by a morpheme word. as well as The POS tags of the intentional utterance are used to tag the part of speech, and the center-word pairs are identified as the candidate intentions, wherein the POS tags of the center words include noun tags, and the POS tags of the word pairs include verb tags.
10. The method of claim 9, wherein selecting the salient intent from the candidate intents comprises: Generate a set of documents having documents corresponding to the corresponding candidate intents in the candidate intents, wherein each document in the set of documents covers an action-object pair defined by the corresponding candidate intent in the candidate intents; A concept group is generated based on the items appearing in the action-object pairs contained in the set of documents; For each concept group in the concept group, a weight value is calculated for each candidate intent in the candidate intent, the weight value measuring the relevance between the candidate intent of a given document in the document and a given concept group in the concept group; as well as Based on a predetermined number of candidate intentions, weight values indicating a higher degree of relevance are generated, and in each of the concept groups, the predetermined number of candidate intentions are selected as the significant intentions.
11. The method of claim 10, wherein grouping the salient intents based on the semantic similarity comprises: An embedding is computed for each of the salient intentions, wherein the embedding includes an encoded representation of the text, and semantically similar texts have similar encoded representations; The calculated embeddings are compared to determine the semantic similarity between pairs of said salient intentions; and The significant intents that have a semantic similarity higher than a predetermined threshold are grouped together.
12. The method of claim 11, wherein the embedding is calculated as the inverse document frequency average of the global vector embeddings of the constituent center-terminal pairs of the salient intent; and The embeddings calculated for the comparison include cosine similarity.
13. The method of claim 8, wherein associating the unrelated intentional utterances from the session data with the salient intent groups comprises repeatedly performing a first process to cover each unrelated intentional utterance among the unrelated intentional utterances associated with each salient intent group in the salient intent groups, wherein, If described relative to an exemplary first case involving a first salient intention group and a second salient intention group, and a first unrelated utterance with intention containing a first candidate intention and a second candidate intention, then the first process includes: Calculate the semantic similarity between each of the first candidate intent and the second candidate intent and each of the intent substitutions in the first salient intent group; Calculate the semantic similarity between each of the first candidate intent and the second candidate intent and each of the intent alternatives in the second salient intent group; Determine which of the stated intention substitutions produce the highest computed semantic similarity; and The first unrelated intentional utterance is associated with one of the intentional alternatives in the first and second salient intention groups that is determined to produce the highest calculated semantic similarity.
14. A system for automating the creation of conversational bots, the system comprising: processor; and A memory, wherein the memory stores instructions that, when executed by the processor, cause the processor to perform the following operations: Receive session data, which includes text derived from sessions, wherein each session is between a customer and a customer service representative; Receive seed intent data including seed intents, each seed intent including a seed intent tag and a sample of intent-laden utterances associated with the seed intent; The intent mining algorithm is used to automatically mine the session data to determine new utterances to be associated with the seed intent; Expand the seed intent data to include newly mined discourses associated with the seed intent; as well as The expanded seed intent data is uploaded to the chatbot, and the chatbot is used to conduct automated conversations with other customers; The intent mining algorithm includes: Analyze the utterances occurring within the session in the session data to identify utterances with intent, wherein: Each of the statements comprises a turn within the session, thereby the customer communicating in the form of customer statements or the customer service representative communicating in the form of customer service representative statements; and Intentional utterances are defined as those utterances that are determined to be more likely to express an intention; The identified intentional utterances are analyzed to identify candidate intentions, wherein each candidate intention is identified as a text phrase appearing within an intentional utterance in the intentional utterance, the text phrase having two parts: an action and an object, the action comprising words or phrases describing a purpose or task, and the object comprising words or phrases describing the object or thing to which the action is performed. For each of the seed intentions, a seed intention substitution is identified from the sample intentional utterances associated with that seed intention, wherein the seed intention substitution is identified as a text phrase appearing within a sample intentional utterance of the sample intentional utterance, the text phrase having two parts: action and object; and The intentional utterance from the session data is associated with the seed intention by determining the semantic similarity between: the candidate intention present in the intentional utterance; and the seed intention replacement belonging to each seed intention tag in the seed intention tags.
15. The system of claim 14, wherein recognizing the intentional utterance comprises: Select the first part of the customer's utterance as the utterance with intent, and discard the second part of the customer's utterance within the session data; and The first portion of the customer utterance is defined as a predetermined number of consecutive customer utterances that appear at the beginning of each session in the session, and the second portion is defined as the remainder of each session in the session.
16. The system of claim 15, wherein recognizing the intentional utterance further comprises: Discard any portion of the customer statement that fails to meet the word count constraint in the first part of the customer statement; The word count constraints mentioned above include: Minimum word count constraint, wherein customer utterances in the first part of the customer utterance that have fewer words than the minimum word count constraint are discarded; and Maximum word count constraint, wherein customer utterances in the first part of the customer utterance that have more words than the maximum word count constraint are discarded.
17. The system of claim 14, wherein identifying candidate intent comprises: A syntactic dependency parser is used to analyze the grammatical structure of the intentional utterance to identify head-morpheme pairs, each head-morpheme pair including a head word modified by a morpheme word. The POS tags of the intentional utterance are used to tag the part of speech, and the center-word pairs are identified as the candidate intentions, wherein the POS tags of the center words include noun tags, and the POS tags of the word pairs include verb tags.
18. The system of claim 14, wherein the identification seed intent substitution comprises: A syntactic dependency parser is used to analyze the grammatical structure of the sample intentional utterance to identify head-to-word pairs, each head-to-word pair including a head word modified by a morpheme. The part-of-speech (POS) tags are used to label the part of speech of the sample with intent, and the center-word pair is identified as the candidate intent, wherein the POS tags of the center word include noun tags and the POS tags of the word pair include verb tags.
19. The system of claim 18, wherein associating the intentioned utterance from the session data with the seed intention comprises repeatedly performing a first process to cover each intentioned utterance of the intentioned utterances associated with each of the seed intentions, wherein, If described relative to an exemplary first case involving a first seed intention and a second seed intention, and a first intentional utterance containing a first candidate intention and a second candidate intention, then the first process includes: Calculate the semantic similarity between each of the first candidate intent and the second candidate intent and each of the intent substitutions in the first seed intent; Calculate the semantic similarity between each of the first candidate intent and the second candidate intent and each of the intent substitutions in the second seed intent; Determine which of the stated intention substitutions produce the highest computed semantic similarity; and The first intentional utterance is associated with one of the intentional substitutes included in the first seed intention and the second seed intention, which is determined to produce the highest calculated semantic similarity.
20. The system of claim 14, further comprising: The intent mining algorithm is used to automatically mine new intents, and each of the new intents mined includes intent tags, intent substitutions, and associated utterances. as well as Expand the seed intent data to include newly mined intents; The intent mining algorithm further includes: According to one or more criteria, a significant intent is selected from unrelated intentional utterances, wherein the unrelated intentional utterances are the candidate intents that exist in the intentional utterances but have not yet been associated with one of the seed intents; Based on the semantic similarity between the salient intentions, the selected salient intentions are grouped into salient intention groups; For each of the salient intent groups, select one salient intent as the intent marker and designate other salient intents as intent substitutes; and The unrelated intentional utterances from the session data are associated with the salient intent groups by determining the semantic similarity between: the candidate intent present in the unrelated intentional utterances; and the intent substitutions within each salient intent group in the salient intent groups.
21. The system of claim 20, wherein each of the newly discovered intentions comprises: Given a salient intent group within the salient intent group, each salient intent group is defined as follows: The salient intent selected as the intent marker from the salient intents; and The other significant intents designated as replacements for the aforementioned intent; and The unrelated intentional utterance becomes associated with the given salient intention group within the salient intention group.
22. The system of claim 21, wherein identifying candidate intent comprises: A syntactic dependency parser is used to analyze the grammatical structure of the intentional utterance to identify head-morpheme pairs, each head-morpheme pair including a head word modified by a morpheme word. The POS tags of the intentional utterance are used to tag the part of speech, and the center-word pairs are identified as the candidate intentions, wherein the POS tags of the center words include noun tags, and the POS tags of the word pairs include verb tags.
23. The system of claim 22, wherein selecting the salient intent from the candidate intents comprises: Generate a set of documents having documents corresponding to the corresponding candidate intents in the candidate intents, wherein each document in the set of documents covers an action-object pair defined by the corresponding candidate intent in the candidate intents; A concept group is generated based on the items appearing in the action-object pairs contained in the set of documents; For each concept group in the concept group, a weight value is calculated for each candidate intent in the candidate intent, the weight value measuring the relevance between the candidate intent of a given document in the document and a given concept group in the concept group; as well as Based on a predetermined number of candidate intentions, weight values indicating a higher degree of relevance are generated, and in each of the concept groups, the predetermined number of candidate intentions are selected as the significant intentions.
24. The system of claim 23, wherein grouping the salient intents based on the semantic similarity comprises: An embedding is computed for each of the salient intentions, wherein the embedding includes an encoded representation of the text, and semantically similar texts have similar encoded representations; The calculated embeddings are compared to determine the semantic similarity between pairs of said salient intentions; and The significant intents that have a semantic similarity higher than a predetermined threshold are grouped together.
25. The system of claim 24, wherein the embedding is calculated as the inverse document frequency average of the global vector embeddings of the constituent center-terminal pairs of the salient intent; and The embeddings calculated for the comparison include cosine similarity.
26. The system of claim 21, wherein associating the unrelated intentional utterances from the session data with the salient intent groups comprises repeatedly performing a first process to cover each unrelated intentional utterance among the unrelated intentional utterances associated with each salient intent group in the salient intent groups, wherein, If described relative to an exemplary first case involving a first salient intention group and a second salient intention group, and a first unrelated utterance with intention containing a first candidate intention and a second candidate intention, then the first process includes: Calculate the semantic similarity between each of the first candidate intent and the second candidate intent and each of the intent substitutions in the first salient intent group; Calculate the semantic similarity between each of the first candidate intent and the second candidate intent and each of the intent alternatives in the second salient intent group; Determine which of the stated intention substitutions produce the highest computed semantic similarity; and The first unrelated intentional utterance is associated with one of the intentional alternatives in the first and second salient intention groups that is determined to produce the highest calculated semantic similarity.
Citation Information
Patent Citations
Computing suggested actions in caller agent phone calls by using real-time speech analytics and real-time desktop analytics
US20150201077A1
System and Method for Customer Experience Automation
US20190037077A1
Systems and methods relating to BOT authoring by mining intents from conversation data via intent seeding
US20220101839A1
Method for efficiently obtaining language materials for identifying dialogue intentions on large scale
CN111078893A