Systems and methods related to predictive analysis using multi-dimensional event representations in customer journeys
By classifying and coding the event attributes in the customer journey and clustering to generate vector embedding sequences, the machine learning model is trained to identify milestone events, solving the problem of inefficient processing of multidimensional event sequences in the existing technology, and achieving more accurate predictive analysis and visualization.
Patent Information
- Application Number
- CN202380082694.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-19
- Filing Date
- 2023-12-19
- Publication Date
- 2025-07-11
AI Technical Summary
The prior art is difficult to effectively process multidimensional event sequences when customers interact with websites, resulting in low predictive analysis accuracy and inefficient storage and processing of dictionary tree data structures when processing complex event sequences.
By dividing event attributes in the customer journey into low cardinality and high cardinality, generating vector embedding sequences using classification coding and clustering techniques, machine learning models are trained to identify milestone events, and identify important events through attention scores to generate concise visual models.
Improves the accuracy and efficiency of customer journey forecasting and provides a concise visual model to help enterprises optimize customer experience and KPI forecasting.
Smart Images

Figure CN120303681A_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 433,536, filed on December 19, 2022, entitled "Systems and Methods Relating to Predictive Analytics Using Multidimensional Event Representation in Customer Journeys". This application also claims the benefit of U.S. Patent Application No. 18 / 545,105, filed on December 19, 2023, also entitled "Systems and Methods Relating to Predictive Analytics Using Multidimensional Event Representation in Customer Journeys". Background of the Invention
[0003] The present invention generally relates to customer relationship services and customer relationship management via a contact center and associated cloud - based systems, including delivering customer assistance via a contact center and Internet - based service options. More specifically, but without limitation, the present invention relates to enabling interactions between websites, mobile applications, analytics, and contact centers by using techniques such as predictive analytics and / or machine learning, including representing customer journeys as multi - dimensional event sequences to enhance performance. Summary of the Invention
[0004] The present invention includes a computer-implemented method that includes the steps of: generating training data samples from corresponding journey data samples via a training data process, where each journey data sample in the journey data samples includes a customer journey represented by data describing the following: a sequence of events; for each event in the event sequence, a value associated with a corresponding event attribute in a list of event attributes; and a journey outcome; and training a machine learning model using the generated training data samples. For each training data sample, the input to the machine learning model includes a sequence of vector embeddings generated via the training data process for the sequence of events; and for each training data sample, the output of the machine learning model includes an associated journey outcome. The training data process includes generating a vector embedding for each event included within a given journey data sample in the journey data samples, the vector embedding capturing the value for each event attribute in the event attributes, by the steps of: partitioning the event attributes in the list of event attributes into low-cardinality groups and high-cardinality groups, where the partitioning is done based on whether the cardinality of the values for a given event attribute included within the training data sample is above or below a predefined cardinality threshold; for each low-cardinality group in the low-cardinality groups, categorically encoding the values included within the low-cardinality group based on the total number of unique values that occur therein; for each high-cardinality group in the high-cardinality groups: clustering the values included within the high-cardinality group to create a plurality of cluster groups; and categorically encoding the values included within the high-cardinality group based on the plurality of cluster groups in which the values reside; and considering each training data sample in the training data samples to include: a sequence of vector embeddings generated from the sequence of events included within the associated journey sample in the journey samples; and the journey outcome of the associated journey sample in the journey samples.
[0005] These and other features of the present application will become more apparent when the following detailed description of the exemplary embodiments is read in conjunction with the accompanying drawings and the appended claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] A more complete appreciation of the present invention will become more apparent when considered in conjunction with the accompanying drawings and better understood by reference to the following detailed description, in which like reference symbols indicate like components, in the drawings:
[0007] Figure 1 A schematic block diagram of a computing device depicting an exemplary embodiment of the present invention and / or an exemplary embodiment that can be implemented or practiced using the same;
[0008] Figure 2 A schematic block diagram of a communication infrastructure or contact center depicting an exemplary embodiment of the present invention and / or an exemplary embodiment that can be implemented or practiced using the same;
[0009] Figure 3 is a simplified flowchart showing the functions of a machine learning model according to an embodiment of the present invention;
[0010] Figure 4 is a schematic diagram of a machine learning model for an exemplary operation according to an embodiment of the present invention; and
[0011] Figure 5 is a method for generating training data to train a machine learning model according to an embodiment of the present invention. Detailed Description
[0012] For the purpose of facilitating an understanding of the principles of the present invention, reference will now be made to exemplary embodiments shown in the accompanying drawings, and these embodiments will be described using specific language. However, it will be apparent to those of ordinary skill in the art that the detailed materials provided in the examples may not be necessary for practicing the present invention. In other cases, well-known materials or methods have not been described in detail to avoid obscuring the present invention. Additionally, as would be commonly thought by those skilled in the art, further modifications to the examples provided herein or the application of the principles of the present invention can be envisioned. In one or more embodiments or examples, specific features, structures, or characteristics may be combined in any suitable combination and / or sub-combination. Those skilled in the art will recognize that the various embodiments may be computers implemented using many different types of data processing equipment, where the embodiments are implemented as apparatus, methods, or computer program products. Thus, the example embodiments may take the form of hardware embodiments, software embodiments, or a combination thereof.
[0013] Computing device
[0014] The present invention may be a computer implemented using different forms of data processing equipment (e.g., digital microprocessors and associated memories) that execute appropriate software programs. Considering the background factors, Figure 1 illustrates a schematic block diagram of an exemplary computing device 100 according to an embodiment of the present invention and / or by which those embodiments may be implemented or practiced.
[0015] The computing device 100 can be implemented, for example, via firmware (e.g., application-specific integrated circuit), hardware, or a combination of software, firmware, and hardware. Each of the servers, controllers, switches, gateways, engines, and / or modules (which may be collectively referred to as servers or modules) in the figures below can be implemented via one or more computing devices in the computing device 100. For example, various servers can be processes running on one or more processors of one or more computing devices 100, which can execute computer program instructions and interact with other systems or modules to perform the various functions described herein. Unless otherwise explicitly restricted, functions described with respect to multiple computing devices can be integrated into a single computing device, or the various functions described with respect to a single computing device can be distributed across several computing devices. Additionally, with respect to the computing systems described in the figures below (such as, for example Figure 2 the contact center 200), the various servers and their computing devices can be located on local computing devices 100 (i.e., on-site or at the same physical location as the contact center agent), on remote computing devices 100 (i.e., off-site or in a cloud computing environment, e.g., in a remote data center connected to the contact center via a network), or some combination thereof. Functions provided by servers on off-site computing devices can be accessed and provided via a virtual private network (VPN), as if such servers were on-site, or software as a service (SaaS) (which uses various protocols to access via the Internet) can be used to provide functions, such as by exchanging data via extensible markup language (XML), JSON, etc.
[0016] As shown in the illustrated example, the computing device 100 can include a central processing unit (CPU) or processor 105 and a main memory 110. The computing device 100 can also include a storage device 115, a removable media interface 120, a network interface 125, an I / O controller 130, and one or more input / output (I / O) devices 135, which, as shown, can include a display device 135A, a keyboard 135B, and a pointing device 135C. The computing device 100 can also include additional elements, such as a memory port 140, a bridge 145, I / O ports, one or more additional input / output devices 135D, 135E, 135F, and a cache memory 150 that communicates with the processor 105.
[0017] Processor 105 can be any logic circuit that responds to and processes instructions fetched from main memory 110. For example, processor 105 can be implemented by an integrated circuit (e.g., a microprocessor, a microcontroller, or a graphics processing unit) or implemented in a field programmable gate array or an application specific integrated circuit. As depicted, processor 105 can communicate directly with cache memory 150 via an auxiliary bus or a backside bus. Main memory 110 can be one or more memory chips capable of storing data and allowing the stored data to be accessed by central processing unit 105. Storage device 115 can provide storage for an operating system that controls scheduling tasks and controls access to system resources, as well as other software. Unless otherwise restricted, computing device 100 can include an operating system and software capable of performing the functions described herein.
[0018] As depicted in the illustrated example, computing device 100 can include various I / O devices 135, one or more of which can be connected via I / O controller 130. Input devices can include, for example, keyboard 135B and pointing device 135C, such as a mouse or a stylus. For example, output devices can include a video display device, speakers, and a printer. More generally, I / O devices 135 can include any conventional devices for performing the functions described herein.
[0019] Unless otherwise restricted, computing device 100 can be any workstation, desktop computer, laptop or notebook computer, server, virtualized machine, mobile phone or smartphone, portable telecommunications device, media playback device, or any other type of computing, telecommunications, or media device capable of (but not limited to) performing the operations and functions described herein. Computing device 100 can include multiple such devices that are network connected or connected to other systems and resources via a network. Unless otherwise restricted, computing device 100 can communicate with other computing devices 100 using any conventional communication protocol via any type of network.
[0020] Contact center
[0021] Now refer to Figure 2, shows a communication infrastructure or contact center system (or simply referred to as "contact center") 200 according to an exemplary embodiment of the present invention and / or an exemplary embodiment that can be implemented or practiced using the same. Considering the background factors, customer service providers generally provide various types of services through contact centers. Such contact centers may be staffed with employees or customer service agents (or simply referred to as "agents"), where the agents act as intermediaries between a company, enterprise, government agency, or organization (hereinafter interchangeably referred to as "organization" or "enterprise") and an individual such as a user, individual, or customer (hereinafter interchangeably referred to as "individual" or "customer"). For example, an agent at a contact center may assist a customer in making a purchase decision, receiving an order, or resolving issues with a received product or service. Within a contact center, such interactions between agents and customers can occur over various communication channels, such as, for example, via voice (e.g., a telephone call or an Internet Protocol voice or VoIP call), video (e.g., a video conference), text (e.g., an email and a text chat), screen sharing, co-browsing, etc.
[0022] In operation, contact centers generally strive to provide high-quality services to customers while minimizing costs. For example, one way to operate a contact center is to handle each customer's interaction with a live agent. While this approach may score well in terms of service quality, it can also be very expensive due to the high cost of agent labor. Therefore, most contact centers utilize automated processes to replace live agents, such as interactive voice response (IVR) systems, interactive media response (IMR) systems, Internet robots or "bots", automated chat modules or "chatbots", etc.
[0023] Specifically referring to Figure 2 , a customer service provider can use the contact center 200 to provide various types of services to customers. For example, the contact center 200 can be used to participate in and manage the interactions of automated processes (or bots) or human agents communicating with customers. The contact center 200 can be an in-house facility of a company or enterprise for performing sales and customer service functions with respect to the products and services available through the enterprise. On the other hand, the contact center 200 can be operated by a service provider contracted to provide customer-related services to a company or organization. Additionally, the contact center 200 can be deployed on equipment dedicated to an enterprise or a third-party service provider and / or deployed in a remote computing environment, such as, for example, a private or public cloud environment having an infrastructure for supporting multiple contact centers for multiple enterprises. The contact center 200 can include software applications or programs that can be executed on-site or remotely or a combination thereof. It should also be understood that the various components of the contact center 200 can be distributed across various geographical locations.
[0024] Unless otherwise explicitly restricted, any one of the computing elements of the present invention can also be implemented in a cloud-based computing environment or a cloud computing environment. As used herein, "cloud computing" (or simply "the cloud") is defined as a model for enabling ubiquitous, convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, servers, storage devices, applications, and services), which can be rapidly provisioned via virtualization and released with minimal management effort or service provider interaction, and then scaled accordingly. Cloud computing can consist of various characteristics (e.g., on-demand self-service, broad network access, resource pooling, rapid elasticity, measurable services, etc.), service models (e.g., software as a service ("SaaS"), platform as a service ("PaaS"), infrastructure as a service ("IaaS")), and deployment models (e.g., private cloud, community cloud, public cloud, hybrid cloud, etc.). The cloud execution model is typically referred to as "serverless architecture", which generally includes a service provider that dynamically manages the allocation and configuration of remote servers to achieve the desired functionality.
[0025] According to Figure 2 the illustrated example, the components or modules of the contact center 200 may include: a plurality of customer devices 205; a communication network (or simply "network") 210; a switch / media gateway 212; a call controller 214; an interactive media response (IMR) server 216; a routing server 218; a storage device 220; a statistics server 226; a plurality of agent devices 230, each agent device having a work pod 232; a multimedia / social media server 234; a knowledge management server 236, which is coupled to a knowledge system 238; a chat server 240; a web server 242; an interaction server 244; a universal contact server (or "UCS") 246; a reporting server 248; a media service server 249; and an analytics module 250. It should be understood that any one of the computer-implemented components, modules, or servers described with respect to Figure 2 or in any of the following figures can be implemented via a computing device (such as Figure 1 the computing device 100). As will be seen, the contact center 200 generally manages resources (e.g., personnel, computers, telecommunications equipment, etc.) to achieve service delivery via telephone, email, chat, or other communication mechanisms. Figure 2Each component, module, and / or server of (and other figures included herein) may each include one or more processors that execute computer program instructions and interact with other system components to perform the various functions described herein. Additionally, the terms "interact" and "communicate" may be used interchangeably and generally refer to any real-time and non-real-time interaction using any communication channel, including but not limited to phone calls (PSTN or VoIP calls), emails, voicemails, videos, chats, screen sharing, text messages, social media messages, WebRTC calls, etc. Access to and control of the components of the contact system 200 may be affected through a user interface (UI) that may be generated on the customer device 205 and / or the agent device 230.
[0026] A customer desiring to receive service from the contact center 200 may initiate an inbound communication (e.g., a phone call, an email, a chat, etc.) to the contact center 200 via the customer device 205. Although Figure 2 two such customer devices are shown, it should be understood that any number of customer devices may exist. The customer device 205 may be, for example, a communication device such as a phone, a smartphone, a computer, a tablet, or a laptop. Depending on the functions described herein, a customer may generally use the customer device 205 to initiate, manage, and conduct communications with the contact center 200, such as phone calls, emails, chats, text messages, web browsing sessions, and other multimedia transactions. Inbound and outbound communications to and from the customer device 205 may traverse the network 210, where the nature of the network generally depends on the type of customer device used and the form of the communication. For example, the network 210 may include communication networks for phone, cellular, and / or data services. The network 210 may be a private or public switched telephone network (PSTN), a local area network (LAN), a private wide area network (WAN), and / or a public WAN such as the Internet. Additionally, the network 210 may include a wireless carrier network that includes a code division multiple access network, a global system for mobile communications (GSM) network, or any wireless network / technology conventional in the art.
[0027] The switch / media gateway 212 can be coupled to the network 210 for receiving and transmitting telephone calls between customers and the contact center 200. The switch / media gateway 212 can include a telephone or communication switch that is configured to act as a central switch for agent routing within the center. The switch can be a hardware switching system or implemented via software. For example, the switch 215 can include an automatic call distributor, a private branch exchange (PBX), an IP-based software switch, and / or any other switch with dedicated hardware and software that is configured to receive interactions from the Internet source and / or the telephone network source from customers and route those interactions to, for example, one of the agent devices 230. Generally speaking, the switch / media gateway 212 establishes a voice connection between the customer and the agent by establishing a connection between the customer device 205 and the agent device 230. The switch / media gateway 212 can be coupled to a call controller 214 that, for example, acts as an adapter or interface between the switch and other routing, monitoring, and communication processing components of the contact center 200. The call controller 214 can be configured to process PSTN calls, VoIP calls, etc. The call controller 214 can include computer telephony integration (CTI) software for interacting with the switch / media gateway and other components. The call controller 214 can extract data about incoming interactions, such as the customer's telephone number, IP address, or email address, and then communicate this data with other contact center components when processing the interaction.
[0028] The Interactive Media Response (IMR) server 216 implements self-service or virtual assistant functions. Specifically, the IMR server 216 can be similar to an Interactive Voice Response (IVR) server, except that the IMR server 216 is not limited to voice and can also cover various media channels. In an example showing voice, the IMR server 216 can be configured with an IMR script for asking customers about their needs. By continuing to interact with the IMR server 216, customers can receive services without speaking to an agent. The IMR server 216 can determine the reason for the customer to contact the contact center in order to route the communication to the appropriate resources.
[0029] The routing server 218 routes incoming interactions. For example, once it is determined that an inbound communication should be handled by an agent, the functionality within the routing server 218 can select the most appropriate agent and route the communication to that agent. This type of functionality can be referred to as predictive routing. This agent selection can be based on which available agent is most suitable for handling the communication. More specifically, the selection of the appropriate agent can be based on a routing strategy or algorithm implemented by the routing server 218. In doing so, the routing server 218 can query data related to the incoming interaction, such as data related to a specific customer, available agents, and interaction type, which, as described in more detail below, can be stored in a specific database. Once an agent is selected, the routing server 218 can interact with the call controller 214 to route (i.e., connect) the incoming interaction to the corresponding agent device 230. As part of this connection, information about the customer can be provided to the selected agent via the agent device 230 of the selected agent, which can enhance the service that the agent can provide.
[0030] Regarding data storage, the contact center 200 can include one or more mass storage devices (generally represented by the storage device 220), which are used to store data in one or more databases. For example, the storage device 220 can store customer data held in the customer database 222. Such customer data can include customer profiles, contact information, service level agreements (SLAs), and interaction histories (e.g., details of previous interactions with a specific customer, including the nature of the previous interactions, disposition data, wait times, handling times, and actions taken by the contact center to resolve customer issues). As another example, the storage device 220 can store agent data in the agent database 223. Agent data maintained by the contact center 200 can include agent availability and agent profiles, schedules, skills, average handling times, and so on. As another example, the storage device 220 can store interaction data in the interaction database 224. Interaction data can include data related to many past interactions between the customer and the contact center. More generally, it should be understood that unless otherwise specified, the storage device 220 can be configured to include databases and / or store data related to any type of information described herein, where these databases and / or data can be accessed by other modules or servers of the contact center 200 in a manner that facilitates the functions described herein. For example, servers or modules of the contact center 200 can query such databases to retrieve data stored therein or transmit data thereto for storage.
[0031] The statistics server 226 can be configured to record and aggregate data related to the performance and operational aspects of the contact center 200. Such information can be compiled by the statistics server 226 and made available to other servers and modules such as the reporting server 248, which can then generate reports that are used to manage the operational aspects of the contact center and perform automated actions in accordance with the functions described herein. Such data can relate to the status of contact center resources, for example, average wait time, abandonment rate, agent occupancy rate, and other data as required by the functions described herein.
[0032] The agent device 230 of the contact center 200 can be a communication device that is configured to interact with the various components and modules of the contact center 200 to facilitate the functions described herein. For example, the agent device 230 can include a telephone suitable for conventional telephone calls or VoIP calls. The agent device 230 can also include a computing device that is configured to communicate with the servers of the contact center 200 in accordance with the functions described herein, perform data processing associated with operations, and interact with customers via voice, chat, email, and other multimedia communication mechanisms. Although only two such agent devices are shown, any number of agent devices can exist.
[0033] The multimedia / social media server 234 can be configured to facilitate media interactions (excluding voice) with the customer device 205 and / or the server 242. Such media interactions can be related to, for example, email, voicemail, chat, video, text messaging, web, social media, co-browsing, etc. The multimedia / social media server 234 can take the form of any IP router in the art that has dedicated hardware and software for receiving, processing, and forwarding multimedia events and communications.
[0034] The knowledge management server 234 can be configured to facilitate interactions between the customer and the knowledge system 238. Generally speaking, the knowledge system 238 can be a computer system capable of receiving questions or queries and providing answers as a response. The knowledge system 238 can include an artificial intelligence computer system that is capable of answering questions posed in natural language by retrieving information from information sources such as encyclopedias, dictionaries, newswire articles, literary works, or other documents submitted to the knowledge system 238 as reference materials, as is known in the art.
[0035] The chat server 240 can be configured to conduct, orchestrate, and manage electronic chat communications with customers. Such chat communications can be conducted by the chat server 240 in a manner where the customer communicates with an automated chatbot, a human agent, or both. The chat server 240 can be used as a chat orchestration server that schedules chat sessions between chatbots and available human agents. In such cases, the processing logic of the chat server 240 can be rule-driven to enable intelligent workload distribution among available chat resources. The chat server 240 can also implement, manage, and facilitate a user interface (also referred to as UI) associated with chat features. The chat server 240 can be configured to transfer a chat between an automated source and a human source within a single chat session with a particular customer. The chat server 240 can be coupled to a knowledge management server 234 and a knowledge system 238 to receive suggestions and answers to queries raised by customers during a chat, such that, for example, links to relevant articles can be provided.
[0036] The web server 242 provides site hosting to various social interaction sites (such as Facebook, Twitter, Instagram, etc.) subscribed to by customers. Although depicted as part of the contact center 200, it should be understood that the web server 242 can be provided and / or remotely maintained by a third party. The web server 242 can also provide web pages for an enterprise or organization being supported by the contact center 200. For example, a customer can browse the web page and receive information about the products and services of a particular enterprise. Within such an enterprise web page, mechanisms can be provided for initiating an interaction with the contact center 200, for example, via web chat, voice, or email. An example of such a mechanism is a desktop applet that can be deployed on a web page or website hosted on the web server 242. As used herein, a desktop applet refers to a user interface component that performs a specific function. In some specific implementations, the desktop applet includes a GUI that overlays a web page displayed to a customer via the Internet. The desktop applet can display information, for example, in a window or text box, or include buttons or other controls that allow a customer to access certain functions such as sharing or opening files or initiating a communication. In some specific implementations, the desktop applet includes a user interface component that has a portable portion of code that can be installed and executed within a separate web page without compilation. Such desktop applets can include additional user interfaces and can be configured to access various local resources (such as calendar or contact information on a customer device) or remote resources via a network (e.g., instant messaging, email, or social network updates).
[0037] The interaction server 244 is configured to manage deferrable activities of the contact center and their routing to agents for completion. As used herein, deferrable activities include back-office work that can be performed offline, such as replying to emails, attending training, and other activities that do not require real-time communication with customers.
[0038] The Universal Contact Server (UCS) 246 can be configured to retrieve information stored in the customer database 222 and / or transmit information thereto for storage therein. For example, the UCS 246 can be used as part of a chat feature to facilitate maintaining a history of how to handle chats with a particular customer, which can then be used as a reference for how to handle future chats. More generally, the UCS 246 can be configured to facilitate maintaining a history of customer preferences, such as preferred media channels and best contact times. To this end, the UCS 246 can be configured to identify data related to the interaction history with each customer, such as data related to comments from agents, customer communication history, etc. Then, each of these data types can be stored in the customer database 222 or on other modules and retrieved as needed according to the functions described herein.
[0039] The reporting server 248 can be configured to generate reports based on data compiled and aggregated by the statistics server 226 or other sources. Such reports can include near-real-time reports or historical reports and relate to the status and performance characteristics of contact center resources, such as, for example, average wait time, abandonment rate, agent occupancy rate. Reports can be generated automatically or in response to a request and are used to manage the contact center according to the functions described herein.
[0040] The media services server 249 provides audio services and / or video services to support contact center features. According to the functions described herein, such features can include prompts for an IVR or IMR system (e.g., playback of audio files), hold music, voicemail / one-way recording, multi-party recording (e.g., multi-party recording of audio and / or video calls), speech recognition, dual-tone multi-frequency (DTMF) recognition, audio and video transcoding, secure real-time transport protocol (SRTP), audio or video conferencing, call analytics, keyword detection, etc.
[0041] The analysis module 250 may be configured to perform analysis on data received from multiple different data sources, as may be required for the functionality described herein. The analysis module 250 may also generate, update, train, and modify predictors or models, such as machine learning model 251 and / or model 253, based on the data collected. To achieve this, the analysis module 250 may access data stored in the storage device 220, including the customer database 222 and the agent database 223. The analysis module 250 may also access the interaction database 224, which stores data related to interactions and interaction content (e.g., audio and transcripts of detected interactions and events therein), interaction metadata (e.g., customer identifier, agent identifier, interaction media, interaction duration, interaction start and end times, department, tagged categories), and application settings (e.g., interaction paths through the contact center). The analysis module 250 may retrieve such data from the storage device 220 for use in developing and training algorithms and models. It should be understood that while the analysis module 250 is depicted as being part of the contact center, the functionality described with respect to this analysis module may also be implemented on the customer system (or, as also used herein, on the "customer side" of the interaction) and used for the benefit of the customer.
[0042] The machine learning model 251 may include one or more machine learning models that may be based on neural networks. In some embodiments, the machine learning model 251 is configured as a deep learning model, which is a type of machine learning based on neural networks where multiple processing layers are used to extract progressively higher-level features from data. For example, the machine learning model 251 may be configured to predict behavior. Such behavior models may be trained to predict the behavior of customers and agents in various situations so that interactions can be tailored to the customer and handled more effectively by the agent. As another example, the machine learning model 251 may be configured to predict aspects related to contact center operations and performance. In other cases, for example, the machine learning model 251 may also be configured to perform natural language processing and, for example, provide intent recognition and the like.
[0043] The analysis module 250 may further include an optimization system 252. The optimization system 252 may include one or more models 253, and the one or more models may include a machine learning model 251 and an optimizer 254. The optimizer 254 may be used in combination with the model 253 to minimize a cost function subject to a set of constraints, where the cost function is a mathematical representation of a desired objective or system operation. Since the model 253 is typically non-linear, the optimizer 254 may be a non-linear programming optimizer. However, it is contemplated that the optimizer 254 may be implemented by using a variety of different types of optimization methods, either alone or in combination, including but not limited to linear programming, quadratic programming, mixed integer non-linear programming, stochastic programming, global non-linear programming, genetic algorithms, particle / swarm techniques, etc. The analysis module 250 may utilize the optimization system 255 as part of an optimization process to optimize or at least enhance various aspects of the contact center performance and operation. For example, this may include aspects related to customer experience, agent experience, interaction routing, natural language processing, intent recognition, system resource allocation, system analysis, or other functions related to automation processes.
[0044] Machine learning model
[0045] Figure 3 An exemplary machine learning model 300 is shown, and the exemplary machine learning model may be included in one or more embodiments of the present invention. The machine learning model 300 may be a component, a module, a computer program, a system, or an algorithm. As described below, some embodiments herein use machine learning to provide predictive analytics for applications in a contact center. The machine learning model 300 may be used as the model that powers these embodiments. The machine learning model 300 is trained with training data samples 306, and the training data samples may include input objects 310 and desired output values 312. For example, the input objects 310 and the desired object values 312 may be tensors. A tensor is an n-dimensional matrix, where n may be any one of 0 (constant), 1 (array), 2 (2D matrix), 3, 4, or more.
[0046] The machine learning model 300 has internal parameters that determine its decision boundary and the output produced by the machine learning model 300. After each training iteration, including inputting the input object 310 of the training data sample into the machine learning model 300, the actual output 308 of the machine learning model 300 for the input object 310 is compared with the expected output value 312. One or more internal parameters 302 of the machine learning model 300 can be adjusted such that when the machine learning model 300 is run with the new parameters, the output 308 produced will be closer to the expected output value 312. If the output 308 produced is already the same as the expected output value 312, the internal parameters 302 of the machine learning model 300 can be adjusted to strengthen and reinforce those parameters that cause the correct output and reduce and weaken the parameters that tend to deviate from the correct output.
[0047] For example, the output of the machine learning model 300 can be a numerical value in the case of regression or an identifier of a class in the case of a classifier. A machine learning model trained to perform regression can be referred to as a regression model, and a machine learning model trained to perform classification can be referred to as a classifier. The aspects of the input object that the machine learning model 300 can consider when making its decision can be referred to as features. After the machine learning model 300 has been trained, a new, unseen input object 320 can be provided as input to the model 300. Then, the machine learning model 300 produces an output representing the predicted target value 304 of the new input object 320 based on its internal parameters 302 learned from the training.
[0048] The machine learning model 300 can be, for example, a neural network, a support vector machine (SVM), a Bayesian network, logistic regression, logistic classification, a decision tree, an ensemble classifier, or other machine learning models. The machine learning model 300 can be supervised or unsupervised. In the unsupervised case, the machine learning model 300 can identify patterns in the unstructured data 340 without training data samples 306. The unstructured data 340 is, for example, the raw data for which an inference process is desired to be performed. The unsupervised machine learning model can generate an output 342 including data identifying the structure or pattern.
[0049] A neural network can be composed of multiple neural network nodes, where each node includes an input value, a set of weights, and an activation function. The neural network node can calculate the activation function based on the input value to generate an output value. The activation function can be a non-linear function calculated based on the weighted sum of the input values plus an optional constant. In some embodiments, the activation function is a logistic function, a sigmoid function, or a hyperbolic tangent function. The neural network nodes can be connected to each other such that the output of one node is the input of another node. Additionally, the neural network nodes can be organized into multiple layers, each layer including one or more nodes. The input layer can include the input to the neural network, and the output layer can include the output of the neural network. The neural network can be trained and its internal parameters updated by using backpropagation, and these internal parameters include the weights of each neural network node.
[0050] In some embodiments, a Convolutional Neural Network (CNN) may be used. A Convolutional Neural Network is a type of neural network and machine learning model. A Convolutional Neural Network may include one or more convolutional filters (also known as kernels) that operate on the output of a previous neural network layer and produce an output for use by a subsequent neural network layer. A convolutional filter may have a window in which it operates. The window can be spatially local. If a node in the previous layer is within the window, the node in the previous layer may be connected to a node in the current layer. If the node is not within the window, the node is not connected. A Convolutional Neural Network is a locally connected neural network, which is a neural network in which neural network nodes are connected to nodes within a spatially local region of the previous layer. Additionally, a Convolutional Neural Network is a sparsely connected neural network, which is a neural network in which most nodes in each hidden layer are connected to less than half of the nodes in the subsequent layer. In other embodiments, a Recurrent Neural Network (RNN) may be used. A Recurrent Neural Network is another type of neural network and machine learning model. A Recurrent Neural Network includes at least one feedback loop, where the output of at least one neural network node is input into a neural network node in a previous layer. A Recurrent Neural Network maintains a state between iterations, such as in the form of a tensor. The state is updated at each iteration, and the state tensor is passed as an input to the Recurrent Neural Network in a new iteration. In other embodiments, the Recurrent Neural Network is a Long Short-Term Memory (LSTM) neural network. In some embodiments, the Recurrent Neural Network is a Bidirectional LSTM neural network. A Feedforward Neural Network is another type of neural network and does not have a feedback loop. In some embodiments, a Feedforward Neural Network may be densely connected, meaning that most neural network nodes in each layer are connected to most neural network nodes in the subsequent layer. In some embodiments, a Feedforward Neural Network is a fully connected neural network, where each neural network node is connected to each neural network node in the subsequent layer. A Gated Graph Sequence Neural Network (GGSNN) is a type of neural network that may be used in some embodiments. In a GGSNN, the input data is a graph, including nodes and edges between nodes, and the neural network outputs a graph. The graph can be directed or undirected. A propagation step is performed to compute a node representation for each node, where the node representation may be based on the features of the node. An output model maps from the node representation and the corresponding label to an output for each node. The output model is defined per node and is a differentiable function that maps to the output. Additionally, an embodiment may include different types or the same type of neural networks that are linked together in a series of sequential or parallel neural networks, where a subsequent neural network receives the output of one or more previous neural networks as input. Backpropagation from the last neural network to the first neural network may be used to train the combination of multiple neural networks end-to-end. As stated, the machine learning model 251 may also be configured as a deep learning model.A deep learning model is a type of machine learning based on neural networks, where multiple processing layers are used to extract increasingly higher-level features from data. Deep learning models are generally better at unsupervised learning.
[0051] Figure 4Illustrated is the execution of inference on an input 360 that includes data-related customer journeys using a machine learning model 300. For example, as discussed below, the input 360 may include customer journey data, i.e., a sequence of multi-dimensional vector embeddings generated to represent a sequence of web events describing a customer journey. The machine learning model 300 then performs inference on the data based on its internal parameters 302 learned through training. The machine learning model 300 generates an output 370 of a prediction result. In an exemplary implementation, the machine learning model 300 may be configured according to the requirements of a specific machine learning algorithm to achieve the functions described herein. For example, the machine learning model 300 may include one or more neural networks. More specifically, the machine learning model 300 may include a recurrent neural network (RNN), which is generally effective in processing sequential data, such as text, audio, or time series data. Such models are designed to remember or "store" information from previous inputs, which allows them to utilize the context and dependencies between time steps. This makes such models useful for tasks such as language translation, speech recognition, and time series prediction. In some implementations, the RNN may include a long short-term memory (LSTM) network or a gated recurrent unit (GRU). Both LSTM and GRU are designed to address the "vanishing gradient" problem in RNNs, which occurs when the gradient of the weights in the network becomes very small and the network has difficulty learning. The LSTM network is a type of RNN that uses special types of memory cells to store and output information. These memory cells are designed to remember information for a long time, and they achieve this by using a set of "gates" that control the flow of information into and out of the cell. The gates in the LSTM network are controlled by a sigmoid activation function, which outputs a value between 0 and 1. The gates allow the network to selectively store or forget information based on the input values and the previous state of the cell. On the other hand, the GRU is a simplified version of the LSTM that uses a single "update gate" to control the flow of information into the memory cell, rather than the three gates used in the LSTM. This makes the GRU easier to train and faster to run than the LSTM, but they may be less effective in storing and accessing long-term dependencies. In other implementations, the machine learning model 520 may be configured as a sequence-to-sequence model that includes a first encoder model and a decoder model. The first encoder may include an RNN, or a convolutional neural network (CNN), or another machine learning model capable of accepting a sequence input. The decoder may include an RNN, CNN, or another machine learning model capable of generating a sequence output. The sequence-to-sequence model may be trained on training data samples, where each training data sample includes a sequence of vector embeddings representing a sequence of customer journey events and a journey result. For example, the sequence-to-sequence model may be trained by inputting the input data into the first encoder model to create a first embedding vector. The first embedding vector may be input into the decoder model to create an output result of the prediction result.The output result can be compared with the actual result, and the parameters of the first encoder and decoder can be adjusted to reduce the difference between the predicted result and the actual result. The parameters can be adjusted through backpropagation. In an embodiment, the sequence-to-sequence model can include a second encoder that receives additional information related to a failed test / interrupt change to create a second embedding vector. The first embedding vector can be combined with the second embedding vector as the input to the decoder. For example, the first embedding vector and the second embedding vector can be combined using concatenation, addition, multiplication, or another function. In another example, features can include statistical information calculated based on the count and order of words or characters in data associated with specific event features (e.g., words associated with a page URL). In other embodiments, the machine learning model for outputting the predicted result can be an unsupervised model, such as a deep learning model. An unsupervised model is not trained but identifies its predictions based on the identification of patterns in the data. In an embodiment, the unsupervised machine learning model can identify common features in the training dataset.
[0052] Multidimensional event representation for predictive analytics
[0053] Deriving actionable insights from the customer journey has been a key area of research focus in the past few years. It should be understood that being able to generate a concise visualization of such a journey is an important step in better understanding the patterns found in the customer journey. For events that occur when a customer interacts with a website (which may be referred to as "web events" or simply "events"), this is extremely difficult due to the lack of structure regarding how such interactions unfold. As will be seen, the present invention provides a way to model the customer journey to identify milestone events related to a specific outcome. The milestone events can then be used to generate a guiding visualization related to the predictive sequence of such web events found in the customer journey. Additionally, through the iterative process enabled by the present invention, as noise events / parameters are identified and removed, the predictive power and accuracy of the customer journey model based on web events can be significantly enhanced. These enhanced models can then be used to provide a more accurate visualization and provide effective recommendations for the next best action.
[0054] More generally, as used herein, a customer journey refers to the sequence of events that occur when a customer interacts with a business or company. This can include the customer interacting with an automated system (e.g., a virtual agent / robot or an IVR system). As noted above, this can also include how the customer interacts with the company's website. A customer journey can be used as a way to organize information about how a customer interacts with a company over time. A customer journey is a discrete, unevenly sampled time series of customer events that contains a heterogeneous set of attributes and characteristics. They may contain both explicit commitment signals (such as buying a new car) and more ambiguous commitment signals (such as a series of monthly credit card purchases or contacts with customer service). It has been observed that the use of customer journey architectures improves machine learning capabilities, whether for recommendation systems or experience customization. The customer journey can then be used to drive sales through these recommendations and customized experiences. By using the customer journey, the customer experience can be seen as a dynamic rather than a static factor. Thus, for example, the customer experience becomes a continuously managed signal or parameter that can be used to decide, recommend, and trigger actions from the company to its customers and potential customers. For example, in the program state of machine learning or operations research applications, modeling the customer journey is used to determine data points (such as one or more new events) that minimize the knowledge gap in the customer journey. These results can be modeled to determine the "next best action" for the company to take with respect to the customer, which may result in a desired outcome (such as closing a sale) or avoid an undesired outcome (such as a customer discontinuing service).
[0055] Considerable progress has been made in effectively leveraging certain types of customer journeys. These include those types of customer journeys that are more structured in terms of how customers are guided through the customer journey, such as, for example, customer journeys involving interactions with agent bots (which may be referred to as bot flows), and customer journeys related to interactions with IVR systems (which may be referred to as IVR paths or journeys). However, beyond this, there is still a great need for similar support related to those customer journeys that describe how customers interact with websites. Progress in this area has proven to be a daunting challenge. The main reason for this is that while bot flows and IVR paths have a fixed structure for how events are sequenced in any given customer journey, the way customers interact with websites does not. This lack of structure results in a wide variety of variations in the possible event sequences (i.e., the ordering of events in a customer journey). For example, the use of a trie data structure has been very effective in modeling bot flows and IVR paths for prediction purposes and generating useful visualizations, but this usefulness has not been extended to customer journeys involving web events. One reason for this is that the inherent characteristics of the trie data structure make it almost impossible to scale for this type of use case. In the case of a trie data structure, a new branch is created for each new prefix of the (event sequence), which in turn results in the creation of a very large and cumbersome trie data structure. Even a slight change in the starting event or event sequence causes the size of the trie data structure to grow rapidly, which leads to a significant increase in memory usage and processing time, making it infeasible.
[0056] Events that occur when interacting with a website are typically multi-dimensional, i.e., the events contain many different attributes, and this increases the difficulty of interpreting customer journeys in this area. To handle such complex situations where the event sequences vary highly, or to consider multiple attributes (i.e., dimensions) of the events when mining for patterns for prediction or visualization purposes, the trie data structure simply does not work. Instead, as proposed in this disclosure, it has been found that more advanced machine learning algorithms (specifically sequence learning algorithms) are more effective. Such large variations in the sequence pose the problem of noisy visualizations, where it becomes crucial to identify which events are important for the visualization. Embodiments of the present invention provide solutions to these several challenges.
[0057] Turning now to specific embodiments, it should be understood that finding patterns or common paths in the customer journey is very valuable to a company, as this allows the company to make predictions that optimize the customer experience and company KPIs. However, as will be understood, events in the customer journey do not have equal importance. (Note that a reference event includes a reference to an event or a web event.) Thus, to identify common patterns in the customer journey, it is crucial to remove unimportant events (noise) and focus only on important events (milestone events) that lead to the achievement of outcomes and have predictive value. Such milestone events are events that can accurately predict the customer's next actions, desires, or needs. Of course, it is often difficult to distinguish between these two types of events, which leads to challenges in customer journey analysis, especially in the context of complex interactions involving websites.
[0058] When customer journey events are multi-dimensional, i.e., each event has several different attributes, current techniques for finding milestones and patterns in the journey simply do not perform well. This is especially true when these event attributes have a high cardinality.
[0059] When mapping the customer journey to interactions with a company website, each event can be scored against several different types of event attributes. For example, in this context, events typically involve a customer moving between different web pages of the website, submitting a search, or interacting with various aspects of a page, where the possible event attributes describing these actions can include a potentially long list of things. For example, event attributes can include characteristics describing the website, the page being visited, the order of visited pages, the search string, the referring website, etc. Preferred embodiments of the present invention can include an event attribute list for describing the attributes associated with each event, which can include web page or page attributes, customer attributes, location attributes, date attributes, attributes of the referring website, the number of event attributes, browser attributes, session attributes, user device attributes, query or search attributes, etc. These attributes can be determined for each web event that occurs within the customer journey, such as for each visited web page, desktop applet interacted with, search entered, etc., when the customer interacts with a particular website. According to an exemplary embodiment, the event attribute list can include each of the following or a subset thereof: geolocation_country; geolocation_locality; geolocation_regionName; visit_date; referrer_domain; session_referrer_url; referrer_keywords; ipAddress; visitReferrer_domain; referrer_name; visitReferrer_pathname; visitReferrer_medium; referrer_pathname; total_EventCount; session_referrer_queryString; session_referrer_hostname; referrer_hostname; browser_fingerprint; browser_featuresWebrtc; browser_viewheight; browser_featuresFlash; session_shortId; browser_featuresJava; browser_lang; device_category; device_screenheight; device_osVersion; page_domain; page_URL; totalPageviewCount; page_lang; session_type; session_pageviewCount; customerIdType; loginId;customerId; session_eventCount; device_type; page_breadcrumb; visitId; session_createdDate; device_isMobile; session_referrer_pathname; session_referrer_fragment; device_osFamily; ipOrganization; mktCampaign_content; outcomeName; visitReferrer_name; browser_family; marketingCampaign_source; referrer_queryString; outcomeId; referrer_url; browser_version; geolocation_longitude; session_id; referrer_fragment; browser_viewWidth; session_secondsSincePrevious; searchQuery; device_fingerprint; session_secondsSinceFirst; session_durationInSeconds; visitReferrer_url; page_title; session_referrer_name; visitReferrer_hostname; geolocation_timezone; geolocation_postalCode; externalContactId; mktCampaign_clickId; page_fragment; geolocation_source; geolocation_latitude; visitReferrer_keywords; page_keywords; page_queryString; session_referrer_keywords; userAgentString; referrer_medium; page_hostname; page_pathname; organizationId; and visitReferrer_fragment.;
[0060] According to an exemplary embodiment, vector embeddings are generated for each event (i.e., web event), which capture (i.e., mathematically account for or reflect) data that describes a score or value for each event attribute of a given event among these event attributes. The generation of such vector embeddings can be done in the following manner.
[0061] First, categorical encoding is performed on low-cardinality event attributes. In such cases, low-cardinality can be defined as a cardinality below a selected threshold (e.g., 10). For example, the country where a customer is located (e.g., "geolocation_country") can be one of the event attributes with low cardinality. Thus, for example, if there are five unique countries that occur as values in this event attribute, each country can be enumerated as 1, 2, 3, 4, 5 and then normalized, which forms the basis of the embedding vector.
[0062] According to an exemplary embodiment, another process can be used to handle event attributes with high cardinality. For event attributes with high cardinality, the unique values can be clustered first. For example, the event attribute associated with a page URL can have high cardinality because the name of the page URL is unique. In an exemplary embodiment, to perform clustering, features need to be extracted from the page URL first. For example, this can include removing special characters so that a string of words is retained. Then, for example, a separate vector embedding can be created for each page URL. As will be understood by those of ordinary skill in the art, this can be done via a trained machine learning model configured for this purpose. The vector embedding is created through a machine learning process in which the model is trained to transform a particular kind of data into a numerical vector. Then, another machine learning model (which is trained to perform clustering) can be used to cluster the vector embeddings based on the similarity of the vector embeddings. For example, neural network-based clustering can be used. This type of clustering uses a neural network to learn the cluster structure of the data. Examples of this method are autoencoders and deep embedding clustering. Other types of clustering algorithms can also be used, such as centroid-based clustering (which uses the mean or median of the points in the cluster as the center or centroid of the cluster), K-means (the most popular centroid-based clustering algorithm); hierarchical clustering (which constructs a hierarchy of clusters where each cluster is a subset of the next higher-level cluster), density-based clustering (which groups together points that are close to each other in the feature space) (such as DBSCAN), and distribution-based clustering (which models the data as a mixture of probability distributions) (such as Gaussian mixture model (GMM)), and spectral clustering (which uses the eigenvectors of the similarity matrix to cluster the data).
[0063] For example, in an exemplary embodiment, words that appear in a given page URL value can be considered NLP words. That is, after removing special characters, the remaining text can be used to generate vector embeddings, which can then be used to cluster similar page URLs. The clustering process can be used to significantly reduce the cardinality that appears in the values of event attributes. For example, in the example of page URLs, each value can now be represented relative to its cluster, which forms the basis of the vector embedding.
[0064] At this point, it might be asked why not generate categorical encodings for all event attributes and then cluster the multi-dimensional vectorized events. The reason for this is that the result would produce a single cluster number for each event and would thus represent the journey as a one-dimensional sequence. This would limit the effectiveness of the associated predictive analysis. The reason for not doing so is also related to interpretability. Although clustering the events as a whole generates a one-dimensional journey sequence, this does not provide an explanation of the nature of these clusters. Thus, any results obtained from path analysis and milestone events would be at the level of "cluster numbers" and would therefore not be interpretable in a meaningful way.
[0065] With the first two steps completed, the customer journey can then be represented as a sequence of events. Specifically, the customer journey is represented as a sequence of vector embeddings for the corresponding sequence of events, as obtained above. Each event in the journey is transformed into a vector embedding that captures the values that the event has in each event attribute in the list of event attributes. According to the exemplary embodiment, this sequence of vector embeddings forms the input for training a predictive model. As for the output, this may be based on whether a target result has been achieved in each particular customer journey. For example, the target result (or outcome) may be whether a sale has been made. Or, for example, the outcome may be whether a website interaction has led the customer to accept the offered chat session with a live agent. In the exemplary embodiment, a binary variable can be used to represent the output, such as "0" indicating that the result has not been achieved and "1" indicating that the result has been achieved. As will be understood, the output represents the target variable for training the predictive model.
[0066] According to the exemplary embodiment, the vector embeddings and the result data are then used to train a predictive model. For example, the model can be an attention-based bidirectional LSTM model or a Transformer model. As will be understood, the model is trained to predict the output given the input and is selected based on its ability to process low-cardinality multi-dimensional event embeddings, which are generated as described above.
[0067] In an alternative embodiment, each tuple of the multi-dimensional event embedding can be classified and encoded to generate a one-dimensional journey. For example, events E1 = (a1, a2, a3, a4, a5), E2 = (b1, b2, b3, b4, b5), E3 = (c1, c2, c3, c4, c5). Then, E1 can be classified and encoded as 1, E2 can be classified and encoded as 2, and E3 can be classified and encoded as 3. The cardinality of this classification and encoding will be equal to the product of the cardinalities of the individual event attributes. The classification and encoding can be done in the following way (without overly increasing the cardinality), since steps have been taken during the embedding process to convert high-cardinality event attributes to low-cardinality representations using clustering.
[0068] As a final step, after the model is trained, attention scores can be calculated for certain events that occur in the selected customer journey. Such customer journeys can be selected as those cases where the trained model successfully predicts the outcome. Such cases can be selected from the training dataset and / or include new cases where the model prediction is accurate. Then, the attention scores can be compared with an attention score threshold, where each event in the events that produce an attention score higher than the threshold is identified as a milestone event. Each event with an attention score lower than the selected threshold can be considered noise and ignored. This can be repeated across many such test cases to determine the events that most frequently occur as milestone events. Then, the customer journey for website interaction can be represented as a sequence of such milestone events. As will be understood, such milestone events can be used to generate a concise and informative visualization of a particular type of customer journey related to the company website. In an exemplary embodiment, the visualization can include the business flow and direction from one milestone event to another, similar to a weighted directed graph.
[0069] As an additional step, the high attention scores can be correlated with the specific event attributes present in the milestone events. This provides the key to understanding which event attributes of the multi-dimensional vector that constitutes the event are more important in determining the outcome. By doing so, additional interpretability is provided. Then, this additional information can be used for additional predictive insights, including "next best action" recommendations. Furthermore, via an iterative process, the length of the list of event attributes may decrease, as some attributes are found to be predictive (or important for why an event is identified as a milestone event), while other attributes are found to only add noise. Then, a more accurate model can be trained by focusing on the fewer event attributes that are found to be strongly related to certain outcomes, while discarding the event attributes that are found to be irrelevant noise.
[0070] Now refer to Figure 5, an exemplary method 500 is shown that illustrates an embodiment of the present invention. Method 500 begins at step 505 by generating training data samples from corresponding journey data samples via a training data process (which is shown as an inset to step 505), where each journey data sample in the journey data samples includes a customer journey represented by data describing the following: a sequence of events; for each event in the event sequence, a value associated with a corresponding event attribute in an event attribute list; and a journey outcome. At step 510, method 500 continues to train a machine learning model using the generated training data samples from the previous step. In training, for each training data sample, the input to the machine learning model includes a sequence of vector embeddings generated via the training data process of the event sequence; and for each training data sample, the output of the machine learning model includes an associated journey outcome.
[0071] Regarding the training data process, as shown, the process can begin at step 515 by generating a vector embedding for each event included in a given journey data sample within the journey data samples, where the vector embedding captures the value for each event attribute in the event attributes. At steps 520, 525, and 520, steps for the training data process to generate vector embeddings are described. At step 520, the training data process includes partitioning the event attributes in the event attribute list into a low-cardinality group and a high-cardinality group, where the partitioning is done based on whether the cardinality of the values included for a given event attribute within the training data sample is above or below a predefined cardinality threshold. At step 520, the training data process includes, for each low-cardinality group in the low-cardinality group, categorically encoding the values included within the low-cardinality group based on the total number of unique values that appear therein. At step 520, the training data process includes, for each high-cardinality group in the high-cardinality group, clustering the values included within the high-cardinality group to create multiple cluster groups; and categorically encoding the values included within the high-cardinality group based on the multiple cluster groups in which the values reside. Thus, each training data sample includes: a sequence of vector embeddings generated from the event sequence included in the associated journey sample in the journey samples; and the journey outcome of the associated journey sample in the journey samples.
[0072] In an exemplary embodiment, the event is a web event. Each web event can include actions taken by a customer when interacting with a company's website, such as actions related to selecting a particular web page of the website for browsing and submitting a search query.
[0073] In an exemplary embodiment, when describing with respect to a first training data sample that represents how each training data sample in the training data samples is used to train a machine learning model, the step of training the machine learning model may include: providing, as an input, a sequence of vector embeddings generated for the first training data sample to the machine learning model; generating a predicted journey result as an output of the machine learning model; comparing the journey result of the first training data sample with the predicted journey result, and via this comparison, determining the difference between them; and adjusting the parameters of the machine learning model to reduce the determined difference.
[0074] In an exemplary embodiment, the journey result may include data indicating whether a binary condition is achieved. The binary condition may be related to a company's performance metric, such as whether a sale is completed or whether a performance metric is met.
[0075] In an exemplary embodiment, the event attributes may include: a plurality of attributes describing the web page that the customer is browsing, including at least the URL address and the keywords associated therewith; a plurality of attributes describing the customer, including at least the customer identifier, the location of the customer, and the attributes describing the customer's device; a plurality of attributes describing the referring website, including at least the keywords associated with the referring website; at least one attribute describing the search query submitted by the customer on the referring website or the company's website; and at least one attribute describing the total number of events in the browsing session.
[0076] In an exemplary embodiment, in the step of clustering the values included in each high-cardinality group among the high-cardinality groups, one or more other trained machine learning models may be used. This step may further include: extracting features from the values included in the high-cardinality group; generating a vector embedding for each of the values for the extracted features; and clustering according to the similarities found in the generated vector embeddings.
[0077] In an exemplary embodiment, the machine learning model is a transformer model. In other embodiments, the machine learning model is an attention-based bidirectional long short-term memory recurrent neural network that computes attention scores for corresponding events included within a customer journey when predicting journey outcomes. In an exemplary embodiment, the method may further include the step of determining milestone events by: receiving an attention score threshold; receiving a subset of training data samples that includes training data samples in which the trained machine learning model accurately predicts whether a binary condition is met; determining the attention scores for the events included within each of the training data samples in the subset of training data samples; for each of these events, comparing the attention score to the attention score threshold; determining whether each of these events is a milestone event based on whether the event's attention score exceeds the attention score threshold; and identifying the customer journey type as a sequence of the determined milestone events.
[0078] In an exemplary embodiment, the method may further include the step of generating a visual representation of the identified customer journey type. The visual representation may include a sequence of connected nodes, where each of these nodes is labeled as one of the milestone events of the customer journey type. In an exemplary embodiment, the generated visual representation includes business flow and direction markings from one of the milestone events to another of the milestone events. This may be in the form of numbers marking the flow of the connector, and the connector may be formed as an arrow of direction. In an exemplary embodiment, the step of receiving an attention score threshold includes receiving a plurality of different attention score thresholds in order to generate a plurality of corresponding customer journey types that have different numbers of identified milestone events according to the plurality of different attention score thresholds. The method may include the step of generating a visual representation of each of the plurality of customer journey types.
[0079] In an exemplary embodiment, the method may further include the step of correlating high attention scores achieved when determining milestone events with one or more specific event attributes present in the milestone events. In an exemplary embodiment, the method may further include the step of outputting one or more recommendations for actions to take when a future customer interacts with the company's website when one or more specific event attributes are detected during an interaction with a future customer. Such recommendations may be executed and / or automated in real time in response to a live interaction. In an exemplary embodiment, the method may further include the steps of determining one or more other specific event attributes as noise based on low correlation when determining milestone events; and outputting one or more recommendations for removing one or more other specific event attributes when training a revised version of the machine learning model.
[0080] Those skilled in the art will understand that many different features and configurations described above in connection with several exemplary embodiments may be further selectively applied to form other possible embodiments of the present invention. For the sake of brevity and in consideration of the capabilities of those of ordinary skill in the art, not every possible iteration in the possible iterations is provided or discussed in detail, but all combinations and possible embodiments included in the following claims or otherwise are intended to be part of this application. In addition, it should be apparent that the foregoing only relates to the described embodiments of this application, and many changes and modifications can be made herein without departing from the spirit and scope of this application as defined by the following claims and their equivalents.
Claims
1. A computer-implemented method, the computer-implemented method comprising the steps of: Generating training data samples from corresponding journey data samples via a training data process, each journey data sample in the journey data samples including a customer journey represented by data describing the following: A sequence of events; For each event in the events, a value associated with a corresponding event attribute in an event attribute list; and A journey result; Wherein the training data process includes generating a vector embedding for each event included in a given journey data sample within the journey data samples by the following steps, the vector embedding capturing the value for each event attribute in the event attributes: Dividing the event attributes in the event attribute list into a low-cardinality group and a high-cardinality group, wherein the division is done based on whether the cardinality of the values for a given event attribute included within the training data sample is above or below a predefined cardinality threshold; For each low-cardinality group in the low-cardinality group, categorically encoding the values included within the low-cardinality group based on the total number of unique values therein; For each high-cardinality group in the high-cardinality group: Clustering the values included within the high-cardinality group to create multiple cluster groups; And Categorically encoding the values included within the high-cardinality group based on the multiple cluster groups in which the values reside; Considering each training data sample in the training data samples to include: A sequence of the vector embeddings generated from the sequence of events included in the associated journey sample in the journey sample; and The journey result of the associated journey sample in the journey sample; Using the generated training data samples to train a machine learning model, wherein: For each training data sample, the input to the machine learning model includes the sequence of the vector embeddings; and For each training data sample, the output of the machine learning model includes an associated journey result.
2. The computer-implemented method according to claim 1, wherein the events include web events, each web event including actions taken by the customer when interacting with the company's website, including actions related to selecting a specific web page of the website for browsing and submitting a search query.
3. The computer-implemented method according to claim 2, wherein when describing a first training data sample regarding how each training data sample in the training data samples is used to train the machine learning model, the step of training the machine learning model includes: Providing the sequence of vector embeddings generated for the first training data sample as an input to the machine learning model; Generating a predicted journey result as the output of the machine learning model; Comparing the journey result of the first training data sample with the predicted journey result, and via the comparison, determining the difference between them; And Adjusting the parameters of the machine learning model to reduce the determined difference.
4. The computer-implemented method according to claim 2, wherein the journey result includes data indicating whether a binary condition is achieved, the binary condition being related to a performance metric of the company.
5. The computer-implemented method according to claim 4, wherein the list of event attributes includes: Multiple attributes describing the web page that the customer is browsing, including at least the URL address and keywords associated therewith; Multiple attributes describing the customer, including at least the customer identifier, the location of the customer, and attributes describing the customer's device; Multiple attributes describing the referring website, including at least keywords associated with the referring website; At least one attribute describing a search query submitted by the customer on the referring website or the company's website; and At least one attribute describing the total number of events in the browsing session.
6. The computer-implemented method according to claim 4, wherein the step of clustering the values included in each high-cardinality group in the high-cardinality groups using one or more other trained machine learning models includes: Extracting features from the values included in the high-cardinality group; Generating a vector embedding for the extracted features for each of the values; Clustering based on the similarities found in the generated vector embeddings.
7. The computer-implemented method according to claim 6, wherein the machine learning model includes a transformer model.
8. The computer-implemented method according to claim 6, wherein the machine learning model includes an attention-based bidirectional long short-term memory recurrent neural network, and the attention-based bidirectional long short-term memory recurrent neural network calculates an attention score for a corresponding event included in the customer journey when predicting the journey result.
9. The computer-implemented method according to claim 8, the computer-implemented method further includes the step of determining milestone events by: Receiving an attention score threshold; Receiving a subset of training data samples, the subset of training data samples including training data samples in which the trained machine learning model accurately predicts whether the binary condition is achieved; Determining the attention score of the event included in each of the training data samples in the subset of training data samples; For each of the events, comparing the attention score with the attention score threshold; Determining whether each of the events includes a milestone event based on whether the attention score of the event exceeds the attention score threshold; And Identifying the customer journey type as a sequence of the determined milestone events.
10. The computer-implemented method according to claim 9, the computer-implemented method further includes the step of generating a visual representation of the identified customer journey type, the visual representation including a sequence of connected nodes, wherein each of the nodes is labeled as one of the milestone events of the customer journey type.
11. The computer-implemented method according to claim 10, wherein the generated visual representation includes business flows and directional markers from one milestone event among the milestone events to another milestone event among the milestone events.
12. The computer-implemented method according to claim 11, wherein the step of receiving the attention score threshold includes receiving a plurality of different attention score thresholds so as to generate a plurality of corresponding customer journey types, the plurality of customer journey types having different numbers of identified milestone events according to the plurality of different attention score thresholds.
13. The computer-implemented method according to claim 12, the computer-implemented method further including the step of generating the visual representation of each of the plurality of customer journey types.
14. The computer-implemented method according to claim 9, the computer-implemented method further including the following steps: Correlate the high attention scores achieved when determining the milestone events with one or more specific event attributes present in the milestone events.
15. The computer-implemented method according to claim 14, the computer-implemented method further including the following steps: When the presence of the one or more specific event attributes is detected during an interaction with a prospective customer, output one or more recommendations regarding actions to be taken when the prospective customer interacts with the company's website.
16. The computer-implemented method according to claim 15, the computer-implemented method further including the following steps: Determine one or more other specific event attributes including noise based on low correlation when determining the milestone events; Output one or more recommendations regarding removing the one or more other specific event attributes when training a revised version of the machine learning model.
17. A system, the system comprising: A processor; And A memory storing instructions that, when executed by the processor, cause the processor to perform the following steps: Generate training data samples from corresponding journey data samples via a training data process, each of the journey data samples including a customer journey represented by data describing the following: A sequence of events; For each of the events, a value associated with a corresponding event attribute in an event attribute list; and A journey outcome; Wherein the training data process includes generating a vector embedding for each of the events included within a given journey data sample among the journey data samples, the vector embedding capturing the value for each of the event attributes, by the following steps: Partition the event attributes in the event attribute list into a low-cardinality group and a high-cardinality group, wherein the partitioning is done according to whether the cardinality of the values for a given event attribute included within the training data sample is above or below a predefined cardinality threshold; For each of the low-cardinality groups within the low-cardinality group, classify and encode the values included within the low-cardinality group according to the total number of unique values therein; For each high-cardinality group in the high-cardinality groups: Cluster the values included within the high-cardinality group to create a plurality of cluster groups; And Classify and code the values included within the high-cardinality group according to the plurality of cluster groups in which the values reside; It is considered that each training data sample in the training data samples includes: A sequence of the vector embeddings generated from the sequence of events included in the associated journey sample in the journey sample; and The journey outcome of the associated journey sample in the journey sample; Use the generated training data samples to train a machine learning model, wherein: For each training data sample, the input to the machine learning model includes the sequence of the vector embeddings; and For each training data sample, the output of the machine learning model includes an associated journey outcome.
18. The system according to claim 17, wherein the event includes a web event, and each web event includes an action taken by the customer when interacting with the company's website, including actions related to selecting a specific web page of the website for browsing and submitting a search query.
19. The system according to claim 18, wherein the machine learning model includes an attention-based bidirectional long short-term memory recurrent neural network, and the attention-based bidirectional long short-term memory recurrent neural network calculates an attention score for a corresponding event included in the customer journey when predicting the journey outcome; and wherein the memory further stores instructions that, when executed by the processor, cause the processor to perform the following steps: Determine milestone events by: Receiving an attention score threshold; Receiving a subset of the training data samples, the subset of the training data samples including the training data samples in which the trained machine learning model accurately predicts whether the binary condition is achieved; Determining the attention score of the event included in each of the training data samples in the subset of the training data samples; For each of the events, comparing the attention score with the attention score threshold; Based on whether the attention score of the event exceeds the attention score threshold, determining whether each of the events includes a milestone event; And Identifying the customer journey type as a sequence of the determined milestone events.
20. The system according to claim 19, wherein the memory further stores instructions that, when executed by the processor, cause the processor to perform the following steps: Generate a visual representation of the identified customer journey type, the visual representation including a sequence of connected nodes, wherein each of the nodes is labeled as one of the milestone events of the customer journey type.
Citation Information
Patent Citations
Systems and methods relating to predictive analytics using multidimensional event representation in customer journeys
US20240202096A1