Systems and methods relating to predictive analytics using multidimensional event representations in customer journeys
By generating training data samples from customer journey data and training machine learning models using vector embeddings and clustering techniques, the complexity of customer interactions is addressed, enhancing predictive accuracy and personalization in contact centers.
Patent Information
- Application Number
- JP2025530700
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-19
- Filing Date
- 2023-12-19
- Publication Date
- 2026-01-06
AI Technical Summary
Existing customer relationship management systems struggle to effectively utilize predictive analytics and machine learning to improve contact center performance, particularly in handling customer interactions and providing personalized services, due to the complexity and variability of customer journeys.
A computer-implemented method for generating training data samples from customer journey data, involving vector embeddings, low-cardinality and high-cardinality group division, clustering, and categorization to train machine learning models, which enhance the predictive capabilities of customer journey models.
Improves the predictive accuracy and efficiency of customer journey models, enabling more effective next-best-action recommendations and personalized customer interactions.
Smart Images

Figure 2026500112000001_ABST
Abstract
Description
[Technical Field]
[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims the benefit of U.S. Provisional Patent Application No. 63 / 433,536, entitled "Systems and Methods Relating to Predictive Analytics Using Multidimensional Event Representation in Customer Journeys," filed with the U.S. Patent and Trademark Office on December 19, 2022. This application also claims the benefit of U.S. Patent Application No. 18 / 545,105, entitled "Systems and Methods Relating to Predictive Analytics Using Multidimensional Event Representation in Customer Journeys," filed concurrently with the U.S. Patent and Trademark Office on December 19, 2023. [Background technology]
[0002] The present invention relates generally to customer relationship services and management via contact centers and associated cloud-based systems, including the delivery of customer assistance via contact centers and internet-based service options. More specifically, but not by way of limitation, the present invention relates to website, mobile app, analytics, and contact center interactions through the use of techniques such as predictive analytics and / or machine learning, including the use of representing customer journeys as sequences of multidimensional events, to improve performance. Summary of the Invention
[0003] The present invention includes a computer-implemented method of generating training data samples from respective journey data samples via a training data process, each journey data sample including a customer journey represented by data describing a sequence of events, and for each of the events, values associated with each event attribute in a list of event attributes, and a journey outcome, and training a machine learning model using the generated training data samples. The input of the machine learning model for each training data sample includes a sequence of vector embeddings generated via the training data process of the sequence of events, and the output of the machine learning model for each training data sample includes an associated journey outcome. The training data process includes dividing the vector embeddings for each of the events contained in a given one of the journey data samples capturing values for each of the event attributes into low-cardinality and high-cardinality groups, where the dividing is done according to whether the values contained in the training data sample for the given event attribute have cardinality above or below a predetermined cardinality threshold, and dividing the low-cardinality groups into high-cardinality groups according to the total number of unique values appearing therein, for each low-cardinality group. generating training data samples by categorizing the values contained within the high cardinality groups; clustering the values contained within the high cardinality groups to create, for each high cardinality group, a plurality of cluster groups, and categorizing the values contained within the high cardinality groups according to the plurality of cluster groups in which the values reside; and regarding each training data sample as including a sequence of vector embeddings generated from a sequence of events contained in an associated journey sample from among the journey samples, and a journey outcome of the associated journey sample from among the journey samples.
[0004] These and other features of the present application will become more apparent from a consideration of the following detailed description of exemplary embodiments taken in conjunction with the drawings and the appended claims. [Brief explanation of the drawings]
[0005] A more complete understanding of the present invention will be more readily apparent as the invention becomes better understood by reference to the following detailed description when considered in conjunction with the accompanying drawings in which like reference characters indicate like elements and in which: [Figure 1] 1 illustrates a schematic block diagram of a computing device according to an exemplary embodiment of the present invention and / or on which an exemplary embodiment of the present invention may be enabled or practiced. [Figure 2] 1 illustrates a schematic block diagram of a communications infrastructure or contact center according to an exemplary embodiment of the present invention and / or in which an exemplary embodiment of the present invention may be enabled or implemented; [Figure 3] FIG. 1 is a simplified flow diagram illustrating the functioning of a machine learning model according to an embodiment of the present invention. [Figure 4] 1 is a schematic diagram of a machine learning model, in accordance with an exemplary operation of an embodiment of the present invention; [Figure 5] 1 is a method for generating training data for training a machine learning model, according to one embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0006] For purposes of promoting an understanding of the principles of the present invention, reference will now be made to exemplary embodiments illustrated in the drawings and specific language will be used to describe the same. However, it will be apparent to those skilled in the art that detailed materials provided in the examples may not be required to practice the present invention. In other instances, well-known materials or methods have not been described in detail to avoid obscuring the invention. Additionally, further modifications in the provided examples or applications of the principles of the present invention as presented herein are contemplated as would normally occur to one skilled in the art. Particular features, structures, or characteristics may be combined in any suitable combination and / or subcombination in one or more embodiments or examples. Those skilled in the art will recognize that various embodiments may be computer-implemented using many different types of data processing equipment, and that embodiments may be implemented as an apparatus, a method, or a computer program product. Thus, exemplary embodiments may take the form of hardware embodiments, software embodiments, or a combination thereof.
[0007] Computing Devices The present invention may be computer-implemented using different forms of data processing equipment, such as a digital microprocessor and associated memory, executing a suitable software program. By way of background, Figure 1 illustrates a schematic block diagram of an exemplary computing device 100 in accordance with and / or on which embodiments of the present invention may be enabled or practiced.
[0008] Computing device 100 may be implemented, for example, via firmware (e.g., an application-specific integrated circuit), hardware, or a combination of software, firmware, and hardware. Each of the servers, controllers, switches, gateways, engines, and / or modules (which may collectively be referred to as servers or modules) in the following figures may be implemented via one or more of computing devices 100. By way of example, various servers may be processes running on one or more processors of one or more computing devices 100 that execute computer program instructions and may interact with other systems or modules to perform various functions described herein. Unless specifically limited otherwise, functionality described in the context of multiple computing devices may be integrated into a single computing device, or various functionality described in the context of a single computing device may be distributed across several computing devices. Furthermore, with respect to the computing systems described in the following figures, such as contact center 200 of FIG. 2, its various servers and computer devices may be located on a local computing device 100 (i.e., on-site or in the same physical location as the contact center agents), on a remote computing device 100 (i.e., off-site or in a cloud computing environment, e.g., in a remote data center connected to the contact center via a network), or on some combination thereof.Functionality provided by servers located on off-site computing devices may be accessed and provided via a virtual private network (VPN) as if such servers were on-site, or functionality may be provided using software as a service (SaaS) accessed over the internet using various protocols, such as by exchanging data via extensible markup language (XML), JSON, etc.
[0009] As shown in the illustrated example, computing device 100 may include a central processing unit (CPU) or processor 105 and main memory 110. Computing device 100 may also include a storage device 115, a removable media interface 120, a network interface 125, an I / O controller 130, and one or more input / output (I / O) devices 135, which may include a display device 135A, a keyboard 135B, and a pointing device 135C, as shown. Computing device 100 may further include additional elements, such as a memory port 140, a bridge 145, an I / O port, one or more additional input / output devices 135D, 135E, 135F, and a cache memory 150 in communication with processor 105.
[0010] Processor 105 may be any logic circuit that responds to and processes instructions fetched from main memory 110. For example, processor 105 may be implemented by an integrated circuit, such as a microprocessor, microcontroller, or graphics processing unit, or in a field programmable gate array or application-specific integrated circuit. As shown, processor 105 may communicate directly with cache memory 150 via a secondary bus or a backside bus. Main memory 110 may be one or more memory chips that store data and allow the stored data to be accessed by central processing unit 105. Storage device 115 may provide storage for an operating system and other software that controls scheduling tasks and access to system resources. Unless otherwise limited, computing device 100 may include an operating system and software capable of performing the functionality described herein.
[0011] As shown in the illustrated embodiment, computing device 100 may include a wide variety of I / O devices 135, one or more of which may be connected via I / O controller 130. Input devices may include, for example, a keyboard 135B and a pointing device 135C (e.g., a mouse or optical pen). Output devices may include, for example, a video display device, speakers, and a printer. More generally, I / O devices 135 may include any conventional devices for performing the functions described herein.
[0012] Unless otherwise limited, computing device 100 may be any workstation, desktop computer, laptop or notebook computer, server machine, virtual machine, mobile or smartphone, portable telecommunications device, media playback device, or any other type of computing, telecommunications, or media device capable of performing the operations and functions described herein, including, but not limited to, computing device 100 may include multiple such devices connected by a network or to other systems and resources via a network. Unless otherwise limited, computing device 100 may communicate with other computing devices 100 over any type of network using any conventional communication protocol.
[0013] Contact Center Referring now to FIG. 2 , a communications infrastructure or contact center system (or simply “contact center”) 200 according to and / or in which exemplary embodiments of the present invention may be enabled or implemented is illustrated. By way of background, customer service providers generally offer many types of services through contact centers. Such contact centers may be staffed with employees or customer service agents (or simply “agents”) who serve as an interface between companies, businesses, government agencies, or organizations (hereinafter interchangeably referred to as “organizations” or “enterprises”) and people, such as users, individuals, or customers (hereinafter interchangeably referred to as “individuals” or “customers”). For example, contact center agents may assist customers in making purchasing decisions, placing orders, or resolving issues with products or services already received. Within a contact center, such interactions between agents and customers may occur via a variety of communication channels, such as via voice (e.g., telephone calls or voice-over-IP or VoIP calls), video (e.g., video conferencing), text (e.g., email and text chat), screen sharing, co-browsing, etc.
[0014] Operationally, contact centers generally strive to provide high-quality service to customers while minimizing costs. For example, one way contact centers operate is to handle all customer interactions with live agents. While this approach may be entirely successful from a service quality perspective, it would likely be prohibitively expensive due to the high cost of agent labor. For this reason, most contact centers utilize automated processes such as interactive voice response (IVR) systems, interactive media response (IMR) systems, internet robots or "bots," and automated chat modules or "chatbots" instead of live agents.
[0015] With specific reference to FIG. 2 , contact center 200 may be used by a customer service provider to provide various types of services to customers. For example, contact center 200 may be used to engage in and manage interactions in which automated processes (or bots) or human agents communicate with customers. Contact center 200 may be an in-house facility of a business or enterprise for performing sales and customer service functions related to products and services available through the enterprise. In another aspect, contact center 200 may be operated by a service provider contracted to provide customer-related services to a business or organization. Furthermore, contact center 200 may be deployed on equipment dedicated to the enterprise or a third-party service provider and / or in a remote computing environment, such as, for example, a private or public cloud environment with infrastructure to support multiple contact centers for multiple enterprises. Contact center 200 may include software applications or programs that may run on-premise, remotely, or some combination thereof. It should further be appreciated that various components of contact center 200 may be distributed across various geographic locations.
[0016] Unless specifically limited otherwise, any of the computing elements of the present invention may be cloud-based or implemented within a cloud computing environment. As used herein, "cloud computing" or simply "cloud" is defined as a model for enabling ubiquitous, convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, servers, storage, applications, and services) that can be rapidly provisioned through virtualization, released with minimal management effort or service provider interaction, and then scaled appropriately. Cloud computing can be comprised of a variety of characteristics (e.g., on-demand self-service, wide area network access, resource pooling, rapid scalability, measurable services, etc.), service models (e.g., Software as a Service ("SaaS"), Platform as a Service ("PaaS"), Infrastructure as a Service ("IaaS"), and deployment models (e.g., private cloud, community cloud, public cloud, etc.). A cloud execution model, often referred to as a "serverless architecture," generally involves a service provider dynamically managing the allocation and provisioning of remote servers to achieve desired functions.
[0017] 2, the components or modules of contact center 200 may include a plurality of customer devices 205, a communications network (or simply “network”) 210, a switch / media gateway 212, a call controller 214, an interactive media response (IMR) server 216, a routing server 218, a storage device 220, a statistics server 226, a plurality of agent devices 230, each having a work bin 232, a multimedia / social media server 234, a knowledge management server 236 coupled to a knowledge system 238, a chat server 240, a web server 242, a dialogue server 244, a universal contact server (or “universal contact server, UCS”) 246, a reporting server 248, a media services server 249, and an analytics module 250. It should be understood that any of the computer-implemented components, modules, or servers described in connection with FIG. 2 or in any of the following figures may be implemented via a computing device, such as computing device 100 of FIG. 1. As will be appreciated, contact center 200 generally manages resources (e.g., employees, computers, telecommunications equipment, etc.) to enable delivery of services via telephone, email, chat, or other communication mechanisms. The various components, modules, and / or servers in FIG. 2 (and other figures contained herein) may each include one or more processors that execute computer program instructions and interact with other system components to perform the various functions described herein. Furthermore, the terms “interaction” and “communication” are used interchangeably and generally refer to any real-time and non-real-time interaction using any communication channel, including, but not limited to, telephone calls (PSTN or VoIP calls), email, voicemail, video, chat, screen sharing, text messages, social media messages, WebRTC calls, etc.Access to and control of components of the contact system 200 may be affected through a user interface (UI), which may be generated on the customer device 205 and / or the agent device 230 .
[0018] Customers desiring to receive service from contact center 200 may initiate inbound communications (e.g., telephone calls, emails, chats, etc.) to contact center 200 via customer devices 205. FIG. 2 shows two such customer devices, but it should be understood that any number may be present. Customer devices 205 may be, for example, communication devices such as telephones, smartphones, computers, tablets, or laptops. According to the functionality described herein, customers may generally use customer devices 205 to initiate, manage, and conduct communications with contact center 200, such as telephone calls, emails, chats, text messages, web browsing sessions, and other multimedia transactions. Inbound and outbound communications to customer devices 205 may traverse network 210, the nature of which typically depends on the type of customer device and the form of communication used. By way of example, network 210 may include communication networks for telephone, mobile phone, and / or data services. Network 210 may be a private or public switched telephone network (PSTN), a local area network (LAN), a private wide area network (WAN), and / or a public WAN such as the Internet. Additionally, network 210 may include wireless carrier networks, including code division multiple access networks, global system for mobile communications (GSM) networks, or any wireless network / technology conventional in the art.
[0019] The switch / media gateway 212 may be coupled to the network 210 for receiving and transmitting telephone calls between customers and the contact center 200. The switch / media gateway 212 may include a telephone or communication switch configured to act as a central switch for agent routing within the center. The switch may be a hardware switching system or may be implemented via software. For example, the switch 215 may include an automatic call distributor, a private branch exchange (PBX), an IP-based software switch, and / or any other switch having dedicated hardware and software configured to receive interactions from customers, from the Internet, and / or from the telephone network, and route those interactions, for example, to one of the agent devices 230. Generally, the switch / media gateway 212 establishes a connection between the customer device 205 and the agent device 230, thereby establishing a voice connection between the customer and the agent. The switch / media gateway 212 may be coupled to a call controller 214, for example, which serves as an adapter or interface between the switch and other routing, monitoring, and communication processing components of the contact center 200. The call controller 214 may be configured to handle PSTN calls, VoIP calls, etc. The call controller 214 may include computer-telephone integration (CTI) software for interfacing with switches / media gateways and other components. The call controller 214 may extract data about incoming interactions, such as the customer's phone number, IP address, or email address, and then communicate them to other contact center components when processing the interaction.
[0020] The interactive media response (IMR) server 216 enables self-help or virtual assistant capabilities. Specifically, the IMR server 216 may be similar to an interactive voice response (IVR) server, except that the IMR server 216 is not limited to voice and may also cover various media channels. In an example illustrating voice, the IMR server 216 may be configured with an IMR script to query the customer about their needs. Through ongoing interaction with the IMR server 216, the customer can receive service without needing to speak with an agent. The IMR server 216 may ascertain the reason the customer is contacting the contact center so as to route the communication to the appropriate resource.
[0021] The routing server 218 routes incoming interactions. For example, once it is determined that an inbound communication should be handled by a human agent, functionality within the routing server 218 may select the most appropriate agent and route the communication to that agent. This type of functionality may be referred to as predictive routing. Such agent selection may be based on which available agent is best suited to handle the communication. More specifically, the selection of the appropriate agent may be based on a routing strategy or algorithm implemented by the routing server 218. In doing so, the routing server 218 may query data related to the incoming interaction, such as data related to the particular customer, available agents, and type of interaction, which may be stored in certain databases as described more below. Once an agent is selected, the routing server 218 may interact with the call controller 214 to route (i.e., connect) the incoming interaction to a corresponding agent device 230. As part of this connection, information about the customer may be provided to the selected agent via their agent device 230, which may improve the service the agent can provide.
[0022] With respect to data storage, contact center 200 may include one or more mass storage devices, generally represented by storage device 220, for storing data in one or more databases. For example, storage device 220 may store customer data maintained in customer database 222. Such customer data may include customer profiles, contact information, service level agreements (SLAs), and interaction history (e.g., details of previous interactions with particular customers, including the nature of the previous interactions, disposition data, wait times, handle times, and actions taken by the contact center to resolve the customer's issues). As another example, storage device 220 may store agent data in agent database 223. Agent data maintained by contact center 200 may include agent availability and agent profiles, schedules, skills, average handle times, etc. As another example, storage device 220 may store interaction data in interaction database 224. The interaction data may include data related to numerous past interactions between customers and the contact center. More generally, unless otherwise specified, it should be understood that storage device 220 may be configured to include databases and / or store data related to any of the types of information described herein, and that those databases and / or data may be accessible to other modules or servers of contact center 200 in a manner that facilitates the functions described herein. For example, a server or module of contact center 200 may query such databases and retrieve data stored therein or send data thereto for storage.
[0023] Statistics server 226 may be configured to record and aggregate data related to the performance and operational aspects of contact center 200. Such information may be compiled by statistics server 226 and made available to other servers and modules, such as reporting server 248, which may then generate reports used to manage operational aspects of the contact center and perform automated actions in accordance with the functionality described herein. Such data may relate to the status of contact center resources, such as average wait times, abandon rates, agent occupancy, and others as may be required by the functionality described herein.
[0024] Agent devices 230 of contact center 200 may be communications devices configured to interact with various components and modules of contact center 200 to facilitate the functionality described herein. For example, agent devices 230 may include telephones adapted for regular telephone calls or VoIP calls. Agent devices 230 may further include computing devices configured to communicate with servers of contact center 200, perform data processing associated with operations, and interface with customers via voice, chat, email, and other multimedia communication mechanisms in accordance with the functionality described herein. While only two such agent devices are shown, any number may be present.
[0025] The multimedia / social media server 234 may be configured to facilitate media interactions (other than voice) with the customer device 205 and / or server 242. Such media interactions may relate to, for example, email, voicemail, chat, video, text messaging, web, social media, collaborative browsing, etc. The multimedia / social media server 234 may take the form of any conventional IP router in the art having dedicated hardware and software for receiving, processing, and forwarding multimedia events and communications.
[0026] The knowledge management server 234 may be configured to facilitate interactions between customers and the knowledge system 238. Generally, the knowledge system 238 may be a computer system capable of receiving questions or queries and providing answers in response. The knowledge system 238 may include an artificial intelligence computer system capable of answering questions posed in natural language by retrieving information from sources such as encyclopedias, dictionaries, newswire articles, literary works, or other documents that are populated into the knowledge system 238 as reference material, as is known in the art.
[0027] Chat server 240 may be configured to conduct, orchestrate, and manage electronic chat communications with customers. Such chat communications may be conducted by chat server 240 in a manner such that customers communicate with automated chatbots, human agents, or both. Chat server 240 may function as a chat orchestration server that allocates chat conversations between chatbots and available human agents. In such cases, the processing logic of chat server 240 may be rules driven to leverage intelligent workload distribution among available chat resources. Chat server 240 may also implement, manage, and facilitate a user interface (also known as a UI) associated with the chat functionality. Chat server 240 may be configured to route chats between automated and human sources within a single chat session with a particular customer. Chat server 240 may be coupled to knowledge management server 234 and knowledge system 238 to receive suggestions and answers to queries posed by customers during chats, such as to provide links to related articles.
[0028] Web server 242 provides site hosts for a wide variety of social interaction sites to which customers subscribe, such as Facebook, Twitter, and Instagram. While illustrated as part of contact center 200, it should be understood that web server 242 may be provided by a third party and / or maintained remotely. Web server 242 may also serve web pages to businesses or organizations supported by contact center 200. For example, customers may view web pages to receive information about a particular business's products and services. Within such business web pages, mechanisms may be provided for initiating interactions with contact center 200, for example, via web chat, voice, or email. One example of such a mechanism is a widget that may be deployed on a web page or website hosted on web server 242. As used herein, a widget refers to a user interface component that performs a particular function. In some embodiments, a widget comprises a GUI that is overlaid on a web page displayed to customers over the Internet. A widget may present information, such as in a window or text box, or include buttons or other controls that allow customers to access a particular function, such as sharing or opening a file or initiating a communication. In some implementations, a widget includes a user interface component having a portable portion of code that can be installed and executed within a separate web page without being compiled. Such widgets may include additional user interfaces and may be configured to access a wide variety of local resources (e.g., calendar or contact information on the customer device) or remote resources over a network (e.g., instant messaging, email, or social networking updates).
[0029] Interaction server 244 can be configured to manage contact center deferrable activities and the routing of those activities to human agents for completion. As used herein, deferrable activities include back-office work that can be performed offline, such as responding to emails, participating in training, and other activities that do not involve real-time communication with customers.
[0030] A universal contact server (UCS) 246 may be configured to retrieve information stored in the customer database 222 and / or send information to the customer database for storage in the customer database. For example, the UCS 246 may be utilized as part of a chat function to facilitate maintaining a history of how chats with particular customers were handled, which may then be used as a reference for how future chats should be handled. More generally, the UCS 246 may be configured to facilitate maintaining a history of customer preferences, such as preferred media channels and best times to contact. To do this, the UCS 246 may be configured to identify data related to each customer's interaction history, such as data regarding comments from agents, customer communication history, etc. Each of these data types may then be stored in the customer database 222 or other modules and retrieved as needed by the functions described herein.
[0031] Reporting server 248 may be configured to generate reports from data compiled and aggregated by statistics server 226 or other sources. Such reports may include near real-time or historical reports and may relate to the status of contact center resources and performance characteristics, such as average wait times, abandonment rates, agent occupancy, etc. Reports may be generated automatically or in response to request and may be used toward managing the contact center in accordance with the functionality described herein.
[0032] Media services server 249 provides audio and / or video services to support contact center features. According to functionality described herein, such features may include prompts for IVR or IMR systems (e.g., playing audio files), music on hold, voicemail / single-party recording, multi-party recording (e.g., of audio and / or video calls), speech recognition, dual tone multi-frequency (DTMF) recognition, audio and video transcoding, secure real-time transport protocol (SRTP), audio or video conferencing, call analytics, keyword spotting, etc.
[0033] Analytics module 250 may be configured to perform analytics on data received from multiple different data sources, as may be required by the functionality described herein. Analytics module 250 may also generate, update, train, and modify predictors or models, such as machine learning model 251 and / or model 253, based on the collected data. To accomplish this, analytics module 250 may have access to data stored in storage device 220, including customer database 222 and agent database 223. Analytics module 250 may also have access to interaction database 224, which stores data related to interactions and interaction content (e.g., audio and transcripts of interactions and events detected therein), interaction metadata (e.g., customer identifier, agent identifier, medium of the interaction, interaction length, interaction start and end times, department, tagged categories), and application settings (e.g., interaction path through the contact center). Analytics module 250 may retrieve such data from storage device 220 to develop and train algorithms and models. Although the analytics module 250 is illustrated as being part of a contact center, it will be understood that the functionality described in this regard may also be implemented on a customer system (or, as used herein, the "customer side" of the interaction) and used for the benefit of the customer.
[0034] Machine learning model 251 may include one or more machine learning models, which may be based on neural networks. In particular embodiments, machine learning model 251 is configured as a deep learning model, which is a type of machine learning based on neural networks that uses multiple processing layers to extract progressively higher-level features from data. As an example, machine learning model 251 may be configured to predict behavior. Such behavior models may be trained to predict customer and agent behavior in a wide variety of situations so that interactions can be personally tailored to the customer and handled more efficiently by agents. As another example, machine learning model 251 may be configured to predict aspects related to contact center operation and performance. In other cases, for example, machine learning model 251 may also be configured to perform natural language processing, providing, for example, intent recognition, and the like.
[0035] The analytics module 250 may further include an optimization system 252. The optimization system 252 may include one or more models 253, which may include machine learning models 251, and an optimizer 254. The optimizer 254 may be used in conjunction with the model 253 to minimize a cost function, which is a mathematical representation of a desired objective or system behavior, under a set of constraints. Because the model 253 is typically nonlinear, the optimizer 254 may be a nonlinear programming optimizer. However, it is contemplated that the optimizer 254 may be implemented using a wide variety of different types of optimization approaches, individually or in combination, including, but not limited to, linear programming, quadratic programming, mixed-integer nonlinear programming, stochastic programming, global nonlinear programming, genetic algorithms, particle / swarm techniques, and the like. The analytics module 250 may utilize the optimization system 255 as part of an optimization process in which aspects of the contact center's performance and operations are optimized or at least improved. This may include, for example, aspects related to customer experience, agent experience, dialogue routing, natural language processing, intent recognition, system resource allocation, system analysis, or other functionality related to automated processes.
[0036] Machine learning models FIG. 3 illustrates an exemplary machine learning model 300 that may be included in one or more embodiments of the present invention. The machine learning model 300 may be a component, module, computer program, system, or algorithm. As described below, some embodiments herein use machine learning to provide predictive analytics for applications in contact centers. The machine learning model 300 may be used as a model to power those embodiments. The machine learning model 300 is trained using training data samples 306, which may include input objects 310 and desired output values 312. For example, the input objects 310 and the desired object values 312 may be tensors. A tensor is an n-dimensional matrix, where n may be 0 (a constant), 1 (an array), 2 (a 2D matrix), 3, 4, or more.
[0037] The machine learning model 300 has internal parameters that determine its decision boundary and therefore the output that the machine learning model 300 produces. After each training iteration, which involves inputting training data sample input objects 310 into the machine learning model 300, the actual output 308 of the machine learning model 300 for the input objects 310 is compared to a desired output value 312. One or more internal parameters 302 of the machine learning model 300 may be adjusted so that, upon running the machine learning model 300 with the new parameters, the produced output 308 is closer to the desired output value 312. If the produced output 308 was already identical to the desired output value 312, the internal parameters 302 of the machine learning model 300 may be adjusted to strengthen and reinforce those parameters that caused the correct output and to reduce and weaken parameters that tended to deviate from the correct output.
[0038] The output of the machine learning model 300 may be, for example, a numerical value in the case of regression or a category identifier in the case of a classifier. A machine learning model trained to perform regression may be referred to as a regression model, and a machine learning model trained to perform classification may be referred to as a classifier. Aspects of an input object that the machine learning model 300 may consider when making its decisions may be referred to as features. After the machine learning model 300 has been trained, a new, unseen input object 320 may be provided as an input to the model 300. The machine learning model 300 then produces an output representing a predicted target value 304 of the new input object 320 based on its internal parameters 302 learned from training.
[0039] The machine learning model 300 may be, for example, a neural network, a support vector machine (SVM), a Bayesian network, logistic regression, logistic classification, a decision tree, an ensemble classifier, or other machine learning model. The machine learning model 300 may be supervised or unsupervised. If unsupervised, the machine learning model 300 may identify patterns in unstructured data 340 without training data samples 306. The unstructured data 340 may be, for example, raw data upon which an inference process is desired to be performed. An unsupervised machine learning model may generate an output 342 that includes data that identifies a structure or pattern.
[0040] A neural network may be composed of multiple neural network nodes, each of which includes input values, a set of weights, and an activation function. The neural network nodes may calculate the activation function on the input values to produce output values. The activation function may be a nonlinear function calculated on a weighted sum of the input values plus an optional constant. In some embodiments, the activation function is a logistic, sigmoid, or hyperbolic tangent function. The neural network nodes may be connected to each other so that the output of one node is the input of another node. Furthermore, the neural network nodes may be organized into layers, each of which includes one or more nodes. The input layer may include the input to the neural network, and the output layer may include the output of the neural network. The neural network may be trained and its internal parameters, including the weights of each neural network node, may be updated by using backpropagation.
[0041] In some embodiments, a convolutional neural network (CNN) may be used. A convolutional neural network is a type of neural network and machine learning model. A convolutional neural network may include one or more convolutional filters, also known as kernels, that operate on the output of a neural network layer preceding the convolutional neural network to produce an output that is consumed by a neural network layer following the convolutional neural network. A convolutional filter may have a window within which it operates. The window may be spatially local. If a node in the preceding layer is within the window, the node in the preceding layer may be connected to a node in the current layer. If a node in the preceding layer is not within the window, it is not connected. A convolutional neural network is a type of locally connected neural network, in which neural network nodes are connected to nodes in the preceding layer within a spatially localized region. Furthermore, a convolutional neural network is a type of sparsely connected neural network, in which most of the nodes in each hidden layer are connected to fewer than half of the nodes in the subsequent layer. In other embodiments, a recurrent neural network (RNN) may be used. A recurrent neural network is another type of neural network and machine learning model. A recurrent neural network includes at least one back loop, where the output of at least one neural network node is input to a neural network node in a previous layer. A recurrent neural network maintains a state between iterations, such as in the form of a tensor. The state is updated at each iteration, and the state tensor is passed as an input to the recurrent neural network at the new iteration. In yet other embodiments, the recurrent neural network is a long short-term memory (LSTM) neural network.In some embodiments, the recurrent neural network is a bidirectional LSTM neural network. A feedforward neural network is another type of neural network that does not have a back loop. In some embodiments, the feedforward neural network may be densely connected, meaning that a majority of the neural network nodes in each layer are connected to a majority of the neural network nodes in subsequent layers. In some embodiments, the feedforward neural network is a fully connected neural network, where each of the neural network nodes is connected to each neural network node in subsequent layers. A gated graph sequence neural network (GGSNN) is one type of neural network that may be used in some embodiments. In a GGSNN, the input data is a graph consisting of nodes and edges between the nodes, and the neural network outputs a graph. The graph may be directed or undirected. A propagation step is performed to calculate a node representation for each node, which may be based on the node's features. An output model maps from the node representation and corresponding label to an output for each node. The output model is a differentiable function defined for each node that maps to an output. Additionally, embodiments may include neural networks of different or the same type linked together in a series of neural networks, either serially or in parallel, where a subsequent neural network accepts as input the output of one or more preceding neural networks. A combination of multiple neural networks may be trained end-to-end using backpropagation from the last neural network to the first. As previously mentioned, machine learning model 251 may be configured as a deep learning model. A deep learning model is a type of machine learning based on neural networks that uses multiple processing layers to extract progressively higher-level features from data. Deep learning models are generally more adept at unsupervised learning.
[0042] FIG. 4 illustrates the use of a machine learning model 300 to perform inference on input 360 containing data related to a customer journey. For example, as discussed below, input 360 may include customer journey data, i.e., a sequence of multidimensional vector embeddings generated to represent a sequence of web events describing a customer journey. Machine learning model 300 then performs inference on the data based on internal parameters 302 learned through training. Machine learning model 300 generates a predicted outcome output 370. In exemplary embodiments, machine learning model 300 may be configured according to the desirability of a particular machine learning algorithm to achieve the functions described herein. As an example, machine learning model 300 may include one or more neural networks. More specifically, machine learning model 300 may include a recurrent neural network (RNN), which is generally effective for processing continuous data such as text, audio, or time-series data. Such models are designed to remember or "remember" information from previous inputs, allowing them to exploit context and dependencies between time steps. This is useful for tasks such as language translation, speech recognition, and time-series forecasting. In some embodiments, the RNN may include a long short-term memory (LSTM) network or a gated recurrent unit (GRU). Both LSTMs and GRUs are designed to address the "vanishing gradient" problem in RNNs, which occurs when the gradient of the weights in the network becomes very small and the network has difficulty learning. LSTM networks are a type of RNN that uses a special type of memory cell to store and output information. These memory cells are designed to remember information for long periods of time and do so by using a set of "gates" that control the flow of information into and out of the cells. The gates in an LSTM network are controlled by a sigmoid activation function, which outputs a value between 0 and 1. The gates allow the network to selectively remember or forget information depending on the value of the input and the previous state of the cell.On the other hand, a GRU is a simplified version of an LSTM that uses a single "update gate" to control the flow of information to memory cells instead of the three gates used in an LSTM. This makes a GRU easier to train and faster to execute than an LSTM, but it may not be as effective at storing and accessing long-term dependencies. In other embodiments, the machine learning model 520 may be configured as a sequence-to-sequence model including a first encoder model and a decoder model. The first encoder may include an RNN, a convolutional neural network (CNN), or another machine learning model capable of accepting sequence inputs. The decoder may include an RNN, a CNN, or another machine learning model capable of generating sequence outputs. The sequence-to-sequence model may be trained with training data samples, each of which includes a sequence of vector embeddings representing a sequence of customer journey events and journey outcomes. For example, the sequence-to-sequence model may be trained by inputting input data into the first encoder model to create a first embedding vector. The first embedding vector may be input into the decoder model to create an output result of a predicted outcome. The output outcome may be compared to the actual outcome, and parameters of the first encoder and decoder may be adjusted to reduce the difference between the predicted outcome and the actual outcome. The parameters may be adjusted through backpropagation. In one embodiment, the sequence-to-sequence model may include a second encoder that takes additional information related to the failed test / breaking change to create a second embedding vector. The first embedding vector may be combined with the second embedding vector as input to the decoder. For example, the first and second embedding vectors may be combined using concatenation, addition, multiplication, or another function. In another example, the features may include statistics calculated based on the count and order of words or characters in data associated with a particular event feature, such as words associated with a page URL. In other embodiments, the machine learning model for outputting the predicted outcome may be an unsupervised model, such as a deep learning model.An unsupervised model is not trained, but instead makes its predictions based on identifying patterns in the data. In one embodiment, an unsupervised machine learning model may identify common features in a training dataset.
[0043] Multidimensional event representation for predictive analytics Deriving actionable insights from customer journeys has been an essential area of research focus over the past few years. It is understood that being able to generate concise visualizations of such journeys is an important step toward better understanding the patterns found within customer journeys. Doing this for events that occur when customers interact with websites, which may be referred to as "web events" or simply "events," is extremely difficult given the lack of structure regarding how such interactions unfold. As will be appreciated, the present invention provides a method for modeling customer journeys to identify milestone events that correlate with specific outcomes. The milestone events can then be used to create informative visualizations regarding the predicted sequence of such web events found in the customer journey. Furthermore, through the iterative process enabled by the present invention, the predictive power and accuracy of customer journey models based on web events can be significantly improved as noisy events / parameters are identified and removed. These improved models can then be used to provide more accurate visualizations as well as effective next-best-action recommendations.
[0044] More generally, as used herein, a customer journey refers to the sequence of events that occur when a customer interacts with a company or business. This may include the customer's interactions with an automated system, such as a virtual agent / bot or an IVR system. As previously mentioned, this may further include how the customer interacts with the business's website. A customer journey can be used as a way to organize information about how customers interact with a business over time. A customer journey is a time series of discrete, unevenly sampled customer events that include a heterogeneous set of attributes and features. They may include both clear commitment signals, such as buying a new car, and more ambiguous commitment signals, such as a series of monthly credit card purchases or customer service contacts. The use of a customer journey framework has been observed to improve machine learning performance, whether for recommendation systems or experience customization. Customer journeys can be used to drive sales through these recommendations and customized experiences. Using customer journeys, a customer's experience can be treated as a dynamic rather than static factor. Thus, for example, customer experience becomes a continuously managed signal or parameter that can be used to determine, recommend, and trigger actions from a business to its customers and prospects. For example, modeling of the customer journey in a programmatic state of machine learning or operational research application can be used to determine data points, such as one or more new events, that minimize knowledge gaps in the customer journey. These outcomes can be modeled to determine the "next best action" the business should take with respect to the customer that is either more likely to produce a desired outcome, such as making a sale, or more likely to avoid an undesirable outcome, such as the customer canceling service.
[0045] Much progress has been made related to the effective use of certain types of customer journeys. These include types of customer journeys that are more structured in terms of how customers are allowed to navigate the journey, such as customer journeys involving interactions with agent bots, which may be referred to as blot flows, and customer journeys related to interactions with IVR systems, which may be referred to as IVR paths or journeys. However, beyond that, there remains a significant need for similar support related to customer journeys that describe how customers interact with websites. Progress in this area has proven difficult. The primary reason for this is that bot flows and IVR paths have a fixed structure for how events can be ordered within any given customer journey, but the manner in which customers interact with websites does not. This lack of structure leads to a large number of variations in possible event sequences (i.e., the ordering of events in a customer journey). As an example, the use of trie data structures has been productive in modeling bot flows and IVR paths for predictive purposes and to generate useful visualizations, but this usefulness has not extended to customer journeys involving web events. One reason for this is that the inherent properties of trie data structures make scaling them for this type of use case nearly impossible: in a trie data structure, a new branch is created for each new prefix (in the event sequence), which leads to creating very large and cumbersome trie data structures. Trie data structures grow in size quickly with even small variations in the starting event or event sequence, which leads to a significant increase in memory usage and processing time that makes trie data structures impractical.
[0046] Events that occur when interacting with a website are generally multidimensional, i.e., events contain many different attributes, which increases the difficulty in deciphering customer journeys in this field. To handle such complex cases of high variability in event sequences or to consider multiple attributes (i.e., dimensions) of events while mining patterns for prediction or visualization purposes, trie data structures simply do not work. Instead, as proposed in this disclosure, more advanced machine learning algorithms (specifically, sequence learning algorithms) have been found to be more effective. Such large variability in sequences poses a noisy visualization problem, making it essential to identify which events are important for visualization. Embodiments of the present invention provide solutions to some of these challenges.
[0047] Now, with reference to specific embodiments, it will be appreciated that discovering patterns or common paths in a customer journey can be extremely valuable to a business because it allows the business to make predictions that optimize customer experience and business KPIs. However, as will be appreciated, events within a customer journey do not have equal importance. (Note that a reference event includes a web event or reference to an event.) Therefore, to identify common patterns in a customer journey, it is essential to filter out insignificant events (noise) and focus only on important events (milestone events) that lead to the achievement of outcomes and have predictive value. Such milestone events are events from which accurate predictions can be made about a customer's next action, want, or need. Not surprisingly, distinguishing between these two types of events is often difficult, which contributes to the challenges surrounding customer journey analysis, especially in the context of complex interactions involving websites.
[0048] When customer journey events are multidimensional, i.e., when each event has several different attributes, current techniques for discovering milestones and patterns in the journey simply do not work well. This is especially true when these event attributes have a high degree of cardinality.
[0049] When a customer journey maps interactions with a business website, each event can be scored relative to several different types of event attributes. For example, in this context, an event generally involves a customer navigating between different web pages of the website, submitting a search, or interacting with an aspect of one of the pages, and possible event attributes describe these actions, including what could be a long list. For example, event attributes may include characteristics that describe the website, the page being visited, the order in which the pages are visited, the search string, the referring website, and many others. A preferred embodiment of the present invention may include a list of event attributes used to describe the attributes associated with each event, which may include web page or page attributes, customer attributes, location attributes, date attributes, referring website attributes, event count attributes, browser attributes, session attributes, user device attributes, query or search attributes, and others. These attributes may be determined relative to each web event that occurs within the customer journey, for example, relative to each web page visited, widget interacted with, search entered, etc. that occurs while a customer is interacting with a particular website. According to an example embodiment, the list of event attributes may include each of the following, or a subset thereof: geolocation_country, geolocation_locality, geolocation_regionName, visit_date, referrer_domain, session_referrer_url, referrer_keywords, ipAddress, visitReferrer_domain, referrer_name, visitReferrer_pathname, visitReferrer_medium, referrer_pathname, total_EventCount, session_referrer_queryString, session_referrer_hostname, referrer_hostname, browser_fingerprint, browser_featuresWebrtc, browser_viewheight,browser_featuresFlash、session_shortId、browser_featuresJava、browser_lang、device_category、device_screenheight、device_osVersion、page_domain、page_URL、totalPageviewCount、page_lang、session_type、session_pageviewCount、customerIdType、loginId、customerId、session_eventCount、device_type、page_breadcrumb、visitId、session_createdDate、device_isMobile、session_referrer_pathname、session_referrer_fragment、device_osFamily、ipOrganization、mktCampaign_content、outcomeName、visitReferrer_name、browser_family、marketingCampaign_source、referrer_queryString、outcomeId、referrer_url、browser_version、geolocation_longitude、session_id、referrer_fragment、browser_viewWidth、session_secondsSincePrevious、searchQuery、device_fingerprint、session_secondsSinceFirst、session_durationInSeconds、visitReferrer_url、page_title、session_referrer_name、visitReferrer_hostname、geolocation_timezone、geolocation_postalCode、externalContactId、mktCampaign_clickId、page_fragment、geolocation_source、geolocation_latitude、visitReferrer_keywords、page_keywords、page_queryString, session_referrer_keywords, userAgentString, referrer_medium, page_hostname, page_pathname, visitReferrer_fragment, and visitReferrer_fragment. ,
[0050] According to an exemplary embodiment, a vector embedding is generated for each event (i.e., web event) that captures (i.e., mathematically accounts for or reflects) data describing the score or value of a given event for each of the event attributes. The generation of such a vector embedding may be done in the following manner.
[0051] First, low-cardinality event attributes are coded categorically. In such cases, low cardinality may be defined as cardinality below a selected threshold, e.g., 10. As an example, the country in which a customer is located (e.g., "geolocation_country") may be one of the event attributes with low cardinality. Thus, for example, if there are five unique countries that appear as values in this event attribute, the countries may be listed as 1, 2, 3, 4, 5 and then normalized, thereby forming the basis of an embedding vector.
[0052] According to an exemplary embodiment, a separate process may be used to handle event attributes with high cardinality. For event attributes with high cardinality, unique values may be clustered first. For example, event attributes associated with page URLs may have high cardinality because the names of page URLs are unique. In an exemplary embodiment, to perform clustering, features first need to be extracted from the page URLs. For example, this may include removing special characters so that strings of words remain. Then, for example, a separate vector embedding can be generated for each page URL. This may be done via a machine learning model trained for this purpose, as will be understood by those skilled in the art. Vector embeddings are created through a machine learning process in which a model is trained to convert specific types of data into numerical vectors. Then, another machine learning model (trained to perform clustering) can be used to cluster the vector embeddings based on their similarity. For example, neural network-based clustering can be used. This type of clustering uses a neural network to learn the cluster structure of the data. Examples of this method are autoencoders and deep embedding clustering. Other types of clustering algorithms may also be used, such as centroid-based clustering (which uses the mean or median of the points of a cluster as the cluster center or centroid), K-means (the most common centroid-based clustering algorithm), hierarchical clustering (which builds a hierarchy of clusters, where each cluster is a subset of the next higher-level cluster), density-based clustering such as DBSCAN (which groups points that are close to each other in feature space together), and distribution-based clustering such as Gaussian mixture model (GMM), which models the data as a mixture of probability distributions, and spectral clustering (which uses the eigenvectors of a similarity matrix to cluster the data).
[0053] For example, in an exemplary embodiment, words appearing in a given page URL value can be treated like NLP words. That is, after special characters are removed, the remaining work can be used to generate a vector embedding, which is then used to cluster similar page URLs. The clustering process can be used to significantly reduce the cardinality appearing in the values of the event attribute. For example, in the example of page URLs, each can be represented relative to its cluster, which forms the basis of the vector embedding.
[0054] At this point, one may ask why categorical coding is not generated for all event attributes and then the multidimensional vectorized events are clustered. The reason is that the result would result in a single cluster number for each event and thus represent the journey as a single-dimensional sequence. This would limit the effectiveness of the associated predictive analytics. The reason this is not done also relates to explainability. Clustering events as a whole would generate a single-dimensional journey sequence but would not provide an explanation for the properties of those clusters. Therefore, the results obtained from path analysis and milestone events are at the level of "cluster number" and therefore cannot be interpreted in a meaningful way.
[0055] Once the first two steps are complete, the customer journey can be represented as a sequence of events. In particular, the customer journey is represented as a sequence of vector embeddings, as obtained above, for each sequence of events. Each event in the journey is converted into a vector embedding that captures the values the event had in each event attribute in the list of event attributes. According to an exemplary embodiment, this sequence of vector embeddings constitutes the input for training a predictive model. With regard to the output, this may be based on whether a target outcome was achieved or not in each particular customer journey. For example, the target outcome (or outcome) may be whether a sale was achieved or not. Or, for example, the outcome may be whether a customer accepted an offered chat session with a live agent as a result of a website interaction. In an exemplary embodiment, the output may be represented using a binary variable, such as "0" for an outcome that was not achieved and "1" for an outcome that was achieved. As will be appreciated, the output represents the target variable for training the predictive model.
[0056] According to an exemplary embodiment, a predictive model is then trained using the vector embeddings and outcome data. For example, the model may be an attention-based bidirectional LSTM model or a Transformer model. As will be appreciated, this model is trained to predict outputs given inputs, and is selected according to its ability to handle the low-cardinality multidimensional event embeddings whose generation is described above.
[0057] In an alternative embodiment, each tuple in a multidimensional event embedding can be coded categorically to generate a single-dimensional journey. For example, events E1 = (a1, a2, a3, a4, a5), E2 = (b1, b2, b3, b4, b5), and E3 = (c1, c2, c3, c4, c5). E1 can then be coded categorically as 1, E2 as 2, and E3 as 3. The cardinality of this categorical coding is equal to the product of the cardinalities of the individual event attributes. Categorical coding can be done in a manner (without unduly increasing cardinality) as a step taken during the embedding process, using clustering to convert high-cardinality event attributes into a lower-cardinality representation.
[0058] As a final step, after the model is trained, attention scores can be calculated for certain events that appear in the selected customer journey. Such customer journeys can be selected as examples for which the trained model successfully predicts outcomes. Such examples can be selected from the training dataset and / or can include new examples for which the model predictions are accurate. The attention scores can then be compared to an attention score threshold, and each event that results in an attention score higher than the threshold is identified as a milestone event. Each event with an attention score lower than the selected threshold can be treated as noise and ignored. This can be repeated across many such test cases to determine the events that most regularly appear as milestone events. A customer journey for a website interaction can be represented as a sequence of such milestone events. As can be appreciated, such milestone events can be used to generate concise and informative visualizations of specific types of customer journeys taken in connection with a business website. In an exemplary embodiment, the visualization can include the amount and direction of traffic from one milestone event to another, similar to a weighted, directed graph.
[0059] As an additional step, high attention scores can be correlated with specific event attributes present in milestone events. This provides a key to understanding which of the event attributes that make up the event's multidimensional vector are more important in determining the outcome. Doing this provides additional explainability. This additional information can then be used for additional predictive insights, including "next best action" recommendations. Furthermore, through an iterative process, the length of the list of event attributes can be reduced, as certain event attributes are found to be predictive (or important to why an event is identified as a milestone event) and others are found to merely add noise. A more accurate model can then be trained by focusing on the fewer event attributes found to be strongly correlated with a particular outcome, while discarding event attributes found to be uncorrelated noise.
[0060] 5, an exemplary method 500 illustrating an embodiment of the present invention is shown. Method 500 begins in step 505 by generating training data samples from respective journey data samples via a training data process (shown as an insert into step 505), each of the journey data samples including a customer journey represented by data describing a sequence of events, and for each of the events, values associated with each event attribute in a list of event attributes, and a journey outcome. In step 510, method 500 continues by training a machine learning model using the generated training data samples from the previous step. In training, the input of the machine learning model for each training data sample includes a sequence of vector embeddings generated via the training data process of the sequence of events, and the output of the machine learning model for each training data sample includes an associated journey outcome.
[0061] With respect to the training data process, as shown, the process may begin in step 515 by generating a vector embedding for each of the events included in a given one of the journey data samples capturing values for each of the event attributes. Steps 520, 525, and 520 describe the steps by which the training data process generates the vector embedding. In step 520, the training data process includes partitioning the event attributes of the list of event attributes into low-cardinality and high-cardinality groups, the partitioning being according to whether the values included in the training data sample for the given event attribute have cardinality above or below a predetermined cardinality threshold. In step 520, the training data process includes, for each low-cardinality group, categorically encoding the values included in the low-cardinality group according to the total number of unique values that appear therein. In step 520, the training data process includes, for each high cardinality group, clustering the values contained within the high cardinality group to create a plurality of cluster groups, and categorically encoding the values contained within the high cardinality group according to the plurality of cluster groups they reside in. Thus, each training data sample includes a sequence of vector embeddings generated from a sequence of events contained in an associated one of the journey samples, and a journey outcome for the associated one of the journey samples.
[0062] In an exemplary embodiment, the events are web events, and each web event may include an action taken by a customer when interacting with a business's website, such as an action related to selecting a particular web page of the website to view and submitting a search query.
[0063] In an example embodiment, when described in connection with a first training data sample representing how each of the training data samples is used to train a machine learning model, the step of training the machine learning model may include providing a sequence of vector embeddings generated for the first training data sample as input to the machine learning model; generating predicted journey outcomes as output of the machine learning model; comparing the journey outcomes of the first training data sample with the predicted journey outcomes and determining, via the comparison, differences therebetween; and adjusting parameters of the machine learning model to reduce the determined differences.
[0064] In an example embodiment, a journey outcome may include data indicating whether a binary condition was achieved or not, which may relate to a business performance metric, for example, whether a sale was made or a performance metric was met.
[0065] In an exemplary embodiment, the event attributes may include a plurality of attributes describing the web page being viewed by the customer, including at least a URL address and associated keywords; a plurality of attributes describing the customer, including at least a customer identifier, a customer location, and attributes describing the customer's device; a plurality of attributes describing the referring website, including at least keywords associated with the referring website; at least one attribute describing a search query submitted by the customer on the referring website or the business's website; and at least one attribute describing the total number of events in the browsing session.
[0066] In an example embodiment, one or more other trained machine learning models may be used in clustering the values contained within each of the high-cardinality groups, which may further include extracting features from the values contained within the high-cardinality groups, generating vector embeddings of the extracted features for each of the values, and clustering according to similarities found in the generated vector embeddings.
[0067] In an exemplary embodiment, the machine learning model is a Transformer model. In other embodiments, the machine learning model is an attention-based, bidirectional, long short-term memory recurrent neural network that calculates an attention score for each event included in the customer journey when predicting the journey outcome. In an exemplary embodiment, the method may further include determining the milestone events by receiving an attention score threshold; receiving a subset of the training data samples, the subset including the training data samples for which the trained machine learning model accurately predicts whether the binary condition is achieved or not; determining an attention score for the events included in each of the training data samples of the subset of training data samples; comparing, for each of the events, the attention score to the attention score threshold; determining whether each of the events includes a milestone event based on whether the attention score for the event exceeds the attention score threshold; and identifying the customer journey type as being a sequence of the determined milestone events.
[0068] In an exemplary embodiment, the method may further include generating a visual representation of the identified customer journey types. The visual representation may include a sequence of connected nodes, each of which is labeled as being one of the milestone events for the customer journey type. In an exemplary embodiment, the generated visual representation includes a volume and direction of traffic labeling from one of the milestone events to another of the milestone events. This may be in the form of numbers labeling volume connectors, with the connectors formed as directional arrows. In an exemplary embodiment, receiving an attention score threshold includes receiving a plurality of different attention score thresholds to generate a plurality of respective customer journey types, the plurality of customer journey types having varying numbers of identified milestone events according to the plurality of different attention score thresholds. The method may include generating a visual representation for each of the plurality of customer journey types.
[0069] In an exemplary embodiment, the method may further include correlating the high attention score achieved in determining the milestone event with one or more specific event attributes present in the milestone event. In an exemplary embodiment, the method may further include outputting one or more recommendations regarding actions to take with a future customer interacting with the business's website when the one or more specific event attributes are detected as present during the interaction with the future customer. Such recommendations may be implemented in real time in response to the live interaction and / or may be executed automatically. In an exemplary embodiment, the method may further include determining one or more other specific event attributes that include noise based on the low correlation in determining the milestone event, and outputting one or more recommendations regarding removing the one or more other specific event attributes when training a revised version of the machine learning model.
[0070] As will be appreciated by those skilled in the art, many of the various features and configurations described above in connection with some exemplary embodiments may be selectively applied to form other possible embodiments of the present invention. For the sake of brevity and in consideration of the capabilities of those skilled in the art, each possible iteration will not be provided or discussed in detail, but all combinations and possible embodiments encompassed by the following several claims or otherwise are intended to be part of this application. Furthermore, it will be apparent that the above relates only to the described embodiments of this application, and that numerous changes and modifications can be made herein without departing from the spirit and scope of this application, as defined by the following claims and their equivalents.
Claims
1. 1. A computer-implemented method comprising: generating training data samples from each journey data sample via a training data process, each of the journey data samples comprising: The sequence of events and for each of said events, a value associated with each event attribute in a list of event attributes; a customer journey represented by data describing a journey outcome; the training data process generating a vector embedding for each of the events included in a given one of the journey data samples capturing the values for each of the event attributes; partitioning the event attributes of the list of event attributes into low cardinality and high cardinality groups, said partitioning being performed according to whether the values contained in the training data sample for a given event attribute have cardinality above or below a predetermined cardinality threshold; for each of the low cardinality groups, categorically encoding the values contained within the low cardinality group according to a total number of unique values appearing therein; For each of the high cardinality groups, clustering the values contained within the high cardinality group to create a plurality of cluster groups; and categorically encoding the values contained within the high cardinality groups according to the plurality of cluster groups in which the values reside; Each of the training data samples is a sequence of the vector embeddings generated from the sequence of events contained in an associated one of the journey samples; generating the journey outcomes of the associated ones of the journey samples; training a machine learning model using the generated training data samples; for each training data sample, the input of the machine learning model includes the sequence of vector embeddings; and The computer-implemented method, wherein for each training data sample, the output of the machine learning model includes the associated journey outcome.
2. 2. The computer-implemented method of claim 1, wherein the events include web events, each web event including an action taken by a customer as the customer interacts with a business's website, the actions including actions related to selecting a particular web page of the website to view and submitting a search query.
3. the step of training the machine learning model, when described with respect to a first training data sample, representing how each of the training data samples is used to train the machine learning model, providing the sequence of vector embeddings generated for the first training data sample as input to the machine learning model; generating a predicted journey outcome as an output of the machine learning model; comparing the journey outcomes of the first training data sample with the predicted journey outcomes and, via said comparison, determining differences therebetween; and adjusting parameters of the machine learning model to reduce the determined difference.
4. 3. The computer-implemented method of claim 2, wherein the journey outcomes include data indicating whether a binary condition was achieved or not, the binary condition relating to a performance metric of the business.
5. The list of event attributes is: a plurality of attributes describing the web page being viewed by the customer, the attributes including at least a URL address and associated keywords; a plurality of attributes describing the customer, the plurality of attributes including at least a customer identifier, a location of the customer, and attributes describing a device of the customer; a plurality of attributes describing the referring website, including at least keywords associated with the referring website; at least one attribute describing the referring website or a search query submitted by the customer on the website of the business; and at least one attribute describing a total number of events in a browsing session.
6. clustering the values contained within each of the high cardinality groups using one or more other trained machine learning models; extracting features from the values contained within the high cardinality group; generating a vector embedding of the extracted features for each of the values; 5. The computer-implemented method of claim 4, comprising clustering according to similarities found in the generated vector embeddings.
7. The computer-implemented method of claim 6 , wherein the machine learning model comprises a Transformer model.
8. 7. The computer-implemented method of claim 6, wherein the machine learning model comprises an attention-based bidirectional long short-term memory recurrent neural network that calculates an attention score for each event included in a customer journey when predicting journey outcomes.
9. Milestone events, receiving an attention score threshold; receiving a subset of training data samples, the subset including training data samples for which the trained machine learning model accurately predicts whether the binary condition is met or not met; determining the attention score for the events contained in each of the training data samples of the subset of training data samples; For each of the events, comparing the attention score to the attention score threshold; determining whether each of the events includes a milestone event based on whether the attention score for the event exceeds the attention score threshold; The computer-implemented method of claim 8 , further comprising determining a customer journey type by: identifying the customer journey type as being a sequence of the determined milestone events.
10. 10. The computer-implemented method of claim 9, further comprising generating a visual representation of the identified customer journey type, the visual representation comprising a sequence of connected nodes, each of the nodes labeled as being one of the milestone events of the customer journey type.
11. The computer-implemented method of claim 10 , wherein the generated visual representation includes an amount and direction of traffic labeling from one of the milestone events to another of the milestone events.
12. 12. The computer-implemented method of claim 11, wherein the step of receiving the attention score threshold comprises receiving a plurality of different attention score thresholds to generate a plurality of respective customer journey types, the plurality of customer journey types having varying numbers of the identified milestone events according to the plurality of different attention score thresholds.
13. The computer-implemented method of claim 12 , further comprising generating the visual representation of each of the plurality of customer journey types.
14. 10. The computer-implemented method of claim 9, further comprising correlating a high attention score achieved in determining the milestone event with one or more particular event attributes present in the milestone event.
15. 15. The computer-implemented method of claim 14, further comprising outputting one or more recommendations regarding actions to take with a prospective customer interacting with the website of the business when the one or more particular event attributes are detected as being present during the interaction with the prospective customer.
16. determining one or more other specific event attributes, including noise, based on their low correlation in determining the milestone event; and outputting one or more recommendations regarding removing the one or more other particular event attributes when training a revised version of the machine learning model.
17. 1. A system comprising: a processor; a memory storing instructions that, when executed by the processor, cause the processor to: generating training data samples from each journey data sample via a training data process, each of the journey data samples comprising: The sequence of events and for each of said events, a value associated with each event attribute in a list of event attributes; a customer journey represented by data describing a journey outcome; a vector embedding for each of the events included in a given one of the journey data samples from which the training data process captures the values for each of the event attributes; partitioning the event attributes of the list of event attributes into low cardinality and high cardinality groups, said partitioning being performed according to whether the values contained in the training data sample for a given event attribute have cardinality above or below a predetermined cardinality threshold; for each of the low cardinality groups, categorically encoding the values contained within the low cardinality group according to a total number of unique values appearing therein; For each of the high cardinality groups, clustering the values contained within the high cardinality group to create a plurality of cluster groups; and categorically encoding the values contained within the high cardinality groups according to the plurality of cluster groups in which the values reside; Each of the training data samples is a sequence of the vector embeddings generated from the sequence of events contained in an associated one of the journey samples; generating the journey outcomes of the associated ones of the journey samples; training a machine learning model using the generated training data samples; for each training data sample, the input of the machine learning model includes the sequence of vector embeddings; and For each training data sample, the output of the machine learning model includes the associated journey outcome.
18. 20. The system of claim 17, wherein the events include web events, each web event including an action taken by a customer as the customer interacts with a business's website, the actions including actions related to selecting a particular web page of the website to view and submitting a search query.
19. the machine learning model includes an attention-based bidirectional long short-term memory recurrent neural network that calculates an attention score for each event included in the customer journey when predicting journey outcomes; The memory further stores instructions that, when executed by the processor, cause the processor to: Milestone events, receiving an attention score threshold; receiving a subset of training data samples, the subset including training data samples for which the trained machine learning model accurately predicts whether the binary condition is met or not met; determining the attention score for the events contained in each of the training data samples of the subset of training data samples; For each of the events, comparing the attention score to the attention score threshold; determining whether each of the events includes a milestone event based on whether the attention score for the event exceeds the attention score threshold; and identifying a customer journey type as being a sequence of the determined milestone events.
20. The memory further stores instructions that, when executed by the processor, cause the processor to:
20. The system of claim 19, further comprising: generating a visual representation of the identified customer journey type, the visual representation comprising a sequence of connected nodes, each of the nodes labeled as being one of the milestone events of the customer journey type.