Explainable classification with self-control using client-independent machine learning models
A client-independent machine learning model accurately classifies IT tickets while avoiding misclassifications and providing explanations, addressing the limitations of general language classifiers in IT ticketing systems.
Patent Information
- Application Number
- JP2024546014
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-04-14
- Filing Date
- 2023-05-23
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2043-05-23
AI Technical Summary
Existing IT ticketing systems struggle with accurately classifying technical issues using general language classifiers, which fail to handle technical data and do not provide explanations for their decisions, leading to potential misclassifications and incorrect automated resolutions.
A client-independent machine learning model is trained to classify IT-related tickets while refraining from misclassifying non-IT domain inputs, providing explanations for its decisions using disjunctive normal forms and positively associated features.
The model ensures accurate classification of IT tickets, prevents misclassifications, and enhances user confidence by explaining its decision-making process, thereby improving the reliability of automated resolution systems.
Smart Images

Figure 2025515542000001_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates generally to computer systems, and more particularly to computer-implemented methods, computer systems, and computer program products constructed and arranged to provide explainable classification with abstention using agnostic machine learning models. [Background technology]
[0002] Information Technology (IT) ticketing systems are tools used to track IT service requests, events, incidents, and alerts that may require further action from the IT department. Ticketing software allows to organize IT issues within them to be resolved by streamlining the resolution process. The elements they manage, so-called tickets, provide context about said issues including details, categories, and any associated tags.
[0003] The ticket often contains additional contextual details and also contains relevant contact information for the individual who created the ticket. Tickets are typically generated by employees, but automated tickets may also be created when a specific incident occurs and is flagged. Once a ticket is created, it is assigned to an IT agent who will resolve it. An effective ticketing system allows for tickets to be submitted through a variety of methods. These include submission through a virtual agent, phone, email, service portal, live agent, walk-up experience, etc.
[0004] In general, automated systems automate the resolution of environmental aspects and problems, event monitoring software monitors components and environments, and incidents are reported via tickets through a ticketing system. A typical system might monitor tickets using natural language and output what the problem is via a general language classifier. Unfortunately, while general language classifiers work well with general text, they do not work well with tickets that contain technical data and do not explain how they arrived at their decision. What is needed is a system that can analyze technical issues, classify technical issues, detect these issues, or combine them, without the actions required for this system. Summary of the Invention
[0005] An embodiment of the present invention is directed to a computer-implemented method for providing explainable classification with restraints using a client-independent machine learning model. A non-limiting computer-implemented method includes classifying a record with a label using a machine learning model by a processor, the machine learning model restraining itself from classifying the given record in response to the given record being outside of an information technology (IT) domain. The computer-implemented method includes generating, by the processor, an explanation of the machine learning model's decision to classify the record with the label. The computer-implemented method includes displaying the explanation in a human-readable format.
[0006] Other embodiments of the present invention implement features of the above methods in computer systems and computer program products.
[0007] Additional technical features and advantages are realized through the techniques of the present invention. Embodiments and aspects of the invention are described in detail herein and are considered part of the subject matter of the claims. For a better understanding, reference should be made to the detailed description and drawings.
[0008] The particulars of the exclusive rights set forth herein are particularly pointed out and distinctly set forth in the claims at the end of this specification. The foregoing and other features and advantages of embodiments of the present invention will become apparent from the following detailed description taken in conjunction with the accompanying drawings. [Brief description of the drawings]
[0009] [Figure 1] FIG. 1 is a block diagram of an example computer system for use with one or more embodiments of the present invention. [Diagram 2] FIG. 1 is a block diagram of an example of a system configured to provide explainable classification with restraint using a client-independent machine learning model in accordance with one or more embodiments of the present invention. [Diagram 3] 1 is a flowchart of a computer-implemented method for training a machine learning model for each classification label in accordance with one or more embodiments of the present invention. [Figure 4A] FIG. 1 is a block diagram illustrating an example of a confusion matrix for a machine learning model, in accordance with one or more embodiments of the present invention. [Figure 4B] FIG. 2 is a block diagram illustrating an example of a coefficient matrix for a machine learning model in accordance with one or more embodiments of the present invention. [Figure 5A] FIG. 1 is an example of a chart illustrating regions of positively associated features determined to contribute positively to each of the predicted labels during classification by a machine learning model, in accordance with one or more embodiments of the present invention. [Figure 5B] 1 is an example of a chart illustrating how certain features contribute to determining a class label, in accordance with one or more embodiments of the present invention. [Figure 5C] 1 is a graph illustrating the contribution of features analyzed by a machine learning model to determine a classification in accordance with one or more embodiments of the present invention. [Figure 6] 1 is a flowchart of a computer-implemented method for computing positively associated features for classification in accordance with one or more embodiments of the present invention. [Figure 7] 1 is a flowchart of a computer-implemented method for explaining machine learning decisions using disjunctive normal forms, in accordance with one or more embodiments of the present invention. [Figure 8] FIG. 1 is a block diagram illustrating converting a linear classification formula into a disjunctive normal form in accordance with one or more embodiments of the present invention. [Figure 9] 1 is a flowchart of a computer-implemented method for explaining machine learning decisions using positively associated features in accordance with one or more embodiments of the present invention. [Figure 10] FIG. 1 is a block diagram illustrating selecting features (e.g., tokens) from ticket data and presenting positively associated features that contribute to a class label for a ticket in accordance with one or more embodiments of the present invention. [Figure 11] 1 is a flowchart of a computer-implemented method for providing explainable classification with restraint using a client-independent machine learning model in accordance with one or more embodiments of the present invention. [Figure 12] 1 is a flowchart of a computer-implemented method for providing explainable classification with restraint using a client-independent machine learning model in accordance with one or more embodiments of the present invention. [Figure 13] FIG. 1 illustrates a cloud computing environment in accordance with one or more embodiments of the present invention. [Figure 14] FIG. 2 illustrates an extraction model layer in accordance with one or more embodiments of the present invention. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0010] One or more embodiments provide explainable classification with restraint using a client-independent machine learning model. For a set of all tickets present in a client's environment, the client-independent machine learning model is configured to determine a next discretionary action to resolve a computer / network issue. The client-independent machine learning model is trained on tickets and solutions from a variety of different clients, including clients across different industries such as cyber security, finance, government, manufacturing, schools, retail, cloud computing, data storage, etc., to be client-independent or industry-independent. According to one or more embodiments, the client-independent machine learning model is trained such that when analyzing a ticket, the client-independent machine learning model knows when to classify a ticket and when not to classify a ticket. This can be achieved by building an agnostic machine learning model that only classifies the areas it understands, such as information technology (IT) domains or IT environments, but for areas it does not understand, the agnostic machine learning model restrains itself from classifying. Furthermore, one or more embodiments are configured to explain how the agnostic machine learning model arrived at a particular decision, which gives a user confidence in the agnostic machine learning model's decisions.
[0011] Incident identification and automated resolution is the process of managing IT service disruptions and restoring services. For example, a monitoring system monitors the IT environment of a client in an industry. The term "IT environment" refers to the infrastructure, hardware, software, and systems that a client (entity or business) relies on every day in the course of using information technology. Some of the commonly used resources in an IT environment include computers, Internet access, peripheral devices, etc. Examples of IT environments include hardware routers, personal computers, servers, switches, and data centers, software user applications that enable and make available hardware connections, web servers, and applications, and networks, i.e., firewalls, cables, and other components that facilitate internal and external communication in a business. Upon detection of a technical event in the IT environment and / or at the request of a user of the IT environment, the monitoring system generates a ticket. The ticket is sent to the automated resolution system and / or IT department for resolution. A ticket is a dedicated document or record that represents an incident, alert, request, and / or event that requires an action from the IT department. A ticket is a historical document that details a service event such as an incident, problem, and / or service request. The ticket governs and controls how the service event is handled.
[0012] A typical system may monitor the environment and attempt to identify problems. However, while general language classifiers handle general text well, they do not handle tickets that contain technical data well and do not explain how they arrived at their decision. Furthermore, classifiers are poor at restraining themselves in their classifications. For example, a ticket stating that "the first time man went into space was in 1961" should not be classified as a technical problem / challenge, but many classifiers will analyze the ticket, detect the word "space" and erroneously classify it as a disk handler problem.
[0013] Technical solutions and benefits include a system that provides a client / industry agnostic model to an IT environment, according to one or more embodiments. Thus, thousands of different agnostic machine learning models are not necessary for thousands of different clients or industries, but the agnostic machine learning model works across a variety of clients in different industries. In one or more embodiments, the agnostic machine learning model is configured to refrain from classifying inputs outside of the IT environment or IT domain. This allows the agnostic machine learning model (e.g., classifier) to avoid misclassifying tickets with labels for automatic resolution by an automatic resolution system when the ticket (as an input) is not in the IT environment or IT domain and therefore should not actually generate a label. Yet another technical solution and benefit may include providing the agnostic machine learning model with a machine learning explanation for decision making using disjunctive normal form (DNF) and / or positively associated features. One or more embodiments provide dimensionality reduction using gradient threshold reduction to extract positive association values as a list of all features that influence the agnostic machine learning model (e.g., classifier). Some embodiments may not have these potential benefits or advantages, and these potential benefits or advantages are not necessarily required for all embodiments.
[0014] One or more embodiments described herein may utilize machine learning techniques to perform tasks such as classifying features of interest. More specifically, one or more embodiments described herein may incorporate and utilize rule-based decision making and artificial intelligence (AI) reasoning to accomplish various operations described herein, i.e., classifying features of interest. The phrase "machine learning" broadly describes the ability of electronic systems to learn from data. A machine learning system, engine, or module may include trainable machine learning algorithms that can be trained, e.g., in an external cloud environment, to learn functional relationships between inputs and outputs, and the resulting model (which may be referred to as a "trained neural network," "trained model," "trained classifier," or "trained machine learning model," or combinations thereof) may be used, for example, to classify features of interest.
[0015] Returning to FIG. 1, a computer system 100 is generally shown in accordance with one or more embodiments of the present invention. The computer system 100 may be an electronic computer framework that includes and / or employs any number and combination of computing devices and networks utilizing various communication technologies as described herein. The computer system 100 may be easily expanded, extended, and modular, and may be turned into different services or reconfigured with some features independent of each other. The computer system 100 may be, for example, a server, a desktop computer, a laptop computer, a tablet computer, or a smartphone. In some examples, the computer system 100 may be a cloud computing node. The computer system 100 may be described in the general context of computer system executable instructions, e.g., program modules, executed by the computer system. Generally, program modules may include routines, programs, objects, components, logic, data structures, etc. that perform particular tasks or implement particular abstract data types. Computer system 100 may also be practiced in a distributed cloud computing environment where tasks are performed by remote processing devices that are linked through a communications network. In a distributed cloud computing environment, program modules may be located in both local and remote computer system storage media including memory storage devices.
[0016] As shown in FIG. 1, computer system 100 has one or more central processing units (CPUs) 101a, 101b, 101c, etc. (collectively or generically referred to as processor 101). Processor 101 may be a single-core processor, a multi-core processor, a computing cluster, or any number of any other configurations. Processor 101, also referred to as processing circuitry, is coupled to system memory 103 and various other components via system bus 102. System memory 103 may include read-only memory (ROM) 104 and random access memory (RAM) 105. ROM 104 is coupled to system bus 102 and may include a basic input / output system (BIOS) or its successors, such as a unified extensible firmware interface (UEFI), that controls certain basic functions of computer system 100. RAM is a read-write memory coupled to system bus 102 for use by processor 101. The system memory 103 provides temporary memory space for the operation of the above instructions during operation. The system memory 103 may include random access memory (RAM), read-only memory, flash memory, or any other suitable memory system.
[0017] Computer system 100 includes an input / output (I / O) adapter 106 and a communications adapter 107 coupled to a system bus 102. I / O adapter 106 may be a small computer system interface (SCSI) adapter that communicates with a hard disk 108 and / or any other similar components. I / O adapter 106 and hard disk 108 are collectively referred to herein as mass storage 110.
[0018] Software 111 for execution on computer system 100 may be stored in mass storage 110. Mass storage 110 is an example of a tangible storage medium readable by processor 101 on which software 111 is stored as instructions for execution by processor 101 to cause computer system 100 to operate, for example, as described herein below with reference to various figures. Examples of computer program products and the execution of such instructions are discussed in further detail herein. Communications adapter 107 interconnects system bus 102 to network 112, which may be an external network, thereby enabling computer system 100 to communicate with other such systems. In one embodiment, portions of system memory 103 and mass storage 110 collectively store an operating system, which may be any suitable operating system that coordinates the functions of the various components depicted in FIG. 1.
[0019] Additional input / output devices are shown connected to the system bus 102 via a display adapter 115 and an interface adapter 116. In one embodiment, adapters 106, 107, 115, and 116 may be connected to one or more I / O buses that are connected to the system bus 102 via an intermediate bus bridge (not shown). A display 119 (e.g., a screen or display monitor) is connected to the system bus 102 by a display adapter 115, which may include a graphics controller and a video controller to enhance performance of graphics intensive applications. A keyboard 121, a mouse 122, a speaker 123, a microphone 124, etc. may be interconnected to the system bus 102 via an interface adapter 116, which may include, for example, a super I / O chip integrated multi-device adapter in a single integrated circuit. Suitable I / O buses for connecting peripheral devices such as hard disk controllers, network adapters, and graphics adapters typically include common protocols such as Peripheral Component Interconnect (PCI) and Peripheral Component Interconnect Express (PCIe), etc. Thus, as configured in Figure 1, computer system 100 includes processing capability in the form of processor 101, storage capabilities including system memory 103 and mass storage 110, input means such as keyboard 121, mouse 122, and microphone 124, and output capabilities including speaker 123 and display 119.
[0020] In some embodiments, the communications adapter 107 can transmit data using any suitable interface or protocol, including an Internet Small Computer System interface. The network 112 can be a cellular network, a radio network, a wide area network (WAN), a local area network (LAN), or the Internet, among others. An external computing device can be connected to the computer system 100 through the network 112. In some examples, the external computing device can be an external web server or a cloud computing node.
[0021] It should be understood that the block diagram of Figure 1 is not intended to indicate that computer system 100 includes all of the components depicted in Figure 1. Rather, computer system 100 may include any suitable fewer or additional components not depicted in Figure 1 (e.g., additional memory components, embedded controllers, modules, additional network interfaces, etc.). Additionally, the embodiments described herein with reference to computer system 100 may be implemented with any suitable logic, which may include any suitable hardware (e.g., a processor, embedded controller, application specific integrated circuit, among others), software (e.g., an application, among others), firmware, or any suitable combination of hardware, software, and firmware, in various embodiments, as referred to herein.
[0022] FIG. 2 illustrates a block diagram of an example of a system 200 configured to provide explainable classification with self-control using a client-independent machine learning model, according to one or more embodiments. The system 200 includes a computer system 202 configured to communicate with a number of different computer systems, such as a computer system 240A for managing an IT environment for one client in one industry, a computer system 240B for managing an IT environment for another client in another industry, through a computer system 240N for managing an IT environment for yet another client in a different industry, over a network 250. The computer systems 240A, 240B, through 240N may be generally referred to as computer systems 240. Each computer system 240 has its own IT management system 244 for monitoring the IT environment for a respective client in a respective industry and storing the respective tickets and solutions in a ticket repertoire 246. The ticket repertoire 246 is operable to store a number of tickets and their respective solutions for the IT environment of the computer system 240. The network 250 may be a wired or wireless communication network.
[0023] The IT management system 244 may include or represent a monitoring ticketing system and an automated resolution system for each of the clients in the industry. By the software application 204 communicating with the computer system 240 over a network 250, which may be a wired or wireless communications network, the software application 204 is configured to retrieve various tickets and their respective solutions in the ticket repertoire 246 from different clients in different industries.
[0024] As illustrated by the dotted lines, in one or more embodiments, computer system 202 may include an IT management system 244 and its ticket repertoire 246 for one or more computer systems 240A-240N for each IT environment of each client. Computer system 202 may manage the client's IT environment for one or more computer systems 240A-240N. Any portion of system 200, including computer system 202 and one or more computer systems 240A-240N, may be part of a cloud computing environment 50 as discussed further herein (shown in FIG. 13).
[0025] In system 200, computer system 202, computer systems 240A-240N, IT management system 244, software application 204, training data 206, machine learning model 220, auto-resolution system 222, rule generation algorithm 224, etc., can include and / or use any of the discussed functions in computer system 100, where computer system 100 includes various hardware components and various software applications, such as software 111, that can execute as instructions on one or more processors 101 to perform actions according to one or more embodiments of the present invention. Software application 204 can include, incorporate, or call portions of various other software, algorithms, application programming interfaces (APIs), etc., that operate as discussed herein. Software application 204 represents numerous software applications.
[0026] The tickets and their respective solutions are stored in a repertoire, such as a storage, as training data 206. The software application 204 filters the training data 206 to ensure that the training data 206 resides only in IT environments, which may also be referred to as an IT domain or IT space. The IT domain encompasses the IT environments of clients across industries. Any tickets that are not related to events that reside in the IT domain (e.g., errors, issues, security breaches, failed computer equipment, etc.) are removed from the training data 206.
[0027] The computer system 202 includes a machine learning model 220, which is a client-independent machine learning model that is trained to classify tickets in an IT environment. In one or more embodiments, the client-independent machine learning model is trained, for example, only to classify tickets in an IT environment. The machine learning model 220 may represent a number of machine learning models 220. The machine learning model 220 classifies the ticket by predicting a label that identifies how to solve a computer problem associated with the ticket. The ticket and its predicted label may be sent to an automated resolution system 222 to automatically solve the computer problem of the ticket according to the predicted label output from the machine learning model 220. In one or more embodiments, the machine learning model 220 is a linear classifier and processes a linear classification algorithm. Terms such as label, class, classification, classification label, class label, and the like may be utilized interchangeably to refer to categories of machine learning.
[0028] This linear classification algorithm uses the features of an object, such as the features of a ticket, to determine which class (or group) it belongs to. A linear classifier accomplishes this by making a classification decision based on the value of a linear combination of the features. The features of the object, also known as feature values, are typically presented to the machine in a vector called a feature vector. Examples of linear classification algorithms and techniques include the Naive Bayes algorithm, Linear Discriminant Analysis (LDA) algorithm, Least Squares algorithm, Support Vector Machine algorithm, Ridge Regression algorithm, Lasso algorithm, Elastic Net algorithm, Least Angle Regression algorithm, Orthogonal Matching Pursuit algorithm, Bayesian Regression algorithm, Logistic Regression algorithm, Linear Regression algorithm, Perceptron algorithm, Passive Aggressive Classification algorithm, etc., as will be understood by those skilled in the art.
[0029] The machine learning model 220 may be configured with a trained linear classification algorithm for each classification label for a ticket. Furthermore, the machine learning model 220 is configured to refrain from classifying tickets that do not belong within the scope of the IT environment or IT domain. In one or more embodiments, there may be a classification label indicated as unclassified / unknown, and the machine learning model 220 may be configured to use the unclassified / unknown label to indicate that the feature vector (i.e., features) of the ticket does not apply to the IT environment or IT domain. By refraining from classifying tickets that are not originating from or related to the IT environment or both as unclassified / unknown, or by classifying such tickets, or both, the machine learning model 220 is configured to prevent misclassified tickets from being erroneously sent to the automated resolution system 222 and correspondingly performing an automatic corrective action by the automated resolution system 222 on the IT environment. In one or more embodiments, the machine learning model 220 may prevent misclassified tickets from being erroneously sent to the automated resolution system 222 and correspondingly performing an incomplete automatic corrective action on the IT environment. For example, one or more software and / or hardware components within the IT environment may be automatically altered by an automated resolution system based on an incorrect classification label of the ticket, thereby causing a malfunction of software and / or hardware components of computer systems within the IT environment.
[0030] 3 is a flowchart of a computer-implemented method 300 for training a machine learning model 220 with a linear classification algorithm for each classification label, resulting in a trained machine learning model 220, according to one or more embodiments. The computer-implemented method 300 is performed by the computer system 202.
[0031] At block 302 of the computer-implemented method 300, the software application 204 is configured to retrieve the training data by compiling data from each ticket into the training data stored in the training data 206. The software application 204 may be configured to parse the training data 206 to determine and filter any tickets along with their solutions that are not in the IT environment. This leaves only tickets that are not in the IT environment that represent the IT domain, such that the machine learning model 220 is trained and learns about tickets and their respective solutions that are in the IT domain. In one or more embodiments, for example, the machine learning model 220 is only trained and learns about tickets and their respective solutions that are in the IT domain. Upon detecting a ticket that is not in the IT domain, the machine learning model 220 may refrain from labeling the ticket and / or label the ticket as unclassified / unknown, such that the ticket is prevented / blocked from being processed by the automated resolution system 222.
[0032] The training data 206 can be refined using cross-fold validation. For the training data 206, tickets are labeled in preparation for training the machine learning model 220. For the training data 206, these labels can be automation playbooks obtained by matching automation executions with the incident tickets being addressed. For example, the machine learning model 220 may be trained on tickets that are resolved by an automation platform (e.g., IT management system 244), such as the RedHat® Ansible® platform, and trained on tickets that are not resolved by the automation platform. For example, the automation platform (e.g., IT management system 244) will log into the system when it receives a ticket, and if it has a playbook, the automation platform will execute the playbook, resolve the ticket, and close the ticket. A closed ticket is considered to be factual or completed. The ticket is utilized to train the machine learning model 220 into one of the known class examples in the IT domain, such as, for example, application down, database space issue, disk handler, network connectivity, file system mount handler, high disk space usage handler, high memory and page file usage, host down handler, service handler, job abend, etc. Of course, the exemplary enumeration of the above classification labels is not meant to be all-inclusive.
[0033] In block 304, the software application 204 is configured to train the machine learning model 220 with the training data 206. For example, the data of each ticket (along with its corresponding label) is input to the machine learning model 220 as a feature vector so that the linear classifier algorithm of the machine learning model 220 learns how to classify the ticket input data. The training data is labeled, meaning that the tickets are pre-labeled to determine the time when the output of the machine learning model 220 predicts the correct label. During the training phase, the predicted labels of the tickets from the machine learning model 220 are compared to the labels of the tickets in the training data 206 to continuously improve the machine learning model 220. This enables the machine learning model 220 to learn the correct classification label for each ticket.
[0034] At block 306, the machine learning model 220 is configured to classify the ticket input data with an explanation. For example, each ticket is classified based on the basis and / or decision made by the machine learning model 220. The machine learning model 220 is configured to generate an explanation to the user in terms of the machine learning rule and / or as the positively associated features on which the machine learning model 220 bases its decision for the predicted label of the input ticket. In one or more embodiments, the machine learning model 220 may comprise and / or employ a rule generation algorithm 224 for the machine learning rule and / or the identified positively associated features and negatively associated features, further details of which are discussed herein.
[0035] At block 308, the software application 204 is configured to validate the classification result by comparing the predicted classification label of the ticket with the labels of the tickets in the training data. At block 310, the software application 204 is configured to verify whether the predicted classification label of the ticket matches the labels of the tickets in the training data.
[0036] In block 312, if there is a match (Yes), the software application 204 is configured to verify whether the machine learning rules and / or explanations of the identified positively associated features are satisfactory. This may include requesting input from a subject matter expert and / or employing a natural language processor (NLP) system. If the explanations are satisfactory (Yes), the software application 204 is configured to terminate training of the machine learning model.
[0037] In block 314, if (No) the classification results are not satisfactory for the decision in block 310, the software application 204 is configured to improve the training data and / or tune the classifier algorithm. Also, if (No) the explanation is not satisfactory for the decision in block 312, the software application 204 is configured to improve the explanation.
[0038] It should be noted that a separate linear classification algorithm is trained for each classification label. In one or more embodiments, the machine learning models 220 can include multiple linear classification algorithms, one for each classification label corresponding to the IT domain. In one or more embodiments, there may be multiple machine learning models 220, one for each classification label corresponding to the IT domain. Any discussion of a single linear classification algorithm and / or a single machine learning model 220 applies by analogy to all linear classification algorithms and / or all machine learning models 220 corresponding to all classification labels of the IT domain.
[0039] As described herein, when an input ticket tells an exemplary classifier that "the first man went into space in 1961," the exemplary classifier attempts to classify the ticket with a label, such as disk or disk handler. However, such a label, in this example, would be a misclassification, potentially leading to an incorrect action being taken by the automated resolution system.
[0040] As a technical benefit and solution, one or more embodiments are configured to refrain from classifying such tickets as indicating that "the first time a human went into space was in 1961" because the machine learning model 220 has been trained to refrain from classifying such tickets. Instead, the machine learning model 220 may output unknown / unclassified, thereby preventing the automated resolution system 222 from modifying one or more software and / or hardware components in the IT environment of an industrial client. Thus, based on the output from the machine learning model 220, the software application 204 may recognize that the ticket is unknown / unclassified in the IT domain, and instead of sending the ticket to the automated resolution system 222, send the ticket to a dedicated IT department for resolution. The software application 204 is configured to send tickets that have been properly labeled by the machine learning model 220 to the automated resolution system 222 for automated processing in the IT environment of the industrial client. According to the labels from the machine learning model 220, the automated solving system 222 is configured to modify software components, hardware components, and / or both software and hardware components of one or more computer systems in the IT environment, resulting in improvements to the computer systems themselves. The modification of the software and / or hardware components on computer systems in the IT environment is a practical application related to the use of the machine learning model 220 to solve technical computer problems.
[0041] One or more embodiments provide an approach to compute positively associated features in a ticket for text classification. The positively associated features are tokens in the ticket. A token can refer to one or more words, phrases, sentences, etc. in the text of the ticket and in a process sometimes called tokenization. The tokens can be utilized as features in the feature vector of the ticket. One or more embodiments extract a list of all positively associated features for all (IT) tickets along with their labels, which can be utilized to further train the machine learning model 220 and for use in the explanations discussed herein.
[0042] Further, one or more embodiments are configured to generate a linear classifier for the machine learning model 220, generate coefficient and confusion matrices for insight into how the linear classifier works at a corpus level, extract positive relevance values for all IT tickets using gradient descent iterative threshold shrinkage, curate the positive relevance values with subject matter experts / IT domain experts to find "true" positive relevance values, extract rules and rule features as new training data refines given labels, and create a rich training model and a list of regions of allowed positive relevance values. Thus, one or more embodiments can receive incoming new incident tickets and extract positive relevance values for a given ticket, and if no positive relevance values are identified, the linear classifier refrains from classifying the ticket.
[0043] Referring to Figure 4A, a block diagram shows an example of a confusion matrix of a machine learning model according to one or more embodiments. Figure 4B illustrates a block diagram of an example of a coefficient matrix of a machine learning model according to one or more embodiments. The ticket and label features of the confusion matrix in Figure 4A are projected into the example coefficient matrix illustrated in Figure 4B.
[0044] By probing the machine learning model 220 to learn from errors, the software application 204 can generate an example confusion matrix that captures how the classifier is confused when learning and classifying the training data in Figure 4A. A confusion matrix is an N x N matrix used to evaluate the performance of a classification model, where N is the number of target classes. This matrix compares the actual target values in the training data 206 to the values predicted by the machine learning model.
[0045] Referring to FIG. 4B, the coefficient matrix includes entries representing the signed coefficients of each feature for the hyperplane in the binary linear classifier for each class. The software application 204 is configured to identify the feature with the largest absolute coefficient, e.g., the largest absolute coefficient value. Using the coefficient matrix, this provides the software application 204 with the features that have the most impact on the classification on a global scale. The coefficient matrix, sometimes called a correlation matrix, is a table that presents correlation coefficients for different variables. The coefficient matrix illustrates the correlation between all possible pairs of values in the table. It summarizes large data sets and identifies and visualizes patterns in the given data.
[0046] In Figure 4B, a feature having a positive value for a given classification label indicates that the corresponding feature positively contributes to the machine learning model 220's decision to classify the given classification label, whereas a feature having a negative value (i.e., a negative sign) for a given class label indicates that the corresponding feature does not contribute (i.e., a negative contribution) to the classification of the given class label.
[0047] Based on all features in the ticket (e.g., in the training data 206) for all class labels, the software application 204 using the machine learning model 220 generates a region of positively associated features for each predicted classification label, as illustrated in FIG. 5A. In FIG. 5A, an example chart illustrates positively associated features, also referred to as positive association values, that are determined to contribute positively to each predicted label during classification by the machine learning model 220. For example, the positively associated features "low space", "disk:c handler", and "file system" are positively associated features that positively contribute to the predicted classification label "disk handler" during classification by the machine learning model 220. This classification label disk handler example is utilized for purposes of illustration and not limitation of various example scenarios. Of course, the embodiments are not limited to the classification label disk handler.
[0048] FIG. 5B is a chart illustrating how an example feature "space" contributes to the determination of a classification label. As seen in FIG. 5B, the feature space has a positive coefficient for the classification label disk handler, meaning that the feature space positively contributes to the machine learning model 220's decision to output the classification label disk handler. Similarly, the feature space has a negative coefficient for some other classification labels, meaning that the feature space negatively contributes or has no effect on the machine learning model 220's decision to output the corresponding class label. Therefore, the feature space can be identified as a negatively associated feature and removed as a feature for the corresponding classification label, in which case the space has a negative coefficient. The software application 204 continues removing features because any feature identified has a negatively associated feature with a negative coefficient for a given classification label, thereby leaving only the positively associated features available in the space of positively associated features for the respective classification label, as illustrated in FIG. 5A.
[0049] Continuing with the example scenario for the classification label disk handler, FIG. 5C is a graph illustrating the contribution of features being analyzed by the machine learning model 220 to make its classification of disk handler, which features are now attested to by a human subject matter expert. For each feature, the graph shows its negative contribution to the decision by the machine learning model 220 to classify the ticket with the label disk handler, and its positive contribution to the decision. To attest by the subject matter expert to determining features with positive association values of, for example, “low space”, “disk c:handler”, and “file system” as positively contributing to the machine learning model 220 decision, the software application 204 identifies the positively associated features for a given class. As identified by the software application 204, the subject matter expert determines that all positively associated features are true positively associated features for their respective classification labels, resulting in a region of positively associated features for those labels, as illustrated in FIG. 5A. In one or more embodiments, the region of positively associated features excludes any negatively associated features. Thus, the positively associated features for their respective classification labels are collected and added to the training data 206 as a new training data set in the training data 206 for further training the machine learning model 220 to classify tickets with the classification labels. In one or more embodiments, the positively associated features are established to be the positively associated features for their respective classification labels. This additional training further refines the ability of the machine learning model 220 to learn to refrain from classifying tickets that are not in the IT domain and improves the accuracy of the machine learning model 220.
[0050] During the inference phase, when the machine learning model 220 receives a ticket and outputs its classification label, the software application 204 can probe the machine learning model 220 to obtain positively associated features for any ticket. Also, for incoming tickets, the machine learning model 220 is configured to learn to extract and recognize positively associated features for a given ticket, and when there are no positively associated features found in the ticket, the machine learning model 220 refrains from classifying the ticket. On the other hand, when positively associated features are recognized in the ticket, the machine learning model 220 is configured to classify the ticket, highlight the positively associated features for display to the user (on the display 119), and extract classifier rules using disjunctive normal forms to explain the machine learning model's decisions.
[0051] 6 is a flowchart of a computer-implemented method 600 for computing positively associated features for text classification, according to one or more embodiments. In one or more embodiments, the software application 204 employs, utilizes, and / or integrates with the machine learning model 220 to perform the computer-implemented method 600. Additionally, the software application 204 may be utilized to explore the machine learning model 220 to perform the computer-implemented method 600.
[0052] At blocks 602, 604, and 606, the software application 204 is configured to input the text of the incident ticket to a preprocessor to generate a feature vector from the input text. The preprocessor extracts input features from the text, which are formed into a feature vector. Known techniques can be used to convert the text into a feature vector. The feature vector is input to a machine learning model 220, which outputs a classification label for the corresponding ticket.
[0053] In block 608, the software application 204 is configured to construct a confusion matrix and generate a coefficient matrix for each label output by the machine learning model 220. An example of a confusion matrix and a coefficient matrix are illustrated in Figures 4A and 4B, respectively.
[0054] In block 610, the software application 204 is configured to provide the feature vector and coefficient matrix to an L1 (and / or L2) regularization problem that models positive associations. An example of a regression model that uses an L1 regularization technique is called least absolute shrinkage and selection operator (lasso) regression, and an example of a regression model that uses an L2 regularization technique is called ridge regression. For L1 regularization, lasso regression adds the absolute magnitude of the coefficients to the loss function as a penalty term. L1 regularization may be an option when there are a large number of features that give a sparse solution. For L2 regularization, ridge regression adds the squared magnitude of the coefficients to the loss function as a penalty term. L2 regression may be used to estimate the significance of predictors and may be based on applying penalties to predictors that are not significant. Furthermore, elastic net is when combining L1 and L2 regularization, and this combination leads to the elastic net method of adding hyperparameters.
[0055] In block 612, the software application 204 is configured to apply an iterative shrinkage / threshold algorithm (ISTA) to the L1 (and / or L2) control problem. ISTA is widely used in solving linear inverse problems due to its simplicity. ISTA may involve bidirectional threshold shrinkage of gradient descent.
[0056] In blocks 614 and 616, the software application 204 is configured to generate a sparse feature vector by using the ISTA method and select positively associated features from the sparse feature vector. The sparse feature vector for a classification label has fewer features than the original feature vector in one or more embodiments. Thus, with all the sparse feature vectors generated for each classification label in the IT domain, the software application 204 selects features for each classification label to generate a list of positively associated features for each class label, which are output in block 618. As mentioned above, FIG. 5A represents a region of positively associated features for each classification label, such as the example class label disk handler.
[0057] To further explain the computation of positively associated features for text classification, the following is an example scenario for illustrative purposes and not limitation. The ISTA algorithm is used to compute a sparse solution to invert a linear problem. A typical example of an inverse linear problem is linear regression. One example to consider is a classification problem, e.g., a text classification problem. For text classification problems in the IT ticket management domain, any linear classification algorithm can be used. Now, consider a ticket T that has been classified into class C (e.g., "Disk Handler") using a linear classification algorithm of the machine learning model 220. It is often useful to show "evidence" of the inner workings of a classifier (e.g., the machine learning model 220) and "explain" why it classified ticket T into class C. The positive association values in a ticket such as T are a small subset of the features of T that are involved in its classification into class C. Such a set of features provides a good explanation of the inner workings of a classifier according to one or more embodiments. As discussed in the above example, the present disclosure formulates the problem of finding positively associated features for a text classification problem as a sparse inverse linear problem. One or more embodiments customize and simplify the ISTA algorithm to make it more efficient for this use case. This customization devise a specific problem formulation, such as the one used internally by the Passive-Aggressive Classifier (PAC), and uses its structural properties to efficiently implement the iterative thresholding step.
[0058] One or more embodiments provide explainability for machine learning model decisions. It explains why the machine learning model 220 classified a given ticket with a given classification label. A typical IT domain has several stakeholders, e.g., IT users, IT workers, service availability managers, incident ticket owners, incident ticket assignees, change owners, etc. Different stakeholders may require explanations with different complexity and depth of reasoning. According to one or more embodiments, the explainability of machine learning decisions can be presented in disjunctive normal form and / or using positively related features. In Boolean logic, disjunctive normal form is the canonical normal form of a logical formula consisting of a disjunction of conjunctions. Disjunctive normal form can be described with the terms "OR", "AND", etc.
[0059] Returning to explainability using disjunctive normal forms, Figure 7 is a flowchart of a computer-implemented method 700 for explaining machine learning decisions using disjunctive normal forms, according to one or more embodiments. Figure 7 is described with reference to Figure 8, which is a block diagram illustrating the conversion of a linear classification formula to disjunctive normal form, according to one or more embodiments. The machine learning model 220 receives an input of a ticket and outputs a classification label, for example, as a disk handler.
[0060] In block 702 of the computer-implemented method 700, the software application 204 is configured to extract a linear classification equation from the machine learning model 220. An example of a linear classification equation is shown in block 802 of FIG. 1i f 1 +β 2i f 2 +…+β ki f k where β ki represents an example coefficient of the kth feature, and f krepresents an example feature of the sum of k features. In FIG. 8, the output classification label is disk handler with coefficients as the weights of each feature (which can be positive and negative weights). In this example, the linear classification equation of the class label for a given ticket is disk handler=0.2(filesystem+0.3(mounted)+0.1(limit)-0.5(database)-0.1(cpu)... as illustrated in block 802 of FIG. 8.
[0061] In block 704, the software application 204 is configured to select coefficients (i.e., weights) of the features with positive signs to be utilized in the disjunctive normal form (i.e., “ANDed” together).
[0062] In block 706, optionally, the software application 204 is configured to select coefficients (i.e., weights) of features with negative signs to be utilized (i.e., “ANDed” together) in the disjunctive normal form while preserving their negative signs. In one or more embodiments, the software application 204 may employ or invoke a natural language processing (NLP) model 228 in performing blocks 704 and 706 to parse and analyze the features and their respective coefficients (negative and positive values) in the linear classification formula. In one or more embodiments, the NLP model 228 may be a pre-trained NLP model that has been further trained on the features and their respective coefficients (negative and positive values) in a known linear classification formula to select the features and their coefficients for use in the disjunctive normal form.
[0063] In block 708, the software application 204 is configured to convert the features with selected coefficients into a disjunctive normal form as rules for the machine learning model 220 decision. In one or more embodiments, the software application 204 may only select features with coefficients with positive values, may select features with coefficients with positive values above a threshold, or may select features with coefficients with positive values along with features with negative coefficient values above a threshold. An example of a rule in disjunctive normal form is described in block 804 of FIG. 8. Specifically, block 804 describes an example of three different rules, each separated by "OR", as seen in FIG. 8. Note that the term database has a negative sign indicating its negative contribution to the classification label disc handler. In some embodiments, the negative sign may be replaced with "NOT" to represent a negative contribution. Example rules are displayed (on the display 119) to a user of the machine learning model 220 to explain the decision-making underlying the class label, e.g. disc handler, output by the machine learning model 220 for a given ticket.
[0064] In one or more embodiments, the software application 204 is configured to employ / invoke a rule generation algorithm 224 to generate rules in disjunctive normal form. In one or more embodiments, the rule generation algorithm 224 may be a rule-based algorithm. One example of a rule-based system is a domain-specific expert system that uses rules to make inferences or choices. A rule-based system includes a set of facts or sources of data relevant to a subject under consideration and a set of rules for manipulating that data. These rules are sometimes referred to as "If statements" because they tend to follow the lines of "IFX happens THEN do Y".
[0065] In one or more embodiments, the rule generation algorithm 224 may be a machine learning algorithm that has been trained on training data. For example, the training data includes a linear classification formula for each classification label in the IT domain of the ticket, and the training data includes features that are positive and negative and their corresponding coefficients. During the training phase of the rule generation algorithm 224, a subject matter expert / IT expert can accept or reject the rules in the disjunctive normal form for each class, thereby improving the rule generation algorithm as a machine learning algorithm. During the inference phase, when the rule generation algorithm 224 is implemented as a trained machine learning algorithm / model, the trained machine learning algorithm receives input classification labels, features, and respective coefficients to output a disjunctive normal form of the features and logical terms (e.g., "AND", "OR", etc.) as illustrated in block 804 of FIG. 8.
[0066] In describing explainability using disjunctive normal form, the machine learning model 220 uses a linear classifier, such as a passive-aggressive classifier that includes unigrams and bigrams as features. The presence or absence of these features, in one embodiment, provides the software application 204 with information about the type of incident, such as the type of ticket. Explaining the behavior of the machine learning model 220 (i.e., the classifier) helps users increase their confidence in the classifier. Therefore, one or more embodiments describe the classification in disjunctive normal form, which is a natural form of representation of knowledge for humans.
[0067] Returning to explainability using positively associated features, Figure 9 is a flowchart of a computer-implemented method 900 for explaining machine learning decisions using positively associated features, according to one or more embodiments. Figure 9 is described with reference to Figure 10, which is a block diagram illustrating selecting features (e.g., tokens) from ticket data, according to one or more embodiments, to present the positively associated features that contribute to a classification label for the ticket. The machine learning model 220 receives an input of a ticket and outputs a classification label, e.g., a disk handler.
[0068] In block 902, the software application 204 is configured to extract, for tickets classified into a given classification label, positively associated features from the machine learning model 220. Any of the techniques discussed herein may be utilized to determine the positively associated features for a classification label.
[0069] In block 904, the software application 204 is configured to select the positively associated features with the highest values, e.g., values above a predetermined threshold. Examples of positively associated features are illustrated in Figures 5A and 5C.
[0070] In block 906, the software application 204 is configured to display (e.g., on the display 119) the selected positively associated features with the highest values that are contributing (i.e., influencing) the decision of the machine learning model 220. For example, the selected positively associated features are configured to show an explanation associated with the respective prediction (i.e., what is it about this particular ticket feature that suggested automating the disk handler). In one or more embodiments, for example, the selected positively associated features are configured to only show an explanation associated with the respective prediction. As displayed to the user (e.g., on the display 119), FIG. 10 shows the machine learning rules in block 1002 that have been extracted in a disjunctive normal form, and shows the machine learning features that influence the decision in block 1004. FIG. 10 also displays the ticket description of the ticket in block 1006, and a bar graph showing the positively associated features that influence the decision of the machine learning model 220 in block 1008. In block 1008, the feature disk has a larger impact on the decision than the feature space, although both are used to explain the decision of the machine learning model 220.
[0071] 11 is a flowchart of a computer-implemented method 1100 for providing explainable classification with restraint using a client-independent machine learning model, according to one or more embodiments. The computer-implemented method 1100 can be performed by the computer system 202. Reference can be made to any of the drawings discussed herein.
[0072] In block 1102, the machine learning model 220 receives an input of a record (e.g., a ticket), where the record is related to an information technology (IT) domain. In block 1102, the machine learning model 220 classifies the record with a label, where the machine learning model refrains from classifying a given record (e.g., another ticket) in response to the given record being outside of the IT domain.
[0073] In one or more embodiments, the machine learning model 220 identifies a given record outside the IT domain as unclassified. The machine learning model 220 is trained on training data 206 in the IT domain. The machine learning model 220 is trained by receiving inputs to training records of the training data 206 having training records (e.g., tickets) and their corresponding labels. The machine learning model 220 is trained to refrain from classifying any records outside the IT domain.
[0074] Further, the records and labels are provided to an automated resolution system 222 that is configured to alter at least one component (e.g., a hardware component, a software component, and / or both hardware and software components) in the IT environment of the industry. A given record that is outside of the IT domain is prevented from being provided to the automated resolution system 222, thereby avoiding any component (e.g., a hardware component, a software component, and / or both hardware and software components) in the IT environment being altered based on a misclassification of the given record.
[0075] The machine learning model 220 includes a linear classifier algorithm, and the record is a technical problem ticket in an IT environment. The linear classifier algorithm has been trained on training data in the IT domain, and the linear classifier algorithm has been trained on positively associated features for the labels without negatively associated features (as illustrated in FIGS. 5A and 5C), and the positively associated features for the labels have been verified by human subject matter experts.
[0076] 12 is a flowchart of a computer-implemented method 1200 for providing explainable classification with restraint using a client-independent machine learning model, according to one or more embodiments. The computer-implemented method 1200 can be performed by the computer system 202. Reference can be made to any of the drawings discussed herein.
[0077] In block 1202, the machine learning model 220 classifies an input record (e.g., a ticket), where the machine learning model refrains from classifying a given record (e.g., another ticket) in response to the given record being outside of the information technology (IT) domain. In block 1204, the software application 204 and / or the machine learning model 220 generates an explanation of the machine learning model's decision to classify the record along with a label. In block 1206, the software application 204 and / or the machine learning model 220 provides a display (e.g., on the display 119) of the explanation in a human-readable format.
[0078] In one or more embodiments, the human readable format includes a disjunctive normal form. The explanation of the decision by the machine learning model 220 is based on a linear classification formula utilized by the machine learning model 220. The explanation of the decision by the machine learning model 220 is based on features and their corresponding respective coefficients derived from the linear classification formula of the machine learning model as illustrated in block 802 of FIG. 8. That is, the features are extracted from the text (e.g., tokenized text) of the record (e.g., ticket) as illustrated in block 1006 of FIG. 10. Furthermore, the human readable format includes a display (e.g., on the display 119) of the positively associated features with a respective contribution of a degree for each positively associated feature to the decision by the machine learning model 220 as illustrated in block 1008 of FIG. 10. Furthermore, the machine learning model is trained on training data in the IT domain. The machine learning model includes a linear classifier algorithm and the record is a ticket of a technical problem in an IT environment.
[0079] In one or more embodiments, the machine learning model 220, the rule generation algorithm 224, and / or the NLP model 228 may include various engines / classifiers and / or may be implemented on neural networks. The features of the engines / classifiers may be implemented by configuring and arranging the computer system 202 to execute the machine learning algorithms. Generally, the machine learning algorithms actually extract features from the received data (e.g., technical computer problem tickets) in order to "classify" the received data. Examples of suitable classifiers include, but are not limited to, neural networks, support vector machines (SVMs), logistic regression, decision trees, hidden Markov models (HMMs), etc. The end result of the classifier's operation, i.e., "classification," is to predict a class (or label) for the data. The machine learning algorithms apply machine learning techniques to the received data to create / train / update a unique "model" over time. The learning or training performed by the engines / classifiers may be supervised, unsupervised, or a hybrid that includes aspects of supervised and unsupervised learning. Supervised learning is when the training data is classified / labeled when it is already available. Unsupervised learning is when the training data is not classified / labeled and therefore needs to be evolved through iterative classifiers. Unsupervised learning can utilize additional learning / training methods, such as clustering, anomaly detection, neural networks, deep learning, etc.
[0080] In one or more embodiments, the engine / classifier is implemented as a neural network (or artificial neural network) that uses connections between pre-neurons and post-neurons and thus represents connection weights. The connections represent, for example, synapses between pre-neurons and post-neurons. Neuromorphic systems are interconnected elements that act as stimulated "neurons" and exchange "messages" with each other. Similar to the so-called "plasticity" of synaptic neurotransmitter connections that carry messages between biological neurons, connections in neuromorphic systems such as neural networks carry electronic messages between stimulated neurons, and these messages are provided with a numerical weight that corresponds to the strength or weakness of a given connection. Those weights can be adjusted or tuned based on experience, allowing the neuromorphic system to adapt to the inputs and learn. After being weighted and transformed by a function (i.e., a transfer function) determined by the designer of the network, the activation of these input neurons is then transferred to other downstream neurons, often called "hidden" neurons. This process is repeated until an output neuron is activated. Thus, the activated output neurons determine (or "learn") and provide an output, or inference about the input.
[0081] A training dataset (e.g., training data 206) can be utilized to train the machine learning algorithm. The training dataset can include historical data of past tickets and corresponding options / suggestions / solutions for each ticket. The option / suggestion labels can be applied to each ticket to train the machine learning algorithm as part of supervised learning. For pre-processing, the raw training dataset can be collected and categorized manually. The classified dataset may be labeled (e.g., using Amazon Web Services® (AWS®) labeling tools such as Amazon Sage Maker® Ground Truth). The training dataset may be split into training, testing, and validation datasets. The training and validation datasets are used for training and evaluation, while the testing dataset is used after training and testing the machine learning model on an unseen dataset. The training dataset may be processed through different data augmentation techniques. Training takes the labeled dataset, base network, loss function, and hyperparameters, and once all of these are created and compiled, training of the neural network occurs, ultimately resulting in a trained machine learning model (e.g., a trained machine learning algorithm). Once the model is trained, it (including the tuned weights) is saved to a file for development of a test dataset and / or further testing.
[0082] Although this disclosure includes a detailed description of cloud computing, it will be understood that implementation of the teachings described herein is not limited to a cloud computing environment. Rather, embodiments of the invention may be implemented in conjunction with any other type of computing environment now known or later developed.
[0083] Cloud computing is a service delivery model for enabling convenient, on-demand network access to a shared pool of configurable computational resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) and for rapidly provisioning and releasing these resources with minimal administrative effort or interaction with the service provider. This cloud model includes at least five characteristics, at least three service models, and at least four deployment models.
[0084] The features are as follows:
[0085] On-Demand Self-Service: Cloud customers can automatically provision server time and computing power, such as network storage, as needed, without the need for unilateral human interaction with the service provider.
[0086] Broad Network Access: The capability is available over the network and can be accessed using standard mechanisms, facilitating use by heterogeneous thin-client or thick-client platforms (e.g., cell phones, laptops, and PDAs).
[0087] Resource Pool: The provider's computing resources are pooled and offered to multiple consumers using a multi-tenant model, with different physical and virtual resources dynamically allocated and reallocated according to demand. Consumers typically have a sense of location independence in that they have no control or knowledge regarding the exact location of the resources offered to them, but can still specify a location (e.g., country, state, or data center) at a higher level of abstraction.
[0088] Rapid Elasticity: Capacity is provisioned quickly and elastically, in some cases automatically, so it can be quickly scaled out and quickly released to quickly scale in. Capacity available for provisioning often appears to the consumer as unlimited, available for purchase in any quantity at any time.
[0089] Metered Services: Cloud systems leverage metering to automatically control and optimize resource usage at a level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, and reported, providing transparency to both providers and consumers of the services utilized.
[0090] The service model is as follows:
[0091] SaaS (Software as a Service): The capability provided to the consumer is the use of the provider's applications running on a cloud infrastructure. Those applications can be accessed from a variety of client devices through thin-client interfaces such as web browsers (e.g., web-based email). The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, storage, or individual application functionality, except for the possibility of limited user-specific application configuration settings.
[0092] PaaS (Platform as a Service): The capability provided to a consumer is to deploy applications that the consumer creates or acquires, written using programming languages and tools supported by the provider, onto a cloud infrastructure. The consumer does not manage or control the underlying cloud infrastructure, including networks, servers, operating systems, or storage, but does have control over the deployed applications and, in some cases, the configuration of the application hosting environment.
[0093] Infrastructure as a Service (IaaS): The capability provided to a consumer is the provisioning of processing, storage, network, and other basic computing resources onto which the consumer can deploy and run any software, which may include operating systems and applications. The consumer does not manage or control the underlying cloud infrastructure, but has control over the operating systems, storage, deployed applications, and in some cases, limited control over selected network components (e.g., host firewalls).
[0094] The deployment model is as follows:
[0095] Private Cloud: The cloud infrastructure is operated exclusively for the organization. The cloud infrastructure may be managed by the organization or a third party and may reside on-premise or off-premise, or both.
[0096] Community Cloud: The cloud infrastructure is shared by multiple organizations to support a specific community with shared concerns (e.g., mission, security requirements, policy, and compliance considerations), and may be managed by those organizations or a third party, and may reside on-premises or off-premises, or both.
[0097] Public Cloud: This cloud infrastructure is available for use by general users or large industry organizations and is owned by an organization that sells cloud services.
[0098] Hybrid cloud: This cloud infrastructure is a combination of two or more clouds (private, community, or public) that remain uniquely connected to each other through standardized or proprietary technologies that allow for data and application portability (e.g., cloud bursting to balance the load between clouds).
[0099] A cloud computing environment is a service-oriented environment with an emphasis on statelessness, loose coupling, modularity, and semantic interoperability. At the heart of cloud computing is an infrastructure that comprises a network of interconnected nodes.
[0100] 13, an exemplary cloud computing environment 50 is shown. As shown, the cloud computing environment 50 includes one or more cloud computing nodes 10 with which local computing devices used by cloud consumers (e.g., personal digital assistants (PDAs) or mobile phones 54A, desktop computers 54B, laptop computers 54C, or automotive computer systems 54N, or combinations thereof) communicate. The nodes 10 communicate with each other. These nodes may be physically or virtually grouped (not shown) in one or more networks, such as a private cloud, a community cloud, a public cloud, or a hybrid cloud, or combinations thereof, as previously described herein. This allows the cloud computing environment 50 to provide an infrastructure, platform, or software-as-a-service, or combinations thereof, without the cloud consumers having to maintain resources on their local computing devices. The types of computing devices 54A-54N illustrated in FIG. 13 are intended to be illustrative only, and it will be understood that the computing node 10 and cloud computing environment 50 may communicate with any type of computer-controlled device via any type of network or network-addressable connection (e.g., a connection using a web browser) or both.
[0101] Referring now to Figure 14, a set of functional abstraction layers provided by cloud computing environment 50 (Figure 13) is shown. It should be understood in advance that the components, layers, and functions shown in Figure 14 are intended to be illustrative only and that embodiments of the present invention are not limited thereto. As shown, the following layers and corresponding functions are provided:
[0102] The hardware and software layer 60 includes hardware and software components. Examples of hardware components include mainframes 61, RISC (Reduced Instruction Set Computer) architecture based servers 62, servers 63, blade servers 64, storage devices 65, and networks and network components 66. In some embodiments, the software components include network application server software 67 and database software 68.
[0103] The virtualization layer 70 comprises an abstraction layer at which virtual entities such as virtual servers 71 , virtual storage 72 , virtual networks including virtual private networks 73 , virtual applications and operating systems 74 , and virtual clients 75 are provided.
[0104] In one example, the management layer 80 may provide the following functions: Resource provisioning 81 provides dynamic procurement of computing and other resources utilized to execute tasks within the cloud computing environment. Metering and pricing 82 provides tracking of costs as resources are utilized within the cloud computing environment and billing or invoicing for the consumption of these resources. In one example, these resources may include application software licenses. Security provides identity verification of cloud consumers and tasks, and protection of data and other resources. User portal 83 provides consumers and system administrators access to the cloud computing environment. Service level management 87 provides allocation and management of cloud computing resources such that required service levels are met. Service level agreement (SLA) planning and achievement 88 provides advance provisioning and procurement of cloud computing resources in anticipation of future demand according to SLAs.
[0105] The Workload Layer 90 shows examples of functions utilized in a cloud computing environment. Examples of workloads and functions provided by this layer include mapping and navigation 91, software development and lifecycle management 92, virtual classroom instruction delivery 93, data analytics processing 94, transaction processing 95, and workloads and functions 96.
[0106] Various embodiments of the present invention are described herein with reference to the associated drawings. Alternative embodiments may be devised without departing from the spirit of the present invention. Although various connections and relationships (e.g., above, below, next to, etc.) are described between elements in the following description and drawings, those skilled in the art will recognize that many of the relationships described herein are independent of direction, as the described functionality is maintained even if the orientation is changed. These connections and / or relationships may be direct or indirect unless otherwise specified, and the present invention is not intended to be limited in this respect. Thus, coupling between entities may refer to direct or indirect coupling, and relationships between entities may refer to direct or indirect relationships. As an example of an indirect relationship, reference herein to forming layer "A" on layer "B" includes the situation where one or more intermediate layers (e.g., layer "C") are between layer "A" and layer "B," so long as the relevant features and functionality of layer "A" and layer "B" are not substantially altered by the intermediate layers.
[0107] For the sake of brevity, conventional techniques for making and using aspects of the present invention may or may not be described in detail herein. In particular, various aspects of computing systems and particular computer programs that implement various technical features described herein are well known. Thus, for the sake of brevity, many conventional implementation details are only briefly described herein or omitted entirely, and details of well-known systems and / or processes are not shown.
[0108] In some embodiments, various functions or acts may be performed at a given location and / or in association with one or more devices or systems. In some embodiments, some given functions or acts may be performed at a first device or location, and remaining functions or acts may be performed at one or more additional devices or locations.
[0109] The terms used herein are for the purpose of describing particular embodiments only and are not intended to be limiting of the invention. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising", as used herein, indicate the presence of stated features, integers, steps, operations, elements, or components, or combinations thereof, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, or groups thereof, or combinations thereof.
[0110] The corresponding structures, materials, acts, and equivalents of all means-plus-function or step-plus-function elements in the following claims are intended to include any structure, material, or act for performing a function in combination with other elements recited in the claims specifically recited in the claims. The present disclosure has been presented for purposes of illustration and description, but is not intended to be exhaustive or limited to the disclosed form. Many modifications and variations will become apparent to those skilled in the art without departing from the scope of the disclosure. The embodiments have been selected and described in order to best explain the principles and practical application of the disclosure and to enable others skilled in the art to understand the disclosure in terms of various embodiments with various modifications as suited to the particular use contemplated.
[0111] The diagrams depicted herein are illustrative. There may be numerous variations to the diagrams or steps (or operations) described herein without departing from the spirit of the disclosure. For example, acts may be performed in a different order, or may be added, deleted, or modified. Also, the term "coupled" depicts having a signal path between two elements and does not imply a direct connection between those elements without an intervening element / connection between them. All of these variations are considered to be part of the disclosure.
[0112] The following definitions and abbreviations shall be used to interpret the claims and the specification. As used herein, the terms "comprises," "comprising," "includes," "including," "has," "having," "contains," or "containing," or any variations thereof, are intended to include a non-exclusive inclusion. For example, a composition, mixture, process, method, article, or device that includes a recitation of elements is not necessarily limited to only those elements, but may include other elements not expressly recited or inherent in such composition, mixture, process, method, article, or device.
[0113] Moreover, the term "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any embodiment or design described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments or designs. The terms "at least one" and "one or more" are understood to include any integer number greater than or equal to one, i.e., 1, 2, 3, 4, etc. The term "multiple" is understood to include any integer number greater than or equal to two, i.e., 2, 3, 4, 5, etc. The term "connected" can include both indirect and direct "connections."
[0114] The terms "about," "substantially," "approximately," and variations thereof are intended to include the degree of error associated with measurement of a particular quantity based on equipment available at the time of filing of this application. For example, "about" can include a range of ±8% or 5%, or 2% of a given value.
[0115] The present invention may be a system, method, or computer program product, or combination thereof, at any possible and technically detailed level of integration. The computer program product may include computer readable storage medium(s) having computer readable program instructions for causing a processor to perform aspects of the present invention.
[0116] A computer readable storage medium may be any tangible device capable of holding and storing instructions for use by an instruction execution device. A computer readable storage medium may be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. A non-exhaustive list of more specific examples of computer readable storage media includes portable floppy disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanically encoded devices such as punch cards or ridge structures in grooves on which instructions are recorded, and any suitable combination thereof. As used herein, a computer-readable storage medium should not be construed as a signal that is itself ephemeral, such as electric waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium (e.g., light pulses passing through a fiber optic cable), or electrical signals transmitted over wires.
[0117] The computer readable program instructions described herein may be downloaded from a computer readable storage medium to each computing device / processing device or to an external computer or storage device over a network (e.g., the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof) that may include copper transmission cables, optical transmission fiber, wireless transmission, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. A network adapter card or network interface within each computing device / processing device receives the computer readable program instructions from the network and transfers the computer readable program instructions for storage on a computer readable storage medium within each computing device / processing device.
[0118] The computer readable program instructions for carrying out the operations of the present invention may be source or object code written in any combination of one or more programming languages, including assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state setting data, configuration data for integrated circuits, or object oriented programming languages such as Smalltalk, C++, and procedural programming languages such as the "C" programming language or similar programming languages. The computer readable program instructions may execute entirely on the user's computer, partially on the user's computer as a standalone software package, partially on the user's computer and on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (e.g., via the Internet using an Internet Service Provider). In some embodiments, to carry out aspects of the invention, electronic circuitry including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may execute computer readable program instructions to customize the electronic circuitry by utilizing state information of the computer readable program instructions.
[0119] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks included in the flowchart illustrations and / or block diagrams, can be implemented by computer readable program instructions.
[0120] These computer readable program instructions may be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, executed by the processor of the computer or other programmable data processing apparatus, produce means for performing the functions / operations specified in one or more blocks of the flowcharts and / or block diagrams. These computer readable program instructions may be stored on a computer readable storage medium and capable of directing a computer, programmable data processing apparatus, or other device, or combination thereof, to function in a particular manner, such that the computer readable storage medium on which the instructions are stored comprises an article of manufacture containing instructions for performing aspects of the functions / operations specified in one or more blocks of the flowcharts and / or block diagrams.
[0121] Computer-readable program instructions may be loaded into a computer, other programmable data processing apparatus, or other device such that the instructions, which execute on the computer, other programmable apparatus, or other device, perform the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams, thereby causing a series of operable steps to be performed on the computer, other programmable apparatus, or other device to produce a computer-implemented process.
[0122] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block of the flowcharts or block diagrams may represent a module, segment, or portion of instructions, which includes one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions shown in the blocks may occur in a different order than that shown in the figures. For example, two blocks shown in succession may in fact be executed substantially simultaneously, or the blocks may sometimes be executed in the reverse order, depending on the functionality involved. It should also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, may be implemented by a special-purpose hardware-based system that performs the specified functions or operations or executes a combination of special-purpose hardware and computer instructions.
[0123] The description of various embodiments of the present invention has been presented for purposes of illustration, but is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will become apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terms used herein have been selected to best explain the principles of the embodiments, practical applications, or technical improvements over the art found in the market, and to enable others skilled in the art to understand the embodiments described herein.
Claims
1. classifying, by a processor, the record with the label using a machine learning model, the machine learning model refraining from classifying the given record in response to the given record being outside of an information technology (IT) domain; generating, by the processor, an explanation of the decision by the machine learning model to classify the record along with the label; displaying said description in a human readable form; and 4. A computer-implemented method comprising:
2. The computer-implemented method of claim 1 , wherein the human readable form comprises a disjunctive normal form.
3. The computer-implemented method of claim 1 , wherein the explanation for the decision by the machine learning model is based on a linear classification formula utilized by the machine learning model.
4. the explanation of the decision by the machine learning model is based on features and coefficients corresponding to the features, the features and the coefficients being obtained from a linear classification formula of the machine learning model; The features are extracted from the text of the record.
10. The computer-implemented method of claim 1.
5. 2. The computer-implemented method of claim 1, wherein the human-readable format includes an indication of positively associated features and a degree of contribution of each of the positively associated features to the decision by the machine learning model.
6. The computer-implemented method of claim 1 , wherein the machine learning model is trained on training data in the IT domain.
7. the machine learning model comprises a linear classifier algorithm; The record is a ticket for a technical problem in an IT environment.
10. The computer-implemented method of claim 1.
8. a memory having computer readable instructions; one or more processors for executing the computer-readable instructions; the computer readable instructions comprising: classifying the records with the labels using a machine learning model, the machine learning model refraining from classifying the given record in response to the given record being outside of an information technology (IT) domain; generating an explanation of the machine learning model's decision to classify the record along with the label; and displaying said description in a human readable form; and and controlling the one or more processors to perform operations including:
9. The system of claim 8 , wherein the human readable form comprises a disjunctive normal form.
10. The system of claim 8 , wherein the explanation for the decision by the machine learning model is based on a linear classification formula utilized by the machine learning model.
11. the explanation of the decision by the machine learning model is based on features and coefficients corresponding to the features, the features and the coefficients being obtained from a linear classification formula of the machine learning model; The features are extracted from the text of the record. The system of claim 8.
12. 10. The system of claim 8, wherein the human-readable form includes an indication of positively associated features and a degree of contribution of each of the positively associated features to the decision by the machine learning model.
13. The system of claim 8 , wherein the machine learning model is trained on training data in the IT domain.
14. the machine learning model comprises a linear classifier algorithm; The record is a ticket for a technical problem in an IT environment. The system of claim 8.
15. 1. A computer program product including a computer readable storage medium having program instructions embodied therein, the program instructions comprising: classifying the records with the labels using a machine learning model, the machine learning model refraining from classifying the given record in response to the given record being outside of an information technology (IT) domain; generating an explanation of the machine learning model's decision to classify the record along with the label; and displaying said description in a human readable form; and a computer program product executable by the one or more processors to cause the one or more processors to perform operations including:
16. 15. The computer program product of claim 14, wherein the human readable form comprises a disjunctive normal form.
17. 15. The computer program product of claim 14, wherein the explanation for the decision by the machine learning model is based on a linear classification formula utilized by the machine learning model.
18. the explanation of the decision by the machine learning model is based on features and coefficients corresponding to the features, the features and the coefficients being obtained from a linear classification formula of the machine learning model; The features are extracted from the text of the record.
15. A computer program product according to claim 14.
19. 15. The computer program product of claim 14, wherein the human readable form includes an indication of positively associated features and a degree of contribution of each of the positively associated features to the decision by the machine learning model.
20. 15. The computer program product of claim 14, wherein the machine learning model is trained on training data in the IT domain.
Citation Information
Patent Citations
Hybrid learning-based ticket classification and response
US20200104752A1
Model agnostic contrastive explanations for structured data
US20200193243A1
Predictive Resolutions for Tickets Using Semi-Supervised Machine Learning
US20210019648A1
Object identification system, operation processing device, vehicle, lighting tool for vehicle, and training method for classifier
WO2020121973A1