Explainable classification with self-control using client-independent machine learning models
A client-agnostic machine learning model with self-restraint capabilities addresses the misclassification issues in IT ticketing systems by accurately classifying IT domain tickets and preventing incorrect automated modifications.
Patent Information
- Application Number
- JP2025526774
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-04-14
- Filing Date
- 2023-05-23
- Publication Date
- 2025-11-14
AI Technical Summary
Existing IT ticketing systems struggle with general language classifiers that fail to handle technical data and do not explain their decision-making processes, leading to misclassifications that can result in incorrect automated modifications of software and hardware components.
A client-agnostic machine learning model is trained to classify IT domain tickets while refraining from classifying non-IT domain tickets, using a restraint mechanism and providing explanations through disjunctive normal forms and positively associated features.
The model prevents misclassifications by accurately classifying IT domain tickets and avoiding incorrect automated modifications, enhancing the reliability of IT system operations.
Smart Images

Figure 2025537280000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates generally to computer systems, and more particularly to computer-implemented methods, computer systems, and computer program products that are constructed and arranged to provide explainable classification with abstention using client-agnostic machine learning models. [Background technology]
[0002] Information technology (IT) ticketing systems are tools used to track IT service requests, events, incidents, and alerts that may require further action from the IT department. Ticketing software allows organizations to organize the IT issues they manage for resolution by streamlining the resolution process. The elements they manage, called tickets, provide context about said issues, including details, categories, and any associated tags.
[0003] This ticket often contains additional contextual details and relevant contact information for the individual who created it. Tickets are typically generated by employees, but automated tickets may also be created when a specific incident occurs and is flagged. Once a ticket is created, it is assigned to an IT agent who will resolve it. An effective ticketing system allows tickets to be submitted through a variety of methods. These include submission through a virtual agent, phone, email, service portal, live agent, walk-up experience, etc.
[0004] Generally, automated systems automate environmental aspects and problem resolution, event monitoring software monitors components and environments, and incidents are reported via tickets through ticketing systems. A typical system might monitor tickets using natural language and output what the problem is via a general language classifier. Unfortunately, while general language classifiers handle general text well, they do not handle tickets containing technical data well and do not explain how the system arrived at its decision. What is needed is a system that can analyze technical issues, classify them, detect these issues, or combine them, without the necessary actions. Summary of the Invention
[0005]
[0003] Embodiments of the present invention are directed to a computer-implemented method for providing explainable classifications with restraint using a client-independent machine learning model. A non-limiting computer-implemented method includes inputting records associated with an information technology (IT) domain into a machine learning model by a processor. The computer-implemented method includes classifying the labeled records by the processor using the machine learning model. The machine learning model utilizes a restraint method for classifying a given record in response to the given record being outside the scope of the IT domain.
[0006] This may provide an improvement over known methods for classification by providing machine learning models with the ability to refrain from classifying irrelevant tickets, thereby preventing tickets from being automatically processed by an automated system to erroneously modify software and / or hardware components of one or more computer systems based on tickets that should not have been classified.
[0007] In addition to one or more of the above or below features, the machine learning model may be used in conjunction with a restraint method to identify a given record that is outside of the IT domain as unclassified, which advantageously enables the machine learning model to avoid classifying a given ticket that is not in the IT domain, thereby preventing an automated system from inadvertently modifying software and / or hardware components of one or more computer systems based on a ticket that should not have been classified.
[0008] In addition to one or more of the above or below features, the machine learning model is trained on training data in the IT domain. Based on this training, the machine learning model becomes more accurate and specific to the IT domain. This advantageously enables the machine learning model to avoid classifying a given ticket that is not in the IT domain, thereby preventing an automated system from inadvertently modifying software and / or hardware components of one or more computer systems based on a ticket that should not have been classified.
[0009] In addition to one or more of the above or below features, the machine learning model is trained by receiving training data input including training records and labels corresponding to the training records, and the classifier employing the machine learning model is temporarily restrained from classifying any records outside the IT domain by utilizing a restraint mechanism. The overall classification becomes more accurate and specific to the IT domain based on this training. This advantageously enables the machine learning model to avoid classifying a given ticket that is not in the IT domain, thereby preventing an automated system from accidentally modifying software and / or hardware components of one or more computer systems based on a ticket that should not have been classified.
[0010] In addition to one or more of the above or below features, the records and labels are provided to an automated resolution system configured to modify at least one component in an IT environment of the industry, thereby advantageously providing automated modifications to software and / or hardware components of one or more computer systems in response to technical problems in the IT environment.
[0011] In addition to one or more of the above or below features, a given record outside the scope of the IT domain is prevented from being provided to an automated resolution system, thereby avoiding modification of any component in the IT environment based on an incorrect classification of a given record. This advantageously provides automated modifications to software and / or hardware components of one or more computer systems in response to technical problems in said IT environment. As a result, this advantageously enables said classification mechanism to avoid classifying a given ticket that is not in the IT domain, thereby preventing an automated system from mistakenly modifying software and / or hardware components of one or more computer systems based on a ticket that should not have been classified.
[0012] According to one or more embodiments, a non-limiting computer-implemented method includes inputting, by a processor, records associated with an information technology (IT) domain into a linear classifier (machine learning algorithm). The non-limiting method includes classifying, by the processor, the labeled records using a linear classification algorithm, where the linear classification algorithm refrains from classifying a given record outside the scope of the IT domain in response to the given record.
[0013] This provides an improvement over known methods for classification by providing a restraining mechanism with the ability to restrain itself from classifying tickets that are not in the IT domain. The classification mechanism becomes more accurate and specific to the IT domain. Furthermore, this prevents tickets from being automatically processed by an automated system to erroneously modify software and / or hardware components of one or more computer systems based on tickets that should not have been classified.
[0014] In addition to one or more of the above or below features, the self-control mechanism (using a linear classification algorithm) identifies a given record that is outside of the IT domain as unclassified, which advantageously enables the machine learning model to avoid classifying a given ticket that is not in the IT domain, thereby preventing an automated system from inadvertently modifying software and / or hardware components of one or more computer systems based on a ticket that should not have been classified.
[0015] In addition to one or more features described above or below, a linear classification algorithm is trained on training data from the IT domain. A self-control mechanism extracts pertinent positive (PP) features and pertinent negative (PN) features associated with the label. These PPs and PNs are validated by a domain expert. This advantageously improves the accuracy of the classifier, allowing the classifier to avoid classifying a given ticket that is not in the IT domain, thereby preventing an automated system from inadvertently modifying software and / or hardware components of one or more computer systems based on a ticket that should not have been classified.
[0016] Other embodiments of the present invention implement features of the above methods in computer systems and computer program products.
[0017] Additional technical features and advantages are realized through the techniques of the present invention. Embodiments and aspects of the invention are described in detail herein and are considered part of the subject matter of the claims. For a better understanding, please refer to the detailed description and drawings.
[0018] The details of the exclusive rights set forth herein are particularly pointed out and distinctly set forth in the claims at the end of this specification. The foregoing and other features and advantages of embodiments of the present invention will become apparent from the following detailed description taken in conjunction with the accompanying drawings. [Brief explanation of the drawings]
[0019] [Figure 1] FIG. 1 is a block diagram of an example computer system for use with one or more embodiments of the present invention. [Figure 2] FIG. 1 is a block diagram of an example of a system configured to provide explainable classification with restraint using a client-independent machine learning model, in accordance with one or more embodiments of the present invention. [Figure 3] 1 is a flowchart of a computer-implemented method for training a machine learning model for each classification label in accordance with one or more embodiments of the present invention. [Figure 4A] FIG. 1 is a block diagram illustrating an example of a confusion matrix for a machine learning model, in accordance with one or more embodiments of the present invention. [Figure 4B] FIG. 1 is a block diagram illustrating an example of a coefficient matrix for a machine learning model, in accordance with one or more embodiments of the present invention. [Figure 5A] 1 is an example of a chart illustrating regions of positively associated features determined to positively contribute to each of the predicted labels during classification by a machine learning model, in accordance with one or more embodiments of the present invention. [Figure 5B] 1 is an example of a chart illustrating how certain features contribute to determining a class label, in accordance with one or more embodiments of the present invention. [Figure 5C]1 is a graph illustrating the contribution of features analyzed by a machine learning model to determine a classification, in accordance with one or more embodiments of the present invention. [Figure 6] 1 is a flowchart of a computer-implemented method for computing positively associated features for classification in accordance with one or more embodiments of the present invention. [Figure 7] 1 is a flowchart of a computer-implemented method for explaining machine learning decisions using disjunctive normal forms, according to one or more embodiments of the present invention. [Figure 8] FIG. 1 is a block diagram illustrating converting a linear classification formula into a disjunctive normal form in accordance with one or more embodiments of the present invention. [Figure 9] 1 is a flowchart of a computer-implemented method for explaining machine learning decisions using positively associated features, in accordance with one or more embodiments of the present invention. [Figure 10] FIG. 1 is a block diagram illustrating selecting features (e.g., tokens) from ticket data and presenting positively associated features that contribute to a class label for a ticket, according to one or more embodiments of the present invention. [Figure 11] 1 is a flowchart of a computer-implemented method for providing explainable classification with restraint using a client-independent machine learning model, in accordance with one or more embodiments of the present invention. [Figure 12] 1 is a flowchart of a computer-implemented method for providing explainable classification with restraint using a client-independent machine learning model, in accordance with one or more embodiments of the present invention. [Figure 13] FIG. 1 illustrates a cloud computing environment in accordance with one or more embodiments of the present invention. [Figure 14] FIG. 1 illustrates an extraction model layer in accordance with one or more embodiments of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0020] One or more embodiments provide explainable classification with self-restraint using a client-agnostic machine learning model. For a set of all tickets present in a client's environment, the client-agnostic machine learning model is configured to determine a next discretionary action to resolve a computer / network issue. The client-agnostic machine learning model is trained on tickets and solutions from a variety of different clients, including clients across different industries such as cybersecurity, finance, government, manufacturing, schools, retail, cloud computing, and data storage, to be client-agnostic or industry-agnostic. According to one or more embodiments, the client-agnostic machine learning model is trained to know when to classify a ticket and when not to classify a ticket when analyzing it. This can be achieved by building an agnostic machine learning model that only classifies areas it understands, such as information technology (IT) domains or IT environments, but refrains from classifying areas it does not understand. Furthermore, one or more embodiments are configured to explain how the agnostic machine learning model arrived at a particular decision, which gives a user confidence in the agnostic machine learning model's decisions.
[0021] Incident identification and automated resolution are processes that manage IT service disruptions and restore service. For example, a monitoring system monitors a client's IT environment in an industry. The term "IT environment" refers to the infrastructure, hardware, software, and systems that a client (entity or business) relies on daily in the course of using information technology. Some commonly used resources in an IT environment include computers, Internet access, and peripheral devices. Examples of an IT environment include hardware routers, personal computers, servers, switches, and data centers; software user applications that enable and enable hardware connections, web servers, and applications; and the network, i.e., firewalls, cables, and other components that facilitate internal and external communications within a business. Upon detecting a technical event in the IT environment and / or at the request of a user of the IT environment, the monitoring system generates a ticket. The ticket is sent to an automated resolution system, the IT department, or both for resolution. A ticket is a dedicated document or record that represents an incident, alert, request, or event, or a combination thereof, that requires action from the IT department. A ticket is a historical document that details a service event, such as an incident, problem, or service request, or a combination of these. The ticket governs and controls how the service event is handled.
[0022] A typical system may monitor the environment and attempt to identify problems. However, while general language classifiers handle general text well, they struggle with tickets containing technical data and do not explain how they arrived at their decision. Furthermore, classifiers are poor at restraining themselves in their classifications. For example, a ticket stating that "the first man in space was in 1961" should not be classified as a technical problem / challenge, but many classifiers will analyze the ticket, detect the word "space," and incorrectly classify it as a disk handler problem.
[0023] Technical solutions and benefits include a system, according to one or more embodiments, that provides a client / industry-agnostic model for an IT environment. Thus, thousands of different agnostic machine learning models are not needed for thousands of different clients or industries, but the agnostic machine learning model operates across a variety of clients in different industries. In one or more embodiments, the agnostic machine learning model is configured to refrain from classifying inputs outside the IT environment or IT domain. This enables the agnostic machine learning model (e.g., a classifier) to avoid misclassifying tickets with labels for automated resolution by an automated resolution system when the ticket (as input) is not in the IT environment or IT domain and therefore should not actually generate a label. Yet another technical solution and benefit may include providing machine learning explanations for decision-making to the agnostic machine learning model using disjunctive normal forms (DNFs) and / or positively associated features. One or more embodiments use gradient threshold reduction to achieve dimensionality reduction and extract positively associated values as a list of all features that influence the agnostic machine learning model (e.g., a classifier). Some embodiments may not have these potential benefits or advantages, and these potential benefits or advantages are not necessarily required for all embodiments.
[0024] One or more embodiments described herein may utilize machine learning techniques to perform tasks such as classifying features of interest. More specifically, one or more embodiments described herein may incorporate and utilize rule-based decision-making and artificial intelligence (AI) reasoning to accomplish various operations described herein, i.e., classifying features of interest. The phrase "machine learning" broadly describes the ability of an electronic system to learn from data. A machine learning system, engine, or module may include trainable machine learning algorithms that can be trained, for example, in an external cloud environment, to learn functional relationships between inputs and outputs, and the resulting model (sometimes referred to as a "trained neural network," "trained model," "trained classifier," or "trained machine learning model," or combinations thereof) can be used, for example, to classify features of interest.
[0025] Returning to FIG. 1 , a computer system 100 is generally illustrated in accordance with one or more embodiments of the present invention. Computer system 100 may be an electronic computer framework that includes and / or employs any number and combination of computing devices and networks utilizing various communication technologies, as described herein. Computer system 100 may be easily expanded, extended, and modular, allowing different services to be configured or some features to be reconfigured independently of one another. Computer system 100 may be, for example, a server, desktop computer, laptop computer, tablet computer, or smartphone. In some examples, computer system 100 may be a cloud computing node. Computer system 100 may be described in the general context of computer-system-executable instructions, e.g., program modules, executed by the computer system. Generally, program modules may include routines, programs, objects, components, logic, data structures, etc. that perform particular tasks or implement particular abstract data types. Computer system 100 may also operate in a distributed cloud computing environment where tasks are performed by remote processing devices that are linked through a communications network. In a distributed cloud computing environment, program modules may be located in both local and remote computer system storage media, including memory storage devices.
[0026] As shown in FIG. 1, computer system 100 includes one or more central processing units (CPUs) 101a, 101b, 101c, etc. (collectively or generically referred to as processor 101). Processor 101 may be a single-core processor, a multi-core processor, a computing cluster, or any number of other configurations. Processor 101, also referred to as processing circuitry, is coupled to system memory 103 and various other components via system bus 102. System memory 103 may include read-only memory (ROM) 104 and random access memory (RAM) 105. ROM 104 is coupled to system bus 102 and may include a basic input / output system (BIOS) or its successor, such as a unified extensible firmware interface (UEFI), that controls certain basic functions of computer system 100. RAM is a read-write memory coupled to system bus 102 for use by processor 101. System memory 103 provides temporary memory space for the operation of the above instructions during operation and may include random access memory (RAM), read-only memory, flash memory, or any other suitable memory system.
[0027] Computer system 100 includes an input / output (I / O) adapter 106 and a communications adapter 107 coupled to system bus 102. I / O adapter 106 may be a small computer system interface (SCSI) adapter that communicates with a hard disk 108 and / or any other similar component. I / O adapter 106 and hard disk 108 are collectively referred to herein as mass storage 110.
[0028] Software 111 for execution on computer system 100 may be stored in mass storage 110. Mass storage 110 is an example of a tangible storage medium readable by processor 101, on which software 111 is stored as instructions for execution by processor 101 to cause computer system 100 to operate, for example, as described herein below with reference to various figures. Examples of computer program products and the execution of such instructions are discussed in further detail herein. Communications adapter 107 interconnects system bus 102 to network 112, which may be an external network, thereby enabling computer system 100 to communicate with other such systems. In one embodiment, portions of system memory 103 and mass storage 110 collectively store an operating system, which may be any suitable operating system that coordinates the functions of the various components shown in FIG. 1.
[0029] Additional input / output devices are shown connected to system bus 102 via display adapter 115 and interface adapter 116. In one embodiment, adapters 106, 107, 115, and 116 may be connected to one or more I / O buses that are connected to system bus 102 via intermediate bus bridges (not shown). Display 119 (e.g., a screen or display monitor) is connected to system bus 102 by display adapter 115, which may include a graphics controller and a video controller to enhance performance of graphics-intensive applications. Keyboard 121, mouse 122, speaker 123, microphone 124, etc., may be interconnected to system bus 102 via interface adapter 116, which may include, for example, a super I / O chip-integrated multi-device adapter within a single integrated circuit. Suitable I / O buses for connecting peripheral devices such as hard disk controllers, network adapters, and graphics adapters typically include common protocols such as Peripheral Component Interconnect (PCI) and Peripheral Component Interconnect Express (PCIe). Thus, as configured in Figure 1, computer system 100 includes processing capabilities in the form of processor 101, storage capabilities including system memory 103 and mass storage 110, input means such as keyboard 121, mouse 122, and microphone 124, and output capabilities including speakers 123 and display 119.
[0030] In some embodiments, communications adapter 107 can transmit data using any suitable interface or protocol, including an Internet small computer system interface. Network 112 can be a cellular network, a radio network, a wide area network (WAN), a local area network (LAN), or the Internet, among others. External computing devices can connect to computer system 100 through network 112. In some examples, the external computing device can be an external web server or a cloud computing node.
[0031] It should be understood that the block diagram of Figure 1 is not intended to indicate that computer system 100 includes all of the components shown in Figure 1. Rather, computer system 100 may include any suitable fewer or additional components not shown in Figure 1 (e.g., additional memory components, embedded controllers, modules, additional network interfaces, etc.). Furthermore, the embodiments described herein with reference to computer system 100 may be implemented with any suitable logic, which, as referred to herein, may include any suitable hardware (e.g., a processor, embedded controller, application specific integrated circuit, among others), software (e.g., an application, among others), firmware, or any suitable combination of hardware, software, and firmware in various embodiments.
[0032] FIG. 2 illustrates a block diagram of an example system 200 configured to provide explainable classification with self-control using a client-independent machine learning model, according to one or more embodiments. System 200 includes a computer system 202 configured to communicate with numerous different computer systems across a network 250, such as computer system 240A for managing an IT environment for one client in one industry, computer system 240B for managing an IT environment for another client in another industry, through computer system 240N for managing an IT environment for yet another client in a different industry. Computer systems 240A, 240B, through 240N may be generally referred to as computer systems 240. Each computer system 240 has its own IT management system 244 for monitoring the IT environment for a respective client in the respective industry and storing the respective tickets and solutions in a ticket repertoire 246. Ticket repertoire 246 is operable to store numerous tickets and their respective solutions for the IT environment of computer system 240. The network 250 can be a wired or wireless communication network.
[0033] The IT management system 244 may include or represent a monitoring ticketing system and an automated resolution system for each client in the industry. By the software application 204 communicating with the computer system 240 over the network 250, which may be a wired or wireless communications network, the software application 204 is configured to retrieve various tickets and their respective solutions in the ticket repertoire 246 from different clients in different industries.
[0034] As illustrated by the dotted lines, in one or more embodiments, computer system 202 may include an IT management system 244 and its ticket repertoire 246 for one or more computer systems 240A-240N in each IT environment of each client. Computer system 202 may manage the client's IT environment for one or more computer systems 240A-240N. Any portion of system 200, including computer system 202 and one or more computer systems 240A-240N, may be part of a cloud computing environment 50 as discussed further herein (shown in FIG. 13).
[0035] In system 200, computer system 202, computer systems 240A-240N, IT management system 244, software application 204, training data 206, machine learning model 220, automated resolution system 222, rule generation algorithm 224, etc., can include and / or use any discussed functionality in computer system 100, where computer system 100 includes various hardware components and various software applications, such as software 111, that can execute as instructions on one or more processors 101 to perform actions according to one or more embodiments of the present invention. Software application 204 can include, incorporate, or call, or any combination of, various other software, algorithms, application programming interfaces (APIs), etc., that operate as discussed herein. Software application 204 represents numerous software applications.
[0036] The tickets and their respective solutions are stored in a repertoire, such as storage, as training data 206. The software application 204 filters the training data 206 to ensure that the training data 206 resides only in IT environments, which may also be referred to as IT domains or IT spaces. The IT domain encompasses the IT environments of clients across industries. Any tickets not related to events residing in the IT domain (e.g., errors, issues, security breaches, faulty computer equipment, etc.) are removed from the training data 206.
[0037] The computer system 202 includes a machine learning model 220, which is a client-independent machine learning model that is trained to classify tickets in an IT environment. In one or more embodiments, the client-independent machine learning model is trained, for example, only to classify tickets in an IT environment. The machine learning model 220 may represent a number of machine learning models 220. The machine learning model 220 classifies tickets by predicting a label that identifies how to solve a computer problem associated with the ticket. The ticket and its predicted label can be sent to an automated resolution system 222 to automatically solve the ticket's computer problem according to the predicted label output from the machine learning model 220. In one or more embodiments, the machine learning model 220 is a linear classifier and processes a linear classification algorithm. Terms such as label, class, classification, classification label, and class label may be used interchangeably to refer to machine learning categories.
[0038] Linear classification algorithms use features of an object, such as the features of a ticket, to determine which class (or group) it belongs to. A linear classifier accomplishes this by making a classification decision based on the value of a linear combination of features. The object's features, also known as feature values, are typically presented to the machine in a vector called a feature vector. Examples of linear classification algorithms and techniques include the naive Bayes algorithm, linear discriminant analysis (LDA) algorithm, least squares algorithm, support vector machine algorithm, ridge regression algorithm, lasso algorithm, elastic net algorithm, least angle regression algorithm, orthogonal matching pursuit algorithm, Bayesian regression algorithm, logistic regression algorithm, linear regression algorithm, perceptron algorithm, passive-aggressive classification algorithm, etc., as will be understood by those skilled in the art.
[0039] The machine learning model 220 can be configured with a trained linear classification algorithm for each classification label for a ticket. Furthermore, the machine learning model 220 is configured to refrain from classifying tickets that do not belong to the IT environment or IT domain. In one or more embodiments, there may be a classification label labeled unclassified / unknown, and the machine learning model 220 may be configured to use the unclassified / unknown label to indicate that the ticket's feature vector (i.e., features) does not apply to the IT environment or IT domain. By refraining from classifying tickets that originate from and / or are not related to the IT environment as unclassified / unknown, or by classifying such tickets, or both, the machine learning model 220 is configured to prevent misclassified tickets from being erroneously submitted to the automated resolution system 222 and corresponding automated corrective actions being taken by the automated resolution system 222 on the IT environment. In one or more embodiments, the machine learning model 220 can prevent misclassified tickets from being erroneously submitted to the automated resolution system 222 and corresponding incomplete automated corrective actions being taken on the IT environment. For example, one or more software and / or hardware components within the IT environment may be automatically altered by an automated resolution system based on an incorrect classification label in a ticket, thereby causing a malfunction of software and / or hardware components of computer systems within the IT environment.
[0040] 3 is a flowchart of a computer-implemented method 300 for training a machine learning model 220 with a linear classification algorithm for each classification label, resulting in a trained machine learning model 220, according to one or more embodiments. The computer-implemented method 300 is performed by a computer system 202.
[0041] At block 302 of the computer-implemented method 300, the software application 204 is configured to retrieve training data by compiling data from each ticket into training data stored in the training data 206. The software application 204 may be configured to parse the training data 206 to determine and filter any tickets along with their solutions that are not in the IT environment. This leaves only tickets that are not in the IT environment that represent the IT domain, so that the machine learning model 220 is trained to learn about tickets and their respective solutions that are in the IT domain. In one or more embodiments, for example, the machine learning model 220 is only trained to learn about tickets and their respective solutions that are in the IT domain. Upon detecting a ticket that is not in the IT domain, the machine learning model 220 refrains from labeling the ticket, or the ticket may be labeled as unclassified / unknown, or both, which results in the ticket being prevented / blocked from processing by the automated resolution system 222.
[0042] The training data 206 can be refined using cross-fold validation. For the training data 206, tickets are labeled in preparation for training the machine learning model 220. For the training data 206, these labels can be automation playbooks obtained by matching automation executions with the incident tickets being addressed. For example, the machine learning model 220 can be trained on tickets that are resolved by an automation platform (e.g., IT management system 244), such as the RedHat® Ansible® platform, and on tickets that are not resolved by the automation platform. For example, when the automation platform (e.g., IT management system 244) receives a ticket, it logs into the system. If it has a playbook, the automation platform runs the playbook, resolves the ticket, and closes the ticket. A closed ticket is considered to be in fact or completed. The tickets are utilized to train the machine learning model 220 into one of the known class examples in the IT domain, such as, for example, application down, database space issue, disk handler, network connectivity, file system mount handler, high disk space usage handler, high memory and page file usage, host down handler, service handler, job abend, etc. Of course, the above exemplary enumeration of classification labels is not meant to be comprehensive.
[0043] In block 304, the software application 204 is configured to train the machine learning model 220 using the training data 206. For example, each ticket's data (along with its corresponding label) is input to the machine learning model 220 as a feature vector so that the linear classifier algorithm of the machine learning model 220 learns how to classify the ticket's input data. The training data is labeled, meaning that the tickets are pre-labeled to determine when the output of the machine learning model 220 predicts the correct label. During the training phase, the predicted labels of the tickets from the machine learning model 220 are compared to the labels of the tickets in the training data 206 to continuously improve the machine learning model 220. This enables the machine learning model 220 to learn the correct classification label for each ticket.
[0044] In block 306, the machine learning model 220 is configured to classify the ticket input data with an explanation. For example, each ticket is classified based on a basis and / or decisions made by the machine learning model 220. The machine learning model 220 is configured to generate an explanation to the user in terms of the machine learning rules and / or as positively associated features on which the machine learning model 220 bases its decision for the predicted label of the input ticket. In one or more embodiments, the machine learning model 220 may comprise and / or employ a rule generation algorithm 224 for the machine learning rules and / or identified positively and negatively associated features, as further details are discussed herein.
[0045] At block 308, the software application 204 is configured to validate the classification result by comparing the predicted classification label of the ticket with the labels of the tickets in the training data. At block 310, the software application 204 is configured to verify whether the predicted classification label of the ticket matches the labels of the tickets in the training data.
[0046] In block 312, if there is a match (Yes), the software application 204 is configured to verify whether the explanation for the machine learning rule and / or the identified positively associated features is satisfactory. This may include requesting input from a subject matter expert, employing a natural language processor (NLP) system, or both. If the explanation is satisfactory (Yes), the software application 204 is configured to terminate training of the machine learning model.
[0047] In block 314, if (No) the classification results are not satisfactory for the decision in block 310, the software application 204 is configured to improve the training data, tune the classifier algorithm, or both. Also, if (No) the explanation is not satisfactory for the decision in block 312, the software application 204 is configured to improve the explanation.
[0048] Note that a separate linear classification algorithm is trained for each classification label. In one or more embodiments, the machine learning models 220 can include multiple linear classification algorithms, one for each classification label corresponding to the IT domain. In one or more embodiments, there may be multiple machine learning models 220, one for each classification label corresponding to the IT domain. Any discussion of a single linear classification algorithm and / or a single machine learning model 220 applies by analogy to all linear classification algorithms and / or all machine learning models 220 corresponding to all classification labels in the IT domain.
[0049] As described herein, when an input ticket tells a typical classifier that "the first man went into space in 1961," the typical classifier attempts to classify the ticket with a label, such as disk or disk handler. However, such a label, in this example, is a misclassification, potentially leading to an incorrect action being taken by the automated resolution system.
[0050] As a technical benefit and solution, one or more embodiments are configured to refrain from classifying such tickets indicating "the first man in space was in 1961" because machine learning model 220 has been trained to refrain from classifying such tickets. Instead, machine learning model 220 may output unknown / unclassified, thereby preventing automated resolution system 222 from modifying one or more software and / or hardware components in the IT environment of an industrial client. Thus, based on the output from machine learning model 220, software application 204 may recognize that the ticket is unknown / unclassified in the IT domain and, instead of sending the ticket to automated resolution system 222, may send the ticket to a specialized IT department for resolution. Software application 204 is configured to send tickets that have been properly labeled by machine learning model 220 to automated resolution system 222 for automatic processing in the IT environment of an industrial client. Following the label from the machine learning model 220, the automated solving system 222 is configured to modify software components, hardware components, or both software and hardware components of one or more computer systems in an IT environment, resulting in improvements to the computer systems themselves. The modification of software components, hardware components, or both, on computer systems in an IT environment to solve technical computer problems is a practical application related to the use of the machine learning model 220.
[0051] One or more embodiments provide an approach for computing positively associated features in a ticket for text classification. The positively associated features are tokens in the ticket. A token can refer to one or more words, phrases, sentences, etc. in the ticket text, and in a process sometimes called tokenization. The tokens can be used as features in the ticket's feature vector. One or more embodiments extract a list of all positively associated features for all (IT) tickets, along with their labels, which can be used to further train the machine learning model 220 and for use in the explanations discussed herein.
[0052] Additionally, one or more embodiments are configured to generate a linear classifier for the machine learning model 220, generate coefficient and confusion matrices for insight into how the linear classifier performs at a corpus level, extract positive relevance values for all IT tickets using gradient descent iterative threshold shrinkage, curate the positive relevance values with subject matter experts / IT domain experts to find "true" positive relevance values, extract rules and rule features as new training data refines a given label, and create a rich training model and a list of regions of allowed positive relevance values. Thus, one or more embodiments can receive incoming new incident tickets and extract positive relevance values for a given ticket; if no positive relevance values are identified, the linear classifier refrains from classifying the ticket.
[0053] Referring to Figure 4A, a block diagram shows an example of a confusion matrix for a machine learning model according to one or more embodiments. Figure 4B illustrates a block diagram of an example of a coefficient matrix for a machine learning model according to one or more embodiments. The ticket and label features of the confusion matrix in Figure 4A are projected into the example coefficient matrix illustrated in Figure 4B.
[0054] By probing the machine learning model 220 to learn from its errors, the software application 204 can generate an example confusion matrix that captures how the classifier confuses itself when learning and classifying the training data in FIG. 4A. A confusion matrix is an N×N matrix used to evaluate the performance of a classification model, where N is the number of target classes. This matrix compares actual target values in the training data 206 with values predicted by the machine learning model.
[0055] Referring to FIG. 4B, the coefficient matrix contains entries representing the signed coefficients of each feature for the hyperplane in the binary linear classifier for each class. The software application 204 is configured to identify the feature with the largest absolute coefficient, e.g., the largest absolute coefficient value. Using the coefficient matrix, this provides the software application 204 with the features that have the greatest impact on the classification on a global scale. The coefficient matrix, sometimes called a correlation matrix, is a table that displays correlation coefficients for different variables. The coefficient matrix illustrates the correlations between all possible pairs of values in the table. It summarizes large data sets and identifies and visualizes patterns in the given data.
[0056] In Figure 4B, a feature with a positive value for a given class label indicates that the corresponding feature positively contributes to the machine learning model 220's decision to classify the given class label, while a feature with a negative value (i.e., a negative sign) for a given class label indicates that the corresponding feature does not contribute (i.e., a negative contribution) to the classification of the given class label.
[0057] Based on all features in the ticket (e.g., in training data 206) for all class labels, software application 204 using machine learning model 220 generates a region of positively associated features for each predicted classification label, as illustrated in FIG. 5A. In FIG. 5A, an example chart illustrates positively associated features, also referred to as positive association values, that were determined to positively contribute to each predicted label during classification by machine learning model 220. For example, the positively associated features "low space," "disk:c handler," and "file system" are positively associated features that positively contribute to the predicted classification label "disk handler" during classification by machine learning model 220. This classification label disk handler example is used for illustrative purposes in various example scenarios and is not limiting. Of course, embodiments are not limited to the classification label disk handler.
[0058] FIG. 5B is a chart illustrating how example feature "spaces" contribute to the determination of a classification label. As seen in FIG. 5B, the feature space has a positive coefficient for the classification label disk handler, meaning that the feature space positively contributes to the machine learning model 220's decision to output the classification label disk handler. Similarly, the feature space has a negative coefficient for some other classification labels, meaning that the feature space negatively contributes or has no influence on the machine learning model 220's decision to output the corresponding class label. Therefore, a feature space can be identified as a negatively associated feature and removed as a feature for the corresponding classification label, in which case the space has a negative coefficient. The software application 204 continues removing features because, as shown in FIG. 5A, any identified feature has a negatively associated feature with a negative coefficient for a given classification label, thereby leaving only the positively associated features available in the space of positively associated features for each classification label.
[0059] Continuing with the example scenario for the classification label disk handler, FIG. 5C is a graph illustrating the contribution of features being analyzed by machine learning model 220 to make its classification of disk handler, which features are now verified by a human subject matter expert. For each feature, the graph shows its negative contribution to the decision by machine learning model 220 to classify the ticket with the label disk handler, as well as its positive contribution to the decision. To verify with the subject matter expert that features with positive association values of, for example, “low space,” “disk handler,” and “file system” are identified as positively contributing to the machine learning model's 220 decision, software application 204 identifies the positively associated features for a given class. As identified by software application 204, the subject matter expert determines that all positively associated features are true positively associated features for their respective classification labels, resulting in a region of positively associated features for those labels, as illustrated in FIG. 5A. In one or more embodiments, the region of positively associated features excludes any negatively associated features. Therefore, the positively associated features for their respective classification labels are collected and added to the training data 206 as a new training data set in the training data 206 for further training the machine learning model 220 to classify tickets with the classification labels. In one or more embodiments, the positively associated features are established to be positively associated features for their respective classification labels. This additional training further refines the ability of the machine learning model 220 to learn to refrain from classifying tickets that are not in the IT domain and improves the accuracy of the machine learning model 220.
[0060] During the inference phase, once the machine learning model 220 receives a ticket and outputs its classification label, the software application 204 can probe the machine learning model 220 to obtain positively associated features for any ticket. For incoming tickets, the machine learning model 220 is configured to learn to extract and recognize positively associated features for a given ticket, and to refrain from classifying the ticket when no positively associated features are found in the ticket. On the other hand, when positively associated features are recognized in the ticket, the machine learning model 220 is configured to classify the ticket, highlight the positively associated features for display to the user (on the display 119), and extract classifier rules using disjunctive normal forms to explain the machine learning model's decisions.
[0061] 6 is a flowchart of a computer-implemented method 600 for computing positively associated features for text classification, according to one or more embodiments. In one or more embodiments, the software application 204 employs, utilizes, integrates with, or combines the machine learning model 220 to perform the computer-implemented method 600. Additionally, the software application 204 may be utilized to explore the machine learning model 220 to perform the computer-implemented method 600.
[0062] In blocks 602, 604, and 606, the software application 204 is configured to input the text of the incident ticket to a preprocessor to generate a feature vector from the input text. The preprocessor extracts input features from the text, which are formed into a feature vector. Known techniques can be used to convert the text into a feature vector. The feature vector is input to a machine learning model 220, which outputs a classification label for the corresponding ticket.
[0063] In block 608, the software application 204 is configured to construct a confusion matrix and generate a coefficient matrix for each label output by the machine learning model 220. Examples of a confusion matrix and a coefficient matrix are illustrated in Figures 4A and 4B, respectively.
[0064] In block 610, the software application 204 is configured to provide the feature vector and coefficient matrix to an L1 (or L2, or both) regularization problem that models positive correlations. An example of a regression model using an L1 regularization technique is called least absolute shrinkage and selection operator (lasso) regression, and an example of a regression model using an L2 regularization technique is called ridge regression. For L1 regularization, lasso regression adds the absolute magnitude of the coefficients to the loss function as a penalty term. L1 regularization may be an option when there are a large number of features that give a sparse solution. For L2 regularization, ridge regression adds the squared magnitude of the coefficients to the loss function as a penalty term. L2 regression can be used to estimate the significance of predictors and may be based on penalizing insignificant predictors. Furthermore, elastic net combines L1 and L2 regularization, resulting in an elastic net method that adds hyperparameters.
[0065] In block 612, the software application 204 is configured to apply an iterative shrinkage / thresholding algorithm (ISTA) to the L1 (and / or L2) control problem. ISTA is widely used in solving linear inverse problems due to its simplicity. ISTA may involve bidirectional threshold shrinkage of gradient descent.
[0066] In blocks 614 and 616, the software application 204 is configured to generate a sparse feature vector by using the ISTA method and select positively associated features from the sparse feature vector. In one or more embodiments, the sparse feature vector for a classification label has fewer features than the original feature vector. Thus, using all of the sparse feature vectors generated for each classification label in the IT domain, the software application 204 selects features for each classification label to generate a list of positively associated features for each class label, which are output in block 618. As mentioned above, FIG. 5A illustrates a region of positively associated features for each classification label, such as an example of a class label disc handler.
[0067] To further explain the computation of positively associated features for text classification, the following is an illustrative, non-limiting example scenario: The ISTA algorithm is used to compute a sparse solution for inverting a linear problem. A typical example of an inverse linear problem is linear regression. One example for consideration is a classification problem, e.g., a text classification problem. For text classification problems in the IT ticket management domain, any linear classification algorithm can be used. Now, consider ticket T that was classified into class C (e.g., "Disk Handler") using a linear classification algorithm in machine learning model 220. It is often useful to provide "evidence" of the inner workings of a classifier (e.g., machine learning model 220) and "explain" why ticket T was classified into class C. The positive association values in a ticket such as T are a small subset of the features in T that are responsible for its classification into class C. Such a set of features provides a good explanation for the inner workings of a classifier according to one or more embodiments. As discussed in the above example, this disclosure formulates the problem of finding positively associated features for a text classification problem as a sparse inverse linear problem. One or more embodiments customize and simplify the ISTA algorithm to make it efficient for this use case. This customization devise a specific problem formulation, such as the one used internally by passive-aggressive classifiers (PACs), and uses its structural properties to efficiently perform the iterative thresholding step.
[0068] One or more embodiments provide explainability for machine learning model decisions, which explains why the machine learning model 220 classified a given ticket with a given classification label. A typical IT domain has several stakeholders, such as IT users, IT workers, service availability managers, incident ticket owners, incident ticket assignees, and change owners. Different stakeholders may require explanations with different levels of complexity and depth of reasoning. According to one or more embodiments, the explainability of machine learning decisions can be presented in disjunctive normal form, using positively related features, or both. In Boolean logic, disjunctive normal form is the canonical normal form of a logical formula consisting of a disjunction of conjunctions. Disjunctive normal form can be described using terms such as "OR" and "AND."
[0069] Returning to explainability using disjunctive normal forms, Figure 7 is a flowchart of a computer-implemented method 700 for explaining machine learning decisions using disjunctive normal forms, according to one or more embodiments. Figure 7 is described with reference to Figure 8, which is a block diagram illustrating the conversion of a linear classification formula to disjunctive normal form, according to one or more embodiments. The machine learning model 220 receives an input of a ticket and outputs a classification label, for example, as a disk handler.
[0070] In block 702 of the computer-implemented method 700, the software application 204 is configured to extract a linear classification formula from the machine learning model 220. An example of a linear classification formula is shown as β in block 802 of FIG. 1i f1+β 2i f2+…+β ki f k where β ki represents the exemplary coefficient of the kth feature, and f krepresents an example feature of the sum of k features. In FIG. 8, the output classification label is disk handler with coefficients as the weights of each feature (which can be positive and negative weights). In this example, the linear classification formula for the class label for a given ticket is disk handler = 0.2(filesystem + 0.3(mounted) + 0.1(limit) - 0.5(database) - 0.1(cpu) as shown in block 802 of FIG. 8.
[0071] In block 704, the software application 204 is configured to select coefficients (i.e., weights) of features with positive signs to be utilized in disjunctive normal form (i.e., "ANDed" together).
[0072] In block 706, optionally, the software application 204 is configured to select coefficients (i.e., weights) of features with negative signs to be utilized (i.e., “ANDed” together) in the disjunctive normal form, while preserving their negative signs. In one or more embodiments, the software application 204 may employ or invoke a natural language processing (NLP) model 228 to parse and analyze the features and their respective coefficients (negative and positive values) in the linear classification formula when performing blocks 704 and 706. In one or more embodiments, the NLP model 228 may be a pre-trained NLP model that has been further trained on features and their respective coefficients (negative and positive values) in a known linear classification formula to select features and their coefficients for use in the disjunctive normal form.
[0073] In block 708, the software application 204 is configured to convert the features with selected coefficients into a disjunctive normal form (DNF) as rules for the machine learning model 220's decisions. In one or more embodiments, the software application 204 may only select features with coefficients with positive values, may select features with coefficients with positive values above a threshold, or may select features with coefficients with positive values along with features with negative coefficient values above a threshold. An example of a rule in DNF is described in block 804 of FIG. 8. Specifically, block 804 describes an example of three different rules, each separated by "OR," as seen in FIG. 8. Note that the term database has a negative sign indicating its negative contribution to the classification label disc handler. In some embodiments, the negative sign may be replaced with "NOT" to represent a negative contribution. Example rules are displayed to a user of the machine learning model 220 (on the display 119) to explain the decision-making underlying the class label, e.g., disc handler, output by the machine learning model 220 for a given ticket.
[0074] In one or more embodiments, the software application 204 is configured to employ / invoke a rule generation algorithm 224 to generate rules in disjunctive normal form. In one or more embodiments, the rule generation algorithm 224 may be a rule-based algorithm. An example of a rule-based system is a domain-specific expert system that uses rules to make inferences or choices. A rule-based system includes a set of facts or sources of data related to a subject matter under consideration and a set of rules for manipulating that data. These rules are sometimes referred to as "If statements" because they tend to follow the lines of "If Y happens THEN do Y."
[0075] In one or more embodiments, the rule generation algorithm 224 may be a machine learning algorithm that has been trained on training data. For example, the training data may include a linear classification formula for each classification label in the IT domain of tickets, with the training data including positive and negative features and their corresponding coefficients. During the training phase of the rule generation algorithm 224, a subject matter expert / IT specialist may accept or reject rules in the disjunctive normal form for each class, thereby improving the rule generation algorithm as a machine learning algorithm. During the inference phase, when the rule generation algorithm 224 is implemented as a trained machine learning algorithm / model, the trained machine learning algorithm receives input classification labels, features, and their respective coefficients to output a disjunctive normal form of the features and logical terms (e.g., “AND,” “OR,” etc.), as illustrated in block 804 of FIG. 8 .
[0076] In describing explainability using disjunctive normal form, the machine learning model 220 uses a linear classifier, such as a passive-aggressive classifier that includes unigrams and bigrams as features. The presence or absence of these features, in one embodiment, provides the software application 204 with information about the type of incident, e.g., ticket type. Explaining the behavior of the machine learning model 220 (i.e., the classifier), helps users increase their confidence in the classifier. Therefore, one or more embodiments describe classifications in disjunctive normal form, which is a natural form of knowledge representation for humans.
[0077] Returning to explainability using positively associated features, Figure 9 is a flowchart of a computer-implemented method 900 for explaining machine learning decisions using positively associated features, according to one or more embodiments. Figure 9 is described with reference to Figure 10, which is a block diagram illustrating selecting features (e.g., tokens) from ticket data and presenting the positively associated features that contribute to a classification label for the ticket, according to one or more embodiments. A machine learning model 220 receives an input of a ticket and outputs a classification label, e.g., a disk handler.
[0078] In block 902, the software application 204 is configured to extract, for tickets classified into a given classification label, positively associated features from the machine learning model 220. Any of the techniques discussed herein may be utilized to determine the positively associated features for a classification label.
[0079] In block 904, the software application 204 is configured to select the positively associated features with the highest values, e.g., values above a predetermined threshold. Examples of positively associated features are illustrated in Figures 5A and 5C.
[0080] In block 906, the software application 204 is configured to display (e.g., on the display 119) the selected positively associated features with the highest values that contributed to (i.e., influenced) the decision of the machine learning model 220. For example, the selected positively associated features are configured to indicate an explanation associated with each prediction (i.e., what it is about the features of this particular ticket that suggested automating the disc handler). In one or more embodiments, for example, the selected positively associated features are configured to only indicate an explanation associated with each prediction. As displayed to the user (e.g., on the display 119), FIG. 10 shows the machine learning rules extracted in disjunctive normal form in block 1002 and the machine learning features that influenced the decision in block 1004. FIG. 10 also displays the ticket description of the ticket in block 1006 and a bar graph showing the positively associated features that influenced the decision of the machine learning model 220 in block 1008. In block 1008, the feature disc has a greater influence on the decision than the feature space, although both are used to explain the decision of the machine learning model 220.
[0081] 11 is a flowchart of a computer-implemented method 1100 for providing explainable classification with restraint using a client-independent machine learning model, according to one or more embodiments. The computer-implemented method 1100 can be performed by the computer system 202. Reference can be made to any of the drawings discussed herein.
[0082] In block 1102, the machine learning model 220 receives an input of a record (e.g., a ticket), where the record is related to an information technology (IT) domain. In block 1102, the machine learning model 220 classifies the record with a label, where the machine learning model refrains from classifying the given record (e.g., another ticket) outside of the IT domain in response to the given record.
[0083] In one or more embodiments, the machine learning model 220 identifies a given record outside the IT domain as unclassified. The machine learning model 220 is trained on training data 206 in the IT domain. The machine learning model 220 is trained by receiving inputs to the training records of the training data 206 having training records (e.g., tickets) and their corresponding labels. The machine learning model 220 is trained to refrain from classifying any records outside the IT domain.
[0084] Further, the records and labels are provided to an automated resolution system 222, which is configured to alter at least one component (e.g., a hardware component, a software component, or both hardware and software components) in the IT environment of the industry. A given record outside the IT domain is prevented from being provided to the automated resolution system 222, thereby avoiding any component (e.g., a hardware component, a software component, or both hardware and software components) in the IT environment from being altered based on a misclassification of the given record.
[0085] The machine learning model 220 includes a linear classifier algorithm, and the record is a technical problem ticket in an IT environment. The linear classifier algorithm has been trained on training data in the IT domain, and the linear classifier algorithm has been trained on positively associated features for the label without negatively associated features (as illustrated in FIGS. 5A and 5C), and the positively associated features for the label have been verified by human subject matter experts.
[0086] 12 is a flowchart of a computer-implemented method 1200 for providing explainable classification with restraint using a client-independent machine learning model, according to one or more embodiments. The computer-implemented method 1200 can be performed by the computer system 202. Reference can be made to any of the drawings discussed herein.
[0087] In block 1202, the machine learning model 220 classifies an input record (e.g., a ticket), where the machine learning model refrains from classifying a given record (e.g., another ticket) in response to the given record being outside the information technology (IT) domain. In block 1204, the software application 204 and / or the machine learning model 220 generates an explanation of the machine learning model's decision to classify the record along with a label. In block 1206, the software application 204 and / or the machine learning model 220 provides a display (e.g., on the display 119) of the explanation in a human-readable format.
[0088] In one or more embodiments, the human-readable format includes a disjunctive normal form. The explanation of the decision by the machine learning model 220 is based on a linear classification formula utilized by the machine learning model 220. The explanation of the decision by the machine learning model 220 is based on features and their corresponding coefficients, which are derived from the linear classification formula of the machine learning model, as illustrated in block 802 of FIG. 8 . That is, features are extracted from the text (e.g., tokenized text) of the record (e.g., ticket), as illustrated in block 1006 of FIG. 10 . Furthermore, the human-readable format includes a display (e.g., on display 119) of the positively associated features along with a respective contribution of each positively associated feature to the decision by the machine learning model 220, as illustrated in block 1008 of FIG. 10 . Furthermore, the machine learning model is trained on training data in the IT domain. The machine learning model includes a linear classifier algorithm, and the record is a ticket for a technical problem in an IT environment.
[0089] In one or more embodiments, the machine learning model 220, the rule generation algorithm 224, or the NLP model 228, or a combination thereof, can include various engines / classifiers and / or can be implemented on neural networks. The features of the engines / classifiers can be implemented by configuring and arranging the computer system 202 to execute the machine learning algorithm. Generally, the machine learning algorithm actually extracts features from received data (e.g., technical computer problem tickets) in order to “classify” the received data. Examples of suitable classifiers include, but are not limited to, neural networks, support vector machines (SVMs), logistic regression, decision trees, hidden Markov models (HMMs), etc. The end result of the classifier's operation, i.e., “classification,” is to predict a class (or label) for the data. The machine learning algorithm applies machine learning techniques to the received data to create / train / update a unique “model” over time. The learning or training performed by the engine / classifier can be supervised, unsupervised, or a hybrid that includes aspects of supervised and unsupervised learning. Supervised learning is when training data is classified / labeled when it is already available. Unsupervised learning is when training data is not classified / labeled and therefore must be evolved through iterative classifiers. Unsupervised learning can utilize additional learning / training methods, such as clustering, anomaly detection, neural networks, and deep learning.
[0090] In one or more embodiments, the engine / classifier is implemented as a neural network (or artificial neural network) that uses connections between pre-neurons and post-neurons and thus represents connection weights. Connections represent, for example, synapses between pre-neurons and post-neurons. Neuromorphic systems are interconnected elements that function as stimulated "neurons" and exchange "messages" with each other. Similar to the so-called "plasticity" of synaptic neurotransmitter connections that carry messages between biological neurons, connections in neuromorphic systems, such as neural networks, carry electronic messages between stimulated neurons, and these messages are provided with numerical weights that correspond to the strength or weakness of a given connection. These weights can be adjusted or tuned based on experience, allowing the neuromorphic system to adapt to inputs and learn. After being weighted and transformed by a function (i.e., a transfer function) determined by the network designer, activation of these input neurons then travels to other downstream neurons, often referred to as "hidden" neurons. This process is repeated until an output neuron is activated. Thus, activated output neurons determine (or "learn") and provide an output, or inference, about the input.
[0091] A training dataset (e.g., training data 206) can be utilized to train a machine learning algorithm. The training dataset can include historical data of past tickets and corresponding options / suggestions / solutions for each ticket. The option / suggestion labels can be applied to each ticket to train the machine learning algorithm as part of supervised learning. For preprocessing, the raw training dataset can be manually collected and categorized. The classified dataset may be labeled (e.g., using Amazon Web Services® (AWS®) labeling tools, such as Amazon Sage Maker® Ground Truth). The training dataset may be divided into training, testing, and validation datasets. The training and validation datasets are used for training and evaluation, while the testing dataset is used after training and testing the machine learning model on an unseen dataset. The training dataset may be processed through different data augmentation techniques. Training takes a labeled dataset, a base network, a loss function, and hyperparameters; once all of these are created and compiled, neural network training occurs, ultimately resulting in a trained machine learning model (e.g., a trained machine learning algorithm). Once the model is trained, it (including the tuned weights) is saved to a file for development of a test dataset and / or further testing.
[0092] Although this disclosure includes a detailed description of cloud computing, it will be understood that implementation of the teachings described herein is not limited to cloud computing environments. Rather, embodiments of the present invention may be implemented in conjunction with any other type of computing environment now known or later developed.
[0093] Cloud computing is a service delivery model for enabling convenient, on-demand network access to a shared pool of configurable computational resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) and for rapidly provisioning and releasing these resources with minimal administrative effort or interaction with the service provider. This cloud model includes at least five characteristics, at least three service models, and at least four deployment models.
[0094] The features are as follows:
[0095] On-demand self-service: Cloud customers can unilaterally and automatically provision computing power, such as server time and network storage, as needed, without the need for human interaction with the service provider.
[0096] Broad Network Access: The capability is available over the network and can be accessed using standard mechanisms, facilitating use by heterogeneous thin-client or thick-client platforms (e.g., mobile phones, laptops, and PDAs).
[0097] Resource Pool: A provider's computing resources are pooled and offered to multiple consumers using a multi-tenant model, with various physical and virtual resources dynamically allocated and reallocated according to demand. There is a sense of location independence in that consumers typically have no control or knowledge regarding the exact location of the resources they are offered, although they can still specify location (e.g., country, state, or data center) at a higher level of abstraction.
[0098] Rapid Elasticity: Capacity is quickly and elastically provisioned, sometimes automatically, and can be quickly scaled out and quickly released to quickly scale in. Capacity available for provisioning often appears to consumers as unlimited, available for purchase in any quantity at any time.
[0099] Metered Services: Cloud systems leverage metering capabilities to automatically control and optimize resource usage at an abstraction level appropriate to the type of service (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, and reported, providing transparency to both providers and consumers of utilized services.
[0100] The service model is as follows:
[0101] SaaS (Software as a Service): The consumer is provided with the ability to use the provider's applications running on a cloud infrastructure. Those applications can be accessed from a variety of client devices through thin-client interfaces such as web browsers (e.g., web-based email). The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, storage, or individual application features, except for the possibility of setting limited user-specific application configuration settings.
[0102] PaaS (Platform as a Service): The ability offered to a consumer is to deploy applications they create or acquire, written using programming languages and tools supported by the provider, onto a cloud infrastructure. The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, or storage, but does have control over the deployed applications and, in some cases, the configuration of the application hosting environment.
[0103] Infrastructure as a Service (IaaS): The capability provided to a consumer is the provisioning of processing, storage, network, and other basic computing resources, upon which the consumer can deploy and run any software, which may include operating systems and applications. The consumer does not manage or control the underlying cloud infrastructure, but does have control over the operating system, storage, deployed applications, and in some cases, limited control over selected network components (e.g., host firewalls).
[0104] The deployment model is as follows:
[0105] Private Cloud: The cloud infrastructure is operated solely for the organization, and may be managed by the organization or a third party, and may reside on-premises or off-premises, or both.
[0106] Community Cloud: The cloud infrastructure is shared by multiple organizations to support a specific community with shared concerns (e.g., mission, security requirements, policy, and compliance considerations). The cloud infrastructure may be managed by these organizations or a third party, and may reside on-premises or off-premises, or both.
[0107] Public Cloud: This cloud infrastructure is available for use by the general public or large industry organizations and is owned by an organization that sells cloud services.
[0108] Hybrid cloud: This cloud infrastructure is a combination of two or more clouds (private, community, or public) that remain distinct but are joined together by standardized or proprietary technologies that allow for data and application portability (e.g., cloud bursting to balance load between clouds).
[0109] A cloud computing environment is a service-oriented environment that emphasizes statelessness, loose coupling, modularity, and semantic interoperability. At the heart of cloud computing is an infrastructure that consists of a network of interconnected nodes.
[0110] Referring now to FIG. 13 , an exemplary cloud computing environment 50 is shown. As shown, the cloud computing environment 50 includes one or more cloud computing nodes 10, among which local computing devices used by cloud consumers (e.g., personal digital assistants (PDAs) or mobile phones 54A, desktop computers 54B, laptop computers 54C, or automotive computer systems 54N, or combinations thereof) communicate. The nodes 10 communicate with each other. These nodes may be physically or virtually grouped (not shown) in one or more networks, such as a private cloud, community cloud, public cloud, or hybrid cloud, or combinations thereof, as previously described herein. This enables the cloud computing environment 50 to provide an infrastructure, platform, or software-as-a-service, or combinations thereof, that does not require cloud consumers to maintain resources on their local computing devices. The types of computing devices 54A-54N shown in FIG. 13 are intended to be illustrative only, and it is understood that computing node 10 and cloud computing environment 50 can communicate with any type of computer-controlled device via any type of network or network-addressable connection (e.g., a connection using a web browser), or both.
[0111] Referring now to Figure 14, a set of functional abstraction layers provided by cloud computing environment 50 (Figure 13) is shown. It should be understood in advance that the components, layers, and functions shown in Figure 14 are intended to be illustrative only, and that embodiments of the present invention are not limited thereto. As shown, the following layers and corresponding functions are provided:
[0112] Hardware and software layer 60 includes hardware and software components. Examples of hardware components include mainframe 61, RISC (Reduced Instruction Set Computer) architecture-based servers 62, servers 63, blade servers 64, storage devices 65, and networks and network components 66. In some embodiments, software components include network application server software 67 and database software 68.
[0113] The virtualization layer 70 comprises an abstraction layer at which virtual entities such as virtual servers 71, virtual storage 72, virtual networks including virtual private networks 73, virtual applications and operating systems 74, and virtual clients 75 are provided.
[0114] In one example, the management layer 80 may provide the following functions: Resource provisioning 81 provides dynamic procurement of computing and other resources utilized to perform tasks within the cloud computing environment. Metering and pricing 82 provides tracking of costs as resources are utilized within the cloud computing environment and billing or invoicing for the consumption of these resources. In one example, these resources may include application software licenses. Security provides identity verification of cloud consumers and tasks, as well as protection of data and other resources. User portal 83 provides consumers and system administrators with access to the cloud computing environment. Service level management 87 provides allocation and management of cloud computing resources to ensure required service levels are met. Service level agreement (SLA) planning and achievement 88 provides proactive provisioning and procurement of cloud computing resources in anticipation of future demand according to SLAs.
[0115] The Workload Layer 90 shows examples of functions utilized in a cloud computing environment. Examples of workloads and functions provided by this layer include mapping and navigation 91, software development and lifecycle management 92, virtual classroom education delivery 93, data analytics processing 94, transaction processing 95, and workloads and functions 96.
[0116] Various embodiments of the present invention are described herein with reference to the associated drawings. Alternate embodiments may be devised without departing from the spirit of the present invention. While various connections and relationships (e.g., above, below, adjacent, etc.) are described between elements in the following description and drawings, those skilled in the art will recognize that many of the relationships described herein are independent of direction, as the described functionality is maintained even when the orientation is changed. These connections and / or relationships may be direct or indirect unless otherwise specified, and the present invention is not intended to be limited in this respect. Accordingly, a connection between entities may refer to a direct or indirect connection, and a relationship between entities may refer to a direct or indirect relationship. As an example of an indirect relationship, a reference herein to forming layer "A" on layer "B" includes the situation where one or more intermediate layers (e.g., layer "C") are present between layer "A" and layer "B," so long as the relevant features and functionality of layer "A" and layer "B" are not substantially altered by the intermediate layers.
[0117] For the sake of brevity, conventional techniques for making and using aspects of the present invention may or may not be described in detail herein. In particular, various aspects of computing systems and particular computer programs that implement various technical features described herein are well known. Thus, for the sake of brevity, many conventional implementation details are only briefly mentioned herein or omitted entirely, and details of well-known systems and / or processes are not shown.
[0118] In some embodiments, various functions or acts may be performed at a given location, or in connection with one or more devices or systems, or both. In some embodiments, some given functions or acts may be performed at a first device or location, and remaining functions or acts may be performed at one or more additional devices or locations.
[0119] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used herein, indicate the presence of stated features, integers, steps, operations, elements, or components, or combinations thereof, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, or groups thereof, or combinations thereof.
[0120] The corresponding structure, material, acts, and equivalents of all means-plus-function or step-plus-function elements within the scope of the following claims are intended to include any structure, material, or acts for performing the function in combination with other elements recited in the claims as specifically recited in the claims. This disclosure has been presented for purposes of illustration and description, but is not intended to be exhaustive or limited to the form disclosed. Many modifications and variations will become apparent to those skilled in the art without departing from the scope of the disclosure. The embodiments were chosen and described to best explain the principles and practical application of the disclosure and to enable others skilled in the art to understand the disclosure in terms of various embodiments with various modifications as may be suited to the particular uses contemplated.
[0121] The diagrams shown herein are illustrative. There may be numerous variations to the diagrams or steps (or operations) described therein without departing from the spirit of this disclosure. For example, actions may be performed in a different order, or elements may be added, deleted, or modified. Also, the term "coupled" depicts having a signal path between two elements and does not imply a direct connection between those elements without an intervening element / connection between them. All of these variations are considered to be part of this disclosure.
[0122] The following definitions and abbreviations will be used to interpret the claims and the specification. As used herein, the terms "comprises," "comprising," "includes," "including," "has," "having," "contains," or "containing," or any variations thereof, are intended to cover a non-exclusive inclusion. For example, a composition, mixture, process, method, article, or device that contains a list of elements is not necessarily limited to only those elements, but may include other elements not expressly listed or inherent in such composition, mixture, process, method, article, or device.
[0123] Moreover, the term "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any embodiment or design described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments or designs. The terms "at least one" and "one or more" are understood to include any integer greater than or equal to one, i.e., 1, 2, 3, 4, etc. The term "plurality" is understood to include any integer greater than or equal to two, i.e., 2, 3, 4, 5, etc. The term "connected" can include both an indirect "connected" and a direct "connected."
[0124] The terms "about," "substantially," "approximately," and variations thereof are intended to include the degree of error associated with measuring a particular quantity based on equipment available at the time of filing. For example, "about" can include a range of ±8%, or 5%, or 2% of a given value.
[0125] The present invention may be a system, method, or computer program product, or combination thereof, at any possible level of technical detail integration. The computer program product may include computer-readable storage medium(s) having computer-readable program instructions for causing a processor to perform aspects of the present invention.
[0126] A computer-readable storage medium may be any tangible device capable of holding and storing instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. A non-exhaustive list of more specific examples of computer-readable storage media includes portable floppy disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanically encoded devices such as punch cards or ridge-in-groove structures on which instructions are recorded, and any suitable combination thereof. As used herein, computer-readable storage media should not be construed as being ephemeral signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium (e.g., light pulses passing through fiber optic cable), or electrical signals transmitted over wires.
[0127] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to each computing / processing device or to an external computer or storage device over a network (e.g., the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof). This network may include copper transmission cables, optical transmission fiber, wireless transmission, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. A network adapter card or network interface within each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage on a computer-readable storage medium within each computing / processing device.
[0128] Computer-readable program instructions for carrying out operations of the present invention may be source or object code written in any combination of one or more programming languages, including assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, configuration data for integrated circuits, or object-oriented programming languages such as Smalltalk, C++, and procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer as a standalone software package, partially on the user's computer and on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be to an external computer (e.g., via the Internet using an Internet Service Provider). In some embodiments, to carry out aspects of the present invention, electronic circuitry including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may execute computer readable program instructions to customize the electronic circuitry by utilizing state information of the computer readable program instructions.
[0129] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0130] These computer-readable program instructions may be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to create a machine, where the instructions, executed by the processor of the computer or other programmable data processing apparatus, create means for performing the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams. These computer-readable program instructions may be stored on a computer-readable storage medium and capable of directing a computer, programmable data processing apparatus, or other device, or combination thereof, to function in a particular manner, such that the computer-readable storage medium on which the instructions are stored comprises an article of manufacture containing instructions that implement aspects of the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams.
[0131] Computer-readable program instructions may be loaded into a computer, other programmable data processing apparatus, or other device such that the instructions, which execute on the computer, other programmable apparatus, or other device, perform the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams, thereby causing a series of operable steps to be performed on the computer, other programmable apparatus, or other device to produce a computer-implemented process.
[0132] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowcharts or block diagrams may represent a module, segment, or portion of instructions, which includes one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions shown in the blocks may occur in a different order than that shown in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending on the functionality involved. It should also be noted that each block of the block diagrams and / or flowchart diagrams, and combinations of blocks in the block diagrams and / or flowchart diagrams, may be implemented by a special-purpose hardware-based system that performs the specified functions or operations or executes a combination of special-purpose hardware and computer instructions.
[0133] The description of various embodiments of the present invention has been presented for purposes of illustration and is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will become apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein has been selected to best explain the principles, practical applications, or technical improvements over techniques found in the marketplace of the embodiments, and to enable others skilled in the art to understand the embodiments described herein.
Claims
1. inputting, by a processor, records associated with an information technology (IT) domain into a machine learning model; classifying, by a processor, the labeled records using the machine learning model, wherein the machine learning model refrains from classifying the given record in response to the given record being outside of the IT domain. Computer-implemented methods.
2. The computer-implemented method of claim 1 , wherein the machine learning model identifies the given record outside the IT domain as unclassified.
3. The computer-implemented method of claim 1 , wherein the machine learning model is trained on training data in the IT domain.
4. the machine learning model is trained by receiving an input of training data including training records and training labels corresponding to the training records; The machine learning model is temporarily restrained from classifying any records outside of the IT domain. The computer-implemented method of claim 1 .
5. The computer-implemented method of claim 1 , wherein the record and the label are provided to an automated resolution system configured to modify at least one component within an industrial IT environment.
6. 6. The computer-implemented method of claim 5, wherein the given record that is outside the scope of the IT domain is prevented from being provided to the automated resolution system, thereby avoiding alteration of any component within the IT environment based on an incorrect classification of the given record.
7. the machine learning model comprises a linear classification algorithm; The record is a ticket for a technical problem in the IT environment. The computer-implemented method of claim 1 .
8. a memory having computer readable instructions; and one or more processors for executing the computer-readable instructions, wherein the computer-readable instructions control the one or more processors to: inputting records associated with an information technology (IT) domain into a machine learning model; classifying the labeled records using the machine learning model, wherein the machine learning model refrains from classifying the given record in response to the given record being outside of the IT domain. system.
9. The system of claim 8 , wherein the machine learning model identifies the given record outside the IT domain as unclassified.
10. The system of claim 8 , wherein the machine learning model is trained on training data within the IT domain.
11. the machine learning model is trained by receiving an input of training data including training records and training labels corresponding to the training records; The machine learning model is temporarily restrained from classifying any records outside of the IT domain. The system of claim 8.
12. The system of claim 8 , wherein the record and the label are provided to an automated resolution system configured to modify at least one component within an industrial IT environment.
13. 13. The system of claim 12, wherein the given record that is outside the scope of the IT domain is prevented from being provided to the automated resolution system, thereby avoiding alteration of any component within the IT environment based on an incorrect classification of the given record.
14. the machine learning model comprises a linear classification algorithm; The record is a ticket for a technical problem in the IT environment. The system of claim 8.
15. A computer program product including a computer-readable storage medium having program instructions embodied therein, said program instructions being operable by one or more processors to: inputting records associated with an information technology (IT) domain into a machine learning model; and classifying labeled records using the machine learning model, wherein the machine learning model refrains from classifying a given record that is outside of the IT domain in response to the given record being outside of the IT domain. computer Program products.
16. 16. The computer program product of claim 15, wherein the machine learning model identifies the given record outside the IT domain as unclassified.
17. 16. The computer program product of claim 15, wherein the machine learning model is trained on training data in the IT domain.
18. the machine learning model is trained by receiving an input of training data including training records and training labels corresponding to the training records; The machine learning model is temporarily restrained from classifying any records outside of the IT domain.
16. A computer program product according to claim 15.
19. 16. The computer program product of claim 15, wherein the record and the label are provided to an automated resolution system configured to modify at least one component within an industrial IT environment.
20. 20. The computer program product of claim 19, wherein the given record that is outside the scope of the IT domain is prevented from being provided to the automated resolution system, thereby avoiding alteration of any component within the IT environment based on an incorrect classification of the given record.
21. the machine learning model comprises a linear classification algorithm; The record is a ticket for a technical problem in the IT environment.
16. A computer program product according to claim 15.
22. inputting records associated with an information technology (IT) domain into a linear classification algorithm by a processor; classifying, by a processor, the labeled records using the linear classification algorithm, wherein the linear classification algorithm refrains from classifying a given record in response to the given record being outside the range of the IT domain. Computer-implemented methods.
23. 23. The computer-implemented method of claim 22, wherein the machine learning model identifies the given record outside the IT domain as unclassified.
24. 23. The computer-implemented method of claim 22, wherein the linear classification algorithm is trained on training data in the IT domain, the linear classification algorithm is further trained on positively associated features for a label without negatively associated features, and the positively associated features for the label are validated.
25. a memory having computer readable instructions; and one or more processors for executing the computer-readable instructions, wherein the computer-readable instructions control the one or more processors to: inputting records associated with an information technology (IT) domain into a linear classification algorithm by a processor; classifying, by a processor, the labeled records using the linear classification algorithm, wherein the linear classification algorithm refrains from classifying a given record in response to the given record being outside the range of the IT domain. system.
Citation Information
Patent Citations
System analysis system and system analysis method
JP2018180759A
Diagnostic device, diagnostic method, program, and recording medium
JP2019101495A
Real-time motion feedback for extended reality
JP2020091836A
Predictive resolutions for tickets using semi-supervised machine learning
US11556843B2
Methods and systems for multi-resource outage detection for a system of networked computing devices and root cause identification
US20220107858A1