Self-controlled and explainable classification using client-independent machine learning models

A client-independent machine learning model for IT ticket classification addresses the challenge of misclassification and lack of explainability by focusing on IT domain issues and providing transparent decisions, improving IT environment management.

JP7850266B2Active Publication Date: 2026-04-22KYNDRYL INC
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
KYNDRYL INC
Filing Date
2023-05-23
Publication Date
2026-04-22

AI Technical Summary

Technical Problem

Existing IT ticket classification systems struggle to handle technical data effectively and fail to provide explainable decision-making, often misclassifying non-IT domain issues and lacking self-restraint in classification.

Method used

A client-independent machine learning model is trained to classify IT-related issues within its domain, refraining from misclassifying non-IT domain tickets and providing explainable decisions using disjunctive normal form and positively related features.

Benefits of technology

The model accurately classifies IT-related issues while preventing misclassification, enhancing decision transparency and reducing incorrect automated actions in IT environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007850266000001
    Figure 0007850266000001
  • Figure 0007850266000002
    Figure 0007850266000002
  • Figure 0007850266000003
    Figure 0007850266000003
Patent Text Reader

Abstract

An embodiment relates to providing explainable classification with restraint using a client-independent machine learning model. The technique includes classifying a record with a label using a machine learning model by a processor, the machine learning model restraining itself from classifying the given record in response to the given record being outside of an information technology (IT) domain. The processor generates an explanation of the machine learning model's decision to classify the record with the label and displays the explanation in a human-readable format. [Representative diagram] Figure 12
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to computer systems, and more specifically to a computer-implemented method, computer system, and computer program product configured and arranged to provide explainable classification with abstention using a client-agnostic machine learning model.

Background Art

[0002] An information technology (IT) ticket issuance system is a tool used to track requests, events, incidents, and alerts for IT services that may require additional action from the IT department. Ticket issuance software makes it possible to organize the IT issues within them to be resolved by rationalizing the resolution process. The elements they manage, so-called tickets, provide context regarding the issues, including details, categories, and any relevant tags.

[0003] This ticket often includes additional context details and also the relevant contact information of the individual who created the ticket. Tickets are usually generated by employees, but automatic tickets may also be created when a specific incident occurs and is flagged. When a ticket is created, it is assigned to an IT agent who is expected to resolve it. An effective ticket issuance system enables tickets to be submitted via various methods. These include submissions via virtual agents, phone, email, service portals, live agents, walk-up experiences, and the like.

[0004] Generally, automated systems automate environmental aspects and problem solving, event monitoring software monitors components and the environment, and incidents are reported via tickets through a ticketing system. A typical system might monitor tickets using natural language and output what the problem is via a general language classifier. Unfortunately, while general language classifiers handle general text well, they do not handle tickets containing technical data well and do not explain how the decisions were made. What is needed is a system that can analyze technical issues, classify those technical issues, detect these issues, or combine them, and this system does not require any action. [Overview of the Initiative]

[0005] Embodiments of the present invention relate to a computer implementation method for providing self-restrained, explainable classifications using a client-independent machine learning model. A non-restrictive computer implementation method includes a processor classifying records together with labels using a machine learning model, the machine learning model including self-restraint in classifying given records in response to records outside the scope of the information technology (IT) domain. The computer implementation method includes the processor generating an explanation of the decisions made by the machine learning model for classifying the records together with labels. The computer implementation method includes displaying the explanation in a human-readable format.

[0006] Another embodiment of the present invention implements the features of the above method in a computer system and a computer program product.

[0007] Additional technical features and benefits are realized through the methods of the present invention. Embodiments and aspects of the present invention are described in detail herein and are considered to be part of the subject matter of the claims. For a better understanding, please refer to the detailed description and drawings.

[0008] Details of the exclusive rights described herein are specifically pointed out and clearly stated in the claims at the end of this specification. The aforementioned and other features, as well as the advantages of embodiments of the present invention, are evident from the following detailed description in conjunction with the accompanying drawings. [Brief explanation of the drawing]

[0009] [Figure 1] This is a block diagram of an example of a computer system for use with one or more embodiments of the present invention. [Figure 2] This is a block diagram of an example of a system configured to provide self-controlled, explainable classifications using a client-independent machine learning model, according to one or more embodiments of the present invention. [Figure 3] This is a flowchart of a computer implementation method for training a machine learning model for each classification label, according to one or more embodiments of the present invention. [Figure 4A] This is a block diagram illustrating an example of a confusion matrix for a machine learning model according to one or more embodiments of the present invention. [Figure 4B] This is a block diagram illustrating an example of a coefficient matrix for a machine learning model according to one or more embodiments of the present invention. [Figure 5A] This is an example of a chart illustrating regions of positively related features that are determined to positively contribute to each of the predicted labels during classification by a machine learning model according to one or more embodiments of the present invention. [Figure 5B] This is an example of a chart illustrating how a particular feature contributes to the determination of a class label according to one or more embodiments of the present invention. [Figure 5C] This graph illustrates the contribution of features analyzed by a machine learning model for determining classification, according to one or more embodiments of the present invention. [Figure 6] This is a flowchart of a computer implementation method for calculating positively related features for classification, according to one or more embodiments of the present invention. [Figure 7] This is a flowchart of a computer implementation method for explaining machine learning decisions using disjunctive normal forms, according to one or more embodiments of the present invention. [Figure 8] This is a block diagram illustrating the conversion of a linear classification formula into a disjunctive standard form according to one or more embodiments of the present invention. [Figure 9] This is a flowchart of a computer implementation method that explains machine learning decisions using positively related features, according to one or more embodiments of the present invention. [Figure 10] This block diagram illustrates, according to one or more embodiments of the present invention, the selection of features (e.g., tokens) from ticket data and the presentation of positively related features that contribute to class labels for tickets. [Figure 11] This is a flowchart of a computer implementation method that provides a self-controlled, explainable classification using a client-independent machine learning model, according to one or more embodiments of the present invention. [Figure 12] This is a flowchart of a computer implementation method that provides a self-controlled, explainable classification using a client-independent machine learning model, according to one or more embodiments of the present invention. [Figure 13] This figure shows a cloud computing environment according to one or more embodiments of the present invention. [Figure 14] This figure shows an extraction model layer according to one or more embodiments of the present invention. [Modes for carrying out the invention]

[0010] One or more embodiments provide an explainable classification with restraint using a client-independent machine learning model. For all sets of tickets present in a client's environment, the client-independent machine learning model is configured to determine the next discretionary action to resolve a computer / network issue. The client-independent machine learning model is trained with tickets and solutions from a variety of different clients to become client-independent or industry-independent, including clients across different industries such as cybersecurity, finance, government, manufacturing, education, retail, cloud computing, and data storage. According to one or more embodiments, the client-independent machine learning model is trained to know when to classify tickets and when not to classify them when analyzing tickets. This can be achieved by building an independent machine learning model that only classifies within its scope of understanding, such as an information technology (IT) domain or IT environment, but restrains its classification in areas it does not understand. Furthermore, one or more embodiments are configured to explain how the independent machine learning model arrived at a particular decision, which gives the user confidence in the independent machine learning model's decisions.

[0011] Incident identification and automated resolution are processes for managing IT service disruptions and restoring services. For example, a monitoring system monitors a client's IT environment in a given industry. The term "IT environment" refers to the infrastructure, hardware, software, and systems that a client (entity or business) relies on daily in the process of using information technology. Some of the resources commonly used in an IT environment include computers, internet access, and peripheral devices. Examples of an IT environment include hardware such as routers, personal computers, servers, switches, and data centers; software such as user applications, web servers, and applications that enable and make hardware connections available; and networks, i.e., firewalls, cables, and other components that facilitate internal and external communications in a business. When a technical event is detected in the IT environment, and / or when a user of the IT environment requests it, the monitoring system generates a ticket. The ticket is sent to the automated resolution system and / or the IT department for resolution. A ticket is a dedicated document or record that represents an incident, alert, request, and / or event that requires action from the IT department. Furthermore, tickets are historical documents that detail service events such as incidents, problems, and / or service requests. Tickets govern and control how service events are handled.

[0012] A typical system might attempt to monitor the environment and identify problems. However, while general language classifiers handle general text well, they don't handle tickets containing technical data well, nor do they explain how the decisions were made. Furthermore, classifiers are poor at exercising self-restraint in their classifications. For example, a ticket stating "The first humans went to space in 1961" should not be classified as a technical problem / issue, but many classifiers will analyze the ticket, detect the word "space," and mistakenly classify it as a disk handler problem.

[0013] Technical solutions and benefits include systems that, according to one or more embodiments, provide a client / industry-independent model in an IT environment. Thus, thousands of different independent machine learning models are not necessary for thousands of different clients or industries, but the independent machine learning models operate across various clients of different industries. In one or more embodiments, the independent machine learning model is configured to refrain from classifying inputs external to the IT environment or IT domain. Thereby, the independent machine learning model (e.g., classifier) can avoid misclassifying tickets along with labels against automatic resolution by an automatic resolution system when the ticket (as input) is not actually supposed to generate a label because it is not in the IT environment or IT domain. As yet another technical solution and benefit, it may be possible to provide a machine learning-based explanation for decision-making to the independent machine learning model using disjunctive normal form (DNF) and / or positively related features. One or more embodiments provide dimensionality reduction using a gradient threshold reduction method and extract positive correlation values as a list of all features that affect the independent machine learning model (e.g., classifier). Some embodiments may not have these potential benefits or advantages, and these potential benefits or advantages are not necessarily required for all embodiments.

[0014] One or more embodiments described herein can utilize machine learning techniques to perform tasks such as classifying desired features. More specifically, one or more embodiments described herein can combine rule-based decision-making and artificial intelligence (AI) reasoning to achieve the various operations described herein, namely the classification of desired features. The term “machine learning” broadly describes the function of electronic systems that learn from data. A machine learning system, engine, or module may include a trainable machine learning algorithm that can be trained, for example, in an external cloud environment, to learn functional relationships between inputs and outputs, and the resulting model (sometimes called a “trained neural network,” “trained model,” “trained classifier,” or “trained machine learning model,” or a combination thereof) can be used, for example, to classify desired features.

[0015] Returning to FIG. 1, computer system 100 is generally shown in accordance with one or more embodiments of the present invention. Computer system 100 can be an electronic computer framework that includes, employs, or both includes and employs any number and combination of computing devices and networks using various communication technologies, as described herein. Computer system 100 can be easily extended, expanded, and modularized, and can be changed to different services or reconfigured some features independently of each other. Computer system 100 can be, for example, a server, a desktop computer, a laptop computer, a tablet computer, or a smartphone. In some examples, computer system 100 can be a cloud computing node. Computer system 100 may be described in the general context of computer system-executable instructions, such as program modules, executed by a computer system. Generally, program modules may include routines, programs, objects, components, logic, data structures, etc. that perform specific tasks or implement specific extracted data types. Computer system 100 may be executed in a distributed cloud computing environment, in which tasks are performed by remote processing devices connected through a communication network. In a distributed cloud computing environment, program modules may be located on both local and remote computer system storage media including memory storage devices.

[0016] As shown in Figure 1, the computer system 100 has one or more central processing units (CPUs) 101a, 101b, 101c, etc. (collectively referred to as the processor 101). The processor 101 can be a single-core processor, a multi-core processor, a computing cluster, or any number of any configuration. The processor 101, also referred to as the processing circuit, is coupled to the system memory 103 and various other components via the system bus 102. The system memory 103 may include read-only memory (ROM) 104 and random-access memory (RAM) 105. The ROM 104 is coupled to the system bus 102 and may include a basic input / output system (BIOS) or its successor, such as a unified extensible firmware interface (UEFI), which controls certain basic functions of the computer system 100. The RAM is a write-read memory coupled to the system bus 102 for use by the processor 101. System memory 103 provides temporary memory space for the operation of the above instructions during the operation. System memory 103 may include random access memory (RAM), read-only memory, flash memory, or any other suitable memory system.

[0017] The computer system 100 includes an input / output (I / O) adapter 106 and a communication adapter 107 coupled to a system bus 102. The I / O adapter 106 may be a small computer system interface (SCSI) adapter that communicates with a hard disk 108 or any other similar component or both. The I / O adapter 106 and the hard disk 108 are collectively referred to as mass storage 110 in this specification.

[0018] Software 111 to run on computer system 100 may be stored in mass storage 110. Mass storage 110 is an example of a tangible storage medium readable by processor 101, in which software 111 is stored as instructions for execution by processor 101 to cause computer system 100 to operate, for example, as described below herein with reference to various figures. Examples of computer program products and the execution of such instructions are discussed in further detail herein. A communication adapter 107 interconnects system bus 102 to network 112, which may be an external network, thereby enabling computer system 100 to communicate with other such systems. In one embodiment, the system memory 103 and the mass storage 110 jointly store an operating system, which may be any suitable operating system that coordinates the functions of the various components shown in Figure 1.

[0019] Additional input / output devices are shown to be connected to the system bus 102 via a display adapter 115 and an interface adapter 116. In one embodiment, adapters 106, 107, 115, and 116 may be connected to one or more I / O buses connected to the system bus 102 via an intermediate bus bridge (not shown). A display 119 (e.g., a screen or display monitor) is connected to the system bus 102 by a display adapter 115, which may include a graphics controller and a video controller for enhancing the performance of graphics-intensive applications. A keyboard 121, mouse 122, speaker 123, microphone 124, etc., can be interconnected to the system bus 102 via an interface adapter 116, which may include, for example, a super I / O chip integrated multi-device adapter within a single integrated circuit. A suitable I / O bus for connecting peripheral devices such as hard disk controllers, network adapters, and graphics adapters typically includes common protocols, such as Peripheral Component Interconnect (PCI) and Peripheral Component Interconnect Express (PCIe). Thus, as configured in Figure 1, the computer system 100 includes processing capabilities in the form of a processor 101, storage capabilities including system memory 103 and mass storage 110, input means such as a keyboard 121, a mouse 122, and a microphone 124, and output capabilities including a speaker 123 and a display 119.

[0020] In some embodiments, the communication adapter 107 can transmit data using any suitable interface or protocol, in particular the Internet Small Computer System Interface. The network 112 may be a cellular network, a radio network, a wide area network (WAN), a local area network (LAN), or the Internet. An external computing device may be connected to the computer system 100 through the network 112. In some examples, the external computing device may be an external web server or a cloud computing node.

[0021] It should be understood that the block diagram in Figure 1 is not intended to indicate that computer system 100 includes all the components shown in Figure 1. Rather, computer system 100 may include any suitable, fewer, or additional components not shown in Figure 1 (e.g., additional memory components, embedded controllers, modules, additional network interfaces, etc.). Furthermore, the embodiments described herein with reference to computer system 100 may be implemented with any suitable logic, which may include any suitable hardware (e.g., among processors, embedded controllers, application-specific integrated circuits), software (e.g., among applications), firmware, or any suitable combination of hardware, software, and firmware in various embodiments, as referenced herein.

[0022] Figure 2 illustrates a block diagram of an example of a system 200 configured to provide self-controlled, explainable classifications using a client-independent machine learning model, in one or more embodiments. System 200 includes a computer system 202 configured to communicate with numerous different computer systems, for example, across a network 250, such as computer system 240A for managing the IT environment for a client in one industry, computer system 240B for managing the IT environment for another client in a different industry, and computer system 240N for managing the IT environment for yet another client in a different industry. Computer systems 240A and 240B through 240N may generally be referred to as computer system 240. Each computer system 240 has its own IT management system 244 for monitoring the IT environment for each client in its respective industry and storing its respective tickets and solutions in a ticket repertoire 246. The ticket repertoire 246 is operable to store numerous tickets and their respective solutions for the IT environment of computer system 240. Network 250 can be a wired or wireless communication network.

[0023] The IT management system 244 may include or represent a monitoring ticketing system and an automated resolution system for each client in the industry. By communicating with the computer system 240 over a network 250 which can be a wired or wireless communication network, the software application 204 is configured to extract various tickets and their respective solutions from a ticket repertoire 246 from different clients in different industries.

[0024] As shown by the dotted lines, in one or more embodiments, the computer system 202 may include an IT management system 244 and its ticket repertoire 246 for one or more computer systems 240A to 240N in each IT environment of each client. The computer system 202 can manage the client's IT environment for one or more computer systems 240A to 240N. Any portion of the system 200 including the computer system 202 and one or more computer systems 240A to 240N may be a portion of a cloud computing environment 50, as further considered herein (illustrated in Figure 13).

[0025] In System 200, computer systems 202, computer systems 240A-240N, IT management system 244, software application 204, training data 206, machine learning model 220, automatic resolution system 222, rule generation algorithm 224, etc., may include, use, or both, any of the considered functions within computer system 100, where computer system 100 includes various hardware components and various software applications such as software 111 that can be executed as instructions on one or more processors 101 to carry out actions according to one or more embodiments of the present invention. Software application 204 may include, incorporate, call, or combine other various software, algorithms, application programming interfaces (APIs), etc., that operate as considered herein. Software application 204 may represent a number of software applications.

[0026] The tickets and their respective solutions are stored in a repertoire, such as storage, as training data 206. A software application 204 filters the training data 206 to ensure that it exists only within an IT environment, which may also be called an IT domain or IT space. An IT domain encompasses the IT environments of clients in each industry. Any tickets unrelated to events existing within an IT domain (e.g., errors, issues, security breaches, faulty computer equipment, etc.) are removed from the training data 206.

[0027] The computer system 202 includes a machine learning model 220, which is a client-independent machine learning model trained to classify tickets within an IT environment. In one or more embodiments, this client-independent machine learning model is trained, for example, solely to classify tickets within an IT environment. The machine learning model 220 may represent a number of machine learning models 220. The machine learning model 220 classifies tickets by predicting labels that identify how to solve the computer problems associated with the tickets. The tickets and their predicted labels are sent to the automated resolution system 222, which can automatically solve the computer problems of the tickets according to the predicted label output obtained from the machine learning model 220. In one or more embodiments, the machine learning model 220 is a linear classifier and processes a linear classification algorithm. Terms such as label, class, classification, classification label, and class label may be used interchangeably to refer to categories of machine learning.

[0028] This linear classification algorithm uses the features of an object, such as the features of a ticket, to determine which class (or group) it belongs to. A linear classifier achieves this by making a classification decision based on the value of a linear combination of features. The features of an object are also known as feature values ​​and are typically presented in a machine in a vector called a feature vector. Examples of linear classification algorithms and techniques include the simple Bayesian algorithm, linear discriminant analysis (LDA) algorithm, least squares algorithm, support vector machine algorithm, ridge regression algorithm, lasso algorithm, elastic net algorithm, minimum angle regression algorithm, orthogonal matching tracking algorithm, Bayesian regression algorithm, logistic regression algorithm, linear regression algorithm, perceptron algorithm, and passive attack classification algorithms, which are familiar to those skilled in the art.

[0029] The machine learning model 220 may be configured to have a trained linear classification algorithm for each classification label for a ticket. Furthermore, the machine learning model 220 may be configured to refrain from classifying tickets that do not belong to the IT environment or IT domain. In one or more embodiments, there may be a classification label labeled as Unclassified / Unknown, and the machine learning model 220 may be configured to use the Unclassified / Unknown label to indicate that the ticket's feature vector (i.e., features) does not apply to the IT environment or IT domain. By refraining from classifying tickets that originate from, are not related to, or both originate from the IT environment as Unclassified / Unknown, or by classifying such tickets, or both, the machine learning model 220 is configured to prevent misclassified tickets from being mistakenly sent to the automated resolution system 222 and consequently from being subjected to automated corrective actions on the IT environment by the automated resolution system 222. In one or more embodiments, the machine learning model 220 can prevent misclassified tickets from being mistakenly sent to the automated resolution system 222 and consequently from being subjected to incomplete automated corrective actions on the IT environment. For example, one or more software or hardware components, or both, within an IT environment may be automatically changed by an automated resolution system based on incorrect classification labels on tickets, thereby causing malfunctions in the software or hardware components, or both, of the computer systems within the IT environment.

[0030] Figure 3 is a flowchart of a computer implementation method 300 for training a machine learning model 220 having a linear classification algorithm for each classification label, according to one or more embodiments, and thereby producing a trained machine learning model 220. The computer implementation method 300 is performed by a computer system 202.

[0031] In block 302 of the computer implementation method 300, the software application 204 is configured to read the training data by compiling the data from each ticket into the training data stored in the training data 206. The software application 204 may also be configured to parse the training data 206 to determine and filter out any tickets along with their solutions that are not in the IT environment. This leaves only tickets that are not in the IT environment and represent the IT domain, and as a result, the machine learning model 220 is trained and learns about the tickets and their individual solutions that are in the IT domain. In one or more embodiments, for example, the machine learning model 220 is trained and learns only about the tickets and their individual solutions that are in the IT domain. When it detects a ticket that is not in the IT domain, the machine learning model 220 refrains from labeling the ticket, and / or the ticket may be labeled as unclassified / unknown, and as a result, the ticket is prevented / blocked from being processed by the automated resolution system 222.

[0032] The training data 206 can be refined using cross-fold validation. Tickets are labeled in preparation for training the machine learning model 220. These labels can be automated playbooks obtained by matching automated executions with addressed incident tickets. For example, the machine learning model 220 may be trained on tickets resolved by an automation platform (e.g., IT management system 244), such as the Red Hat® Ansible® platform, and also on tickets not resolved by the automation platform. For example, when the automation platform (e.g., IT management system 244) receives a ticket, it logs into the system, and if it has a playbook, the automation platform executes the playbook, resolves the ticket, and closes the ticket. A closed ticket is considered factual or completed. The tickets are used to train the machine learning model 220 to one of the known class examples within the IT domain, such as application down, database space issue, disk handler, network connectivity, file system mount handler, high disk space usage handler, high memory and page file usage, host down handler, service handler, and job abend. Naturally, this illustrative list of classification labels is not meant to be exhaustive.

[0033] In block 304, the software application 204 is configured to train the machine learning model 220 using the training data 206. For example, the data for each ticket (along with its corresponding label) is input to the machine learning model 220 as a feature vector so that the linear classifier algorithm of the machine learning model 220 learns how to classify the input data of tickets. The training data is labeled, which means that the tickets are pre-labeled to determine the time it takes for the output of the machine learning model 220 to predict the correct label. During the training phase, the predicted labels of the tickets from the machine learning model 220 are compared with the labels of the tickets in the training data 206 to continuously improve the machine learning model 220. This allows the machine learning model 220 to learn the correct classification label for each ticket.

[0034] In block 306, the machine learning model 220 is configured to classify the ticket input data along with descriptions. For example, each ticket is classified based on the foundation and / or decisions made by the machine learning model 220. The machine learning model 220 is configured to generate descriptions for the user with respect to the machine learning rules and / or as positively related features on which the machine learning model 220 bases its decisions on the predicted labels for the input tickets. In one or more embodiments, the machine learning model 220 may include and / or employ a rule generation algorithm 224 for the machine learning rules and / or identified positively related and negatively related features, further details of which are discussed herein.

[0035] In block 308, the software application 204 is configured to validate the classification results by comparing the predicted classification labels of tickets with the labels of tickets in the training data. In block 310, the software application 204 is configured to check whether the predicted classification labels of tickets match the labels of tickets in the training data.

[0036] In block 312, if (Yes) a match is found, the software application 204 is configured to check whether the machine learning rule and / or the identified positively related feature description is satisfactory. This may involve soliciting input from a subject area expert and / or employing a natural language processing (NLP) system. If (Yes) the description is satisfactory, the software application 204 is configured to terminate training the machine learning model.

[0037] In block 314, if the classification result in (No) block 310 is not satisfactory for the decision, the software application 204 is configured to improve the training data and / or fine-tune the classifier algorithm. Also, if the explanation in (No) block 312 is not satisfactory for the decision, the software application 204 is configured to improve the explanation.

[0038] It should be noted that a separate linear classification algorithm is trained for each classification label. In one or more embodiments, the machine learning model 220 may include a number of linear classification algorithms, one for each classification label corresponding to an IT domain. In one or more embodiments, there may be a number of machine learning models 220, one for each classification label corresponding to an IT domain. Any consideration of a single linear classification algorithm and / or a single machine learning model 220 is more similar to all linear classification algorithms and / or all machine learning models 220 corresponding to all classification labels of an IT domain.

[0039] As described herein, when an input ticket labels a typical classifier as "The first humans went into space in 1961," that typical classifier would attempt to classify a labeled ticket, such as a disk or disk handler. However, such a label is misclassified in this example, which could lead to incorrect actions being taken by the automated resolution system.

[0040] As a technical benefit and solution, one or more embodiments are configured to refrain from classifying such tickets that indicate "The first time humans went into space was in 1961," because the machine learning model 220 is trained to refrain from classifying such tickets. Rather, the machine learning model 220 may output unknown / unclassified, thereby preventing the automated resolution system 222 from modifying one or more software and / or hardware components in the IT environment of a client in a certain industry. Thus, based on the output from the machine learning model 220, the software application 204 can recognize that the ticket is unknown / unclassified in the IT domain and send the ticket to a dedicated IT department for resolution instead of sending it to the automated resolution system 222. The software application 204 is configured to send tickets that have been properly labeled by the machine learning model 220 to the automated resolution system 222 for automated processing in the IT environment of a client in a certain industry. According to the label from the machine learning model 220, the automated resolution system 222 is configured to modify software components, hardware components, and / or both software and hardware components of one or more computer systems in an IT environment, thereby improving the computer system itself. The modification of software and / or hardware components is a practical application related to the use of the machine learning model 220, solving technical computer problems on computer systems in an IT environment.

[0041] One or more embodiments provide a method for computing positively related features within tickets for text classification. Positively related features are tokens within the ticket. Tokens can refer to one or more words, phrases, sentences, etc., in the text of the ticket and in the process referred to as tokenization. Tokens can be used as features in the ticket's feature vector. One or more embodiments extract a list of all positively related features for all (IT) tickets, along with their labels, which can be used to further train a machine learning model 220 and for use in the descriptions considered herein.

[0042] Furthermore, one or more embodiments are configured to generate a linear classifier for the machine learning model 220, generate coefficient and confusion matrices for insights into how the linear classifier works at the corpus level, extract positive association values ​​for all IT tickets using gradient descent iterative thresholding, have the positive association values ​​reviewed by subject matter experts / IT domain experts to find "true" positive association values, extract rules and rule features as new training data aligns with a given label, and create a full-fledged trained model and a list of acceptable positive association value regions. Thus, one or more embodiments can receive incoming new incident tickets and extract positive association values ​​for a given ticket, and if no positive association values ​​are identified, the linear classifier restrains classifying the ticket.

[0043] Referring to Figure 4A, the block diagram shows an example of a confusion matrix for a machine learning model according to one or more embodiments. Figure 4B illustrates a block diagram of an example of a coefficient matrix for a machine learning model according to one or more embodiments. The ticket and label features of the confusion matrix in Figure 4A are projected into the example coefficient matrix illustrated in Figure 4B.

[0044] To learn from errors, by exploring the machine learning model 220, the software application 204 can generate an example of a confusion matrix that helps understand how the classifier gets confused when learning and classifying the training data in Figure 4A. The confusion matrix is ​​an N×N matrix used to evaluate the performance of the classification model, where N is the number of target classes. This matrix compares the actual target values ​​in the training data 206 with the values ​​predicted by the machine learning model.

[0045] Referring to Figure 4B, the coefficient matrix contains entries representing the signed coefficients of each feature for the hyperplane in a two-way linear classifier for each class. Software application 204 is configured to identify the feature with the largest absolute coefficient, e.g., the feature with the largest absolute coefficient value. Using the coefficient matrix, this provides software application 204 with the features that have the greatest influence on classification on a global scale. The coefficient matrix is ​​sometimes called a correlation matrix and is a table that presents correlation coefficients for different variables. The coefficient matrix illustrates the correlations between pairs of all possible values ​​in the table. It summarizes large datasets and identifies and visualizes patterns in given data.

[0046] In Figure 4B, features with a positive number for a given classification label indicate that the corresponding feature positively contributes to the machine learning model 220's decision to classify to the given classification label. On the other hand, features with a negative number (i.e., a negative sign) for a given class label indicate that the corresponding feature does not contribute to the classification to the given class label (i.e., a negative contribution).

[0047] Based on all features in the tickets for all class labels (e.g., in the training data 206), the software application 204 using the machine learning model 220 generates a region of positively associated features for each predicted classification label, as illustrated in Figure 5A. In Figure 5A, the example chart shows positively associated features, also called positive association values, which are determined to positively contribute to each predicted label during classification by the machine learning model 220. For example, the positively associated features "low space," "disk:c handler," and "file system" are positively associated features that positively contribute to the predicted classification label "disk handler" in the classification by the machine learning model 220. This example of the classification label disk handler is used for illustrative purposes and is not limited to various scenario examples. Naturally, the embodiments are not limited to the classification label disk handler.

[0048] Figure 5B is a chart illustrating how an example of a feature “space” contributes to the determination of a classification label. As seen in Figure 5B, the feature space has a positive coefficient for the classification label disk handler, meaning that the feature space positively contributes to the determination of the machine learning model 220 to output the classification label disk handler. Similarly, the feature space has negative coefficients for several other classification labels, meaning that the feature space negatively contributes to or does not influence the determination of the machine learning model 220 to output the corresponding class labels. Therefore, the feature space can be identified as a negatively related feature and removed as a feature for the corresponding classification label, in which case the space has a negative coefficient. The software application 204 continues the feature removal because, as illustrated in Figure 5A, any identified feature has a negatively related feature with a negative coefficient for a given classification label, thereby leaving only positively related features available in the region of positively related features for each classification label.

[0049] Continuing with the example scenario for the classification label disk handler, Figure 5C is a graph illustrating the contributions of features analyzed by the machine learning model 220 to classify the disk handler, and these features are now validated by a human subject area expert. For each feature, the graph shows its negative and positive contributions to the decision by the machine learning model 220 for classifying the ticket together with the label disk handler. To validate by the subject area expert that the positively associated features are, for example, "low space," "disk c: handler," and "file system" as positively contributing to the decision of the machine learning model 220, the software application 204 identifies positively associated features for a given class. As identified by the software application 204, the subject area expert confirms that all positively associated features are true positively associated features for their respective classification labels, resulting in a region of positively associated features for those labels, as illustrated in Figure 5A. In one or more embodiments, the region of positively associated features excludes any negatively associated features. Therefore, the positively relevant features for each of those classification labels are collected and added to the training data 206 as a new training dataset to further train the machine learning model 220 to classify tickets together with the classification labels. In one or more embodiments, the positively relevant features are verified to be the positively relevant features for each of those classification labels. This additional training further refines the machine learning model 220's ability to learn to restrain itself from classifying tickets that are not in the IT domain, thereby improving the accuracy of the machine learning model 220.

[0050] During the inference phase, when the machine learning model 220 receives a ticket and outputs its classification label, the software application 204 can explore the machine learning model 220 to obtain positively relevant features for any given ticket. Furthermore, for incoming tickets, the machine learning model 220 learns to extract and recognize positively relevant features for a given ticket, and is configured to refrain from classifying the ticket if no positively relevant features are found in the ticket. On the other hand, when positively relevant features are recognized in the ticket, the machine learning model 220 is configured to classify the ticket, highlight the positively relevant features to the user on the display (on display 119), and explain the machine learning model's decision by extracting classifier rules using disjunctive normal form.

[0051] Figure 6 is a flowchart of a computer implementation method 600 for calculating positively relevant features for text classification, according to one or more embodiments. In one or more embodiments, a software application 204 employs, utilizes, and / or integrates with a machine learning model 220 to perform the computer implementation method 600. Furthermore, the software application 204 may be used to explore the machine learning model 220 and perform the computer implementation method 600.

[0052] In blocks 602, 604, and 606, the software application 204 is configured to input the text of an incident ticket to a preprocessor in order to generate a feature vector from the input text. The preprocessor extracts the input features from the text, and the features are formed into a feature vector. Known methods can be used to convert the text into a feature vector. The feature vector is input to the machine learning model 220, which outputs a classification label for the corresponding ticket.

[0053] In block 608, the software application 204 is configured to construct a confusion matrix and generate a coefficient matrix for each label output using the machine learning model 220. Examples of the confusion matrix and coefficient matrix are illustrated in Figures 4A and 4B, respectively.

[0054] In block 610, software application 204 is configured to supply feature vectors and coefficient matrices to L1 (and / or L2) regularization problems that model positive associations. An example of a regression model using an L1 regularization method is called least absolute shrinkage and selection operator (Lasso) regression, and an example of a regression model using an L2 regularization method is called ridge regression. For L1 regularization, Lasso regression adds the "absolute magnitude" of the coefficients as a penalty term to the loss function. L1 regularization can be an option when there are many features to give a sparse solution. For L2 regularization, ridge regression adds the "square of the magnitude" of the coefficients as a penalty term to the loss function. L2 regression can be used to estimate the significance of predictors and may be based on applying a penalty to non-significant predictors. Furthermore, elastic networks are a combination of L1 and L2 regularization, and this combination becomes an elastic network method that adds hyperparameters.

[0055] In block 612, software application 204 is configured to apply an iterative reduction / thresholding algorithm (ISTA) to an L1 (and / or L2) control problem. Due to its simplicity, ISTA is widely used when solving linear inverse problems. ISTA may include bidirectional threshold reduction of gradient descent.

[0056] In blocks 614 and 616, the software application 204 is configured to generate sparse feature vectors using the ISTA method and to select positively relevant features from these sparse feature vectors. In one or more embodiments, the sparse feature vectors for classification labels have fewer features than the original feature vector. Therefore, using all the sparse feature vectors generated for each classification label in the IT domain, the software application 204 selects features for each classification label to generate a list of positively relevant features for each class label, and these positively relevant features are output in block 618. As described above, Figure 5A represents the area of ​​positively relevant features for each classification label, such as an example of a class label disk handler.

[0057] To further illustrate the calculation of positively relevant features for text classification, the following is an example of a scenario for illustrative purposes, and is not limiting. We use the ISTA algorithm to compute a sparse solution for inversely reversing a linear problem. A typical example of an inverse linear problem is linear regression. An example for consideration is a classification problem, e.g., a text classification problem. Any linear classification algorithm can be used for a text classification problem in the IT ticket management domain. Here, consider a ticket T classified into class C (e.g., "disk handler") using the linear classification algorithm of machine learning model 220. It is often useful to show "evidence" of the inner workings of the classifier (e.g., machine learning model 220) and "explain" why ticket T was classified into class C. Positively relevant values ​​in a ticket like T are a small subset of features of T that are involved in its classification into class C. Such a set of features provides a good explanation of the inner workings of the classifier by one or more embodiments. As discussed in the example above, this disclosure systematically describes the problem of finding positively relevant features for a text classification problem as a sparse inverse linear problem. One or more embodiments customize and simplify the ISTA algorithm to make it more efficient for this use case. This customization involves formulating a specific problem, such as a formulation used internally by a passive-attack classifier (PAC), and using its structural properties to efficiently perform the iterative thresholding step.

[0058] One or more embodiments provide explainability for decisions of a machine learning model. This explains why the machine learning model 220 classified a given ticket along with a given classification label. A typical IT domain has several stakeholders, e.g., IT users, IT workers, service availability managers, incident ticket owners, incident ticket assignees, change owners, etc. Different stakeholders may require explanations with different levels of complexity and depth of reasoning. According to one or more embodiments, the explainability of a machine learning decision can be presented in disjunctive normal form and / or positively related features. In Boolean logic, disjunctive normal form is the canonical normal form of a logical formula consisting of disjunctions of conjunctions. Disjunctive normal form can be denoted by terms such as "(multiple) OR," "(multiple) AND," etc.

[0059] Returning to explainability using disjunctive normals, Figure 7 is a flowchart of a computer implementation method 700 for illustrating machine learning decisions using disjunctive normals, in one or more embodiments. Figure 7 is illustrated with reference to Figure 8, which is a block diagram illustrating the conversion from a linear classification formula to a disjunctive normal in one or more embodiments. The machine learning model 220 receives a ticket as input and outputs a classification label, for example, as a disk handler.

[0060] In block 702 of the computer implementation method 700, the software application 204 is configured to extract a linear classification formula from the machine learning model 220. An example of a linear classification formula is shown in block 802 of Figure 8. 1i f1+β 2i f² + ... + β ki f k This is illustrated as follows, and in the formula, β ki represents an exemplary coefficient of the kth feature, and f krepresents an exemplary feature representing the sum of k features. In Figure 8, the output classification label is a disk handler with coefficients as the weights (which can be positive and negative) of each feature. In this example, the linear classification formula for the class label for a given ticket is disk handler = 0.2(filesystem + 0.3(mounted) + 0.1(limit) - 0.5(database) - 0.1(cpu) ..., as illustrated in block 802 of Figure 8.

[0061] In block 704, the software application 204 is configured to select coefficients (i.e., weights) of positively signed features that are used in the disjunctive normal form (i.e., "AND" together).

[0062] In block 706, optionally, the software application 204 is configured to select the coefficients (i.e., weights) of negatively signed features that are used (i.e., "ANDed") in the disjunctive normal form, while preserving their negative sign. In one or more embodiments, when executing blocks 704 and 706, the software application 204 may employ or invoke a natural language processing (NLP) model 228 to parse and analyze the features and their respective coefficients (negative and positive values) in the linear classification. In one or more embodiments, the NLP model 228 may be a pre-trained NLP model that has been further trained on the features and their respective coefficients (negative and positive values) in known linear classifications to select the features and their coefficients for use in the disjunctive normal form.

[0063] In block 708, the software application 204 is configured to transform features with selected coefficients into a disjunctive normal form as rules for determining the machine learning model 220. In one or more embodiments, the software application 204 may select only features with coefficients that have positive values, or features with coefficients that have positive values ​​above a threshold, or features with coefficients that have negative values ​​above a threshold along with features with negative coefficient values ​​above a threshold. An example of a rule in the disjunctive normal form is illustrated in block 804 of Figure 8. Specifically, block 804 illustrates three different rule examples, each separated by "OR", as seen in Figure 8. Note that the term database has a negative sign to indicate its negative contribution to the classification label disk handler. In some embodiments, the negative sign may be replaced with "NOT" to represent a negative contribution. Examples of rules are displayed to the user of the machine learning model 220 (on display 119) to explain the decision-making underlying the output of the machine learning model 220 for a given ticket, such as a class label, e.g., a disk handler.

[0064] In one or more embodiments, the software application 204 is configured to employ / call a rule generation algorithm 224 to generate rules in disjunctive normal form. In one or more embodiments, the rule generation algorithm 224 may be a rule-based algorithm. An example of a rule-based system is a domain-specific expert system that uses rules to make inferences or choices. A rule-based system includes a set of facts or data sources related to the subject being captured, and a set of rules for manipulating that data. These rules are sometimes referred to as "if statements" because they tend to follow a line of "if happens then do Y".

[0065] In one or more embodiments, the rule generation algorithm 224 may be a machine learning algorithm that has been trained on training data. For example, the training data may include a linear classification formula for each classification label within the IT domain of tickets, and the training data may include positive and negative features and their corresponding coefficients. During the training phase of the rule generation algorithm 224, a subject area expert / IT expert can accept or reject rules in the disjunctive normal form for each class, thereby improving the rule generation algorithm as a machine learning algorithm. During the inference phase, once the rule generation algorithm 224 is implemented as a trained machine learning algorithm / model, this trained machine learning algorithm receives input classification labels, features, and their respective coefficients to output the disjunctive normal form of features and logical terms (e.g., "AND", "OR", etc.) as shown in block 804 of Figure 8.

[0066] When describing explainability using disjunctive normal form, the machine learning model 220 uses a linear classifier, such as a passive-aggressive classifier that includes unigrams and bigrams as features. In one embodiment, the presence or absence of these features provides the software application 204 with information about the type of incident, such as the type of ticket. Describing the behavior of the machine learning model 220 (i.e., the classifier) ​​helps the user to increase confidence in the classifier. Therefore, one or more embodiments describe classification in disjunctive normal form, which is a natural form of knowledge representation for humans.

[0067] Returning to explainability using positively related features, Figure 9 is a flowchart of a computer implementation method 900 for explaining machine learning decisions using positively related features, according to one or more embodiments. Figure 9 is illustrated with reference to Figure 10, which is a block diagram illustrating the selection of features (e.g., tokens) from ticket data, according to one or more embodiments, and presenting positively related features that contribute to a classification label for a ticket. The machine learning model 220 takes a ticket as input and outputs a classification label, e.g., disk handler.

[0068] In block 902, the software application 204 is configured to extract positively relevant features from the machine learning model 220 for tickets classified under a given classification label. Any method considered herein may be used to determine positively relevant features for the classification labels.

[0069] In block 904, the software application 204 is configured to select positively related features that have the highest value, for example, a value exceeding a predetermined threshold. Examples of positively related features are illustrated in Figures 5A and 5C.

[0070] In block 906, the software application 204 is configured to display (e.g., on display 119) selected positively related features with the highest values ​​contributing to (i.e., influencing) the decisions of the machine learning model 220. For example, the selected positively related features are configured to show explanations related to individual predictions (i.e., what is related to the features of this particular ticket that suggested the automation of the disk handler). In one or more embodiments, for example, the selected positively related features are configured to show only explanations related to individual predictions. As displayed to the user (e.g., on display 119), Figure 10 shows the machine learning rules extracted in the disjunctive normal form in block 1002 and the machine learning features that influence the decisions in block 1004. Figure 10 also displays the ticket descriptions for the tickets in block 1006 and a bar graph showing the positively related features that influence the decisions of the machine learning model 220 in block 1008. In block 1008, the feature disk has a greater influence on the decisions than the feature space, although both are used to explain the decisions of the machine learning model 220.

[0071] Figure 11 is a flowchart of a computer implementation method 1100 for providing self-explainable classification using a client-independent machine learning model, according to one or more embodiments. The computer implementation method 1100 can be performed by a computer system 202. Any drawings considered herein can be referenced.

[0072] In block 1102, the machine learning model 220 receives input in the form of a record (e.g., a ticket) that relates to the information technology (IT) domain. In block 1102, the machine learning model 220 classifies the record along with a label, and the machine learning model refrains from classifying a given record (e.g., another ticket) in response to a given record that is outside the scope of the IT domain.

[0073] In one or more embodiments, the machine learning model 220 identifies a given record outside the IT domain as unclassified. The machine learning model 220 is trained on training data 206 within the IT domain. The machine learning model 220 is trained by receiving input to the training data 206, which has training records (e.g., tickets) and their corresponding labels. The machine learning model 220 is trained to refrain from classifying any records outside the IT domain.

[0074] Furthermore, the records and labels are provided to an automated resolution system 222, which is configured to modify at least one component in the industry's IT environment (e.g., a hardware component, a software component, and / or both a hardware component and a software component). A given record outside the scope of the IT domain is prevented from being provided to the automated resolution system 222, thereby preventing any component in the IT environment (e.g., a hardware component, a software component, and / or both a hardware component and a software component) from being modified based on a misclassification of a given record.

[0075] Machine learning model 220 includes a linear classifier algorithm, and the record is a ticket for a technical problem in an IT environment. The linear classifier algorithm is trained on training data in the IT domain, and the linear classifier algorithm is trained on positively relevant features for labels without negatively relevant features (as illustrated in Figures 5A and 5C), and the positively relevant features for labels have been validated by human subject matter experts.

[0076] Figure 12 is a flowchart of a computer implementation method 1200 for providing self-explainable classification using a client-independent machine learning model, according to one or more embodiments. The computer implementation method 1200 can be performed by a computer system 202. Any drawings considered herein can be referenced.

[0077] In block 1202, the machine learning model 220 classifies input records (e.g., tickets), and the machine learning model refrains from classifying a given record (e.g., another ticket) in response to a given record that is outside the information technology (IT) domain. In block 1204, the software application 204 and / or the machine learning model 220 classify the records along with labels by generating an explanation of the decisions made by the machine learning model. In block 1206, the software application 204 and / or the machine learning model 220 provide the explanation in a human-readable format to the display (e.g., on display 119).

[0078] In one or more embodiments, the human-readable format includes a disjunctive normal form. The explanation of the decision by the machine learning model 220 is based on a linear classification formula used by the machine learning model 220. The explanation of the decision by the machine learning model 220 is based on features and their respective coefficients, which are derived from the linear classification formula of the machine learning model, as shown in block 802 of Figure 8. That is, the features are extracted from the text (e.g., tokenized text) of the record (e.g., a ticket), as shown in block 1006 of Figure 10. Furthermore, the human-readable format includes a display of the positively related features (e.g., on a display 119) with some respective contribution for each positively related feature to the decision by the machine learning model 220, as shown in block 1008 of Figure 10. Furthermore, the machine learning model is trained on training data in the IT domain. The machine learning model includes a linear classifier algorithm, and the record is a ticket for a technical problem in the IT environment.

[0079] In one or more embodiments, the machine learning model 220, the rule generation algorithm 224, and / or the NLP model 228 may include various engines / classifiers and / or be implemented on a neural network. Features of the engines / classifiers can be implemented by configuring and deploying a computer system 202 that runs the machine learning algorithms. Generally, machine learning algorithms extract features from the received data (e.g., tickets for technical computer problems) in order to “classify” the received data. Examples of suitable classifiers include, but are not limited to, neural networks, support vector machines (SVMs), logistic regression, decision trees, and hidden Markov models (HMMs). The final result of the classifier's operation, i.e., “classification,” is to predict a class (or label) about the data. Machine learning algorithms apply machine learning techniques to the received data to create / train / update a unique “model” over time. The learning or training performed by the engines / classifiers can be supervised, unsupervised, or a hybrid including aspects of supervised and unsupervised learning. Supervised learning is when training data is already available and is classified / labeled. Unsupervised learning is when training data is not classified / labeled and therefore needs to be progressed through iterative classifiers. Unsupervised learning can utilize additional learning / training methods, such as clustering, anomaly detection, neural networks, and deep learning.

[0080] In one or more embodiments, the engine / classifier is implemented as a neural network (or artificial neural network) that uses connections between preneurons and postneurons to represent connection weights. These connections represent, for example, synapses between preneurons and postneurons. A neuromorphic system is an interconnected system of elements that function as stimulated “neurons” and exchange “messages” with one another. Similar to the so-called “plasticity” of synaptic neurotransmitter connections that carry messages between biological neurons, connections within a neuromorphic system, such as a neural network, carry electronic messages between stimulated neurons, and these messages are provided with numerical weights corresponding to the strength or weakness of a given connection. These weights can be adjusted or regulated based on experience, thereby allowing the neuromorphic system to adapt to inputs and learn. After being weighted and deformed by a function (i.e., transport function) determined by the network designer, the activation of these input neurons then moves to other downstream neurons, often also called “hidden” neurons. This process is repeated until an output neuron is activated. Therefore, activated output neurons determine (or "learn") and provide reasoning about the output or input.

[0081] A training dataset (e.g., trainingdata206) can be used to train a machine learning algorithm. The training dataset can include historical data of past tickets and corresponding choices / suggestions / solutions for each ticket. Labels for choices / suggestions can be applied to each ticket to train the machine learning algorithm as part of supervised learning. For preprocessing, the raw training dataset may be manually collected and classified. The classified datasets may be labeled (for example, using Amazon Web Services® labeling tools such as Amazon SageMaker® Ground Truth). The training dataset may be divided into training, test, and validation datasets. The training and validation datasets are used for training and evaluation, while the test dataset is used after training and testing the machine learning model on unseen datasets. The training dataset may be processed through different data augmentation techniques. Training takes the labeled dataset, base network, loss function, and hyperparameters, and once all of these are created and compiled, the neural network is trained, ultimately producing a trained machine learning model (e.g., a trained machine learning algorithm). Once the model is trained, it is saved to a file (including the adjusted weights) for development of the test dataset and / or further testing.

[0082] Although this disclosure includes a detailed description of cloud computing, it will be understood that implementations of the teachings described herein are not limited to cloud computing environments. Rather, embodiments of the present invention may be implemented in combination with any other type of computing environment that is currently known or may be developed in the future.

[0083] Cloud computing is a service delivery model that enables convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services), allowing these resources to be provisioned and released quickly with minimal administrative effort or interaction with service providers. This cloud model includes at least five features, at least three service models, and at least four deployment models.

[0084] The features are as follows:

[0085] On-demand self-service: Cloud users can unilaterally and automatically provision server time and computing power such as network storage as needed, without requiring human interaction with the service provider.

[0086] Broad network access: This capability is available over a network and can be accessed using standard mechanisms, thus facilitating use by heterogeneous thin-client or thick-client platforms (e.g., mobile phones, laptops, and PDAs).

[0087] Resource Pooling: A provider's computing resources are pooled and delivered to multiple users using a multi-tenant model, with various physical and virtual resources dynamically allocated and reallocated as needed. Users typically have a sense of location independence in that they have neither control nor know the exact location of the resources provided, although they can still specify a location (e.g., country, state, or data center) at a higher level of abstraction.

[0088] Rapid Adaptability: Capabilities can be provisioned quickly and flexibly, sometimes automatically, scale out rapidly, and be released quickly to scale in rapidly. The capacity available for provisioning often appears to the user as if they can purchase any amount at any time without limit.

[0089] Measured Services: Cloud systems leverage metering capabilities to automatically control and optimize resource usage at an appropriate level of abstraction for each service type (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, and reported, providing transparency to both service providers and users.

[0090] The service model is as follows:

[0091] SaaS (Software as a Service): The capability provided to the user is the use of the provider's applications running on cloud infrastructure. These applications can be accessed from various client devices via thin client interfaces such as web browsers (e.g., web-based email). Users do not manage or control the underlying cloud infrastructure, including the network, servers, operating system, storage, or individual application functions, except for the possibility of making limited user-specific application configuration settings.

[0092] PaaS (Platform as a Service): The ability provided to the user is to deploy applications created or acquired by the user, using programming languages ​​and tools supported by the provider, onto a cloud infrastructure. The user does not manage or control the underlying cloud infrastructure, including the network, servers, operating system, or storage, but can control the configuration of the deployed application and, in some cases, the application hosting environment.

[0093] IaaS (Infrastructure as a Service): The capabilities provided to users include provisioning of processing, storage, networking, and other basic computing resources, allowing users to deploy and run any software, including operating systems and applications. Users do not manage or control the underlying cloud infrastructure, but they can control the operating system, storage, and deployed applications, and in some cases, have limited control over selected network components (e.g., host firewalls).

[0094] The deployment model is as follows:

[0095] Private Cloud: This cloud infrastructure is operated solely for the organization. This cloud infrastructure is managed by this organization or a third party, or resides on-premises, off-premises, or both.

[0096] Community Cloud: This cloud infrastructure is shared by multiple organizations and supports specific communities that share common interests (e.g., missions, security requirements, policies, and compliance considerations). This cloud infrastructure may be managed by these organizations or third parties, or it may reside on-premises, off-premises, or both.

[0097] Public Cloud: This cloud infrastructure is available for use by general users or large industry groups and is owned by the organization that sells the cloud service.

[0098] Hybrid Cloud: This cloud infrastructure is a combination of two or more clouds (private, community, or public) that are joined together while retaining their own distinct entities, through standardized or proprietary technologies that enable the portability of data and applications (e.g., cloud bursting to adjust load balancing between clouds).

[0099] Cloud computing environments are service-oriented environments that emphasize statelessness, loose coupling, modularity, and semantic interoperability. At the heart of cloud computing is an infrastructure consisting of a network of interconnected nodes.

[0100] Referring here to Figure 13, an exemplary cloud computing environment 50 is shown. As illustrated, the cloud computing environment 50 includes one or more cloud computing nodes 10 on which local computing devices used by cloud users (e.g., personal digital assistants (PDAs) or mobile phones 54A, desktop computers 54B, laptop computers 54C, or automotive computer systems 54N, or a combination thereof) communicate with each other. The nodes 10 communicate with each other. These nodes are grouped physically or virtually (not illustrated) within one or more networks, such as private clouds, community clouds, public clouds, or hybrid clouds, or a combination thereof, as described herein. This allows the cloud computing environment 50 to provide infrastructure, platforms, or software-as-services, or a combination thereof, that cloud users do not need to maintain resources on their local computing devices. The types of computing devices 54A to 54N shown in Figure 13 are intended for illustrative purposes only, and it is understood that the computing node 10 and the cloud computing environment 50 can communicate with any type of computer-controlled device via any type of network or network-addressable connection (e.g., a connection using a web browser) or both.

[0101] Referring now to Figure 14, a set of functional abstraction layers provided by the cloud computing environment 50 (Figure 13) is shown. It should be understood in advance that the components, layers, and functions shown in Figure 14 are intended for illustrative purposes only, and embodiments of the present invention are not limited thereto. As illustrated, the following layers and corresponding functions are provided:

[0102] The hardware and software layer 60 includes hardware components and software components. Examples of hardware components include a mainframe 61, RISC (Reduced Instruction Set Computer) architecture-based servers 62, 63, blade servers 64, storage devices 65, and networks and network components 66. In some embodiments, software components include network application server software 67 and database software 68.

[0103] The virtualization layer 70 includes an abstraction layer that provides virtual entities such as virtual servers 71, virtual storage 72, virtual networks 73 including virtual private networks, virtual applications and operating systems 74, and virtual clients 75.

[0104] For example, the management layer 80 may provide the following functions: Resource provisioning 81 provides dynamic procurement of computing and other resources used to perform tasks within the cloud computing environment. Measurement and pricing 82 provides tracking of costs as resources are used within the cloud computing environment and billing or invoicing for the consumption of these resources. For example, these resources may include application software licenses. Security provides identity verification of cloud consumers and tasks, as well as protection of data and other resources. User portal 83 provides consumers and system administrators with access to the cloud computing environment. Service level management 87 provides allocation and management of cloud computing resources to ensure that the required service levels are met. Service Level Agreement (SLA) planning and achievement 88 provides proactive preparation and procurement of cloud computing resources in accordance with the SLA, where future demands are anticipated.

[0105] Workload Layer 90 provides examples of functions used in a cloud computing environment. Examples of workloads and functions provided by this layer include mapping and navigation 91, software development and lifecycle management 92, virtual classroom education delivery 93, data analytics processing 94, transaction processing 95, and workloads and functions 96.

[0106] Various embodiments of the present invention are described herein with reference to the relevant drawings. Alternative embodiments can be invented without departing from the spirit of the invention. Various connections and positional relationships (e.g., above, below, adjacent, etc.) are described between elements in the following description and drawings, but those skilled in the art will recognize that many of the positional relationships described herein are independent of direction, as the described functionality is maintained even if the orientation is changed. These connections and / or positional relationships can be direct or indirect unless otherwise specified, and the invention is not intended to be limited in this respect. Thus, the connection between entities can refer to a direct or indirect connection, and the positional relationship between entities can refer to a direct or indirect positional relationship. As an example of an indirect positional relationship, the reference herein to forming layer "A" on layer "B" includes the state in which one or more intermediate layers (e.g., layer "C") are between layer "A" and layer "B", as long as the relevant features and functionality of layers "A" and "B" are not substantially altered by the intermediate layer.

[0107] For the sake of brevity, prior art relating to carrying out and using aspects of the present invention may or may not be described in detail herein. Specifically, various aspects of computing systems and particular computer programs implementing the various technical features described herein are well known. Therefore, for the sake of brevity, many details of prior art implementations are either briefly stated herein or omitted entirely, and details of well known systems and / or processes are not shown.

[0108] In some embodiments, various functions or actions may be performed at a given location and / or in connection with one or more devices or systems. In some embodiments, some given functions or actions may be performed at a first device or location, and the remaining functions or actions may be performed at one or more additional devices or locations.

[0109] The terms used herein are for the sole purpose of describing specific embodiments and are not intended to limit the invention. Where used herein, the singular forms “a,” “an,” and “the” are intended to include the plural form unless otherwise explicitly indicated in the context. It will be further understood that the terms “equipped with” or “equipped with” or both, where used herein, indicate the presence of a described function, integer, step, operation, element, or component, or a combination thereof, but do not exclude the presence or addition of one or more other functions, integers, steps, operations, elements, components, or groups thereof, or combinations thereof.

[0110] The corresponding structures, materials, actions, and equivalents of all means-plus-function elements or step-plus-function elements within the following claims are intended to include any structures, materials, or actions for performing a function in combination with other elements described in the claims, specifically as described in the claims. While this disclosure has been presented for illustrative and explanatory purposes, it is not intended to be exhaustive or to limit the disclosed forms. Many modifications and variations will be apparent to those skilled in the art without departing from the scope of this disclosure. Embodiments have been selected and described to best illustrate the principles and practical applications of this disclosure, and to enable others skilled in the art to understand this disclosure in relation to various embodiments with various modifications suitable for a particular intended use.

[0111] The diagrams shown herein are illustrative. Numerous variations are possible to the diagrams or steps (or actions) described herein without departing from the spirit of this disclosure. For example, actions may be performed in a different order, and may be added, deleted, or modified. Also, the term “coupled” describes the presence of a signal path between two elements and does not imply a direct connection between those elements without an intervening element / connection. All of these variations are considered to be part of this disclosure.

[0112] The following definitions and abbreviations are to be used to interpret the claims and this specification. As set forth herein, the terms “comprises,” “comprising,” “includes,” “including,” “has,” “having,” “contains,” or “containing,” or any variation thereof, are intended to include non-exclusive inclusion. For example, a construct, mixture, process, method, article, or apparatus that includes an enumeration of elements may include other elements that are not explicitly enumerated or that are not inherently present in such construct, mixture, process, method, article, or apparatus, but are not necessarily limited to those elements.

[0113] Furthermore, the term “exemplary” is used herein to mean “serving as an example, illustration, or explanatory role.” Any embodiment or design described herein as “exemplary” is not necessarily construed to be preferable or advantageous to other embodiments or designs. The terms “at least one” and “one or more” are understood to include one or more any integers, i.e., 1, 2, 3, 4, etc. The term “multiple” is understood to include two or more any integers, i.e., 2, 3, 4, 5, etc. The term “connection” may include both indirect and direct “connections.”

[0114] The terms “about,” “substantially,” “approximately,” and their variations are intended to include the degree of error associated with measuring specific quantities based on instruments available at the time of filing this application. For example, “about” may include a range of ±8%, 5%, or 2% of a given value.

[0115] The present invention may be any possible and technically detailed system, method, or computer program product, or combination thereof, of integrated levels. The computer program product may include a computer-readable storage medium(s) having computer-readable program instructions for causing a processor to perform aspects of the present invention.

[0116] A computer-readable storage medium can be a tangible device capable of holding and storing instructions for use by an instruction execution device. A computer-readable storage medium may, but is not limited to, electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. A non-exhaustive list of further specific examples of computer-readable storage media includes portable floppy disks, hard disks, random-access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random-access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disks (DVDs), memory sticks, floppy disks, mechanically encoded devices such as punched cards or grooved raised structures on which instructions are recorded, and any suitable combination thereof. When used herein, computer-readable storage media should not be interpreted as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmitting media (e.g., light pulses passing through optical fiber cables), or electrical signals transmitted through wires.

[0117] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to each computing device / processing device, or to an external computer or external storage device via a network (e.g., the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof). This network may include copper transmission cables, optical transmission fibers, wireless transmitters, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. A network adapter card or network interface within each computing device / processing device receives computer-readable program instructions from the network and transfers those computer-readable program instructions for storage on a computer-readable storage medium within each computing device / processing device.

[0118] The computer-readable program instructions for performing the operation of the present invention may be assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, configuration data for integrated circuits, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk and C++, and procedural programming languages ​​such as the C programming language or similar programming languages. The computer-readable program instructions can be executed as a whole on the user's computer, partially as a standalone software package on the user's computer, partially on the user's computer and on a remote computer, respectively, or entirely on a remote computer or on a server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or wide area network (WAN), or the connection may be made to an external computer (for example, via the Internet using an Internet service provider). In some embodiments, to carry out aspects of the present invention, an electronic circuit including, for example, a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA) may execute computer-readable program instructions to customize the electronic circuit by utilizing state information of computer-readable program instructions.

[0119] Aspects of the present invention will be described herein with reference to flowcharts or block diagrams, or both, of methods, apparatus (systems), and computer program products, according to embodiments of the present invention. It will be understood that each block in a flowchart or block diagram, or both, and combinations of blocks contained in a flowchart or block diagram, or both, can be implemented by computer-readable program instructions.

[0120] These computer-readable program instructions may be provided to a general-purpose computer, a dedicated computer, or a processor of another programmable data processing device to create a machine, so that instructions executed via the processor of a computer or other programmable data processing device can create means to perform functions / operations specified in one or more blocks of a flowchart or block diagram or both. These computer-readable program instructions may be stored on a computer-readable storage medium containing instructions that include instructions to perform modes of functions / operations specified in one or more blocks of a flowchart or block diagram or both, and can instruct a computer, a programmable data processing device, or other device, or a combination thereof, to function in a particular manner.

[0121] Computer-readable program instructions may be read into a computer, another programmable data processing device, or other device so that instructions executed on a computer, another programmable device, or other device perform functions / operations specified in one or more blocks of a flowchart or block diagram or both, thereby causing a series of operable steps to be executed on a computer, another programmable device, or other device that generates a computer implementation process.

[0122] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing a specified logical function(s). In some alternative implementations, the functions shown within a block may occur in a different order than that shown in the drawings. For example, two consecutively shown blocks may actually be executed substantially simultaneously, and these blocks may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart diagram, or both, and any combination of blocks in a block diagram or flowchart diagram, or both, may be implemented by a purpose-specific hardware-based system that performs a specified function or operation, or executes a particular combination of purpose-specific hardware and computer instructions.

[0123] While various embodiments of the present invention have been presented for illustrative purposes, they are not intended to be comprehensive or to limit the disclosed embodiments. Many modifications and changes will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein has been selected to best describe the principles, practical applications, or technical advancements beyond the technology available on the market of the embodiments, and to enable other those skilled in the art to understand the embodiments described herein.

Claims

1. The process involves classifying tickets along with labels using a machine learning model via a processor, wherein the machine learning model excludes a given ticket that falls outside the scope of the information technology (IT) domain in the classification process. To generate an explanation of the decision to output a classification label by the machine learning model for classifying the tickets together with the aforementioned label, To display the above explanation in a human-readable format and Computer implementation methods, including those mentioned above.

2. The computer implementation method according to claim 1, wherein the aforementioned human-readable form includes disjunctive normal form.

3. The computer implementation method according to claim 1, wherein the explanation of the decision by the machine learning model is based on a linear classification formula used by the machine learning model.

4. The explanation of the decision by the machine learning model is based on features and coefficients corresponding to those features, and the features and coefficients are obtained from the linear classification formula of the machine learning model. The aforementioned features are extracted from the text of the ticket. The computer implementation method according to claim 1.

5. The computer implementation method according to claim 1, wherein the human-readable format includes a representation of positively related features and the degree of contribution of each of the positively related features to the decision made by the machine learning model.

6. The computer implementation method according to claim 1, wherein the machine learning model is trained on training data within the IT domain.

7. The aforementioned machine learning model includes a linear classifier algorithm, The aforementioned ticket is a ticket for a resource-related issue in the IT environment. The computer implementation method according to claim 1.

8. Memory with computer-readable instructions, One or more processors that execute the aforementioned computer-readable instructions A system including, wherein the computer-readable instruction is, The method involves classifying tickets along with labels using a machine learning model, wherein the machine learning model classifies tickets excluding those tickets in response to a given ticket that falls outside the scope of the information technology (IT) domain. To generate an explanation of the decision to output a classification label by the machine learning model for classifying the tickets together with the aforementioned label, To display the above explanation in a human-readable format and A system that controls one or more processors to perform operations including those described above.

9. The system according to claim 8, wherein the aforementioned human-readable form includes a disjunctive normal form.

10. The system according to claim 8, wherein the explanation of the decision by the machine learning model is based on a linear classification formula used by the machine learning model.

11. The explanation of the decision by the machine learning model is based on features and coefficients corresponding to those features, and the features and coefficients are obtained from the linear classification formula of the machine learning model. The aforementioned features are extracted from the text of the ticket. The system according to claim 8.

12. The system according to claim 8, wherein the human-readable format includes a representation of positively related features and the degree of contribution of each of the positively related features to the decision made by the machine learning model.

13. The system according to claim 8, wherein the machine learning model is trained on training data within the IT domain.

14. The aforementioned machine learning model includes a linear classifier algorithm, The aforementioned ticket is a ticket for a resource-related issue in the IT environment. The system according to claim 8.

15. A computer program including a computer-readable storage medium in which program instructions are embodied, wherein the program instructions are The method involves classifying tickets along with labels using a machine learning model, wherein the machine learning model classifies tickets excluding those tickets in response to a given ticket that falls outside the scope of the information technology (IT) domain. To generate an explanation of the decision to output a classification label by the machine learning model for classifying the tickets together with the aforementioned label, To display the above explanation in a human-readable format and A computer program that is executable by one or more processors to perform an operation including the following.

16. The computer program according to claim 15, wherein the aforementioned human-readable form includes disjunctive normal form.

17. The computer program according to claim 15, wherein the explanation of the decision by the machine learning model is based on a linear classification formula used by the machine learning model.

18. The explanation of the decision by the machine learning model is based on features and coefficients corresponding to those features, and the features and coefficients are obtained from the linear classification formula of the machine learning model. The aforementioned features are extracted from the text of the ticket. The computer program according to claim 15.

19. The computer program according to claim 15, wherein the human-readable format includes representations of positively related features and the extent of each of the positively related features' contributions to the decision made by the machine learning model.

20. The computer program according to claim 15, wherein the machine learning model is trained on training data within the IT domain.

Citation Information

Patent Citations

  • Hybrid learning-based ticket classification and response

    US20200104752A1

  • Model agnostic contrastive explanations for structured data

    US20200193243A1

  • Predictive Resolutions for Tickets Using Semi-Supervised Machine Learning

    US20210019648A1

  • Object identification system, operation processing device, vehicle, lighting tool for vehicle, and training method for classifier

    WO2020121973A1