IT resource management method and device, storage medium and processor
By combining real-time data collection and hash value comparison with an improved PageRank algorithm and causal graph model, the problems of response latency and low fault location efficiency in IT resource management are solved, achieving efficient IT resource operation and maintenance management.
Patent Information
- Application Number
- CN202511032410.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-11-07
AI Technical Summary
Existing IT resource management methods suffer from response delays and low fault location efficiency, making it difficult to meet the real-time requirements of modern systems. Furthermore, traditional polling mechanisms and full data comparisons result in low operational efficiency.
By collecting multi-dimensional IT resource attribute information in real time, converting it into JSON strings and calculating hash values, and automatically comparing hash values to detect changes, combined with an improved PageRank algorithm and cause-effect graph model, alarm correlation and root cause analysis are achieved to generate alarm strategies.
It achieves change detection capabilities at the second or even millisecond level, accurately pinpoints the source of anomalies, significantly shortens troubleshooting time, improves operation and maintenance efficiency, and meets the real-time requirements of modern IT systems.
Smart Images

Figure CN120909827A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, in particular to an IT resource management method and device, a storage medium and a processor. BACKGROUND
[0002] With the rapid development of information technology, the scale and complexity of enterprise IT systems are constantly improving, and the types of IT resources involved are increasingly diverse, including servers, network devices, application systems, databases, and various middleware, etc. In the daily operation process of the system, abnormal alarms caused by non-operation and non-transaction fluctuations frequently occur, and a part of the problem roots often come from abnormal changes of IT resource information, such as server configuration changes, software version upgrades, and operation responsibility person adjustments, etc. However, in the traditional operation system, such changes are often ignored or difficult to be found in time, resulting in low efficiency of system fault troubleshooting, and even causing more serious system failure risks.
[0003] At present, the existing IT resource management mostly adopts a polling mechanism-based detection method, which has obvious response delay, and can usually only achieve minute-level change awareness, which is difficult to meet the real-time requirements of modern systems. In addition, the existing technology needs to compare full data when locating the root cause, which seriously restricts the efficiency of fault location and root cause analysis.
[0004] Therefore, there is an urgent need for an efficient IT resource management method to improve the automation and intelligent level of IT resource change management, thereby improving the operation and management efficiency of IT resources. SUMMARY
[0005] Based on the above problems, the present application provides an IT resource management method, device, storage medium and processor, which aims to realize the operation and management closed loop from abnormal awareness to root cause repair of IT resources in an automated manner, thereby improving the operation and management efficiency of IT resources.
[0006] The embodiments of the present application disclose the following technical solutions:
[0007] In a first aspect, the present application provides an IT resource management method, which comprises:
[0008] Collecting multi-dimensional attribute information corresponding to the IT resources of an enterprise; the multi-dimensional attribute information includes data corresponding to resource attributes of multiple dimensions respectively;
[0009] Converting the multi-dimensional attribute information into a JSON string, and calculating hash values corresponding to each of the resource attributes based on the JSON string;
[0010] comparing each of the hash values with a corresponding historical hash value to obtain a comparison result; the comparison result is used to indicate whether data corresponding to the resource attribute has changed, and the historical hash value is a hash value calculated from data updated last time by the resource attribute;
[0011] if the comparison result indicates that the data corresponding to the resource attribute has changed, associating changed data of the resource attribute with a corresponding alarm event to obtain alarm association information; the alarm event is a system abnormal operation event caused by the data change of the resource attribute;
[0012] calculating a PageRank value of a node related to the alarm association information by a modified PageRank algorithm from a resource attribute causal graph model, and determining an alarm root cause matching the alarm event according to the PageRank value; the PageRank value is used to reflect the importance of the node in the entire system abnormal propagation network;
[0013] generating a corresponding alarm strategy according to the alarm root cause.
[0014] In an optional implementation, the comparing each of the hash values with a corresponding historical hash value to obtain a comparison result comprises:
[0015] comparing whether each of the hash values and the corresponding historical hash value are the same;
[0016] if the hash value and the corresponding historical hash value are the same, determining that the data corresponding to the resource attribute has not changed, and generating a corresponding comparison result;
[0017] if the hash value and the corresponding historical hash value are not the same, determining that the data corresponding to the resource attribute has changed, and generating a corresponding comparison result.
[0018] In an optional implementation, before the associating the changed data of the resource attribute with the corresponding alarm event to obtain the alarm association information, the method further comprises:
[0019] detecting whether the data change of the resource attribute causes the system abnormal operation;
[0020] if the data change of the resource attribute does not cause the system abnormal operation, updating a configuration management database based on the changed data;
[0021] if the data change of the resource attribute causes the system abnormal operation, determining an alarm event corresponding to the data change according to a preset alarm triggering rule.
[0022] In an optional implementation, the resource attribute causal graph model is obtained by the following process:
[0023] Obtaining historical abnormal operation data of the system from a fault database;
[0024] Identifying causal dependency relationships between a plurality of random variables based on the historical abnormal operation data; the plurality of random variables at least include fault types, resource states, and configuration changes;
[0025] Mapping the causal dependency relationships into a directed acyclic graph to construct an initial causal graph model; nodes of the initial causal graph model represent the random variables, and edges of the initial causal graph model represent the causal dependency relationships;
[0026] Obtaining a sample data set; the sample data set includes a plurality of training samples and label data corresponding to the training samples, the training samples include historical alarm correlation information of the system, and the label data includes actual alarm root causes corresponding to the historical alarm correlation information;
[0027] Inputting data in the training sample set into the initial causal graph model to obtain a prediction result output by the initial causal graph model; the prediction result includes alarm root causes corresponding to the historical alarm correlation information;
[0028] Adjusting parameters of the initial causal graph model based on differences between the alarm root causes corresponding to the historical alarm correlation information and the actual alarm root causes, and iteratively training the adjusted initial causal graph model by using training samples and corresponding label data in the sample data set that are not used, until a training termination condition is met, to obtain the resource attribute causal graph model.
[0029] In an optional implementation, the inputting data in the training sample set into the initial causal graph model to obtain a prediction result output by the initial causal graph model includes:
[0030] Inputting data in the sample data set into the initial causal graph model, and calculating PageRank values of nodes related to the historical alarm correlation information by the initial causal graph model through an improved PageRank algorithm; the improved PageRank algorithm is used to quantify importance of the nodes to convert abnormal propagation paths of the system during abnormal operation into calculable weights;
[0031] Predicting alarm root causes matching the historical alarm correlation information based on the PageRank values, and generating corresponding prediction results.
[0032] In an optional implementation, the generating, according to the alarm root cause, of the corresponding alarm strategy comprises:
[0033] determining, according to a preset level allocation rule, an alarm level corresponding to the alarm root cause;
[0034] determining, based on the alarm level, an alarm strategy corresponding to the alarm event; the alarm strategy comprises an alarm response time and an alarm notification mode.
[0035] In a second aspect, the application provides an IT resource management device, which comprises:
[0036] a collection module configured to collect multi-dimensional attribute information corresponding to IT resources of an enterprise; the multi-dimensional attribute information comprises data corresponding to resource attributes of multiple dimensions respectively;
[0037] a calculation module configured to convert the multi-dimensional attribute information into a JSON string, and calculate hash values corresponding to the resource attributes based on the JSON string;
[0038] a comparison module configured to compare each of the hash values with a corresponding historical hash value to obtain a comparison result; the comparison result is used to indicate whether the data corresponding to the resource attributes has been changed, and the historical hash value is a hash value calculated by using data updated last time by the resource attributes;
[0039] an association module configured to associate changed data of the resource attributes with a corresponding alarm event to obtain alarm association information, if the comparison result indicates that the data corresponding to the resource attributes has been changed; the alarm event is a system abnormal operation event caused by the change of the data of the resource attributes;
[0040] a determination module configured to calculate, by using an improved PageRank algorithm, a PageRank value of a node related to the alarm association information according to a resource attribute causal graph model, and determine an alarm root cause matched with the alarm event according to the PageRank value; the PageRank value is used to reflect an importance degree of the node in an abnormal propagation network of the entire system;
[0041] a generation module configured to generate a corresponding alarm strategy according to the alarm root cause.
[0042] Optionally, the comparison module comprises:
[0043] a comparison unit configured to compare whether each of the hash values and the corresponding historical hash value are the same;
[0044] The first determining unit is configured to determine that the data corresponding to the resource attribute has not changed if the hash value is the same as the corresponding historical hash value, and to generate a corresponding comparison result.
[0045] The second determining unit is configured to determine that the data corresponding to the resource attribute has changed if the hash value is not the same as the corresponding historical hash value, and to generate a corresponding comparison result.
[0046] In a third aspect, the present application provides a computer readable storage medium, and the computer readable storage medium stores a computer program. When the computer program is run by a processor, the IT resource management method described above is implemented.
[0047] In a fourth aspect, the present application provides a processor for running a computer program. When the computer program is run, the IT resource management method described above is executed.
[0048] Compared with the prior art, the present application has the following beneficial effects:
[0049] In the technical solution of the present application, the multi-dimensional attribute information corresponding to the IT resources of an enterprise is collected in real time. The multi-dimensional attribute information includes data corresponding to resource attributes of multiple dimensions, which avoids the system response delay problem caused by the traditional polling mechanism for detecting IT resources, significantly reduces the system response delay, and improves the data processing efficiency.
[0050] Subsequently, each hash value is automatically compared with the corresponding historical hash value to obtain a comparison result. The comparison result is used to indicate whether the data corresponding to the resource attribute has changed. The historical hash value is a hash value calculated from the data updated last time by the resource attribute. Therefore, it is not necessary to perform full data comparison to determine whether the resource attribute has changed, which significantly shortens the change detection time, thereby realizing the change perception ability of seconds or even milliseconds, and meeting the strict real-time requirements of modern IT systems. If the comparison result indicates that the data corresponding to the resource attribute has changed, the change data of the resource attribute is automatically associated with the corresponding alarm event to obtain alarm association information. The alarm event is a system abnormal operation event caused by the data change of the resource attribute. Therefore, the alarm association information can directly lock the abnormal source, and provides data preparation for subsequent alarm root cause analysis.
[0051] The pre-trained resource attribute causal graph model subsequently calculates the PageRank value of the node related to the alarm associated information through an improved PageRank algorithm, and determines the alarm root cause matched with the alarm event according to the PageRank value, wherein the PageRank value is used to reflect the importance of the node in the abnormal propagation network of the whole system, and the importance of the node in the abnormal propagation network (such as the incompatible problem of dependent libraries) can be accurately evaluated through the PageRank value, thereby avoiding the limitations of traditional static analysis, significantly shortening the troubleshooting time, and improving the efficiency of alarm root cause positioning; finally, the corresponding alarm strategy is generated according to the alarm root cause, realizing the operation and maintenance closed loop of IT resource from abnormal perception to root cause repair, and significantly improving the operation and maintenance efficiency of IT resource. BRIEF DESCRIPTION OF DRAWINGS
[0052] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor under the premise of not paying creative labor.
[0053] Figure 1 The flow chart of an IT resource management method provided by the embodiments of the present application;
[0054] Figure 2 The flow chart of another IT resource management method provided by the embodiments of the present application;
[0055] Figure 3 The flow chart of the training process of a resource attribute causal graph model provided by the embodiments of the present application;
[0056] Figure 4 The flow chart of the training process of another resource attribute causal graph model provided by the embodiments of the present application;
[0057] Figure 5 The structural schematic diagram of an IT resource management device provided by the embodiments of the present application. DETAILED DESCRIPTION
[0058] As described above, at present, the existing IT resource management mostly adopts the polling mechanism-based detection mode, which has obvious response delay, and can only realize minute-level change awareness, which is difficult to meet the real-time requirements of modern systems. In addition, the existing technology generally adopts full-amount log storage mode when recording resource change logs, which not only occupies a large amount of storage space, but also needs to compare full-amount data when locating the root cause, which seriously restricts the efficiency of fault positioning and root cause analysis. Therefore, an efficient IT resource management method is needed to improve the automation and intelligent level of IT resource change management, thereby improving the operation and management efficiency of IT resources.
[0059] The inventors have proposed an IT resource management method, which avoids the system response delay problem caused by the traditional polling mechanism detection of IT resources, significantly reduces the system response delay, by real-time collection of multi-dimensional attribute information corresponding to the enterprise's IT resources, wherein the multi-dimensional attribute information includes data corresponding to resource attributes of multiple dimensions respectively; then automatically converts the multi-dimensional attribute information into a JSON string, and calculates the hash value corresponding to each resource attribute based on the JSON string, which realizes the replacement of the original data by the hash value for subsequent comparison, greatly reduces the storage space occupation, and reduces the calculation complexity, thereby improving the data processing efficiency.
[0060] Subsequently, each hash value is automatically compared with the corresponding historical hash value to obtain a comparison result, wherein the comparison result is used to indicate whether the data corresponding to the resource attribute has changed, and the historical hash value is a hash value calculated by the data updated last time by the resource attribute, so that whether the resource attribute has changed can be determined without full-amount data comparison, which significantly shortens the change detection time, thereby realizing the change awareness ability of seconds or even milliseconds, and meeting the strict real-time requirements of modern IT systems; if the comparison result indicates that the data corresponding to the resource attribute has changed, the change data of the resource attribute is automatically associated with the corresponding alarm event to obtain alarm association information, wherein the alarm event is a system abnormal operation event caused by the data change of the resource attribute, so that the alarm association information can directly lock the abnormal source, and provide data preparation for subsequent alarm root cause analysis.
[0061] Then, the pre-trained resource attribute causal graph model calculates the PageRank value of the node related to the alarm associated information by the improved PageRank algorithm, and determines the alarm root cause matched with the alarm event according to the PageRank value, wherein the PageRank value is used to reflect the importance degree of the node in the abnormal propagation network of the whole system, and the importance of the node in the abnormal propagation network (such as the incompatible problem of dependent library) can be accurately evaluated through the PageRank value, avoiding the limitation of traditional static analysis, significantly shortening the troubleshooting time, and improving the efficiency of alarm root cause positioning; finally, the corresponding alarm strategy is generated according to the alarm root cause, realizing the operation and maintenance closed loop from abnormal perception to root cause repair of the IT resource, and significantly improving the operation and maintenance efficiency of the IT resource.
[0062] In order to enable personnel in the art to better understand the present application scheme, the technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0063] Key word definition:
[0064] Improved PageRank algorithm: PageRank algorithm is a classic algorithm for measuring the importance of nodes in a graph, which is used to evaluate the weight of a web page in the Internet. In the resource attribute causal graph model of the present application, the improved PageRank algorithm is used to locate the root cause node of resource anomaly, and the complex abnormal propagation link is converted into a calculable weight by quantifying the influence of the node.
[0065] Method embodiment
[0066] The embodiments of the IT resource management method provided by the present application need to be explained that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from here.
[0067] Referring to Figure 1 , the figure is a flowchart of an IT resource management method provided by an embodiment of the present application, as Figure 1 shown, the method comprises the following steps:
[0068] Step S101, collecting the multi-dimensional attribute information corresponding to the enterprise's IT resource.
[0069] In an alternative embodiment, an IT resource management system can serve as an execution subject of the IT resource management method of the embodiments of the present application. For the convenience of description, the IT resource management system is referred to as the system hereinafter.
[0070] In step S101, the data corresponding to each of the resource attributes of multiple dimensions in the multi-dimensional attribute information; for example, the data corresponding to the resource attributes of servers, network devices, application systems, databases, operation and maintenance contacts, departments, versions, Ant states, and various types of middleware.
[0071] In the embodiments of the present application, the system collects the multi-dimensional attribute information corresponding to the IT resources of the enterprise shown in FIG. 1 in real time through multiple protocols such as SSH (i.e., Secure Shell, used to obtain version information by executing relevant commands), SNMP (full name: Simple Network Management Protocol, used to collect hardware indicators), RESTful API (used to obtain cloud resource states), and the like. Figure 2 The multi-dimensional attribute information corresponding to the IT resources of the enterprise shown in FIG. 1 avoids the problem of system response delay caused by the traditional polling mechanism for detecting IT resources, and significantly reduces the system response delay.
[0072] It should be noted that the collection frequency of the system when collecting multi-dimensional attribute information is divided into active polling and event-driven. Among them, the default interval of active polling is 60s, which can also be configured according to requirements; event-driven captures change events related to configuration fields in real time through file listening.
[0073] Step S102, convert the multi-dimensional attribute information into a JSON string, and calculate the hash value corresponding to each resource attribute based on the JSON string.
[0074] In the embodiments of the present application, the system converts the multi-dimensional attribute information into a JSON string, and calculates the hash value corresponding to each resource attribute based on the JSON string through the SHA-256 hash algorithm, to ensure the uniqueness of the data, so that the hash value is used instead of the original data for subsequent comparison, greatly reducing the storage space occupation, while reducing the calculation complexity, and improving the data processing efficiency.
[0075] Optionally, after calculating the hash value, the system can use Redis as an in-memory cache to store the current hash value: (Key: resource: <resourceid>Value: <current_hash>).
[0076] In step S103, each hash value is compared with the corresponding historical hash value to obtain a comparison result.
[0077] In step S103, the comparison result is used to indicate whether the data corresponding to the resource attribute has changed, and the historical hash value is a hash value calculated by the data updated last time by the resource attribute.
[0078] In the embodiments of the present application, the system compares each hash value with the corresponding historical hash value to obtain a comparison result. Specifically, the system determines whether the data corresponding to the resource attribute has changed by comparing whether each hash value is the same as the corresponding historical hash value. If the hash value is the same as the corresponding historical hash value, the system can determine that the data corresponding to the resource attribute has not changed, and generates a corresponding comparison result. If the hash value is not the same as the corresponding historical hash value, the system can determine that the data corresponding to the resource attribute has changed, and generates a corresponding comparison result. Thus, without full data comparison, it can be determined whether the resource attribute has changed, the change detection time is significantly shortened, and the change perception ability of seconds or even milliseconds is realized, meeting the strict requirements of modern IT systems for real-time performance.
[0079] Optionally, the comparison logic of the hash value is specifically as follows:
[0080] def detect_change(resource):
[0081] current_attrs = fetch_attributes(resource) # Collect the latest attributes
[0082] current_hash = sha256(json.dumps(current_attrs)).hexdigest()
[0083] last_hash = redis.get(f"resource:{resource.id}")
[0084] if current_hash!= last_hash:
[0085] diff = compute_diff(last_attrs, current_attrs) # Difference calculation
[0086] trigger_logging(resource.id, diff)
[0087] redis.set(f"resource:{resource.id}",current_hash)
[0088] Step S104: If the comparison result indicates that the data corresponding to the resource attribute has changed, then associate the changed data of the resource attribute with the corresponding alarm event to obtain alarm association information.
[0089] In step S104, the alarm event is a system malfunction event caused by changes in resource attribute data. For example, alarm events such as persistent Ant status errors or deployment failures after version changes.
[0090] In this embodiment of the application, if the comparison result indicates that the data corresponding to the resource attribute has changed, the system can automatically associate the changed data of the resource attribute with the corresponding alarm event to obtain alarm association information. Thus, the source of the anomaly can be directly locked through the alarm association information, providing data preparation for subsequent alarm root cause analysis.
[0091] Optionally, before associating the changed resource attribute data with the corresponding alarm events to obtain alarm association information, such as... Figure 2 As shown, the system can detect whether changes to resource attribute data cause abnormal system operation (e.g., system crash after version upgrade); if changes to resource attribute data do not cause abnormal system operation, the system updates the configuration management database (e.g., CMDB, Configuration Management Database: used to provide static data such as departments and maintenance contacts) based on the changed data; if changes to resource attribute data cause abnormal system operation, the system determines the alarm event corresponding to the data change according to the preset alarm triggering rules.
[0092] In this embodiment of the application, the preset alarm triggering rules specifically include the following rules:
[0093] (1) Threshold rule: Ant state error lasts for more than 5 minutes; the rule is defined as follows:
[0094] rule("Ant deployment timeout alert"){
[0095] when{
[0096] resource.ant_status.runtime=="error"
[0097] &&duration(resource.ant_status.runtime)>"5m"
[0098] }
[0099] then{
[0100] triggerAlert("P1","Ant deployment timed out, please contact the person in charge")
[0101] }};
[0102] (2) Association rule: Deployment failure occurs within 30 minutes after version change;
[0103] (3) Composite rule: Deployment failure occurs within 30 minutes after version change.
[0104] It should be noted that by automatically detecting whether changes to resource attributes cause abnormal system operation, it can quickly identify whether changes have caused anomalies (such as performance degradation, service unavailability, etc.), avoid the problem from escalating, reduce the risk of human intervention, and reduce the risk of system failure due to misoperation or omission. It can take action before anomalies occur or in the early stages (such as after resource changes) to prevent the system from entering an unrecoverable failure state.
[0105] Step S105: The resource attribute cause-effect graph model calculates the PageRank value of the node related to the alarm association information using the improved PageRank algorithm, and determines the alarm root cause matching the alarm event based on the PageRank value.
[0106] In step S105, the PageRank value is used to reflect the importance of a node in the anomaly propagation network of the entire system.
[0107] In this embodiment, the system uses a pre-trained resource attribute causal graph model and an improved PageRank algorithm to calculate the PageRank values of nodes (e.g., root cause nodes, intermediate nodes, and symptom nodes) related to alarm association information. Based on the PageRank values, the system determines the root cause of the alarm that matches the alarm event. The PageRank value reflects the importance of a node in the entire system's anomaly propagation network. The higher the PageRank value of a node, the greater the probability that it will be identified as the root cause of the alarm event. Thus, the importance of a node in the anomaly propagation network (e.g., dependency library incompatibility issues) can be accurately assessed through the PageRank value, avoiding the limitations of traditional static analysis, significantly shortening the fault diagnosis time, and improving the efficiency of alarm root cause localization.
[0108] To improve the accuracy of alarm root cause identification, in this embodiment, the system can construct an initial causal graph model (e.g., a Bayesian network model) based on historical abnormal operation data in the fault database, and train the initial causal graph model using data from the sample dataset to obtain a resource attribute causal graph model. Specifically, see... Figure 3 FIG. 3 is a flowchart of a training process of a resource attribute causal graph model according to an embodiment of the present application. The process includes the following steps:
[0109] In step S31, historical abnormal operation data of the system is obtained from a fault database.
[0110] In step S32, a causal dependency relationship between a plurality of random variables is identified based on the historical abnormal operation data.
[0111] In step S32, the plurality of random variables at least include a fault type, a resource state, and a configuration change. In addition, the plurality of random variables can also include a network delay and the like.
[0112] In the embodiment of the present application, the system can extract a plurality of random variables such as a fault type, a resource state, a configuration change, and a network delay from historical abnormal operation data and operation and maintenance experience, and define a value space thereof. A specific example can be shown in Table 1. Then, the system can identify a causal dependency relationship between the plurality of random variables, so as to quickly track an abnormal propagation path according to the causal dependency relationship and avoid misjudgment of a traditional correlation analysis.
[0113] Table 1
[0114]
[0115] It should be noted that the operation and maintenance experience includes an understanding of a causal relationship between resource attributes (such as "server downtime → service unavailable") by an expert, which can be used as a prior constraint of a PC (Peter-Clark) algorithm to avoid missing a key causal relationship due to insufficient data or noise, thereby providing an accurate data basis for subsequent model construction.
[0116] In step S33, the causal dependency relationship is mapped to a directed acyclic graph to construct an initial causal graph model. A node of the initial causal graph model represents a random variable, and an edge of the initial causal graph model represents a causal dependency relationship.
[0117] In the embodiment of the present application, the system can identify a potential dependency relationship in historical abnormal operation data and operation and maintenance experience through a PC algorithm, and further determine a direction of an edge in combination with operation and maintenance experience (such as an implicit causal rule) to map the causal dependency relationship to a directed acyclic graph and construct an initial causal graph model.
[0118] Before constructing the initial causal graph model, the system can define a node of the initial causal graph model according to a plurality of random variables. A specific example can be shown in Table 2.
[0119] Table 2
[0120]
[0121] In Table 2, the root cause node represents a root cause that can cause a failure or a problem in the system, is the starting point of the causal dependency relationship, and is used to analyze the influence on the subsequent nodes as the source of the causal chain. The intermediate node represents an intermediate link between the root cause and the symptom, depicts an indirect path of the causal dependency relationship, and is used to reflect the process of the root cause indirectly causing the symptom through intermediate factors (such as dependency conflicts or resource exhaustion). The symptom node represents an observable failure performance or result in the system, is the end point of the causal chain, and is used to verify the accuracy of the root cause analysis as the final observed phenomenon.
[0122] Optionally, the system can also construct a conditional probability table (CPT) according to the historical abnormal operation data, operation and maintenance experience, and causal dependency relationship, and construct an initial causal graph model based on the conditional probability table and the directed acyclic graph. An example of the CPT of the Dependency_Conflict node can be as shown in Table 3.
[0123] Table 3
[0124]
[0125] In step S34, the sample data set is obtained.
[0126] In step S34, the sample data set includes a plurality of training samples and label data corresponding to the training samples, the training samples include historical alarm correlation information of the system, and the label data includes actual alarm root causes corresponding to the historical alarm correlation information.
[0127] In step S35, the data in the training sample set is input into the initial causal graph model to obtain a prediction result output by the initial causal graph model; the prediction result includes an alarm root cause corresponding to the historical alarm correlation information.
[0128] In the embodiment of the present application, the system inputs the data in the training sample set into the initial causal graph model, and the initial causal graph model can predict the alarm root cause corresponding to the historical alarm correlation information by using the improved PageRank algorithm. Specifically, referring to Figure 4 The figure is a flowchart of another training process of a resource attribute causal graph model provided by the embodiment of the present application, and the process includes the following steps:
[0129] In step S351, the data in the sample data set is input into the initial causal graph model, and the initial causal graph model calculates the PageRank value of the node related to the historical alarm correlation information by using the improved PageRank algorithm.
[0130] In step S351, the improved PageRank algorithm is used to quantify the importance of the node to convert the abnormal propagation path of the system abnormal operation into a calculable weight.
[0131] In the embodiment of the present application, the system can quantify the importance of the nodes related to the historical alarm associated information by the improved PageRank algorithm based on the initial causal graph model, and convert the abnormal propagation path of the system in abnormal operation into a calculable weight (i.e. PageRank value).
[0132] In step S352, the alarm root cause matching the historical alarm associated information is predicted based on the PageRank value, and the corresponding prediction result is generated.
[0133] In the embodiment of the present application, the system can predict the alarm root cause matching the historical alarm associated information based on the PageRank value of the initial causal graph model, and generate the corresponding prediction result. For example, the nodes related to the historical alarm associated information are Middleware_Version_Change, Dependency_Conflict and Deployment_Failure respectively, and the PageRank values are 0.9, 0.7 and 0.6 respectively. At this time, the system can determine Middleware_Version_Change as the alarm root cause according to the PageRank value.
[0134] In step S36, the parameters of the initial causal graph model are adjusted based on the difference between the alarm root cause corresponding to the historical alarm associated information and the actual alarm root cause, and the adjusted initial causal graph model is iteratively trained by using the training samples not used in the sample data set and the corresponding label data until the training termination condition is met, to obtain the resource attribute causal graph model.
[0135] It should be noted that the system can obtain a resource attribute causal graph model with higher recognition accuracy by adjusting the parameters of the initial causal graph model based on the difference between the alarm root cause corresponding to the historical alarm associated information and the actual alarm root cause, and iteratively training the adjusted initial causal graph model by using the training samples not used in the sample data set and the corresponding label data, thereby improving the recognition accuracy of the alarm root cause and significantly improving the operation and maintenance efficiency of the IT resource.
[0136] In step S106, the corresponding alarm strategy is generated according to the alarm root cause.
[0137] In the embodiment of the present application, the system can determine the alarm level corresponding to the alarm root cause according to the preset level allocation rule, and determine the alarm strategy corresponding to the alarm event based on the alarm level, wherein the alarm strategy includes the alarm response time and the alarm notification method, thereby realizing the operation and maintenance closed loop from abnormal perception to root cause repair of the IT resource, and significantly improving the operation and maintenance efficiency of the IT resource. The specific example can be shown in Table 4.
[0138] Table 4
[0139]
[0140]
[0141] In one alternative embodiment, the system can merge repeated alarms for the same resource within 10 minutes into a single alarm, and prioritize root cause alarms over derived alarms (e.g., only notify "version incompatibility" instead of all downstream anomalies). The system can also automatically go silent during maintenance windows (e.g., every Tuesday from 03:00 to 04:00).
[0142] In another alternative embodiment, the system can integrate scattered data such as departments, middleware versions, and Ant status based on a unified resource description model (JSON Schema + OWL semantic extension) to construct a global resource graph, supporting contextual association analysis. For example, a middleware version upgrade event can be automatically associated with its department, maintenance personnel, and historical change records to form a complete event chain, thereby enabling real-time awareness of data changes in resource attributes.
[0143] The IT resource management method provided in this application achieves real-time collection of multi-dimensional attribute information corresponding to enterprise IT resources, avoiding the system response delay problem caused by the traditional polling mechanism for detecting IT resources, and significantly reducing system response latency; it achieves the use of hash values to replace the original data for subsequent comparison, greatly reducing storage space occupation, while reducing computational complexity and improving data processing efficiency; it achieves the ability to determine whether resource attributes have changed without full data comparison, significantly shortening change detection time, and thus achieving second-level or even millisecond-level change perception capability, meeting the stringent real-time requirements of modern IT systems; it achieves the ability to directly pinpoint the source of anomalies through alarm correlation information, providing data preparation for subsequent alarm root cause analysis; it achieves the ability to accurately assess the importance of nodes in the anomaly propagation network (such as dependency library incompatibility issues) through PageRank values, avoiding the limitations of traditional static analysis, significantly shortening fault troubleshooting time, and improving the efficiency of alarm root cause location; finally, it generates corresponding alarm strategies based on alarm root causes, realizing a closed-loop operation and maintenance system for IT resources from anomaly perception to root cause repair, significantly improving the efficiency of IT resource operation and maintenance management.
[0144] Device Examples
[0145] This application provides an IT resource management device, wherein... Figure 5 This is a schematic diagram of a battery heating control device provided in an embodiment of this application, as shown below. Figure 5 As shown, the device includes: a data acquisition module 11, a calculation module 12, a comparison module 13, an association module 14, a determination module 15, and a generation module 16. From Figure 5 The connection relationship between several modules can be seen.
[0146] The collection module 11 is configured to collect multi-dimensional attribute information corresponding to the IT resources of the enterprise.
[0147] The computing module 12 is configured to convert the multi-dimensional attribute information into a JSON string and calculate a hash value corresponding to each resource attribute based on the JSON string.
[0148] The comparison module 13 is configured to compare each hash value with a corresponding historical hash value to obtain a comparison result.
[0149] The association module 14 is configured to associate changed data of the resource attribute with a corresponding alarm event to obtain alarm association information if the comparison result indicates that the data corresponding to the resource attribute has changed.
[0150] The determination module 15 is configured to calculate a PageRank value of a node related to the alarm association information by an improved PageRank algorithm based on a resource attribute causal graph model, and determine an alarm root cause matching the alarm event according to the PageRank value.
[0151] The generation module 16 is configured to generate a corresponding alarm strategy according to the alarm root cause.
[0152] Optionally, the comparison module includes:
[0153] The comparison unit is configured to compare whether each hash value and the corresponding historical hash value are the same.
[0154] The first determination unit is configured to determine that the data corresponding to the resource attribute has not changed and generate a corresponding comparison result if the hash value and the corresponding historical hash value are the same.
[0155] The second determination unit is configured to determine that the data corresponding to the resource attribute has changed and generate a corresponding comparison result if the hash value and the corresponding historical hash value are not the same.
[0156] Optionally, the IT resource management device further includes:
[0157] The detection module is configured to detect whether the data change of the resource attribute causes a system abnormal operation before associating the changed data of the resource attribute with the corresponding alarm event to obtain the alarm association information.
[0158] an updating module configured to update the configuration management database based on the changed data if the change of the data of the resource attribute does not cause abnormal operation of the system;
[0159] an event determining module configured to determine an alarm event corresponding to the change of the data according to a preset alarm triggering rule if the change of the data of the resource attribute causes abnormal operation of the system.
[0160] Optionally, the IT resource management apparatus further comprises:
[0161] an extracting module configured to acquire historical abnormal operation data of the system from a fault database;
[0162] an identifying module configured to identify a causal dependency relationship between a plurality of random variables based on the historical abnormal operation data; the plurality of random variables at least include a fault type, a resource state and a configuration change;
[0163] a constructing module configured to map the causal dependency relationship into a directed acyclic graph to construct an initial causal graph model; a node of the initial causal graph model represents a random variable, and an edge of the initial causal graph model represents the causal dependency relationship;
[0164] an acquiring module configured to acquire a sample data set; the sample data set includes a plurality of training samples and label data corresponding to the training samples, the training samples include historical alarm correlation information of the system, and the label data includes an actual alarm root cause corresponding to the historical alarm correlation information;
[0165] a predicting module configured to input data in the training sample set into the initial causal graph model to obtain a prediction result output by the initial causal graph model; the prediction result includes an alarm root cause corresponding to the historical alarm correlation information;
[0166] an iterative training module configured to adjust parameters of the initial causal graph model based on a difference between the alarm root cause corresponding to the historical alarm correlation information and the actual alarm root cause, and perform iterative training on the adjusted initial causal graph model through training samples and corresponding label data in the sample data set that are not used, until a training stop condition is met, to obtain the resource attribute causal graph model.
[0167] Optionally, the predicting module comprises:
[0168] a calculating unit configured to input data in the sample data set into the initial causal graph model, and calculate a PageRank value of a node related to the historical alarm correlation information by the initial causal graph model through an improved PageRank algorithm; the improved PageRank algorithm is used to quantify an importance degree of the node to convert an abnormal propagation path when the system is abnormally operated into a calculable weight;
[0169] The generating unit is configured to predict an alarm root cause matched with the historical alarm associated information based on the PageRank value, and generate a corresponding prediction result.
[0170] Optionally, the generating module comprises:
[0171] The first determining unit is configured to determine an alarm level corresponding to the alarm root cause according to a preset level allocation rule.
[0172] The second determining unit is configured to determine an alarm strategy corresponding to the alarm event based on the alarm level, wherein the alarm strategy comprises an alarm response time and an alarm notification manner.
[0173] The storage medium embodiment
[0174] The computer readable storage medium provided by the embodiment of the present application stores a program, wherein when the program is executed by a processor, part or all steps of the IT resource management method introduced in the foregoing method embodiment are implemented. The storage medium can be a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0175] The processor embodiment
[0176] The processor provided by the embodiment of the present application is used to run a program, wherein when the program is running, part or all steps of the IT resource management method introduced in the foregoing method embodiment are executed.
[0177] It should be noted that each embodiment in the present specification is described in a progressive manner, and the same and similar parts of each embodiment can be referred to each other. Each embodiment focuses on the difference from other embodiments. Especially, the device embodiment is basically similar to the method embodiment, so the description is relatively simple, and the related parts can be referred to the part of the method embodiment. The device embodiment described above is only illustrative, and the units described as separate components can be or can not be physically separated, and the components indicated as units can be or can not be physical components, that is, they can be located in one place or distributed on multiple network components. According to actual needs, part or all of the modules can be selected to achieve the purpose of the embodiment. Those skilled in the art can understand and implement without creative labor.
[0178] The above merely provides one specific embodiment of the present application, but the protection scope of the present application is not limited thereto, any changes or replacements within the technical scope disclosed by the present application, which can be easily thought by any person skilled in the art, should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.< / resourceid>
Claims
1. An IT resource management method, characterized by, The method comprises the following steps: collecting multi-dimensional attribute information corresponding to IT resources of an enterprise; the multi-dimensional attribute information comprises data corresponding to resource attributes of multiple dimensions respectively; converting the multi-dimensional attribute information into a JSON string, and calculating hash values corresponding to the resource attributes based on the JSON string; comparing each hash value with a corresponding historical hash value to obtain a comparison result; the comparison result is used to indicate whether the data corresponding to the resource attribute has changed, and the historical hash value is a hash value calculated by data updated last time on the resource attribute; if the comparison result indicates that the data corresponding to the resource attribute has changed, associating the changed data of the resource attribute with a corresponding alarm event to obtain alarm association information; the alarm event is a system abnormal operation event caused by the data change of the resource attribute; calculating the PageRank values of nodes related to the alarm association information by an improved PageRank algorithm based on a resource attribute causal graph model, and determining an alarm root cause matching the alarm event according to the PageRank values; the PageRank values are used to reflect the importance of the nodes in the entire abnormal propagation network of the system; generating a corresponding alarm strategy according to the alarm root cause.
2. The method of claim 1, wherein, The comparison of each hash value with a corresponding historical hash value to obtain a comparison result comprises the following steps: comparing whether each hash value and a corresponding historical hash value are the same; if the hash value and the corresponding historical hash value are the same, determining that the data corresponding to the resource attribute has not changed, and generating a corresponding comparison result; if the hash value and the corresponding historical hash value are not the same, determining that the data corresponding to the resource attribute has changed, and generating a corresponding comparison result.
3. The method of claim 1, wherein, Before associating the changed data of the resource attribute with a corresponding alarm event to obtain alarm association information, the method further comprises the following steps: detecting whether the data change of the resource attribute causes the system abnormal operation; if the data change of the resource attribute does not cause the system abnormal operation, updating a configuration management database based on the changed data; if the data change of the resource attribute causes the system abnormal operation, determining an alarm event corresponding to the data change according to a preset alarm triggering rule.
4. The method of claim 1, wherein, The resource attribute causal graph model is obtained by the following process: obtaining historical abnormal operation data of the system from a fault database; identifying causal dependency relationships between multiple random variables based on the historical abnormal operation data; the multiple random variables at least include fault types, resource states, and configuration changes; mapping the causal dependency relationships into a directed acyclic graph to construct an initial causal graph model; the nodes of the initial causal graph model represent the random variables, and the edges of the initial causal graph model represent the causal dependency relationships; acquire a sample data set; the sample data set includes a plurality of training samples and label data corresponding to the training samples, the training samples include historical alarm association information of the system, and the label data includes an actual alarm root cause corresponding to the historical alarm association information; input data in the training sample set into the initial causal graph model to obtain a prediction result output by the initial causal graph model; the prediction result includes an alarm root cause corresponding to the historical alarm association information; based on a difference between the alarm root cause corresponding to the historical alarm association information and the actual alarm root cause, adjust parameters of the initial causal graph model, and iteratively train the adjusted initial causal graph model through training samples and corresponding label data in the sample data set that are not used, until a training stop condition is met, to obtain the resource attribute causal graph model.
5. The method of claim 4, wherein, The inputting of the data in the training sample set into the initial causal graph model to obtain the prediction result output by the initial causal graph model comprises: inputting data in the sample data set into the initial causal graph model, and calculating a PageRank value of a node related to the historical alarm association information by the initial causal graph model through an improved PageRank algorithm; the improved PageRank algorithm is used to quantify the importance of the node, so as to convert an abnormal propagation path of the system during abnormal operation into a calculable weight; predicting an alarm root cause matching the historical alarm association information based on the PageRank value, and generating a corresponding prediction result.
6. The method as claimed in claim 1, wherein, The generation of the alarm strategy corresponding to the alarm root cause comprises: determining an alarm level corresponding to the alarm root cause according to a preset level allocation rule; determining an alarm strategy corresponding to the alarm event based on the alarm level; the alarm strategy includes an alarm response time and an alarm notification mode.
7. An IT resource management apparatus characterized by comprising: comprise: a collection module configured to collect multi-dimensional attribute information corresponding to IT resources of an enterprise; the multi-dimensional attribute information includes data corresponding to resource attributes of multiple dimensions respectively; a calculation module configured to convert the multi-dimensional attribute information into a JSON string, and calculate hash values corresponding to the resource attributes based on the JSON string; a comparison module configured to compare each hash value with a corresponding historical hash value to obtain a comparison result; the comparison result is used to indicate whether the data corresponding to the resource attribute has changed, and the historical hash value is a hash value calculated by data updated last time on the resource attribute; an association module configured to, if the comparison result indicates that the data corresponding to the resource attribute has changed, associate changed data of the resource attribute with a corresponding alarm event to obtain alarm association information; the alarm event is a system abnormal operation event caused by the data change of the resource attribute; The determining module is configured to calculate a PageRank value of a node related to the alarm-associated information by using an improved PageRank algorithm of the resource attribute causal graph model, and determine an alarm root cause matched with the alarm event according to the PageRank value; the PageRank value is used to reflect an importance degree of the node in an abnormal propagation network of the whole system. The generating module is configured to generate a corresponding alarm strategy according to the alarm root cause.
8. The apparatus of claim 7, wherein, The comparing module comprises: A comparing unit is configured to compare whether each hash value is same as a corresponding historical hash value; A first determining unit is configured to determine that data corresponding to the resource attribute has not changed, and generate a corresponding comparison result, if the hash value is same as the corresponding historical hash value; A second determining unit is configured to determine that data corresponding to the resource attribute has changed, and generate a corresponding comparison result, if the hash value is not same as the corresponding historical hash value.
9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and when the computer program is run by the processor, the IT resource management method in any one of claims 1-6 is implemented.
10. A processor, comprising: The computer program is used to run, and when the computer program is run, the IT resource management method in any one of claims 1-6 is executed.