Enterprise operation and maintenance management platform

By designing an enterprise operation and maintenance management platform, using the combination of information collection, monitoring and control modules, the problem of low efficiency of traditional operation and maintenance management is solved, and efficient and intelligent operation and maintenance management is achieved.

CN119473804BActive Publication Date: 2025-05-16SHIJIAZHUANG MAITEDA ELECTRONICS TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510072091.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-16
Estimated Expiration
2045-01-17

AI Technical Summary

Technical Problem

Traditional operation and maintenance management methods are inefficient, difficult to detect network security problems, and cannot meet the high requirements of modern enterprises for information equipment management.

Method used

An enterprise operation and maintenance management platform is designed, including information collection module, monitoring module and control module. The information collection module collects the status information data of the information equipment in real time. The monitoring module uses the target decision tree model to process these data, outputs the target monitoring results, and the control module formulates operation and maintenance strategies based on the monitoring results.

Benefits of technology

It improves the efficiency of operation and maintenance management, reduces the potential risks caused by misjudgment or misjudgment, realizes the intelligence and automation of operation and maintenance work, reduces the work pressure of operation and maintenance personnel, and improves the level and quality of operation and maintenance management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119473804B_ABST
    Figure CN119473804B_ABST
Patent Text Reader

Abstract

The present disclosure provides an enterprise operation and maintenance management platform, which belongs to the field of operation and maintenance management technology. The enterprise operation and maintenance management platform includes: an information collection module, which is used to collect status information data of information equipment; a monitoring module, which is connected to the information collection module, and is used to input the status information data into a target decision tree model to obtain a target monitoring result of the information equipment; and a control module, which is connected to the monitoring module and is used to determine a target operation and maintenance strategy according to the target monitoring result. The present disclosure can solve the problem of low efficiency of operation and maintenance management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of operation and maintenance management technology, and in particular to an enterprise operation and maintenance management platform. Background Art

[0002] With the continuous improvement of enterprise informatization, intelligent operation and maintenance management systems have emerged. There are many resources in the internal architecture of an enterprise, such as servers, network equipment, and operation and maintenance processes. Traditional operation and maintenance management methods rely on manual inspections, which are not only inefficient, but also difficult to detect network security issues, and cannot meet the high requirements of modern enterprises for information equipment management.

[0003] Therefore, there is an urgent need for an efficient and controllable enterprise operation and maintenance management platform. Summary of the invention

[0004] The embodiments of the present disclosure provide an enterprise operation and maintenance management platform to solve the problem of low operation and maintenance management efficiency.

[0005] The present disclosure provides an enterprise operation and maintenance management platform, including:

[0006] An information collection module is used to collect status information data of information equipment;

[0007] The monitoring module is connected to the information collection module and is used to input the status information data into the target decision tree model to obtain the target monitoring result of the information equipment;

[0008] The control module is connected to the monitoring module and is used to determine the target operation and maintenance strategy according to the target monitoring results.

[0009] In an exemplary embodiment of the present disclosure, the enterprise operation and maintenance management platform further includes:

[0010] An operation and maintenance module is used to determine operation and maintenance control parameters based on the target operation and maintenance strategy, and to maintain information equipment based on the operation and maintenance control parameters;

[0011] The operation and maintenance module is connected to the control module.

[0012] In an exemplary embodiment of the present disclosure, the monitoring module includes:

[0013] A training unit, used to train a decision tree model based on the historical state information data and the historical monitoring results corresponding to the historical state information data to obtain a target decision tree model;

[0014] A monitoring unit, used to determine a target monitoring result of the information device based on a target decision tree model;

[0015] The monitoring unit is connected with the information collection module, the training unit and the control module respectively.

[0016] In an exemplary embodiment of the present disclosure, the training unit is specifically configured to:

[0017] determining a plurality of state characteristics of the historical state information data;

[0018] The historical state information data is divided based on multiple state features to obtain multiple state information subsets; the multiple state features correspond to the multiple state information subsets one by one;

[0019] Establishing a recursive strategy based on each state feature and a subset of state information corresponding to each state feature;

[0020] The target decision tree model is determined based on a recursive strategy.

[0021] In an exemplary embodiment of the present disclosure, the control module is specifically configured to:

[0022] Process the target monitoring results and status information data based on the reinforcement learning algorithm to determine the target operation and maintenance strategy;

[0023] The control module is connected to the information collection module.

[0024] In an exemplary embodiment of the present disclosure, the control module is further configured to:

[0025] Determine the monitoring status of the information equipment based on the target monitoring results;

[0026] Determine the value function of the reinforcement learning algorithm based on the monitoring state and state information data;

[0027] Determine the target operation and maintenance strategy based on the feedback of the value function.

[0028] In an exemplary embodiment of the present disclosure, the enterprise operation and maintenance management platform further includes:

[0029] Security management module, used to monitor the security status of information equipment;

[0030] The security management module is connected to the control module.

[0031] In an exemplary embodiment of the present disclosure, the security management module is specifically configured to:

[0032] Determine multiple evaluation factors of information equipment and evaluation levels corresponding to the multiple evaluation factors;

[0033] Determine the membership degree of the evaluation level corresponding to each evaluation factor based on each evaluation factor;

[0034] Determine the fuzzy evaluation matrix according to the membership degree of the evaluation level corresponding to each evaluation factor;

[0035] Determine a weight set according to the weight of each evaluation factor;

[0036] Based on the fuzzy evaluation matrix and the weight set, the security evaluation level is determined, and the security evaluation level is the security status of the information equipment.

[0037] In an exemplary embodiment of the present disclosure, the enterprise operation and maintenance management platform further includes: a storage module;

[0038] The storage module is connected to the information collection module and the monitoring module respectively;

[0039] The storage module is used to store the status information data of the information equipment and the monitoring results of the monitoring module.

[0040] In an exemplary embodiment of the present disclosure, the enterprise operation and maintenance management platform further includes: an alarm module;

[0041] The alarm module is connected with the monitoring module;

[0042] The alarm module is used to generate an alarm in response to a fault detected by the monitoring module.

[0043] The beneficial effects of the enterprise operation and maintenance management platform provided by the embodiments of the present disclosure are:

[0044] Through the information collection module, the operation and maintenance management platform of the present disclosure can collect the status information data of information equipment in real time and accurately, providing a basis for subsequent monitoring and analysis. Secondly, the monitoring module uses the target decision tree model to process the status information data, which can efficiently and accurately judge the operating status of the information equipment, thereby outputting reliable target monitoring results, which not only improves the efficiency of operation and maintenance work, but also reduces the potential risks caused by misjudgment or missed judgment. Finally, the control module formulates the target operation and maintenance strategy according to the target monitoring results, realizes the intelligence and automation of operation and maintenance work, not only reduces the work pressure of operation and maintenance personnel, but also improves the level and quality of operation and maintenance management. Therefore, the present disclosure can improve the efficiency of operation and maintenance management. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0046] Figure 1 It is a structural diagram of an enterprise operation and maintenance management platform provided by an embodiment of the present disclosure;

[0047] Figure 2 It is a structural diagram of another enterprise operation and maintenance management platform provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0048] In order to enable people in the technical field to better understand the present solution, the technical solution in the embodiment of the present solution will be clearly described below in conjunction with the drawings in the embodiment of the present solution. Obviously, the described embodiment is an embodiment of a part of the present solution, not all of the embodiments. Based on the embodiments in the present solution, all other embodiments obtained by ordinary technicians in the field without creative work should fall within the scope of protection of the present solution.

[0049] The term "including" and any other variations in the specification and claims of this solution and the above drawings mean "including but not limited to", and is intended to cover non-exclusive inclusions and is not limited to the examples listed in the text. In addition, the terms "first" and "second" are used to distinguish different objects, not to describe a specific order.

[0050] The following is a detailed description of the implementation of the present disclosure in conjunction with the specific drawings:

[0051] Figure 1 This is a schematic diagram of the structure of an enterprise operation and maintenance management platform provided by an embodiment of the present disclosure. Figure 1 , the enterprise operation and maintenance management platform includes:

[0052] The information collection module 101 is used to collect status information data of information equipment;

[0053] The monitoring module 102 is connected to the information collection module 101 and is used to input the status information data into the target decision tree model to obtain the target monitoring result of the information device;

[0054] The control module 103 is connected to the monitoring module 102 and is used to determine a target operation and maintenance strategy according to the target monitoring result.

[0055] In this embodiment, the information collection module 101 is used to collect various data related to information equipment, and its purpose is to obtain various raw data that can reflect the operating status of information equipment for subsequent analysis and processing. Information equipment is equipment involved in the enterprise information operation system and has certain information processing or transmission functions. For example, information equipment can be internal servers, network switches, routers, storage devices, and various office computers, printers, etc.

[0056] Status information data reflects the status of information equipment, and may include hardware indicators, software running status, network connection status, and other aspects of information equipment. Among them, hardware indicators of information equipment may include CPU usage, memory usage, disk remaining space, temperature of each device, etc.; software running status may include operating system process status, whether the application is running normally, and whether there is any error information, etc.; network connection status may include network bandwidth utilization, network packet loss rate, and connection stability, etc.

[0057] For example, in an operation and maintenance management platform of an e-commerce enterprise, the information collection module 101 can regularly obtain hardware performance indicator data such as CPU usage and memory usage from multiple servers of the enterprise through the Simple Network Management Protocol (SNMP); use Windows Management Instrumentation (WMI) to collect software-related information such as the number of running processes of the operating system and whether there are suspicious background programs on office computers; and use network monitoring tools to collect network connection status data such as traffic data of each port of the network switch and network packet loss rate, and summarize these data of different dimensions as the basis for subsequent analysis.

[0058] The monitoring module 102 is a functional module that performs further analysis and judgment based on the data collected by the information acquisition module 101. It can connect with the information acquisition module 101, receive the collected data, and use the target decision tree model to analyze the status information data, and then obtain the monitoring results about the information equipment, so as to judge whether the operation status of the information equipment is normal, whether there are potential problems, and whether maintenance is required.

[0059] The monitoring module 102 can obtain the status information data collected by the information collection module 101 in real time, and provide the above data as input to the pre-built target decision tree model. The target decision tree model can determine whether the information equipment is currently in different status categories such as normal operation, performance warning, and failure, and the final output is the target monitoring result of the information equipment, which can intuitively reflect the current status of the information equipment.

[0060] For example, assume that an enterprise has built a decision tree model (i.e., a target decision tree model) for monitoring server status, which uses the server's CPU usage, memory usage, disk I / O read / write speed, etc. as input features. After the monitoring module 102 receives the relevant data of a server from the information collection module 101 (such as CPU usage of 85%, memory usage of 70%, and normal disk I / O read / write speed), it inputs these data into the decision tree model, and after the node judgment inside the model, for example, if the CPU usage is greater than 80%, it further judges the memory usage and other conditions, and finally outputs the target monitoring result that the server is in "performance warning", indicating that the current resource usage of the server is close to the critical value, and the operation and maintenance personnel need to pay attention and may take corresponding measures.

[0061] The control module 103 is used to determine the operation and maintenance strategy for the information equipment based on the target monitoring results obtained by the monitoring module 102, and decide what specific operations to take in the future to ensure the normal operation of the information equipment, optimize performance or solve problems that arise. The control module 103 can obtain the target monitoring results output by the monitoring module 102 in real time. Then, based on the target monitoring results, combined with some rules, policy libraries or empirical knowledge pre-set by the enterprise, determine the specific target operation and maintenance strategy. For example, if the target monitoring results show that the information equipment is operating normally, maintain the status quo or perform some conventional inspection strategies; if the target monitoring results indicate that the information equipment has a fault, a corresponding fault repair strategy can be formulated, such as restarting the equipment, calling backup resources, arranging maintenance personnel to visit, etc.

[0062] When the monitoring module 102 outputs the target monitoring result that a certain network switch has abnormal port traffic, the control module 103 analyzes and judges according to the above results and the pre-set policy library. If the port traffic occasionally exceeds the threshold but quickly returns to normal, the target operation and maintenance strategy that can be determined is to first record the abnormal situation and continue to observe for a period of time; if the port traffic is abnormal for a long time and affects the stability of the network connection, the control module 103 will determine a series of target operation and maintenance strategies such as reconfiguring port parameters, checking whether the connection line is damaged, and replacing the network port module when necessary, and convey these strategies to the relevant operation and maintenance execution personnel or the automated operation and maintenance system for execution.

[0063] It can be concluded from the above that, through the information collection module 101, the operation and maintenance management platform of the present disclosure can collect the status information data of information equipment in real time and accurately, providing a basis for subsequent monitoring and analysis. Secondly, the monitoring module 102 uses the target decision tree model to process the status information data, which can efficiently and accurately judge the operating status of the information equipment, thereby outputting reliable target monitoring results, which not only improves the efficiency of operation and maintenance work, but also reduces the potential risks caused by misjudgment or missed judgment. Finally, the control module 103 formulates the target operation and maintenance strategy according to the target monitoring results, realizes the intelligence and automation of operation and maintenance work, not only reduces the work pressure of operation and maintenance personnel, but also improves the level and quality of operation and maintenance management. Therefore, the present disclosure can improve the efficiency of operation and maintenance management.

[0064] In one embodiment of the present disclosure, reference Figure 2 , the enterprise operation and maintenance management platform also includes:

[0065] The operation and maintenance module 104 is used to determine the operation and maintenance control parameters based on the target operation and maintenance strategy, and to maintain the information equipment based on the operation and maintenance control parameters;

[0066] The operation and maintenance module 104 is connected to the control module 103 .

[0067] In this embodiment, the operation and maintenance module 104 is used to perform actual operation and maintenance of the target operation and maintenance strategy, receive the target operation and maintenance strategy instructions from the control module 103, and determine specific operation and maintenance control parameters based on the above instructions, and then perform corresponding operation and maintenance operations to ensure the normal operation of the information equipment, performance optimization or troubleshooting.

[0068] According to the target operation and maintenance strategy output by the control module 103, the specific operation control details are further clarified, that is, the operation and maintenance control parameters are determined, which determine the specific mode and degree of the operation and maintenance operation. For example, if the target operation and maintenance strategy is to adjust the CPU frequency of the server, the operation and maintenance control parameters may include the specific frequency value to be adjusted, the time point of adjustment, the allowable error range during the adjustment process, etc. Afterwards, the operation and maintenance module 104 performs the actual operation and maintenance operation according to the determined parameters and performs corresponding processing on the information device.

[0069] For example, in an Internet company, its server cluster bears a large amount of business traffic. When the control module 103 determines the target operation and maintenance strategy based on the monitoring results to optimize the resources of certain servers with overload, the operation and maintenance module 104 starts working after receiving the strategy. First, the operation and maintenance control parameters are determined. For example, for the optimization of server memory resources, it is decided to adjust the memory cache mechanism to prioritize caching high-frequency access data, and at the same time adjust the memory pre-allocation threshold from the current 80% to 70% to avoid excessive occupation of memory resources. After determining the above parameters, the operation and maintenance module 104 connects to the target server through an automated script or remote management tool, and adjusts the memory cache mechanism and pre-allocation threshold according to the set parameters to complete the operation and maintenance optimization of the server.

[0070] It can be concluded from the above that the operation and maintenance module 104 can accurately determine the operation and maintenance control parameters based on the target operation and maintenance strategy provided by the control module 103, and automatically perform the operation and maintenance tasks. This embodiment not only improves the pertinence and accuracy of the operation and maintenance work, but also realizes the automation and intelligence of the operation and maintenance process. In this embodiment, through the real-time intervention and precise control of the operation and maintenance module 104, the enterprise can better maintain the stable operation of information equipment, reduce the occurrence of failures and downtime, and thus ensure the continuity and stability of the business.

[0071] In one embodiment of the present disclosure, reference Figure 2 , the monitoring module 102 includes:

[0072] A training unit 201 is used to train a decision tree model based on historical state information data and historical monitoring results corresponding to the historical state information data to obtain a target decision tree model;

[0073] A monitoring unit 202, configured to determine a target monitoring result of the information device based on a target decision tree model;

[0074] The monitoring unit 202 is connected to the information collection module 101 , the training unit 201 and the control module 103 respectively.

[0075] In this embodiment, the training unit 201 learns a large amount of historical state information data and corresponding historical monitoring results, so that the decision tree model can capture the rules and patterns in the historical state information data, thereby being able to accurately monitor and judge the new state information data. The training unit 201 uses historical data to train the decision tree model, and can output the trained decision tree model to the monitoring unit 202. The historical state information data includes various state information of the information device at different time points in the past, such as the CPU usage rate, memory occupancy rate, disk read and write rate and other data of the server in the past, as well as the flow rate, packet loss rate and other information of the network device. The historical monitoring result is the judgment result obtained by analyzing the above state information at that time, such as whether the information device is operating normally, whether there are performance problems or failures, etc. The training unit 201 inputs the above historical data into the decision tree model, and through a certain iterative process, continuously adjusts the structure and parameters of the decision tree model, so that the model can accurately predict the corresponding monitoring results based on the input state information data, and finally obtains the target decision tree model after evaluation. The target decision tree model is a model that can be used for actual information equipment monitoring after training.

[0076] For example, taking the network equipment operation and maintenance of a telecommunications operator as an example, the training unit 201 collects historical status information data of network routers in the past year, including port traffic, CPU usage, memory usage, network delay and other data at different time periods every day. At the same time, the corresponding historical monitoring results are also collected, such as network congestion caused by excessive traffic in certain time periods, and the monitoring result is that the network performance is reduced; the equipment operates normally in certain time periods, and the monitoring result is normal at this time. The training unit 201 organizes the above data in a certain format and inputs it into the decision tree module for training. During the training process, the decision tree model can automatically generate a series of judgment nodes and branches according to the characteristics and distribution of the status information data. For example, first, the first layer of nodes is divided according to whether the port traffic exceeds a certain threshold. If it exceeds the threshold, it is further judged whether the CPU usage is too high. Through multiple iterations and optimizations, a target decision tree model that can accurately judge the current state of the network router is finally obtained. For example, when the status information data such as the port traffic and CPU usage of the current router is input, the model can accurately judge whether the network is in a normal operating state, whether there is a potential performance risk, etc.

[0077] The monitoring unit 202 can obtain the real-time status information data of the information device from the information acquisition module 101, and then input the above data into the trained target decision tree model. After the judgment of the model, the target monitoring result about the information device is output, and the result is output to the control module 103 so that the corresponding operation and maintenance strategy can be adopted later. The monitoring unit 202 uses the target decision tree model to monitor the status of the information device. When the information acquisition module 101 transmits the real-time status information data of the information device, the monitoring unit 202 uses the above data as the input of the target decision tree model. The target decision tree model analyzes and judges the input status information data according to the patterns and rules learned in the previous training, and finally outputs a judgment result about the current status of the information device, that is, the target monitoring result.

[0078] It can be concluded from the above that this embodiment can accurately construct a target decision tree model by using historical data to train the decision tree model through the training unit 201, thereby improving the accuracy and reliability of monitoring. Secondly, the monitoring unit 202 quickly determines the target monitoring results of the information device based on the target decision tree model, thereby realizing an efficient and real-time monitoring function.

[0079] In one embodiment of the present disclosure, reference Figure 2 , the training unit 201 is specifically used for:

[0080] determining a plurality of state characteristics of the historical state information data;

[0081] The historical state information data is divided based on multiple state features to obtain multiple state information subsets; the multiple state features correspond to the multiple state information subsets one by one;

[0082] Establishing a recursive strategy based on each state feature and a subset of state information corresponding to each state feature;

[0083] The target decision tree model is determined based on a recursive strategy.

[0084] In this embodiment, the state characteristics are various indicators that can reflect the operating status of information equipment, such as the memory usage characteristics of the server, the traffic characteristics of the network equipment, etc. The above characteristics are key elements for analyzing and judging the status of information equipment. Different state characteristics can reflect the current operation of information equipment from different angles. The state information subset is a partial data set obtained by dividing the historical state information data according to the state characteristics, that is, based on the different values ​​of a certain state characteristic, the overall historical state information data is divided into different subsets, and the data in each subset has similarity under the state characteristic.

[0085] The recursive strategy is used to recursively determine the node division, growth and final structure of the tree during the decision tree construction process. The recursive strategy selects the best state features as the basis for node division according to certain rules, and continues to recursively construct child nodes based on the divided data subsets until the stopping condition is met, thereby constructing a complete decision tree model.

[0086] From the large amount of historical status information data collected by information devices, various attributes that are important for judging the status of information devices are selected as status features. This feature will serve as the basis for subsequent data division and decision tree construction. For example, for the historical operation data of an enterprise's server, CPU usage, memory usage, disk space remaining, average response time, etc. can be determined as status features. By analyzing the values ​​of the above different features, we can understand whether the server's past operating status was normal, busy, or faulty.

[0087] For each selected state feature, the overall historical state information data is divided into different parts according to its different values ​​or value ranges, and each part is a state information subset. And each state feature has its corresponding divided state information subset, that is, the different values ​​of a state feature determine the composition of the corresponding subset. For example, taking the CPU usage state feature as an example, if the CPU usage is divided according to the three value ranges of less than 30% (indicating a low load state), 30%-70% (indicating a normal load state), and greater than 70% (indicating a high load state), then the historical state information data can be divided into three subsets, each containing data records within the corresponding CPU usage range. Other state features are also similarly operated according to their respective division rules to generate their respective corresponding subsets.

[0088] When building a decision tree, it is necessary to determine how to select the best feature for node division based on each state feature and its corresponding state information subset, as well as how to continue to build subnodes based on the divided subsets. This process is achieved by establishing a recursive strategy. The recursive strategy involves some measurement indicators, such as information gain, information gain ratio, or Gini index. By calculating the values ​​of these indicators for each state feature on the corresponding state information subset, it is determined which state feature is the best basis for dividing the current node, and then the various levels of the decision tree are continuously recursively constructed according to this rule. For example, the information gain corresponding to each state feature is calculated, and the state feature with the largest information gain is selected as the division feature of the root node. Then, for the divided subset, the information gain is repeatedly calculated in the remaining state features, and the best feature is selected to divide the subnodes, and so on.

[0089] This embodiment gradually builds a complete decision tree by continuously performing feature selection, data set division, and sub-node construction according to the rules specified by the recursive strategy. The final decision tree that can be used to accurately monitor and judge the status of information equipment is the target decision tree model. After learning and training historical data, the model has mastered the intrinsic relationship between different state features and the actual operating status of the equipment, so that it can effectively analyze and judge the newly input status information data.

[0090] Specifically, the construction process of the recursive strategy is as follows:

[0091] First, establish the Gini index. The Gini index is used to measure the impurity of a data set. The smaller the Gini index, the higher the purity of the data set.

[0092] Assume that the historical status information data is , the state characteristics included are , the total number of historical status information data is , the number of samples belonging to the state feature category is , then the Gini index of historical state information data is The calculation formula is as follows:

[0093]

[0094] in, The state feature features (categories), .

[0095] Second, establish the conditional Gini index.

[0096] Set state characteristics have Different value ranges , according to the state characteristics The value of Divide into State feature subset ,in Status Features The value is Then in the state feature under conditions The conditional Gini index The calculation formula is:

[0097]

[0098] Third, set the dynamic weight adjustment factor. In order to consider the correlation between different state features and the dynamic changes of state information data, the dynamic weight adjustment factor is introduced. For each state feature Each value of , weight adjustment factor Calculated according to the following formula:

[0099] ;

[0100] ;

[0101] ;

[0102] ;

[0103] in, Indicates state characteristics and other status characteristics The correlation coefficient of ; Indicates state characteristics The value of Indicates other status characteristics The value of Indicates state characteristics Value The mean of Indicates other status characteristics Value The mean of and Represents the adjustment parameter, which is used to control the amplitude and sensitivity of weight adjustment; Indicates the time difference between the current data collection time and the last model update time to reflect the timeliness of the data; .

[0104] Fourth, construct dynamic Gini gain.

[0105] Dynamic Gini Gain The calculation formula is:

[0106]

[0107] The larger the dynamic Gini gain is, the more effective the state feature is in partitioning the data set after considering feature relevance and data timeliness.

[0108] Fifth, the decision tree recursive strategy is determined based on the dynamic Gini gain.

[0109] First, determine the stop condition. If the stop condition is met, mark the current node as a leaf node and return. The stop condition is the historical status information data. All samples in the same information device state or state characteristics The collection is empty;

[0110] Then, all state features are traversed, the dynamic Gini gain of each state feature is calculated, and the state feature with the largest dynamic Gini gain is selected as the best partition feature of the current node, that is, the best feature;

[0111] Finally, the historical state information data is divided according to the different values ​​of the best feature, the corresponding nodes are established, and each state subset is processed. For non-empty subsets, the subtree is recursively constructed and connected to the corresponding branch of the current node.

[0112] The decision tree constructed in the above manner in this embodiment can better adapt to the complex and changeable data environment in enterprise operation and maintenance management, and improve the accuracy and timeliness of information equipment status judgment. For example, when judging whether a server has a risk of failure, considering the correlation between CPU usage and memory usage and the change trend of recent data, the decision tree is made more reasonable in selecting the division features through dynamic weight adjustment factors, thereby improving the accuracy of fault prediction.

[0113] It can be concluded from the above that this embodiment can more comprehensively capture the inherent laws and characteristics of the data by determining multiple state features of the historical state information data. Secondly, the state information subsets are divided based on the state features, which improves the accuracy and robustness of the model. At the same time, the establishment of a recursive strategy can flexibly respond to complex and changeable decision-making problems, making the model more adaptive.

[0114] In one embodiment of the present disclosure, reference Figure 2 , the control module 103 is specifically used for:

[0115] Process the target monitoring results and status information data based on the reinforcement learning algorithm to determine the target operation and maintenance strategy;

[0116] The control module 103 is connected to the information collection module 101 .

[0117] In one embodiment of the present disclosure, reference Figure 2 , the control module 103 is further configured to:

[0118] Determine the monitoring status of the information equipment based on the target monitoring results;

[0119] Determine the value function of the reinforcement learning algorithm based on the monitoring state and state information data;

[0120] Determine the target operation and maintenance strategy based on the feedback of the value function.

[0121] In this embodiment, the reinforcement learning algorithm interacts with the environment through the agent, continuously tries different behaviors, and learns the optimal behavior strategy based on the rewards fed back by the environment. In enterprise operation and maintenance management, the agent can be understood as the control module 103, the environment is the state information data and target monitoring results of the information equipment, the behavior is various operation and maintenance strategies, and the reward is the evaluation of the state change after the operation and maintenance strategy is executed.

[0122] The control module 103 takes the target monitoring results given by the monitoring module 102 and the status information data collected by the information acquisition module 101 as input, and uses the reinforcement learning algorithm to analyze and calculate. Through the reinforcement learning algorithm, the intelligent agent (control module 103) tries different operation and maintenance strategies (behaviors) under the current environmental state (i.e., the target monitoring results and status information data), and learns and selects the optimal target operation and maintenance strategy based on the rewards of environmental feedback, so as to achieve the purpose of improving the operating status of information equipment and improving performance.

[0123] The control module 103 clarifies the current specific state of the information device, such as normal state, performance warning state, fault state, etc., according to the target monitoring result output by the monitoring module 102.

[0124] The value function is used to evaluate the expectation of long-term cumulative rewards for taking a certain action in a certain state. The control module 103 determines the value function in the reinforcement learning algorithm based on the monitoring state of the information device and the relevant state information data. The determination of the value function needs to consider the future rewards that may be brought by different operation and maintenance strategies (behaviors) in the current state, which is the key basis for the reinforcement learning algorithm to learn the optimal strategy.

[0125] The control module 103 selects the operation and maintenance strategy that can maximize the value function as the target operation and maintenance strategy based on the calculated value function feedback. The target operation and maintenance strategy is most likely to bring long-term, positive system state improvement and rewards based on the evaluation of the reinforcement learning algorithm in the current state.

[0126] The core goal of the reinforcement learning algorithm is to learn an optimal strategy so that the behavior selected in any state can maximize the value function. During the training process, the reinforcement learning algorithm continuously adjusts the strategy to achieve an optimal matching relationship between the strategy and the value function, that is, to guide the optimization and update of the strategy based on the feedback of the value function.

[0127] Assume the state of the information device is , including various performance indicators, resource usage, and failure impact, etc.; the operation and maintenance strategy adopted is , the new state reached after executing the operation and maintenance strategy is ;

[0128] Value Function The calculation formula is as follows:

[0129]

[0130] in, Indicates new status For example, if the performance is measured by server response time, Status The average server response time under , which reflects the optimization effect of the operation and maintenance strategy on the server response time. It is the weight coefficient of the performance indicator, which is set according to the degree of business dependence on performance. For example, for businesses with high real-time requirements, A larger value can be set.

[0131] Indicates new status The optimization of resource utilization of information equipment is used to measure the optimization of resource utilization. Taking server memory resources as an example, Status The actual utilization of memory is as follows: ,but ,When the memory utilization is close to the middle value of the reasonable range, the value of this part is higher. It is the weight coefficient of resource utilization, which is set according to the enterprise's attention to resource costs.

[0132] In the new state For example, through historical data statistics and real-time monitoring, the probability of failure of the server in the current state in the future can be evaluated. is the fault impact weight coefficient. For services that are seriously affected by faults, A larger value should be set.

[0133] It can be concluded from the above that, by real-time monitoring of the status of information equipment and adjusting the value function accordingly, the control module 103 can dynamically optimize the operation and maintenance strategy to ensure the efficiency and pertinence of the operation and maintenance activities. Secondly, the application of the reinforcement learning algorithm enables the control module 103 to continuously learn and improve the strategy from actual operation and maintenance, thereby improving the adaptive ability and intelligence level of this embodiment. The control module 103 effectively improves the operation and maintenance efficiency, reduces the operation and maintenance costs, and provides a strong guarantee for the stable operation of information equipment.

[0134] In one embodiment of the present disclosure, reference Figure 2 , the enterprise operation and maintenance management platform also includes:

[0135] The security management module 105 is used to monitor the security status of information equipment;

[0136] The security management module 105 is connected to the control module 103 .

[0137] In one embodiment of the present disclosure, reference Figure 2 , the security management module 105 is specifically used for:

[0138] Determine multiple evaluation factors of information equipment and evaluation levels corresponding to the multiple evaluation factors;

[0139] Determine the membership degree of the evaluation level corresponding to each evaluation factor based on each evaluation factor;

[0140] Determine the fuzzy evaluation matrix according to the membership degree of the evaluation level corresponding to each evaluation factor;

[0141] Determine a weight set according to the weight of each evaluation factor;

[0142] Based on the fuzzy evaluation matrix and the weight set, the security evaluation level is determined, and the security evaluation level is the security status of the information equipment.

[0143] In this embodiment, the security management module 105 is used to monitor, evaluate and manage the security status of information equipment, collect data related to the security of information equipment through the information collection module 101, judge the security situation of information equipment, and provide decision support for the enterprise's information security protection.

[0144] Evaluation factors are used to measure various indicators or influencing factors of the security status of information equipment. The above factors cover the security attributes of information equipment in different aspects. By analyzing them, we can fully understand the security risks faced by information equipment. The evaluation level is a classification and grading of the status of each evaluation factor, which is used to intuitively represent the security level of the factor. The evaluation level is divided into multiple different levels, such as high risk, medium risk, low risk, or serious, relatively serious, general, good, etc.

[0145] The membership degree indicates the degree to which a certain evaluation factor belongs to a certain evaluation level, and its value range is between 0 and 1. The membership degree reflects the fuzzy relationship between the evaluation factor and the evaluation level, because in actual situations, some evaluation factors may not belong to a certain level completely and clearly, but have the characteristics of multiple levels to a certain extent.

[0146] The fuzzy evaluation matrix is ​​composed of the membership of each evaluation factor to different evaluation levels. The number of rows in the matrix is ​​equal to the number of evaluation factors, and the number of columns is equal to the number of evaluation levels. It fully reflects the fuzzy relationship between all evaluation factors and evaluation levels, and is an important basis for subsequent fuzzy comprehensive evaluation.

[0147] The weight set is a vector composed of the weights of each evaluation factor. The weights reflect the relative importance of each evaluation factor in the overall security evaluation. The sum of the weights is 1. By setting the weights reasonably, the factors that have a greater impact on the security status of information equipment can be highlighted.

[0148] The security management module 105 obtains a comprehensive evaluation vector by performing an operation (such as a fuzzy synthesis operation) on the fuzzy evaluation matrix and the weight set, and the vector represents the comprehensive membership of the information device to each evaluation level. Then, according to certain principles (such as the maximum membership principle), the security evaluation level of the information device is determined from the comprehensive evaluation vector, and the level is the security status of the information device.

[0149] Specifically, the relevant steps of this embodiment are as follows:

[0150] First, determine the evaluation factors and evaluation levels.

[0151] Assume that the evaluation factor set of information equipment is ;

[0152] in Indicates evaluation factors, such as It could be the vulnerability of the device. It is the network access control situation, etc.

[0153] Assume that the evaluation level set is ;

[0154] in Indicates evaluation factors, such as Indicates a high risk level. Indicates medium risk level. Indicates a low risk level, etc.

[0155] Second, determine the degree of affiliation.

[0156] For each evaluation factor , which corresponds to the evaluation level The membership degree is recorded as , Indicates evaluation factors Belongs to the evaluation level degree, and There are many ways to determine the degree of membership, such as fuzzy statistics.

[0157] Third, determine the fuzzy evaluation matrix.

[0158] According to the membership determined above , we can construct the fuzzy evaluation matrix , it is a The matrix is ​​in the following form:

[0159]

[0160] Fourth, determine the weight matrix.

[0161] Evaluation factors The weight is , weight set , and satisfies , . Weight The determination of can also be determined through various methods such as hierarchical analysis method.

[0162] Fifth, determine the safety assessment level.

[0163] This embodiment uses fuzzy synthesis operation to convert the weight set and fuzzy evaluation matrix Perform calculations to obtain a comprehensive evaluation vector ,in Represents the Hadamard product.

[0164] Finally, the safety evaluation level is determined according to the maximum membership principle, that is, The value of is taken as the safety evaluation level. For example, the value range of low risk level is , the value range of medium risk level is , the value range of high risk level is .

[0165] It can be concluded from the above that this embodiment achieves accurate monitoring of the security status of information equipment through detailed evaluation factors and corresponding evaluation levels, greatly improving the accuracy and efficiency of security management. Secondly, by combining the fuzzy evaluation matrix and the weight set, the importance and influence of each evaluation factor are comprehensively considered, making the security evaluation more comprehensive. This embodiment not only reflects the overall security level of information equipment, but also can timely discover potential safety hazards, providing strong decision-making support for operation and maintenance personnel.

[0166] In one embodiment of the present disclosure, reference Figure 2 , the enterprise operation and maintenance management platform further includes: a storage module 106;

[0167] The storage module 106 is connected to the information collection module 101 and the monitoring module 102 respectively;

[0168] The storage module 106 is used to store the status information data of the information device and the monitoring results of the monitoring module 102.

[0169] In this embodiment, the storage module 106 is used for data storage, and saves the state information data generated during the operation of the information device, so as to facilitate subsequent query, analysis, and as a basis for the work of other modules. For example, the historical state information data and the corresponding historical monitoring results stored in the storage module 106 can be provided to the training unit 201, so the storage module 106 is also connected to the training unit 201.

[0170] From the above, it can be concluded that this embodiment realizes centralized storage of information equipment status information data and monitoring results through storage module 106, effectively improving the convenience and efficiency of data management, ensuring the integrity and traceability of data, and providing comprehensive data support for operation and maintenance personnel, which helps to timely discover and resolve equipment failures, optimize operation and maintenance processes, and improve overall operation and maintenance efficiency and quality.

[0171] In one embodiment of the present disclosure, reference Figure 2 , the enterprise operation and maintenance management platform also includes: an alarm module 107;

[0172] The alarm module 107 is connected to the monitoring module 102;

[0173] The alarm module 107 is used to generate an alarm in response to the monitoring result of the monitoring module 102 being a fault.

[0174] In this embodiment, the alarm module 107 is used to issue an alarm when the monitoring result is a fault, so as to promptly notify relevant personnel (such as operation and maintenance engineers) that the information equipment has a condition that requires attention and processing, ensuring that potential problems can be quickly known and corresponding measures can be taken to resolve them.

[0175] A failure occurs when an information device has an abnormal situation during operation, causing it to be unable to perform its functions normally, or a problem that may affect the normal operation of the business. For example, sudden server crashes, network connection interruptions, application crashes, etc. all fall into the category of failures.

[0176] From the above, it can be concluded that this embodiment realizes immediate response and alarm to information equipment failures through the alarm module 107, greatly enhancing the early warning capability of the platform, and can quickly notify operation and maintenance personnel when equipment failures occur, effectively shortening the fault response time, reducing the risk of business interruption caused by equipment failures, and improving overall operation and maintenance efficiency.

[0177] The above embodiments are only used to illustrate the technical solutions of the present disclosure, rather than to limit them. Although the present disclosure has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present disclosure.

Claims

1. An enterprise operation and maintenance management platform, characterized in that: include: An information collection module is used to collect status information data of information equipment; A monitoring module, connected to the information acquisition module, for inputting the status information data into a target decision tree model to obtain a target monitoring result of the information device; The monitoring module includes: a training unit, which is used to train a decision tree model based on historical status information data and historical monitoring results corresponding to the historical status information data to obtain a target decision tree model; The training unit is specifically used for: Determining a plurality of state characteristics of the historical state information data; Dividing the historical state information data based on the multiple state characteristics to obtain multiple state information subsets; the multiple state characteristics correspond to the multiple state information subsets one by one; Establishing a recursive strategy based on each state feature and a state information subset corresponding to each state feature; determining the recursive strategy based on a dynamic Gini gain; The calculation formula of the dynamic Gini gain is: in, is the dynamic Gini gain, is the Gini index, is the historical status information data, is the state feature, The state feature Features, , state characteristics have Different value ranges , according to the state characteristics The value of the historical status information data Divide into State feature subset , Status Features The value is A subset of samples, , is the total number of historical status information data, is the number of samples of the state feature category, is the dynamic weight adjustment factor, , and These are adjustment parameters. Indicates state characteristics and other status characteristics The correlation coefficient of , It is the time difference between the current data collection time and the last model update time; Determining the target decision tree model based on the recursive strategy; A control module, connected to the monitoring module, for determining a target operation and maintenance strategy according to the target monitoring result; The control module is specifically used for: The target monitoring results and the status information data are processed based on a reinforcement learning algorithm to determine a target operation and maintenance strategy; the control module is connected to the information acquisition module; The control module is further specifically used for: Determining a monitoring status of the information device based on the target monitoring result; The value function of the reinforcement learning algorithm is determined based on the monitoring state and the state information data; the calculation formula of the value function is: in, is the value function, The operation and maintenance strategy adopted is: is the new state achieved after executing the operation and maintenance strategy. For new status The degree of improvement in the performance of information equipment. is the weight coefficient of the performance index, For new status The optimization of resource utilization of information equipment is as follows: is the weight coefficient of resource utilization, In the new state The probability of failure occurring is is the fault impact weight coefficient; A target operation and maintenance strategy is determined based on feedback of the value function.

2. The enterprise operation and maintenance management platform according to claim 1, characterized in that: Also includes: An operation and maintenance module, used to determine operation and maintenance control parameters based on the target operation and maintenance strategy, and to maintain the information equipment based on the operation and maintenance control parameters; The operation and maintenance module is connected to the control module.

3. The enterprise operation and maintenance management platform according to claim 1, characterized in that: The monitoring module comprises: A monitoring unit, configured to determine a target monitoring result of the information device based on the target decision tree model; The monitoring unit is connected to the information acquisition module, the training unit and the control module respectively.

4. The enterprise operation and maintenance management platform according to claim 1, characterized in that: Also includes: A security management module, used to monitor the security status of the information equipment; The security management module is connected to the control module.

5. The enterprise operation and maintenance management platform according to claim 4, characterized in that: The security management module is specifically used for: Determining a plurality of evaluation factors of the information device and evaluation levels corresponding to the plurality of evaluation factors; Determine, based on each evaluation factor, the degree of membership of the evaluation level corresponding to each evaluation factor; Determine the fuzzy evaluation matrix according to the membership degree of the evaluation level corresponding to each evaluation factor; Determine a weight set according to the weight of each evaluation factor; Based on the fuzzy evaluation matrix and the weight set, a security evaluation level is determined, where the security evaluation level is the security status of the information device.

6. The enterprise operation and maintenance management platform according to claim 1, characterized in that: Also included: a storage module; The storage module is connected to the information acquisition module and the monitoring module respectively; The storage module is used to store the status information data of the information device and the monitoring results of the monitoring module.

7. The enterprise operation and maintenance management platform according to claim 1, characterized in that: Also includes: Alarm module; The alarm module is connected to the monitoring module; The alarm module is used to generate an alarm in response to a fault detected by the monitoring module.

Citation Information

Patent Citations

  • Power failure processing method, device and equipment for electric power guarantee area and medium

    CN118504991A

  • Intelligent power management and distribution method and system for integrated computer cabinet

    CN118868267A

  • Customer resource monitoring for versatile scaling service scaling policy recommendations

    US10409642B1