An IT informatization dual-model operation and maintenance management method based on a multi-layer data lake

By adopting multi-layer data lake architecture and dual-model operation and maintenance management methods in IT operation and maintenance management, the problems of multi-source heterogeneous data integration, system problem identification and resource optimization are solved, and efficient, accurate and intelligent IT operation and maintenance management are achieved.

CN119323367BActive Publication Date: 2025-06-20ZHUHAI ZHILIAN SIXUN TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411876316.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-19
Publication Date
2025-06-20
Estimated Expiration
2044-12-19

AI Technical Summary

Technical Problem

The existing IT operation and maintenance management methods are difficult to integrate multi-source heterogeneous data, have low accuracy in identifying system performance problems, insufficient resource allocation optimization efficiency, and limited real-time support capabilities.

Method used

The IT information dual-model operation and maintenance management method based on multi-layer data lakes is adopted. By collecting IT operation and maintenance data in real time from multiple data sources, a multi-layer operation and maintenance data network is formed and embedded into the data lake to quickly access and retrieve large-scale data. Use the operation and maintenance identification model and operation and maintenance adjustment model to conduct big data analysis and resource allocation optimization on data.

Benefits of technology

It realizes efficient integration and management of multi-source heterogeneous data, accurately identifying system problems and performance bottlenecks, intelligent resource optimization configuration and real-time operation and maintenance monitoring, significantly improving operation and maintenance efficiency and system stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119323367B_ABST
    Figure CN119323367B_ABST
Patent Text Reader

Abstract

The present invention discloses an IT informatization dual-model operation and maintenance management method based on a multi-layer data lake, comprising the following steps: collecting IT operation and maintenance related data in real time from multiple data sources, and after performing correlation processing, forming a multi-layer operation and maintenance data network; embedding the multi-layer operation and maintenance data network into the data lake by means of a layering technique to form a multi-layer operation and maintenance data lake, and performing rapid access and retrieval of large-scale operation and maintenance data through multi-layer interfaces; performing big data analysis on the retrieved data to complete precise filtering of the data, extraction of key features and deep combination, and obtaining a number of key performance evaluation status quantity sets; inputting the status quantity sets into an operation and maintenance identification model to identify potential factors that may cause system problems and performance bottlenecks; automatically adjusting IT resource allocation based on the potential factors through an operation and maintenance adjustment model to optimize system performance and improve resource utilization rate; and performing dynamic display of the identification results and adjustment results. The present invention significantly improves operation and maintenance efficiency and system stability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent operation and maintenance management, and particularly relates to an IT informatization dual-model operation and maintenance management method based on a multi-layer data lake. Background Art

[0002] With the rapid development of information technology and the promotion of enterprise digital transformation, the complexity of IT systems has been continuously increasing, and the traditional IT operation and maintenance management methods can no longer meet the requirements of the rapid growth of current multi-source and multi-level data. The main challenges faced in IT operation and maintenance management include: diversified data sources, huge data volume and high real-time requirements, great difficulty in identifying system performance bottlenecks, and low efficiency in optimizing resource allocation. The existence of these problems makes it difficult to improve the efficiency and reliability of IT operation and maintenance management, seriously affecting the stability and operation efficiency of enterprise informatization systems.

[0003] The existing IT operation and maintenance management methods mainly focus on the analysis and adjustment based on rules or single models. Although they can identify and solve some system failures and performance bottlenecks, there are the following problems:

[0004] (1) Difficulty in integrating multi-source heterogeneous data: The sources of IT operation and maintenance data are complex, including both structured data (such as performance indicators in databases) and unstructured data (such as log files and operation and maintenance documents). The existing methods have insufficient ability to integrate multi-source heterogeneous data and cannot efficiently analyze using all data.

[0005] (2) Low accuracy in identifying system performance problems: Traditional operation and maintenance management tools mainly rely on empirical rules or simple models for the analysis and identification of performance problems, lacking in-depth excavation of potential factors of system problems, resulting in insufficient accuracy and timeliness of problem diagnosis.

[0006] (3) Unintelligent optimization of resource allocation: Most of the existing operation and maintenance management methods rely on static resource allocation strategies, lacking the ability to dynamically perceive and adjust the system operation state, and it is difficult to maximize resource utilization.

[0007] (4) Lack of scalability and real-time support: With the expansion of business scale and the explosion of data volume, the processing capacity of traditional IT operation and maintenance management tools is difficult to expand and cannot meet the requirements of real-time operation and maintenance management. Summary of the Invention

[0008] To overcome the deficiencies of the prior art, the present invention provides an IT informatization dual-model operation and maintenance management method based on a multi-layer data lake, which is used to solve the technical problems existing in the existing IT operation and maintenance management, such as the difficulty in integrating multi-source heterogeneous data, the low accuracy of system problem identification, the insufficient efficiency of resource configuration optimization, and the limited real-time support ability, so as to form an efficient, accurate, and intelligent IT informatization operation and maintenance management method, significantly improving the operation and maintenance efficiency and system stability.

[0009] To solve the above problems, the technical solutions adopted by the present invention are as follows:

[0010] An IT informatization dual-model operation and maintenance management method based on a multi-layer data lake includes the following steps:

[0011] Collect IT operation and maintenance related data in real time from multiple data sources, and after performing correlation processing, form a multi-layer operation and maintenance data network;

[0012] Embed the multi-layer operation and maintenance data network into the data lake by using a layering technology to form a multi-layer operation and maintenance data lake, and perform rapid access and retrieval of large-scale operation and maintenance data through the multi-layer interfaces of the multi-layer operation and maintenance data lake;

[0013] Perform big data analysis on the retrieved large-scale operation and maintenance data, complete the precise filtering of the data, extraction of key features, and in-depth combination to obtain several key performance evaluation status quantity sets;

[0014] Input the several key performance evaluation status quantity sets into an operation and maintenance identification model to identify potential factors that may cause system problems and performance bottlenecks;

[0015] Automatically adjust the IT resource configuration based on the potential factors through an operation and maintenance adjustment model to optimize the system performance and improve the resource utilization rate;

[0016] Dynamically display the identification results and adjustment results by associating the operation and maintenance identification model and the operation and maintenance adjustment model to support operation and maintenance decision-making;

[0017] Among them, when collecting IT operation and maintenance related data in real time from multiple data sources, receive data through a Kafka operation and maintenance cluster, divide data reception groups, and complete HDFS storage by configuring nodes.

[0018] As a preferred embodiment of the present invention, when receiving data through a Kafka operation and maintenance cluster, dividing data reception groups, and completing HDFS storage by configuring nodes, it includes:

[0019] Transmit the collected IT operation and maintenance related data to the Kafka operation and maintenance cluster through Flume;

[0020] Multiple instances in the Kafka operation and maintenance cluster respectively correspond to different data sources, which are used to receive IT operation and maintenance related data of each data source, and the multiple instances are divided into multiple data receiving groups according to a preset quantity;

[0021] Equip each data receiving group with a corresponding node, and converge the collected data into multiple nodes;

[0022] Equip each node with a corresponding Producer, and send the collected data back to the Kafka operation and maintenance cluster again through the Producer for HDFS storage.

[0023] As a preferred embodiment of the present invention, when forming a multi-layer operation and maintenance data network, it includes:

[0024] Adopt a data storage method of [time_stamp][key1][value1]…[keyN][valueN] for the time-series data and text data in the Kafka operation and maintenance cluster;

[0025] Make the number of key-value pairs [key][value] expandable at each time point corresponding to each [time_stamp], so as to store the time-series data and text data into different-level table structures;

[0026] Assign combined query conditions to the time-series data and text data in the table structure, and perform aggregation calculations to form the multi-layer operation and maintenance data network.

[0027] As a preferred embodiment of the present invention, when forming a multi-layer operation and maintenance data lake, it includes:

[0028] Encode the data in different-level table structures into numerical data sequences with the same structural form according to preset semantic logic, and upload them to the cloud server for dimensionless processing and standardization processing, and then converge and upload them to the cloud database through the gateway;

[0029] Parse and continuously monitor the data sequences uploaded by the gateway through the cloud server;

[0030] When it is detected that the data sequence stream is interrupted, extract the data sequence through the cloud server and compare it with a preset data format;

[0031] When the comparison passes, connect to the corresponding database interface according to the data type, assign an identification number, convert it into JSON format to obtain the actual data value, and convert it into a triple through the mapping rule;

[0032] Correspond the triple with concepts and entities in the data lake database connected to the corresponding database interface to obtain a data lake middle platform, and provide multi-layer API interfaces for the data lake middle platform.

[0033] As a preferred embodiment of the present invention, when performing fast access and retrieval, it includes:

[0034] Determine the search area and target conditions according to the access and retrieval conditions, and obtain the top-level data block that is the smallest and can completely contain the search area;

[0035] Judge whether the top-level data block meets the target conditions. If it meets, insert the data block index number into the linked list. If it does not meet, continue to search the branches of the top-level data block until all data blocks that meet the target conditions are found;

[0036] Insert the indexes of all data blocks that meet the target conditions into the linked list in the order from left to right and from top to bottom, and perform fast access and retrieval of large-scale operation and maintenance data through the linked list.

[0037] As a preferred embodiment of the present invention, when completing precise filtering of data, extraction of key features, and deep combination, it includes:

[0038] Use filtering rules to filter the data to form a key matrix model for identifying the operating state of the system;

[0039] Use the Apriori algorithm to obtain frequent multiple sets, and mine the association rules between data types and fault types in the key matrix model to establish the set of several key performance evaluation status quantities;

[0040] Among them, when mining association rules, it includes:

[0041] Perform clustering analysis on the relationships between different data types and fault types in the key matrix model, mine the shape coefficient and silhouette coefficient of the associated data curve, and extract the collection quantities of relevant data types to obtain the set of key performance evaluation status quantities for different fault types.

[0042] As a preferred embodiment of the present invention, when identifying potential factors, it includes:

[0043] The operation and maintenance identification model provides: the antecedent of the association rule and the consequent of the association rule , the antecedent of the association rule represents the attribute characteristics of IT informatization operation and maintenance data, and the consequent of the association rule represents whether there are system problems and performance bottlenecks, represents system problems, represents performance bottlenecks;

[0044] Based on the antecedent and consequent of the association rule and , obtain the support degree of the performance bottleneck occurrence ;

[0045] Based on the antecedent and consequent of the association rule and , obtain the support degree of the system problem occurrence ;

[0046] Take the attribute features greater than the threshold as the first potential factor;

[0047] Take the attribute features greater than the threshold as the second potential factor.

[0048] As a preferred embodiment of the present invention, when automatically adjusting the IT resource configuration, it includes:

[0049] The operation and maintenance adjustment model determines the first control coefficient and the second control coefficient according to the first potential factor and the second potential factor, and controls the corresponding types of IT resource flows through the first control coefficient and the second control coefficient, as shown in Formula 3 and Formula 4:

[0050] (3);

[0051] In the formula, is the first control amount of the first potential factor corresponding to the type of IT resources during the planned period ; is the th th type of IT resource flow during the planned period is the number of the type of IT resource flows during the planned period is the th first control coefficient of the

[0052] (4);

[0053] In the formula, is the second control amount of the second potential factor corresponding to the type of IT resources during the planned period ; is the th th type of IT resource flow during the planned period is The quantity of class IT resource flows For the planned time period The th The second control coefficient of class IT resources

[0054] As a preferred embodiment of the present invention, when automatically adjusting the IT resource configuration, it further includes:

[0055] Maximize the approximation between the first control quantity and the second control quantity to obtain the first IT resource optimal configuration result and the second IT resource optimal configuration result, as shown in Formula 5 and Formula 6:

[0056] (5);

[0057] In the formula, Is the actual control quantity of class IT resources corresponding to the first potential factor in the planned time period corresponding to The actual control quantity of class IT resources Is the number of IT resource types corresponding to the first potential factor Is the total number of the first planned time periods Is the first IT resource optimal configuration result;

[0058] (6);

[0059] In the formula, Is the actual control quantity of class IT resources corresponding to the second potential factor in the planned time period corresponding to The actual control quantity of class IT resources Is the number of IT resource types corresponding to the second potential factor Is the total number of the second planned time periods Is the second IT resource optimal configuration result

[0060] As a preferred embodiment of the present invention, when performing dynamic display, it includes:

[0061] Define a graph, and use the identified first potential factor, second potential factor, first IT resource optimal configuration result, and second IT resource optimal configuration result to determine the vertices of the recognition result and the adjustment result by the host IP, collection directory, and file name template respectively, and set the attribute information of the vertices;

[0062] Use the vertex as the source end, and match the corresponding transmission target end. Take the transmission target end as the vertex of the next layer of the dynamic display graph one by one, set its attribute information, and connect it with the upper layer;

[0063] Match the source end information of the next layer, and use the matched source end information as the vertex of the next lower layer until no source end information can be matched

[0064] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0065] The present invention proposes an IT informatization dual-model operation and maintenance management method based on a multi-layer data lake. By constructing a multi-layer data lake architecture and combining an operation and maintenance identification model with an operation and maintenance adjustment model, it realizes the efficient integration of multi-source heterogeneous data in complex IT systems, the accurate identification of potential problems, and the dynamic optimal allocation of resources, and has the following beneficial effects:

[0066] (1) Efficient integration and management of multi-source heterogeneous data

[0067] By adopting a multi-layer operation and maintenance data lake architecture and a hierarchical embedding method of the data network, the present invention can uniformly store and manage structured data and unstructured data from multiple sources. Using multi-layer API interfaces, it realizes the rapid access and retrieval of large-scale data, significantly improving the processing efficiency and expansion ability of IT operation and maintenance data.

[0068] (2) Accurate identification of system problems and performance bottlenecks

[0069] Through big data analysis of large-scale operation and maintenance data by the operation and maintenance identification model of the present invention, using association rules and clustering algorithms to mine the key factors of potential problems and performance bottlenecks in the system, and combining the generation of a key performance evaluation status quantity set, it can accurately identify the root causes that may trigger system problems, improving the accuracy and timeliness of problem diagnosis.

[0070] (3) Intelligent resource optimization configuration

[0071] The operation and maintenance adjustment model dynamically determines the control coefficients of IT resources according to the identified potential factors, and realizes the intelligent scheduling and configuration of IT resources through an optimization algorithm. The optimization results effectively improve the resource utilization rate, reduce the risk of system performance bottlenecks, and at the same time reduce manual intervention in the operation and maintenance process.

[0072] (4) Real-time operation and maintenance monitoring and dynamic decision-making support

[0073] By combining the operation and maintenance identification model and the operation and maintenance adjustment model, the present invention can realize the dynamic display of operation and maintenance identification results and adjustment results. By adopting a visual dynamic display graph, operation and maintenance personnel can intuitively understand the system operation status and adjustment effects, providing efficient support for real-time operation and maintenance decision-making.

[0074] (5) Strong scalability and real-time performance

[0075] The present invention combines a multi-layer data lake architecture with a cloud server to achieve dimensionless processing and standardized storage of data, as well as continuous monitoring and dynamic analysis of real-time data streams. When data anomalies are detected, it can quickly locate problems and execute corresponding operations, with good real-time performance and scalability, meeting the dynamic operation and maintenance requirements of large-scale IT systems.

[0076] (6)Fully support the intelligent operation and maintenance of complex IT systems

[0077] Based on the multi-layer data lake architecture, the present invention makes full use of big data analysis and intelligent algorithms to form a full-process operation and maintenance management solution from data collection, analysis to resource optimization, which is applicable to a variety of complex IT operation and maintenance scenarios, such as cloud computing platforms, distributed storage systems, large enterprise data centers, etc., effectively improving the stability and operation and maintenance efficiency of the system.

[0078] The following further elaborates on the present invention in conjunction with the accompanying drawings and specific embodiments. Description of the Drawings

[0079] Figure 1 It is a flowchart of the IT informatization dual-model operation and maintenance management method based on a multi-layer data lake provided by the present invention. Specific Embodiments

[0080] The IT informatization dual-model operation and maintenance management method based on a multi-layer data lake provided by the present invention, as Figure 1 shown, includes the following steps:

[0081] Step S1: Real-time collect IT operation and maintenance related data from multiple data sources, and after correlation processing, form a multi-layer operation and maintenance data network;

[0082] Step S2: Embed the multi-layer operation and maintenance data network into the data lake using hierarchical technology to form a multi-layer operation and maintenance data lake, and perform rapid access and retrieval of large-scale operation and maintenance data through the multi-layer interfaces of the multi-layer operation and maintenance data lake;

[0083] Step S3: Perform big data analysis on the retrieved large-scale operation and maintenance data to complete accurate filtering of data, extraction of key features, and in-depth combination, and obtain several key performance evaluation status quantity sets;

[0084] Step S4: Input several key performance evaluation status quantity sets into the operation and maintenance identification model to identify potential factors that may cause system problems and performance bottlenecks;

[0085] Step S5: Automatically adjust the IT resource configuration based on the potential factors through the operation and maintenance adjustment model to optimize the system performance and improve the resource utilization rate;

[0086] Step S6: Dynamically display the recognition results and adjustment results through the associated operation and maintenance recognition model and operation and maintenance adjustment model to support operation and maintenance decision-making.

[0087] In the above Step S1, when collecting IT operation and maintenance related data in real time from multiple data sources, it includes:

[0088] Transmit the collected IT operation and maintenance related data to the Kafka operation and maintenance cluster through Flume;

[0089] Multiple instances in the Kafka operation and maintenance cluster respectively correspond to different data sources, which are used to receive IT operation and maintenance related data from each data source, and divide the multiple instances into multiple data receiving groups according to a preset quantity;

[0090] Equip each data receiving group with a corresponding node, and converge the collected data into multiple nodes;

[0091] Equip each node with a corresponding Producer, and use the Producer to send the collected data back to the Kafka operation and maintenance cluster again for HDFS storage;

[0092] Among them, the configuration of Flume is implemented through SpoolDirectorySource, and the type of Flume is defined as spooldir, and the data input path for receiving IT operation and maintenance related data is specified;

[0093] The Kafka operation and maintenance cluster manages the cluster configuration through Zookeeper and performs rebalance when electing a leader or when the data receiving group changes.

[0094] Furthermore, when performing HDFS storage, it includes:

[0095] When the data collected by the Producer enters the Kafka operation and maintenance cluster, the data is split through the HDFS Client, with the NameNode as the master node, and by interacting with the NameNode, manage the namespace and data block mapping information of HDFS, configure the replication policy, process requests, and generate file location information;

[0096] Use the DataNode as the slave node to store data, read and write data, and report storage information to the NameNode;

[0097] Assist the NameNode to run through the Secondary NameNode, regularly merge the fsimage and fsedits and push them to the NameNode, and assist the NameNode in recovery.

[0098] In the above step S1, when forming a multi-layer operation and maintenance data network, it includes:

[0099] For time-series data and text data in the Kafka operation and maintenance cluster, adopt a data storage method of [time_stamp][key1][value1][key2][value2][key3][value3]…[keyN][valueN];

[0100] Make the number of key-value pairs [key][value] scalable at each time point corresponding to [time_stamp], so as to store time-series data and text data into table structures at different levels;

[0101] Assign combined query conditions to the time-series data and text data in the table structure, and perform aggregation calculations to form a multi-layer operation and maintenance data network;

[0102] Among them, the multi-layer operation and maintenance data network adopts a consistent or non-consistent method in the data writing form, and through the cooperation of the consistent and non-consistent methods, realizes data display in different dimensions and forms;

[0103] Aggregation calculations include:

[0104] SELECT key2, sum(value1) / sum(value2) FROM

[0105] metric WHERE key1 = “A” AND timestamp >= B

[0106] AND timestamp < C GROUP BY key2。

[0107] Specifically, [time_stamp] is a time stamp, which is used to identify the recording time of data. Each pair of [key][value] represents an attribute and a value, and the number of key-value pairs is dynamic and can be extended to [keyN][valueN] according to actual needs.

[0108] In the above step S2, when forming a multi-layer operation and maintenance data lake, it includes:

[0109] Encode the data in table structures at different levels into numerical data sequences with the same structural form according to preset semantic logic, and upload them to the cloud server;

[0110] Perform dimensionless processing and standardization processing on the numerical data sequences through the cloud server;

[0111] After the dimensionless and standardized data sequences are aggregated and uploaded to the cloud database through the gateway, the cloud server parses the data sequences uploaded by the gateway and continuously monitors the data sequence stream information in real time;

[0112] When it is detected that the data sequence stream is interrupted, the cloud server extracts the data sequence and compares it with a preset data format;

[0113] After the data format comparison passes, connect to the corresponding database interface according to the data type, assign an identification number, convert it to the JSON format to obtain the actual data value, and convert it to a triple through the mapping rule;

[0114] Correspond the triples with the concepts and entities in the data lake database connected to the corresponding database interface to obtain the data lake middle platform, and provide multiple layers of API interfaces for the data lake middle platform for subsequent big data analysis and dynamic display.

[0115] Further, when performing dimensionless processing, it includes:

[0116] For data that needs to use the PCA algorithm for dimensionality reduction or data that needs to use distance to measure similarity, use a method including zero-mean normalization to make the data with different dimensions dimensionless;

[0117] When performing standardization processing, it includes:

[0118] For data that does not involve distance measurement, covariance calculation, and does not conform to the normal distribution, use a method including maximum-minimum normalization to limit the matrix value within a preset range.

[0119] In the above step S2, when performing fast access and retrieval of large-scale operation and maintenance data, it includes:

[0120] Determine the search area and target conditions according to the access and retrieval conditions, and obtain the top-level data block that is the smallest and can completely contain the search area;

[0121] Judge whether the top-level data block meets the target conditions. If it meets, insert the data block index number into the linked list. If it does not meet, continue to search the branches of the top-level data block until all data blocks that meet the target conditions are found;

[0122] Insert the indexes of all data blocks that meet the target conditions into the linked list in the order from left to right and from top to bottom, and perform fast access and retrieval of large-scale operation and maintenance data through the linked list.

[0123] In the above step S3, when completing the precise filtering of data, extraction of key features, and deep combination, it includes:

[0124] Filter the data using filtering rules to form a key matrix model for identifying the operating state of the system;

[0125] Use the Apriori algorithm to obtain frequent multi-sets, and mine the association rules between data types and fault types in the key matrix model to establish several key performance evaluation state variable sets;

[0126] Among them, when filtering the data using filtering rules, it includes:

[0127] Let be the set of operation and maintenance data, providing a transaction database composed of a series of transactions with unique identifiers , and each transaction corresponds to an operation and maintenance data subset on the set , that is , then the filtering rule is expressed as a logical implication formula of , where , ;

[0128] Obtain the percentage of transactions containing in the transaction database, filter out the operation and maintenance data subsets with a percentage less than the first preset value to obtain a preliminary filtered operation and maintenance data set;

[0129] Obtain the percentage of transactions containing and transactions containing in the transaction database, filter out the operation and maintenance data subsets with a percentage less than the second preset value, and form a key matrix model;

[0130] When mining association rules, it includes:

[0131] Conduct cluster analysis on the relationship between different data types and fault types in the key matrix model, mine the shape coefficient and silhouette coefficient of the associated data curve, and extract the collection amounts of relevant data types to obtain key performance evaluation state variable sets for different fault types;

[0132] In the key matrix model, the abscissa is the data type and the ordinate is the collection amount.

[0133] In step S4 above, when identifying potential factors, it includes:

[0134] The operation and maintenance identification model provides: the antecedent of the association rule and the consequent of the association rule , the antecedent of the association rule represents the attribute characteristics of IT information-based operation and maintenance data, and the consequent of the association rule represents whether system problems and performance bottlenecks occur, Indicates a system problem, Indicates a performance bottleneck;

[0135] Based on the antecedent of the association rule and the consequent of the association rule , obtain the support degree of the occurrence of performance bottlenecks , as shown in Formula 1:

[0136] (1);

[0137] In the formula, is the support degree of the th category attribute feature of the th feature factor for a certain type of performance bottleneck, is the th category attribute feature of the th feature factor, is the total number of times that the th category attribute feature of the th feature factor causes the operation to have a performance bottleneck, is the total number of operations;

[0138] Based on the antecedent of the association rule and the consequent of the association rule , obtain the support degree of the occurrence of system problems , as shown in Formula 2:

[0139] (2);

[0140] In the formula, is the support degree of the th category attribute feature of the th feature factor for a certain type of system problem, is the total number of times that the th category attribute feature of the th feature factor causes the operation to have a system problem;

[0141] Take the attribute features with the support degree of the occurrence of performance bottlenecks greater than the threshold as the first potential factor;

[0142] Take the attribute features with the support degree of the occurrence of system problems greater than the threshold as the second potential factor.

[0143] Specifically, the attribute features of IT information operation and maintenance data specifically include: system performance data, log data, event and alarm data, configuration information data, user behavior data, data stream and network traffic data, monitoring and analysis data, automated operation and maintenance data, exception and problem diagnosis data, operation and maintenance management data, security and compliance related data.

[0144] In the above step S5, when automatically adjusting the IT resource configuration, it includes:

[0145] The operation and maintenance adjustment model determines the first control coefficient and the second control coefficient according to the first latent factor and the second latent factor, and controls the corresponding type of IT resource flow through the first control coefficient and the second control coefficient, as shown in Formula 3 and Formula 4:

[0146] (3);

[0147] In the formula, is the first control amount of the type of IT resources corresponding to the first latent factor during the planned period ; is the th article type of IT resource flow, is the quantity of the type of IT resource flow, is the first control coefficient of the th article type of IT resources during the planned period

[0148] (4);

[0149] In the formula, is the second control amount of the type of IT resources corresponding to the second latent factor during the planned period ; is the th article type of IT resource flow, is the quantity of the type of IT resource flow, is the second control coefficient of the th article type of IT resources during the planned period

[0150] Maximize the proximity between the first control amount and the second control amount to obtain the first IT resource optimized configuration result and the second IT resource optimized configuration result, as shown in Formula 5 and Formula 6:

[0151] (5);

[0152] In the formula, is the first latent factor corresponding to the during the planned period The actual control amount of IT resources of a certain type is the number of types of IT resources corresponding to the first potential factor is the total number of the first planned time periods is the optimization and allocation result of the first IT resources

[0153] (6);

[0154] In the formula is the actual control amount of IT resources of a certain type corresponding to the second potential factor during the planned time period is the number of types of IT resources corresponding to the second potential factor is the total number of the second planned time periods is the optimization and allocation result of the second IT resources

[0155] Specifically, the configuration of various IT resources includes: computing resource configuration, storage resource configuration, network resource configuration, software resource configuration, security resource configuration, scheduling and automation resource configuration, energy resource configuration, user resource configuration, high-availability and disaster recovery resource configuration.

[0156] Furthermore, the computing resource configuration includes: server resources, container resources, cloud computing resources.

[0157] The storage resource configuration includes: storage capacity, distributed storage, backup storage, database storage.

[0158] The network resource configuration includes: bandwidth, IP address, load balancing, firewall rules, VPN configuration.

[0159] The software resource configuration includes: operating system, middleware, database, application deployment.

[0160] The security resource configuration includes: identity authentication, permission management, security policy, logging and auditing.

[0161] The scheduling and automation resource configuration includes: task scheduling, resource allocation strategy, automated operation and maintenance.

[0162] The energy resource configuration includes: power distribution, heat dissipation resource configuration.

[0163] The user resource configuration includes: account allocation, quota limit, user policy, terminal device management.

[0164] The high-availability and disaster recovery resource configuration includes: redundant resources, backup strategy, disaster recovery system.

[0165] In the above step S6, when dynamically displaying the recognition result and the adjustment result, it includes:​​

[0166] Define a graph, and use the host IP, collection directory, and file name template to determine the vertices of the recognition result and adjustment result for the identified first potential factor, second potential factor, first IT resource optimization configuration result, and second IT resource optimization configuration result respectively, and set the attribute information of the vertices.

[0167] Use the vertices as the source ends, and match the corresponding transmission target ends. Take each transmission target end as the vertex of the next layer of the dynamic display graph, set its attribute information, and connect it to the previous layer.

[0168] Match the source end information of the next layer, and use the matched source end information as the vertices of the layer after the next layer until no source end information can be matched.

[0169] Among them, the nodes generated by the first potential factor are associated with the nodes generated by the first IT resource optimization configuration result, and the nodes generated by the second potential factor are associated with the nodes generated by the second IT resource optimization configuration result. The attribute information includes: the font, color, and shape of the vertices.

[0170] The above embodiments are only the preferred embodiments of the present invention, and cannot be used to limit the protection scope of the present invention. Any non-substantive changes and substitutions made by those skilled in the art based on the present invention belong to the protection scope required by the present invention.

Claims

1. A dual-model operation and maintenance management method for IT informatization based on a multi-layer data lake, characterized in that: The following steps are involved: Collect IT operation and maintenance related data from multiple data sources in real time, and form a multi-layer operation and maintenance data network after correlation processing; The multi-layer operation and maintenance data network is embedded into the data lake using layered technology to form a multi-layer operation and maintenance data lake, and the multi-layer interfaces of the multi-layer operation and maintenance data lake are used to quickly access and retrieve large-scale operation and maintenance data; Perform big data analysis on the retrieved large-scale operation and maintenance data, complete accurate data filtering, extract key features and deep integration, and obtain several key performance evaluation state quantity sets; Inputting the plurality of key performance evaluation state quantity sets into the operation and maintenance identification model to identify potential factors that may cause system problems and performance bottlenecks; Automatically adjust IT resource configuration based on the potential factors through an operation and maintenance adjustment model; Dynamically displaying identification results and adjustment results by associating the operation and maintenance identification model with the operation and maintenance adjustment model to support operation and maintenance decision-making; Among them, when collecting IT operation and maintenance related data from multiple data sources in real time, the Kafka operation and maintenance cluster receives the data, divides the data receiving groups and configures nodes to complete HDFS storage; When identifying potential factors, include: The operation and maintenance identification model provides: association rule antecedent and the association rule , association rule antecedent Represents the attribute characteristics of IT information operation and maintenance data, and the latter term of the association rule Indicates whether there are system problems or performance bottlenecks. Indicates a system problem. Indicates a performance bottleneck; Based on the antecedent and consequent items of association rules and , get the support of performance bottleneck ; Based on the antecedent and consequent items of association rules and , get the support level of system problems ; Will The attribute features greater than the threshold are taken as the first latent factor; Will Attribute features greater than the threshold are used as the second latent factor; When automatically adjusting IT resource configuration, include: The operation and maintenance adjustment model determines the first control coefficient and the second control coefficient according to the first potential factor and the second potential factor, and calculates the first control quantity according to formula (3), and calculates the second control quantity according to formula (4): (3); In the formula, In the first planning period The first latent factor corresponds to The first control quantity of IT resources; In the first planning period No. strip IT resource flow, for The number of IT resource flows, In the first planning period No. strip The first control coefficient of IT resources; (4); In the formula, For the second planning period The second latent factor corresponds to The second control quantity of IT-like resources; For the second planning period No. strip IT resource flow, for The number of IT resource flows, For the second planning period No. strip The second control coefficient of IT-like resources; According to formula (5), the maximum approximation of the first control quantity is performed to obtain the first IT resource optimization configuration result: (5); In the formula, In the first planning period The first latent factor corresponds to The actual amount of IT resources controlled, is the number of IT resource types corresponding to the first potential factor, is the total number of time periods in the first plan, Optimize the configuration results for the first IT resource; According to formula (6), the second control quantity is approached to the maximum extent, and the second IT resource optimization configuration result is obtained: (6); In the formula, For the second planning period The second latent factor corresponds to The actual amount of IT resources controlled, is the number of IT resource types corresponding to the second potential factor, is the total number of time periods in the second plan, Optimize configuration results for secondary IT resources.

2. The IT informatization dual-model operation and maintenance management method based on a multi-layer data lake according to claim 1 is characterized in that: When receiving data through the Kafka operation and maintenance cluster, dividing the data receiving groups and configuring nodes to complete HDFS storage, it includes: The collected IT operation and maintenance related data is transmitted to the Kafka operation and maintenance cluster through Flume; By respectively corresponding multiple instances in the Kafka operation and maintenance cluster to different data sources for receiving IT operation and maintenance related data from each data source, the multiple instances are divided into multiple data receiving groups according to a preset number; Equip each data receiving group with a corresponding node and aggregate the collected data into multiple nodes; Each node is equipped with a corresponding Producer, and the collected data is sent back to the Kafka operation and maintenance cluster through the Producer for HDFS storage.

3. The IT information dual-model operation and maintenance management method based on a multi-layer data lake according to claim 2 is characterized in that: When forming a multi-layer operation and maintenance data network, it includes: The time series data and text data in the Kafka operation and maintenance cluster are stored in the data storage method of [time_stamp][key1][value1]…[keyN][valueN]; The number of key-value pairs [key][value] at each time point corresponding to [time_stamp] is expandable, so that time series data and text data can be stored in table structures at different levels; Combined query conditions are assigned to the time series data and text data in the table structure, and aggregation calculations are performed to form the multi-layer operation and maintenance data network.

4. The IT information dual-model operation and maintenance management method based on a multi-layer data lake according to claim 3 is characterized in that: When forming a multi-layer operation and maintenance data lake, it includes: Encode the data in the table structures at different levels into numerical data sequences with the same structural form according to the preset semantic logic, upload them to the cloud server for dimensionless processing and standardization, and upload them to the cloud database through the gateway; The data sequence uploaded by the gateway is parsed and continuously monitored in real time through the cloud server; When the data sequence flow is detected to be interrupted, the data sequence is extracted through the cloud server and compared with the preset data format; When the comparison is passed, the corresponding database interface is connected according to the data type, and an identification number is assigned and converted into JSON format to obtain the actual data value, which is then converted into a triple through the mapping rules; The triples are matched with concepts and entities in the data lake database connected to the corresponding database interface to obtain a data lake middle platform, and a multi-layer API interface is provided for the data lake middle platform.

5. The IT informatization dual-model operation and maintenance management method based on a multi-layer data lake according to any one of claims 1 to 4, characterized in that: When performing quick access and retrieval, it includes: Determine the search area and target conditions according to the access and retrieval conditions, and obtain the smallest top-level data block that can completely contain the search area; Determine whether the top-level data block meets the target condition, if so, insert the data block index number into the linked list, if not, continue to search for branches of the top-level data block until all data blocks that meet the target condition are found; The indexes of all data blocks that meet the target conditions are inserted into a linked list in order from left to right and from top to bottom, and large-scale operation and maintenance data are quickly accessed and retrieved through the linked list.

6. The IT informatization dual-model operation and maintenance management method based on a multi-layer data lake according to any one of claims 1 to 4, characterized in that: When completing accurate data filtering, key feature extraction and deep integration, it includes: Filter the data using filtering rules to form a key matrix model for identifying the system's operating status; Using the Apriori algorithm to obtain frequent multinomial sets, and mining the association rules between data types and fault types in the key matrix model, to establish the several key performance evaluation state quantity sets; Among them, when mining association rules, it includes: A cluster analysis is performed on the relationship between different data types and fault types in the key matrix model, the shape coefficient and silhouette coefficient of the associated data curve are mined, the collection amount of related data types is extracted, and a key performance evaluation state quantity set for different fault types is obtained.

7. The IT informatization dual-model operation and maintenance management method based on a multi-layer data lake according to claim 1 is characterized in that: When performing dynamic display, it includes: A graph is defined, and the vertices of the identified first potential factor, the second potential factor, the first IT resource optimization configuration result, and the second IT resource optimization configuration result are determined by the host IP, the collection directory, and the file name template, respectively, and the attribute information of the vertex is set; Taking the vertices as source ends and matching corresponding transmission target ends, taking the transmission target ends one by one as vertices of the next layer of the dynamic display graph, setting their attribute information, and connecting them with the previous layer; The next layer of source information is matched, and the matched source information is used as the vertex of the next layer until no source information is matched.

Citation Information

Patent Citations

  • Big data workbench for intelligent operation and maintenance

    CN117131034A

  • Government affair system operation and maintenance method and system based on multi-source data fusion

    CN118796647A