Method and device for realizing cross-platform high-compatibility automatic unified operation and maintenance monitoring data fusion, processor and readable storage medium thereof
Through automated acquisition and machine learning, cross-platform operation and maintenance monitoring data is analyzed, the problem of poor cross-platform compatibility is solved, efficient operation and maintenance monitoring and fault location is achieved, and operation and maintenance costs and management complexity is reduced.
Patent Information
- Application Number
- CN202510620151.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-08-19
AI Technical Summary
The existing operation and maintenance monitoring technology lacks cross-platform compatibility and relies on manual configuration to achieve effective data integration and comprehensive analysis, resulting in low operation and maintenance efficiency and difficulty in fault location.
Through automated acquisition of cross-platform multi-dimensional operation and maintenance monitoring data, machine learning algorithms are used to build prediction models, conduct joint analysis of alarm information and root cause inference, and combine knowledge base search solutions to achieve high compatibility of cross-platform unified operation and maintenance monitoring.
It realizes the unified operation and maintenance monitoring of multi-platforms, reduces operation and maintenance costs and management complexity, optimizes resource configuration, improves the accuracy and efficiency of fault location, and reduces redundant alarm information.
Smart Images

Figure CN120508471A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of IT operation and maintenance, and in particular to the field of operation and maintenance monitoring technology, and specifically refers to a method, device, processor and computer-readable storage medium thereof for realizing cross-platform highly compatible automated unified operation and maintenance monitoring data fusion. Background Art
[0002] An enterprise's IT operation and maintenance environment usually contains a variety of incompatible systems, equipment, and platforms, including Windows, Linux, network equipment, servers, and cloud platforms. Traditional operation and maintenance monitoring technologies are often limited to a specific platform. The operation and maintenance monitoring technologies of different platforms are usually independent of each other, with different data formats, monitoring indicators, and alarm methods. There is a lack of cross-platform compatibility, making it difficult to conduct comprehensive operation and maintenance monitoring of the enterprise's IT system as a whole, and it is difficult to achieve effective data integration and comprehensive analysis.
[0003] Existing operation and maintenance monitoring methods mostly rely on manual configuration and operation, and configure alarm strategies based on experience. When abnormal situations occur in operation and maintenance monitoring, the accident has often already occurred, making it impossible to make advance predictions and conduct capacity planning in advance.
[0004] In actual operation and maintenance monitoring, a failure in a business module may trigger alarms in multiple modules and generate a large amount of alarm information. These alarms contain a lot of redundant information, which reduces the efficiency of operation and maintenance personnel. Summary of the Invention
[0005] The purpose of the present invention is to overcome the shortcomings of the above-mentioned existing technologies and provide a method, device, processor and computer-readable storage medium for realizing cross-platform highly compatible automated unified operation and maintenance monitoring data fusion that meets high operation and maintenance monitoring efficiency, high management level and a wide range of applications.
[0006] To achieve the above objectives, the present invention provides a method, device, processor, and computer-readable storage medium for realizing cross-platform highly compatible automated unified operation and maintenance monitoring data fusion as follows:
[0007] The method for realizing cross-platform highly compatible automated unified operation and maintenance monitoring data fusion is mainly characterized in that the method comprises the following steps:
[0008] (1) Automatically and uniformly acquire multi-dimensional operation and maintenance monitoring data across platforms, and collect and aggregate multi-dimensional operation and maintenance monitoring data from different platforms and multiple operating systems;
[0009] (2) Based on the historical operation and maintenance monitoring data obtained, a prediction model is constructed with the help of machine learning algorithms to predict the trend of operation and maintenance monitoring indicators;
[0010] (3) Perform root cause analysis on a series of alarm data of a fault, conduct joint analysis on the alarm data to classify the alarm information, search the classified alarm information through the knowledge base, and infer the root cause of the alarm.
[0011] Preferably, the automated unified acquisition of cross-platform multi-dimensional operation and maintenance monitoring data in step (1) includes automated acquisition of physical monitoring data, automated acquisition of network monitoring data, automated acquisition of system monitoring data, and automated acquisition of application monitoring data.
[0012] Preferably, the automatic acquisition of physical monitoring data specifically includes: regularly collecting the power supply, port connectivity, and CPU, memory, and disk performance indicators of physical devices in the platform, aggregating them through a unified operation and maintenance monitoring data acquisition interface, and storing the data;
[0013] The automated acquisition of network monitoring data specifically includes: regularly collecting network link connectivity, port activity status, network performance, and network link flow indicators of physical devices within the platform, aggregating them through a unified operation and maintenance monitoring data acquisition interface, and storing the data;
[0014] The automated acquisition of system monitoring data specifically includes: regularly collecting system CPU, memory, disk usage, system operating status, system open port connectivity, and network bandwidth indicators within the platform, aggregating them through a unified operation and maintenance monitoring data acquisition interface, and storing the data;
[0015] The automated acquisition of application monitoring data specifically includes: regularly collecting application traffic conditions, process status, response time, distributed tracking request flow and detailed process information indicators within the platform, aggregating them through a unified operation and maintenance monitoring data acquisition interface, and storing the data.
[0016] Preferably, the step (2) specifically includes the following steps:
[0017] (2.1) Build and train neural network models;
[0018] (2.2) Prediction and analysis based on neural network models.
[0019] Preferably, the step (2.1) specifically includes the following steps:
[0020] (2.1.1) Obtain historical operation and maintenance monitoring data;
[0021] (2.1.2) Convert historical operation and maintenance monitoring data into the format required by the LSTM model and divide the data into training and validation sets;
[0022] (2.1.3) Initialize the parameters of the LSTM model and build the LSTM structure as needed;
[0023] (2.1.4) Train and verify the model to generate an LSTM model.
[0024] Preferably, the step (2.2) specifically includes the following steps:
[0025] (2.2.1) Obtain real-time operation and maintenance monitoring data;
[0026] (2.2.2) Convert real-time operation and maintenance monitoring data into the format required by the LSTM model;
[0027] (2.2.3) Input the converted data into the model and output the prediction results of the corresponding indicators;
[0028] (2.2.4) Compare the prediction results and if they reach or exceed the set threshold, an alarm will be triggered.
[0029] Preferably, the step (3) specifically includes the following steps:
[0030] (3.1) Extracting alarm data feature information and extracting segmentation features from the alarm data according to the optimal segmentation basis;
[0031] (3.2) Establish a joint analysis model for alarm data;
[0032] (3.3) Generate root cause alarm information based on the knowledge base.
[0033] Preferably, the step (3.3) is specifically as follows:
[0034] Alarms are classified according to the joint analysis model of alarm data. Based on the classification results, the knowledge base is searched to match the same or similar fault cases and solutions to form root cause alarms.
[0035] The device for realizing cross-platform highly compatible automated unified operation and maintenance monitoring data fusion has the following main features:
[0036] a processor configured to execute computer-executable instructions;
[0037] The memory stores one or more computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the above-mentioned method for realizing cross-platform highly compatible automated unified operation and maintenance monitoring data fusion are implemented.
[0038] The main feature of the processor for realizing cross-platform highly compatible automated unified operation and maintenance monitoring data fusion is that the processor is configured to execute computer-executable instructions. When the computer-executable instructions are executed by the processor, the various steps of the above-mentioned method for realizing cross-platform highly compatible automated unified operation and maintenance monitoring data fusion are implemented.
[0039] The main feature of this computer-readable storage medium is that a computer program is stored thereon, and the computer program can be executed by a processor to implement the various steps of the above-mentioned method for achieving cross-platform highly compatible automated unified operation and maintenance monitoring data fusion.
[0040] The method, device, processor and computer-readable storage medium for realizing cross-platform highly compatible automated unified operation and maintenance monitoring data fusion of the present invention are adopted. Through the automated acquisition of cross-platform multi-dimensional operation and maintenance monitoring data, unified operation and maintenance monitoring of multiple operating systems and hardware platforms is realized, and the operation and maintenance monitoring cost and management complexity are reduced; through the automated unified acquisition interface, multi-dimensional operation and maintenance monitoring data acquisition of different platforms and devices is realized, solving the problem of poor cross-platform compatibility; through multi-dimensional operation and maintenance monitoring data fusion, intelligent algorithms are used to deeply mine data, and historical and real-time operation and maintenance monitoring data trends are analyzed, so as to predict resource requirements for different applications, optimize resource allocation, and ensure efficient operation of the system; for the alarm information generated by operation and maintenance monitoring, alarm information classification is realized through alarm joint analysis, the number and type of active alarms are reduced, and the root causes and solutions of alarms are inferred by combining the knowledge base, reducing redundant information, solving the problems of low manual handling efficiency and low system availability under a large number of alarms, improving the accuracy of fault location, and improving the efficiency of precise fault location. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 This is a flow chart of the method for realizing cross-platform highly compatible automated unified operation and maintenance monitoring data fusion of the present invention.
[0042] Figure 2 This is a schematic diagram of the automated unified acquisition of operation and maintenance monitoring data for the method of realizing cross-platform high-compatibility automated unified operation and maintenance monitoring data fusion of the present invention.
[0043] Figure 3 This is a schematic diagram of the fusion prediction of operation and maintenance monitoring data of the method for realizing cross-platform high-compatibility automated unified operation and maintenance monitoring data fusion of the present invention.
[0044] Figure 4 A schematic diagram of root cause inference based on alarm joint analysis of the method for realizing cross-platform highly compatible automated unified operation and maintenance monitoring data fusion of the present invention. DETAILED DESCRIPTION
[0045] In order to more clearly describe the technical content of the present invention, further description is given below in conjunction with specific embodiments.
[0046] The method of realizing cross-platform highly compatible automated unified operation and maintenance monitoring data fusion of the present invention comprises the following steps:
[0047] (1) Automatically and uniformly acquire multi-dimensional operation and maintenance monitoring data across platforms, and collect and aggregate multi-dimensional operation and maintenance monitoring data from different platforms and multiple operating systems;
[0048] (2) Based on the historical operation and maintenance monitoring data obtained, a prediction model is constructed with the help of machine learning algorithms to predict the trend of operation and maintenance monitoring indicators;
[0049] (3) Perform root cause analysis on a series of alarm data of a fault, conduct joint analysis on the alarm data to classify the alarm information, search the classified alarm information through the knowledge base, and infer the root cause of the alarm.
[0050] As a preferred embodiment of the present invention, the automated unified acquisition of cross-platform multi-dimensional operation and maintenance monitoring data described in step (1) includes automated acquisition of physical monitoring data, automated acquisition of network monitoring data, automated acquisition of system monitoring data, and automated acquisition of application monitoring data.
[0051] As a preferred embodiment of the present invention, the automated acquisition of physical monitoring data specifically includes: regularly collecting the power supply, port connectivity, and CPU, memory, and disk performance indicators of physical devices in the platform, aggregating the data through a unified operation and maintenance monitoring data acquisition interface, and storing the data;
[0052] The automated acquisition of network monitoring data specifically includes: regularly collecting network link connectivity, port activity status, network performance, and network link flow indicators of physical devices within the platform, aggregating them through a unified operation and maintenance monitoring data acquisition interface, and storing the data;
[0053] The automated acquisition of system monitoring data specifically includes: regularly collecting system CPU, memory, disk usage, system operating status, system open port connectivity, and network bandwidth indicators within the platform, aggregating them through a unified operation and maintenance monitoring data acquisition interface, and storing the data;
[0054] The automated acquisition of application monitoring data specifically includes: regularly collecting application traffic conditions, process status, response time, distributed tracking request flow and detailed process information indicators within the platform, aggregating them through a unified operation and maintenance monitoring data acquisition interface, and storing the data.
[0055] As a preferred embodiment of the present invention, the step (2) specifically includes the following steps:
[0056] (2.1) Build and train neural network models;
[0057] (2.2) Prediction and analysis based on neural network models.
[0058] As a preferred embodiment of the present invention, the step (2.1) specifically includes the following steps:
[0059] (2.1.1) Obtain historical operation and maintenance monitoring data;
[0060] (2.1.2) Convert historical operation and maintenance monitoring data into the format required by the LSTM model and divide the data into training and validation sets;
[0061] (2.1.3) Initialize the parameters of the LSTM model and build the LSTM structure as needed;
[0062] (2.1.4) Train and verify the model to generate an LSTM model.
[0063] As a preferred embodiment of the present invention, the step (2.2) specifically includes the following steps:
[0064] (2.2.1) Obtain real-time operation and maintenance monitoring data;
[0065] (2.2.2) Convert real-time operation and maintenance monitoring data into the format required by the LSTM model;
[0066] (2.2.3) Input the converted data into the model and output the prediction results of the corresponding indicators;
[0067] (2.2.4) Compare the prediction results and if they reach or exceed the set threshold, an alarm will be triggered.
[0068] As a preferred embodiment of the present invention, the step (3) specifically includes the following steps:
[0069] (3.1) Extracting alarm data feature information and extracting segmentation features from the alarm data according to the optimal segmentation basis;
[0070] (3.2) Establish a joint analysis model for alarm data;
[0071] (3.3) Generate root cause alarm information based on the knowledge base.
[0072] As a preferred embodiment of the present invention, the step (3.3) is specifically as follows:
[0073] Alarms are classified according to the joint analysis model of alarm data. Based on the classification results, the knowledge base is searched to match the same or similar fault cases and solutions to form root cause alarms.
[0074] The device for realizing cross-platform highly compatible automated unified operation and maintenance monitoring data fusion of the present invention comprises:
[0075] a processor configured to execute computer-executable instructions;
[0076] The memory stores one or more computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the above-mentioned method for realizing cross-platform highly compatible automated unified operation and maintenance monitoring data fusion are implemented.
[0077] The processor of the present invention is used to realize cross-platform highly compatible automated unified operation and maintenance monitoring data fusion, wherein the processor is configured to execute computer executable instructions. When the computer executable instructions are executed by the processor, the various steps of the above-mentioned method for realizing cross-platform highly compatible automated unified operation and maintenance monitoring data fusion are realized.
[0078] The computer-readable storage medium of the present invention stores a computer program thereon, and the computer program can be executed by a processor to implement the various steps of the above-mentioned method for achieving cross-platform highly compatible automated unified operation and maintenance monitoring data fusion.
[0079] In a specific embodiment of the present invention, a cross-platform, highly compatible, automated, unified operation and maintenance monitoring fusion method is provided to address the problems of existing operation and maintenance monitoring technologies, such as lack of cross-platform compatibility, reliance on manual configuration and operation, inability to make advance predictions, and low efficiency in handling operation and maintenance alarms.
[0080] like Figure 1 As shown, the cross-platform highly compatible automated unified operation and maintenance monitoring fusion method of the present invention includes:
[0081] S100, automatic and unified acquisition of cross-platform multi-dimensional operation and maintenance monitoring data, collection and aggregation of multi-dimensional operation and maintenance monitoring data of different platforms and multiple operating systems.
[0082] Automatic acquisition of cross-platform and multi-dimensional operation and maintenance monitoring data Figure 2 As shown, it includes the automated acquisition of physical monitoring data, network monitoring data, system monitoring data, and application monitoring data.
[0083] Automated acquisition of physical monitoring data refers to the periodic collection of power supply, port connectivity, and CPU, memory, and disk performance indicators of physical devices within the platform, which are aggregated through a unified operation and maintenance monitoring data acquisition interface and stored.
[0084] In a specific embodiment, the automated acquisition of physical monitoring data refers to regularly collecting the power supply, port connectivity, and CPU, memory, and disk performance indicators of physical devices in the platform through protocols or logs such as snmp, IPMI, and lldp, and calling tool APIs, aggregating the data through a unified operation and maintenance monitoring data acquisition interface, and storing the data in a time series format in clickhouse or other time series databases.
[0085] Automated acquisition of network monitoring data refers to the regular collection of network link connectivity, port activity status, network performance, and network link flow indicators of physical devices within the platform, which are aggregated through a unified operation and maintenance monitoring data acquisition interface and stored.
[0086] In a specific embodiment, the automated acquisition of network monitoring data refers to regularly collecting network link connectivity, port activity status, network performance, and network link traffic indicators of physical devices in the platform through protocols or logs such as snmp, IPMI, and lldp, and calling tool APIs, aggregating the data through a unified operation and maintenance monitoring data acquisition interface, and storing the data in a time series format in clickhouse or other time series databases.
[0087] The automated acquisition of system monitoring data refers to the periodic collection of system CPU, memory, disk usage, system operating status, system open port connectivity and network bandwidth indicators within the platform, which are aggregated through a unified operation and maintenance monitoring data acquisition interface and stored.
[0088] In a specific embodiment, the automated acquisition of system monitoring data refers to regularly collecting system CPU, memory, disk usage, system operating status, system open port connectivity, and network bandwidth indicators by deploying node_exporter or directly calling the tool API, aggregating the data through a unified operation and maintenance monitoring data acquisition interface, and storing the data in a time series format in clickhouse or other time series databases.
[0089] Automated acquisition of application monitoring data refers to the periodic collection of application traffic conditions, process status, response time, distributed tracing request flows, and detailed process information indicators within the platform, which are aggregated through a unified operation and maintenance monitoring data acquisition interface and stored.
[0090] In a specific embodiment, the automated acquisition of application monitoring data refers to regularly collecting application traffic, process status, response time, distributed tracing request flow, and detailed process information indicators by deploying an agent for the application development language or directly calling the tool API, aggregating the data through a unified operation and maintenance monitoring data acquisition interface, and storing the data in a time series format in ClickHouse or other time series databases.
[0091] S110, Operation and maintenance monitoring data fusion prediction, refers to the process of building a model and making predictions based on the acquired historical operation and maintenance monitoring data with the help of machine learning algorithms; using the operation and maintenance monitoring data to build a prediction model to predict the trend of operation and maintenance monitoring indicators, and triggering an alarm when the indicator threshold is reached or exceeded.
[0092] In a preferred embodiment, the operation and maintenance monitoring data fusion prediction is as follows Figure 3 As shown, it includes building and training neural network models, and making predictions and analyses based on the models.
[0093] In a specific embodiment, building and training a neural network model involves preprocessing and feature engineering the acquired historical operation and maintenance monitoring data to form a training sample set, training the model using a neural network algorithm, and outputting a neural network model. In a preferred embodiment, the neural network algorithm used is a long short-term memory (LSTM) algorithm, which specifically includes the following steps:
[0094] (1) Obtaining historical operation and maintenance monitoring data, which may be historical operation and maintenance monitoring data of one or more indicators, including CPU usage, memory usage, disk usage, application traffic, application response time, network link traffic, etc.
[0095] (2) Convert the historical operation and maintenance monitoring data into the format required by the LSTM model and divide the data into a training set and a validation set.
[0096] (3) Initialize the parameters of the LSTM model and build the LSTM structure as needed.
[0097] (4) Train and verify the model to complete the generation of the LSTM model.
[0098] In a specific embodiment, model-based prediction and analysis involves preprocessing and feature engineering acquired real-time operation and maintenance monitoring data, inputting it into a neural network model for prediction and outputting prediction results. Analysis is then performed based on the prediction results, and an alert is issued if an outlier is detected. An outlier refers to a predicted metric reaching or exceeding a threshold. The process specifically includes the following steps.
[0099] (1) Obtaining real-time operation and maintenance monitoring data, which can be historical operation and maintenance monitoring data of one or more indicators, including CPU usage, memory usage, disk usage, application traffic, application response time, network link traffic, etc.
[0100] (2) Convert real-time operation and maintenance monitoring data into the format required by the LSTM model.
[0101] (3) Input the converted data into the model and output the prediction results of the corresponding indicators.
[0102] (4) Compare the prediction results and trigger an alarm if the set threshold is reached or exceeded.
[0103] S120 root cause inference based on joint alarm analysis refers to the process of performing root cause analysis on a series of alarm data of a fault to identify the root cause alarm. The alarm data is jointly analyzed and combined with knowledge base search to infer the root cause of the alarm.
[0104] In a preferred embodiment, the root cause inference based on the joint analysis of alarms is as follows: Figure 4 As shown, it includes extraction of alarm data feature information, establishment of alarm data joint analysis model, and generation of root cause alarm information based on the knowledge base.
[0105] In a specific embodiment, alarm data feature information extraction involves extracting partitioning features from the alarm data based on optimal partitioning criteria. Features include alarm time, alarm type, alarm device, and alarm level. The partitioning criteria use metrics such as information gain, information gain ratio, and Gini index to assess the importance of each feature in distinguishing different alarm associations. The feature with the highest information gain and information gain ratio and the lowest Gini index is selected as the partitioning feature for the decision tree.
[0106] In one specific embodiment, establishing a joint analysis model for alarm data involves building a decision tree model based on selected feature information. Each internal node in the decision tree represents an alarm feature, branches represent feature values, and leaf nodes represent associated results. To reduce the risk of overfitting the decision tree, strategies such as limiting the maximum number of branches and the minimum number of leaf node samples can be used to proactively remove some branches.
[0107] In a specific embodiment, knowledge-based root cause alarm generation involves classifying alarms based on a joint analysis model of alarm data. Based on the classification results, a knowledge base search is performed to match identical or similar fault cases and solutions to generate root cause alarms. A knowledge base refers to various operational knowledge, events, and problem solutions collected and stored in the system according to specific classification rules to address various operational issues. Knowledge base search refers to the process of finding key knowledge through techniques such as exact or fuzzy keyword matching, full-text search, and semantic search.
[0108] (1) The sources of the knowledge base include knowledge from external channels such as manufacturer documents, technical forums, industry reports, as well as the experience accumulated by the operation and maintenance team in their work and records of problem solving.
[0109] (2) The classification of knowledge is carried out by adopting multi-category label classification and storage in the database, and the knowledge system, business system, knowledge structure, etc. to which the knowledge belongs are marked.
[0110] (3) Search the knowledge base through keyword precision, fuzzy matching, and full-text search.
[0111] The key step in the technical solution of this invention is to automatically collect operation and maintenance monitoring data, not to collect operation and maintenance knowledge. The core step of the technical solution of this invention involves analyzing and predicting the collected operation and maintenance monitoring data indicators based on a neural network model. This core step is based on root cause inference based on joint alarm analysis. This first classifies large amounts of alarm data using a decision tree model, and then conducts root cause inference by combining it with a knowledge base search.
[0112] The automated collection and aggregation of multi-dimensional operation and maintenance monitoring data is a key prerequisite for subsequent fusion analysis.
[0113] The acquired historical operation and maintenance monitoring data is preprocessed and feature-engineered to form a training sample set. A neural network algorithm is then used for training and outputting a neural network model. Real-time operation and maintenance monitoring data is preprocessed and feature-engineered, input into the neural network model for prediction, and the prediction results are output. The prediction results are analyzed, and if any abnormal values are detected, an alert is issued. This step allows resource demand to be predicted in advance based on the predicted values, facilitating more rational resource allocation.
[0114] Establishing a joint analysis model for alarm data to classify monitoring alarm data. In actual operations and maintenance monitoring, a resource anomaly may trigger a series of related alarms. Directly identifying faults for each alarm based on a knowledge graph without prioritizing these alarms can lead to misjudgments and fail to improve operation and maintenance efficiency, as the number of alarms requiring resolution remains unreduced. Therefore, this solution classifies alarms based on the joint analysis model for alarm data. Based on the classification results, searching the knowledge base to match identical or similar fault cases and solutions is performed, resulting in root cause alarms, which is both more effective and necessary.
[0115] The technical solution of the present invention combines the above monitoring characteristics to realize multi-dimensional unified monitoring, and can achieve fusion analysis of monitoring data of different dimensions. The present invention proposes a prediction method for constructing a neural network model to realize learning and training of the collected monitoring data so as to predict the future trend of operation and maintenance monitoring indicators. The root cause inference based on alarm joint analysis provided by the present invention can greatly reduce the number of alarms to be processed by classifying alarm data, and infer the root cause of the alarm based on the knowledge base after classification. The specific implementation scheme of this embodiment can be found in the relevant description in the above embodiment, and will not be repeated here.
[0116] It can be understood that the same or similar parts of the above embodiments can be referenced to each other, and the contents not described in detail in some embodiments can refer to the same or similar contents in other embodiments.
[0117] It should be noted that, in the description of the present invention, the terms "first", "second", etc. are used for descriptive purposes only and should not be understood as indicating or implying relative importance. In addition, in the description of the present invention, unless otherwise specified, the meaning of "plurality" is at least two.
[0118] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present invention includes alternative implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present invention pertain.
[0119] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution device. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0120] Those skilled in the art will understand that all or part of the steps in the method for implementing the above-mentioned embodiment can be completed by instructing related hardware through a program, and the corresponding program can be stored in a computer-readable storage medium. When the program is executed, it includes one of the steps of the method embodiment or a combination thereof.
[0121] Furthermore, the functional units in the various embodiments of the present invention may be integrated into a single processing module, each unit may exist physically separately, or two or more units may be integrated into a single module. The aforementioned integrated modules may be implemented in the form of hardware or software functional modules. If the integrated modules are implemented in the form of software functional modules and sold or used as independent products, they may also be stored in a computer-readable storage medium.
[0122] The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk, etc.
[0123] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0124] The method, device, processor and computer-readable storage medium for realizing cross-platform highly compatible automated unified operation and maintenance monitoring data fusion of the present invention are adopted. Through the automated acquisition of cross-platform multi-dimensional operation and maintenance monitoring data, unified operation and maintenance monitoring of multiple operating systems and hardware platforms is realized, and the operation and maintenance monitoring cost and management complexity are reduced; through the automated unified acquisition interface, multi-dimensional operation and maintenance monitoring data acquisition of different platforms and devices is realized, solving the problem of poor cross-platform compatibility; through multi-dimensional operation and maintenance monitoring data fusion, intelligent algorithms are used to deeply mine data, and historical and real-time operation and maintenance monitoring data trends are analyzed, so as to predict resource requirements for different applications, optimize resource allocation, and ensure efficient operation of the system; for the alarm information generated by operation and maintenance monitoring, alarm information classification is realized through alarm joint analysis, the number and type of active alarms are reduced, and the root causes and solutions of alarms are inferred by combining the knowledge base, reducing redundant information, solving the problems of low manual handling efficiency and low system availability under a large number of alarms, improving the accuracy of fault location, and improving the efficiency of precise fault location.
[0125] In this specification, the present invention has been described with reference to specific embodiments thereof. However, it will be apparent that various modifications and variations may be made without departing from the spirit and scope of the present invention. Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive.
Claims
1. A method for realizing cross-platform highly compatible automated unified operation and maintenance monitoring data fusion, characterized in that: The method comprises the following steps: (1) Automatically and uniformly acquire multi-dimensional operation and maintenance monitoring data across platforms, and collect and aggregate multi-dimensional operation and maintenance monitoring data from different platforms and multiple operating systems; (2) Based on the historical operation and maintenance monitoring data obtained, a prediction model is constructed with the help of machine learning algorithms to predict the trend of operation and maintenance monitoring indicators; (3) Perform root cause analysis on a series of alarm data of a fault, conduct joint analysis on the alarm data to classify the alarm information, search the classified alarm information through the knowledge base, and infer the root cause of the alarm.
2. The method for realizing cross-platform highly compatible automated unified operation and maintenance monitoring data fusion according to claim 1 is characterized in that: The automated unified acquisition of cross-platform and multi-dimensional operation and maintenance monitoring data described in step (1) includes automated acquisition of physical monitoring data, automated acquisition of network monitoring data, automated acquisition of system monitoring data, and automated acquisition of application monitoring data.
3. The method for realizing cross-platform highly compatible automated unified operation and maintenance monitoring data fusion according to claim 2 is characterized in that: The automated acquisition of physical monitoring data specifically includes: regularly collecting the power supply, port connectivity, and CPU, memory, and disk performance indicators of physical devices in the platform, aggregating them through a unified operation and maintenance monitoring data acquisition interface, and storing the data; The automated acquisition of network monitoring data specifically includes: regularly collecting network link connectivity, port activity status, network performance, and network link flow indicators of physical devices within the platform, aggregating them through a unified operation and maintenance monitoring data acquisition interface, and storing the data; The automated acquisition of system monitoring data specifically includes: regularly collecting system CPU, memory, disk usage, system operating status, system open port connectivity, and network bandwidth indicators within the platform, aggregating them through a unified operation and maintenance monitoring data acquisition interface, and storing the data; The automated acquisition of application monitoring data specifically includes: regularly collecting application traffic conditions, process status, response time, distributed tracking request flow and detailed process information indicators within the platform, aggregating them through a unified operation and maintenance monitoring data acquisition interface, and storing the data.
4. The method for realizing cross-platform highly compatible automated unified operation and maintenance monitoring data fusion according to claim 1 is characterized in that: The step (2) specifically includes the following steps: (2.1) Build and train neural network models; (2.2) Prediction and analysis based on neural network models.
5. The method for realizing cross-platform highly compatible automated unified operation and maintenance monitoring data fusion according to claim 4 is characterized in that: The step (2.1) specifically includes the following steps: (2.1.1) Obtain historical operation and maintenance monitoring data; (2.1.2) Convert historical operation and maintenance monitoring data into the format required by the LSTM model and divide the data into training and validation sets; (2.1.3) Initialize the parameters of the LSTM model and build the LSTM structure as needed; (2.1.4) Train and verify the model to generate an LSTM model.
6. The method for realizing cross-platform highly compatible automated unified operation and maintenance monitoring data fusion according to claim 4 is characterized in that: The step (2.2) specifically includes the following steps: (2.2.1) Obtain real-time operation and maintenance monitoring data; (2.2.2) Convert real-time operation and maintenance monitoring data into the format required by the LSTM model; (2.2.3) Input the converted data into the model and output the prediction results of the corresponding indicators; (2.2.4) Compare the prediction results and if they reach or exceed the set threshold, an alarm will be triggered.
7. The method for realizing cross-platform highly compatible automated unified operation and maintenance monitoring data fusion according to claim 1 is characterized in that: The step (3) specifically includes the following steps: (3.1) Extracting alarm data feature information and extracting segmentation features from the alarm data according to the optimal segmentation basis; (3.2) Establish a joint analysis model for alarm data; (3.3) Generate root cause alarm information based on the knowledge base.
8. The method for realizing cross-platform highly compatible automated unified operation and maintenance monitoring data fusion according to claim 7 is characterized in that: The step (3.3) is specifically as follows: Alarms are classified according to the joint analysis model of alarm data. Based on the classification results, the knowledge base is searched to match the same or similar fault cases and solutions to form root cause alarms.
9. A device for realizing cross-platform highly compatible automated unified operation and maintenance monitoring data fusion, characterized in that: The device comprises: a processor configured to execute computer-executable instructions; A memory storing one or more computer-executable instructions, wherein when the computer-executable instructions are executed by the processor, the steps of the method for realizing cross-platform highly compatible automated unified operation and maintenance monitoring data fusion as described in any one of claims 1 to 8 are implemented.
10. A processor for realizing cross-platform highly compatible automated unified operation and maintenance monitoring data fusion, characterized in that: The processor is configured to execute computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the method for realizing cross-platform highly compatible automated unified operation and maintenance monitoring data fusion described in any one of claims 1 to 8 are implemented.
11. A computer-readable storage medium, characterized in that A computer program is stored thereon, and the computer program can be executed by a processor to implement the various steps of the method for realizing cross-platform highly compatible automated unified operation and maintenance monitoring data fusion as described in any one of claims 1 to 8.