IT fault detection method and system based on operation and maintenance AI large model
By applying the operation and maintenance AI model in IT fault detection, extracting and analyzing abnormal data from the IT system, the problem of detecting failure types in the existing technology has been solved, and higher fault detection accuracy and completeness are achieved.
Patent Information
- Application Number
- CN202510101882.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-16
AI Technical Summary
The existing IT fault detection methods rely on the rules engine and are difficult to detect unexperienced fault types, resulting in reduced detection accuracy.
The IT fault detection method based on the operation and maintenance AI model is adopted. By obtaining the system operation data of the IT system, abnormal data is extracted, and feature extraction and fault analysis is used to perform feature extraction and fault analysis. Combined with equipment fault analysis, the fault detection results of the IT system are generated.
It improves the accuracy of IT fault detection, can detect unexperienced fault types, and avoids the problem of omissions in fault analysis.
Smart Images

Figure CN120011123A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of IT system fault detection, and in particular to an IT fault detection method and system based on an operation and maintenance AI big model. Background Art
[0002] IT is a series of technologies and methods that use electronic devices and computer technology to store, process, transmit and manage information. Currently, IT is widely used in various industries and fields, including enterprise, government, education, medical care, communications, finance, etc. Since there is important information data in the fields of IT application, it is necessary to perform fault detection on IT to ensure the stability of IT.
[0003] The existing IT fault detection method adopts the rule engine method: according to the characteristics of the IT system, a series of rules are formulated, and the system operation status and parameters are monitored and matched with the rules to determine whether a fault has occurred. For example, the CPU usage exceeds the threshold, the disk space is insufficient, the CPU utilization, memory usage, network traffic, etc. The collected monitoring data is matched with the rule set to determine whether the conditions of any rule are met. If a rule is met, it can be considered that the system has failed. However, this method is only applicable to the fault types that have occurred, and cannot detect and identify the fault types that have not occurred, resulting in reduced accuracy of IT fault detection. Summary of the invention
[0004] The present invention provides an IT fault detection method and system based on an operation and maintenance AI big model, the main purpose of which is to improve the accuracy of IT fault detection.
[0005] To achieve the above objectives, the present invention provides an IT fault detection method based on an operation and maintenance AI big model, comprising:
[0006] Acquire an IT system to be fault-checked, the IT system comprising: a software system and an equipment system, collect system operation data of each operation subsystem in the software system, and extract abnormal operation data from the system operation data;
[0007] Input the abnormal operation data as input data into the trained operation and maintenance AI big model, use the feature extraction layer in the operation and maintenance AI big model to extract features from the abnormal operation data to obtain abnormal features, use the timing convolution layer in the operation and maintenance AI big model to analyze the timing relationship between the abnormal features, and use the loop analysis layer in the operation and maintenance AI big model to analyze the system fault corresponding to the software system based on the timing relationship and the abnormal features;
[0008] Collecting the equipment operation data corresponding to each device in the equipment system in real time, and analyzing the equipment failure corresponding to each device in the equipment system by using the long short-term memory network in the operation and maintenance AI big model according to the equipment operation data;
[0009] The system failure and the equipment failure are combined to obtain a failure detection result corresponding to the IT system.
[0010] Optionally, extracting abnormal operation data from the system operation data includes:
[0011] Dispatching data logs corresponding to the system operation data;
[0012] Calculating the operation error rate corresponding to the system operation data according to the data log;
[0013] Performing mapping processing on the system operation data to obtain a data mapping value;
[0014] Calculating the data variance value corresponding to the system operation data according to the data mapping value;
[0015] Abnormal operation data in the system operation data is identified according to the data variance value and the operation error rate.
[0016] Optionally, calculating the data variance value corresponding to the system operation data according to the data mapping value includes:
[0017] The data variance value corresponding to the system operation data is calculated by the following formula:
[0018]
[0019] Among them, A represents the data variance value corresponding to the system operation data, a represents the sequence number corresponding to the system operation data, β represents the number of data of the system operation data, and B a represents the expected value of the ath data in the system operation data, d a Indicates the data mapping value of the ath data in the system operation data. Indicates the average mapping value of system operation data.
[0020] Optionally, the extracting features of the abnormal operation data using the feature extraction layer in the operation and maintenance AI big model to obtain abnormal features includes:
[0021] Using the input neural unit in the feature extraction layer to perform data optimization processing on the abnormal operation data to obtain optimized abnormal data;
[0022] Using the convolutional neural unit in the feature extraction layer to extract features from the optimized abnormal data to obtain initial abnormal features;
[0023] Performing dimensionality reduction processing on the initial abnormal features by using the pooling function in the feature extraction layer to obtain reduced-dimensional abnormal features;
[0024] Calculating the feature gain value corresponding to the dimension reduction abnormal feature using the entropy function in the feature extraction layer;
[0025] According to the feature gain value, the output neural unit in the feature extraction layer is used to perform output processing on the dimension reduction abnormal feature to obtain the abnormal feature.
[0026] Optionally, the calculating the feature gain value corresponding to the dimension reduction abnormal feature by using the entropy function in the feature extraction layer includes:
[0027] The entropy function includes:
[0028]
[0029] Where D represents the feature gain value corresponding to the dimension reduction anomaly feature, E(c) represents the occurrence probability corresponding to the c-th feature in the dimension reduction anomaly feature, c represents the feature sequence number corresponding to the dimension reduction anomaly feature, |Fc| represents the feature subset containing the c-th feature in the dimension reduction anomaly feature, |ω| represents all feature subsets in the dimension reduction anomaly feature, and φc represents the number of feature subsets containing the c-th feature in the dimension reduction anomaly feature.
[0030] Optionally, analyzing the temporal relationship between the abnormal features by using the temporal convolution layer in the operation and maintenance AI big model includes:
[0031] Using the recognition function in the temporal convolution layer to identify the temporal dimension corresponding to the abnormal feature;
[0032] According to the time series dimension, a time series window corresponding to the abnormal feature is set, and a convolution step corresponding to the convolution kernel in the time series convolution layer is defined;
[0033] According to the convolution step size and the time series window, the abnormal feature is convolved using the convolution kernel to obtain a convolution time series feature;
[0034] Calculating the correlation coefficient between the convolutional time series features using the linear function in the time series convolution layer;
[0035] The time series relationship between the abnormal features is determined according to the convolution time series features and the correlation coefficient.
[0036] Optionally, analyzing the system failure corresponding to the software system by using the loop analysis layer in the operation and maintenance AI big model according to the time sequence relationship and the abnormal characteristics includes:
[0037] According to the time series relationship, using the hierarchical relationship analysis network in the cyclic analysis layer to identify the cascade abnormality features in the abnormality features;
[0038] Query the corresponding system functions in the software system and analyze the functional attributes corresponding to the system functions;
[0039] Analyzing system dependencies between the software systems based on the functional attributes;
[0040] According to the system dependency, the traceability neural network in the loop analysis layer is used to analyze the triggering abnormality features in the cascade abnormality features;
[0041] Analyzing the characteristic fault corresponding to the triggering abnormal feature by using the fault diagnosis network in the loop analysis layer;
[0042] According to the characteristic fault, a system fault corresponding to the software system is obtained.
[0043] Optionally, analyzing the device failure corresponding to each device in the device system using the long short-term memory network in the operation and maintenance AI big model according to the device operation data includes:
[0044] Using the clustering function in the long short-term memory network to cluster the device operation data, so as to obtain clustered operation data;
[0045] Using the attention mechanism in the long short-term memory network to identify data variables in the clustering operation data;
[0046] Calculating variable weights corresponding to the data variables, and extracting key variables from the data variables according to the variable weights;
[0047] Calculating data values corresponding to the clustering operation data using a linear regression function in the long short-term memory network;
[0048] Combining the key variables and the data values, constructing an operation trend graph corresponding to each device in the device system;
[0049] According to the operation trend diagram, using the recursive neural unit in the long short-term memory network to analyze the fault variables in the key variables;
[0050] The fault factors corresponding to the fault variables are analyzed, and fault merging processing is performed on the fault factors to obtain the device fault corresponding to each device in the device system.
[0051] Optionally, calculating the variable weight corresponding to the data variable includes:
[0052] The variable weight corresponding to the data variable is calculated by the following formula:
[0053]
[0054] Among them, H represents the variable weight corresponding to the data variable, t represents the number of data variables, j represents the variable sequence number corresponding to the data variable, and M j+1 Indicates the effect value between the j+1th variable and other variables in the data variable, M j Represents the effect value between the j-th variable and other variables in the data variable.
[0055] An IT fault detection system based on an operation and maintenance AI big model, characterized in that the system comprises:
[0056] A data extraction module is used to obtain an IT system to be detected for faults, wherein the IT system includes a software system and an equipment system, collect system operation data of each operation subsystem in the software system, and extract abnormal operation data from the system operation data;
[0057] A system fault analysis module, used to input the abnormal operation data as input data into the trained operation and maintenance AI big model, use the feature extraction layer in the operation and maintenance AI big model to extract features from the abnormal operation data to obtain abnormal features, use the timing convolution layer in the operation and maintenance AI big model to analyze the timing relationship between the abnormal features, and use the loop analysis layer in the operation and maintenance AI big model to analyze the system fault corresponding to the software system based on the timing relationship and the abnormal features;
[0058] An equipment failure analysis module is used to collect equipment operation data corresponding to each equipment in the equipment system in real time, and analyze the equipment failure corresponding to each equipment in the equipment system using the long short-term memory network in the operation and maintenance AI big model according to the equipment operation data;
[0059] The detection result generating module is used to obtain a fault detection result corresponding to the IT system by combining the system fault and the equipment fault.
[0060] The present invention can obtain abnormal data in the system operation data by extracting abnormal operation data from the system operation data, so as to facilitate the subsequent analysis of system failures based on the abnormal operation data. The present invention can obtain data-specific attributes in the abnormal operation data by using the feature extraction layer in the operation and maintenance AI big model to extract features from the abnormal operation data, so as to facilitate the subsequent improvement of the accuracy of system failure analysis. The present invention analyzes the device failure corresponding to each device in the device system based on the device operation data and the device parameters using the long short-term memory network in the operation and maintenance AI big model, and can find the device problem corresponding to each device in the device system, and make corresponding repair processing as soon as possible to reduce the impact caused by the failure. The present invention combines the system failure and the device failure to obtain the fault detection result corresponding to the IT system, and performs fault detection on the software and hardware of the IT system, thereby improving the accuracy of fault analysis and avoiding the problem of omission of fault analysis. Therefore, the IT fault detection method and system based on the operation and maintenance AI big model provided by the embodiment of the present invention can improve the accuracy of IT fault detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 A flowchart of an IT fault detection method based on an operation and maintenance AI big model provided by an embodiment of the present invention;
[0062] Figure 2 A functional module diagram of an IT fault detection system based on an operation and maintenance AI big model provided in one embodiment of the present invention.
[0063] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings in conjunction with the embodiments. DETAILED DESCRIPTION
[0064] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0065] The embodiment of the present application provides an IT fault detection method based on an operation and maintenance AI big model. In the embodiment of the present application, the execution subject of the IT fault detection method based on the operation and maintenance AI big model includes but is not limited to at least one of the electronic devices such as the server, the terminal, etc. that can be configured to execute the method provided by the embodiment of the present application. In other words, the IT fault detection method based on the operation and maintenance AI big model can be executed by software or hardware installed on the terminal device or the server device, and the software can be a blockchain platform. The server includes but is not limited to: a single server, a server cluster, a cloud server or a cloud server cluster, etc. The server can be an independent server, or it can be a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (Content Delivery Network, CDN), and basic cloud computing services such as big data and artificial intelligence platforms.
[0066] Reference Figure 1 As shown, it is a flow chart of an IT fault detection method based on an operation and maintenance AI big model provided by an embodiment of the present invention. In this embodiment, the IT fault detection method based on an operation and maintenance AI big model includes steps S1-S4.
[0067] S1. Acquire an IT system to be fault-checked, the IT system comprising: a software system and an equipment system, collect system operation data of each operation subsystem in the software system, and extract abnormal operation data from the system operation data.
[0068] The present invention can obtain abnormal data in the system operation data by extracting abnormal operation data in the system operation data, thereby facilitating subsequent analysis of system failures based on the abnormal operation data, wherein the IT system refers to an integrated system composed of hardware, software, network and data, etc., which is mainly used for processing, storing, transmitting and managing information, the system operation data is various data records generated during the operation of each operation subsystem in the software system, such as user activity record data, and the abnormal operation data is data with abnormalities in the system operation data. Optionally, the collection of system operation data of each operation subsystem in the software system can be achieved by a data collector, and the data collector is compiled by a script language.
[0069] As an embodiment of the present invention, the extraction of abnormal operation data in the system operation data includes: scheduling a data log corresponding to the system operation data, calculating an operation error rate corresponding to the system operation data according to the data log, mapping the system operation data to obtain a data mapping value, calculating a data variance value corresponding to the system operation data according to the data mapping value, and identifying the abnormal operation data in the system operation data according to the data variance value and the operation error rate.
[0070] Among them, the data log is a record of the changes in the system operation data and the operation process, the operation error rate represents the error frequency corresponding to the system operation data, the data mapping value is the expression value corresponding to the system operation data, and the data variance value represents the degree of discreteness corresponding to the system operation data.
[0071] Optionally, the data log corresponding to the system operation data can be scheduled by a log recorder to identify the data operation records recorded in the data log, count the total number of operations and the number of operation errors in the data operation records, calculate the ratio of the number of operation errors to the total number of operations, and obtain the operation error rate corresponding to the system operation data according to the ratio. Mapping processing of the system operation data can be achieved through a mapping function, such as a mapping function, which compares the data variance value and the operation error rate with the corresponding preset variance threshold and preset error rate threshold respectively. When the data variance value is greater than the preset variance threshold and the operation error rate is greater than the preset error rate threshold, abnormal operation data in the system operation data is identified.
[0072] Further, as an optional embodiment of the present invention, calculating the data variance value corresponding to the system operation data according to the data mapping value includes:
[0073] The data variance value corresponding to the system operation data is calculated by the following formula:
[0074]
[0075] Among them, A represents the data variance value corresponding to the system operation data, a represents the sequence number corresponding to the system operation data, β represents the number of data of the system operation data, and B a represents the expected value of the ath data in the system operation data, d a Indicates the data mapping value of the ath data in the system operation data. Indicates the average mapping value of system operation data.
[0076] S2. Input the abnormal operation data as input data into the trained operation and maintenance AI big model, use the feature extraction layer in the operation and maintenance AI big model to extract features from the abnormal operation data to obtain abnormal features, use the timing convolution layer in the operation and maintenance AI big model to analyze the timing relationship between the abnormal features, and according to the timing relationship and the abnormal features, use the loop analysis layer in the operation and maintenance AI big model to analyze the system fault corresponding to the software system.
[0077] The present invention uses the feature extraction layer in the operation and maintenance AI big model to extract features from the abnormal operation data, so as to obtain data-specific attributes in the abnormal operation data, thereby facilitating the subsequent improvement of the accuracy of system fault analysis. The operation and maintenance AI big model is a huge model that uses artificial intelligence technology to perform operation and maintenance management and detect faults. The feature extraction layer is a neural network used for feature extraction in the operation and maintenance AI big model, and the abnormal feature is a data-specific attribute corresponding to the abnormal operation data.
[0078] As an embodiment of the present invention, the feature extraction layer in the operation and maintenance AI big model is used to extract features from the abnormal operation data to obtain abnormal features, including: using the input neural unit in the feature extraction layer to perform data optimization processing on the abnormal operation data to obtain optimized abnormal data, using the convolution neural unit in the feature extraction layer to perform feature extraction on the optimized abnormal data to obtain initial abnormal features, using the pooling function in the feature extraction layer to perform dimensionality reduction processing on the initial abnormal features to obtain reduced dimensionality abnormal features, using the entropy function in the feature extraction layer to calculate the feature gain value corresponding to the reduced dimensionality abnormal features, and according to the feature gain value, using the output neural unit in the feature extraction layer to output the reduced dimensionality abnormal features to obtain abnormal features.
[0079] Among them, the input neural unit is a neural network used to adjust the abnormal operation data, so as to improve the quality of the data, the optimized abnormal data is the data obtained after processing the invalid or repeated data in the abnormal operation data, the convolution neural unit is a neural network used to extract the data representation of the optimized abnormal data, the initial abnormal feature is the data representation corresponding to the optimized abnormal data, the pooling function is a function used to reduce the dimension of the initial abnormal feature, such as the maximum pooling function, the feature gain value represents the contribution corresponding to the reduced dimension abnormal feature, so as to judge the importance of the reduced dimension abnormal feature, and the output neural unit is a neural network used to output important features from the reduced dimension abnormal feature.
[0080] Optionally, the abnormal operation data can be optimized by an optimization function in an input neural unit, the optimization function is compiled by a scripting language, and features of the optimized abnormal data can be extracted by an activation function in the convolution neural unit, the activation function includes a Sigmoid function, and according to the feature gain value, the reduced dimensionality abnormal features can be output by an output function in the output neural unit, the output function includes a softmax function.
[0081] Optionally, as an optional embodiment of the present invention, the calculating the feature gain value corresponding to the dimension reduction abnormal feature by using the entropy function in the feature extraction layer includes:
[0082] The entropy function includes:
[0083]
[0084] Where D represents the feature gain value corresponding to the dimension reduction anomaly feature, E(c) represents the occurrence probability corresponding to the c-th feature in the dimension reduction anomaly feature, c represents the feature sequence number corresponding to the dimension reduction anomaly feature, |Fc| represents the feature subset containing the c-th feature in the dimension reduction anomaly feature, |ω| represents all feature subsets in the dimension reduction anomaly feature, and φc represents the number of feature subsets containing the c-th feature in the dimension reduction anomaly feature.
[0085] The present invention utilizes the temporal convolution layer in the operation and maintenance AI big model to analyze the temporal relationship between the abnormal features, so as to obtain the temporal correlation between the abnormal features, thereby facilitating the subsequent analysis of the system failure of the software system in combination with the abnormal features, wherein the temporal relationship is the temporal correlation between the abnormal features, that is, the change of one feature may be affected by the past values of other features.
[0086] As an embodiment of the present invention, the use of the timing convolution layer in the operation and maintenance AI big model to analyze the timing relationship between the abnormal features includes: using the recognition function in the timing convolution layer to identify the timing dimension corresponding to the abnormal feature, according to the timing dimension, setting the timing window corresponding to the abnormal feature, and defining the convolution step corresponding to the convolution kernel in the timing convolution layer, according to the convolution step and the timing window, using the convolution kernel to convolve the abnormal feature to obtain a convolution timing feature, using the linear function in the timing convolution layer to calculate the correlation coefficient between the convolution timing features, and determining the timing relationship between the abnormal features according to the convolution timing feature and the correlation coefficient.
[0087] Among them, the identification function is a function used to identify the dimension of time information corresponding to the abnormal feature, the identification function is compiled by JAVA language, the timing dimension is the dimension of time-related information corresponding to the abnormal feature, the timing window is the window size corresponding to the abnormal feature when a subsequent convolution operation is performed, the convolution step size is the distance moved by the subsequent convolution kernel convolution processing, the convolution timing feature is the time-related feature corresponding to the abnormal feature, the correlation coefficient represents the correlation between the convolution timing features, and the linear function includes a linear function.
[0088] Optionally, the unit time series length corresponding to the time series dimension can be calculated, and the periodicity corresponding to the abnormal feature can be analyzed. According to the unit time series length and the periodicity, a time series window corresponding to the abnormal feature is set, such as a sliding window or a rolling window. The convolution step corresponding to the convolution kernel can be defined according to the window size information of the time series window, and the time series corresponding to the convolution time series feature can be determined. The correlation coefficient is compared with a preset threshold. If the correlation coefficient is greater than the preset threshold and the corresponding sequence in the time series is at the front, the relationship between the time series of the abnormal features is determined. The time series of feature a is before the time series of feature b, and the correlation coefficient is greater than the preset threshold, which indicates that the time series relationship between feature a and feature b is a sequential relationship, that is, feature b appears after feature a appears.
[0089] The present invention analyzes the system failure corresponding to the software system according to the timing relationship and the abnormal characteristics by utilizing the loop analysis layer in the operation and maintenance AI big model, thereby improving the accuracy of the system failure analysis of the software system, wherein the system failure is a failure corresponding to the software system.
[0090] As an embodiment of the present invention, the system failure corresponding to the software system is analyzed by using the loop analysis layer in the operation and maintenance AI big model according to the timing relationship and the abnormal feature, including: according to the timing relationship, using the hierarchical relationship analysis network in the loop analysis layer to identify the cascade abnormal feature in the abnormal feature, querying the corresponding system function in the software system, analyzing the functional attributes corresponding to the system function, analyzing the system dependencies between the software systems according to the functional attributes, and according to the system dependencies, using the traceability neural network in the loop analysis layer to analyze the triggering abnormal feature in the cascade abnormal feature, using the fault diagnosis network in the loop analysis layer to analyze the characteristic fault corresponding to the triggering abnormal feature, and obtaining the system failure corresponding to the software system according to the characteristic fault.
[0091] Among them, the hierarchical relationship analysis network is a network used to analyze the subordinate relationship between variables, the cascade abnormality feature is a feature with a cascade effect in the abnormality feature, the system function is the corresponding function in the software system, the functional attribute is the functional property corresponding to the system function, the system dependency represents the degree of dependency between the software systems, the traceability neural network is a neural network used to trace the origin of an event, the triggering abnormality feature is the triggering source feature in the cascade feature, the fault diagnosis network is a neural network used to diagnose the fault corresponding to the event, and the characteristic fault is the fault condition corresponding to the triggering abnormality feature.
[0092] Optionally, the cascading abnormal features in the abnormal features can be identified through the hierarchical clustering function in the hierarchical relationship analysis network, the corresponding system functions in the software system can be obtained by querying the system library of the IT system, and the analysis of the functional attributes corresponding to the system functions can be achieved through an attribute analysis tool. The attribute analysis tool is compiled in a programming language, and the degree of data interaction between each system in the software system can be understood from the functional attributes. The system dependency between the software systems is analyzed according to the dependency of the degree of data interaction. The triggering abnormal features in the cascade abnormal features can be analyzed through the recurrent neural unit in the traceability neural network, and the triggering abnormal features can be decoded and encoded by the autoencoder in the fault diagnosis network to analyze the abnormal factors corresponding to the triggering abnormal features, and query the faults corresponding to the abnormal factors to obtain the characteristic faults.
[0093] S3. Collect the equipment operation data corresponding to each device in the equipment system in real time, and analyze the equipment failure corresponding to each device in the equipment system using the long short-term memory network in the operation and maintenance AI big model based on the equipment operation data.
[0094] The present invention analyzes the equipment failure corresponding to each device in the equipment system according to the equipment operation data and the equipment parameters by using the long short-term memory network in the operation and maintenance AI large model, so as to find the equipment problem corresponding to each device in the equipment system and make corresponding repair processing as soon as possible to reduce the impact caused by the failure, wherein the equipment operation data are various data records generated during the operation of each device in the equipment system, and the equipment parameters are parameters under the normal working state corresponding to each device in the equipment system. Optionally, real-time collection of the equipment operation data corresponding to each device in the equipment system can be achieved through a data collector, such as temperature, pressure, humidity and other sensors, and the equipment parameters corresponding to each device in the equipment system can be obtained by referring to the equipment manual.
[0095] As an embodiment of the present invention, according to the equipment operation data, the long short-term memory network in the operation and maintenance AI big model is used to analyze the equipment fault corresponding to each device in the equipment system, including: using the clustering function in the long short-term memory network to cluster the equipment operation data to obtain clustered operation data, using the attention mechanism in the long short-term memory network to identify data variables in the clustered operation data, calculating the variable weights corresponding to the data variables, extracting the key variables in the data variables according to the variable weights, using the linear regression function in the long short-term memory network to calculate the data values corresponding to the clustered operation data, combining the key variables and the data values to construct an operation trend graph corresponding to each device in the equipment system, according to the operation trend graph, using the recursive neural unit in the long short-term memory network to analyze the fault variables in the key variables, and analyzing the fault factors corresponding to the fault variables, performing fault merging processing on the fault factors, and obtaining the equipment fault corresponding to each device in the equipment system.
[0096] Among them, the clustering function is a function used to cluster data, such as the K-means clustering function, the clustered operation data is data obtained by clustering data with the same attributes in the equipment operation data, the data variable is a changing factor in the clustered operation data, the variable weight represents the importance corresponding to the data variable, the linear regression function is a function used to calculate the numerical value corresponding to the data, such as a polynomial regression function, the operation trend chart is a visualization chart of the operation status corresponding to each device in the equipment system, the recursive neural unit is a neural unit used to analyze the trend of the operation trend chart, the fault variable is a variable with abnormalities in the key variables, and the fault factor is the fault cause corresponding to the fault variable.
[0097] Optionally, the data variables in the clustered operation data can be identified by the variable identification function in the attention mechanism, and the variable identification function is compiled by a programming language. The extraction of key variables in the data variables can be achieved by the left function, and the construction of the operation trend chart corresponding to each device in the device system can be achieved by the Visio drawing tool. According to the operation trend chart, the recursive neural unit can calculate the rate of change of the trend direction corresponding to the operation trend chart. If the rate of change is greater than the preset rate of change, it indicates that the corresponding key variable has a fault, so as to analyze the fault variables in the key variables. The device components corresponding to the fault variables can be queried, and the fault factors can be determined according to the device components. Fault merging processing of the fault factors can be achieved by a merge function, such as the merge() function.
[0098] Optionally, as an optional embodiment of the present invention, the calculating the variable weight corresponding to the data variable includes:
[0099] The variable weight corresponding to the data variable is calculated by the following formula:
[0100]
[0101] Among them, H represents the variable weight corresponding to the data variable, t represents the number of data variables, j represents the variable sequence number corresponding to the data variable, and M j+1 Indicates the effect value between the j+1th variable and other variables in the data variable, M j Represents the effect value between the j-th variable and other variables in the data variable.
[0102] S4. Combining the system failure and the equipment failure, obtain a fault detection result corresponding to the IT system.
[0103] The present invention combines the system failure and the equipment failure to obtain a fault detection result corresponding to the IT system, performs fault detection on the software and hardware of the IT system, improves the accuracy of fault analysis, and avoids the problem of omissions in fault analysis. Optionally, the system failure and the equipment failure are summarized to obtain a summary failure, and the fault point corresponding to the summary failure is located. The fault detection result corresponding to the IT system is obtained by combining the fault point and the summary failure.
[0104] The present invention can obtain abnormal data in the system operation data by extracting abnormal operation data from the system operation data, so as to facilitate the subsequent analysis of system failures based on the abnormal operation data. The present invention can obtain data-specific attributes in the abnormal operation data by using the feature extraction layer in the operation and maintenance AI big model to extract features from the abnormal operation data, so as to facilitate the subsequent improvement of the accuracy of system failure analysis. The present invention analyzes the device failure corresponding to each device in the device system based on the device operation data and the device parameters using the long short-term memory network in the operation and maintenance AI big model, and can find the device problem corresponding to each device in the device system, and make corresponding repair processing as soon as possible to reduce the impact caused by the failure. The present invention combines the system failure and the device failure to obtain the fault detection result corresponding to the IT system, and performs fault detection on the software and hardware of the IT system, thereby improving the accuracy of fault analysis and avoiding the problem of omission of fault analysis. Therefore, the IT fault detection method based on the operation and maintenance AI big model provided in the embodiment of the present invention can improve the accuracy of IT fault detection.
[0105] like Figure 2As shown, it is a functional module diagram of an IT fault detection system based on an operation and maintenance AI big model provided by one embodiment of the present invention.
[0106] The IT fault detection system 100 based on the operation and maintenance AI big model described in the present invention can be installed in an electronic device. According to the functions implemented, the IT fault detection system 100 based on the operation and maintenance AI big model can include a data extraction module 101, a system fault analysis module 102, an equipment fault analysis module 103 and a detection result generation module 104. The module described in the present invention can also be called a unit, which refers to a series of computer program segments that can be executed by an electronic device processor and can complete fixed functions, which are stored in the memory of the electronic device.
[0107] In this embodiment, the functions of each module / unit are as follows:
[0108] The data extraction module 101 is used to obtain an IT system to be detected for faults, wherein the IT system includes: a software system and an equipment system, collect system operation data of each operation subsystem in the software system, and extract abnormal operation data from the system operation data;
[0109] The system fault analysis module 102 is used to input the abnormal operation data as input data into the trained operation and maintenance AI big model, use the feature extraction layer in the operation and maintenance AI big model to extract features from the abnormal operation data to obtain abnormal features, use the timing convolution layer in the operation and maintenance AI big model to analyze the timing relationship between the abnormal features, and use the loop analysis layer in the operation and maintenance AI big model to analyze the system fault corresponding to the software system based on the timing relationship and the abnormal features;
[0110] The equipment failure analysis module 103 is used to collect equipment operation data corresponding to each equipment in the equipment system in real time, and analyze the equipment failure corresponding to each equipment in the equipment system using the long short-term memory network in the operation and maintenance AI big model according to the equipment operation data;
[0111] The detection result generating module 104 is used to obtain a fault detection result corresponding to the IT system by combining the system fault and the device fault.
[0112] In detail, each module described in the IT fault detection system 100 based on the operation and maintenance AI big model described in the embodiment of the present application adopts the same Figure 1 The IT fault detection method based on the operation and maintenance AI big model described in the article has the same technical means and can produce the same technical effects, so I will not go into details here.
[0113] In the several embodiments provided by the present invention, it should be understood that the provided methods and systems can be implemented in other ways. For example, the method embodiments described above are only illustrative, for example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.
[0114] It is obvious to those skilled in the art that the present invention is not limited to the details of the above exemplary embodiments, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0115] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand artificial intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.
[0116] In addition, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or systems stated in a system claim can also be implemented by one unit or system through software or hardware. The words first, second, etc. are used to indicate names, and do not indicate any particular order.
[0117] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the spirit and scope of the technical solution of the present invention.
Claims
1. An IT fault detection method based on an operation and maintenance AI big model, characterized in that: The method comprises: Acquire an IT system to be fault-checked, the IT system comprising: a software system and an equipment system, collect system operation data of each operation subsystem in the software system, and extract abnormal operation data from the system operation data; Input the abnormal operation data as input data into the trained operation and maintenance AI big model, use the feature extraction layer in the operation and maintenance AI big model to extract features from the abnormal operation data to obtain abnormal features, use the timing convolution layer in the operation and maintenance AI big model to analyze the timing relationship between the abnormal features, and use the loop analysis layer in the operation and maintenance AI big model to analyze the system fault corresponding to the software system based on the timing relationship and the abnormal features; Collecting the equipment operation data corresponding to each device in the equipment system in real time, and analyzing the equipment failure corresponding to each device in the equipment system by using the long short-term memory network in the operation and maintenance AI big model according to the equipment operation data; The system failure and the equipment failure are combined to obtain a failure detection result corresponding to the IT system.
2. The IT fault detection method based on the operation and maintenance AI big model according to claim 1 is characterized in that: The extracting abnormal operation data from the system operation data includes: Dispatching data logs corresponding to the system operation data; Calculating the operation error rate corresponding to the system operation data according to the data log; Performing mapping processing on the system operation data to obtain a data mapping value; Calculating the data variance value corresponding to the system operation data according to the data mapping value; Abnormal operation data in the system operation data is identified according to the data variance value and the operation error rate.
3. The IT fault detection method based on the operation and maintenance AI big model as claimed in claim 2 is characterized in that: Calculating the data variance value corresponding to the system operation data according to the data mapping value includes: The data variance value corresponding to the system operation data is calculated by the following formula: Among them, A represents the data variance value corresponding to the system operation data, a represents the sequence number corresponding to the system operation data, β represents the number of system operation data, and B a represents the expected value of the ath data in the system operation data, d a Indicates the data mapping value of the ath data in the system operation data. Indicates the average mapping value of system operation data.
4. The IT fault detection method based on the operation and maintenance AI big model according to claim 1 is characterized in that: The feature extraction layer in the operation and maintenance AI big model is used to extract features from the abnormal operation data to obtain abnormal features, including: Using the input neural unit in the feature extraction layer to perform data optimization processing on the abnormal operation data to obtain optimized abnormal data; Using the convolutional neural unit in the feature extraction layer to extract features from the optimized abnormal data to obtain initial abnormal features; Performing dimensionality reduction processing on the initial abnormal features by using the pooling function in the feature extraction layer to obtain reduced-dimensional abnormal features; Calculating the feature gain value corresponding to the dimension reduction abnormal feature using the entropy function in the feature extraction layer; According to the feature gain value, the output neural unit in the feature extraction layer is used to perform output processing on the dimension reduction abnormal feature to obtain the abnormal feature.
5. The IT fault detection method based on the operation and maintenance AI big model as claimed in claim 4 is characterized in that: The calculating the feature gain value corresponding to the dimension reduction abnormal feature by using the entropy function in the feature extraction layer includes: The entropy function includes: Where D represents the feature gain value corresponding to the dimension reduction anomaly feature, E(c) represents the occurrence probability corresponding to the c-th feature in the dimension reduction anomaly feature, c represents the feature sequence number corresponding to the dimension reduction anomaly feature, |Fc| represents the feature subset containing the c-th feature in the dimension reduction anomaly feature, and |ω| represents all feature subsets in the dimension reduction anomaly feature. Indicates the number of feature subsets that contain the cth feature in the dimension reduction anomaly feature.
6. The IT fault detection method based on the operation and maintenance AI big model according to claim 1 is characterized in that: The analyzing the temporal relationship between the abnormal features by using the temporal convolution layer in the operation and maintenance AI big model includes: Using the recognition function in the temporal convolution layer to identify the temporal dimension corresponding to the abnormal feature; According to the time series dimension, a time series window corresponding to the abnormal feature is set, and a convolution step corresponding to the convolution kernel in the time series convolution layer is defined; According to the convolution step size and the time series window, the abnormal feature is convolved using the convolution kernel to obtain a convolution time series feature; Calculating the correlation coefficient between the convolutional time series features using the linear function in the time series convolution layer; The time series relationship between the abnormal features is determined according to the convolution time series features and the correlation coefficient.
7. The IT fault detection method based on the operation and maintenance AI big model according to claim 1 is characterized in that: The analyzing the system failure corresponding to the software system by using the loop analysis layer in the operation and maintenance AI big model according to the time sequence relationship and the abnormal characteristics includes: According to the time series relationship, using the hierarchical relationship analysis network in the cyclic analysis layer to identify the cascade abnormality features in the abnormality features; Query the corresponding system functions in the software system and analyze the functional attributes corresponding to the system functions; Analyzing system dependencies between the software systems based on the functional attributes; According to the system dependency, the traceability neural network in the loop analysis layer is used to analyze the triggering abnormality features in the cascade abnormality features; Analyzing the characteristic fault corresponding to the triggering abnormal feature by using the fault diagnosis network in the loop analysis layer; According to the characteristic fault, a system fault corresponding to the software system is obtained.
8. The IT fault detection method based on the operation and maintenance AI big model according to claim 1 is characterized in that: The analyzing the equipment failure corresponding to each device in the equipment system using the long short-term memory network in the operation and maintenance AI big model according to the equipment operation data includes: Using the clustering function in the long short-term memory network to cluster the device operation data, so as to obtain clustered operation data; Using the attention mechanism in the long short-term memory network to identify data variables in the clustering operation data; Calculating variable weights corresponding to the data variables, and extracting key variables from the data variables according to the variable weights; Calculating data values corresponding to the clustering operation data using a linear regression function in the long short-term memory network; Combining the key variables and the data values, constructing an operation trend graph corresponding to each device in the device system; According to the operation trend diagram, using the recursive neural unit in the long short-term memory network to analyze the fault variables in the key variables; The fault factors corresponding to the fault variables are analyzed, and fault merging processing is performed on the fault factors to obtain the device fault corresponding to each device in the device system.
9. The IT fault detection method based on the operation and maintenance AI big model as claimed in claim 8, characterized in that: The calculating the variable weight corresponding to the data variable includes: The variable weight corresponding to the data variable is calculated by the following formula: Among them, H represents the variable weight corresponding to the data variable, t represents the number of data variables, j represents the variable sequence number corresponding to the data variable, and M j+1 Indicates the effect value between the j+1th variable and other variables in the data variable, M j Represents the effect value between the j-th variable and other variables in the data variable.
10. An IT fault detection system based on an operation and maintenance AI big model, characterized in that: The system comprises: A data extraction module is used to obtain an IT system to be detected for faults, wherein the IT system includes: a software system and an equipment system, collect system operation data of each operation subsystem in the software system, and extract abnormal operation data from the system operation data; A system fault analysis module, used to input the abnormal operation data as input data into the trained operation and maintenance AI big model, use the feature extraction layer in the operation and maintenance AI big model to extract features from the abnormal operation data to obtain abnormal features, use the timing convolution layer in the operation and maintenance AI big model to analyze the timing relationship between the abnormal features, and use the loop analysis layer in the operation and maintenance AI big model to analyze the system fault corresponding to the software system based on the timing relationship and the abnormal features; An equipment failure analysis module is used to collect equipment operation data corresponding to each equipment in the equipment system in real time, and analyze the equipment failure corresponding to each equipment in the equipment system using the long short-term memory network in the operation and maintenance AI big model based on the equipment operation data; The detection result generating module is used to obtain a fault detection result corresponding to the IT system by combining the system fault and the equipment fault.