A method and system for detecting computer system faults based on artificial intelligence
By building a fault identification model based on convolutional neural network, the problems of data noise and subjectivity checking in computer system fault detection are solved, and efficient automatic fault detection and accurate fault prediction are achieved.
Patent Information
- Application Number
- CN202411280362.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-13
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-09-13
AI Technical Summary
The prior art faces data noise, missing and incompleteness problems in computer system fault detection, which affects model training and result accuracy, and relies on manual inspection to be subjective and inefficient.
By obtaining the initial computer system data, eliminating invalid data, and classifying hardware status data and system operating status data according to the fault prediction indicators, building a two-dimensional difference matrix and multi-modal samples, using a convolutional neural network to build a fault identification model, and performing unsupervised pre-training and fully supervised fine-tuning.
It realizes automatic fault detection of computer system data, improves the scalability and classification accuracy of data collection, reduces the overhead of data transmission and storage, and obtains high accuracy and fault prediction advance time.
Smart Images

Figure CN119226018B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer fault diagnosis, and particularly relates to a computer system fault detection method and system based on artificial intelligence. Background Art
[0002] With the continuous deepening of digital transformation, enterprises' dependence on information technology systems is increasing day by day. The traditional IT operation and maintenance mode, which relies on manual monitoring and manual problem handling, can no longer meet the modern complex and changeable business needs. As cyber threat actors develop more and more sophisticated strategies, cutting-edge cybersecurity has become a necessity for industry organizations and government agencies. A large number of new threat strategies have overwhelmed many of the most advanced cybersecurity models. With the progress of new computing, storage, and energy efficiency having evolved to spawn new technological breakthroughs, such as the Internet of Things (IoT), the development of security faces the challenge of matching the progress of new threats and new technologies. IoT technology enables data collection, processing, and communication in autonomous vehicles, smart cities, smart grids, and e-health applications. Given the numerous functions and low cost of IoT devices, they are usually distributed on a large scale rather than operating in a controlled and security-enhanced environment. These uncontrolled variables expose opportunities for physical, network, and application attacks, usually using updated and more easily exploitable protocols with limited protection.
[0003] At the present stage, intelligent fault diagnosis relies on AI systems, and AI systems require a large amount of high-quality data to provide accurate diagnosis and repair suggestions; however, in reality, the data often has noise, missing, or incomplete situations, which will affect the accuracy of model training and results; in the AI fault detection environment, ensuring that the data used for training the AI is clear and unambiguous often relies on manual inspection, and manual inspection is often subjective, and different inspectors may have different judgments on the same inspection results; among them, the model-based diagnosis applying AI technology depends on the accuracy of the system model as well as the quality and integrity of the data. With the large-scale and complexity of industrial systems, the fault detection and prediction methods of computer systems rely on empirical models and rules, facing problems such as poor adaptability and low efficiency;
[0004] At the present stage, big data is needed to explore various security scenarios, and there is a lack of available representative datasets for detection. Most work shows great effectiveness in detecting network traffic, system logs, and other types of data anomalies, but the centralized learning method needs to collect data from multiple edge computing systems, which violates the data privacy principle. Summary of the Invention
[0005] To solve the above problems existing in the prior art, the present invention provides a computer system fault detection method based on artificial intelligence,
[0006] The object of the present invention can be achieved by the following technical solutions:
[0007] A method for detecting computer system faults based on artificial intelligence, comprising:
[0008] S1: Obtain the initial computer system data, eliminate the invalid data, incomplete data and duplicate data in the initial computer system data, set the corresponding attribute classification of the initial computer system data according to the fault prediction index, and extract the hardware status data and system operation status data in the initial computer system data according to the attribute classification;
[0009] S2: Construct a two-dimensional difference matrix according to the data categories of the hardware status data and the system operation status data, and construct a multi-modal sample according to the two-dimensional difference matrix;
[0010] S3: Construct a fault recognition model, and train the fault recognition model according to the multi-modal sample, and the training process includes unsupervised pre-training and full-supervised fine-tuning;
[0011] S4: Automatically perform modal matching on the observed values of the computer system data collected in real time according to the fault recognition model to obtain the system fault recognition status.
[0012] Specifically, the hardware status data includes power supply current parameters, fan speed parameters, and hardware temperature parameters, and the system operation status data includes CPU occupancy rate, interrupt statistics parameter values, queue and load parameter values, memory statistics parameter values, page statistics parameter values, network statistics and register values, and I / O statistics and register values.
[0013] Specifically, the two-dimensional difference matrix is the difference between the hardware status data and the heterogeneous data attributes of the system operation status data, and the expression formula is:
[0014] ,
[0015] Wherein, G is the two-dimensional difference matrix, n is the total number of data attributes of the system operation status data, m is the total number of data attributes of the hardware status data, a 1 , a 2 , a m are the first type, the second type and the m-th type of attribute values in the hardware status data, a 1,2 , a 2,1 are the difference and the opposite number of the attribute values of the first type and the second type in the hardware status data,a 1,m and a m,1 is the difference value and the opposite number of the attribute values of the first type and the mth type, a 2,m and a m,2 is the difference value and the opposite number of the attribute values of the second type and the mth type; b 1 and b 2 and b n are the attribute values of the first type, the second type and the nth type in the system operation status data, b 1,2 and b 2,1 is the difference value and the opposite number of the attribute values of the first type and the second type in the system operation status data, b 1,n and b n,1 is the difference value and the opposite number of the attribute values of the first type and the nth type in the system operation status data, b 2,n and b n,2 is the difference value and the opposite number of the attribute values of the second type and the nth type in the system operation status data.
[0016] Specifically, the fault recognition model is based on a convolutional neural network. The two-dimensional difference matrix is used as the network input data. A weight sharing layer is added after the convolutional layer. The convolutional layer extracts the feature map of the input data through a local receptive field. The feature map is used as the input of the weight sharing layer. According to the similarity between the convolutional kernel and the matched filter structure, the convolutional filters with generalization ability for different parameters are constructed by using the convolutional kernel weights of different convolutional layers; the weight sharing layer includes a square sampling layer and a square root sampling layer.
[0017] Specifically, the unsupervised pre-training is to learn the weights of the square sampling layer in the weight sharing layer. The weights of the square root sampling layer are fixedly set and hard-coded to represent the topological structure of the square sampling layer. After the input data is given, the weights of the square sampling layer are learned by characterizing the coefficient features of the input data in the square root sampling layer. The calculation formula is:
[0018] ,
[0019] where W is the weight matrix of the square sampling layer, W T is the transposed matrix of the weight matrix of the square sampling layer, and I is the identity matrix, Tis the sampling period, m is the number of hidden units, n is the input data dimension size, V ik is the weight of the square root sampling layer, W kj is the weight of the square sampling layer, is the local feature of the input data.
[0020] Specifically, the fully supervised fine-tuning calculates the gradient through activation function softmax regression, and back-propagates the gradient error from the pooled unit output end of the convolutional neural network to gradually update the weight of the weight sharing unit of the convolutional neural network.
[0021] A computer system fault detection system based on artificial intelligence, comprising: a data acquisition module, a data encoding module, a model training module, and a model application module;
[0022] The data acquisition module is used to obtain initial computer system data, remove invalid data, incomplete data and duplicate data in the initial computer system data, set attribute classification corresponding to the initial computer system data according to the fault prediction index, and extract hardware status data and system operation status data in the initial computer system data according to the attribute classification;
[0023] The data encoding module is used to construct a two-dimensional difference matrix according to the data categories of the hardware status data and the system operation status data, and to construct a multimodal sample according to the two-dimensional difference matrix;
[0024] The model training module is used to construct a fault recognition model, and train the fault recognition model according to the multimodal samples, wherein the training process includes unsupervised pre-training and fully supervised fine-tuning;
[0025] The model application module is used to perform automatic modal matching on computer system data observation values collected in real time according to the fault identification model to obtain a system fault identification state.
[0026] The beneficial effects of the present invention are:
[0027] Through data-driven fault prediction, the initial computer system data is collected, clustered, and then the pattern is extracted, which can reduce redundancy without losing important information; the distributed data collection method can obtain more comprehensive operating status data of computing nodes including hardware and software, effectively utilize communication resources, reduce I / O overhead, further reduce data transmission and storage overhead, and improve the scalability of data collection methods.
[0028] A method is proposed to construct and update a fault identification model. Through unsupervised training, the sample classification prediction can better adapt to data fluctuations and concept drifts in prediction. During supervised fine-tuning, real-time collected data samples of the computing node hardware and system operating status are classified, achieving a high classification accuracy. Based on the classification prediction of real-time collected data samples, good accuracy and fault prediction lead time are obtained. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] For the convenience of those skilled in the art to understand, the present invention will be further described below in conjunction with the accompanying drawings.
[0030] Figure 1 It is a schematic flowchart of a computer system fault detection method based on artificial intelligence according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0031] To further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following will describe in detail the specific embodiments, structures, features, and their effects of the present invention in conjunction with the accompanying drawings and preferred embodiments.
[0032] Please refer to Figure 1 , a computer system fault detection method based on artificial intelligence, including:
[0033] S1: Obtain initial computer system data, remove invalid data, incomplete data, and duplicate data in the initial computer system data, set the corresponding attribute classification for the initial computer system data according to the fault prediction index, and extract the hardware status data and system operating status data from the initial computer system data according to the attribute classification;
[0034] S2: Construct a two-dimensional difference matrix according to the data categories of the hardware status data and system operating status data, and construct multimodal samples according to the two-dimensional difference matrix;
[0035] S3: Construct a fault identification model, and train the fault identification model according to the multimodal samples. The training process includes unsupervised pre-training and full supervised fine-tuning;
[0036] S4: Automatically perform modal matching on the real-time collected computer system data observation values according to the fault identification model to obtain the system fault identification status.
[0037] Specifically, the hardware status data includes power supply current parameters, fan speed parameters, and hardware temperature parameters, and the system operating status data includes CPU occupancy rate, interrupt statistic parameter values, queue and load parameter values, memory statistic parameter values, page statistic parameter values, network statistic and register values, I / O statistic and register values.
[0038] In this embodiment, the method for obtaining and aggregating hardware environment status data is that the block is monitored through the SMC (System Management Controller) board. Based on the maintenance control network and the SMC system, a client-server structure is adopted to obtain the hardware environment status data records of 16 computing nodes in each chassis maintained by the SMC board server of the chassis at one time. Using TCP / IP sockets, the multi-threaded method is used to collect SMC data of multiple chassis in parallel, reducing the resource occupation of the management node during multi-node data collection and avoiding the bottleneck problem of the management node. The system operation status data is obtained by using the / proc file system method; by analyzing the relevant files in the files or folders such as cpuinfo, meminfo, slabinfo, uptime, net / , sys / , scsi / in / proc, the status information of system operations including CPU, memory, network, and I / O can be obtained.
[0039] Specifically, the two-dimensional difference matrix is the difference between the hardware status data and the data attributes of its own heterogeneous data of the system operation status data, and the expression formula is:
[0040] ,
[0041] where, G is the two-dimensional difference matrix, n is the total number of data attributes of the system operation status data, m is the total number of data attributes of the hardware status data, a 1 , a 2 , a m are the first, second, and m-th category attribute values in the hardware status data, a 1,2 , a 2,1 are the difference and opposite number of the attribute values of the first and second categories in the hardware status data, a 1,m , a m,1 are the difference and opposite number of the attribute values of the first and m-th categories, a 2,m , a m,2 are the difference and opposite number of the attribute values of the second and m-th categories; b 1 , b 2 , b nare the first-class, second-class, and nth-class attribute values in the system operation status data, b 1,2 and b 2,1 are the difference and the opposite of the difference between the first-class and second-class attribute values in the system operation status data, b 1,n and b n,1 are the difference and the opposite of the difference between the first-class and nth-class attribute values in the system operation status data, b 2,n and b n,2 are the difference and the opposite of the difference between the second-class and nth-class attribute values in the system operation status data.
[0042] Specifically, the fault recognition model is based on a convolutional neural network. The two-dimensional difference matrix is used as the network input data. A weight sharing layer is added after the convolutional layer. The convolutional layer extracts the feature map of the input data through a local receptive field and uses the feature map as the input of the weight sharing layer. According to the similarity between the convolutional kernel and the matched filter structure, the convolutional filters with generalization ability for different parameters are constructed using the convolutional kernel weights of different convolutional layers. The weight sharing layer includes a square sampling layer and a square root sampling layer.
[0043] Specifically, in the unsupervised pre-training, the weights of the square sampling layer in the weight sharing layer are learned. The weights of the square root sampling layer are fixedly set and hard-coded to represent the topological structure of the square sampling layer. After the input data is given, the weights of the square sampling layer are learned by characterizing the coefficient features of the input data in the square root sampling layer. The calculation formula is:
[0044] ,
[0045] where W is the weight matrix of the square sampling layer, W T is the transpose matrix of the weight matrix of the square sampling layer, I is the identity matrix, T is the sampling period, m is the number of hidden units, n is the size of the input data dimension, V ik is the weight of the square root sampling layer, W kj is the weight of the square sampling layer, is the local feature of the input data.
[0046] Specifically, the full-supervised fine-tuning calculates the gradient through the activation function softmax regression, and backpropagates the error of the gradient from the output end of the pooling unit of the convolutional neural network to gradually update the weights of the weight-sharing unit of the convolutional neural network.
[0047] A computer system fault detection system based on artificial intelligence, comprising: a data acquisition module, a data encoding module, a model training module, and a model application module;
[0048] The data acquisition module is used to obtain the initial computer system data, eliminate the invalid data, incomplete data, and duplicate data in the initial computer system data, set the corresponding attribute classification for the initial computer system data according to the fault prediction index, and extract the hardware status data and system operation status data in the initial computer system data according to the attribute classification;
[0049] The data encoding module is used to construct a two-dimensional difference matrix according to the data categories of the hardware status data and the system operation status data, and construct a multi-modal sample according to the two-dimensional difference matrix;
[0050] The model training module is used to construct a fault recognition model and train the fault recognition model according to the multi-modal sample. The training process includes unsupervised pre-training and full-supervised fine-tuning;
[0051] The model application module is used to automatically perform modal matching on the observed values of the computer system data collected in real time according to the fault recognition model to obtain the system fault recognition status.
[0052] The computer storage medium of the embodiment of the present invention can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0053] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such propagated data signals may take many forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.
[0054] The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wire, optical fiber cable, RF, etc., or any suitable combination of the foregoing. The computer program code for performing the operations of the present invention may be written in one or more programming languages or combinations thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partially on the user's computer, execute as a stand-alone software package, execute partially on the user's computer and partially on a remote computer, or execute entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0055] As described above, the above are only preferred embodiments of the present invention and do not impose any formal limitations on the present invention. Although the present invention has been disclosed above with preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or refinements to equivalent embodiments by using the disclosed technical content within the scope of the technical solution of the present invention. However, as long as it does not depart from the content of the technical solution of the present invention, any simple modification, equivalent change, and refinement made to the above embodiments based on the technical essence of the present invention still fall within the scope of the technical solution of the present invention.
Claims
1. A computer system fault detection method based on artificial intelligence, characterized in that: include: S1: Acquire initial computer system data, remove invalid data, incomplete data and duplicate data in the initial computer system data, set attribute classification corresponding to the initial computer system data according to the fault prediction index, and extract hardware status data and system operation status data in the initial computer system data according to the attribute classification; S2: constructing a two-dimensional difference matrix according to the data categories of the hardware status data and the system operation status data, and constructing a multimodal sample according to the two-dimensional difference matrix; S3: constructing a fault recognition model, and training the fault recognition model according to the multimodal samples, wherein the training process includes unsupervised pre-training and fully supervised fine-tuning; The fault recognition model is based on a convolutional neural network, and uses the two-dimensional difference matrix as network input data. A weight sharing layer is added after the convolution layer. The convolution layer extracts a feature map of the input data through a local receptive field, and uses the feature map as the input of the weight sharing layer. According to the similarity between the convolution kernel and the matched filter structure, the convolution kernel weights of different convolution layers are used to construct a convolution filter with generalization ability for different parameters; the weight sharing layer includes a square sampling layer and a square root sampling layer; The unsupervised pre-training is performed by learning the weight of the square sampling layer in the weight sharing layer, the weight of the square root sampling layer is fixedly set and hard-coded to characterize the topological structure of the square sampling layer, and after the input data is given, the weight of the square sampling layer is learned by characterizing the coefficient features of the input data in the square root sampling layer, and the calculation formula is: Among them, W is the weight matrix of the square sampling layer, W T is the transposed matrix of the weight matrix of the square sampling layer, I is the identity matrix, T is the sampling period, m is the number of hidden units, n is the input data dimension size, V ik is the weight of the square root sampling layer, W kj is the weight of the square sampling layer, is the local feature of the input data; S4: Automatically perform modal matching on the real-time collected computer system data observation values according to the fault identification model to obtain a system fault identification state.
2. The method according to claim 1, characterized in that The hardware status data includes power supply current parameters, fan speed parameters, and hardware temperature parameters; the system operation status data includes CPU occupancy, interrupt statistical parameter values, queue and load parameter values, memory statistical parameter values, page statistical parameter values, network statistics and register values, and I / O statistics and register values.
3. The method according to claim 2, characterized in that The two-dimensional difference matrix is the difference between the hardware status data and system operation status data and the heterogeneous data attributes thereof, and the expression formula is: Among them, G is a two-dimensional difference matrix, n is the total number of data attributes of system operation status data, m is the total number of data attributes of hardware status data, a1, a2, a m are the first, second, and mth attribute values in the hardware status data, a 1,2 、a 2,1 is the difference and opposite number of the attribute values of the first and second categories in the hardware status data, a 1,m 、a m,1 is the difference and opposite number of the attribute values of the first and mth categories, a 2,m 、a m,2 is the difference and opposite number of the attribute values of the second category and the mth category; b1, b2, b n are the first, second, and nth attribute values in the system operation status data, b 1,2 , b 2,1 is the difference and opposite number of the attribute values of the first and second categories in the system operation status data, b 1,n , b n,1 is the difference and opposite number of the attribute values of the first and nth categories in the system operation status data, b 2,n , b n,2 It is the difference and opposite number of the attribute values of the second category and the nth category in the system operation status data.
4. The method according to claim 1, characterized in that: The fully supervised fine-tuning calculates the gradient through activation function softmax regression, and back-propagates the gradient error from the pooled unit output end of the convolutional neural network to gradually update the weight of the weight sharing unit of the convolutional neural network.
5. A fault detection system using the artificial intelligence-based computer system fault detection method according to claim 1, characterized in that: include: Data acquisition module, data encoding module, model training module, model application module; The data acquisition module is used to obtain initial computer system data, remove invalid data, incomplete data and duplicate data in the initial computer system data, set attribute classification corresponding to the initial computer system data according to the fault prediction index, and extract hardware status data and system operation status data in the initial computer system data according to the attribute classification; The data encoding module is used to construct a two-dimensional difference matrix according to the data categories of the hardware status data and the system operation status data, and to construct a multimodal sample according to the two-dimensional difference matrix; The model training module is used to construct a fault recognition model, and train the fault recognition model according to the multimodal samples, wherein the training process includes unsupervised pre-training and fully supervised fine-tuning; The model application module is used to perform automatic modal matching on computer system data observation values collected in real time according to the fault identification model to obtain a system fault identification state.
6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the artificial intelligence-based computer system fault detection method as described in any one of claims 1 to 4 is implemented.
7. A storage medium containing computer executable instructions, characterized in that: The computer executable instructions are used to execute the artificial intelligence-based computer system fault detection method as described in any one of claims 1 to 4 when executed by a computer processor.
Citation Information
Patent Citations
Computer equipment, fault detection system and method and readable storage medium
CN117194163A
Multi-modal learning method based on self-supervised knowledge distillation strategy
CN117973568A