System fault prediction method, device, equipment and medium

By collecting and integrating multi-source data features of healthy systems and using neural network models to classify faults, the problem of low prediction accuracy in the existing technology is solved, and higher fault prediction accuracy and automated repair are achieved.

CN120602318APending Publication Date: 2025-09-05PING AN PAY ELECTRONIC PAYMENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510586993.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

In the prior art, the health system failure prediction method relies on preset rules and thresholds, making it difficult to cope with dynamic changes in the system and complex scenarios, resulting in low prediction accuracy.

Method used

Collect the operation source data, monitoring source data, log source data, network traffic source data and service grid source data of the healthy system, and extract and fusion features, including operation timing characteristics, monitoring timing characteristics, log statistical characteristics, network characteristics and dependency characteristics, and use neural network models to classify faults.

Benefits of technology

It improves the accuracy and reliability of system failure classification, can better deal with dynamic changes in the system and complex scenarios, reduce information uncertainty, and realize automated repair.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120602318A_ABST
    Figure CN120602318A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, in particular to a system fault prediction method and device, equipment and a medium. The method is applied to the medical field, and comprises the following steps: performing feature extraction on operation source data to obtain operation time sequence features, performing feature extraction on monitoring source data to obtain monitoring time sequence features, performing feature extraction on log source data to obtain log statistical features, performing feature extraction on network flow source data to obtain network features; the method comprises the steps of performing feature extraction on service grid source data to obtain dependency relationship features, performing feature fusion on operation time sequence features, monitoring time sequence features, log statistical features, network features and the dependency relationship features to obtain fusion features, and performing fault classification on a to-be-detected system according to the fusion features. And the accuracy and reliability of system fault classification are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a system fault prediction method, device, equipment and medium. Background Art

[0002] In the medical field, you can check your health status by logging into your account in the health system. Since the health system involves sensitive information such as basic customer information and health information, the security requirements are relatively strict. Therefore, the health system needs to be monitored and predicted in real time to prevent failures in the health system. With the rapid development of distributed systems and cloud computing technologies, the complexity and scale of health systems continue to expand. Therefore, in large-scale systems, the occurrence of failures has gradually become a normal phenomenon. Once some unknown failures occur during the operation of the health system, it will cause significant losses to the maintenance and operation of the entire health system. The fault prediction methods in the existing technology mainly rely on preset rules and thresholds, and the prediction accuracy is low. It is difficult to cope with dynamic changes in the system and complex scenarios. Therefore, in the process of predicting system failures, how to improve the prediction accuracy has become an urgent problem to be solved. Summary of the Invention

[0003] In view of this, embodiments of the present invention provide a system fault prediction method, apparatus, device, and medium to solve the problem of low prediction accuracy in the process of predicting system faults.

[0004] In a first aspect, an embodiment of the present invention provides a system fault prediction method, the system fault prediction method comprising: Collect the operation source data, monitoring source data, log source data, network traffic source data and service grid source data of the system to be tested; Performing feature extraction on the operation source data to obtain operation sequence features, performing feature extraction on the monitoring source data to obtain monitoring sequence features, performing feature extraction on the log source data to obtain log statistical features, performing feature extraction on the network traffic source data to obtain network features, and performing feature extraction on the service grid source data to obtain dependency features; Fusing the runtime features, the monitoring time series features, the log statistics features, the network features, and the dependency features to obtain fused features; According to the fusion features, fault classification is performed on the system to be detected to obtain a fault classification result.

[0005] In a second aspect, an embodiment of the present invention provides a system fault prediction device, the system fault prediction device comprising: The acquisition module is used to collect the operation source data, monitoring source data, log source data, network traffic source data and service grid source data of the system to be tested; An extraction module is configured to perform feature extraction on the operation source data to obtain operation sequence features, perform feature extraction on the monitoring source data to obtain monitoring sequence features, perform feature extraction on the log source data to obtain log statistical features, perform feature extraction on the network traffic source data to obtain network features, and perform feature extraction on the service grid source data to obtain dependency features; A fusion module, configured to fuse the runtime features, the monitoring time series features, the log statistics features, the network features, and the dependency features to obtain fused features; The classification module is used to classify the faults of the system to be detected according to the fusion features to obtain a fault classification result.

[0006] In a third aspect, an embodiment of the present invention provides a computer device, comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor implements the system fault prediction method as described in the first aspect when executing the computer program.

[0007] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the system fault prediction method described in the first aspect is implemented.

[0008] Compared with the prior art, the present invention has the following beneficial effects: In this application, feature extraction is performed on the operation source data to obtain the operation sequence feature, feature extraction is performed on the monitoring source data to obtain the monitoring sequence feature, feature extraction is performed on the log source data to obtain the log statistical feature, feature extraction is performed on the network traffic source data to obtain the network feature, feature extraction is performed on the service grid source data to obtain the dependency feature, and the operation sequence feature, monitoring sequence feature, log statistical feature, network feature and dependency feature are fused to obtain the fused feature. According to the fused feature, the fault classification of the system to be detected is performed, and various factors of the system are taken into consideration, thereby improving the accuracy and reliability of the system fault classification. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0010] Figure 1This is a schematic diagram of an application environment of a system fault prediction method provided by an embodiment of the present invention; Figure 2 This is a flow chart of a system fault prediction method provided by one embodiment of the present invention; Figure 3 This is a schematic structural diagram of a system fault prediction device provided by one embodiment of the present invention; Figure 4 It is a structural diagram of a computer device provided by one embodiment of the present invention. DETAILED DESCRIPTION

[0011] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0012] In the following description, specific details such as particular system structures and techniques are provided for purposes of illustration and not limitation to facilitate a thorough understanding of the embodiments of the present invention. However, it will be apparent to those skilled in the art that the present invention may be practiced in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, systems, circuits, and methods are omitted so as not to obscure the description of the present invention with unnecessary detail.

[0013] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.

[0014] It will also be understood that the term "and / or" used in the present description and appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0015] As used in the present specification and the appended claims, the term "if" may be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" may be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.

[0016] In addition, in the description of the present specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0017] References to "one embodiment" or "some embodiments" in the present specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present invention. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in yet other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0018] Embodiments of the present invention can acquire and process relevant data based on artificial intelligence technology. Artificial Intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results.

[0019] Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0020] It should be understood that the order of execution of the steps in the following embodiments does not necessarily mean the order in which they are executed. The order in which each process is executed should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0021] In order to illustrate the technical solution of the present invention, specific embodiments are provided below.

[0022] An embodiment of the present invention provides a system fault prediction method that can be applied in the following situations: Figure 1In an application environment, clients communicate with servers. Clients include, but are not limited to, PDAs, desktop computers, laptops, ultra-mobile personal computers (UMPCs), netbooks, cloud computing devices, personal digital assistants (PDAs), and other computer devices. Servers can be standalone servers or cloud servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0023] See also Figure 2 , is a flow chart of a system fault prediction method provided by an embodiment of the present invention, the system fault prediction method can be applied to Figure 1 The server in Figure 2 As shown, the system fault prediction method may include the following steps.

[0024] S201: Collect the operation source data, monitoring source data, log source data, network traffic source data and service grid source data of the system to be detected.

[0025] In step S201, operation source data refers to the basic original data set that supports the operation of the system. Its source, structure and management method directly affect the system operation efficiency and decision accuracy; monitoring source data is the core original data set that supports the operation of the monitoring system. Its composition, collection and processing methods directly affect the real-time and accuracy of monitoring; log source data is data that records the execution status of business logic (such as interface calls and transaction flows), which contains structured information such as event ID, error code, and operation context; the system's network traffic source data is the underlying record of network communication behavior, covering core information such as data packet transmission path, content characteristics and performance indicators. Its composition and analysis method directly affect network performance optimization, security monitoring and business decision-making efficiency; service grid source data is the core information set used to manage, monitor and optimize communication between microservices in the service grid architecture, covering multi-dimensional data such as traffic, configuration, security and observability. It realizes full life cycle management through the collaborative mechanism of the data plane and the control plane, supporting efficient governance and operation and maintenance of microservices.

[0026] In this embodiment, the system to be tested can be a system in the medical field, such as a health system. Each person can check their health status by logging into their account. Because the health system involves sensitive information such as basic customer information and health information, it has strict security requirements. Therefore, real-time monitoring and prediction of the health system are required to prevent health system failures. Health system failures may include batch failures, online failures, database failures, HTTP direct interface failures, page failures, service interface failures, data loss, disk array damage, program logic errors, HTTP connection failures, service interface registration failures, and other failures.

[0027] When collecting operational source data from the system to be tested, real-time data synchronization can be achieved through IoT sensors and software interfaces (such as database connection pools). When collecting monitoring source data from the system to be tested, data can be reported in real time through protocols such as MQTT and Modbus. Monitoring tools can capture the system's metrics through client SDKs or agents. When collecting log source data from the system to be tested, multi-source log collection (servers, databases, and middleware) can be achieved through a unified agent. When collecting network traffic source data from the system to be tested, raw data packets (RAW data) can be non-invasively copied through switch port mirroring or splitters, supporting full traffic analysis. NetFlow / sFlow protocols can also be used to sample and count data streams. When collecting service mesh source data from the system to be tested, proxies such as Envoy can be deployed to intercept and process inbound and outbound traffic.

[0028] It should be noted that when collecting operational source data from the system under test, Hubble data can be used. Hubble data collection supports client-side tracking, server-side tracking, no-tracking solutions, and third-party database access, forming a unified data collection platform. It covers dimensions such as user behavior, business metrics, link calls, logs, and infrastructure performance, and penetrates vertical monitoring layers such as hosts, containers, and networks to achieve full lifecycle data collection. When collecting monitoring source data from the system under test, Prometheus data can be collected. When collecting Prometheus data, target data can be actively captured through HTTP endpoints, or indirectly collected through the Pushgateway component. When collecting log source data from the system under test, ELK Stack log data can be collected. When collecting ELK Stack log data, container log collection and preprocessing can be implemented through DaemonSet. When collecting network traffic source data from the system under test, Wireshark data can be collected. Through the WinPCAP interface, Wireshark directly interacts with the network card, supporting real-time capture of network traffic for dozens of protocols, including HTTP, TCP, UDP, and DNS, covering both wired and wireless network environments. When collecting service mesh source data of the system to be tested, it can be obtained by collecting Istio data. Traffic can be intercepted by injecting Envoy proxy in each Pod, and the original data of communication between services can be captured in real time, including HTTP requests, TCP connections and gRPC streams. Envoy hijacks the traffic in and out of the container in transparent proxy mode, generates detailed network layer and application layer indicators, and automatically records four types of golden indicators (latency, traffic, errors, and saturation), such as request latency, error rate, and traffic distribution.

[0029] It should be noted that the operation source data at least includes memory usage data, memory occupancy value data, network delay data and other data; the monitoring source data at least includes container status data, error rate data, resource quota data and other data; the log source data at least includes program log data, system log data and security log data and other data; the network traffic source data at least includes network packet capture data, protocol analysis data and traffic statistics data and other data; the service grid source data at least includes dependency data between services, call delay data and traffic distribution data.

[0030] It should be noted that when collecting the operating source data, monitoring source data, log source data, network traffic source data and service grid source data of the system to be tested, security policies such as data encryption and permission classification must be followed, especially in scenarios involving user privacy. In this embodiment, multi-source data such as operation source data, monitoring source data, log source data, network traffic source data and service grid source data of the system to be detected are collected to facilitate fault prediction based on the multi-source data and improve the reliability of fault prediction.

[0031] S202: Perform feature extraction on the operation source data to obtain the operation sequence feature, perform feature extraction on the monitoring source data to obtain the monitoring sequence feature, perform feature extraction on the log source data to obtain the log statistical feature, perform feature extraction on the network traffic source data to obtain the network feature, perform feature extraction on the service grid source data to obtain the dependency feature.

[0032] In step S202, feature extraction is performed on the operation source data, monitoring source data, log source data, network traffic source data, and service grid source data respectively to obtain features of the corresponding source data.

[0033] In this embodiment, before feature extraction is performed on the operational, monitoring, log, network traffic, and service grid source data, the collected operational, monitoring, log, network traffic, and service grid source data are preprocessed. Specifically, each source data is cleansed and consolidated. During the cleaning process, missing data is filled in, and abnormal and duplicate data is deleted.

[0034] It should be noted that when filling missing data, linear interpolation methods, time series prediction models, or context-based filling can be used, and this embodiment does not limit this. When deleting abnormal data, statistical methods (such as Z-score, IQR) or machine learning methods (such as Isolation Forest) can be used to detect abnormal data and delete the abnormal data. Deduplication algorithms (such as hash deduplication or timestamp-based deduplication) can be used to remove duplicate data to ensure data uniqueness and accuracy.

[0035] After preprocessing the collected operational, monitoring, log, network traffic, and service mesh source data, timestamps from different sources are aligned to ensure that data from multiple sources at the same point in time can be mapped to a unified time series. For example, a sliding window technique is used to align timestamps of Hubble data, Prometheus data, Wireshark data, and other data to the second or millisecond level. The different source data is then formatted in a unified format, such as JSON or CSV, to facilitate subsequent processing and storage. The processed data from these different sources is stored in a time series database, such as InfluxDB or TimescaleDB, to support efficient time range queries and time series analysis. Each data point is annotated with information such as the timestamp and data source.

[0036] Feature extraction is performed on operation source data, monitoring source data, log source data, network traffic source data, and service grid source data. It should be noted that when performing time series feature extraction on operation source data and monitoring source data, the mean, standard deviation, maximum value, minimum value, slope, and periodicity of the corresponding source data are extracted. When performing statistical feature extraction on log source data, the frequency, distribution of virtual types, and number of keyword occurrences of the log source data are extracted. Network feature extraction is performed on network traffic source data, including packet size, transmission protocol, packet loss rate, latency, and bandwidth utilization. Dependency feature extraction is performed on service grid source data, including service call frequency, call latency, error rate, and weight distribution between services.

[0037] In this embodiment, feature extraction is performed on the operation source data, monitoring source data, log source data, network traffic source data and service grid source data respectively. According to the properties of different source data, features of different dimensions are extracted to make the extracted features more reliable and reasonable, so as to facilitate the use of the corresponding extracted features for system fault prediction.

[0038] S203: Fusing the runtime features, monitoring time series features, log statistics features, network features, and dependency features to obtain fused features.

[0039] In step S203, the runtime features, monitoring time series features, log statistics features, network features and dependency features are fused, that is, the features of different source data are combined to obtain fused features to enhance the representation capability of the features.

[0040] In this embodiment, when the runtime features, monitoring timing features, log statistical features, network features and dependency features are fused, splicing and fusion can be performed, that is, splicing is performed in the order of the above runtime features, monitoring timing features, log statistical features, network features and dependency features to obtain the corresponding feature vector, and the feature vector is determined as the fusion feature.

[0041] In this embodiment, runtime, monitoring, log statistics, network, and dependency features are fused to create fused features. By integrating the complementary features of multi-source data, the error and noise interference of a single data source can be effectively reduced, information uncertainty can be reduced, and data credibility can be improved. Data from different sources are complementary, and after fusion, they can cover a wider range of information dimensions and reduce data blind spots.

[0042] Optionally, the runtime features, monitoring time series features, log statistics features, network features, and dependency features are fused to obtain fused features, including: Obtain the correlation between any two source data among the operation source data, monitoring source data, log source data, network traffic source data, and service grid source data, and determine the correlation weight matrix based on the correlation relationship; According to the association weight matrix, the runtime characteristics, monitoring timing characteristics, log statistical characteristics, network characteristics and dependency characteristics are weighted to obtain weighted runtime characteristics, weighted monitoring timing characteristics, weighted log statistical characteristics, weighted network characteristics and weighted dependency characteristics; The weighted runtime features, weighted monitoring time series features, weighted log statistical features, weighted network features and weighted dependency features are concatenated and fused to obtain fused features.

[0043] In this embodiment, when the runtime features, monitoring timing features, log statistical features, network features and dependency features are fused, weighted fusion can be performed, that is, different weight values ​​are set for the runtime features, monitoring timing features, log statistical features, network features and dependency features.

[0044] In this embodiment, according to the association relationship between any two source data among the operation source data, monitoring source data, log source data, network traffic source data and service grid source data, different weight values ​​are set for the association relationship when the operation sequence characteristics, monitoring sequence characteristics, log statistical characteristics, network characteristics and dependency characteristics are set. Among them, the association relationship is a relationship affected by other source data, that is, the change of other source data has an impact on the source data. For example, the size of the memory occupancy value data in the operation source data has an impact on the resource quota in the monitoring source data. When the memory occupancy value data is large, the corresponding resource quota will generally be large. When the memory occupancy value data is small, the corresponding resource quota will generally be small. Therefore, there is an association relationship between the operation source data and the monitoring source data.

[0045] Obtain the association relationship between any two source data points, including operational source data, monitoring source data, log source data, network traffic source data, and service grid source data. For each source data point, determine the number of associations between the source data point and the other source data points based on the association relationship. Based on the number of associations between the source data point and the other source data points, calculate the association weight matrix. The association weight matrix contains the weights for different source data points.

[0046] For example, if there are three source data items associated with operational source data, three source data items associated with monitoring source data, two source data items associated with log source data, one source data item associated with network traffic source data, and one source data item associated with service grid source data, then the weight of each operational source data item is (3 / (3+3+2+1+1)). This means the weight of each operational source data item is 0.3, the weight of each monitoring source data item is 0.3, the weight of each log source data item is 0.2, the weight of each network traffic source data item is 0.1, and the weight of each service grid source data item is 0.1. This is the sum of the number of associations with each source data item divided by the total number of associations within each source data item. After calculating the weight of each source data item, an association weight matrix is ​​constructed. The association weight matrix is ​​a 1*N-dimensional matrix, where N is the number of source data items.

[0047] Based on the association weight matrix, the runtime features, monitoring time series features, log statistical features, network features, and dependency features are weighted. This means the weight values ​​of the corresponding source data in the association weight matrix are multiplied by the features of the corresponding source data to obtain the weighted runtime features, weighted monitoring time series features, weighted log statistical features, weighted network features, and weighted dependency features. The weighted runtime features, weighted monitoring time series features, weighted log statistical features, weighted network features, and weighted dependency features are then concatenated and fused to obtain the fused features.

[0048] In this embodiment, the weight value of the corresponding source data is determined according to the correlation relationship between any two source data, that is, the corresponding weight value is determined according to the importance of the source data in system fault prediction. According to the correlation weight matrix, the runtime characteristics, monitoring timing characteristics, log statistical characteristics, network characteristics and dependency characteristics are weighted, and the weighted runtime characteristics, weighted monitoring timing characteristics, weighted log statistical characteristics, weighted network characteristics and weighted dependency characteristics are spliced ​​and fused to obtain fused characteristics, thereby improving the accuracy of the fused characteristics.

[0049] Optionally, determining an association weight matrix based on the association relationship includes: For any source data, based on the association relationship, determine the number of associations between the source data and other source data, as well as the number of associations between each data in the source data and each data in other source data; The association weight matrix is ​​calculated based on the number of associations between source data and other source data, and the number of associations between each data in the source data and each data in other source data.

[0050] In this embodiment, for any source data, based on the association relationship, the number of source data with which the source data has an association relationship with other source data, as well as the number of data in the source data with which the data in the other source data has an association relationship are determined, the number of multi-source data with which the source data has an association relationship with other source data, as well as the number of data in the source data with which the data in the other source data have an association relationship are added, and the sum is used as the final total number of data with an association relationship with the corresponding source data, and based on the total number of each source data, the ratio of the total number of each source data to the sum of the total number of all source data is calculated, and the corresponding ratio is determined as the weight value of the corresponding source data.

[0051] It should be noted that the number of associations between each data in the source data and each data in the other source data is the number of associations between each data in each source data and each data in the other source data. For example, the size of the memory usage value data in the running source data has an impact on the resource quota in the monitoring source data. When the memory usage value data is large, the corresponding resource quota will generally be large. When the memory usage value data is small, the corresponding resource quota will generally be small. Therefore, there is an association between the memory usage value data and the resource quota data. If the running source data includes 3 data, calculate the number of associations between each data and each data in the other data respectively, and then add up the number of associations between each data and each data in the other data, that is, add up the number of associations between each data in the running source data and each data in the other source data, and obtain the number of associations between each data in the running source data and each data in the other source data.

[0052] In this embodiment, the association relationship of each data in each source data is taken into consideration to improve the accuracy of determining the size of the association relationship between any two source data, thereby improving the accuracy of the fusion feature.

[0053] Optionally, an association weight matrix is ​​calculated based on the number of associations between source data and other source data, and the number of associations between each data in the source data and each data in other source data, including: Calculate the first weight value of each source data according to the number of associations between the source data and other source data; Calculate a second weight value of the source data according to the number of associations between each data in the source data and each data in other source data; The first weight value and the second weight value are added to obtain the weight value of the corresponding source data, all source data are traversed to obtain the weight value of each source data, and the associated weight matrix is ​​obtained according to the weight value of each source data.

[0054] In this embodiment, the first weight value of each source data is calculated based on the number of source data that have an association relationship with other source data, that is, the ratio of the number of source data that have an association relationship with other source data to the total number is calculated, where the total number is the sum of the number of source data that have an association relationship with other source data.

[0055] The second weight of the source data is calculated based on the number of associations between each data item in the source data and each data item in other source data items. Specifically, based on the number of associations between each data item in the source data and each data item in other source data items, the sum of the number of associations between all data items in each source data item and each data item in other source data items is calculated to obtain the first number of associations for each source data item. The first number of associations for each source data item is summed to obtain the total number of associations for each data item. The ratio of the first number of associations for each source data item to the total number of associations is calculated, and the corresponding ratio is used as the second weight.

[0056] The first weight value and the second weight value are added to obtain the weight value of the corresponding source data, all source data are traversed to obtain the weight value of each source data, and the associated weight matrix is ​​obtained according to the weight value of each source data.

[0057] In this embodiment, the association weight values ​​between the source data and the association weight values ​​between the individual data are calculated separately, and then the association weight values ​​between the data and the association weight values ​​between the individual data are added to obtain the weight value of the corresponding source data, thereby reducing the mutual influence between the individual data and the source data and improving the accuracy of calculating the source data weight.

[0058] S204: Classify the faults of the system to be detected based on the fusion features to obtain a fault classification result.

[0059] In step S204, fault classification is performed on the system to be detected to obtain a fault classification result, wherein the fault classification result includes multiple fault results.

[0060] In this embodiment, fault classification is performed on the system to be detected based on the fusion features to obtain fault classification results, wherein the fault classification results include normal results, network delay fault results, service unavailable fault results, etc.

[0061] In this embodiment, the fusion feature is a fusion of multiple source data features. According to the fusion feature, the fault classification of the system to be detected is performed to improve the accuracy of system fault classification.

[0062] Optionally, based on the fusion features, fault classification is performed on the system to be detected to obtain a fault classification result, including: Get the trained classification model; Input the fused features into the trained classification model and output the fault classification results.

[0063] In this embodiment, a trained classification model is obtained, wherein the trained classification model is a classification model including multiple classifiers, wherein the number of classifiers is equal to the number of fault types in the fault classification result, the fusion features are input into the trained classification model, and the fault classification result is output.

[0064] It should be noted that a trained classification model is a neural network model that includes an input layer, an LSTM layer, a dropout layer, a fully connected layer, and an output layer. The input layer receives the fused input features. The LSTM layer contains multiple LSTM units, which are used to capture long-term dependencies in time series data. The number of neurons in each LSTM unit can be adjusted based on the data size and complexity (e.g., 64, 128, or 256). The dropout layer prevents model overfitting by randomly dropping a certain percentage of neurons (e.g., 0.2 or 0.5). The fully connected layer maps the output of the LSTM layer to the fault category space, with the output dimension being the number of fault categories (e.g., two categories: normal / faulty, or multiple fault categories). The output layer uses the softmax function to output a probability distribution for the fault category, or the sigmoid function for binary classification.

[0065] It should be noted that when obtaining a trained classification model, you can use GridSearch or Random Search to optimize model hyperparameters (such as the number of LSTM units, Dropout ratio, and learning rate). Evaluate the effects of different hyperparameter combinations using validation set performance (such as accuracy and F1 score) to select the optimal parameter combination.

[0066] In this embodiment, a trained classification model is used for classification, thereby improving the classification efficiency of system fault classification.

[0067] Optionally, after fault classification is performed on the system to be detected based on the fusion features and the fault classification result is obtained, the following steps are further included: According to the fault classification result, a matching self-healing strategy that matches the fault classification result is determined, and the matching self-healing strategy is executed.

[0068] In this embodiment, after determining the fault classification result of the system to be detected, the system to be detected is repaired. When repairing the system to be detected, a matching self-healing strategy matching the fault classification result is determined according to the fault classification result, and the matching self-healing strategy is executed.

[0069] This embodiment supports multiple self-healing strategies, including those for scenarios such as high CPU usage, memory leaks, service unavailability, excessive network latency, and abnormal service mesh call links. After determining the matching self-healing strategy that matches the fault classification results, the strategy is continuously optimized and improved by dynamically adjusting its parameters and incorporating reinforcement learning algorithms. When executing the matching self-healing strategy, automated tools can call the Kubernetes API to perform repair operations.

[0070] In this example, a matching self-healing strategy is determined based on the fault classification results, and the matching self-healing strategy is executed to achieve automated repair of the detected system, reduce manual intervention, and lower operation and maintenance costs.

[0071] In this application, feature extraction is performed on the operation source data to obtain the operation sequence feature, feature extraction is performed on the monitoring source data to obtain the monitoring sequence feature, feature extraction is performed on the log source data to obtain the log statistical feature, feature extraction is performed on the network traffic source data to obtain the network feature, feature extraction is performed on the service grid source data to obtain the dependency feature, and the operation sequence feature, monitoring sequence feature, log statistical feature, network feature and dependency feature are fused to obtain the fused feature. According to the fused feature, the fault classification of the system to be detected is performed, and various factors of the system are taken into consideration, thereby improving the accuracy and reliability of the system fault classification.

[0072] See also Figure 3 , Figure 3 This is a schematic diagram of the structure of a system fault prediction device provided by an embodiment of the present invention. The system fault prediction device corresponds to the system fault prediction method in the above embodiment. Figure 2 For the convenience of explanation, only the parts related to this embodiment are shown. Figure 3 The system fault prediction device 30 includes: a collection module 31, an extraction module 32, a fusion module 33, and a classification module 34.

[0073] The acquisition module 31 is used to collect the operation source data, monitoring source data, log source data, network traffic source data and service grid source data of the system to be detected.

[0074] The extraction module 32 is used to perform feature extraction on the operation source data to obtain the operation sequence features, perform feature extraction on the monitoring source data to obtain the monitoring sequence features, perform feature extraction on the log source data to obtain the log statistical features, perform feature extraction on the network traffic source data to obtain the network features, and perform feature extraction on the service grid source data to obtain the dependency features.

[0075] The fusion module 33 is used to fuse the runtime sequence features, monitoring sequence features, log statistics features, network features and dependency features to obtain fused features.

[0076] The classification module 34 is used to classify the faults of the system to be detected based on the fusion features and obtain a fault classification result.

[0077] Optionally, the fusion module 33 includes: The determination unit is used to obtain the association relationship between any two source data among the operation source data, monitoring source data, log source data, network traffic source data and service grid source data, and determine the association weight matrix based on the association relationship.

[0078] The weighting unit is used to perform weighted processing on the runtime characteristics, monitoring timing characteristics, log statistical characteristics, network characteristics and dependency characteristics according to the associated weight matrix to obtain weighted runtime characteristics, weighted monitoring timing characteristics, weighted log statistical characteristics, weighted network characteristics and weighted dependency characteristics.

[0079] The fusion unit is used to splice and fuse the weighted runtime features, the weighted monitoring time series features, the weighted log statistical features, the weighted network features and the weighted dependency features to obtain fused features.

[0080] Optionally, the determining unit includes: The determination subunit is used to determine, for any source data, the number of associations between the source data and other source data, and the number of associations between each data in the source data and each data in other source data based on the association relationship.

[0081] The calculation subunit is used to calculate the association weight matrix according to the number of associations between source data and other source data, and the number of associations between each data in the source data and each data in other source data.

[0082] Optionally, the determining unit includes: The first calculation submodule is used to calculate a first weight value of each source data according to the number of associations between the source data and other source data.

[0083] The second calculation submodule is used to calculate the second weight value of the source data according to the number of associations between each data in the source data and each data in other source data.

[0084] A submodule is obtained, which is used to add the first weight value and the second weight value to obtain the weight value of the corresponding source data, traverse all source data, obtain the weight value of each source data, and obtain the associated weight matrix according to the weight value of each source data.

[0085] Optionally, the classification module 34 includes: The acquisition unit is used to obtain the trained classification model.

[0086] The output unit is used to input the fusion features into the trained classification model and output the fault classification results.

[0087] Optionally, the system failure prediction device 30 further includes: The execution module is used to determine a matching self-healing strategy that matches the fault classification result according to the fault classification result, and execute the matching self-healing strategy.

[0088] It should be noted that the information interaction, execution process and other contents between the above-mentioned units are based on the same concept as the embodiment of the method of the present invention. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.

[0089] Figure 4 This is a schematic diagram of the structure of a computer device provided by an embodiment of the present invention. Figure 4 As shown, the computer device of this embodiment includes: at least one processor ( Figure 4 Only one is shown), a memory, and a computer program stored in the memory and executable on at least one processor, wherein when the processor executes the computer program, the steps of any of the above-mentioned system fault prediction method embodiments are implemented.

[0090] The computer device may include, but is not limited to, a processor and a memory. It will be understood by those skilled in the art that Figure 4 The above is merely an example of a computer device and does not constitute a limitation on the computer device. The computer device may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, it may also include a network interface, a display screen, and an input device.

[0091] The processor may be a CPU, other general-purpose processors, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0092] Memory includes readable storage media, internal memory, and the like. Internal memory can be the internal memory of a computer device, providing an environment for the operation of the operating system and computer-readable instructions stored in the readable storage medium. The readable storage medium can be the computer device's hard drive. In other embodiments, it can also be an external storage device, such as a plug-in hard drive, a Smart Media Card (SMC), a Secure Digital (SD) card, or a flash memory card. Furthermore, memory can include both the computer device's internal storage unit and external storage devices. Memory is used to store the operating system, application programs, boot loaders, data, and other programs, such as the program code of computer programs. Memory can also be used to temporarily store data that has been output or is about to be output.

[0093] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other and are not used to limit the scope of protection of the present invention. The specific working process of the units and modules in the above-mentioned device can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention can implement all or part of the process steps in the above-mentioned method embodiments by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When executed by a processor, the computer program can implement the steps of the above-mentioned method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file, or some intermediate form. Computer-readable media can include at least: any entity or device capable of carrying computer program code, recording media, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signals, telecommunications signals, and software distribution media. Examples include USB flash drives, removable hard drives, magnetic disks, or optical disks. In some jurisdictions, based on legislation and patent practice, computer-readable media cannot be electric carrier signals or telecommunications signals.

[0094] The present invention may implement all or part of the processes in the above-mentioned method embodiments, and may also be completed through a computer program product. When the computer program product runs on a computer device, the computer device can implement the steps in the above-mentioned method embodiments when executing the computer program product.

[0095] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0096] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.

[0097] In the embodiments provided by the present invention, it should be understood that the disclosed apparatus / computer equipment and methods can be implemented in other ways. For example, the apparatus / computer equipment embodiments described above are merely illustrative. For example, the division of modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0098] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0099] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.

Claims

1. A system failure prediction method, characterized in that: The system failure prediction method includes: Collect the operation source data, monitoring source data, log source data, network traffic source data and service grid source data of the system to be tested; Performing feature extraction on the operation source data to obtain operation sequence features, performing feature extraction on the monitoring source data to obtain monitoring sequence features, performing feature extraction on the log source data to obtain log statistical features, performing feature extraction on the network traffic source data to obtain network features, and performing feature extraction on the service grid source data to obtain dependency features; Fusing the runtime features, the monitoring time series features, the log statistics features, the network features, and the dependency features to obtain fused features; According to the fusion features, fault classification is performed on the system to be detected to obtain a fault classification result.

2. The system failure prediction method according to claim 1, wherein: The operation source data includes at least memory usage data, memory occupancy value data, and network delay data; The monitoring source data includes at least container status data error rate data and resource quota data; The log source data includes at least program log data, system log data and security log data; The network traffic source data at least includes network packet capture data, protocol analysis data and traffic statistics data; The service grid source data includes at least dependency data between services, call delay data and traffic distribution data.

3. The system failure prediction method according to claim 1, wherein: The step of fusing the runtime features, the monitoring time series features, the log statistics features, the network features, and the dependency features to obtain fused features includes: Obtaining an association relationship between any two of the operation source data, the monitoring source data, the log source data, the network traffic source data, and the service grid source data, and determining an association weight matrix based on the association relationship; According to the association weight matrix, the runtime characteristics, the monitoring timing characteristics, the log statistical characteristics, the network characteristics and the dependency characteristics are weighted to obtain weighted runtime characteristics, weighted monitoring timing characteristics, weighted log statistical characteristics, weighted network characteristics and weighted dependency characteristics; The weighted runtime features, the weighted monitoring time series features, the weighted log statistical features, the weighted network features and the weighted dependency features are concatenated and fused to obtain fused features.

4. The system failure prediction method according to claim 3, wherein: Determining an association weight matrix according to the association relationship includes: For any source data, based on the association relationship, determine the number of associations between the source data and other source data, and the number of associations between each data in the source data and each data in other source data; An association weight matrix is ​​calculated based on the number of associations between source data and other source data, and the number of associations between each data in the source data and each data in other source data.

5. The system failure prediction method according to claim 4, characterized in that: The association weight matrix is ​​calculated based on the number of associations between source data and other source data, and the number of associations between each data in the source data and each data in other source data, including: Calculate the first weight value of each source data according to the number of associations between the source data and other source data; Calculating a second weight value of the source data according to the number of associations between each data in the source data and each data in other source data; The first weight value and the second weight value are added to obtain the weight value of the corresponding source data, all source data are traversed to obtain the weight value of each source data, and the associated weight matrix is ​​obtained according to the weight value of each source data.

6. The system failure prediction method according to claim 1, wherein: The step of performing fault classification on the system to be detected based on the fusion features to obtain a fault classification result includes: Get the trained classification model; The fusion features are input into the trained classification model, and the fault classification result is output.

7. The system failure prediction method according to claim 1, wherein: After performing fault classification on the system to be detected based on the fusion features and obtaining a fault classification result, the method further includes: According to the fault classification result, a matching self-healing strategy that matches the fault classification result is determined, and the matching self-healing strategy is executed.

8. A system failure prediction device, characterized in that: The system failure prediction device includes: The acquisition module is used to collect the operation source data, monitoring source data, log source data, network traffic source data and service grid source data of the system to be tested; An extraction module is configured to perform feature extraction on the operation source data to obtain operation sequence features, perform feature extraction on the monitoring source data to obtain monitoring sequence features, perform feature extraction on the log source data to obtain log statistical features, perform feature extraction on the network traffic source data to obtain network features, and perform feature extraction on the service grid source data to obtain dependency features; A fusion module, configured to fuse the runtime features, the monitoring time series features, the log statistics features, the network features, and the dependency features to obtain fused features; The classification module is used to classify the faults of the system to be detected according to the fusion features to obtain a fault classification result.

9. A computer device, characterized in that: The computer device includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the system fault prediction method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the system failure prediction method according to any one of claims 1 to 7 is implemented.