Fault diagnosis method, device, medium and product

By acquiring and encoding multi-dimensional data and matching it with a historical fault case library, the problems of low fault diagnosis efficiency and poor accuracy in existing technologies are solved, and fast and accurate fault location and repair are achieved.

CN120492503BActive Publication Date: 2025-09-19INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510969885.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-15
Publication Date
2025-09-19
Estimated Expiration
2045-07-15

AI Technical Summary

Technical Problem

The fault diagnosis method in the existing technology relies on the analysis of a single data type, resulting in low diagnostic efficiency and poor accuracy, and unable to quickly and effectively locate the cause of the fault.

Method used

By obtaining multi-dimensional data of the fault object, including text logs, time series data, topological structure, operation records and environmental parameters, and performing feature encoding, the system matches the data in the historical fault case library. By utilizing the complementarity of multi-dimensional data and the auxiliary diagnosis of the historical case library, the diagnostic conclusion and repair plan can be directly obtained.

Benefits of technology

It improves the accuracy and efficiency of fault diagnosis, avoids the limitations of single data type analysis, shortens diagnosis time, is applicable to various types of fault objects, and has good scalability and adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492503B_ABST
    Figure CN120492503B_ABST
Patent Text Reader

Abstract

The present application discloses a fault diagnosis method, device, medium and product, which relate to the field of data processing technology. By acquiring multi-dimensional data such as text logs, time series data, topological structure, operation records and environmental parameters, various factors that may be involved when a fault occurs are included. Different types of data can complement and verify each other, accurately describe the characteristics of the fault, avoid the limitations of single data type analysis, and effectively improve the accuracy of fault diagnosis. The different types of multi-dimensional data are then converted into target feature representations, and the target feature representations are matched with the data in the historical fault case library. When a historical fault with a matching degree greater than a preset matching threshold is found, the diagnostic conclusion and repair plan of the historical fault can be determined as the diagnostic conclusion and repair plan of the current fault. There is no need to analyze and reason for each new fault, which shortens the time for fault diagnosis and improves the efficiency of fault diagnosis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a fault diagnosis method, device, medium and product. Background Art

[0002] With the rapid development of the Industrial Internet and intelligent operations and maintenance technologies, complex systems such as data centers, intelligent manufacturing equipment, and distributed energy networks are prone to system failures. Rapidly locating the cause of the failure and remediating it are key to ensuring stable system operation. Related technologies often rely on sensor time series data or manual judgment, resulting in low diagnostic efficiency and poor accuracy. Summary of the Invention

[0003] The present application provides a fault diagnosis method, device, medium and product to at least solve the problems of low fault diagnosis efficiency and poor diagnostic accuracy in related technologies.

[0004] In a first aspect, the present application provides a fault diagnosis method, comprising:

[0005] Obtain target multi-dimensional data of the fault object; target multi-dimensional data includes text logs, time series data, topology structure, operation records and environmental parameters;

[0006] Perform feature encoding on the target multi-dimensional data to obtain target feature representation;

[0007] Matching is performed in a historical fault case library based on target feature representation; the historical fault case library includes a variety of historical faults;

[0008] When the matching degree between the target feature representation and the historical fault is greater than or equal to a preset matching threshold, the diagnostic conclusion and repair solution corresponding to the matched historical fault are obtained as the diagnostic conclusion and repair solution of the fault object.

[0009] In a second aspect, the present application further provides a fault diagnosis device, comprising:

[0010] An acquisition module is used to obtain target multi-dimensional data of the fault object; the target multi-dimensional data includes text logs, time series data, topology structure, operation records and environmental parameters;

[0011] A feature encoding module is used to perform feature encoding on the target multi-dimensional data to obtain target feature representation;

[0012] A matching module is used to perform matching in a historical fault case library based on target feature representation; the historical fault case library includes multiple historical faults;

[0013] The acquisition module is also used to obtain the diagnostic conclusion and repair plan corresponding to the matched historical fault as the diagnostic conclusion and repair plan of the fault object when the matching degree between the target feature representation and the historical fault is greater than or equal to a preset matching threshold.

[0014] In a third aspect, the present application further provides an electronic device, comprising:

[0015] memory for storing computer programs;

[0016] A processor is configured to implement the steps of the method according to the first aspect when executing a computer program.

[0017] In a fourth aspect, the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein when the computer program is executed by a processor, the steps of the method of the first aspect are implemented.

[0018] In a fifth aspect, the present application also provides a computer program product, comprising a computer program, which implements the steps of the method of the first aspect when executed by a processor.

[0019] The present application provides a fault diagnosis method, device, medium and product. By acquiring multi-dimensional data such as text logs, time series data, topological structures, operation records and environmental parameters, it includes various factors that may be involved when a fault occurs. Different types of data can complement and verify each other, more accurately describing the characteristics of the fault, avoiding the limitations of single data type analysis, and effectively improving the accuracy of fault diagnosis. The different types of multi-dimensional data are then converted into target feature representations; then, the target feature representations are matched with the data in the historical fault case library. When a historical fault with a matching degree greater than a preset matching threshold is found, the diagnostic conclusion and repair plan of the historical fault can be determined as the diagnostic conclusion and repair plan of the current fault. There is no need to analyze and reason for each new fault, which shortens the time of fault diagnosis and improves the efficiency of fault diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0021] Figure 1 An application scenario diagram corresponding to a fault diagnosis method provided in an embodiment of the present application;

[0022] Figure 2 A flowchart of a fault diagnosis method provided in one embodiment of the present application;

[0023] Figure 3 A schematic diagram of a data clock synchronization process according to an embodiment of the present application;

[0024] Figure 4 A flowchart of a fault diagnosis method provided in another embodiment of the present application;

[0025] Figure 5 A schematic diagram of the structure of a fault diagnosis device provided in one embodiment of the present application;

[0026] Figure 6 A schematic diagram of the structure of an electronic device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0027] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0028] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0029] With the rapid development of the industrial Internet and intelligent operation and maintenance technologies, data centers, intelligent manufacturing equipment, distributed energy networks and other systems often experience system failures. Quickly locating the cause of the failure and repairing it is the key to ensuring the stable operation of the system. In related technologies, fault diagnosis methods usually rely on sensor time series data or manual experience judgment, and often only analyze a single type of data, for example, only analyzing the time series data of equipment operation, or only checking the system's text logs. However, since the occurrence of a fault may be the result of the combined action of multiple factors, different types of data are interrelated. Relying solely on a single data may lead to missed diagnosis or misdiagnosis, and it is impossible to quickly and effectively formulate accurate diagnostic conclusions and repair plans. Therefore, there are problems with low fault diagnosis efficiency and poor diagnostic accuracy.

[0030] Therefore, when addressing the technical issues of the aforementioned related technologies, the first consideration is that fault diagnosis requires the integration of multiple aspects of information. Therefore, during fault diagnosis, multidimensional data of the fault object can be obtained, including text logs, time series data, topological structures, operation records, and environmental parameters, to comprehensively account for all possible factors involved in the fault. Next, because directly processing multidimensional data is difficult, feature encoding is performed on the multidimensional data, converting it into a feature representation. This not only reduces data complexity but also extracts key feature information, facilitating subsequent analysis and processing. After obtaining the feature representation, historical fault cases are used to assist in diagnosis. The feature representation is matched against the historical fault cases. When the feature representation matches the historical fault case, the diagnostic conclusion and repair plan for the matching historical fault are directly obtained as the solution to the current fault object. Thus, by integrating multidimensional data and leveraging historical fault cases, the problem of fault diagnosis based on a single data set in related technologies can be avoided, improving the efficiency and accuracy of fault diagnosis.

[0031] Figure 1 This is an application scenario diagram corresponding to the fault diagnosis method provided in one embodiment of the present application, such as Figure 1 As shown, the application scenario provided by this embodiment includes: a device to be repaired 10 and a fault diagnosis device 40, wherein the device to be repaired 10 is in communication with the fault diagnosis device 40. Specifically, when it is necessary to perform fault diagnosis on the device to be repaired 10, the fault diagnosis device 40 obtains multi-dimensional data of the device to be repaired 10, which includes text logs, time series data, topological structures, operation records, and environmental parameters; then performs feature encoding on the multi-dimensional data to obtain a feature representation; the feature representation is searched and matched in a historical case library, which stores a variety of historical fault cases; after retrieving similar historical cases, the diagnostic conclusion and repair plan of the similar historical case are obtained as the diagnostic conclusion and repair plan of the device to be repaired 10.

[0032] The following specific embodiments describe in detail the technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0033] Figure 2 A flowchart of a fault diagnosis method provided in one embodiment of the present application is shown as follows: Figure 2As shown, the execution subject of this embodiment is a fault diagnosis device, which can be implemented by a computer program or a medium storing a relevant computer program, such as a USB flash drive and / or an optical disk; or it can be implemented by a physical device integrated or installed with a relevant computer program, such as a chip, an electronic device, etc. The electronic device can be a computer or a server, etc. A fault diagnosis method provided in this embodiment includes the following steps:

[0034] S201. Obtain target multi-dimensional data of the fault object; the target multi-dimensional data includes text logs, time series data, topology structure, operation records and environmental parameters.

[0035] Text logs are data stored in text format, including various events, error messages, and status changes recorded during the operation of a faulty device or system. For example, when a computer system crashes, it will record the crash time, program module name, error code, and other information in a log. Text logs detail the abnormalities during system operation and provide important clues for fault diagnosis.

[0036] Time series data is a series of data points recorded in chronological order, reflecting the changes in various parameters over time during the operation of a faulty device or system. For example, in industrial production, parameters such as temperature, pressure, and speed of faulty equipment will change over time. By collecting time series data, we can analyze the trends in the faulty equipment's operating status and determine whether there are any abnormal fluctuations.

[0037] The topology structure describes the connections and interactions between components within a faulty object. For example, in a computer network system, a topology diagram can show the physical connections and data transmission paths between nodes in the network. By analyzing the topology structure, we can clearly understand the fault's propagation path and impact range within the system, helping to pinpoint the specific location of the fault.

[0038] Operation logs record various user or system operations on faulty objects, including information such as the time, type, and object of the operation. For example, in a database system, operations such as inserting, deleting, and modifying a table are all recorded. Operation logs can help determine whether a fault is caused by improper operation, providing important evidence for fault diagnosis.

[0039] Environmental parameters refer to the external environmental conditions of the faulty device, such as temperature, humidity, and voltage. It's important to note that changes in these parameters can cause device failures. For example, electronic equipment operating in high-temperature environments may experience performance degradation or even freeze. Therefore, obtaining these environmental parameters is crucial for analyzing the cause of a fault.

[0040] S202: Perform feature encoding on the target multi-dimensional data to obtain target feature representation.

[0041] It should be noted that since different types of data have different formats and characteristics, in order to be able to uniformly analyze and process the data, it is necessary to feature encode the target multi-dimensional data.

[0042] Alternatively, natural language processing techniques, such as word embedding, can be used on text logs to map each word in the text log to a low-dimensional vector representation, thereby converting the text log into vector-based feature data. By extracting and encoding key words and phrases in the text log, the semantic information in the text can be converted into numerical features that can be processed by computers.

[0043] Alternatively, signal processing techniques such as Fourier transform and wavelet transform can be used to convert time series data from the time domain to the frequency domain to extract the frequency characteristics and trend characteristics of the data. Alternatively, a long short-term memory network can be used to learn the sequence information of time series data to obtain feature vectors that can reflect the law of data change.

[0044] Optionally, a graph neural network can be used for the topological graph structure to encode the information of the nodes and edges in the graph structure, learn the feature representation of the nodes and the relationship features between the nodes, and thus obtain the feature vector of the topological graph structure.

[0045] Optionally, different operation types may be numerically encoded for the operation records, and combined with information such as the operation time to construct a feature vector.

[0046] Optionally, for environmental parameters, their values ​​can be directly normalized and preprocessed to serve as feature data.

[0047] Finally, the feature data of each dimension are fused to obtain the target feature representation that can fully reflect the state of the fault object.

[0048] S203 : Matching is performed in a historical fault case library based on the target feature representation; the historical fault case library includes a variety of historical faults.

[0049] The historical fault case library is a knowledge base that stores a variety of historical fault cases. The historical fault case library is established by analyzing and summarizing a large number of faults that have occurred. It contains the characteristics of various faults, as well as the corresponding diagnostic conclusions and repair solutions.

[0050] Optionally, each historical fault in the historical fault case library stores a corresponding feature representation, diagnostic conclusion, and repair solution. Alternatively, when matching the target feature representation in the historical fault case library, the historical fault description is first converted into a feature representation, and then the similarity between the target feature representation and the historical fault feature representation is calculated.

[0051] Optionally, methods such as cosine similarity and Euclidean distance may be used to calculate the similarity between the target feature representation and the historical fault feature representation.

[0052] S204: When the matching degree between the target feature representation and the historical fault is greater than or equal to a preset matching threshold, the diagnostic conclusion and repair solution corresponding to the matched historical fault are output as the diagnostic conclusion and repair solution of the fault object.

[0053] The preset matching threshold is a similarity value preset according to actual application scenarios and requirements.

[0054] Specifically, when the target feature representation matches a historical fault to a certain degree and reaches or exceeds the preset matching threshold, it indicates that the state of the current fault object is similar to the historical fault. Therefore, the diagnostic conclusion and repair plan of the historical fault can be directly used to quickly resolve the current fault of the fault object.

[0055] The fault diagnosis method provided by the embodiment of the present application, by acquiring multi-dimensional data such as text logs, time series data, topological structures, operation records, and environmental parameters, includes various factors that may be involved when a fault occurs. Different types of data can complement and verify each other, more accurately describing the characteristics of the fault, avoiding the limitations of single data type analysis, and effectively improving the accuracy of fault diagnosis. The different types of multi-dimensional data are then converted into target feature representations; then, the target feature representations are matched with the data in the historical fault case library. When a historical fault with a matching degree greater than a preset matching threshold is found, the diagnostic conclusion and repair plan of the historical fault can be determined as the diagnostic conclusion and repair plan of the current fault. There is no need to analyze and reason for each new fault, which shortens the time for fault diagnosis and improves the efficiency of fault diagnosis.

[0056] Furthermore, the fault diagnosis method provided by the embodiments of this application is applicable to a wide variety of faulty objects, whether industrial equipment, computer systems, or other systems. As long as the corresponding multi-dimensional data can be obtained, the fault diagnosis method provided by this embodiment can be used for fault diagnosis. Furthermore, as the historical fault case library is updated and expanded, new fault modes and diagnostic methods can be continuously learned, adapting to new fault situations, demonstrating excellent scalability and adaptability.

[0057] In order to facilitate a better understanding of the fault diagnosis method provided in the embodiment of the present application, two specific cases are also provided in this embodiment for illustration.

[0058] Case 1:

[0059] In a computer server cluster operation and maintenance scenario of a certain enterprise, when one of the servers experiences a fault such as slow operation or frequent crashes, the fault diagnosis method provided in this embodiment is activated.

[0060] First, obtain multi-dimensional data of the server, including system logs (text logs), which record various error messages that occur during system operation, such as memory overflow error prompts; data such as CPU and memory usage that change over time (time series data); the connection relationship of the server in the network topology (topology structure); recent operation records of software installation, configuration modification, etc. performed by the system administrator on the server; and environmental parameters such as temperature and humidity in the server room.

[0061] Next, we use word embedding technology to encode system logs, converting log text into vector features. We use long-short-term memory networks to process time series data and extract its time series features. We use graph neural networks to encode topological structures. We numerically encode operation records and normalize environmental parameters. We then fuse the features from each dimension to obtain the target feature representation.

[0062] Finally, the target feature representation is matched with data in the historical fault case library, and the matching degree is calculated using cosine similarity. If a historical fault case with a matching degree greater than a preset matching threshold (e.g., 0.8) is found, the corresponding diagnostic conclusion is obtained as "Insufficient memory leads to system performance degradation and crash," and the repair solution is "Increase server memory and optimize the system memory allocation strategy." This is used as the diagnostic conclusion and repair solution for the current server fault.

[0063] Case 2:

[0064] In a factory's production line, a CNC machine tool experienced a malfunction that caused a decrease in machining accuracy.

[0065] First, the operation logs (text logs) of CNC machine tools were collected, which recorded the alarm information during the operation of the machine tools; parameters such as spindle speed and feed rate that change over time (time series data); the connection relationship between the various components of the machine tools (topological structure); the operator's recent operation records such as parameter adjustment and tool replacement on the machine tools; and environmental parameters such as temperature and voltage in the workshop.

[0066] Secondly, natural language processing technology is used to encode the operation logs, the frequency characteristics of the time series data are extracted through Fourier transform, and the topological graph structure is encoded using graph neural network. After corresponding preprocessing and encoding of the operation records and environmental parameters, the target feature representation is obtained by fusion.

[0067] Finally, a match is performed within the historical fault case library, and the Euclidean distance between the target feature representation and each historical fault feature representation is calculated. When a historical fault is found with a match that meets a preset matching threshold (e.g., 0.75), the diagnostic conclusion is "cutting tool wear leading to reduced machining accuracy," and the repair solution is "replace the tool and recalibrate the machine parameters."

[0068] As an optional implementation, based on any of the above embodiments, feature encoding is performed on the target multi-dimensional data to obtain a target feature representation, including:

[0069] Specifically, semantic features and fault keywords are extracted from the text log, and feature vectors of the semantic features and fault keywords are fused to obtain feature representation of the text log.

[0070] Optionally, a pre-trained language model in natural language processing, such as BERT (Bidirectional Encoder Representations from Transformers), can be used to input the text log into the pre-trained language model to obtain the semantic feature vector of the text. A keyword extraction algorithm, such as the TF-IDF (Term Frequency-Inverse Document Frequency) algorithm, can be used to extract fault keywords and map them into vector form. Finally, the semantic feature vector and the fault keyword vector can be concatenated or weighted summed to obtain the feature representation of the text log.

[0071] Specifically, feature extraction is performed on the time series data according to multiple preset time scales, and the mean and variance of the time series data at different time scales are calculated. The means and variances of multiple time scales are filled into a preset matrix in chronological order to obtain the feature representation of the time series data.

[0072] The preset time scale is a pre-set time scale for dividing data. Optionally, the preset time scale can be minute-level, hour-level, day-level, etc. For example, for temperature monitoring time series data of a device in an industrial production process, if the preset time scale is minute-level, hour-level, and day-level, then the data can be analyzed per minute, per hour, and per day, respectively.

[0073] For each preset time scale, calculate the mean and variance of the time series data within that time scale. The mean reflects the average level of the time series data within that time scale, while the variance reflects the degree of fluctuation within that time scale. Taking the minute-level time scale as an example, calculating the mean and variance of all temperature monitoring data within each minute can capture the stability and drastic changes in temperature within that minute.

[0074] The structure of the preset matrix can be designed based on the time scale and the order of the time series data. For example, the preset matrix can arrange the means and variances at different time scales in chronological order. Assuming that the time scales include minutes, hours, and days, then for each time point, the mean and variance at the minute level can be recorded first, then the mean and variance at the hour level, and finally the mean and variance at the day level, and then filled in the corresponding positions of the preset matrix in sequence, thereby constructing the feature matrix of the time series data.

[0075] Specifically, a first preset encoding algorithm is used to perform feature encoding on the node attributes, edge attributes and full-graph features of the topological graph structure to obtain a feature representation of the topological graph structure.

[0076] Optionally, the first preset encoding algorithm can use a graph attention network to calculate the attention weights between nodes, learn the feature representations of node attributes and edge attributes, and aggregate node features to obtain the feature representation of the entire graph, thereby converting the topological graph structure into a feature vector form.

[0077] Specifically, a second preset coding algorithm is used to perform feature coding and feature fusion on the operation type, operation time, and risk level of the operation in the operation record to obtain a feature representation of the operation record.

[0078] Optionally, the second preset encoding algorithm can be to perform one-hot encoding on the operation type, convert the operation time into a timestamp and normalize it, numerically encode the risk level of the operation, and then splice or otherwise fuse the three-part feature vectors to obtain a feature representation of the operation record.

[0079] Specifically, the environmental parameters are normalized and the data dimension is reduced, and feature extraction is performed through an autoencoder to obtain a feature representation of the environmental data.

[0080] Optionally, a normalization method such as minimum-maximum normalization is first used to map the numerical values ​​of the environmental parameters to the interval [0,1], and then data dimensionality reduction is performed through principal component analysis to reduce the data dimension. Finally, the processed data is input into an autoencoder, and the feature vector of the environmental parameter is extracted through the encoder part of the autoencoder.

[0081] Figure 3A schematic diagram of a data clock synchronization process provided in an embodiment of the present application; as an optional implementation, based on any of the above embodiments, Figure 3 As shown in , before encoding the target multi-dimensional data to obtain the target feature representation, the following is also included:

[0082] The target multi-dimensional data is clock synchronized to obtain initial synchronization data; the initial synchronization data is software error compensated to obtain intermediate synchronization data; the intermediate synchronization data is aligned across sampling rates using a dynamic time warping algorithm to obtain the target multi-dimensional data after time synchronization.

[0083] It should be noted that because the target multidimensional data may originate from different sensors, devices, or systems, the clocks of their data sources may not be completely consistent, resulting in time stamp deviations in the target multidimensional data. Therefore, the purpose of clock synchronization is to unify the data from various data sources onto a common time base, eliminating time deviations caused by clock discrepancies.

[0084] Clock synchronization can be achieved through various methods, such as Network Time Protocol (NTP)-based synchronization and GPS-based synchronization. In industrial environments, specialized hardware synchronization devices or methods can also be used to meet high-precision time requirements. After clock synchronization, the resulting initial synchronized data has a higher degree of consistency in the time dimension.

[0085] Software error compensation is the process of correcting systematic errors that may be introduced during the acquisition, transmission, and processing of target multi-dimensional data. Errors can arise from a variety of factors, including sensor accuracy limitations, deviations in data acquisition equipment, and interference in communication links.

[0086] Specifically, software error compensation typically uses mathematical models or algorithms to adjust data. For example, if a sensor's measurement value has a certain offset error, it can be compensated by adding or subtracting the corresponding offset during data processing. Linear errors can be corrected through methods such as linear fitting. Intermediate synchronized data that has undergone software error compensation is more accurate and reliable.

[0087] It should be noted that in practical applications, different data sources may have different sampling rates, which can lead to inconsistencies in the data along the time axis. To facilitate subsequent analysis, the Dynamic Time Warping (DTW) algorithm is used to align the intermediate synchronized data across sampling rates.

[0088] For example, taking time series data and environmental parameters as an example, if the sampling rate of time series data is high and the sampling rate of environmental parameters is low, the DTW algorithm dynamically bends the time axis by calculating the similarity between data sequences, aligning data with different sampling rates to the same time scale, and finally obtaining the target multi-dimensional data after time synchronization.

[0089] Among them, the DTW algorithm is a cross-sampling rate alignment method that calculates the optimal nonlinear matching path between two time series so that data with different sampling rates can be aligned on the time axis.

[0090] The fault diagnosis method provided in this embodiment of the application effectively eliminates deviations in timestamps and measurements of target multidimensional data through clock synchronization and software error compensation, improving the accuracy and consistency of the target multidimensional data. Cross-sampling rate alignment enables multidimensional data of different sampling rates to be fused on a unified time scale, avoiding information loss or mismatching caused by sampling rate differences.

[0091] As an optional implementation manner, based on any of the above embodiments, the method further includes the following steps:

[0092] First, when the matching degree between the target feature representation and the historical fault is less than the preset matching threshold, a fast causal discovery algorithm is used to mine the causal relationship between different features from the target feature representation to construct a causal graph and the confidence level corresponding to the causal graph; and feature fusion and attention calculation are performed on the target feature representation through cross-attention to determine the comprehensive feature vector representation and attention weight.

[0093] Optionally, the fast causal discovery algorithm can be the PCFast algorithm or the TETRAD algorithm, which mines causal relationships between different features within the target feature representation. Different features refer to different feature representations within the target feature representation. For example, analysis reveals a causal relationship between persistently high CPU usage and insufficient system memory. A causal graph is then constructed, and the confidence level corresponding to the causal graph is calculated algorithmically. Simultaneously, a cross-attention mechanism is used to perform feature fusion and attention calculation on the target feature representation, determining a comprehensive feature vector representation and attention weights to highlight the impact of key features on fault diagnosis.

[0094] It should be noted that the rapid causal discovery algorithm is a specialized algorithm for mining causal relationships in data. Its principle is based on causal assumptions and statistical dependencies. Through a series of conditional independence tests and causal direction judgments, it gradually constructs a causal graph structure between features, presenting the causal connections between features in an intuitive graphical manner. During this process, the rapid causal discovery algorithm evaluates the confidence level of each causal relationship, which quantifies the credibility of the causal relationship. For example, in industrial production process monitoring, the rapid causal discovery algorithm can mine the causal chains between interrelated process parameters (such as temperature, pressure, and flow). For example, temperature changes lead to pressure changes, which in turn affect flow, helping to understand the propagation path of faults between different parameters.

[0095] It should be noted that the core idea of ​​the cross-attention mechanism is to establish an attention relationship between two different feature sequences to capture the mutual influence and dependence between them. By processing the target feature representation through cross-attention, on the one hand, it is possible to achieve deep feature fusion and organically integrate features from different sources or different types; on the other hand, it is possible to calculate the relative importance weight of each feature in the diagnostic task, that is, the attention weight. For example, when processing multi-dimensional features containing equipment operation data and environmental data, cross-attention can determine which operation data features are closely related to environmental data features, and highlight those features that are more critical to fault diagnosis. For example, when the ambient temperature change has a significant impact on the temperature anomaly of a certain component of the equipment, cross-attention will give these two features a higher attention weight.

[0096] Secondly, the comprehensive feature vector is mapped to a basic probability distribution based on a neural network to obtain an initial basic probability distribution. Based on the causal graph, confidence level, and attention weight, the initial basic probability distribution is modified to obtain the target basic probability distribution.

[0097] Based on a neural network, the comprehensive feature vector is mapped into a basic probability distribution to obtain an initial basic probability distribution. This initial basic probability distribution is then modified based on the causal graph, confidence level, and attention weight to obtain a target basic probability distribution that better reflects the actual fault situation.

[0098] It should be noted that neural networks have powerful nonlinear mapping capabilities and can learn complex feature-to-probability mapping relationships. In this embodiment, a neural network is used as a mapping tool for comprehensive feature vectors to basic probability distributions, and common network structures such as multilayer perceptrons can be used. By designing a suitable network architecture and training strategy, the neural network can learn the relationship between the comprehensive feature vector and the basic probability distribution. Specifically, during the training process, a large amount of known fault case data, including feature vectors and corresponding fault types and their basic probability distributions, is used to supervise the learning of the neural network. The trained neural network can output an initial basic probability distribution based on the input comprehensive feature vector, reflecting the preliminary probability estimates of each possible fault hypothesis. For example, for a motor fault diagnosis scenario, the trained neural network can map the comprehensive feature vectors of the motor's vibration characteristics, current characteristics, etc. into the initial basic probability distributions of different motor fault types (such as bearing faults, winding faults, rotor faults, etc.).

[0099] Then, based on the causal graph, confidence level, and attention weight, the initial basic probability distribution is modified to obtain the target basic probability distribution.

[0100] Among them, the causal graph provides causal structure information between features, and the confidence and attention weight adjust the probability distribution from the perspectives of the credibility of the causal relationship and the importance of the features, respectively, so that the corrected target basic probability distribution is more in line with the actual fault situation.

[0101] Specifically, the initial probability distribution can be transferred from the cause feature to the result feature according to the causal relationship path in the causal graph. For example, if feature A is the cause of feature B, then the probability distribution of feature A can be partially transferred to feature B. During the probability transfer process, the transfer strength needs to be adjusted according to the confidence of the causal relationship. Causal relationships with higher confidence should transfer more probabilities, while causal relationships with lower confidence should transfer less probability. After the probability transfer is completed, the probability ratio of each feature needs to be adjusted according to the attention weight. Features with higher attention weights should be assigned a larger probability ratio, while features with lower attention weights should be assigned a smaller probability ratio. Through the above steps, the corrected target basic probability distribution can be obtained.

[0102] Thirdly, evidence theory is used to reason about the target basic probability distribution to obtain the confidence values ​​of multiple fault hypotheses.

[0103] Evidence theory is a mathematical theory for dealing with uncertainty and multi-source information fusion. Confidence values ​​not only consider the basic probability distribution itself but also incorporate evidential information contained in factors such as causal diagrams, confidence levels, and attention weights, thereby more accurately reflecting the credibility of each fault hypothesis in the current situation. For example, in fault diagnosis of a chemical production process, there are multiple possible causes (such as pipeline leaks, pump failures, and reactor failures). Evidence theory reasoning can derive a comprehensive confidence level for each fault hypothesis based on its modified basic probability and other evidential information, thereby determining the most likely cause.

[0104] Specifically, evidence theory is used to reason about the target basic probability distribution to obtain the confidence values ​​of multiple fault hypotheses. Evidence theory can effectively integrate evidence from different sources and calculate the comprehensive confidence of each fault hypothesis, providing a more reliable basis for the final fault diagnosis.

[0105] Finally, based on the confidence values ​​of each fault hypothesis and the target multi-dimensional data, the Wright algorithm is used to match the expert rules in the expert rule library. If the matching confidence level is greater than the preset matching confidence level, the diagnostic conclusion and repair plan corresponding to the successfully matched expert rule are output.

[0106] For example, after obtaining the confidence levels of multiple fault hypotheses such as "insufficient memory" and "hardware failure," the Wright algorithm is used to match the expert rules in the expert rule base based on the confidence values ​​of each fault hypothesis and the target multi-dimensional data. If the matching confidence level of the "insufficient memory" fault hypothesis and the rule in the expert rule base "When the CPU usage rate is continuously too high and the memory usage rate is at a high level for a long time, it is highly likely that insufficient memory has caused system performance to degrade" is greater than the preset matching confidence level (e.g., 0.7), the diagnostic conclusion corresponding to the expert rule, "Insufficient memory has caused system performance to degrade and freeze," and the repair solution, "Increase server memory and optimize the system memory allocation strategy," are output.

[0107] Among them, the Wright algorithm is an efficient rule matching algorithm that can quickly find rules that match the current data in a large number of rules. Among them, the expert rule base is a series of fault diagnosis rules summarized by domain experts based on long-term practical experience. The expert rule base contains information such as various known fault modes, diagnostic logic, and corresponding repair solutions. In this embodiment, the fault hypothesis confidence value and target multi-dimensional data obtained based on evidence theory reasoning are used as input, and the Wright algorithm is used to match the expert rules in the expert rule base. When the matching credibility reaches the preset matching credibility requirement, the diagnostic conclusion and repair solution corresponding to the successfully matched expert rule can be directly output, providing clear guidance for on-site fault handling personnel, thereby improving the efficiency and accuracy of fault diagnosis.

[0108] For example, in automobile fault diagnosis, when the evidence theory deduces that a certain fault hypothesis (such as engine ignition system failure) has a high confidence level, the Wright algorithm finds a rule in the expert rule base that matches the fault hypothesis and related vehicle data (such as fault indicator light status, engine speed, ignition signal, etc.), and then outputs the diagnostic conclusion contained in the rule (such as ignition coil failure, spark plug failure, etc.) and the corresponding repair plan (such as replacing the ignition coil, cleaning the spark plug, etc.).

[0109] Optionally, the diagnosis conclusion and repair solution corresponding to the successfully matched expert rule are stored in the historical fault case library, and the update of the historical fault case library is completed.

[0110] The fault diagnosis method provided in the embodiment of the present application can analyze the target feature representation and multi-dimensional data more comprehensively and deeply through the rapid causal discovery algorithm to mine the causal relationship between features, the cross-attention mechanism to determine the importance and relevance of features, neural network mapping and evidence theory reasoning, so as to obtain more accurate fault diagnosis conclusions. Compared with the traditional method that relies only on a single method, it can effectively reduce the probability of misdiagnosis and missed diagnosis. The matching ability based on the Wright algorithm combined with the expert rule base can quickly and accurately draw diagnostic conclusions and repair plans that match the current fault situation, providing a reliable basis for fault repair. In addition, the method of this embodiment is applicable to a variety of different types of fault diagnosis scenarios. Whether it is industrial production equipment, electronic systems or other systems, as long as the corresponding multi-dimensional data can be obtained, the method can be applied for fault diagnosis.

[0111] As an optional implementation manner, based on any of the above embodiments, the method further includes the following:

[0112] Based on the diagnostic conclusions and repair plans corresponding to historical faults, verification is carried out in a sandbox environment. If the verification is successful, the target fault diagnosis information is output; the target fault diagnosis information includes the cause of the fault, the impact of the fault and the repair plan.

[0113] The sandbox environment is an isolated, controlled virtual operating environment that simulates real-world failure scenarios without impacting actual production systems or equipment. By deploying historical failure diagnosis and repair solutions in the sandbox environment, the effectiveness, feasibility, and potential risks of the repair solutions can be tested and evaluated.

[0114] For example, in software system fault diagnosis, after proposing a fix for a historical fault, a sandbox environment is first built with a software architecture and data environment similar to the actual operating environment. The fix is ​​then implemented, and the system's operation is observed in the sandbox to verify whether the fault has been resolved and whether new problems may arise. In industrial automation control systems, the sandbox environment can be a virtual control system simulation platform. The fix for the historical fault is applied to the control system simulation platform, simulating various operating conditions to verify whether the fix accurately resolves the fault.

[0115] Specifically, the target fault diagnosis information is only output if the sandbox environment is successfully verified. This sandbox verification mechanism effectively prevents erroneous or incomplete diagnostic conclusions and repair solutions from being directly applied in practice, thereby reducing the risk of secondary failures caused by incorrect repairs.

[0116] Specifically, the target fault diagnosis information includes the cause of the fault, the impact of the fault, and the repair plan. Among them, the cause of the fault is a detailed description of the factors that caused the fault, and the cause of the fault is derived based on a series of analysis steps performed above, such as causal relationship mining, feature fusion, and attention calculation. The cause of the fault helps users understand the root cause of the fault. Among them, the impact of the fault describes the adverse consequences of the fault on the system, equipment, production process, etc., including but not limited to performance degradation, functional abnormality, safety hazards, etc. The impact of the fault can help users assess the severity of the fault. Among them, the repair plan is a specific solution proposed for the cause of the fault. The repair plan that has been verified in the sandbox environment has higher reliability and feasibility, and users can perform fault repair operations according to the repair plan.

[0117] The fault diagnosis method provided in the embodiment of the present application can actually test the diagnostic conclusions and repair solutions through the verification link of the sandbox environment. Only the diagnostic conclusions that have been verified to be valid by the sandbox environment will be output, further ensuring that the output fault causes, fault impacts and repair solutions are consistent with the actual situation, ensuring that they can effectively solve problems in actual applications, avoiding the failure of fault repair or the occurrence of new problems due to unverified solutions, and improving the reliability and stability of the fault diagnosis conclusions.

[0118] As an optional implementation manner, based on any of the above embodiments, the method further includes the following steps:

[0119] First, the fault location and version information of the fault object are determined based on the target multi-dimensional data.

[0120] The fault location can be a specific part of the device, the location of a component, the code location in the software, etc., and the version information can help determine the version status of the fault object.

[0121] For example, for a piece of industrial equipment, the target multi-dimensional data might include the operating status of each component, sensor data, and maintenance records. Comprehensive analysis of this data can pinpoint a specific component of the equipment as the fault location. Furthermore, based on the device's update history and software system version information, the version status of the faulty object can be determined.

[0122] Secondly, based on the fault location and version information of the fault object, a query is performed in a preset mapping table to determine the target data collection tool; wherein the preset mapping table includes a mapping relationship between the fault location and version information of the fault object and the data collection tool.

[0123] The preset mapping table is a mapping relationship table pre-established based on a large amount of practical experience and technical knowledge, and includes a mapping relationship between the fault location and version information of the fault object and the data collection tool.

[0124] Optionally, the query process of the preset mapping table can be implemented by database query, hash table search, etc., and the corresponding appropriate data collection tool can be quickly located in the preset mapping table according to the determined fault location and version information.

[0125] Finally, the target data collection tool is obtained from the tool acquisition platform according to the preset priority; the preset priority is that the official tool acquisition platform is greater than the third-party tool acquisition platform, and the target data collection tool is used to obtain component log information at the fault location of the fault object.

[0126] It should be noted that the setting of preset priorities is based on considerations of tool reliability and security. Official tools have usually undergone more rigorous testing and verification and have better compatibility and stability.

[0127] For example, official tool acquisition platforms can include software download centers provided by device manufacturers, officially certified tool libraries, etc.; third-party tool acquisition platforms can include unofficial open source software distribution platforms, tool libraries, etc. The target data collection tool is obtained from official tool acquisition platforms first according to preset priorities. If the official platform does not have the tool or it does not work properly, third-party platforms are considered.

[0128] The target data acquisition tool is used to obtain component log information at the fault location of the fault object. Component log information is a very important data source in the fault diagnosis process. It records detailed information such as the operating status, operation history, and abnormal events of the component at the fault location.

[0129] It should be noted that collecting log information at the fault location of the fault object through the target data collection tool can serve as a supplement to the target multi-dimensional data, which is helpful for finding the root cause of the fault at the component location.

[0130] Alternatively, if the faulty object is not connected to the internet, a local tool can be used to search for the faulty object first. If the local tool is not installed or the search fails, the diagnostic device's built-in tool can be used to collect component log information. Optionally, the diagnostic device's built-in tool can be downloaded and installed into the diagnostic device after determining the location and version information of the faulty object's components.

[0131] The fault diagnosis method provided in the embodiments of the present application uses target multi-dimensional data to determine the fault location and version information, enabling rapid and accurate location of the specific fault location and version status. By querying a mapping table, the corresponding data acquisition tool is found and prioritized from the official tool acquisition platform according to a preset priority. This ensures the reliability and security of the data acquisition tool and improves the reliability and credibility of the fault diagnosis process. Furthermore, the acquired component log information can provide more detailed and richer data for fault diagnosis, helping to more accurately determine the cause of the fault and the repair plan, thereby improving the accuracy of the fault diagnosis results.

[0132] As an optional implementation manner, based on any of the above embodiments, before obtaining the target multi-dimensional data of the fault object, the method further includes the following steps:

[0133] First, a diagnosis script input by a user is obtained, and a binary intermediate code is generated based on the diagnosis script.

[0134] Among them, the diagnostic script is a script file written by the user based on the preliminary understanding of the fault object, diagnostic requirements and diagnostic logic. The diagnostic script usually contains a series of instructions, operating steps and logical judgments to guide the fault diagnosis process.

[0135] The process of generating binary intermediate code involves compiling or interpreting the diagnostic script. Binary intermediate code is a form of code that is relatively independent of the specific execution environment. Optionally, using network assembly language technology and a low-level virtual machine toolchain, the diagnostic script can be compiled into binary intermediate code that supports x86-64, ARMv8, and PowerPC architectures. This binary intermediate code is compatible with formats such as Linux ELF and Windows PE, while retaining the underlying hardware operating capabilities, enabling the compilation of diagnostic scripts on different architectures and systems.

[0136] Secondly, a sandbox environment is created in the fault object, the binary intermediate code is input into the sandbox environment, and the system abstraction layer is loaded in the sandbox environment; wherein the system abstraction layer encapsulates the underlying functions of the system.

[0137] The sandbox environment is an isolated, controlled runtime environment that prevents diagnostic scripts from causing unnecessary impact or damage to the actual system of the fault target during execution. The system abstraction layer encapsulates the underlying system functions of the fault target and provides a unified, simplified interface for upper-layer applications. This allows diagnostic scripts to call the underlying system functions of the fault target to perform fault diagnosis operations without directly accessing the underlying system details of the fault target.

[0138] For example, the system abstraction layer can encapsulate low-level hardware access functions, system call interfaces, and data acquisition interfaces, abstracting and encapsulating these functions and providing them to diagnostic scripts through a set of common interfaces. This improves the versatility and portability of diagnostic scripts while also facilitating the management and control of low-level system functions.

[0139] Finally, a communication interface is determined based on the system type of the fault object; the communication interface is used to obtain target multi-dimensional data of the fault object.

[0140] Specifically, different system types have different data communication methods and interface specifications, so it's important to determine the appropriate communication interface based on the specific system type. For example, for a network-based distributed system, the communication interface might be a TCP / IP-based network interface; for an embedded system, it might be a serial communication interface or a CAN bus interface; and for a local computer system, it might be a file access interface or a shared memory interface. Determining the communication interface is crucial for subsequently accurately and efficiently acquiring the target multi-dimensional data of the faulty object.

[0141] The fault diagnosis method provided in the embodiment of the present application can effectively isolate the execution of the diagnostic script from the actual system environment by converting the diagnostic script input by the user into a binary intermediate code and running it in a sandbox environment, thereby preventing potential malicious code or erroneous operations in the diagnostic script from causing damage to the system and ensuring the security of the fault diagnosis process. The system abstraction layer is used to encapsulate the underlying functions of the system, so that the diagnostic script does not rely on the specific underlying implementation details of the system, thereby improving the versatility and portability of the diagnostic script on different system types. It is conducive to reusing the same diagnostic script or a diagnostic script that has been simply modified on a variety of different fault objects. Finally, the communication interface is accurately determined according to the system type of the fault object, which can ensure that the target multi-dimensional data is obtained in an efficient and stable manner.

[0142] Figure 4 A flowchart of a fault diagnosis method provided by another embodiment of the present application is provided; in order to better understand the fault diagnosis method provided by the present application, as Figure 4As shown in , the embodiment of the present application provides a complete fault diagnosis method. The execution subject of the fault diagnosis method provided in the embodiment of the present application is a fault diagnosis device. Specifically, the fault diagnosis method includes the following steps:

[0143] S301: Obtain a diagnosis script input by a user, and generate a binary intermediate code based on the diagnosis script.

[0144] S302: Create a sandbox environment in the fault object, input the binary intermediate code into the sandbox environment, and load the system abstraction layer in the sandbox environment.

[0145] Among them, the system abstraction layer encapsulates the underlying system functions.

[0146] S303: Determine a communication interface based on the system type of the fault object.

[0147] The communication interface is used to obtain target multi-dimensional data of the fault object.

[0148] S304: Obtain target multi-dimensional data of the fault object.

[0149] S305: Determine the fault location and version information of the fault object based on the target multi-dimensional data.

[0150] S306: Based on the fault location and version information of the fault object, query in a preset mapping table to determine a target data collection tool.

[0151] The preset mapping table includes a mapping relationship between the fault location of the fault object and version information and the data collection tool.

[0152] S307: Acquire the target data acquisition tool from the tool acquisition platform according to the preset priority.

[0153] Among them, the preset priority is that the official tool acquisition platform is higher than the third-party tool acquisition platform, and the target data collection tool is used to obtain component log information at the fault location of the fault object.

[0154] S308: Perform clock synchronization on the target multi-dimensional data to obtain initial synchronization data.

[0155] S309: Perform software error compensation on the initial synchronization data to obtain intermediate synchronization data.

[0156] S310: Use a dynamic time warping algorithm to align the intermediate synchronization data across sampling rates to obtain target multi-dimensional data after time synchronization.

[0157] S311 . Feature encoding is performed on the time-synchronized target multi-dimensional data to obtain target feature representation.

[0158] S312: Matching is performed in a historical fault case library based on the target feature representation, and it is determined whether the matching degree between the target feature representation and the historical fault is greater than or equal to a preset matching threshold.

[0159] Among them, the historical fault case library includes a variety of historical faults.

[0160] S313: When the matching degree between the target feature representation and the historical fault is greater than or equal to a preset matching threshold, the diagnosis conclusion and repair solution corresponding to the matched historical fault are obtained as the diagnosis conclusion and repair solution of the fault object.

[0161] S314. Verify the diagnostic conclusions and repair solutions corresponding to historical faults in a sandbox environment.

[0162] S315: If the verification is successful, output target fault diagnosis information.

[0163] The target fault diagnosis information includes the fault cause, fault impact and repair plan.

[0164] S316. When the matching degree between the target feature representation and the historical fault is less than the preset matching threshold, a fast causal discovery algorithm is used to mine the causal relationship between different features from the target feature representation to construct a causal graph and the confidence level corresponding to the causal graph; and feature fusion and attention calculation are performed on the target feature representation through cross-attention to determine the comprehensive feature vector representation and attention weight.

[0165] S317. Map the comprehensive feature vector to a basic probability distribution based on a neural network to obtain an initial basic probability distribution; based on the causal graph, confidence level, and attention weight, modify the initial basic probability distribution to obtain a target basic probability distribution.

[0166] S318. Use evidence theory to reason about the target basic probability distribution to obtain confidence values ​​of multiple fault hypotheses.

[0167] 319. Based on the confidence values ​​of each fault hypothesis and the target multi-dimensional data, the Wright algorithm is used to match the expert rules of the expert rule base.

[0168] 320. When the matching credibility is greater than the preset matching credibility, the diagnostic conclusion and repair plan corresponding to the expert rule that is successfully matched are output.

[0169] Optionally, when the matching reliability is less than the preset matching reliability, steps S316 to S319 may be repeated, or matching failure information may be directly output.

[0170] Optionally, before outputting the diagnosis conclusion and repair plan corresponding to the successfully matched expert rule, verification in a sandbox environment may be performed. For specific methods, refer to steps S314 to S315.

[0171] Figure 5 A schematic diagram of the structure of a fault diagnosis device provided in one embodiment of the present application is shown in FIG. Figure 5 As shown, the fault diagnosis device provided in this embodiment is located in an electronic device. The fault diagnosis device 40 provided in this embodiment includes: an acquisition module 41 , a feature encoding module 42 and a matching module 43 .

[0172] An acquisition module 41 is used to acquire target multi-dimensional data of a fault object; the target multi-dimensional data includes text logs, time series data, topological structures, operation records, and environmental parameters; a feature encoding module 42 is used to perform feature encoding on the target multi-dimensional data to obtain a target feature representation; a matching module 43 is used to perform matching in a historical fault case library based on the target feature representation; the historical fault case library includes a variety of historical faults; the acquisition module 41 is also used to obtain the diagnostic conclusion and repair plan corresponding to the matched historical fault as the diagnostic conclusion and repair plan of the fault object when the matching degree between the target feature representation and the historical fault is greater than or equal to a preset matching threshold.

[0173] Optionally, the feature encoding module 42, when performing feature encoding on the target multi-dimensional data to obtain the target feature representation, is specifically used to: extract semantic features and fault keywords in the text log, and perform feature vector fusion on the semantic features and fault keywords to obtain the feature representation of the text log; perform feature extraction on the time series data according to multiple preset time scales, and calculate the mean and variance of the time series data at different time scales, and fill the mean and variance of multiple time scales into a preset matrix in chronological order to obtain the feature representation of the time series data; use a first preset encoding algorithm to perform feature encoding on the node attributes, edge attributes and full-graph features of the topological graph structure to obtain the feature representation of the topological graph structure; use a second preset encoding algorithm to perform feature encoding on the operation type, operation time and risk level of the operation in the operation record and perform feature fusion to obtain the feature representation of the operation record; normalize and reduce the dimension of the environmental parameters, and perform feature extraction through an autoencoder to obtain the feature representation of the environmental data.

[0174] Optionally, the fault diagnosis device further includes a time alignment module.

[0175] Optionally, the time alignment module, before feature encoding the target multidimensional data to obtain the target feature representation, is specifically used to: perform clock synchronization on the target multidimensional data to obtain initial synchronization data; perform software error compensation on the initial synchronization data to obtain intermediate synchronization data; and use a dynamic time warping algorithm to align the intermediate synchronization data across sampling rates to obtain the target multidimensional data after time synchronization.

[0176] Optionally, the fault diagnosis device further includes a determination module and an output module.

[0177] Optionally, the determination module is specifically used to: when the matching degree between the target feature representation and the historical fault is less than a preset matching threshold, use a fast causal discovery algorithm to mine the causal relationship between different features from the target feature representation to construct a causal graph and the confidence corresponding to the causal graph; and perform feature fusion and attention calculation on the target feature representation through cross-attention to determine the comprehensive feature vector representation and attention weight; map the comprehensive feature vector to a basic probability distribution based on a neural network to obtain an initial basic probability distribution; based on the causal graph, confidence and attention weight, modify the initial basic probability distribution to obtain a target basic probability distribution; use evidence theory to infer the target basic probability distribution to obtain confidence values ​​of multiple fault hypotheses; the matching module 43 is also used to: based on the confidence values ​​of each fault hypothesis and the target multi-dimensional data, use the Wright algorithm to match with the expert rules of the expert rule library; the output module is used to: when the matching credibility is greater than the preset matching credibility, output the diagnostic conclusion and repair plan corresponding to the expert rule that successfully matched.

[0178] Optionally, the fault diagnosis device further includes a verification module.

[0179] Optionally, the verification module is specifically used to: verify in a sandbox environment based on the diagnostic conclusions and repair solutions corresponding to historical faults; if the verification is successful, output target fault diagnosis information; wherein the target fault diagnosis information includes the cause of the fault, the impact of the fault and the repair solution.

[0180] Optionally, the determination module is also used to: determine the fault location and version information of the fault object based on the target multi-dimensional data; based on the fault location and version information of the fault object, query in a preset mapping table to determine the target data collection tool; wherein the preset mapping table includes the mapping relationship between the fault location and version information of the fault object and the data collection tool; obtain the target data collection tool from the tool acquisition platform according to the preset priority; wherein the preset priority is that the official tool acquisition platform is greater than the third-party tool acquisition platform, and the target data collection tool is used to obtain the component log information at the fault location of the fault object.

[0181] Optionally, the fault diagnosis device further includes a creation module.

[0182] Optionally, before obtaining the target multi-dimensional data of the fault object, the acquisition module 41 is also used to: obtain the diagnostic script input by the user, and generate a binary intermediate code based on the diagnostic script; the creation module is specifically used to: create a sandbox environment in the fault object, input the binary intermediate code into the sandbox environment, and load the system abstraction layer in the sandbox environment; wherein the system abstraction layer encapsulates the underlying system functions; the determination module is also used to: determine the communication interface based on the system type of the fault object; the communication interface is used to obtain the target multi-dimensional data of the fault object.

[0183] It should be noted that the technical effects of the fault diagnosis device provided in this embodiment have been explained in the embodiment of the above-mentioned fault diagnosis method, and therefore, they will not be repeated in this embodiment.

[0184] Figure 6 A schematic diagram of the structure of an electronic device provided in one embodiment of the present application is shown in FIG. Figure 6 As shown, the electronic device 50 provided in an embodiment of the present application includes: a memory 51 and a processor 52.

[0185] The memory 51 stores a computer program, and the processor 52 is configured to run the computer program to execute the steps in any one of the above-mentioned fault diagnosis method embodiments.

[0186] In this embodiment, the processor 52 and the memory 51 are connected via a bus. The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. Buses can be categorized as address buses, data buses, control buses, etc. For ease of presentation, Figure 6 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0187] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps of any of the above-mentioned fault diagnosis method embodiments when running.

[0188] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0189] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any of the above-mentioned fault diagnosis method embodiments are implemented.

[0190] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above-mentioned fault diagnosis method embodiments are implemented.

[0191] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0192] The above is a detailed introduction to a fault diagnosis method, device, medium and product provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core ideas of the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.

Claims

1. A fault diagnosis method, characterized in that: include: Obtain target multi-dimensional data of the fault object; The target multi-dimensional data includes text logs, time series data, topological graph structures, operation records and environmental parameters; Performing feature encoding on the target multi-dimensional data to obtain a target feature representation; Matching is performed in a historical fault case library based on the target feature representation; the historical fault case library includes multiple historical faults; When the matching degree between the target feature representation and the historical fault is greater than or equal to a preset matching threshold, obtaining the diagnostic conclusion and repair solution corresponding to the matched historical fault as the diagnostic conclusion and repair solution of the fault object; The method further comprises: When the matching degree between the target feature representation and the historical fault is less than a preset matching threshold, a fast causal discovery algorithm is used to mine the causal relationship between different features from the target feature representation to construct a causal graph and the confidence level corresponding to the causal graph; and feature fusion and attention calculation are performed on the target feature representation through cross-attention to determine a comprehensive feature vector representation and attention weight. Mapping the comprehensive feature vector to a basic probability distribution based on a neural network to obtain an initial basic probability distribution; Based on the causal graph, the confidence level, and the attention weight, modifying the initial basic probability distribution to obtain a target basic probability distribution; Reasoning the target basic probability distribution using evidence theory to obtain confidence values ​​of multiple fault hypotheses; Based on the confidence values ​​of each fault hypothesis and the target multi-dimensional data, the Wright algorithm is used to match the expert rules of the expert rule base; When the matching credibility is greater than the preset matching credibility, the diagnosis conclusion and repair plan corresponding to the expert rule that matched successfully are output.

2. The method according to claim 1, characterized in that The feature encoding of the target multi-dimensional data to obtain a target feature representation includes: Extracting semantic features and fault keywords from the text log, and performing feature vector fusion on the semantic features and the fault keywords to obtain a feature representation of the text log; Extract features from the time series data according to multiple preset time scales, calculate the mean and variance of the time series data at different time scales, and fill the means and variances of the multiple time scales into a preset matrix in chronological order to obtain a feature representation of the time series data; Using a first preset encoding algorithm to perform feature encoding on the node attributes, edge attributes, and full-graph features of the topological graph structure to obtain a feature representation of the topological graph structure; Using a second preset coding algorithm to feature encode the operation type, operation time, and risk level of the operation in the operation record and perform feature fusion to obtain a feature representation of the operation record; The environmental parameters are normalized and the data dimension is reduced, and feature extraction is performed through an autoencoder to obtain a feature representation of the environmental data.

3. The method according to claim 1, characterized in that Before performing feature encoding on the target multi-dimensional data to obtain target feature representation, the method further includes: Performing clock synchronization on the target multi-dimensional data to obtain initial synchronization data; performing software error compensation on the initial synchronization data to obtain intermediate synchronization data; A dynamic time warping algorithm is used to align the intermediate synchronization data across sampling rates to obtain the target multi-dimensional data after time synchronization.

4. The method according to any one of claims 1 to 3, characterized in that The method also includes: Verify the diagnostic conclusions and repair solutions corresponding to the historical faults in a sandbox environment; If the verification is successful, target fault diagnosis information is output; wherein the target fault diagnosis information includes the cause of the fault, the impact of the fault and the repair plan.

5. The method according to any one of claims 1 to 3, characterized in that The method also includes: Determine the fault location and version information of the fault object based on the target multi-dimensional data; Based on the fault location and version information of the fault object, a query is performed in a preset mapping table to determine the target data collection tool; wherein the preset mapping table includes a mapping relationship between the fault location and version information of the fault object and the data collection tool; A target data acquisition tool is obtained from a tool acquisition platform according to a preset priority; wherein the preset priority is that the official tool acquisition platform is higher than the third-party tool acquisition platform, and the target data acquisition tool is used to obtain component log information at the fault location of the fault object.

6. The method according to any one of claims 1 to 3, characterized in that Before obtaining the target multi-dimensional data of the fault object, the method further includes: Obtaining a diagnosis script input by a user, and generating a binary intermediate code based on the diagnosis script; Creating a sandbox environment in the fault object, inputting the binary intermediate code into the sandbox environment, and loading a system abstraction layer in the sandbox environment; wherein the system abstraction layer encapsulates underlying system functions; A communication interface is determined based on the system type of the fault object; the communication interface is used to obtain target multi-dimensional data of the fault object.

7. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the fault diagnosis method according to any one of claims 1 to 6 when executing the computer program.

8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the fault diagnosis method according to any one of claims 1 to 6.

9. A computer program product, characterized in that The invention comprises a computer program, which implements the steps of the fault diagnosis method according to any one of claims 1 to 6 when the computer program is executed by a processor.

Citation Information

Patent Citations

  • GIS fault diagnosis method based on cases and fault reasoning

    CN112329937A