Equipment fault identification method and system and storage medium
Through dimensionality reduction processing, density clustering algorithm, SMOTE oversampling technology, adaptive weight allocation and parallel computing framework optimization simulation annealing algorithm, the rare fault identification capabilities and low computing efficiency in equipment fault identification are solved, and efficient and accurate fault diagnosis and real-time maintenance support are achieved.
Patent Information
- Application Number
- CN202510084472.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-20
AI Technical Summary
In the existing equipment fault identification technology, analog annealing algorithms are difficult to accurately identify rare faults, and their calculation efficiency is low, which cannot meet the real-time requirements.
Fault characteristics are extracted through dimensionality reduction processing, and fault patterns are classified using density clustering algorithm to identify rare faults; SMOTE oversampling technology is used to balance the data set, adaptive weight allocation strategy calculates the importance weight of the fault pattern, and integrates the weight information into the improved objective function; a parallel computing framework is used to optimize the simulation annealing algorithm to improve calculation efficiency.
Effectively identify and handle rare faults, improve the accuracy and efficiency of fault diagnosis, provide timely and precise decision-making support for equipment maintenance, and improve the overall equipment operation reliability.
Smart Images

Figure CN119939277A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of equipment fault identification, and in particular to an equipment fault identification method, system and storage medium. Background Art
[0002] In equipment fault identification, the simulated annealing algorithm faces the problem of insufficient discrimination ability. Due to the large differences in the characteristics of different fault data sets, directly inputting them into the algorithm equally will lead to low discrimination accuracy for certain fault modes. Although the importance of different data sets can be distinguished by assigning weights, how to reasonably set the weights itself is also a thorny issue. In addition, the current fault feature extraction method and objective function design are not accurate enough to fully mine the information contained in the fault data, which further limits the algorithm's discrimination ability.
[0003] At the same time, in actual applications, the types and causes of equipment failures are complex, including common problems such as wear and fatigue, as well as some rare special failures. For these rare failure modes, the existing simulated annealing algorithm is difficult to give accurate judgment results. How to take into account common and rare failures in the algorithm and balance judgment accuracy and generalization ability is an issue worthy of in-depth discussion.
[0004] In addition, equipment fault identification often has high requirements for real-time performance, requiring the algorithm to quickly give identification results so that maintenance measures can be taken in a timely manner. However, the simulated annealing algorithm involves complex optimization calculations and faces the problem of high time costs when processing large-scale data. How to improve the algorithm's computational efficiency while ensuring the accuracy of the identification is also a technical challenge that needs to be solved urgently. Summary of the invention
[0005] The present invention provides an equipment fault identification method, system and storage medium, which can effectively identify and handle rare faults, improve the accuracy and efficiency of fault diagnosis, provide timely and accurate decision support for equipment maintenance, and thus improve the overall equipment operation reliability.
[0006] In a first aspect, in order to solve the above technical problem, the present invention provides a method for identifying a device fault, comprising: The original fault data set in the equipment operation data is obtained, and the principal component analysis method is used to perform dimension reduction processing to obtain the fault feature set after dimension reduction; According to the fault feature set after dimension reduction, the density clustering algorithm is used to classify the fault mode. The fault mode is judged by the difference of feature distribution. If the number of samples in a certain category in the clustering result is lower than the preset early stopping mechanism threshold, it is judged as a rare fault mode, otherwise it is a common fault mode. According to the rare fault mode, a synthetic sample is generated by using SMOTE oversampling technology, and a balanced data set distribution operation is performed through feature correlation analysis to obtain a balanced fault data set; According to the balanced fault data set, an adaptive weight allocation strategy is adopted to calculate the importance weight of each fault mode; The importance weights are integrated into the preset objective function. By optimizing the algorithm selection and batch size setting, a weighted objective function is obtained to optimize the discrimination ability of the simulated annealing algorithm. According to the weighted objective function, the parallel computing framework is used to optimize the calculation process of the simulated annealing algorithm. The number of parallel computing nodes is dynamically adjusted through node load balancing and communication delay optimization. If the calculation time exceeds the preset time threshold, the parallel computing nodes are increased through task scheduling strategy and resource allocation algorithm. The fault identification results are obtained from the optimized simulated annealing algorithm. If there are rare fault modes in the identification results, the early warning mechanism is triggered and maintenance recommendations are generated according to the node fault recovery strategy.
[0007] In a second aspect, the present invention provides a device fault identification system, comprising: The dimension reduction processing module is used to obtain the original fault data set in the equipment operation data, and use the principal component analysis method to perform dimension reduction processing to obtain the fault feature set after dimension reduction; The fault classification module is used to classify the fault mode using the density clustering algorithm according to the fault feature set after dimensionality reduction. If the number of samples in a certain category in the clustering result is lower than the preset early stopping mechanism threshold, it is judged as a rare fault mode, otherwise it is a common fault mode. A data balancing module, used to generate synthetic samples using SMOTE oversampling technology according to the rare fault mode, and perform a balanced data set distribution operation through feature correlation analysis to obtain a balanced fault data set; A weight allocation module is used to calculate the importance weight of each fault mode according to the balanced fault data set using an adaptive weight allocation strategy; Function optimization module, which is used to integrate importance weights into the preset objective function. By optimizing algorithm selection and batch size setting, a weighted objective function is obtained to optimize the discrimination ability of the simulated annealing algorithm. The parallel optimization module is used to optimize the calculation process of the simulated annealing algorithm using a parallel computing framework according to the weighted objective function, dynamically adjust the number of parallel computing nodes through node load balancing and communication delay optimization, and increase parallel computing nodes through task scheduling strategies and resource allocation algorithms if the calculation time exceeds the preset time threshold; The fault identification module is used to obtain the fault identification results from the optimized simulated annealing algorithm. If there are rare fault modes in the identification results, the early warning mechanism is triggered and maintenance suggestions are generated according to the node fault recovery strategy.
[0008] In a third aspect, the present invention further provides an electronic device, comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor implements any one of the above-described device fault identification methods when executing the computer program.
[0009] In a fourth aspect, the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute any one of the above-mentioned device fault identification methods.
[0010] Compared with the prior art, the present invention has the following beneficial effects: The present invention discloses a method for identifying equipment faults. The method first performs dimensionality reduction processing on the original fault data to extract key features; then uses a density clustering algorithm to classify fault modes and identify rare faults; uses oversampling technology to balance the data set for rare faults; then calculates the importance weight of each fault mode through an adaptive weight allocation strategy, and incorporates the weight information into the improved objective function; then uses a parallel computing framework to optimize the simulated annealing algorithm to improve computational efficiency; finally, generates maintenance recommendations based on the discrimination results, and uses real-time data stream processing technology to update the equipment operating status. The present invention can effectively identify and process rare faults, improve the accuracy and efficiency of fault diagnosis, provide timely and accurate decision support for equipment maintenance, and thus improve the overall equipment operation reliability. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 is a schematic flow chart of a method for identifying equipment failures provided by a first embodiment of the present invention; Figure 2 It is a schematic diagram of the structure of the equipment fault identification system provided by the second embodiment of the present invention. DETAILED DESCRIPTION
[0012] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0013] Reference Figure 1The first embodiment of the present invention provides a method for identifying a device fault, comprising the following steps: S11, obtaining an original fault data set in the equipment operation data, performing dimensionality reduction processing using a principal component analysis method, and obtaining a fault feature set after dimensionality reduction; S12, based on the fault feature set after dimensionality reduction, a density clustering algorithm is used to classify the fault mode, and the fault mode is judged by the difference in feature distribution. If the number of samples in a certain category in the clustering result is lower than the preset early stopping mechanism threshold, it is judged as a rare fault mode, otherwise it is a common fault mode; S13, according to the rare fault mode, using SMOTE oversampling technology to generate synthetic samples, and performing a balanced data set distribution operation through feature correlation analysis to obtain a balanced fault data set; S14, calculating the importance weight of each fault mode using an adaptive weight allocation strategy according to the balanced fault data set; S15, integrating the importance weight into the preset objective function, and obtaining the weighted objective function by optimizing the algorithm selection and batch size setting, which is used to optimize the discrimination ability of the simulated annealing algorithm; S16, according to the weighted objective function, a parallel computing framework is used to optimize the calculation process of the simulated annealing algorithm, and the number of parallel computing nodes is dynamically adjusted through node load balancing and communication delay optimization. If the calculation time exceeds a preset time threshold, the parallel computing nodes are increased through a task scheduling strategy and a resource allocation algorithm; S17, obtaining the fault identification result from the optimized simulated annealing algorithm. If there is a rare fault mode in the identification result, the early warning mechanism is triggered, and maintenance suggestions are generated according to the node fault recovery strategy.
[0014] In step S11, the original fault data set in the equipment operation data is obtained, and the principal component analysis method is used to perform dimensionality reduction processing to obtain a fault feature set after dimensionality reduction, including: Obtain the original fault data set based on the equipment operation data; For the original fault data set, the principal component analysis method is used to perform dimensionality reduction processing through feature weight initialization method and regularization term coefficient to obtain the dimensionality reduction processing result; Extract key features of the data set from the dimensionality reduction processing result, and obtain a fault feature set after dimensionality reduction based on the extracted key features; For the fault feature set after dimension reduction, determine whether its feature redundancy is lower than a preset redundancy threshold; If the feature redundancy is lower than the preset redundancy threshold, the fault feature set after dimension reduction is output; If the feature redundancy is higher than the preset redundancy threshold, the process returns to the step of re-performing the dimensionality reduction process by initializing the feature weights and the regularization term coefficients until the feature redundancy meets the requirements.
[0015] Specifically, equipment operation data is a vital source of information in industrial production. By analyzing this data, potential fault risks can be discovered in time. Obtaining the original fault data set is the first step in fault diagnosis and prediction. For example, in a large chemical plant, various sensor data such as temperature, pressure, flow, etc. can be collected, which constitute the original fault data set. Principal component analysis (PCA) is a commonly used dimensionality reduction method that can effectively reduce the complexity of data. When applying PCA, the selection of feature weight initialization and regularization term coefficient is crucial. For example, for the data of a chemical plant, higher weights can be given to temperature and pressure based on the initial feature weights, while lower weights are given to some minor parameters. The regularization term coefficient can be determined by methods such as cross-validation to balance the complexity and generalization ability of the model. After PCA dimensionality reduction processing, the key features of the data set can be extracted. In the example of the chemical plant, it is found that temperature changes, pressure fluctuations, and the concentration of certain chemicals are the most important features. These key features constitute the fault feature set after dimensionality reduction, which greatly simplifies the subsequent analysis work. Feature redundancy is an important indicator for measuring the quality of feature sets. If the feature redundancy is lower than the preset threshold, it means that the dimensionality reduction effect is good, and this feature set can be used directly for model training. For example, if the original data has 100 features, and only 10 features are retained after dimensionality reduction, and these 10 features can explain 95% of the data variance, then the redundancy of this feature set is very low. However, if the feature redundancy is higher than the preset threshold, it is necessary to return to adjust the feature weights and regularization term coefficients. For example, you can try to increase the regularization strength, or adjust the initial weight distribution to give higher weights to some previously neglected features. This process requires multiple iterations until a satisfactory result is obtained. The advantage of this method is that it can effectively process high-dimensional data and extract the most critical information. In industrial production, this can help engineers identify potential problems more quickly and improve the accuracy and efficiency of fault diagnosis.
[0016] For example, in a chemical plant, if temperature anomaly is found to be the most critical feature, then temperature-related parameters can be monitored and analyzed first, so that problems can be discovered and solved more quickly. In addition, this method is highly adaptable. Different industrial scenarios have different feature importances. By adjusting feature weights and regularization parameters, the model can be better adapted to specific application scenarios. For example, in a system dominated by pressure control, pressure-related features need to be given a higher weight. In general, this PCA-based fault feature extraction method, through iterative optimization, can significantly reduce data dimensions while retaining key information, laying a solid foundation for subsequent fault diagnosis and prediction.
[0017] In step S12, based on the fault feature set after dimension reduction, a density clustering algorithm is used to classify the fault modes. By judging the difference in feature distribution, if the number of samples in a certain category in the clustering result is lower than the preset early stopping mechanism threshold, it is determined to be a rare fault mode, otherwise it is a common fault mode, including: For the fault feature set after dimension reduction, cluster analysis is performed on the fault features according to the density connectivity between samples to obtain clustering results of multiple fault modes; Analyze the distribution of the number of samples in each category in the clustering results, and determine whether there is a rare failure mode by comparing it with the preset rare failure threshold; If the number of samples in a certain category is lower than the preset early stopping mechanism threshold, the category is determined to be a rare failure mode; if the number of samples in a certain category is not lower than the preset early stopping mechanism threshold, the category is determined to be a common failure mode.
[0018] In this embodiment, density clustering algorithms such as DBSCAN can effectively identify clusters of different shapes. This feature is particularly important in fault diagnosis. For example, a certain type of bearing fault presents an irregular distribution in the feature space. By setting the neighborhood radius and the minimum number of samples, DBSCAN can divide these density-connected points into a cluster and identify the noise points. The identification of rare faults is a difficult point in fault diagnosis. Suppose that in 1000 samples, a certain fault only occurs 5 times. By setting a suitable threshold, such as the number of samples is less than 1% of the total samples, this type of fault can be marked as a rare fault.
[0019] In step S13, according to the rare fault mode, a synthetic sample is generated by using the SMOTE oversampling technology, and a balanced data set distribution operation is performed through feature correlation analysis to obtain a balanced fault data set.
[0020] Specifically, in the field of fault diagnosis, dealing with rare fault modes is an important challenge. In order to improve the ability to identify rare faults, the SMOTE oversampling technique can be used to balance the data set. For example, suppose there is a fault data set containing 1,000 samples, of which only 20 samples belong to a certain rare fault type. Through the SMOTE algorithm, interpolation can be performed between minority class samples to generate new synthetic samples, increasing the number of rare fault samples to 100, making the data set more balanced. Feature extraction is a key step in fault diagnosis. Taking industrial equipment as an example, multidimensional features such as vibration, temperature, and pressure can be extracted. By calculating the correlation coefficients between these features, a feature correlation matrix can be constructed. Assume that the correlation coefficient of vibration and temperature features is 0.8, while the correlation coefficient of vibration and pressure is only 0.2, which indicates that there is a certain amount of information redundancy between vibration and temperature. In order to reduce feature redundancy, the minimum redundancy maximum correlation criterion can be used for feature selection. In the above example, you can choose to retain the vibration and pressure features and discard the temperature feature. Doing so can not only reduce the computational complexity, but also improve the generalization ability of the model.
[0021] In order to continuously improve the performance of the system, a dynamic update mechanism can also be established. Whenever the system detects a new fault case, especially a rare fault case, the data will be added to the fault knowledge base. For example, if the system successfully predicts and confirms 3 new rare fault cases within a month, the characteristic data of these cases will be used to expand the original data set. Through regular updates, the system's ability to identify various fault modes can be continuously improved, especially for those rare faults that are difficult to capture.
[0022] In step S14, based on the balanced fault data set, an adaptive weight allocation strategy is adopted to calculate the importance weight of each fault mode, including: According to each fault mode in the balanced fault data set, its occurrence frequency in the historical data is counted to obtain a fault mode frequency list; According to the fault mode frequency list, an adaptive weight allocation strategy is adopted to initialize the importance weight of each fault mode through the learning rate adjustment strategy and weight decay factor; For each fault mode, determine whether its occurrence frequency is higher than a preset frequency threshold; if higher than the frequency threshold, multiply the importance weight of the fault mode by a factor greater than 1 to obtain a first adjustment weight; If it is lower than the frequency threshold, the importance weight of the fault mode is multiplied by a factor less than 1 to obtain a second adjustment weight; The first adjusted weight or the second adjusted weight is output as a final importance weight of each failure mode.
[0023] Specifically, in the fault diagnosis system, the next step after balancing the data set is to analyze the fault mode frequency. This step is crucial to improving the accuracy of diagnosis because the frequency of occurrence of different fault modes often varies significantly. For example, on the production line of a factory, bearing wear is a high-frequency fault, while motor burnout is a low-frequency fault. By statistically analyzing historical data, a list of fault mode frequencies can be obtained. Assume that in the past year, bearing wear occurred 200 times, accounting for 40% of the total number of faults; motor burnout only occurred 5 times, accounting for 1%. Based on this frequency list, an adaptive weight allocation strategy is adopted. The core idea of this strategy is to dynamically adjust the model's attention to different faults according to the fault frequency.
[0024] Specifically, a frequency threshold can be set, such as 10%. For fault modes above the threshold, such as bearing wear, their initial weights are multiplied by a factor greater than 1, such as 1.5; while for fault modes below the threshold, such as motor burnout, they are multiplied by a factor less than 1, such as 0.8. The purpose of this is to ensure that the model has a good recognition ability for high-frequency faults while not ignoring low-frequency but serious faults. In the weight adjustment process, the learning rate and weight decay factor play an important role. The learning rate determines the step size of the parameter update at each iteration, while the weight decay factor helps prevent the model from overfitting. For example, a larger learning rate, such as 0.1, can be set for high-frequency faults to converge quickly; while a smaller learning rate, such as 0.01, can be set for low-frequency faults to avoid overfitting a small number of samples. At the same time, for low-frequency faults, the weight decay factor can be increased, such as from 0.001 to 0.005, to enhance the generalization ability of the model. In this way, a more balanced and practical fault diagnosis model can be trained. This model can not only accurately identify common faults, but also remain sensitive to rare but serious faults. In practical applications, this method can significantly improve the accuracy and timeliness of fault warnings, thereby reducing equipment downtime, reducing maintenance costs, and improving production efficiency.
[0025] In step S15, the importance weights are integrated into the preset objective function, and a weighted objective function is obtained by optimizing the algorithm selection and batch size setting to optimize the discrimination ability of the simulated annealing algorithm, including: Construct an initial objective function, introduce the importance weight as a parameter into the initial objective function, and obtain an objective function expression containing weight information; The gradient descent algorithm is used to optimize the initial objective function and obtain the optimal weight parameters; Substitute the optimal weight parameters into the simulated annealing algorithm, and dynamically adjust the acceptance probability and annealing strategy in the simulated annealing process according to the weight information; In the simulated annealing iteration process, the weight parameters and weight update frequency are adaptively adjusted according to the quality of the current solution and the historical optimal solution, and the weights are appropriately scaled and regularized according to the sensitivity of the weight parameters; The optimal hyperparameter combination is selected through cross-validation to obtain the weighted objective function; Among them, the weighted objective function is expressed as:
[0026] In the formula, represents the weighted objective function, Indicates The weight of each sub-goal, Indicates sub-objective function, is the decision variable vector, is the number of sub-goals, is the regularization parameter.
[0027] Specifically, this formula represents the objective function containing weight information. The first term represents the sum of weighted sub-objective functions, and the second term is the L1 regularization term, which is used to control the sparsity of weights. Using optimization algorithms such as gradient descent, the weighted objective function is optimized and solved by setting appropriate batch size and number of iterations to obtain the optimal solution of weight parameters. The optimal weight parameters obtained by the solution are substituted into the simulated annealing algorithm, and the acceptance probability and annealing strategy in the simulated annealing process are dynamically adjusted according to the weight information to improve the discrimination ability of the algorithm. In the simulated annealing iteration process, the weight parameters are adaptively adjusted according to the quality of the current solution and the historical optimal solution to accelerate the weight convergence speed and avoid the algorithm from falling into the local optimum. For the weight update frequency, an adaptive weight update strategy is designed to dynamically adjust the weight update frequency according to the algorithm convergence and problem scale, while ensuring the convergence speed and reducing the computational overhead. In the algorithm iteration process, the changes of weight parameters are monitored in real time, and the weights are appropriately scaled and regularized according to the sensitivity of the weight parameters to improve the generalization ability and robustness of the algorithm.
[0028] In step S16, according to the weighted objective function, the parallel computing framework is used to optimize the calculation process of the simulated annealing algorithm, and the number of parallel computing nodes is dynamically adjusted through node load balancing and communication delay optimization. If the calculation time exceeds the preset time threshold, the parallel computing nodes are increased through the task scheduling strategy and resource allocation algorithm, including: According to the weighted objective function, a parallel computing model of the simulated annealing algorithm is constructed, and a master-slave architecture is used to design a parallel computing framework, in which the master node is responsible for task scheduling and global parameter update, and the slave node is responsible for local computing tasks; Obtain the load status and communication delay data of each parallel computing node, and dynamically adjust the computing task allocation of each node through the load balancing algorithm; During the parallel computing process, the computing time of each node is continuously monitored. If the computing time is found to exceed the preset time threshold, the dynamic node adjustment mechanism is triggered to dynamically increase the number of parallel computing nodes according to the task scheduling strategy and resource allocation algorithm.
[0029] Specifically, constructing a parallel computing model for the simulated annealing algorithm is an important means to improve the efficiency of the algorithm. The master-slave architecture is used to design a parallel computing framework, which can make full use of distributed computing resources. For example, in a large-scale path optimization problem, the master node can be responsible for dividing the entire path network into multiple sub-areas and assigning these sub-areas to slave nodes for local optimization. The master node is also responsible for collecting the optimization results of each slave node and updating the global parameters. In order to achieve efficient parallel computing, factors such as load balancing and communication delay need to be considered. By real-time monitoring of indicators such as CPU usage and memory usage of each node, task allocation can be dynamically adjusted. For example, if it is found that the CPU usage of a slave node continues to exceed 90%, while other nodes are only about 50%, some tasks can be transferred from high-load nodes to low-load nodes. At the same time, by optimizing the communication strategy between nodes, such as using asynchronous communication or batch communication, the impact of communication delay on computing performance can be effectively reduced. In the parallel computing process, the dynamic node adjustment mechanism can effectively deal with computing bottlenecks. Suppose that in a complex financial model optimization task, 10 computing nodes are initially set, but as the computing scale increases, it is found that the average computing time exceeds the preset 5-minute threshold. At this point, the system can automatically trigger the node expansion mechanism and dynamically add 5 computing nodes according to the current load and available resources, thereby reducing the average computing time to an acceptable range. The adaptive annealing temperature adjustment strategy is the key to improving the efficiency of the simulated annealing algorithm. During the optimization process, the annealing temperature can be dynamically adjusted according to the change in the solution quality of multiple consecutive iterations. For example, if the quality of the solution does not improve significantly after 100 consecutive iterations, the annealing temperature can be appropriately increased to increase the probability of the algorithm jumping out of the local optimum; conversely, if the solution quality continues to improve, the annealing temperature can be reduced to accelerate the convergence speed. An efficient parallel random number generation algorithm is crucial to ensure the randomness and parallelism of the simulated annealing algorithm. A hybrid random number generation method based on timestamp and node ID can be used to ensure that the random number sequences generated by each node are both independent of each other and have good statistical properties. This can avoid different nodes generating the same random number sequence, thereby improving the algorithm's exploration ability and the quality of the solution. The incremental global parameter update strategy can effectively reduce the communication overhead between nodes. In a large-scale image recognition task, each slave node can independently process a batch of image data and calculate local gradients. After processing every 1,000 images, the slave node sends the accumulated gradient information to the master node. The master node aggregates the gradient information of all slave nodes, updates the global model parameters, and then broadcasts the updated parameters to all slave nodes. This approach ensures the consistency of the model and reduces the frequency of communication.
[0030] In addition, designing an adaptable fault-tolerant mechanism is crucial to ensure the reliability of parallel computing. Periodic checkpoint technology can be used to save the computing state to persistent storage. If a node fails, the system can quickly restore the computing state from the most recent checkpoint and reallocate the tasks of the failed node to other available nodes. In addition, by setting up a heartbeat mechanism, the master node can detect abnormal nodes in time and take corresponding fault-tolerant measures, such as restarting nodes or reallocating tasks, thereby ensuring the robustness and reliability of the entire parallel computing process.
[0031] In step S17, the fault identification result is obtained from the optimized simulated annealing algorithm. If there is a rare fault mode in the identification result, the early warning mechanism is triggered, and maintenance suggestions are generated according to the node fault recovery strategy, including: The optimized simulated annealing algorithm is used to analyze the equipment operation data and obtain the fault identification results; Interpret the key features in the fault identification results to determine whether there are rare fault modes; If there are rare failure modes in the identification results, the early warning mechanism will be triggered and a warning message will be sent to the system administrator; Generate targeted node failure recovery plans based on early warning information and pre-established node failure recovery strategy knowledge base; Send the generated node fault recovery plan to the faulty node to guide the faulty node to perform self-repair; If the node fault recovery plan fails to execute, a node maintenance recommendation is generated through a decision tree algorithm based on the node operation log and fault information.
[0032] Specifically, the optimized simulated annealing algorithm can be used to analyze system operation data and obtain fault identification results. The algorithm simulates the metal cooling process to find the global optimal solution in the search space. For example, for a large data center, the algorithm can be used to analyze the performance indicators of the server cluster, such as CPU usage, memory usage, network latency, etc., to identify potential faulty nodes. Feature interpretability analysis technology helps to explain the key features in the fault identification results and determine whether there are rare fault modes. Taking the data center as an example, if the analysis results show that the CPU temperature of a server is abnormally high, while other indicators are normal, this indicates that the cooling system is faulty. Through interpretability analysis, it can be determined that temperature anomaly is the key factor causing the failure. When a rare fault mode is found, the early warning mechanism will send an early warning message to the system administrator. In the data center scenario, the early warning message contains key information such as the ID of the faulty server, abnormal indicators, and fault type. In this way, the administrator can quickly locate the problem and take corresponding measures. The case-based reasoning algorithm can be used to generate targeted node fault recovery solutions. The algorithm generates solutions by matching similar cases based on a pre-established node fault recovery strategy knowledge base. For example, for a server with abnormal CPU temperature, the system will recommend specific steps such as checking the cooling fan, cleaning dust, and replacing thermal paste. The generated node fault recovery plan will be sent to the faulty node to guide it to perform self-repair.
[0033] In a data center environment, this involves operations such as remotely executing scripts, adjusting system parameters, or restarting specific services. Automated repair not only improves efficiency, but also reduces human errors. If automatic repair fails, the decision tree algorithm can be used to generate node maintenance recommendations. The algorithm builds a logical decision process by analyzing node operation logs and fault information. For example, for persistent CPU temperature anomalies, the decision tree will recommend deeper solutions such as checking hardware integrity, replacing the CPU, or considering server location adjustments. Finally, node maintenance recommendations are sent to operations and maintenance personnel to guide them in manual maintenance. These recommendations include detailed troubleshooting steps, lists of required tools and parts, safety precautions, etc. By providing a structured maintenance guide, it can ensure that even complex faults can be effectively handled, thereby ensuring the normal operation of the system.
[0034] In summary, the present invention discloses a method for identifying equipment faults. The method first performs dimensionality reduction processing on the original fault data to extract key features; then uses a density clustering algorithm to classify fault modes and identify rare faults; uses oversampling technology to balance the data set for rare faults; then calculates the importance weight of each fault mode through an adaptive weight allocation strategy, and incorporates the weight information into the improved objective function; then uses a parallel computing framework to optimize the simulated annealing algorithm to improve computational efficiency; finally, generates maintenance recommendations based on the discrimination results, and uses real-time data stream processing technology to update the equipment operating status. The present invention can effectively identify and process rare faults, improve the accuracy and efficiency of fault diagnosis, provide timely and accurate decision support for equipment maintenance, and thus improve the overall equipment operation reliability.
[0035] Reference Figure 2 A second embodiment of the present invention provides a device fault identification system, comprising: The dimension reduction processing module is used to obtain the original fault data set in the equipment operation data, and use the principal component analysis method to perform dimension reduction processing to obtain the fault feature set after dimension reduction; The fault classification module is used to classify the fault mode using the density clustering algorithm according to the fault feature set after dimensionality reduction. If the number of samples in a certain category in the clustering result is lower than the preset early stopping mechanism threshold, it is judged as a rare fault mode, otherwise it is a common fault mode. A data balancing module, used to generate synthetic samples using SMOTE oversampling technology according to the rare fault mode, and perform a balanced data set distribution operation through feature correlation analysis to obtain a balanced fault data set; A weight allocation module is used to calculate the importance weight of each fault mode according to the balanced fault data set using an adaptive weight allocation strategy; Function optimization module, which is used to integrate importance weights into the preset objective function. By optimizing algorithm selection and batch size setting, a weighted objective function is obtained to optimize the discrimination ability of the simulated annealing algorithm. The parallel optimization module is used to optimize the calculation process of the simulated annealing algorithm using a parallel computing framework according to the weighted objective function, dynamically adjust the number of parallel computing nodes through node load balancing and communication delay optimization, and increase parallel computing nodes through task scheduling strategies and resource allocation algorithms if the calculation time exceeds the preset time threshold; The fault identification module is used to obtain the fault identification results from the optimized simulated annealing algorithm. If there are rare fault modes in the identification results, the early warning mechanism is triggered and maintenance suggestions are generated according to the node fault recovery strategy.
[0036] It should be noted that an equipment fault identification system provided in an embodiment of the present invention is used to execute all process steps of an equipment fault identification method in the above embodiment, and the working principles and beneficial effects of the two correspond one to one, so they will not be repeated here.
[0037] The embodiment of the present invention further provides an electronic device. The electronic device includes: a processor, a memory, and a computer program stored in the memory and executable on the processor, such as a device fault identification program. When the processor executes the computer program, the steps in the above-mentioned device fault identification method embodiments are implemented, such as Figure 1 Alternatively, when the processor executes the computer program, the functions of the modules / units in the above-mentioned system embodiments are realized, such as the fault identification module.
[0038] Exemplarily, the computer program may be divided into one or more modules / units, which are stored in the memory and executed by the processor to implement the present invention. The one or more modules / units may be a series of computer program instruction segments capable of implementing specific functions, which are used to describe the execution process of the computer program in the electronic device.
[0039] The electronic device may be a computing device such as a desktop computer, a notebook, a PDA, and a smart tablet. The electronic device may include, but is not limited to, a processor and a memory. Those skilled in the art will appreciate that the above components are merely examples of electronic devices and do not constitute a limitation on the electronic device. The electronic device may include more or fewer components than the above components, or may combine certain components, or different components. For example, the electronic device may also include input and output devices, network access devices, buses, etc.
[0040] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the electronic device, and uses various interfaces and lines to connect various parts of the entire electronic device.
[0041] The memory can be used to store the computer program and / or module, and the processor realizes various functions of the electronic device by running or executing the computer program and / or module stored in the memory, and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area can store data created according to the use of the mobile phone (such as audio data, a phone book, etc.), etc. In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0042] Wherein, if the module / unit integrated in the electronic device is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the present invention implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Wherein, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or system that can carry the computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.
[0043] It should be noted that the system embodiment described above is merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, in the drawings of the system embodiment provided by the present invention, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines. A person of ordinary skill in the art may understand and implement it without paying any creative effort.
[0044] The specific embodiments described above further illustrate the purpose, technical solutions and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. It is particularly pointed out that for those skilled in the art, any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for identifying equipment failure, characterized in that: The method comprises: The original fault data set in the equipment operation data is obtained, and the principal component analysis method is used to perform dimension reduction processing to obtain the fault feature set after dimension reduction; According to the fault feature set after dimension reduction, the density clustering algorithm is used to classify the fault mode. The fault mode is judged by the difference of feature distribution. If the number of samples in a certain category in the clustering result is lower than the preset early stopping mechanism threshold, it is judged as a rare fault mode, otherwise it is a common fault mode. According to the rare fault mode, a synthetic sample is generated by using SMOTE oversampling technology, and a balanced data set distribution operation is performed through feature correlation analysis to obtain a balanced fault data set; According to the balanced fault data set, an adaptive weight allocation strategy is adopted to calculate the importance weight of each fault mode; The importance weights are integrated into the preset objective function. By optimizing the algorithm selection and batch size setting, a weighted objective function is obtained to optimize the discrimination ability of the simulated annealing algorithm. According to the weighted objective function, the parallel computing framework is used to optimize the calculation process of the simulated annealing algorithm. The number of parallel computing nodes is dynamically adjusted through node load balancing and communication delay optimization. If the calculation time exceeds the preset time threshold, the parallel computing nodes are increased through task scheduling strategy and resource allocation algorithm. The fault identification results are obtained from the optimized simulated annealing algorithm. If there are rare fault modes in the identification results, the early warning mechanism is triggered and maintenance recommendations are generated according to the node fault recovery strategy.
2. The method according to claim 1, characterized in that The original fault data set in the equipment operation data is obtained, and a principal component analysis method is used to perform dimensionality reduction processing to obtain a fault feature set after dimensionality reduction, including: Obtain the original fault data set based on the equipment operation data; For the original fault data set, the principal component analysis method is used to perform dimensionality reduction processing through feature weight initialization method and regularization term coefficient to obtain the dimensionality reduction processing result; Extract key features of the data set from the dimensionality reduction processing result, and obtain a fault feature set after dimensionality reduction based on the extracted key features; For the fault feature set after dimension reduction, determine whether its feature redundancy is lower than a preset redundancy threshold; If the feature redundancy is lower than the preset redundancy threshold, the fault feature set after dimension reduction is output; If the feature redundancy is higher than the preset redundancy threshold, the process returns to the step of re-performing the dimensionality reduction process by initializing the feature weights and the regularization term coefficients until the feature redundancy meets the requirements.
3. The method according to claim 1, characterized in that The fault mode is classified by using a density clustering algorithm based on the fault feature set after dimensionality reduction. If the number of samples in a certain category in the clustering result is lower than the preset early stopping mechanism threshold, it is determined to be a rare fault mode, otherwise it is a common fault mode, including: For the fault feature set after dimension reduction, cluster analysis is performed on the fault features according to the density connectivity between samples to obtain clustering results of multiple fault modes; Analyze the distribution of the number of samples in each category in the clustering results, and determine whether there is a rare failure mode by comparing it with the preset rare failure threshold; If the number of samples in a certain category is lower than the preset early stopping mechanism threshold, the category is determined to be a rare failure mode; if the number of samples in a certain category is not lower than the preset early stopping mechanism threshold, the category is determined to be a common failure mode.
4. The method according to claim 1, characterized in that: The method uses an adaptive weight allocation strategy based on the balanced fault data set to calculate the importance weight of each fault mode, including: According to each fault mode in the balanced fault data set, its occurrence frequency in the historical data is counted to obtain a fault mode frequency list; According to the fault mode frequency list, an adaptive weight allocation strategy is adopted to initialize the importance weight of each fault mode through the learning rate adjustment strategy and weight decay factor; For each fault mode, determine whether its occurrence frequency is higher than a preset frequency threshold; if higher than the frequency threshold, multiply the importance weight of the fault mode by a factor greater than 1 to obtain a first adjustment weight; If it is lower than the frequency threshold, the importance weight of the fault mode is multiplied by a factor less than 1 to obtain a second adjustment weight; The first adjusted weight or the second adjusted weight is output as a final importance weight of each failure mode.
5. The method according to claim 1, characterized in that: The importance weight is integrated into the preset objective function, and a weighted objective function is obtained by optimizing the algorithm selection and batch size setting, which is used to optimize the discrimination ability of the simulated annealing algorithm, including: Construct an initial objective function, introduce the importance weight as a parameter into the initial objective function, and obtain an objective function expression containing weight information; The gradient descent algorithm is used to optimize the initial objective function and obtain the optimal weight parameters; Substitute the optimal weight parameters into the simulated annealing algorithm, and dynamically adjust the acceptance probability and annealing strategy in the simulated annealing process according to the weight information; In the simulated annealing iteration process, the weight parameters and weight update frequency are adaptively adjusted according to the quality of the current solution and the historical optimal solution, and the weights are appropriately scaled and regularized according to the sensitivity of the weight parameters; The optimal hyperparameter combination is selected through cross-validation to obtain the weighted objective function; Among them, the weighted objective function is expressed as: : In the formula, represents the weighted objective function, Indicates The weight of each sub-goal, Indicates sub-objective function, is the decision variable vector, is the number of sub-goals, is the regularization parameter.
6. The method according to claim 1, characterized in that The method uses a parallel computing framework to optimize the calculation process of the simulated annealing algorithm according to the weighted objective function, dynamically adjusts the number of parallel computing nodes through node load balancing and communication delay optimization, and increases parallel computing nodes through task scheduling strategies and resource allocation algorithms if the calculation time exceeds a preset time threshold, including: According to the weighted objective function, a parallel computing model of the simulated annealing algorithm is constructed, and a master-slave architecture is used to design a parallel computing framework, in which the master node is responsible for task scheduling and global parameter update, and the slave node is responsible for local computing tasks; Obtain the load status and communication delay data of each parallel computing node, and dynamically adjust the computing task allocation of each node through the load balancing algorithm; During the parallel computing process, the computing time of each node is continuously monitored. If the computing time is found to exceed the preset time threshold, the dynamic node adjustment mechanism is triggered to dynamically increase the number of parallel computing nodes according to the task scheduling strategy and resource allocation algorithm.
7. The method according to claim 1, characterized in that The fault identification result is obtained from the optimized simulated annealing algorithm. If there is a rare fault mode in the identification result, the early warning mechanism is triggered, and maintenance suggestions are generated according to the node fault recovery strategy, including: The optimized simulated annealing algorithm is used to analyze the equipment operation data and obtain the fault identification results; Interpret the key features in the fault identification results to determine whether there are rare fault modes; If there are rare failure modes in the identification results, the early warning mechanism will be triggered and a warning message will be sent to the system administrator; Generate targeted node failure recovery plans based on early warning information and pre-established node failure recovery strategy knowledge base; Send the generated node fault recovery plan to the faulty node to guide the faulty node to perform self-repair; If the node fault recovery plan fails to execute, a node maintenance recommendation is generated through a decision tree algorithm based on the node operation log and fault information.
8. A device fault identification system, characterized in that: include: The dimension reduction processing module is used to obtain the original fault data set in the equipment operation data, and use the principal component analysis method to perform dimension reduction processing to obtain the fault feature set after dimension reduction; The fault classification module is used to classify the fault mode using the density clustering algorithm according to the fault feature set after dimensionality reduction. If the number of samples in a certain category in the clustering result is lower than the preset early stopping mechanism threshold, it is judged as a rare fault mode, otherwise it is a common fault mode. A data balancing module, used to generate synthetic samples using SMOTE oversampling technology according to the rare fault mode, and perform a balanced data set distribution operation through feature correlation analysis to obtain a balanced fault data set; A weight allocation module is used to calculate the importance weight of each fault mode according to the balanced fault data set using an adaptive weight allocation strategy; Function optimization module, which is used to integrate importance weights into the preset objective function. By optimizing algorithm selection and batch size setting, a weighted objective function is obtained to optimize the discrimination ability of the simulated annealing algorithm. The parallel optimization module is used to optimize the calculation process of the simulated annealing algorithm using a parallel computing framework according to the weighted objective function, dynamically adjust the number of parallel computing nodes through node load balancing and communication delay optimization, and increase parallel computing nodes through task scheduling strategies and resource allocation algorithms if the calculation time exceeds the preset time threshold; The fault identification module is used to obtain the fault identification results from the optimized simulated annealing algorithm. If there are rare fault modes in the identification results, the early warning mechanism is triggered and maintenance suggestions are generated according to the node fault recovery strategy.
9. An electronic device, characterized in that: The method comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor implements the device fault identification method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored computer program, wherein when the computer program is executed, the device where the computer-readable storage medium is located is controlled to execute the device fault identification method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Fault diagnosis method of proton exchange membrane fuel cell for tramcars
CN110579709A
RBF fault diagnosis method and system based on PCA data processing, terminal and computer storage medium
CN111191725A
Rotational molding machine remote maintenance device and method with automatic fault diagnosis and repair functions
CN118395161A
Forward proxy tunnel monitoring method and system and storage medium
CN119182573A
Interactive manual fault diagnosis system based on artificial intelligence
CN119201527A
Cited By
Fault feature extraction method, electronic equipment and storage medium
CN120336990A
Intelligent equipment health management method and system based on AI
CN122086335A
An ai-based device health management method and system
CN122086335B