Hardware equipment fault self-healing processing method
By dividing the fault categories of hardware equipment into first-level and second-level faults, and using the fault case library and feature vector for automated processing, the long response time problem caused by manual intervention in the existing technology is solved, and the self-healing treatment of hardware equipment failures is realized, and the stability and service quality of the system are improved.
Patent Information
- Application Number
- CN202510044394.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-11
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-01-11
AI Technical Summary
In the face of sudden emergency failures, the existing technology relies on manual intervention to cause a long response time, which is unable to prevent potential system crash risks in time, affecting system performance and reliability.
By dividing the fault categories of hardware equipment into first-level faults and second-level faults, using the fault case library and feature vectors for precise positioning and automated processing, the fault self-healing is achieved.
It realizes automation of the entire process of faults from monitoring and judgment to disposal, significantly improving the stability and service quality of the system, and reducing the time and operation and maintenance costs of human intervention.
Smart Images

Figure CN120066832A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of hardware device fault self-healing, and specifically provides a method for processing hardware device fault self-healing. Background Art
[0002] Current situation of related technologies in this field: With the acceleration of informatization construction, large data centers have an increasing demand for the stability and response speed of IT infrastructure. Although existing operation and maintenance mechanisms can monitor device status in real time, in the face of sudden emergency faults, service interruptions often occur due to delayed manual intervention, affecting user experience and business continuity. Solutions of the prior art corresponding to this technical solution: Traditional IT operation and maintenance strategies mainly include setting threshold alarms and arranging personnel on duty. Once an abnormal device is detected, a warning is immediately issued, and technicians are waiting to arrive at the scene for troubleshooting and taking measures. Defects of the prior art: The current fault handling highly depends on manual decision-making and intervention, with a long response time in case of emergencies at night or on holidays, and the potential risk of system collapse cannot be timely prevented, thus affecting the performance and reliability of the entire system.
[0003] Therefore, this solution proposes a method for processing hardware device fault self-healing, which solves the problems raised in the background art. Summary of the Invention
[0004] The present invention provides a method for processing hardware device fault self-healing, which helps to solve the problems mentioned in the above background art.
[0005] The present invention provides the following technical solution: A method for processing hardware device fault self-healing, including: Classifying the fault categories of hardware devices into primary faults and secondary faults; By classifying faults into primary faults and secondary faults, it helps to accurately locate the severity of problems and ensure that key problems are processed first.
[0006] When a hardware device fails, determine whether the type of the fault is a primary fault; If the type of the fault is a primary fault, extract the feature vector of the fault , where represents the keyword frequency of device logs and performance indicators; Set up a fault case library , the fault case library contains all historical fault cases, and each fault case is represented as , where is the fault feature vector, is the fault solution, is the solution result, It is the timestamp of the fault case; by introducing a fault case library and matching the feature vector with historical data, it is ensured that the solution has data support and the accuracy is improved.
[0007] According to the feature vector X and the fault case library C, execute the classification model prediction strategy to predict the probability that the fault solution of each fault case can successfully solve the current fault; Obtain the probabilities of all fault solutions, calculate the mean value of the probabilities, and record the result as the average probability; Obtain all types of fault solutions with probabilities greater than the average probability, and record them as the candidate type set; According to the feature vector X and the candidate type set, execute the sorting model optimization strategy to optimize the priority of each fault solution in the candidate type set; The probability screening and sorting mechanism of the candidate solution set avoids blind attempts and reduces the risk of processing failure.
[0008] Obtain the fault solution with the highest priority, and record it as the target solution ; Apply the target solution Repair the current fault, and record the feedback result as ; According to the feedback result, dynamically optimize the model for recommending fault solutions; Judge whether the device with a fault has been successfully repaired. If it is successfully repaired, the hardware device that has repaired a first-level fault is determined to have a second-level fault; When the hardware device with a first-level fault is repaired, a maintenance plan is dynamically generated according to multi-dimensional factors; If the repair fails, each element in the candidate type set excluding the target solution is used to repair the hardware device in turn.
[0009] Preferably, the step of executing the classification model prediction strategy according to the feature vector X and the fault case library C to predict the probability that the fault solution of each fault case can successfully solve the current fault includes: Obtain all different fault solutions in the fault case library to form a solution category set; For any one fault solution in the solution category set, set the score Z of the fault solution; Obtain the feature vector ; Calculate , where W is the weight matrix, is the parameter of the model, and b is the bias term; Execute the following formula to convert the score of the fault solution into a probability, specifically: Obtain the number of elements in the solution category set, and record it as the total number of categories K; Number the fault solutions in the solution category set from 1 to K; Calculate , where j = 1, 2...K, is the probability of the fault solution j; Calculate the cross-entropy loss: , where n is the number of samples, K is the number of categories, is the one-hot encoding of the true label, is the probability; Record the result of calculating the cross-entropy loss as the loss value; Set a loss threshold, which is used to judge the accuracy of the prediction result; When the loss value ≤ the loss threshold, end the classification model prediction strategy; When the loss value > the loss threshold, update the parameters W and b using the gradient descent method, specifically: Calculate , as the new W; Calculate , as the new b, where, is the learning rate; Take the new W and b as parameters and execute the classification model prediction strategy again.
[0010] Preferably, according to the feature vector X and the candidate category set, execute the sorting model optimization strategy to optimize the priority of each fault solution in the candidate category set, including: Obtain the number of all elements in the candidate category set, denoted as F; Sort each fault solution in the candidate category set from 1 to F; For any one fault solution in the candidate category set, set the score sco of the fault solution; Obtain the feature vector ; +b, where, is used to extract the features of the feature vector X and the candidate solution , W is the parameter of the model, and b is the bias term; Obtain the scores of each fault solution in the candidate category set, and sort the fault solutions according to the scores; Arbitrarily obtain two fault solutions in the candidate category set and denote them as and , where has a higher score than ; Calculate , and record the result as the sorting value; Compare the sorted value with 1: If the sorted value > 1, stop executing the sorting model optimization strategy; If the sorted value ≤ 1, update the parameters W and b using the gradient descent method, specifically: Calculate as the new W; Calculate as the new b, where is the learning rate; Use the new W and b as parameters to execute the sorting model optimization strategy again.
[0011] By optimizing the candidate solution priorities using the sorting model, the most likely problem-solving solution can be quickly located, improving efficiency. Dynamically adjusting the parameters W and 𝑏 makes the model more adaptable to new problems and ensures the accuracy of fault handling decisions.
[0012] Preferably, the model for dynamically optimizing the recommended fault solution according to the feedback result includes: The feedback result = where indicates that the fault is resolved, -1 indicates that the fault is not resolved, is the maximum allowable processing time; Define the policy indicating the probability distribution of executing the fault solution when the fault feature is where 0 ≤ u ≤ n; Set the discount factor γ, 0 ≤ γ ≤ 1, for weighing immediate rewards and long-term rewards; Calculate the cumulative reward ; Set the reward threshold; When the cumulative reward ≤ the reward threshold, optimize θ using the policy gradient method: = .
[0013] Optimize the model through feedback learning to avoid repeatedly executing inefficient or incorrect solution strategies. The rapid repair and self-healing of faults reduce equipment downtime and maximize equipment availability. Feed back the actual fault solution result (success or failure) to the model to continuously evolve the classification and sorting strategies. Introduce a cumulative reward mechanism to balance short-term and long-term effects and improve the intelligence level of decision-making.
[0014] Preferably, when repairing the hardware device with a level 1 fault, a maintenance plan is dynamically generated according to multi-dimensional factors, including: Establish an equipment status model: ∙ ∙ ∙ ; wherein, is the health status of the hardware device at time t, 0 indicates a fault, and 1 indicates normal operation; is the initial health status; the workload intensity of the hardware device at time t; the cumulative running time of the hardware device at time t; is the intensity of the maintenance measure for the hardware device at time t; , , are the weight factors of the workload, service life, and maintenance measures on the health status, respectively.
[0015] Preferably, after repairing the hardware device with a primary fault, a maintenance plan is dynamically generated according to multi-dimensional factors, including: generating a priority score for the next maintenance plan for each hardware device: + + , and the result is recorded as the maintenance priority score; wherein, , , are weight parameters; , , are the workload, service life, and health status of each hardware device, respectively; set a maintenance threshold; obtain the hardware device with a maintenance priority score greater than the maintenance threshold, and perform maintenance on this hardware device preferentially.
[0016] By dynamically generating a maintenance plan, the service life of the device is extended, and the replacement cost caused by faults is reduced. According to the workload, service life, and health status of the device, a maintenance plan is dynamically generated to accurately meet the maintenance needs of the device. Using the health status model and priority scoring, the maintenance strategy of high-priority devices is clarified, and potential risks are reduced. After repairing the primary fault, preventive maintenance measures are formulated through multi-dimensional factor analysis to reduce the probability of secondary faults. Key devices are preferentially maintained to reduce the risk of large-scale device damage caused by the fault chain reaction.
[0017] The present invention has the following beneficial effects: 1. The hardware device fault self-healing processing method realizes the full automation of the entire process from fault monitoring, determination to disposal because of adopting an intelligent emergency fault diagnosis and processing process, greatly reducing the time cost required for human intervention, and thus significantly improving the overall stability and service quality of the system.
[0018] 2. The hardware device fault self-healing processing method can accurately screen the most suitable emergency measures for the current situation because of adding intelligent decision support based on historical data, greatly reducing the trial-and-error cost and improving the first repair success rate.
[0019] 3. The hardware device fault self-healing processing method combines a regular preventive maintenance mechanism to form a forward-looking operation and maintenance management mode, effectively avoiding potential crises that may occur during long-term operation, ensuring that the device is in good condition for a long time, and indirectly saving a large amount of maintenance expenses.
[0020] 4. The hardware device fault self-healing processing method divides the fault handling process into precise steps through classification model prediction and sorting optimization to ensure that first-level faults are resolved first and avoid wasting resources due to improper handling priorities. The classification model quickly matches historical cases based on feature vectors to screen out possible solutions; the sorting model further optimizes the priorities of the solutions so that the highest-priority solutions can be quickly applied. Gradient descent dynamically adjusts the model parameters to ensure real-time adaptation to the actual operating conditions. These methods effectively reduce the fault handling time, especially in complex systems, and can significantly improve the overall operation and maintenance efficiency.
[0021] 5. The hardware device fault self-healing processing method, the establishment of a fault case library and its matching mechanism with feature vectors, bases fault handling on historical data and scientific models. By calculating the success probability of the solutions, the system only selects candidate solutions with a probability greater than the average probability to reduce the interference of inefficient or even incorrect solutions. Sorting optimization further ensures the priority use of the best solutions through feature extraction and scoring of candidate solutions. The introduction of a feedback mechanism ensures that the model is continuously optimized during use, and its adaptive ability enables it to accurately handle new problems, thus greatly improving the accuracy and reliability of the solutions.
[0022] 6. For the traditional device fault handling, it usually relies on the experience and skills of professional personnel and requires a lot of time and energy for fault troubleshooting and repair. This solution greatly reduces the workload of the operation and maintenance team through automatic classification, sorting and prediction, especially providing intelligent assistance when dealing with complex problems. In addition, the recommended solutions optimized by the system reduce the technical threshold, and even maintenance personnel with less experience can quickly find an efficient solution path, reducing the dependence on professional knowledge and improving the overall work efficiency of the team. Description of the Drawings
[0023] Figure 1 This is a schematic diagram of the process of the present invention.
[0024] Figure 2 This is a schematic diagram of the method of the present invention.
[0025] Figure 3 This is a schematic diagram of the prediction strategy process of the classification model of the present invention.
[0026] Figure 4 This is a schematic diagram of the optimization strategy process of the sorting model of the present invention. Detailed implementation manners
[0027] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0028] Embodiment 1. Refer to Figure 2 , a method for self-healing processing of hardware device failures, including: Classify the failure categories of the hardware device into primary failures and secondary failures; When a hardware device fails, determine whether the type of the failure is a primary failure; If the type of the failure is a primary failure, extract the feature vector of the failure , where represents the keyword frequency of the device log and the performance index; Set up a failure case library , the failure case library contains all the failure cases that have occurred in history, and each failure case is represented as , where is the failure feature vector, is the failure solution, is the solution result, is the timestamp of the failure case; By introducing the failure case library and feature vector analysis, the system realizes the structured storage and processing of failure information. Each failure has detailed classification, solution, result feedback and timestamp, constituting a traceable historical record. This transparent management method makes the failure handling process clearer and is convenient for relevant personnel to review and analyze. In addition, through standardized processing processes such as classification prediction and sorting optimization, the inconsistency of processing quality caused by differences in human experience is avoided, laying a foundation for realizing the standardized management of equipment operation and maintenance.
[0029] According to the feature vector X and the fault case library C, execute the classification model prediction strategy to predict the probability that the fault solution of each fault case can successfully solve the current fault; By predicting the success rate of the fault solution through the classification model, quickly screen out the potentially effective solutions and reduce the trial-and-error time.
[0030] Obtain the probabilities of all fault solutions, calculate the mean value of the probabilities, and record the result as the average probability; Obtain all types of fault solutions with probabilities greater than the average probability, and record them as the candidate type set; According to the feature vector X and the candidate type set, execute the sorting model optimization strategy to optimize the priority of each fault solution in the candidate type set; Obtain the fault solution with the highest priority and record it as the target solution ; Apply the target solution Repair the current fault, and record the feedback result as ; According to the feedback result, dynamically optimize the model for recommending fault solutions; Judge whether the faulty device has been successfully repaired. If it is successfully repaired, then identify the hardware device that has repaired the primary fault as having a secondary fault; When the hardware device with the primary fault is repaired, generate a maintenance plan dynamically according to multi-dimensional factors; If the repair fails, then use each element in the candidate type set excluding the target solution to repair the hardware device in turn.
[0031] In this embodiment, with reference to Figure 1 , the key technologies of the present invention are shown in detail.
[0032] The implementation of this technical solution is deeply integrated with the Internet of Things (IoT) and big data analysis technologies in modern industry, providing a technical foundation for the digital transformation of enterprises. By collecting and analyzing device data in real time, the system realizes the full-process intelligent management from fault handling to preventive maintenance. Combining historical data analysis with machine learning, the system can gradually achieve a higher level of adaptability and automation, providing strong support for enterprises to introduce an artificial intelligence-driven management model, thereby promoting enterprises to move towards a new stage of intelligent upgrading.
[0033] The execution of the classification model prediction strategy according to the feature vector X and the fault case library C to predict the probability that the fault solution of each fault case can successfully solve the current fault includes: Obtain all different fault solutions in the fault case library to form a solution category set; For any one fault solution in the solution category set, set the score Z of the fault solution; Obtain the feature vector ; Calculate , where W is the weight matrix, is the parameter of the model, and b is the bias term; Execute the following formula to convert the score of the fault solution into a probability, specifically: Obtain the number of elements in the solution category set, denoted as the total number of categories K; Number the fault solutions in the solution category set from 1 to K; Calculate , where j = 1, 2...K, is the probability of the fault solution j; Calculate the cross-entropy loss: , where n is the number of samples, K is the number of categories, is the one-hot encoding of the true label, is the probability; Record the result of calculating the cross-entropy loss as the loss value; Set a loss threshold, which is used to judge the accuracy of the prediction result; When the loss value ≤ the loss threshold, end the classification model prediction strategy; When the loss value > the loss threshold, update the parameters W and b using the gradient descent method, specifically: Calculate , as the new W; Calculate , as the new b, where, is the learning rate; Use the new W and b as parameters to execute the classification model prediction strategy again.
[0034] In this embodiment, referring to Figure 3 , it is the flow of the classification model prediction strategy.
[0035] The sorting model optimization strategy is executed according to the feature vector X and the candidate category set to optimize the priority of each fault solution in the candidate category set, including: Obtain the number of all elements in the candidate category set, denoted as F; Sort each fault solution in the candidate category set from 1 to F; For any one fault solution in the candidate category set, set the score sco of the fault solution; Obtain the feature vector ; +b, where, is used to extract the feature vector X and the candidate solution The feature, where W is the parameter of the model and b is the bias term; Obtain the scores of each fault solution in the candidate solution set, and sort the fault solutions according to the scores; Arbitrarily obtain two fault solutions in the candidate solution set and denote them as and , where has a higher score than ; Calculate , and denote the result as the sorting value; Compare the size relationship between the sorting value and 1: If the sorting value > 1, stop executing the sorting model optimization strategy; If the sorting value ≤ 1, update the parameters W and b using the gradient descent method. Specifically: Calculate , as the new W; Calculate , as the new b, where is the learning rate; Use the new W and b as parameters and execute the sorting model optimization strategy again.
[0036] In this embodiment, referring to Figure 4 , it is the process of the sorting model optimization strategy.
[0037] Through the intelligent classification and sorting model, this solution avoids the trial-and-error cost in traditional fault handling methods. Precise prediction and optimization reduce the time and resource consumption of repeatedly trying wrong solutions, and reduce the operation and maintenance cost. In addition, quick repair shortens the equipment downtime and improves the equipment utilization rate, thus reducing the indirect economic losses caused by production interruption. The system also reduces the equipment replacement or major repair cost caused by faults through preventive maintenance, and fundamentally realizes the optimization of costs and the efficient utilization of resources.
[0038] The model for dynamically optimizing and recommending fault solutions according to the feedback results includes: The feedback result = , where represents that the fault is solved, -1 represents that the fault is not solved, is the maximum allowable processing time; Define the policy , indicating the probability distribution of executing the fault solution when the fault feature is , where 0 ≤ u ≤ n; Set the discount factor γ, 0 ≤ γ ≤ 1, which is used to balance the immediate reward and the long-term reward; Calculate the cumulative reward ; Set the reward threshold; When the cumulative reward ≤ the reward threshold, optimize θ using the policy gradient method: = 。
[0039] The system continuously optimizes the classification and ranking models through a feedback learning mechanism, achieving real-time adaptive adjustment. Each repair result during the fault resolution process is incorporated into the system to guide the update of subsequent fault handling strategies. The cumulative reward mechanism further balances short-term and long-term benefits, enabling the system to not only handle current faults but also provide support for long-term optimization. The model dynamically adjusts parameters W and 𝑏 to adapt to changes in the device operating environment, thus ensuring that effective solutions can still be provided when new problems arise, significantly enhancing the intelligence level of the system.
[0040] When the hardware device with a first-level fault is repaired, a maintenance plan is dynamically generated according to multi-dimensional factors, including: Establish a device status model: ∙ ∙ ∙ ; Among them, is the health status of the hardware device at time t, 0 represents a fault, and 1 represents normal operation; is the initial health status; The workload intensity of the hardware device at time t; The cumulative running time of the hardware device at time t; is the intensity of the maintenance measure for the hardware device at time t; , , are the weight factors of the workload, service life, and maintenance measures on the health status, respectively.
[0041] When the hardware device with a first-level fault is repaired, a maintenance plan is dynamically generated according to multi-dimensional factors, including: Generate a priority score for the next maintenance plan for each hardware device: + + , and the result is recorded as the maintenance priority score; Among them, , , are the weight parameters; , , are respectively the workload, service life, and health status of each hardware device; Set the maintenance threshold; Obtain the hardware devices whose maintenance priority scores are greater than the maintenance threshold, and perform maintenance on these hardware devices preferentially.
[0042] After the system repairs a fault, it automatically generates a personalized maintenance plan, and uses multi-dimensional factors such as the workload, service life, and health status of the device to accurately predict and formulate a maintenance plan. By means of priority scoring, the maintenance strategy for high-risk devices is clarified to ensure that resources are concentrated on key issues. Preventive maintenance measures effectively reduce the device failure rate, lower the risk of sudden shutdowns, and at the same time extend the service life of the device. This method not only protects the health status of the device, but also realizes the scientific management of the entire life cycle of the device.
[0043] By optimizing the health status and operating efficiency of the device, the system reduces the energy waste caused by faults or performance degradation. For example, devices operating efficiently usually consume less energy, and the decrease in the failure frequency also means a significant reduction in unnecessary resource consumption. In addition, by replacing key components in a timely manner through a personalized maintenance plan, it is possible to effectively avoid the release of potentially harmful emissions during device failures. This green operation and maintenance method not only improves the economic benefits of the device, but also meets the enterprise's environmental protection and sustainable development goals.
[0044] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.
[0045] The above are only the preferred embodiments of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A method for self-healing hardware device failure, characterized in that: include: The fault categories of hardware devices are divided into primary faults and secondary faults; When a hardware device fails, determine whether the fault type is a primary fault; If the fault type is a primary fault, extract the fault feature vector ,in, Indicates the device log keyword frequency and performance indicators; Set up a fault case library C, which contains all the fault cases that have occurred in history, and represent each fault case as ,in, is the fault feature vector, For troubleshooting, To solve the results, is the timestamp of the fault case; According to the feature vector X and the fault case library C, the classification model prediction strategy is executed to predict the probability that the fault solution of each fault case can successfully solve the current fault; Get the probabilities of all fault solutions, calculate the mean of the probabilities, and record the result as the average probability; Obtain all types of fault solutions with probabilities greater than the average probability, recorded as the candidate type set; According to the feature vector X and the candidate category set, a sorting model optimization strategy is executed to optimize the priority of each fault solution in the candidate category set; Get the fault solution with the highest priority and record it as the target solution ; Application Target Solutions Fix the current fault and record the feedback result as ; Dynamically optimize the model for recommending fault solutions based on feedback results; Determine whether the faulty device has been repaired successfully. If the repair is successful, the hardware device whose primary fault has been repaired is considered to have a secondary fault. After repairing the hardware equipment with primary fault, a maintenance plan is dynamically generated based on multi-dimensional factors; If the repair fails, each element of the target solution is removed from the candidate set and used in turn to repair the hardware device.
2. A method for self-healing hardware device failure according to claim 1, characterized in that: The classification model prediction strategy is executed according to the feature vector X and the fault case library C to predict the probability of the fault solution of each fault case successfully solving the current fault, including: Obtain all different fault solutions in the fault case library to form a solution category set; For any fault solution in the solution category set, set the score Z of the fault solution; Get feature vector ; calculate , where W is the weight matrix, is the parameter of the model, and b is the bias term; Execute the following formula to convert the score of the fault solution into a probability, specifically: Get the number of elements in the solution category set, recorded as the total number of categories K; Number the fault solutions in the solution category set from 1 to K; calculate , where j = 1, 2…K, is the probability of fault solution j; Calculate the cross entropy loss: , where n is the number of samples and K is the number of categories. is the one-hot encoding of the true label, is probability; The result of calculating the cross entropy loss is recorded as the loss value; Setting a loss threshold, where the loss threshold is used to determine the accuracy of the prediction result; When the loss value ≦ the loss threshold, the classification model prediction strategy ends; When the loss value > loss threshold, the gradient descent method is used to update the parameters W and b, specifically: calculate , as the new W; calculate , as the new b, where is the learning rate; The classification model prediction strategy is executed again using the new W and b as parameters.
3. A method for self-healing hardware device failure according to claim 1, characterized in that: The method of executing a sorting model optimization strategy according to the feature vector X and the candidate category set to optimize the priority of each fault solution in the candidate category set includes: Get the number of all elements in the candidate category set, denoted as F; Sort each fault solution in the candidate set from 1 to F; For any fault solution in the candidate set, set the score sco of the fault solution; Get feature vector ; ,in, is the fault solution j, 1≦j≦F, Used to extract feature vector X and candidate solutions The feature of , W is the parameter of the model, and b is the bias term; Obtain the score of each fault solution in the candidate set, and sort the fault solutions according to the score; Arbitrarily obtain two fault solutions in the candidate category set and record them as and ,in Scored higher than ; calculate , and record the result as the sort value; Compare the ranking value with 1: If the ranking value is greater than 1, stop executing the ranking model optimization strategy; If the ranking value is ≤ 1, the gradient descent method is used to update the parameters W and b, specifically: calculate , as the new W; calculate , as the new b, where is the learning rate; The sorting model optimization strategy is executed again using the new W and b as parameters.
4. A method for self-healing hardware device failure according to claim 1, characterized in that: The model for dynamically optimizing and recommending fault solutions according to the feedback results includes: The feedback results ,in, Indicates that the fault is solved, -1 indicates that the fault is not solved. is the maximum allowed processing time; Defining policies , indicating that the fault characteristics are Execute troubleshooting when The probability distribution of , where 0≦u≦n; Set the discount factor γ, 0≦γ≦1, to balance immediate rewards and long-term rewards; Calculating cumulative rewards ; Setting reward thresholds; When the cumulative reward ≦ reward threshold, the policy gradient method is used to optimize θ: 。 5. A method for self-healing hardware device failure according to claim 1, characterized in that: After the hardware device with the first-level fault is repaired, a maintenance plan is dynamically generated based on multi-dimensional factors, including: Build a device status model: ; in, is the health status of the hardware device at time t, 0 indicates failure and 1 indicates normal operation; is the initial health state; The workload intensity of the hardware device at time t; The accumulated running time of the hardware device at time t; is the strength of maintenance measures for hardware equipment at time t; , , are the weight factors of the impact of workload, service life and maintenance measures on health status.
6. A method for self-healing hardware device failure according to claim 1, characterized in that: After the hardware device with the first-level fault is repaired, a maintenance plan is dynamically generated based on multi-dimensional factors, including: Generate a priority score for the next maintenance plan for each hardware device: ,The result is recorded as the maintenance priority score; in, , , is the weight parameter; , , The workload, age, and health status of each hardware device; Set maintenance thresholds; Obtain hardware devices whose maintenance priority scores are greater than the maintenance threshold, and prioritize maintenance for these hardware devices.
Citation Information
Patent Citations
Neural network-based aero-engine sensor fault self-diagnosis method
CN114330517A
Power operation inspection fault text classification method, system and equipment
CN115994216A
Intelligent operation and maintenance system and method for digital twin substation
CN118172040A
Operation and maintenance system of automatic industrial control system
CN118348872A
Automatic fault detection and repair integrated circuit design method based on neural network
CN119149283A