Troubleshooting methods, equipment, media, and program products
By utilizing fault information and predetermined test examples in large-scale data centers, the fault simulation strategy accurately locates and repairs fault scenarios, solves the efficiency and cost issues of traditional fault handling modes, and achieves efficient and accurate fault handling.
Patent Information
- Application Number
- CN202511030153.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-07-25
AI Technical Summary
The traditional single fault handling mode is difficult to meet the complex and diverse fault scenarios of large-scale data centers, resulting in extended equipment downtime, increased maintenance costs, and the inability to handle unknown fault modes in a timely manner, posing system risks.
By responding to fault notifications, determining fault information, and utilizing fault simulation strategies and simulation processing strategies in predetermined test examples, historical fault scenarios are simulated and repaired. Based on causal reasoning and fault causal graphs, faults are accurately located, target test examples are generated, and fault repair is performed.
It improves the accuracy of fault diagnosis and processing efficiency, significantly reduces maintenance costs, enhances equipment reliability, provides intelligent support for fault handling, and reduces fault investigation time and equipment downtime.
Smart Images

Figure CN120524182B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a fault handling method, device, medium, and program product. Background Art
[0002] In complex operating environments, high-load operating conditions, or with delicate internal structures, equipment is prone to complex and diverse failure types. Traditional single-fault handling models are no longer sufficient to address the diverse range of fault scenarios corresponding to these types of faults, resulting in extended equipment downtime, increased repair costs, and even greater system risks due to untimely fault handling. Summary of the Invention
[0003] In view of the above problems, the present application provides a fault handling method, apparatus, device, medium and program product.
[0004] According to a first aspect of the present application, a fault handling method is provided, comprising: in response to a first fault notification for a first target device, determining fault information based on the first fault notification, the fault information indicating at least one fault scenario triggering the first fault notification; based on the fault information, determining at least one predetermined test example that matches at least one fault scenario from a plurality of predetermined test examples, the predetermined test examples comprising a fault simulation strategy and a simulation processing strategy that matches the fault simulation strategy, the fault simulation strategy being used to simulate historical faults, and the simulation processing strategy being used to repair historical faults; determining a target test example based on a fault repair situation for a historical fault scenario obtained using the fault simulation strategy in at least one predetermined test example; and performing fault repair on the first target device based on the target processing strategy of the target test example.
[0005] The second aspect of the present application provides a fault handling device, including: a first determination module, for determining fault information in response to a first fault notification for a first target device based on the first fault notification, the fault information indicating at least one fault scenario that triggers the first fault notification; a second determination module, for determining at least one predetermined test example that matches at least one fault scenario from a plurality of predetermined test examples based on the fault information, the predetermined test examples including a fault simulation strategy and a simulation processing strategy that matches the fault simulation strategy, the fault simulation strategy being used to simulate historical faults, and the simulation processing strategy being used to repair historical faults; a third determination module, for determining a target test example based on the fault repair status of the historical fault scenario obtained using the fault simulation strategy in at least one predetermined test example; and a first execution module, for performing fault repair on the first target device based on the target processing strategy of the target test example.
[0006] The third aspect of the present application provides an electronic device, comprising: one or more processors; a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the above-mentioned fault handling method.
[0007] The fourth aspect of the present application further provides a computer-readable storage medium having a computer program or instruction stored thereon, which implements the steps of the above-mentioned fault handling method when the above-mentioned computer program or instruction is executed by a processor.
[0008] The fifth aspect of the present application further provides a computer program product, comprising a computer program or instructions, which implement the steps of the above-mentioned fault handling method when executed by a processor.
[0009] According to the embodiments of the present application, since at least one fault scenario that triggers the first fault notification indicated by the fault information matches at least one predetermined test example, the fault can be clearly located. Furthermore, based on the fault repair status of the historical fault scenario obtained by the fault simulation strategy in the at least one predetermined test example, a more optimal target processing strategy can be determined based on the at least one predetermined test example, effectively improving the accuracy and processing efficiency of fault diagnosis, significantly reducing maintenance costs, and enhancing the reliability of the first target device. This is particularly suitable for the complex and diverse fault scenarios in large-scale data centers. Furthermore, it provides strong support for intelligent fault processing and helps accumulate fault processing experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The above contents and other objects, features and advantages of the present application will become more apparent through the following description of the embodiments of the present application with reference to the accompanying drawings.
[0011] Figure 1 An application scenario diagram of the fault handling method, apparatus, device, medium, and program product according to an embodiment of the present application is shown.
[0012] Figure 2 A flowchart of a fault handling method according to an embodiment of the present application is shown.
[0013] Figure 3 A schematic diagram of determining a predetermined test example according to an embodiment of the present application is shown.
[0014] Figure 4 A schematic diagram of determining a predetermined test example according to another embodiment of the present application is shown.
[0015] Figure 5 A schematic diagram showing an example of updating a scheduled test according to an embodiment of the present application is shown.
[0016] Figure 6A schematic diagram of updating a predetermined time interval according to an embodiment of the present application is shown.
[0017] Figure 7 A structural block diagram of a fault handling device according to an embodiment of the present application is shown.
[0018] Figure 8 A block diagram of an electronic device suitable for implementing a fault handling method according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0019] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present application. In the detailed description below, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present application. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present application.
[0020] The terms used herein are only for describing specific embodiments and are not intended to limit the present application. The terms "comprise," "include," etc. used herein indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0021] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0022] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).
[0023] During the implementation of the present application, we discovered that as data center server capacity surpasses tens of millions, system crashes caused by hardware failures have become the primary factor impacting service reliability. Traditional single-step fault handling models, such as predicting faults based on historical logs, building virtual test environments, and pre-setting test scripts for individual components, are no longer able to meet the needs of large-scale data centers for efficient and accurate fault handling.
[0024] For example, fault prediction based on historical logs relies on training models based on statistical features of historical logs. This only captures correlations between faults and operating parameters, such as the co-occurrence of hard drive bad sectors and system freezes, leading to ambiguous fault location. Building a virtual test environment can only predict fault types that have already appeared in historical data, lacking coverage for unknown failure modes such as new hardware coupling failures and extreme boundary conditions, such as simultaneous failure of multiple components. Test scenario generation relies on manual experience, resulting in insufficient robustness.
[0025] In view of this, an embodiment of the present application provides a fault handling method, including: in response to a first fault notification for a first target device, determining fault information based on the first fault notification, the fault information indicating at least one fault scenario triggering the first fault notification; based on the fault information, determining at least one predetermined test example that matches at least one fault scenario from multiple predetermined test examples, the predetermined test example including a fault simulation strategy and a simulation processing strategy that matches the fault simulation strategy, the fault simulation strategy is used to simulate historical faults, and the simulation processing strategy is used to repair historical faults; determining a target test example based on the fault repair situation of the historical fault scenario obtained using the fault simulation strategy in at least one predetermined test example; and performing fault repair on the first target device based on the target processing strategy of the target test example.
[0026] Figure 1 An application scenario diagram of the fault handling method, apparatus, device, medium, and program product according to an embodiment of the present application is shown.
[0027] like Figure 1 As shown, the application scenario 100 according to this embodiment may include a first target device 110 and an electronic device 120. A network may be used as a medium for providing a communication link between the first target device 110 and the electronic device 120. The network may include various connection types, such as wired or wireless communication links or fiber optic cables.
[0028] The first target device 110 interacts with the electronic device 120 via a network to receive or send messages, etc. The first target device 110 may be any hardware device that has specific functions and may fail during operation, such as a server, a smart mobile device, etc.
[0029] The electronic device 120 can be any device with data processing and control functions, such as, but not limited to, test equipment, servers, industrial computers, etc. For example, the server can be a server that provides various services, such as a background management server that can process the first fault notification of the first target device 110 (only as an example). The background management server can analyze and process the received first fault notification, and perform fault repair on the first target device 110 based on the processing results (such as the determined target test example, etc.). The test equipment may include a digital twin linkage module for the first target device 110. The real-time interaction between the digital twin linkage module and the first target device 110 realizes the precise mapping and optimized management of the first target device 110. This application does not specifically limit the construction of the digital twin linkage module.
[0030] It should be understood that Figure 1 The number of the first target devices and electronic devices in the embodiment is merely illustrative. Any number of the first target devices and electronic devices may be provided according to implementation requirements.
[0031] The following will be based on Figure 1 The scene described by Figures 2 to 6 The fault handling method of the application embodiment is described in detail.
[0032] Figure 2 A flowchart of a fault handling method according to an embodiment of the present application is shown.
[0033] like Figure 2 As shown, the fault handling method of this embodiment includes operations S210 to S240.
[0034] In operation S210 , in response to a first failure notification for a first target device, failure information is determined based on the first failure notification.
[0035] In operation S220 , at least one predetermined test example that matches each of the at least one fault scenarios is determined from a plurality of predetermined test examples based on the fault information.
[0036] In operation S230 , a target test case is determined based on a fault repair situation for a historical fault scenario obtained by using a fault simulation strategy in at least one predetermined test case.
[0037] In operation S240 , fault repair is performed on the first target device based on the target processing policy of the target test example.
[0038] In the embodiment of the present application, the fault information indicates at least one fault scenario that triggers the first fault notification.
[0039] The first fault notification is used to represent a signal sent by the first target device after detecting an abnormal state.
[0040] For example, the content of the first fault notification can be parsed, and real-time operating data of the first target device can be collected based on the content of the first fault notification. Then, based on the content of the first fault notification and the real-time operating data of the first target device, at least one fault scenario that triggered the first fault notification can be inferred. The content of the first fault notification can include, but is not limited to, an error code, a description, and a timestamp. The real-time operating data can include, but is not limited to, operation records and sensor data.
[0041] Fault scenarios can be used to characterize faults that occur during the operation of the first target device. For example, when the first target device is a server, the fault scenarios may include but are not limited to the hard disk bad sector rate increasing to a critical value, the network delay increasing by 300ms, the cooling fan speed being forced to decrease by 50%, etc. When the first target device is a smart mobile device, the smart mobile device may include but is not limited to aerospace equipment, smart cars, etc. For aerospace equipment, the fault scenarios may include but are not limited to engine vibration, circuit board solder joint fatigue, signal transmission errors, etc. For smart cars, the fault scenarios may include but are not limited to low temperature environments, reduced battery life, limited computing resources, and perception algorithm failure.
[0042] The predetermined test example may include a fault simulation strategy and a simulation processing strategy matching the fault simulation strategy, wherein the fault simulation strategy is used to simulate a historical fault, and the simulation processing strategy is used to repair the historical fault.
[0043] For example, historical faults that can be simulated can be determined based on the fault simulation strategy. When the historical fault matches the fault represented by the fault scenario, it can be determined that the fault scenario matches the fault simulation strategy, and then it can be determined that the fault scenario matches a predetermined test example consisting of the fault model strategy and the simulation processing strategy that matches the fault simulation strategy.
[0044] Since the same fault simulation strategy can be matched with one or more simulation processing strategies, there can be one or more predetermined test examples matching one fault scenario.
[0045] For example, a predetermined test example or all predetermined test examples in which a fault is successfully repaired among multiple predetermined test examples may be determined as target test examples. Alternatively, a predetermined test example or a predetermined test example in which a fault is successfully repaired and takes the shortest time to repair may be determined as target test examples.
[0046] Because at least one fault scenario triggering the first fault notification, as indicated by the fault information, matches at least one predetermined test example, the fault can be clearly located. Furthermore, based on the fault repair results of historical fault scenarios derived from the fault simulation strategy in the at least one predetermined test example, a more optimal target processing strategy can be determined based on the at least one predetermined test example. This effectively improves the accuracy and efficiency of fault diagnosis, significantly reduces maintenance costs, and enhances the reliability of the first target device. This approach is particularly suitable for the complex and diverse fault scenarios found in large-scale data centers. Furthermore, it provides strong support for intelligent fault processing and helps accumulate fault processing experience.
[0047] During the implementation of the embodiments of the present application, it was also found that fault prediction based on historical logs relies on a training model based on the statistical features of historical logs, which only captures the correlation between faults and operating parameters, such as the co-occurrence of hard disk bad sectors and system freezes. It is unable to further determine the root cause of the hard disk bad sectors, resulting in ambiguous fault location.
[0048] Based on this, according to the embodiment of the present application, the fault scenario may include multiple, and the fault information may include a fault chain. Figure 2 Operation S210, in response to a first fault notification for a first target device, may include determining fault information based on the first fault notification. This may include performing causal reasoning on the first fault notification to obtain at least one fault scenario leading to the first fault notification. Determining, from a fault causal graph, historical fault scenarios that match each of the at least one fault scenarios. Determining a fault chain based on other historical fault scenarios that have a causal relationship with each of the at least one fault scenarios, as well as the causal strength of the causal relationship.
[0049] Exemplarily, a large language model may be used to perform causal reasoning on the first fault notification.
[0050] A fault causal graph can be pre-built based on causal reasoning from historical fault notifications. Nodes in the fault causal graph represent historical fault scenarios, and edges between nodes indicate the causal relationship between them. Causal strength quantifies the degree to which one historical fault scenario (cause) affects another historical fault scenario (effect). For example, the causal strength quantifying the causal relationship between a drop in power supply output voltage (cause) and fluctuations in the motherboard power supply circuit (effect) can be 0.92. The causal strength quantifying the causal relationship between fluctuations in the motherboard power supply circuit (cause) and an increase in memory parity error rate (effect) can be 0.78. The causal strength quantifying the causal relationship between memory error accumulation (cause) and the triggering of processor frequency reduction protection can be 0.65, and so on.
[0051] Historical fault scenarios that match at least one fault scenario and other historical fault scenarios can be linked in descending order of causal strength to form a fault chain. For example, a fault chain can be, but is not limited to, output voltage drop → motherboard power supply circuit fluctuation → memory check error rate increase → processor frequency reduction protection triggering.
[0052] Causal reasoning on the first fault notification can track hidden faults and improve fault location accuracy. Based on the fault scenario identified by causal reasoning, the fault causal graph can be used to match historical fault scenarios, quickly identifying historical fault examples similar to the current fault, thereby providing multiple verified solutions for troubleshooting. This approach not only reduces troubleshooting time but also improves the success rate of handling strategies, reduces downtime and repair costs for the first target device, and enhances the reliability and stability of the first target device.
[0053] In another embodiment of the present application, the fault handling method may include the following: Figure 2 In addition to operations S210 to S240, the system may also include generating multiple initial fault simulation strategies based on multiple historical fault scenarios indicated by each of the multiple historical fault chains. For any of the multiple initial fault simulation strategies, determining multiple initial simulation processing strategies that match the initial fault simulation strategy; performing fault repairs sequentially based on the multiple initial simulation processing strategies for the historical fault scenarios obtained using any of the initial fault simulation strategies to obtain multiple initial simulation results; and determining predetermined test examples based on the multiple initial simulation results, the initial fault simulation strategies, and the multiple initial simulation processing strategies.
[0054] In the embodiments of the present application, a causal relationship exists between the multiple historical fault scenarios indicated by the historical fault chain. Multiple initial fault simulation strategies can be sorted based on the causal strength of the causal relationship, and fault simulations can be performed sequentially according to the order of the multiple initial fault simulation strategies. After completing a fault simulation, the corresponding fault repair must be performed, which can be performed multiple times or once.
[0055] Figure 3 A schematic diagram of determining a predetermined test example according to an embodiment of the present application is shown.
[0056] like Figure 3 As shown, a historical fault chain 310 is used as an example: historical fault 1 → historical fault 2 → ... → historical fault N. A first initial fault simulation strategy 321 can be determined based on historical fault 1, a second initial fault simulation strategy 322 can be determined based on historical fault 2, and so on. A second initial fault simulation strategy 32N can be determined based on historical fault N. N is an integer greater than 2.
[0057] The causal strength of the causal relationship between historical fault 1 and historical fault 2 is greater than the causal strength of the causal relationship between historical fault 2 and historical fault 3, and so on. According to the order of the first initial fault simulation strategy 321, the second initial fault simulation strategy 322, ..., and the Nth initial fault simulation strategy 32N, a fault simulation may be performed based on the first initial fault simulation strategy 321, and then fault repair may be performed sequentially based on multiple initial simulation processing strategies matching the first initial fault simulation strategy 321 to obtain multiple initial simulation results. This process may be repeated to perform fault simulation and fault repair to obtain multiple initial simulation results.
[0058] After any initial fault simulation strategy is determined, a large language model may be used or a search tool may be called to determine multiple initial simulation processing strategies that match the initial fault simulation strategy.
[0059] For the first initial fault simulation strategy 321 , the second initial fault simulation strategy 322 , . . . , the Nth initial fault simulation strategy 32N, a first predetermined test example 351 , a second predetermined test example 352 , . . . , the Nth predetermined test example 35N can be obtained respectively.
[0060] Taking the second initial fault simulation strategy 322 as an example, the search function of the large language model can be used based on the second initial fault simulation strategy 322 to obtain the first initial simulation processing strategy 331, ..., and the Mth initial simulation processing strategy 33M that match the second initial fault simulation strategy 322. M is an integer greater than or equal to 1. A fault simulation is performed based on the second initial fault simulation strategy 322, and the fault is repaired based on the first initial simulation processing strategy 331, to obtain a first initial simulation result 341. Similarly, a fault simulation is performed based on the second initial fault simulation strategy 322, and the fault is repaired based on the Mth initial simulation processing strategy 33M, to obtain an Mth initial simulation result 34M. A second predetermined test example 352 can be determined based on the first initial simulation results 341, ..., and the Mth initial simulation result 34M. The second predetermined test example 352 can be one or more. Generating an initial fault simulation strategy based on a historical fault chain can accurately reproduce fault scenarios and improve the coverage of fault handling test scenarios. On this basis, the initial simulation processing strategy is determined. After fault simulation and repair verification, its effectiveness and feasibility can be ensured, the accuracy and efficiency of fault handling can be improved, and a reusable solution can be provided for subsequent similar faults.
[0061] In one example, the first target device may include multiple components. A fault causal graph is constructed based on causal reasoning of multiple historical fault notifications. Multiple historical fault chains are determined based on the causal strength of the causal relationship and the fault causal graph.
[0062] The nodes of the fault causal graph may indicate the historical fault scenarios of the respective components of the multiple components associated with the multiple historical fault notifications. The edges of the fault causal graph may indicate the causal relationship between adjacent historical fault scenarios.
[0063] For example, the first target device may be a server, and its components may include a power supply, memory, processor, hard drive, controller, and so on. Of the two historical fault scenarios with a high causal strength, the historical fault scenario that is the cause can be used as the root node, and the historical fault scenario that is the effect can be used as the leaf node. For example, if the historical fault scenario determined based on a historical fault notification is an increase in cooling system pressure, the components associated with the notification may be the processor, cooling system, fan, and the entire machine. The determined historical fault chain may be abnormal processor power consumption → increased cooling system pressure → increased fan speed → increased decibel noise level of the entire machine, etc.
[0064] Based on the causal strength of causal relationships and the fault causal graph, multiple historical fault chains are determined, which can accurately identify the root cause of the fault and its propagation path, thereby achieving rapid positioning and efficient repair.
[0065] In another example, the initial simulation results may indicate a failure recovery situation for a historical failure scenario.
[0066] Figure 4 A schematic diagram of determining a predetermined test example according to another embodiment of the present application is shown.
[0067] Determining predetermined test cases based on a plurality of initial simulation results, an initial fault simulation strategy, and a plurality of initial simulation processing strategies may include: Figure 4 Operations S401 to S408 are shown.
[0068] In operation S401, it is determined whether at least one simulation result among a plurality of initial simulation results indicates that the fault for the historical fault scenario has been repaired. If so, operations S402 to S403 are performed. If not, operations S404 to S405 or operations S406 to S408 are performed.
[0069] In operation S402 , an initial fault simulation strategy is determined as a fault simulation strategy.
[0070] In operation S403 , based on the at least one simulation result, a simulation processing strategy matching the fault simulation strategy is determined from at least one initial simulation processing strategy corresponding to the at least one simulation result.
[0071] For example, if only one simulation result exists among the multiple initial simulation results, the initial simulation processing strategy corresponding to the simulation result can be determined as the simulation processing strategy that matches the fault simulation strategy. If multiple simulation results exist among the multiple initial simulation results, the simulation result with the shortest repair time can be screened from the multiple simulation results based on the repair time, and the initial simulation processing strategy corresponding to the simulation result with the shortest repair time can be determined as the simulation processing strategy that matches the fault simulation strategy.
[0072] Since the simulation processing strategy that matches the fault simulation strategy is determined based on the fault repair situation of historical fault scenarios, the effectiveness of the simulation processing strategy can be ensured. Furthermore, if there are multiple simulation results, the simulation processing strategy that matches the fault simulation strategy is determined based on the repair time, which can ensure the efficiency of fault repair.
[0073] In operation S404 , a policy update request is sent to the client so that the client updates the initial fault simulation policy and / or the multiple initial simulation processing policies.
[0074] In operation S405 , in response to receiving the updated initial fault simulation strategy and / or the updated initial simulation processing strategy returned by the client, a predetermined test instance is determined based on the updated initial fault simulation strategy and / or the updated initial simulation processing strategy.
[0075] The policy update request may include an initial fault simulation policy and multiple initial simulation processing policies. Upon receiving the policy update request, the client may modify the initial fault simulation policy and / or multiple initial simulation processing policies. Since the fault repair was unsuccessful during fault simulation and repair, interaction with the client may enhance the reference value of the predetermined test example.
[0076] In operation S406 , state information of simulating historical fault scenarios based on the initial fault simulation strategy is recorded.
[0077] In operation S407 , the initial fault simulation strategy is updated based on the status information.
[0078] In operation S408 , a predetermined test case is determined based on the updated initial fault simulation strategy.
[0079] The status information can represent a fault condition associated with the operating status of the associated component when simulating a historical fault scenario based on the initial fault simulation strategy. A new fault scenario can be determined based on the fault condition, and a new fault simulation strategy can be determined based on the new fault scenario. The initial fault simulation strategy can then be updated to the new fault simulation strategy. At least one new simulation processing strategy matching the new fault simulation strategy can be determined based on the new fault simulation strategy. A fault can be simulated based on the new fault simulation strategy, and the fault can be repaired based on the at least one new simulation processing strategy, resulting in at least one new simulation result. A predetermined test case can then be determined based on the at least one new simulation result.
[0080] It should be noted that, based on at least one newly added simulation result, the predetermined test example can be determined as described above. Figure 4 The method shown is similar and will not be repeated here.
[0081] Since the unrepaired faults in historical fault scenarios can be considered as new faults that may occur in the actual operation of the first target device, by recording status information, new fault scenarios can be continuously determined based on the fault conditions that occur after the faults are not repaired, thereby improving the test coverage of the fault scenarios.
[0082] In another embodiment of the present application, the fault handling method may include the following: Figure 2 In addition to the operations S210 to S240 shown, the method may further include an operation of updating a plurality of initial simulation results based on sensor data of the environment where the first target device is located, acquired at predetermined time intervals.
[0083] The predetermined time interval may be an empirical value determined according to actual conditions, such as 30 seconds. The sensor data may include but is not limited to at least one of temperature, voltage, and clock signal.
[0084] For example, the integrated data of the digital twin linkage module can be updated in real time according to the sensor data, and then the fault simulation and fault repair can be re-executed based on the updated integrated data to obtain multiple new simulation results, and multiple initial simulation results can be updated to multiple new simulation results.
[0085] Figure 5 A schematic diagram showing an example of updating a scheduled test according to an embodiment of the present application is shown.
[0086] For example, Figure 5 As shown, taking the predetermined time interval △T as an example, a fault simulation can be performed based on the sensor data 510 of the environment in which the first target device is located obtained every △T based on the fault simulation strategy, and the simulated fault can be repaired based on the simulation processing strategy to obtain a new simulation result, and the initial simulation result 520 can be updated using the new simulation result.
[0087] In response to the updating of the plurality of initial simulation results 520 , the predetermined test cases 530 may be updated.
[0088] Determine a new predetermined test example based on the new simulation result, and then replace the predetermined test example with the new predetermined test example to achieve the update. Determine a new predetermined test example based on the new simulation result, and the same method as above can be used. Figure 4 The method shown is similar and will not be repeated here.
[0089] As the sensor data of the environment in which the first target device is located changes, the initial simulation results of fault repair after simulating historical fault scenarios also change. Therefore, through real-time updates, a test result that is more in line with the actual operation of the first target device can be obtained.
[0090] During the sequential fault repair process for historical fault scenarios generated using any initial fault simulation strategy, that is, during the fault simulation and fault repair process, the point at which the test device approaches a critical point of system failure can be recorded. If simulation results exist, but the test device approaches a critical point of system failure during the simulation results, the corresponding fault handling strategy is invalid and feedback is required to the client to confirm whether to modify the fault handling strategy.
[0091] For example, for the above Figure 2 In operation S230, the at least one predetermined test example includes multiple predetermined test examples. Determining a target test example based on a fault repair situation for a historical fault scenario obtained using a fault simulation strategy in the at least one predetermined test example may include the following operations: sorting the plurality of predetermined test examples according to the fault repair situations corresponding to the plurality of predetermined test examples to obtain sorted predetermined test examples; and sorting the plurality of simulation processing strategies in the plurality of predetermined test examples according to the sorted predetermined test examples, so that fault repair can be sequentially performed on the first target device based on the sorting of the plurality of simulation processing strategies.
[0092] After the fault repair is performed on the first target device, it can also be determined whether to update the multiple predetermined test examples based on the fault repair situation of the first target device. If there is a situation where the fault repair is not performed on the first target device, the status information generated by the repair process can be recorded, and then a new fault scenario is determined based on the status information, and then a new test example is determined. The method of determining a new test example based on a new fault scenario can be the same as the above-mentioned method. Figure 2 The method shown is similar and will not be described in detail here. Thus, the predetermined test examples can be continuously expanded.
[0093] The sensor data may include temperature. The fault handling method may further include the steps of: assigning a weight value to any predetermined test example from the plurality of predetermined test examples based on initial simulation results corresponding to the plurality of predetermined test examples; determining an associated test example associated with the temperature from the plurality of predetermined test examples when it is determined that the difference between temperatures obtained between two adjacent predetermined time intervals is greater than or equal to a predetermined difference; and updating the weight value of the associated test example.
[0094] A weight value may be assigned to any of the plurality of predetermined test examples based on the fault repair status for the historical fault scenario indicated by the initial simulation result. The closer the fault repair is to being successfully repaired, the greater the weight value of the corresponding predetermined test example.
[0095] The predetermined difference value may be an empirical value determined according to actual conditions. The weight value of the associated test example may be updated to the largest value among the plurality of predetermined test examples.
[0096] For example, when there are multiple predetermined test examples and multiple target test examples are determined, the weights of the target test examples can be re-determined based on the updated weights of the associated test examples. Then, fault repair can be performed on the first target device in order of the weights of the target test examples. For example, if the temperature in the computer room rises by 5°C, cooling system-related tests will be prioritized.
[0097] Since the difference in temperature obtained between two adjacent predetermined time intervals is greater than or equal to the predetermined difference, the change in temperature may cause a change in the actual operating condition of the first target device. In this case, if the fault is related to temperature, the target processing strategy related to temperature can be executed preferentially, thereby further accelerating the efficiency of fault processing.
[0098] Considering that during the actual operation of the first target device, the wear and tear of each component of the first target device will have an impact on the actual fault repair.
[0099] Based on this, during the fault simulation and fault repair process, based on multiple initial simulation processing strategies, fault repair is sequentially performed for historical fault scenarios obtained using any of the initial fault simulation strategies. That is, during the fault simulation and fault repair process, the actual loss of the associated components due to the test intensity is determined based on the test intensity of the associated components corresponding to the initial simulation processing strategy or the initial fault simulation strategy. If the actual loss does not meet the predetermined conditions, the test intensity of the associated components is adjusted. If the actual loss meets the predetermined conditions, the test intensity of the associated components is not adjusted.
[0100] Test intensity can include the number of simulations or simulation durations for the associated component, or the number of fault repairs or repair durations for the associated component. The actual wear and tear of the associated component can be calculated based on a material fatigue model. The material fatigue model can be a mapping relationship between test intensity and wear and tear.
[0101] For example, for every 100 memory bit flip tests, the actual wear and tear on the memory could be a 0.02% loss in lifespan. For a power module full-load test lasting more than one hour, the actual wear and tear on the power supply could be a tripling of the capacitor aging rate. The predefined conditions can be determined based on the wear and tear limits of the associated components. For example, if the capacitor aging rate increases by a factor of three, exceeding the wear and tear limit of the power supply, the predefined conditions are not met, and the power supply test intensity needs to be reduced.
[0102] Since the actual damage to the components of the first target device caused by the test intensity is taken into account during the fault simulation and fault repair process, the test strategy can be formulated more accurately, avoiding device damage caused by excessive testing and ensuring the effectiveness and reliability of the test.
[0103] For example, the first target device may include multiple components, and the historical failure scenario simulated in any predetermined test example is for associated components associated with the historical failure scenario among the multiple components. The actual loss condition may include an actual loss percentage.
[0104] Figure 6 A schematic diagram of updating a predetermined time interval according to an embodiment of the present application is shown.
[0105] like Figure 6 As shown, updating the predetermined time interval may further include operations S601 to S604.
[0106] In operation S601 , based on the number of simulations for any predetermined test case, an actual loss percentage of an associated component corresponding to any predetermined test case is determined.
[0107] In operation S602, it is determined whether the actual loss percentage of the associated component corresponding to any predetermined test example is greater than a predetermined threshold. If yes, operation S603 is executed; if not, operation S604 is executed.
[0108] In operation S603, the predetermined time interval is updated.
[0109] In operation S604, the predetermined time interval is not updated.
[0110] The predetermined threshold value may be determined based on the wear limit of the associated component itself.
[0111] For example, based on the actual operation of the associated component, a relationship curve between the number of simulations and the loss rate of the associated component can be constructed. Then, based on the number of simulations and the relationship curve, the actual loss percentage of the associated component can be calculated.
[0112] In another embodiment of the present application, the fault handling method may include the following: Figure 2 In addition to operations S210 to S240 shown, operations may also be included: in response to a second fault notification from a second target device, upon determining that there is a predetermined test example that complies with preset rules among multiple predetermined test examples, performing fault repair on the second target device based on a simulation processing strategy of the predetermined test example that complies with the preset rules.
[0113] In this embodiment, the second target device has a different configuration from the first target device. For example, the second target device and the first target device are manufactured by different manufacturers, resulting in a different configuration for the second target device and the first target device.
[0114] The preset rule may include that the fault scenario indicated by the fault information determined based on the second fault notification is independent of the configuration of the second target device, that is, there is a predetermined test example matching the fault scenario among the multiple predetermined test examples, and the predetermined test example complies with the preset rule.
[0115] If it is determined that no predetermined test example in the predetermined test example set meets the preset rules, other fault scenarios that caused the second fault notification are determined based on the second fault notification. The other fault scenarios are sent to the client, so that the client can configure a fault repair policy for the other fault scenario based on the device manufacturer of the second target device. Fault repair is performed on the second target device based on the fault repair policy provided by the client.
[0116] When it is determined that there is no predetermined test example that meets the preset rule in the predetermined test example set, the second target device may also be managed based on other test devices.
[0117] Since the configuration of the second target device is different from that of the first target device, by determining whether there is a predetermined test example that meets the preset rules, the migration of test examples across devices can be achieved, thereby improving the fault handling efficiency of the second target device.
[0118] It should be noted that the test device and other test devices can form a distributed node and share the encrypted multiple predetermined test examples determined by each. In the case of a common fault simulation strategy, the debugging of the predetermined test examples or test devices can be reduced.
[0119] Based on the above fault handling method, this application also provides a fault handling device. Figure 7 The device is described in detail.
[0120] Figure 7 The structural block diagram of the fault handling device according to an embodiment of the present application is schematically shown.
[0121] like Figure 7 As shown, the fault handling device 700 of this embodiment includes a first determining module 710 , a second determining module 720 , a third determining module 730 and a first executing module 740 .
[0122] The first determination module 710 is configured to, in response to a first fault notification for a first target device, determine fault information based on the first fault notification, where the fault information indicates at least one fault scenario that triggered the first fault notification. In one embodiment, the first determination module 710 may be configured to perform operation S210 described above, which will not be further described herein.
[0123] Second determination module 720 is configured to determine, based on the fault information, at least one predetermined test example from a plurality of predetermined test examples that matches at least one fault scenario. The predetermined test example includes a fault simulation strategy and a simulation processing strategy that matches the fault simulation strategy. The fault simulation strategy is used to simulate a historical fault, and the simulation processing strategy is used to repair the historical fault. In one embodiment, second determination module 720 can be configured to perform operation S220 described above, and will not be further described here.
[0124] The third determination module 730 is used to determine a target test case based on the fault repair situation of the historical fault scenario obtained by using the fault simulation strategy in at least one predetermined test case. In one embodiment, the third determination module 730 can be used to perform the operation S230 described above, which will not be repeated here.
[0125] The first execution module 740 is used to perform fault repair on the first target device based on the target processing strategy of the target test example. In one embodiment, the first execution module 740 can be used to perform the operation S240 described above, which will not be repeated here.
[0126] According to an embodiment of the present application, the fault handling device 700 further includes: a generation module and an example determination module. The generation module is used to generate multiple initial fault simulation strategies based on multiple historical fault scenarios indicated by multiple historical fault chains. The example determination module is used to determine, for any one of the multiple initial fault simulation strategies, multiple initial simulation processing strategies that match the initial fault simulation strategy; based on the multiple initial simulation processing strategies, sequentially perform fault repair on the historical fault scenarios obtained using any one of the initial fault simulation strategies to obtain multiple initial simulation results; and determine a predetermined test example based on the multiple initial simulation results, the initial fault simulation strategy, and the multiple initial simulation processing strategies.
[0127] According to an embodiment of the present application, an initial simulation result indicates a fault repair status for a historical fault scenario. Determining a predetermined test example based on multiple initial simulation results, an initial fault simulation strategy, and multiple initial simulation processing strategies includes: determining the initial fault simulation strategy as the fault simulation strategy when determining that at least one simulation result among the multiple initial simulation results indicates that the fault for the historical fault scenario has been repaired; and determining, based on the at least one simulation result, a simulation processing strategy that matches the fault simulation strategy from at least one initial simulation processing strategy corresponding to the at least one simulation result.
[0128] According to an embodiment of the present application, the fault handling apparatus 700 further includes a sending module and a fourth determination module. The sending module is configured to, upon determining that no simulation result exists in the multiple initial simulation results, send a policy update request to the client, so that the client updates the initial fault simulation policy and / or the multiple initial simulation processing policies. The fourth determination module is configured to, in response to receiving the updated initial fault simulation policy and / or the updated initial simulation processing policy returned by the client, determine a predetermined test example based on the updated initial fault simulation policy and / or the updated initial simulation processing policy.
[0129] According to an embodiment of the present application, the fault handling apparatus 700 further includes a recording module, a first updating module, and a fifth determining module. The recording module is configured to record status information of a historical fault scenario simulated based on the initial fault simulation strategy if no simulation result is found in the multiple initial simulation results. The first updating module is configured to update the initial fault simulation strategy based on the status information. The fifth determining module is configured to determine a predetermined test example based on the updated initial fault simulation strategy.
[0130] According to an embodiment of the present application, the fault handling apparatus 700 further includes a second updating module configured to update the plurality of initial simulation results based on sensor data of the environment of the first target device acquired at predetermined time intervals.
[0131] According to an embodiment of the present application, the fault handling apparatus 700 further includes a third updating module configured to update the predetermined test example based on the updated initial simulation results in response to the update of the multiple initial simulation results.
[0132] According to an embodiment of the present application, the sensor data includes temperature. The fault handling device 700 also includes an assignment module, a sixth determination module, and a fourth update module. The assignment module is used to assign a weight value to any predetermined test example in the multiple predetermined test examples based on the initial simulation results corresponding to the multiple predetermined test examples. The sixth determination module is used to determine an associated test example associated with the temperature from the multiple predetermined test examples when it is determined that the difference between the temperatures obtained at two adjacent predetermined time intervals is greater than or equal to the predetermined difference. The fourth update module is used to update the weight value of the associated test example.
[0133] According to an embodiment of the present application, the first target device includes multiple components, and the historical failure scenario simulated in any predetermined test example is for an associated component associated with the historical failure scenario among the multiple components. The fault handling device 700 also includes a seventh determination module and a fifth update module. The seventh determination module is configured to determine the actual loss percentage of the associated component corresponding to any predetermined test example based on the number of simulations for any predetermined test example. The fifth update module is configured to update the predetermined time interval if the actual loss percentage of the associated component corresponding to any predetermined test example is greater than a predetermined threshold.
[0134] According to an embodiment of the present application, the first target device includes multiple components. The fault handling apparatus 700 further includes a construction module and a historical fault chain determination module. The construction module is configured to construct a fault causal graph based on causal reasoning of multiple historical fault notifications. The nodes of the fault causal graph indicate historical fault scenarios for each of the multiple associated components associated with the multiple historical fault notifications, and the edges of the fault causal graph indicate causal relationships between adjacent historical fault scenarios. The historical fault chain determination module is configured to determine multiple historical fault chains based on the causal strength of the causal relationships and the fault causal graph.
[0135] According to an embodiment of the present application, the fault information includes a fault chain. The first determination module 710 may include an inference unit, a matching unit, and a sub-determination unit.
[0136] The reasoning unit is configured to perform causal reasoning on the first fault notification to obtain at least one fault scenario leading to the first fault notification. The matching unit is configured to determine, from the fault causal graph, historical fault scenarios that match each of the at least one fault scenarios. The sub-determination unit is configured to determine a fault chain based on other historical fault scenarios that have a causal relationship with each of the at least one fault scenarios, as well as the causal strength of the causal relationship.
[0137] According to an embodiment of the present application, the fault handling apparatus 700 further includes a second execution module. The second execution module is configured to, in response to a second fault notification from a second target device, perform fault repair on the second target device based on a simulation processing strategy for the predetermined test example that meets the preset rule, if it is determined that a predetermined test example that meets the preset rule exists among the plurality of predetermined test examples. The second target device has a different configuration from the first target device.
[0138] According to embodiments of the present application, any multiple modules among the first determination module 710, the second determination module 720, the third determination module 730, and the first execution module 740 may be combined into a single module, or any one of these modules may be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules may be combined with at least part of the functionality of other modules and implemented in a single module. According to embodiments of the present application, at least one of the first determination module 710, the second determination module 720, the third determination module 730, and the first execution module 740 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or may be implemented in hardware or firmware through any other reasonable means of circuit integration or packaging, or may be implemented in any one of the three implementation methods of software, hardware, and firmware, or any appropriate combination of any of these. Alternatively, at least one of the first determination module 710 , the second determination module 720 , the third determination module 730 and the first execution module 740 may be at least partially implemented as a computer program module, which may perform corresponding functions when executed.
[0139] Figure 8 A block diagram of an electronic device suitable for implementing a fault handling method according to an embodiment of the present application is shown.
[0140] like Figure 8 As shown, an electronic device 800 according to an embodiment of the present application includes a processor 801, which can perform various appropriate actions and processes based on a program stored in a read-only memory (ROM) 802 or a program loaded from a storage unit 808 into a random access memory (RAM) 803. The processor 801 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or a related chipset and / or a dedicated microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 801 may also include onboard memory for caching purposes. The processor 801 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present application.
[0141] Various programs and data required for the operation of the electronic device 800 are stored in the RAM 803. The processor 801, ROM 802, and RAM 803 are connected to each other via a bus 804. The processor 801 performs various operations of the method flow according to the embodiment of the present application by executing the programs in the ROM 802 and / or RAM 803. It should be noted that the programs can also be stored in one or more memories other than the ROM 802 and RAM 803. The processor 801 can also perform various operations of the method flow according to the embodiment of the present application by executing the programs stored in one or more memories.
[0142] According to an embodiment of the present application, electronic device 800 may further include an input / output (I / O) interface 805, which is also connected to bus 804. Electronic device 800 may also include one or more of the following components connected to I / O interface 805: an input section 806 including a keyboard, mouse, etc.; an output section 807 including devices such as a cathode ray tube (CRT), liquid crystal display (LCD), and speakers; a storage section 808 including a hard disk; and a communication section 809 including a network interface card such as a LAN card or modem. Communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to I / O interface 805 as needed. Removable media 811, such as a magnetic disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed in drive 810 as needed, so that computer programs read from the removable media can be installed into storage section 808 as needed.
[0143] This application also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments, or may exist independently and not be incorporated into the device / apparatus / system. The computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of this application is implemented.
[0144] According to an embodiment of the present application, a computer-readable storage medium may be a non-volatile computer-readable storage medium, and may include, for example, but not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present application, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present application, a computer-readable storage medium may include the ROM 802 and / or RAM 803 described above and / or one or more memories other than ROM 802 and RAM 803.
[0145] The embodiments of the present application also include a computer program product, which includes a computer program containing program code for executing the method shown in the flowchart. When the computer program product is run in a computer system, the program code is used to enable the computer system to implement the method provided in the embodiments of the present application.
[0146] The computer program executes the above functions defined in the system / device of the embodiment of the present application when the processor 801 executes the computer program. According to the embodiment of the present application, the system, device, module, unit, etc. described above can be implemented by a computer program module.
[0147] In one embodiment, the computer program may be stored on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may be transmitted and distributed in the form of a signal on a network medium, downloaded and installed via the communication portion 809, and / or installed from a removable medium 811. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to wireless, wired, or any suitable combination thereof.
[0148] In such an embodiment, the computer program can be downloaded and installed from the network via the communication section 809, and / or installed from the removable medium 811. When the computer program is executed by the processor 801, the above-mentioned functions defined in the system of the embodiment of the present application are performed. According to the embodiment of the present application, the systems, devices, means, modules, units, etc. described above can be implemented by computer program modules.
[0149] According to an embodiment of the present application, the program code for executing the computer program provided by the embodiment of the present application can be written in any combination of one or more programming languages. Specifically, these computer programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, Python, "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect via the Internet).
[0150] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of the boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0151] Those skilled in the art will appreciate that the features described in the various embodiments of this application may be combined and / or coupled in various ways, even if such combinations or couplings are not explicitly described in this application. In particular, the features described in the various embodiments of this application may be combined and / or coupled in various ways without departing from the spirit and teachings of this application. All such combinations and / or couplings fall within the scope of this application.
[0152] The embodiments of the present application have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present application. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be advantageously used in combination. Without departing from the scope of the present application, those skilled in the art may make various substitutions and modifications, and these substitutions and modifications should all fall within the scope of the present application.
Claims
1. A fault handling method, characterized in that: The method comprises: In response to a first fault notification for a first target device, determining fault information based on the first fault notification, the fault information indicating at least one fault scenario triggering the first fault notification; Based on the fault information, determining at least one predetermined test example that matches at least one fault scenario from a plurality of predetermined test examples, the predetermined test example comprising a fault simulation strategy and a simulation processing strategy that matches the fault simulation strategy, the fault simulation strategy being used to simulate a historical fault, and the simulation processing strategy being used to repair the historical fault; determining a target test example based on a fault repair situation for a historical fault scenario obtained using a fault simulation strategy in at least one predetermined test example; performing fault repair on the first target device based on the target processing strategy of the target test example; The predetermined test example is obtained based on the following operations: generating a plurality of initial fault simulation strategies based on a plurality of historical fault scenarios indicated by the plurality of historical fault chains; For any of multiple initial fault simulation strategies: Determining a plurality of initial simulation processing strategies that match the initial fault simulation strategy; Based on the multiple initial simulation processing strategies, fault repair is sequentially performed for historical fault scenarios obtained using the initial fault simulation strategies to obtain multiple initial simulation results, wherein the initial simulation results indicate fault repair status for the historical fault scenarios; In a case where it is determined that at least one simulation result among the plurality of initial simulation results indicates that the fault for the historical fault scenario has been repaired, determining the initial fault simulation strategy as the fault simulation strategy; Based on at least one of the simulation results, the simulation processing strategy that matches the fault simulation strategy is determined from at least one of the initial simulation processing strategies corresponding to the at least one simulation result.
2. The method according to claim 1, characterized in that The method further comprises: When it is determined that the simulation result does not exist in the multiple initial simulation results, sending a policy update request to the client so that the client updates the initial fault simulation policy and / or the multiple initial simulation processing policies; In response to receiving the updated initial fault simulation strategy and / or the updated initial simulation processing strategy returned by the client, the predetermined test example is determined based on the updated initial fault simulation strategy and / or the updated initial simulation processing strategy.
3. The method according to claim 1, characterized in that The method further comprises: In a case where it is determined that the simulation result does not exist in the plurality of initial simulation results, recording state information of simulating historical fault scenarios based on the initial fault simulation strategy; Based on the state information, updating the initial fault simulation strategy; The predetermined test example is determined based on an updated initial fault simulation strategy.
4. The method according to claim 1, wherein The method further comprises: The plurality of initial simulation results are updated based on sensor data of the environment in which the first target device is located acquired at predetermined time intervals.
5. The method according to claim 4, characterized in that The method further comprises: In response to the plurality of initial simulation results being updated, the predetermined test case is updated based on the updated plurality of initial simulation results.
6. The method according to claim 4, characterized in that The sensor data includes temperature; The method further comprises: assigning a weight value to any one of the plurality of predetermined test examples based on initial simulation results corresponding to the plurality of predetermined test examples; When it is determined that the difference between the temperatures acquired at two adjacent predetermined time intervals is greater than or equal to a predetermined difference, determining an associated test example associated with the temperature from the plurality of predetermined test examples; The weight value of the association test example is updated.
7. The method according to claim 5, characterized in that The first target device includes a plurality of components, and the historical failure scenario simulated in any one of the predetermined test examples is for an associated component associated with the historical failure scenario among the plurality of components; The method further comprises: determining an actual loss percentage of an associated component corresponding to any one of the predetermined test cases based on the number of simulations for any one of the predetermined test cases; When the actual loss percentage of the associated component corresponding to any one of the predetermined test examples is greater than a predetermined threshold, the predetermined time interval is updated.
8. The method according to claim 1, characterized in that The first target device includes a plurality of components; The method further comprises: Based on causal reasoning of the plurality of historical fault notifications, a fault causal graph is constructed, wherein nodes of the fault causal graph indicate respective historical fault scenarios of a plurality of associated components associated with the plurality of historical fault notifications among the plurality of components, and edges of the fault causal graph indicate causal relationships between adjacent historical fault scenarios; Based on the causal strength of the causal relationship and the fault causal graph, a plurality of the historical fault chains are determined.
9. The method according to claim 8, characterized in that The fault information includes a fault chain; In response to a first fault notification for a first target device, determining fault information based on the first fault notification includes: Performing causal reasoning on the first fault notification to obtain at least one fault scenario causing the first fault notification; Determining, from the fault cause-effect graph, historical fault scenarios that match at least one of the fault scenarios; The fault chain is determined based on other historical fault scenarios that have a causal relationship with each of the at least one fault scenario and the causal strength of the causal relationship.
10. The method according to claim 1, characterized in that The method further comprises: In response to a second fault notification from a second target device, when it is determined that there is a predetermined test example that complies with preset rules among multiple predetermined test examples, fault repair is performed on the second target device based on a simulation processing strategy of the predetermined test example that complies with the preset rules, and the configuration of the second target device is different from that of the first target device.
11. An electronic device comprising: one or more processors; a memory for storing one or more computer programs, It is characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 10.
12. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.
13. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Power system fault automatic repair method and device based on multi-station cooperation
CN112366694A
Fault processing method, fault processing device, electronic equipment and storage medium
CN115190008A