System fault adaptive emergency disposal method and device, medium and computer equipment
By dynamically determining the fault level and type through the collection of multi-source system data, matching the optimal handling strategy and making adaptive adjustments, the problem of time-consuming and labor-intensive manual handling methods is solved, and efficient and accurate handling of server faults is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BAIHANG CREDIT INFORMATION CO LTD
- Filing Date
- 2026-04-29
- Publication Date
- 2026-05-29
AI Technical Summary
In existing technologies, server fault handling relies on manual determination of methods, which is time-consuming, labor-intensive, and ineffective. It is also limited by the uneven technical level of staff, which affects the effectiveness of fault handling.
By collecting multi-source system data from the target service system, including fault identification feature data, current load data, and fault propagation feature data, the fault level and type are dynamically determined, the optimal handling strategy is matched, and the indicator values are monitored in a rotating manner within the fault handling observation window to automatically upgrade the handling strategy to achieve adaptive adjustment.
It improves the accuracy and efficiency of fault handling, ensures the pertinence and effectiveness of fault handling, avoids delays in handling due to unreasonable observation window settings, and shortens fault recovery time.
Smart Images

Figure CN122111740A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of information technology, and in particular to a system failure adaptive emergency response method, apparatus, medium, and computer equipment. Background Technology
[0002] Modern IT systems are growing in scale, and the dependencies between components are becoming increasingly complex. As the core infrastructure supporting business applications, the stability of servers directly affects business continuity and user experience. Therefore, when servers encounter failures, taking effective measures to handle the faults becomes particularly important.
[0003] Currently, in the process of troubleshooting server systems, the single fault handling method is usually determined manually. However, manually determining the fault handling method is time-consuming and labor-intensive, and due to the varying technical skills of staff or negligence, the fault handling method may be determined inappropriately, thus affecting the effectiveness of fault handling for the service system. Summary of the Invention
[0004] This invention provides a system fault adaptive emergency response method, device, medium, and computer equipment, which mainly improves the fault handling effect and efficiency of the service system.
[0005] According to a first aspect of the present invention, a system fault adaptive emergency response method is provided, comprising: In response to a fault handling signal from the target service system, multi-source system data of the target service system is collected, wherein the multi-source system data includes fault identification feature data, current load data, and fault propagation feature data; Based on the fault identification feature data, the current fault level and fault type of the target service system are determined, and based on the current fault level and fault type, a current fault handling strategy for the target service system is determined, and the target service system is handled using the current fault handling strategy; Based on the fault type, obtain historical similar fault characteristic data of the target service system, and dynamically determine the fault handling observation window based on the historical similar fault characteristic data, the current load data, and the fault propagation characteristic data; Within the fault handling observation window, the fault indicator values during the fault handling process of the target service system are polled. Based on the polled fault indicator values, it is determined whether the current fault handling strategy needs to be upgraded. If so, the current fault handling strategy is upgraded to a target-level fault handling strategy, and the target-level fault handling strategy is used to handle the fault of the target service system. Otherwise, the current fault handling strategy is used to continue to handle the fault of the target service system.
[0006] Optionally, determining the current fault level of the target service system based on the fault identification feature data includes: Extract key feature fields related to system fault identification from the fault identification feature data, wherein the key feature fields include fault index values. Fault duration The number of nodes affected by the fault in the target service system and the total number of system nodes ; Dynamically determine the deterioration weight coefficient of each indicator. Time duration weighting coefficient and the weighting coefficient of the scope of influence ; Based on the deterioration weighting coefficient of the aforementioned indicator The time duration weighting coefficient The influence range weighting coefficient The fault index value The duration of the fault The number of nodes affected by the fault and the total number of nodes in the system Determine the severity of the fault in the target service system. , ,in, For historical baseline index values, This refers to the maximum duration of the fault. Based on the severity of the fault Determine the current fault level of the target service system.
[0007] Optionally, the deterioration weighting coefficient of each indicator can be dynamically determined. Time duration weighting coefficient and the weighting coefficient of the scope of influence ,include: Determine the fault index value The relevant indicator is the duration of the fault. Rate of change of indicators within and maximum index value Based on the fault index value The rate of change of the aforementioned indicator The historical baseline index value and the maximum index value Determine the weighting coefficient for the deterioration of the indicator. , ,in, and These correspond to the indicator adjustment coefficient and the rate of change adjustment coefficient; Determine the baseline time of failure Based on the duration of the fault and the fault duration reference time Determine the time duration weighting coefficient , ,in, For time adjustment factor; Determine the system service priority corresponding to the node affected by the failure. Based on the system service priority and the number of nodes affected by the fault Determine the weighting coefficients for the scope of influence , ,in, This refers to the adjustment factor for nodes affected by faults.
[0008] Optionally, the fault identification feature data includes hardware performance index data, software operation index data, and network connection index data; Determining the current fault level of the target service system based on the fault identification feature data includes: Determine the hardware performance feature vector corresponding to the hardware performance index data, the software operation feature vector corresponding to the software operation index data, and the network connection feature vector corresponding to the network connection index data, and determine the vector length of the hardware performance feature vector, the software operation feature vector, and the network connection feature vector, respectively. Based on the current load data of the target service system, the initial weights corresponding to the hardware performance feature vector, the software operation feature vector, and the network connection feature vector are determined respectively. The initial weights are adjusted based on the vector length to obtain the weight coefficients of the hardware performance feature vector, the software operation feature vector, and the network connection feature vector. The hardware performance feature vector, the software operation feature vector, and the network connection feature vector are weighted and fused based on the weight coefficients to obtain a fused feature vector. The current fault level of the target service system is determined based on the fused feature vector.
[0009] Optionally, a fault handling observation window is dynamically determined based on the historical similar fault characteristic data, the current load data, and the fault propagation characteristic data, including: The information uncertainty of determining the historical similar fault recovery time distribution of the target service system based on the historical similar fault characteristic data. The load sensitivity coefficient of the target service system is determined based on the current load data. Based on the fault propagation characteristic data, determine the degree of impact of the current fault of the target service system on the critical path of the system. ; Information uncertainty based on the historical recovery time distribution of similar faults The load sensitivity coefficient and the degree of impact of the current fault on the system's critical path. Dynamically determine the fault handling observation window , ,in, Basic fault handling observation window This is a random disturbance term.
[0010] Optionally, the historical similar fault characteristic data includes the recovery time intervals of multiple historical similar faults, and the fault propagation characteristic data includes the set of critical system paths of the target service system and the set of total system paths affected by the current fault; The information uncertainty of determining the historical similar fault recovery time distribution of the target service system based on the historical similar fault characteristic data. ,include: Different time intervals are determined based on the recovery time intervals of multiple historical similar faults. Probability of internal fault recovery Based on different time intervals Probability of internal fault recovery Information uncertainty in determining the recovery time distribution of similar historical faults , ,in, This represents the total number of items in the time interval. Determine the load sensitivity coefficient of the target service system based on the current load data. ,include: Determine the length of time elapsed from the time the target service system malfunctioned to the current moment. And determine the maximum load that the target service system can withstand under normal operating conditions. Based on the current load data The time length and the maximum load Determine the load sensitivity coefficient of the target service system. , ,in, This is the time decay factor; Based on the fault propagation characteristic data, determine the degree of impact of the current fault of the target service system on the system's critical path. ,include: Based on the set of critical system paths and the total set of system paths affected by the current fault, the number of critical system paths affected by the current fault is determined. The ratio of the number of critical system paths affected by the current fault to the total number of critical system paths in the set of critical system paths is used as the degree of impact of the current fault on the system's critical paths. .
[0011] Optionally, before upgrading the current fault handling strategy to a target-level fault handling strategy, the method further includes: The method for determining the adjustment range of the current fault level based on the patrol fault index value includes: determining whether the patrol fault index value is less than or equal to the fault index value in the fault identification feature data; if so, setting the adjustment range to a level one adjustment range; otherwise, determining the adjustment range based on the difference between the patrol fault index value and the fault index value in the fault identification feature data. Based on the aforementioned adjustment range, the current fault level is adjusted to the target fault level; The target-level fault handling strategy is determined based on the target fault level and the fault type.
[0012] Optionally, the patrol failure index values include the patrol failure duration and the number of nodes affected by the patrol failure; The adjustment range for the current fault level is determined based on the patrol fault index value, including: Determine whether the number of nodes affected by the round-robin fault is less than a first preset threshold. If so, determine the adjustment range of the current fault level based on the duration of the round-robin fault; otherwise, determine the adjustment range of the current fault level based on the number of nodes affected by the round-robin fault. The method for determining the adjustment range of the current fault level based on the number of nodes affected by the round-robin fault includes: If the number of nodes affected by the round-trip fault is greater than a second preset number threshold, the adjustment range of the current fault level is set to the highest level adjustment range. Otherwise, the adjustment range of the current fault level is determined based on the number range to which the number of nodes affected by the round-trip fault belongs, and the second preset number threshold is greater than the first preset number threshold.
[0013] Optionally, the target service system is fault-handled using the current fault-handling strategy, including: The execution complexity of the current fault handling strategy is determined, wherein the method for determining the execution complexity of the current fault handling strategy includes: determining the execution process information, resource dependency information and execution environment constraint information of the current fault handling strategy, and determining the execution complexity of the current fault handling strategy based on the execution process information, resource dependency information and execution environment constraint information; Based on the execution complexity, a target execution method matching the current fault handling strategy is determined, and the current fault handling strategy is executed in the target service system using the target execution method to achieve fault handling of the target service system.
[0014] Optionally, determining the target execution method matching the current fault handling strategy based on the execution complexity includes: If the execution complexity is less than a preset complexity threshold, then the target execution method matched by the current fault handling strategy is the built-in execution method of the target service system API; If the execution complexity is greater than or equal to the preset complexity threshold, then the target execution method matched by the current fault handling strategy is the execution method issued by the custom script.
[0015] Optionally, executing the current fault handling strategy using the target execution method in the target service system includes: Obtain the fault handling instruction template and the system attribute information of the target service system, parse the strategy parameters from the current fault handling strategy, and fill the strategy parameters and system attribute information into the corresponding variable positions in the fault handling instruction template to generate the current fault handling instruction package; The current fault handling instruction package is sent to the target service system, and the current fault handling instruction package is executed in the target service system using the target execution method to realize the fault handling of the target service system; Before executing the current fault handling instruction package using the target execution method in the target service system, the method further includes: Obtain the environmental status information of the target service system, wherein the environmental status information includes the remaining disk capacity, the running status information of dependent services, and the operation permission information of the user terminal executing the current fault handling instruction package; Based on the environmental state information, it is determined whether the current fault handling instruction package meets the preset execution constraints. If so, the current fault handling instruction package is executed in the target service system using the target execution method; otherwise, the execution of the current fault handling instruction package is blocked.
[0016] According to a second aspect of the present invention, a system fault adaptive emergency response device is provided, comprising: The acquisition unit is used to acquire multi-source system data of the target service system in response to the fault handling signal of the target service system, wherein the multi-source system data includes fault identification feature data, current load data, and fault propagation feature data; The fault handling unit is used to determine the current fault level and fault type of the target service system based on the fault identification feature data, and to determine the current fault handling strategy for the target service system based on the current fault level and fault type, and to handle the fault of the target service system using the current fault handling strategy; The window determination unit is used to obtain historical similar fault feature data of the target service system based on the fault type, and dynamically determine the fault handling observation window based on the historical similar fault feature data, the current load data and the fault propagation feature data. The strategy upgrade unit is used to poll the fault indicator values during the fault handling process of the target service system within the fault handling observation window, and determine whether the current fault handling strategy needs to be upgraded based on the polled fault indicator values. If so, the current fault handling strategy is upgraded to a target-level fault handling strategy, and the target-level fault handling strategy is used to handle the fault of the target service system. Otherwise, the current fault handling strategy is used to continue to handle the fault of the target service system.
[0017] According to a third aspect of the present invention, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the above-described adaptive emergency response method for system failures.
[0018] According to a fourth aspect of the present invention, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the above-described adaptive emergency response method for system failures.
[0019] The present invention provides a system fault adaptive emergency response method, apparatus, medium, and computer equipment. Compared with the current method of manually determining a single fault response method for server systems, the present invention comprehensively analyzes system operating status information by collecting multi-source system data from the target service system, thereby improving the accuracy of system fault determination. It matches the optimal fault response strategy based on the current fault level and type, improving the targeting and effectiveness of fault response. By comprehensively considering factors such as historical similar faults, fault propagation, and current load, the observation window can be dynamically adjusted according to the real-time status and fault characteristics of the system. This flexibility allows for the rapid determination of the most suitable observation window when a fault occurs, timely detection of key changes in the fault, and provides a guarantee for rapid fault response, avoiding delays in response due to unreasonable observation window settings. During the fault response process, by cyclically monitoring fault indicator values within the fault response observation window, the response effect can be dynamically evaluated. When the current response strategy is found to be ineffective, the strategy can be automatically upgraded to a target-level fault response strategy, achieving adaptive adjustment of the fault response strategy and improving the success rate and efficiency of fault response. Attached Figure Description
[0020] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings: Figure 1 A flowchart of a system fault adaptive emergency response method provided by an embodiment of the present invention is shown; Figure 2 This invention provides a flowchart of another system fault adaptive emergency response method according to an embodiment of the invention. Figure 3 This diagram illustrates the structure of a system fault adaptive emergency response device provided in an embodiment of the present invention. Figure 4 This invention provides a schematic diagram of the structure of another system fault adaptive emergency response device. Figure 5 A schematic diagram of the physical structure of a computer device provided in an embodiment of the present invention is shown. Detailed Implementation
[0021] The present invention will be described in detail below with reference to the accompanying drawings and embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in the present application can be combined with each other.
[0022] Currently, manually determining a single fault handling method for the server system is time-consuming and labor-intensive. Furthermore, due to varying skill levels or oversights among staff, the determination of the fault handling method may be unreasonable, thus affecting the effectiveness of fault handling for the service system.
[0023] To address the aforementioned problems, embodiments of the present invention provide a system fault adaptive emergency response method, such as... Figure 1 As shown, the method includes: 101. In response to the fault handling signal of the target service system, collect multi-source system data of the target service system, including fault identification feature data, current load data, and fault propagation feature data.
[0024] The fault identification characteristic data includes unstructured business logs (such as application logs, virtual machine stack information, load balancing logs, etc.), system operation logs, and server operation indicators. Server operation indicators include, but are not limited to, CPU temperature and frequency, average CPU load, memory utilization, disk utilization, disk I / O throughput, network bandwidth, and number of transmission control protocols. Current load data refers to the resource usage and business processing pressure of the target service system at the current moment, including hardware resource load data such as network bandwidth utilization and disk I / O utilization, and software business load data such as the number of users or clients connected to the target service system and the number of requests processed by the system. Fault propagation characteristic data includes, but is not limited to, the set of critical system paths and the set of total system paths affected by the current fault. The set of critical system paths refers to the set of paths in the target service system that play a crucial role in the normal operation of the system and the implementation of core business. These paths cover multiple aspects such as hardware connection paths, software call paths, and data transmission paths. The set of total system paths affected by the current fault refers to the set of all system paths directly or indirectly affected by the current fault in the target service system. It includes not only paths that cannot operate normally directly due to the fault, but also paths that are indirectly affected by the fault propagation.
[0025] In this embodiment of the invention, when the monitoring module in the system detects faults such as abnormal stoppage of critical services or performance indicators exceeding thresholds, it immediately issues a fault handling signal. Upon receiving a fault handling signal for the target service system, it uses hardware monitoring tools, such as hardware sensors built into the server, to collect hardware operating status information. For example, it obtains data such as CPU temperature and frequency through CPU sensors and collects log files in real time using a lightweight log collector. It uses system performance monitoring tools, such as Task Manager, to collect hardware resource load data such as CPU utilization, memory utilization, disk I / O utilization, and network bandwidth utilization. It directly obtains software business load data such as concurrent connections, request processing rate, and transaction volume from within the system. Based on the system architecture design document and actual operation, it identifies paths crucial to the operation of the system's core functions, including hardware connection paths (such as network links between servers), software call paths (such as the calling order of modules in the business process), and data transmission paths (such as the data interaction path between the database and the application), forming a set of critical system paths. It uses fault injection and simulated propagation methods, combined with the system topology and dependencies, to analyze the paths that the current fault may affect. Thus, multi-source system data of the service system can be obtained through the above methods. The fault level of the service system is then determined through a comprehensive analysis of the aforementioned multi-source system data. This embodiment of the invention improves the accuracy of system fault determination by comprehensively analyzing system operating status information.
[0026] 102. Determine the current fault level and fault type of the target service system based on fault identification feature data, and determine the current fault handling strategy for the target service system based on the current fault level and fault type, and use the current fault handling strategy to handle the fault of the target service system.
[0027] In this embodiment of the invention, multi-source system data is cleaned and standardized, including but not limited to outlier removal, missing value filling, and unified timestamp formatting, to ensure the accuracy of fault duration calculation. Then, key feature fields such as fault index values, fault duration, number of affected nodes, and total number of service system nodes are extracted from the fault identification feature data using regular expressions. Further, the fault type of the target service system is matched against a preset fault type configuration table based on the key feature fields. This table records the fault types corresponding to various key feature fields. For example, performance bottleneck fault type: average CPU load > core count × 0.8, and memory utilization > 90%; resource exhaustion fault type: disk utilization > 95%, or disk I / O throughput consistently 0; network anomaly fault type: network bandwidth utilization > 90%, accompanied by inter-node communication timeouts; cascading fault type: a single node failure triggers errors in multiple dependent nodes. If the current key feature field values are: CPU utilization 92%, duration 8 minutes, number of affected nodes 3, total number of nodes 10, CPU utilization > 80%, and duration > 5 minutes, then the performance bottleneck fault type is triggered. Meanwhile, based on key feature fields, the fault level of the service system is matched against a preset fault level configuration table. This table stores fault levels corresponding to different key feature fields. For example, if multiple core indicators (such as CPU + memory) simultaneously exceed their corresponding thresholds, and the fault duration is greater than 10 minutes, it is determined to be a Level 1 fault, i.e., a severe fault. If a single indicator briefly exceeds its threshold, or if the affected node data represents a small proportion of the total number of nodes, it is determined to be a Level 3 fault, i.e., a lower-level fault. This embodiment of the invention, by extracting key feature fields related to system fault identification from fault identification feature data, helps to quickly locate the cause of the fault and reduce troubleshooting time.
[0028] In this embodiment of the invention, optionally, the current fault level of the target service system can also be predicted using a preset fault level prediction model. Based on this, the method includes: inputting fault index values, fault duration, number of affected nodes, and total number of system nodes into the preset fault level prediction model for fault level prediction, thereby obtaining the fault level of the target service system. To improve the prediction accuracy of the preset fault level prediction model, it is first necessary to train and construct the preset fault level prediction model. Based on this, the method includes: constructing a preset initial fault level prediction model; obtaining a sample dataset, wherein the sample dataset includes fault index values, fault duration, number of affected nodes, and total number of system nodes of sample servers with fault level labeling information; dividing the sample dataset into a training set and a test set; training the preset initial fault level prediction model using the training set; and testing the trained preset initial fault level prediction model using the test set; finally, using the trained preset initial fault level prediction model that meets the test conditions as the preset fault level prediction model.
[0029] Specifically, during model training, a pre-defined initial level prediction model is first constructed, followed by the acquisition of a sample dataset. The dataset must contain all necessary files. The data is then converted to a format understandable by the pre-defined initial level prediction model. Finally, the model is trained and tested. Specifically, the dataset can be divided first: using random or specific strategies (such as stratified sampling), the sample dataset is divided into training and testing sets. The training set is then used to train the model, and the testing set is used to test the trained model and evaluate its performance on unseen data. Precision, recall, and other metrics on the testing set are calculated and recorded. If the model performance does not meet the requirements, it can return to the training phase for further iterations or adjustments. This process yields a pre-defined level prediction model that meets the requirements. This pre-defined level prediction model may include an input layer, a feature extraction layer, and a level prediction layer. When predicting the fault level, the fault index value, fault duration, number of affected nodes, and total number of system nodes are first input into the feature extraction layer through the input layer for feature extraction. The features output from the feature extraction layer are then input into the level prediction layer for level prediction, resulting in the current fault level of the target service system.
[0030] Furthermore, a pre-defined fault handling strategy library is obtained, which stores different fault handling strategies corresponding to different fault levels and types. For example, the fault handling strategy corresponding to a Level 1 fault (severe) and a performance bottleneck fault type is: immediately trigger load balancing switching, divert traffic to the backup node, and simultaneously start core node expansion; the fault handling strategy corresponding to a Level 1 fault (severe) and a resource exhaustion fault type is: automatically release resources occupied by non-critical processes and trigger resource expansion; the fault handling strategy corresponding to a Level 1 fault (severe) and a cascading fault type is: isolate the faulty node, terminate its communication with dependent nodes to prevent fault propagation, and start backup link recovery services. The fault handling strategy corresponding to a Level 2 fault (important) and a single metric abnormality fault type is: restrict non-critical business requests and prioritize core services; push alarms to maintenance personnel and require manual confirmation; the fault handling strategy corresponding to a Level 2 fault (important) and a local network congestion fault type is: adjust the network bandwidth allocation strategy and prioritize critical data transmission. Based on the real-time determined fault level and type, the optimal handling strategy is matched from the strategy library. For example, if the fault level is level one and the fault type is "performance bottleneck," then the fault handling strategy of "load balancing switchover + core node expansion" is selected. Furthermore, the target service system is handled according to the determined fault handling strategy. This embodiment of the invention matches handling strategies through a dual dimension of fault level and fault type, avoiding the use of uniform fixed strategies, such as directly restarting the service for minor faults, leading to business interruption, or "over-defensive" handling, such as triggering full-link degradation for local network jitter, thereby improving the targeting and effectiveness of fault handling.
[0031] In another embodiment of the present invention, optionally, the fault identification feature data includes hardware performance index data, software operation index data, and network connection index data. Based on this, the current fault level of the target service system can also be determined in the following manner: The hardware performance feature vector corresponding to the hardware performance index data, the software operation feature vector corresponding to the software operation index data, and the network connection feature vector corresponding to the network connection index data are determined respectively, and the vector lengths of the hardware performance feature vector, the software operation feature vector, and the network connection feature vector are determined respectively; Initial weights corresponding to the hardware performance feature vector, the software operation feature vector, and the network connection feature vector are determined based on the current load data of the target service system; the initial weights are adjusted based on the vector lengths to obtain weight coefficients for the hardware performance feature vector, the software operation feature vector, and the network connection feature vector; the hardware performance feature vector, the software operation feature vector, and the network connection feature vector are weighted and fused based on the weight coefficients to obtain a fused feature vector; and the current fault level of the target service system is determined based on the fused feature vector.
[0032] The hardware performance metrics include, but are not limited to, the CPU temperature, fan speed, and hard drive health status of the service system; the software operation metrics include, but are not limited to, the process response time, service availability, and log error frequency of the service system; and the network connectivity metrics include, but are not limited to, the network bandwidth utilization, packet loss rate, and network latency of the service system.
[0033] Specifically, firstly, the hardware performance metrics, software performance metrics, and network connectivity metrics are standardized. Then, feature extraction models, such as CNN models, are used to extract feature vectors corresponding to the hardware performance metrics, software performance metrics, and network connectivity metrics, respectively. Next, the vector magnitudes of the hardware performance feature vector, software performance feature vector, and network connectivity feature vector are determined as the vector lengths. The relative proportions of the vector magnitudes are then determined according to the following formula:
[0034]
[0035]
[0036] in, , , These are hardware performance feature vectors. vector magnitude The relative proportions and software operation feature vectors vector magnitude The relative proportions and network connection feature vectors vector magnitude The relative proportion.
[0037] Furthermore, the current load data of the service system is determined, namely the current load level L. The current load level L is a quantitative description of the overall load of the target service system at the current moment, reflecting the system's load status in terms of hardware resource utilization, software operation activity, and network connectivity. For example, the current load level L can represent the saturation of server CPU and memory hardware resources, the activity level of software request processing, and network bandwidth usage. Then, based on the current load level L, the initial weights corresponding to the hardware performance feature vector, software operation feature vector, and network connectivity feature vector are determined. For example, a linear interpolation strategy can be used to determine the initial weights of the corresponding feature vectors. , , The specific formula is shown below:
[0038]
[0039]
[0040] in, This is a weight adjustment factor set according to actual needs. Further, the initial weights are adjusted based on the relative proportion of the vector magnitude, specifically according to the following formula:
[0041]
[0042]
[0043] in, The adjusted weighting coefficients for the hardware performance feature vector. The weighting coefficients for adjusting the feature vectors of software operation. These are the adjusted weight coefficients for the network connection feature vectors. These are adjustment coefficients set according to actual needs. Further, the various adjustment weight coefficients are normalized, i.e., the sum of all adjustment weight coefficients is determined, and each adjustment weight coefficient is divided by the sum of the adjustment weight coefficients to obtain the weight coefficients for the hardware performance feature vector, software operation feature vector, and network connection feature vector. Further, the hardware performance feature vector, software operation feature vector, and network connection feature vector are multiplied by their respective weight coefficients, and then the three weighted vectors are summed to obtain the fused feature vector. Pre-defined ranges of values for the fused feature vector corresponding to different fault levels are established. The calculated fused feature vector is compared with these preset ranges to determine the current fault level of the target service system. For example, if the value of the fused feature vector falls within the range of minor faults, the target service system is determined to be at a minor fault level (the lowest level fault, such as level four).
[0044] This invention, by considering indicators from three aspects—hardware performance, software operation, and network connectivity—and integrating them into a comprehensive feature vector, can fully and accurately reflect the overall operating status of the target service system, avoiding the one-sidedness of single-indicator evaluation. Simultaneously, this invention determines initial weights based on current load data and adjusts these weights according to the vector length, making the weight allocation more scientific and reasonable. This allows for dynamic adjustment of the importance of each indicator based on the actual system situation, improving the accuracy of fault assessment.
[0045] 103. Obtain historical similar fault characteristic data of the target service system based on fault type, and dynamically determine the fault handling observation window based on historical similar fault characteristic data, current load data and fault propagation characteristic data.
[0046] In this embodiment of the invention, historical fault records of the same type are retrieved from the system's historical fault database, and relevant feature data of these historical faults are extracted, including fault occurrence time, duration, and time required to restore normal operation. Then, a comprehensive analysis is performed on historical similar fault feature data, current load data, and fault propagation feature data to dynamically determine the fault handling observation window. Finally, the fault handling status of the service system is observed within this window. This embodiment of the invention dynamically determines the observation window by integrating multiple data sources, avoiding the problems of a fixed observation window being too long or too short. An excessively long observation window may lead to fault deterioration and missing the optimal handling opportunity, while an excessively short observation window may increase system burden and costs due to frequent handling. Therefore, the dynamic observation window of this invention can be flexibly adjusted according to the actual situation to ensure action is taken at the appropriate time. At the same time, reasonably setting the fault handling observation window can avoid unnecessary resource waste. If the observation window is too large, the system needs to invest more resources in fault monitoring and analysis, which may affect the operation of normal business; if the observation window is too small, it may lead to frequent false alarms and unnecessary handling operations. Dynamically determining the observation window can reasonably allocate monitoring resources according to the actual fault situation and system load, improving the utilization efficiency of system resources.
[0047] 104. In the fault handling observation window, poll the fault indicator values during the fault handling process of the target service system. Based on the polled fault indicator values, determine whether the current fault handling strategy needs to be upgraded. If so, upgrade the current fault handling strategy to the target-level fault handling strategy and use the target-level fault handling strategy to handle the fault of the target service system. Otherwise, continue to handle the fault of the target service system using the current fault handling strategy.
[0048] In this embodiment of the invention, the fault handling observation window refers to a preset time interval, such as 3 minutes, after the current fault handling strategy begins execution. Within this observation window, fault indicator values of the target service system are collected in a polling manner at preset time intervals, such as every 5 or 10 seconds (polling fault indicator values). These indicator values may include, but are not limited to, system response latency, error rate, CPU utilization, memory usage, or business throughput. Through this polling mechanism, dynamic data flow during the fault handling process can be obtained. The collected polling fault indicator values are compared and analyzed with a preset strategy upgrade threshold to determine whether to upgrade the handling strategy. For example, if the polling indicator values do not improve or even deteriorate within several consecutive cycles, exceeding a preset safety threshold (e.g., response latency consistently higher than 1000ms), the current fault handling strategy is determined to be ineffective or insufficient. At this time, the strategy upgrade mechanism is triggered to upgrade the currently executing fault handling strategy, such as a "level 1 restart strategy," to a target-level fault handling strategy, such as a "level 2 rollback strategy" or a "level 3 isolation strategy." Then, the target-level fault handling strategy is used to continue fault handling of the target service system. If the index values collected during the round-robin monitoring show an improving trend and are within a safe range, the current strategy is deemed effective, and the current fault handling strategy will continue to be maintained until the fault is recovered or the observation window ends. This embodiment of the invention, by setting an observation window and performing round-robin monitoring, can determine the next action based on the actual effect of fault handling, avoiding wasting time on ineffective strategies, preventing the expansion of the fault's impact, significantly shortening fault recovery time, and improving the accuracy of fault handling.
[0049] The adaptive emergency response method for system faults provided by this invention, compared with the current method of manually determining a single fault response method for server systems, improves the accuracy of fault identification by collecting multi-source system data from the target service system and comprehensively analyzing system operating status information. It matches the optimal fault response strategy based on the current fault level and type, enhancing the targeting and effectiveness of fault response. By comprehensively considering factors such as historical similar faults, fault propagation, and current load, the observation window can be dynamically adjusted according to the real-time status and fault characteristics of the system. This flexibility allows for the rapid determination of the most suitable observation window when a fault occurs, timely detection of key changes in the fault, and provides a guarantee for rapid fault response, avoiding delays due to unreasonable observation window settings. During fault response, by cyclically monitoring fault indicator values within the fault response observation window, the effectiveness of the response can be dynamically evaluated. When the current response strategy is found to be ineffective, the strategy can be automatically upgraded to a target-level fault response strategy, achieving adaptive adjustment of the fault response strategy and improving the success rate and efficiency of fault response.
[0050] Furthermore, to better illustrate the above process of handling system failures, as a refinement and extension of the above embodiments, this invention provides another adaptive emergency response method for system failures, such as... Figure 2 As shown, the method includes: 201. In response to the fault handling signal of the target service system, collect multi-source system data of the target service system, including fault identification feature data, current load data, and fault propagation feature data.
[0051] Specifically, after collecting multi-source system data, the data is aligned according to timestamps. Then, the multi-source system data can be pushed to the database for storage via a message queue. A multi-dimensional indicator vector corresponding to the multi-source system data can be generated, and operations such as matching fault levels and fault types can be performed based on this multi-dimensional indicator vector.
[0052] 202. Determine the current fault level and fault type of the target service system based on fault identification feature data.
[0053] In this embodiment of the invention, to accurately select a fault handling strategy, it is first necessary to determine the current fault level of the system. Based on this, step 202 specifically includes: extracting key feature fields related to system fault identification from the fault identification feature data, wherein the key feature fields include fault index values. Fault duration The number of nodes affected by the fault in the target service system and the total number of system nodes ; Dynamically determine the deterioration weight coefficient of each indicator. Time duration weighting coefficient and the weighting coefficient of the scope of influence Based on the deterioration weighting coefficient of the aforementioned indicator The time duration weighting coefficient The influence range weighting coefficient The fault index value The duration of the fault The number of nodes affected by the fault and the total number of nodes in the system Determine the severity of the fault in the target service system. , ,in, For historical baseline index values, The maximum duration of the fault; based on the severity of the fault. Determine the current fault level of the target service system. The method for dynamically determining the indicator deterioration weight coefficient, time duration weight coefficient, and impact scope weight coefficient includes: determining the fault indicator value. The rate of change of the relevant index during the fault duration t and maximum index value Based on the fault index value The rate of change of the aforementioned indicator The historical baseline index value and the maximum index value Determine the weighting coefficient for the deterioration of the indicator. , ,in, and These correspond to the indicator adjustment coefficient and the rate of change adjustment coefficient; determine the fault duration baseline time. Based on the duration of the fault and the fault duration reference time Determine the time duration weighting coefficient , ,in, As a time adjustment factor; determine the system service priority corresponding to the node affected by the fault. Based on the system service priority and the number of nodes affected by the fault Determine the weighting coefficients for the scope of influence , ,in, This refers to the adjustment factor for nodes affected by faults.
[0054] Among them, the historical baseline index value is the reference value of the index when the target service system is running normally and without faults; the maximum index value is the maximum value reached by the fault index during the fault duration; the fault duration baseline time is a time reference value pre-set based on the system's historical fault data. It is determined based on the statistical analysis of the duration of similar faults in the past and the business's acceptability of the fault duration; the system business priority is the priority level pre-set for the business corresponding to the fault-affected node based on factors such as the importance and scope of the business.
[0055] Specifically, regular expressions are used to identify fault indicators such as CPU utilization and system memory in the fault identification feature data, as well as key feature fields such as fault duration, number of affected nodes, and total number of system nodes. Then, these key feature fields are used to match the fault type of the service system in a preset fault type configuration table, and the current fault level of the service system is determined based on these key feature fields. When determining the current fault level, the indicator deterioration weight coefficient, time duration weight coefficient, and impact range weight coefficient are dynamically set in advance based on parameters such as fault indicator value, fault duration, and number of affected nodes. Then, to facilitate subsequent calculations, these weight coefficients are normalized so that their sum is 1. Finally, the severity of the fault in the target service system is determined based on the fault indicator value, fault duration, number of affected nodes in the target service system, total number of system nodes, indicator deterioration weight coefficient, time duration weight coefficient, and impact range weight coefficient. The current fault level is then determined based on the fault severity; for example, the greater the fault severity, the higher the corresponding current fault level. This invention determines the severity of a fault by comprehensively considering multiple key characteristic fields such as fault indicator values, fault duration, and the number of nodes affected by the fault. This enables a more comprehensive and accurate assessment of the fault level of the target service system, avoiding the limitations of single-indicator evaluation. Furthermore, by dynamically determining the indicator deterioration weighting coefficient, time duration weighting coefficient, and impact range weighting coefficient, this invention allows emergency response strategies to be adjusted in real time based on the actual fault situation.
[0056] 203. Determine the current fault handling strategy for the target service system based on the current fault level and fault type, and use the current fault handling strategy to handle the fault in the target service system.
[0057] In this embodiment of the invention, a fault handling strategy matching the current fault level and fault type is determined from a preset fault handling strategy library as the current fault handling strategy. Then, the current fault handling strategy is used to handle the fault in the target service system. Based on this, step 203 specifically includes: determining the execution complexity of the current fault handling strategy. The method for determining the execution complexity of the current fault handling strategy includes: determining the execution flow information, resource dependency information, and execution environment constraint information of the current fault handling strategy; determining the execution complexity of the current fault handling strategy based on the execution flow information, resource dependency information, and execution environment constraint information; determining the target execution method matching the current fault handling strategy based on the execution complexity; and executing the current fault handling strategy in the target service system using the target execution method to achieve fault handling of the target service system.
[0058] The execution process information refers to the execution steps and sequence of the current fault handling strategy. For example, if the handling strategy is "first isolate the faulty node, then select a new node from the backup node pool for replacement, and finally configure and initialize the new node," these specific steps and their logical relationships are recorded to form the execution process information. Resource dependency information refers to the resources that the handling strategy depends on during execution, including hardware resources (such as servers and storage devices), software resources (such as specific applications and configuration files), and human resources (such as the operating permissions and skill requirements of maintenance personnel). For example, in a node replacement strategy, it is necessary to specify the number of available nodes in the backup node pool and the software version required for configuring the new node, etc., as resource dependencies. Execution environment constraint information refers to the environmental restrictions that the handling strategy is subject to during execution, such as network bandwidth limitations, system compatibility requirements, and security policy restrictions. For example, when configuring a new node, it may be restricted by network firewall rules and can only communicate through specific ports. These are all execution environment constraint information.
[0059] Specifically, the process score is obtained by scoring the execution process information, the resource score by scoring the resource dependency information, and the constraint score by scoring the execution environment constraint information. Weighting coefficients are then determined for each of the process, resource, and constraint scores based on actual needs. These weighted coefficients are then used to weight and sum the process, resource, and constraint scores. The execution complexity of the current fault handling strategy is determined based on the weighted sum. Specifically, in scoring the execution process information, the number and complexity of the execution steps can be considered; a larger number of steps and greater complexity result in a higher process score. Similarly, in scoring resource dependency information, the diversity and difficulty of obtaining resources can be considered; greater diversity and difficulty of obtaining resources result in a higher resource score. In scoring the execution environment constraint information, the strictness of the constraints can be considered; stricter network bandwidth limitations and higher system compatibility requirements result in a higher constraint score. Finally, the larger the weighted sum of all scores, the higher the execution complexity of the current fault handling strategy. Next, the target execution method of the current fault handling strategy needs to be determined based on the execution complexity. Based on this, the method includes: if the execution complexity is less than a preset complexity threshold, the target execution method matched by the current fault handling strategy is the built-in execution method of the target service system API; if the execution complexity is greater than or equal to the preset complexity threshold, the target execution method matched by the current fault handling strategy is the execution method issued by the custom script.
[0060] The preset complexity threshold is set based on actual needs. Specifically, a complexity threshold is preset based on factors such as the target service system's historical fault handling experience, system resource status, and business tolerance. For example, through analysis of numerous past fault handling cases, it was found that when the execution complexity value is within a certain range, using a specific execution method can achieve better handling results. The specific value of this threshold is determined by considering these factors. After determining the execution complexity of the current fault handling strategy, it is compared with the preset complexity threshold. If the execution complexity is less than the preset complexity threshold, it indicates that the fault handling strategy is relatively simple. In this case, the built-in execution method of the target service system's API is selected as the target execution method. For example, the target service system provides API interfaces for network connectivity detection and repair, such as the checkNetworkConnection() interface, which is used to detect whether the network connection between a specified node and a critical server is normal. This is done by sending probe packets and waiting for a response to determine the network status. Because the built-in execution method is a standardized execution process pre-integrated into the system, for simple fault handling strategies, the relevant functional modules can be called quickly and efficiently to complete the handling operation without additional complex configuration and development. When the execution complexity is greater than or equal to a preset complexity threshold, it indicates that the fault handling strategy is relatively complex, and the built-in execution method may not meet its diverse needs. In this case, the execution method of issuing a custom script is selected as the target execution method. By writing a custom script, the execution steps, logic, and operations can be flexibly defined according to the specific requirements of the fault handling strategy, making full use of the system's various resources and functions to achieve precise handling of complex faults. If the built-in execution method of the target service system API is selected, the system directly calls the corresponding API interface and executes the handling operation according to the built-in process. If the execution method of issuing a custom script is selected, the script written is issued to the target service system based on the current fault handling strategy, and the system parses and executes the instructions in the script to complete the fault handling task. In this embodiment of the invention, the method for executing a current fault handling strategy using a target execution method in the target service system includes: obtaining a fault handling instruction template and system attribute information of the target service system; parsing strategy parameters from the current fault handling strategy; filling the strategy parameters and the system attribute information into the corresponding variable positions in the fault handling instruction template to generate a current fault handling instruction package; sending the current fault handling instruction package to the target service system; and executing the current fault handling instruction package using the target execution method in the target service system to achieve fault handling for the target service system.
[0061] Specifically, various fault handling instruction templates are designed and stored in advance based on different types of fault handling scenarios and operational requirements. For example, corresponding instruction templates are designed for different types of faults such as network fault repair, service process restart, and storage space cleanup. These templates are stored in specific formats (such as text files, database records, etc.) for quick retrieval and retrieval when needed. System attribute information of the target service system is collected through system monitoring tools, configuration management databases, etc., including IP address, hostname, service name, port number, and the role of resources in the system (such as primary server, backup server, etc.). At the same time, the current fault handling strategy is analyzed in detail to extract the strategy parameters contained therein. For example, if the current fault handling strategy is "immediately trigger load balancing switch, divert traffic to the backup node, and simultaneously start core node expansion," the parsed strategy parameters include the identifier of the backup node and the specific specifications of the core node expansion (such as the number of CPU cores, memory capacity, etc.); if the current fault handling strategy is "automatically release resources occupied by non-critical processes and trigger resource expansion," the parsed strategy parameters include the identification criteria of non-critical processes and the target value of resource expansion, etc. The parsing process can employ appropriate parsing methods based on the policy description format (such as structured text, specific markup languages, etc.). Then, the parsed policy parameters and acquired system attribute information are filled into the corresponding variable positions in the fault handling instruction template according to the variable correspondence defined in the template. For example, in the load balancing switchover instruction template, the IP address, hostname, and other attribute information of the standby node, along with relevant switchover parameters, are filled in to generate a complete, accurate, and targeted current fault handling instruction package. Based on the target service system's communication method and network architecture, a suitable transmission protocol is selected to send the generated current fault handling instruction package to the target service system. Upon receiving the instruction package, the target service system parses it according to the previously determined target execution method and executes the fault handling operations step by step according to the instruction requirements. For example, if the built-in execution method is used, the system directly calls the relevant API interfaces to complete load balancing switchover or resource release and expansion operations; if a custom script delivery method is used, the corresponding script program is run to implement the fault handling function in the instruction package, ultimately achieving fault repair and performance optimization of the target service system. This invention generates instruction packages by accurately parsing policy parameters and combining them with system attribute information, enabling the handling instructions to closely match the actual situation of the target service system and the specific needs of the current fault. By using pre-designed instruction templates, the standardization and normalization of the fault handling process are ensured. Regardless of the fault scenario, instructions are generated according to a unified template, making the handling steps clear and logically coherent, which is easy for operation and maintenance personnel to understand and operate, and also beneficial for subsequent fault auditing and experience summarization.
[0062] In this embodiment of the invention, in order to handle service system faults smoothly and stably, an environmental pre-check of the service system is required before fault handling. Based on this, the method includes: obtaining environmental status information of the target service system, wherein the environmental status information includes disk remaining capacity, running status information of dependent services, and operation permission information of the user terminal executing the current fault handling instruction package; based on the environmental status information, determining whether the current fault handling instruction package meets preset execution constraints; if so, executing the current fault handling instruction package in the target service system using the target execution method; otherwise, blocking the execution of the current fault handling instruction package.
[0063] Specifically, within the target service system, the remaining capacity information of each disk partition is queried using the system's built-in disk management tools or related monitoring interfaces. Other services that the target service system depends on are identified, and their operational status (such as startup status, running status, network connection status, protocol connection status, etc.) is obtained using corresponding service management commands or interfaces. The executing client of the current fault handling instruction package is determined, and its operational permissions within the target service system are queried; for example, whether the client has read, write, or execute permissions for specific files, directories, or system functions is checked. Simultaneously, constraints for the execution of the current fault handling instruction package are pre-set based on the target service system's business requirements, security policies, and resource limitations. For example, the remaining disk capacity must be greater than a certain value to ensure sufficient storage space during instruction execution; critical dependent services must be in normal operating condition, otherwise the instruction may fail due to missing dependencies; and the executing client must possess specific operational permissions to ensure the security of the handling operation. The obtained environmental status information, including remaining disk capacity, dependent service operational status information, and executing client operational permission information, is then compared one by one with the pre-set execution constraints. If all conditions are met—namely, sufficient remaining disk capacity, normal operation of all critical dependent services, and appropriate permissions for the executing client—then the current fault handling instruction package is deemed to meet the preset execution constraints. If any condition is not met, the execution constraints are deemed not met. If the judgment result indicates that the current fault handling instruction package meets the preset execution constraints, it is executed in the target service system according to the previously determined target execution method to complete the fault handling operation. If it is determined that the current fault handling instruction package does not meet the preset execution constraints, to prevent problems such as handling failure, data loss, or security risks due to unmet environmental requirements, the execution of the current fault handling instruction package is immediately blocked, and relevant blocking information and environmental status information can be recorded for subsequent analysis and processing. This embodiment of the invention improves the reliability and success rate of fault handling by comprehensively considering environmental status information to determine whether to execute the instruction package, ensuring that the target service system can recover from faults faster and more effectively.
[0064] 204. Obtain historical similar fault characteristic data of the target service system based on fault type, and dynamically determine the fault handling observation window based on historical similar fault characteristic data, current load data and fault propagation characteristic data.
[0065] In this embodiment of the invention, historical similar fault characteristic data is obtained. Then, based on this historical similar fault characteristic data and other data, the fault handling observation window needs to be dynamically determined. Therefore, step 204 specifically includes: determining the information uncertainty of the historical similar fault recovery time distribution of the target service system based on the historical similar fault characteristic data. The load sensitivity coefficient of the target service system is determined based on the current load data. Based on the fault propagation characteristic data, determine the degree of impact of the current fault of the target service system on the critical path of the system. Information uncertainty based on the historical recovery time distribution of similar faults The load sensitivity coefficient and the degree of impact of the current fault on the system's critical path. Dynamically determine the fault handling observation window , ,in, Basic fault handling observation window This is a random disturbance term. Among them, the information uncertainty in determining the recovery time distribution of similar historical faults is... Load sensitivity coefficient And the extent of the impact of the current failure on the system's critical path. The methods include: determining different time intervals based on the recovery time intervals of multiple historical similar faults. Probability of internal fault recovery Based on different time intervals Probability of internal fault recovery Information uncertainty in determining the recovery time distribution of similar historical faults , ,in, The total number of time intervals; determine the length of time elapsed from the time the target service system malfunctioned to the current moment. And determine the maximum load that the target service system can withstand under normal operating conditions. Based on the current load data The time length and the maximum load Determine the load sensitivity coefficient of the target service system. , ,in, The time decay factor is used; based on the set of critical system paths and the total set of system paths affected by the current fault, the number of critical system paths affected by the current fault is determined, and the ratio of the number of critical system paths affected by the current fault to the total number of critical system paths in the set of critical system paths is used as the degree of impact of the current fault on the system's critical paths. .
[0066] Specifically, historical fault characteristic data includes recovery time intervals for multiple historical faults of the same type, and fault propagation characteristic data includes the set of critical system paths of the target service system and the set of total system paths affected by the current fault. Load can be any of the following load data: CPU utilization, memory usage, or network bandwidth utilization; the maximum load data and the current load data represent the same type of load. Simultaneously, load sensitivity coefficients can be calculated separately for different loads, and then the average of these sensitivity coefficients is used as the final load sensitivity coefficient for determining the observation window. In determining the impact of the current fault on the system's critical paths, the intersection of the set of critical system paths and the set of total system paths affected by the current fault is pre-determined, and the number of paths in the intersection is taken as the number of critical system paths affected by the current fault. The impact of the current fault on the system's critical paths is then determined based on this number. This invention, by comprehensively considering factors such as the uncertainty of information in the recovery time distribution of similar historical faults, the load sensitivity coefficient corresponding to the current load data, and the degree of impact of the fault on the critical path of the system, can flexibly adjust the fault handling observation window according to the real-time status of the system and the characteristics of the fault. This avoids the problem of a fixed observation window being too long or too short, ensuring that appropriate measures can be taken in a timely manner when a fault occurs, preventing the fault from worsening, reducing the impact time and scope of the fault on the system, and improving the stability and reliability of the system.
[0067] 205. In the fault handling observation window, poll the fault index values of the target service system during the fault handling process. Based on the polled fault index values, determine whether the current fault handling strategy needs to be upgraded. If so, determine the adjustment range of the current fault level based on the polled fault index values.
[0068] 206. Adjust the current fault level to the target fault level based on the level adjustment range.
[0069] Specifically, based on the length of the fault handling observation window and the real-time requirements of the target service system, a reasonable polling cycle is set. For example, if the observation window is 10 minutes, the polling cycle can be set to once every 1 minute. The fault indicators to be polled are determined. These indicators should reflect the status and performance changes of the target service system during fault handling. Polled fault indicators include system response time, error rate, and resource utilization. According to the set polling cycle, monitoring tools or custom scripts are used to poll the target service system and obtain the current values of the selected fault indicators. For example, using the system's built-in performance monitoring tool, the system's CPU utilization and memory utilization are obtained every 1 minute as polled fault indicator values. Furthermore, for each selected key fault indicator, corresponding judgment thresholds are set based on historical data, business needs, and system performance requirements. For example, for system response time, a threshold of 3 seconds is set; when the polled response time exceeds 3 seconds, it indicates a performance problem in the system. For error rate, a threshold of 1% is set; when the error rate exceeds 1%, it indicates that the system still has a fault. The key fault indicator values obtained in each round of inspection are compared with pre-set judgment thresholds. If at least one key fault indicator value exceeds its corresponding threshold, it is determined that the current fault handling strategy needs to be upgraded; otherwise, the current fault handling strategy continues to be used to handle the fault in the target service system. During the upgrade of the current fault handling strategy, the adjustment range of the current fault level needs to be determined in advance. Based on this, the method includes: determining whether the inspected fault indicator value is less than or equal to the fault indicator value in the fault identification feature data; if so, setting the adjustment range to a first-level adjustment range; otherwise, determining the adjustment range based on the difference between the inspected fault indicator value and the fault indicator value in the fault identification feature data.
[0070] Specifically, first, it is determined whether the patrol fault indicator value is less than or equal to the fault indicator value. If so, it indicates that the current fault handling strategy has some effect, meaning the fault has improved, but not significantly. The level adjustment should be set to a single adjustment level, for example, increasing the fault level slightly from level three (intermediate) to level two (intermediate-high). If the patrol fault indicator value is greater than the fault indicator value, the difference between the patrol fault indicator value and the fault indicator value is calculated. Based on the magnitude of the difference, the level adjustment is determined according to pre-set rules. For example, if the difference is within 0-10%, the level adjustment is one level increase; if the difference is within 10%-30%, the level adjustment is two levels increase; and if the difference exceeds 30%, the level adjustment is three levels increase.
[0071] In this embodiment of the invention, optionally, the patrol fault index value includes the patrol fault duration and the number of nodes affected by the patrol fault; the method for determining the level adjustment range of the current fault level based on the patrol fault index value includes: determining the level adjustment range of the current fault level based on the patrol fault index value, including: determining whether the number of nodes affected by the patrol fault is less than a first preset quantity threshold; if so, determining the level adjustment range of the current fault level based on the patrol fault duration; otherwise, determining the level adjustment range of the current fault level based on the number of nodes affected by the patrol fault; wherein, the method for determining the level adjustment range of the current fault level based on the number of nodes affected by the patrol fault includes: determining whether the number of nodes affected by the patrol fault is greater than a second preset quantity threshold; if so, setting the level adjustment range of the current fault level to the highest level level adjustment range; otherwise, determining the level adjustment range of the current fault level based on the quantity range to which the number of nodes affected by the patrol fault belongs, wherein the second preset quantity threshold is greater than the first preset quantity threshold.
[0072] The first and second preset quantity thresholds are set according to actual needs; the polling fault duration refers to the time interval between the occurrence of a fault and the polling time; the polling fault affected nodes refer to the total number of nodes affected by the system fault during the polling period.
[0073] Specifically, if the number of nodes affected by the round-robin fault is less than the first preset threshold, it indicates that the impact of the fault is relatively small. In this case, the adjustment range of the current fault level is mainly determined based on the duration of the round-robin fault. If the duration of the round-robin fault exceeds the preset time threshold (a value set according to actual needs) and the fault continues to worsen, the fault level is upgraded to level one, i.e., the adjustment range is level one. If the number of nodes affected by the round-robin fault is greater than or equal to the first preset threshold, it indicates that the impact of the fault is large. In this case, the adjustment range of the current fault level needs to be determined based on the number of nodes affected by the round-robin fault. If the number of nodes affected by the round-robin fault is greater than the second preset threshold, it means that the impact of the fault is extremely widespread. The adjustment range of the current fault level is directly set to the highest level adjustment range to attract high attention and take rapid countermeasures. If the number of nodes affected by the round-robin fault is not greater than the second preset threshold, the adjustment range of the current fault level is determined according to the number range to which the number of nodes affected by the round-robin fault belongs. Multiple number ranges can be pre-divided, and a corresponding adjustment range can be set for each range. The corresponding adjustment range is determined according to the range in which the actual number of nodes belongs. The embodiments of the present invention determine the adjustment range of the fault level through a hierarchical and multi-condition judgment method, which can more accurately reflect the severity of the fault and its impact on the system. Appropriate measures can be taken in a timely manner based on the reasonable fault level adjustment, thereby improving fault handling efficiency and ensuring the stable operation of the system.
[0074] Finally, based on the determined adjustment range, the current fault level is adjusted to the target fault level. For example, if the current fault level is a general fault, and the calculated adjustment range is to increase it by two levels, then the general fault is adjusted to a severe fault. Detailed records are kept of the time, reason, and adjustment range of the fault level adjustment for subsequent fault analysis and summarization.
[0075] 207. Determine the target-level fault handling strategy based on the target fault level and fault type, upgrade the current fault handling strategy to the target-level fault handling strategy, and use the target-level fault handling strategy to handle the fault in the target service system; otherwise, continue to handle the fault in the target service system using the current fault handling strategy.
[0076] Specifically, based on the adjusted target fault level and fault type, a corresponding target-level fault handling strategy is selected from a pre-defined fault handling strategy library. For example, for a severe network connection fault, the handling strategy of restarting the network device, checking the network configuration, and notifying the network administrator is selected. Then, the current fault handling strategy is upgraded to a target-level fault handling strategy, and the strategy is automatically distributed to the relevant handling nodes. The target service system is then handled according to the target-level fault handling strategy. For example, the corresponding operations are executed sequentially according to the strategy of restarting the network device, checking the network configuration, and notifying the network administrator. During the fault handling process using the target-level fault handling strategy, the status of the target service system and changes in key fault indicator values are continuously monitored. If the handling effect is unsatisfactory, the handling strategy can be further upgraded or adjusted, or other measures can be taken. For example, if the network connection remains unstable after restarting the network device, the network line can be further checked or the network service provider can be contacted. This embodiment of the invention determines the level adjustment range according to specific circumstances, enabling timely adjustments to fault handling based on the actual situation of the fault. This improves the timeliness and accuracy of fault handling and avoids the problems of fault expansion and prolonged system unavailability caused by delayed or inaccurate handling.
[0077] In an embodiment of the present invention, optionally, during the process of handling faults in the target service system using fault handling strategies, the strategy is not executed on all faulty nodes at once, but is executed on one node first. After the execution on that node is successful, fault handling is performed on multiple nodes in batches. If the execution on that node fails, only that node is rolled back and different strategies are tried to avoid full trial and error.
[0078] According to another system fault adaptive emergency response method provided by the present invention, compared with the current method of manually determining a single fault response method for the server system, the present invention can comprehensively analyze the system operation status information by collecting multi-source system data of the target service system, thereby improving the accuracy of system fault determination; it matches the optimal fault response strategy according to the current fault level and fault type, improving the pertinence and effectiveness of fault response; by comprehensively considering factors such as historical similar faults, fault propagation, and current load, the observation window can be dynamically adjusted according to the real-time status of the system and fault characteristics. This flexibility allows for the rapid determination of the most suitable observation window when a fault occurs, timely detection of key changes in the fault, and provides a guarantee for rapid fault response, avoiding delays in response due to unreasonable observation window settings; during the fault response process, by cyclically monitoring fault indicator values within the fault response observation window, the response effect can be dynamically evaluated. When the current response strategy is found to be ineffective, the strategy can be automatically upgraded to a target-level fault response strategy, realizing adaptive adjustment of the fault response strategy and improving the success rate and efficiency of fault response.
[0079] Furthermore, as Figure 1 In specific implementation, embodiments of the present invention provide a system fault adaptive emergency response device, such as... Figure 3 As shown, the device includes: a data acquisition unit 31, a fault handling unit 32, a window determination unit 33, and a strategy upgrade unit 34.
[0080] The acquisition unit 31 can be used to acquire multi-source system data of the target service system in response to the fault handling signal of the target service system. The multi-source system data includes fault identification feature data, current load data, and fault propagation feature data.
[0081] The fault handling unit 32 can be used to determine the current fault level and fault type of the target service system based on the fault identification feature data, and to determine the current fault handling strategy for the target service system based on the current fault level and fault type, and to handle the fault of the target service system using the current fault handling strategy.
[0082] The window determination unit 33 can be used to obtain historical similar fault feature data of the target service system based on the fault type, and dynamically determine the fault handling observation window based on the historical similar fault feature data, the current load data and the fault propagation feature data.
[0083] The strategy upgrade unit 34 can be used to poll fault indicator values during the fault handling process of the target service system within the fault handling observation window, and determine whether the current fault handling strategy needs to be upgraded based on the polled fault indicator values. If so, the current fault handling strategy is upgraded to a target-level fault handling strategy, and the target-level fault handling strategy is used to handle the fault of the target service system. Otherwise, the current fault handling strategy is used to continue to handle the fault of the target service system.
[0084] In specific application scenarios, in order to determine the current fault level, such as Figure 4 As shown, the fault handling unit 32 includes a field extraction module 321 and a first determination module 322.
[0085] The field extraction module 321 can be used to extract key feature fields related to system fault identification from the fault identification feature data, wherein the key feature fields include fault index values. Fault duration The number of nodes affected by the fault in the target service system and the total number of system nodes .
[0086] The first determining module 322 can be used to dynamically determine the indicator deterioration weight coefficients respectively. Time duration weighting coefficient and the weighting coefficient of the scope of influence .
[0087] The first determining module 322 can also be used to determine the index deterioration weight coefficient. The time duration weighting coefficient The influence range weighting coefficient The fault index value The duration of the fault The number of nodes affected by the fault and the total number of nodes in the system Determine the severity of the fault in the target service system. , ,in, For historical baseline index values, This represents the maximum duration of the fault.
[0088] The first determining module 322 can also be used to determine the severity of the fault. Determine the current fault level of the target service system.
[0089] In specific application scenarios, in order to dynamically determine the weighting coefficient for indicator deterioration Time duration weighting coefficient and the weighting coefficient of the scope of influence The first determining module 322 can specifically be used to determine the fault index value. The relevant indicator is the duration of the fault. Rate of change of indicators within and maximum index value Based on the fault index value The rate of change of the aforementioned indicator The historical baseline index value and the maximum index value Determine the weighting coefficient for the deterioration of the indicator. , ,in, and These correspond to the indicator adjustment coefficient and the rate of change adjustment coefficient; determine the fault duration baseline time. Based on the duration of the fault and the fault duration reference time Determine the time duration weighting coefficient , ,in, As a time adjustment factor; determine the system service priority corresponding to the node affected by the fault. Based on the system service priority and the number of nodes affected by the fault Determine the weighting coefficients for the scope of influence , ,in, This refers to the adjustment factor for nodes affected by faults.
[0090] In specific application scenarios, the fault identification feature data includes hardware performance index data, software operation index data, and network connection index data; in order to determine the current fault level of the target service system, the fault handling unit 32 also includes a weight adjustment module 323 and a vector fusion module 324.
[0091] The first determining module 322 can also be used to determine the hardware performance feature vector corresponding to the hardware performance index data, the software operation feature vector corresponding to the software operation index data, and the network connection feature vector corresponding to the network connection index data, and to determine the vector lengths of the hardware performance feature vector, the software operation feature vector, and the network connection feature vector, respectively.
[0092] The weight adjustment module 323 can be used to determine the initial weights corresponding to the hardware performance feature vector, the software operation feature vector, and the network connection feature vector based on the current load data of the target service system, and adjust the initial weights based on the vector length to obtain the weight coefficients of the hardware performance feature vector, the software operation feature vector, and the network connection feature vector.
[0093] The vector fusion module 324 can be used to perform weighted fusion of the hardware performance feature vector, the software operation feature vector and the network connection feature vector based on the weight coefficient to obtain a fused feature vector, and to determine the current fault level of the target service system based on the fused feature vector.
[0094] In specific application scenarios, in order to dynamically determine the fault handling observation window, the window determination unit 33 can be specifically used to determine the information uncertainty of the historical similar fault recovery time distribution of the target service system based on the historical similar fault characteristic data. The load sensitivity coefficient of the target service system is determined based on the current load data. Based on the fault propagation characteristic data, determine the degree of impact of the current fault of the target service system on the critical path of the system. Information uncertainty based on the historical recovery time distribution of similar faults The load sensitivity coefficient and the degree of impact of the current fault on the system's critical path. Dynamically determine the fault handling observation window , ,in, Basic fault handling observation window This is a random disturbance term.
[0095] In specific application scenarios, the historical similar fault characteristic data includes recovery time intervals for multiple historical similar faults, and the fault propagation characteristic data includes the set of critical system paths of the target service system and the set of total system paths affected by the current fault. To determine the information uncertainty, load sensitivity coefficient, and impact of the current fault on the system's critical paths based on the recovery time distribution of historical similar faults, the window determination unit 33 can specifically be used to determine different time intervals based on the recovery time intervals of multiple historical similar faults. Probability of internal fault recovery Based on different time intervals Probability of internal fault recovery Information uncertainty in determining the recovery time distribution of similar historical faults , ,in, The total number of time intervals; determine the length of time elapsed from the time the target service system malfunctioned to the current moment. And determine the maximum load that the target service system can withstand under normal operating conditions. Based on the current load data The time length and the maximum load Determine the load sensitivity coefficient of the target service system. , ,in, The time decay factor is used; based on the set of critical system paths and the total set of system paths affected by the current fault, the number of critical system paths affected by the current fault is determined, and the ratio of the number of critical system paths affected by the current fault to the total number of critical system paths in the set of critical system paths is used as the degree of impact of the current fault on the system's critical paths. .
[0096] In specific application scenarios, in order to upgrade the current fault handling strategy to a target-level fault handling strategy, the strategy upgrade unit 34 includes a second determination module 341, a level adjustment module 342, and a strategy upgrade module 343.
[0097] The second determining module 341 can be used to determine the level adjustment range of the current fault level based on the patrol fault index value. The method for determining the level adjustment range of the current fault level based on the patrol fault index value includes: determining whether the patrol fault index value is less than or equal to the fault index value in the fault identification feature data; if so, setting the level adjustment range to a level one adjustment range; otherwise, determining the level adjustment range based on the difference between the patrol fault index value and the fault index value in the fault identification feature data.
[0098] The level adjustment module 342 can be used to adjust the current fault level to the target fault level based on the level adjustment range.
[0099] The strategy upgrade module 343 can be used to determine the target-level fault handling strategy based on the target fault level and the fault type.
[0100] In specific application scenarios, the patrol fault index values include the patrol fault duration and the number of nodes affected by the patrol fault. To determine the adjustment range of the current fault level, the second determining module 341 can be used to determine whether the number of nodes affected by the patrol fault is less than a first preset quantity threshold. If so, the adjustment range of the current fault level is determined based on the patrol fault duration; otherwise, the adjustment range of the current fault level is determined based on the number of nodes affected by the patrol fault. The method for determining the adjustment range of the current fault level based on the number of nodes affected by the patrol fault includes: determining whether the number of nodes affected by the patrol fault is greater than a second preset quantity threshold. If so, the adjustment range of the current fault level is set to the highest level adjustment range; otherwise, the adjustment range of the current fault level is determined based on the quantity range to which the number of nodes affected by the patrol fault belongs. The second preset quantity threshold is greater than the first preset quantity threshold.
[0101] In specific application scenarios, in order to use the current fault handling strategy to handle faults in the target service system, the first determining module 322 can also be used to determine the execution complexity of the current fault handling strategy. The method for determining the execution complexity of the current fault handling strategy includes: determining the execution process information, resource dependency information, and execution environment constraint information of the current fault handling strategy; determining the execution complexity of the current fault handling strategy based on the execution process information, resource dependency information, and execution environment constraint information; determining the target execution method matched with the current fault handling strategy based on the execution complexity; and executing the current fault handling strategy in the target service system using the target execution method to achieve fault handling of the target service system.
[0102] In specific application scenarios, in order to determine the target execution method matching the current fault handling strategy based on the execution complexity, the first determining module 322 can be specifically used to determine the target execution method matching the current fault handling strategy as the built-in execution method of the target service system API if the execution complexity is less than a preset complexity threshold; and to determine the target execution method matching the current fault handling strategy as the execution method issued by the custom script if the execution complexity is greater than or equal to the preset complexity threshold.
[0103] In a specific application scenario, in order to execute the current fault handling strategy in the target service system using the target execution method, the first determining module 322 can be used to obtain the fault handling instruction template and the system attribute information of the target service system, parse the strategy parameters from the current fault handling strategy, fill the corresponding variable positions in the fault handling instruction template with the strategy parameters and the system attribute information to generate the current fault handling instruction package; send the current fault handling instruction package to the target service system, and execute the current fault handling instruction package in the target service system using the target execution method to realize the fault handling of the target service system.
[0104] In specific application scenarios, in order to pre-check the service system policy execution environment, the first determining module 322 can be used to obtain the environment status information of the target service system. The environment status information includes the remaining disk capacity, the running status information of dependent services, and the operation permission information of the user terminal executing the current fault handling instruction package. Based on the environment status information, it is determined whether the current fault handling instruction package meets the preset execution constraints. If so, the current fault handling instruction package is executed in the target service system using the target execution method; otherwise, the execution of the current fault handling instruction package is blocked.
[0105] It should be noted that other corresponding descriptions of the functional modules involved in the system fault adaptive emergency response device provided in this embodiment of the invention can be found in [reference]. Figure 1 The corresponding description of the method shown will not be repeated here.
[0106] Based on the above, Figure 1The method shown, correspondingly, also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, performs the following steps: responding to a fault handling signal of a target service system, collecting multi-source system data of the target service system, wherein the multi-source system data includes fault identification feature data, current load data, and fault propagation feature data; determining the current fault level and fault type of the target service system based on the fault identification feature data, and determining a current fault handling strategy for the target service system based on the current fault level and fault type, and using the current fault handling strategy to handle the fault in the target service system; based on the... The system acquires historical similar fault characteristic data of the target service system, dynamically determines a fault handling observation window based on the historical similar fault characteristic data, the current load data, and the fault propagation characteristic data, and polls the fault indicator values of the target service system during the fault handling process within the fault handling observation window. Based on the polled fault indicator values, it determines whether the current fault handling strategy needs to be upgraded. If so, the current fault handling strategy is upgraded to a target-level fault handling strategy, and the target-level fault handling strategy is used to handle the fault in the target service system. Otherwise, the current fault handling strategy is used to continue to handle the fault in the target service system.
[0107] Based on the above, Figure 1 The method shown and as Figure 3 The embodiment of the device shown in the invention also provides a physical structure diagram of a computer device, such as... Figure 5As shown, the computer device includes: a processor 41, a memory 42, and a computer program stored in the memory 42 and executable on the processor. Both the memory 42 and the processor 41 are mounted on a bus 43. When the processor 41 executes the program, it performs the following steps: In response to a fault handling signal from the target service system, it collects multi-source system data of the target service system, including fault identification feature data, current load data, and fault propagation feature data; Based on the fault identification feature data, it determines the current fault level and fault type of the target service system, and based on the current fault level and fault type, it determines a current fault handling strategy for the target service system; and uses the current fault handling strategy to... The target service system performs fault handling; based on the fault type, it obtains historical similar fault characteristic data of the target service system, and dynamically determines the fault handling observation window based on the historical similar fault characteristic data, the current load data, and the fault propagation characteristic data; within the fault handling observation window, it polls the fault indicator values of the target service system during the fault handling process, and determines whether the current fault handling strategy needs to be upgraded based on the polled fault indicator values. If so, it upgrades the current fault handling strategy to a target-level fault handling strategy and uses the target-level fault handling strategy to handle the fault in the target service system; otherwise, it continues to handle the fault in the target service system using the current fault handling strategy.
[0108] Through the technical solution of this invention, by collecting multi-source system data from the target service system, this invention can comprehensively analyze the system's operating status information, thereby improving the accuracy of system fault determination; by matching the optimal fault handling strategy according to the current fault level and fault type, it improves the pertinence and effectiveness of fault handling; by comprehensively considering factors such as historical similar faults, fault propagation, and current load, the observation window can be dynamically adjusted according to the real-time status of the system and fault characteristics. This flexibility allows for the rapid determination of the most suitable observation window when a fault occurs, timely detection of key changes in the fault, and provides a guarantee for rapid fault handling, avoiding delays in handling due to unreasonable observation window settings; during the fault handling process, by cyclically monitoring fault indicator values within the fault handling observation window, the handling effect can be dynamically evaluated. When it is found that the current handling strategy is ineffective, the strategy can be automatically upgraded to a target-level fault handling strategy, realizing adaptive adjustment of the fault handling strategy and improving the success rate and efficiency of fault handling.
[0109] It is obvious to those skilled in the art that the modules or steps of the present invention described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. Optionally, they can be implemented using computer-executable program code, thereby storing them in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those presented herein, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.
[0110] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A system fault adaptive emergency response method, characterized in that, include: In response to a fault handling signal from the target service system, multi-source system data of the target service system is collected, wherein the multi-source system data includes fault identification feature data, current load data, and fault propagation feature data; Based on the fault identification feature data, the current fault level and fault type of the target service system are determined, and based on the current fault level and fault type, a current fault handling strategy for the target service system is determined, and the target service system is handled using the current fault handling strategy; Based on the fault type, obtain historical similar fault characteristic data of the target service system, and dynamically determine the fault handling observation window based on the historical similar fault characteristic data, the current load data, and the fault propagation characteristic data; Within the fault handling observation window, the fault indicator values during the fault handling process of the target service system are polled. Based on the polled fault indicator values, it is determined whether the current fault handling strategy needs to be upgraded. If so, the current fault handling strategy is upgraded to a target-level fault handling strategy, and the target-level fault handling strategy is used to handle the fault of the target service system. Otherwise, the current fault handling strategy is used to continue to handle the fault of the target service system.
2. The system fault adaptive emergency response method according to claim 1, characterized in that, Determining the current fault level of the target service system based on the fault identification feature data includes: Extract key feature fields related to system fault identification from the fault identification feature data, wherein the key feature fields include fault index values. Fault duration The number of nodes affected by the fault in the target service system and the total number of system nodes ; Dynamically determine the deterioration weight coefficient of each indicator. Time duration weighting coefficient and the weighting coefficient of the scope of influence ; Based on the deterioration weighting coefficient of the aforementioned indicator The time duration weighting coefficient The influence range weighting coefficient The fault index value The duration of the fault The number of nodes affected by the fault and the total number of nodes in the system Determine the severity of the fault in the target service system. , ,in, For historical baseline index values, This refers to the maximum duration of the fault. Based on the severity of the fault Determine the current fault level of the target service system.
3. The adaptive emergency response method for system faults according to claim 2, characterized in that, Dynamically determine the deterioration weight coefficient of each indicator. Time duration weighting coefficient and the weighting coefficient of the scope of influence ,include: Determine the fault index value The relevant indicator is the duration of the fault. Rate of change of indicators within and maximum index value Based on the fault index value The rate of change of the aforementioned indicator The historical baseline index value and the maximum index value Determine the weighting coefficient for the deterioration of the indicator. , ,in, and These correspond to the indicator adjustment coefficient and the rate of change adjustment coefficient; Determine the baseline time for fault duration Based on the duration of the fault and the fault duration reference time Determine the time duration weighting coefficient , ,in, For time adjustment factor; Determine the system service priority corresponding to the node affected by the failure. Based on the system service priority and the number of nodes affected by the fault Determine the weighting coefficients for the scope of influence. , ,in, This refers to the adjustment factor for nodes affected by faults.
4. The system fault adaptive emergency response method according to claim 1, characterized in that, The fault identification feature data includes hardware performance index data, software operation index data, and network connection index data; Determining the current fault level of the target service system based on the fault identification feature data includes: Determine the hardware performance feature vector corresponding to the hardware performance index data, the software operation feature vector corresponding to the software operation index data, and the network connection feature vector corresponding to the network connection index data, and determine the vector length of the hardware performance feature vector, the software operation feature vector, and the network connection feature vector, respectively. Based on the current load data of the target service system, the initial weights corresponding to the hardware performance feature vector, the software operation feature vector, and the network connection feature vector are determined respectively. The initial weights are adjusted based on the vector length to obtain the weight coefficients of the hardware performance feature vector, the software operation feature vector, and the network connection feature vector. The hardware performance feature vector, the software operation feature vector, and the network connection feature vector are weighted and fused based on the weight coefficients to obtain a fused feature vector. The current fault level of the target service system is determined based on the fused feature vector.
5. The adaptive emergency response method for system faults according to claim 1, characterized in that, Based on the historical similar fault characteristic data, the current load data, and the fault propagation characteristic data, a fault handling observation window is dynamically determined, including: The information uncertainty of determining the historical similar fault recovery time distribution of the target service system based on the historical similar fault characteristic data. The load sensitivity coefficient of the target service system is determined based on the current load data. Based on the fault propagation characteristic data, determine the degree of impact of the current fault of the target service system on the critical path of the system. ; Information uncertainty based on the historical recovery time distribution of similar faults The load sensitivity coefficient and the degree of impact of the current fault on the system's critical path. Dynamically determine the fault handling observation window , ,in, Basic fault handling observation window This is a random disturbance term.
6. The system fault adaptive emergency response method according to claim 5, characterized in that, The historical similar fault characteristic data includes the recovery time intervals of multiple historical similar faults, and the fault propagation characteristic data includes the set of critical system paths of the target service system and the set of total system paths affected by the current fault. The information uncertainty of determining the historical similar fault recovery time distribution of the target service system based on the historical similar fault characteristic data. ,include: Different time intervals are determined based on the recovery time intervals of multiple historical similar faults. Probability of internal fault recovery Based on different time intervals Probability of internal fault recovery Information uncertainty in determining the recovery time distribution of similar historical faults , ,in, This represents the total number of items in the time interval. Determine the load sensitivity coefficient of the target service system based on the current load data. ,include: Determine the length of time elapsed from the time the target service system malfunctioned to the current moment. And determine the maximum load that the target service system can withstand under normal operating conditions. Based on the current load data The time length and the maximum load Determine the load sensitivity coefficient of the target service system. , ,in, This is the time decay factor; Based on the fault propagation characteristic data, determine the degree of impact of the current fault of the target service system on the system's critical path. ,include: Based on the set of critical system paths and the total set of system paths affected by the current fault, the number of critical system paths affected by the current fault is determined. The ratio of the number of critical system paths affected by the current fault to the total number of critical system paths in the set of critical system paths is used as the degree of impact of the current fault on the system's critical paths. .
7. The adaptive emergency response method for system faults according to claim 1, characterized in that, Before upgrading the current fault handling strategy to a target-level fault handling strategy, the method further includes: The method for determining the adjustment range of the current fault level based on the patrol fault index value includes: determining whether the patrol fault index value is less than or equal to the fault index value in the fault identification feature data; if so, setting the adjustment range to a level one adjustment range; otherwise, determining the adjustment range based on the difference between the patrol fault index value and the fault index value in the fault identification feature data. Based on the aforementioned adjustment range, the current fault level is adjusted to the target fault level; The target-level fault handling strategy is determined based on the target fault level and the fault type.
8. The system fault adaptive emergency response method according to claim 7, characterized in that, The patrol failure index values include the patrol failure duration and the number of nodes affected by the patrol failure. The adjustment range for the current fault level is determined based on the patrol fault index value, including: Determine whether the number of nodes affected by the round-robin fault is less than a first preset threshold. If so, determine the adjustment range of the current fault level based on the duration of the round-robin fault; otherwise, determine the adjustment range of the current fault level based on the number of nodes affected by the round-robin fault. The method for determining the adjustment range of the current fault level based on the number of nodes affected by the round-robin fault includes: If the number of nodes affected by the round-trip fault is greater than a second preset number threshold, the adjustment range of the current fault level is set to the highest level adjustment range. Otherwise, the adjustment range of the current fault level is determined based on the number range to which the number of nodes affected by the round-trip fault belongs, and the second preset number threshold is greater than the first preset number threshold.
9. The adaptive emergency response method for system faults according to claim 1, characterized in that, The target service system is handled using the current fault handling strategy, including: The execution complexity of the current fault handling strategy is determined, wherein the method for determining the execution complexity of the current fault handling strategy includes: determining the execution process information, resource dependency information and execution environment constraint information of the current fault handling strategy, and determining the execution complexity of the current fault handling strategy based on the execution process information, resource dependency information and execution environment constraint information; Based on the execution complexity, a target execution method matching the current fault handling strategy is determined, and the current fault handling strategy is executed in the target service system using the target execution method to achieve fault handling of the target service system.
10. The system fault adaptive emergency response method according to claim 9, characterized in that, Determining the target execution method matching the current fault handling strategy based on the execution complexity includes: If the execution complexity is less than a preset complexity threshold, then the target execution method matched by the current fault handling strategy is the built-in execution method of the target service system API; If the execution complexity is greater than or equal to the preset complexity threshold, then the target execution method matched by the current fault handling strategy is the execution method issued by the custom script.
11. The system fault adaptive emergency response method according to claim 9, characterized in that, The current fault handling strategy is executed using the target execution method in the target service system, including: Obtain the fault handling instruction template and the system attribute information of the target service system, parse the strategy parameters from the current fault handling strategy, and fill the strategy parameters and system attribute information into the corresponding variable positions in the fault handling instruction template to generate the current fault handling instruction package; The current fault handling instruction package is sent to the target service system, and the current fault handling instruction package is executed in the target service system using the target execution method to realize the fault handling of the target service system; Before executing the current fault handling instruction package using the target execution method in the target service system, the method further includes: Obtain the environmental status information of the target service system, wherein the environmental status information includes the remaining disk capacity, the running status information of dependent services, and the operation permission information of the user terminal executing the current fault handling instruction package; Based on the environmental state information, it is determined whether the current fault handling instruction package meets the preset execution constraints. If so, the current fault handling instruction package is executed in the target service system using the target execution method; otherwise, the execution of the current fault handling instruction package is blocked.
12. A system fault adaptive emergency response device, characterized in that, include: The acquisition unit is used to acquire multi-source system data of the target service system in response to the fault handling signal of the target service system, wherein the multi-source system data includes fault identification feature data, current load data, and fault propagation feature data; The fault handling unit is used to determine the current fault level and fault type of the target service system based on the fault identification feature data, and to determine the current fault handling strategy for the target service system based on the current fault level and fault type, and to handle the fault of the target service system using the current fault handling strategy; The window determination unit is used to obtain historical similar fault feature data of the target service system based on the fault type, and dynamically determine the fault handling observation window based on the historical similar fault feature data, the current load data and the fault propagation feature data. The strategy upgrade unit is used to poll the fault indicator values during the fault handling process of the target service system within the fault handling observation window, and determine whether the current fault handling strategy needs to be upgraded based on the polled fault indicator values. If so, the current fault handling strategy is upgraded to a target-level fault handling strategy, and the target-level fault handling strategy is used to handle the fault of the target service system. Otherwise, the current fault handling strategy is used to continue to handle the fault of the target service system.
13. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 11.
14. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 11.