Fault loss stopping method and device, electronic equipment and storage medium
By orchestrating the fault stop loss business process corresponding to the fault scenario, automatic fault identification and stop loss are realized, which solves the problem of slow response speed for operation and maintenance personnel and improves the fault stop loss efficiency.
Patent Information
- Application Number
- CN202510373026.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-06-27
AI Technical Summary
During the operation of business services, operation and maintenance personnel are unable to monitor the system status in real time, resulting in slow response speed of failures, affecting stop loss efficiency, and may lead to escalation of failures.
Automatic fault identification and stop loss operation process by orchestrating the corresponding fault scenarios, including data collection tasks, abnormal judgment rules, fault identification rules and fault stop loss actions.
The fault can be quickly discovered and the stop loss action can be performed without the participation of operation and maintenance personnel, which improves the fault detection efficiency and fault stop loss efficiency.
Smart Images

Figure CN120216248A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of fault handling, and particularly to a fault loss prevention method, apparatus, electronic device, and storage medium. Background Art
[0002] During the operation of business services, in order to monitor the operation status of business services, it is often necessary to collect various data or events related to the operation status, and perform fault discovery and handling by analyzing the data or events. In the current system operation and maintenance and fault handling processes, relying on operation and maintenance personnel to make abnormal judgments and stop-loss decisions, that is, the manual stop-loss method is still an important means to deal with abnormal situations. However, since operation and maintenance personnel cannot monitor the system status at any time and place, the response speed is slow, which affects the stop-loss efficiency and leads to fault escalation. Summary of the Invention
[0003] This application provides a fault loss prevention method, apparatus, electronic device, and storage medium, which can improve the fault loss prevention efficiency. The technical solution is as follows:
[0004] According to one aspect of this application, a fault loss prevention method is provided. The method includes:
[0005] Receiving a configuration operation for a fault loss prevention business process, and determining target configuration information corresponding to the fault loss prevention business process, where the target configuration information at least includes a data collection task, an abnormal judgment rule, a fault identification rule for a target fault, and a fault loss prevention action;
[0006] After triggering an execution operation for the fault loss prevention business process, loading the data collection task corresponding to the fault loss prevention business process, and obtaining target collection data from a target device based on the data collection task;
[0007] Performing fault identification on the target collection data based on the abnormal judgment rule and / or the fault identification rule to obtain a fault identification result;
[0008] If the fault identification result indicates the existence of the target fault, automatically executing the fault loss prevention action corresponding to the target fault.
[0009] According to another aspect of this application, a fault loss prevention apparatus is provided. The apparatus includes:
[0010] A receiving module, configured to receive a configuration operation for a fault loss prevention business process, and determine target configuration information corresponding to the fault loss prevention business process, where the target configuration information at least includes a data collection task, an abnormal judgment rule, a fault identification rule for a target fault, and a fault loss prevention action;
[0011] An acquisition module, configured to, after triggering the execution operation of the fault stop-loss business process, load the data acquisition task corresponding to the fault stop-loss business process, and obtain target acquisition data from a target device based on the data acquisition task;
[0012] A fault identification module, configured to perform fault identification on the target acquisition data based on the exception judgment rule and / or the fault identification rule, and obtain a fault identification result;
[0013] A fault stop-loss module, configured to, if the fault identification result indicates the existence of the target fault, automatically execute the fault stop-loss action corresponding to the target fault.
[0014] According to one aspect of the present application, there is provided an electronic device, including: a processor and a memory storing a program, the program including instructions that, when executed by the processor, cause the processor to execute the fault stop-loss method as described above.
[0015] According to another aspect of the present application, there is provided a non-transitory computer-readable storage medium storing computer instructions, the computer instructions being used to cause the computer to execute the fault stop-loss method as described above.
[0016] According to another aspect of the present application, there is provided a computer program product, the computer program product including computer instructions, the computer instructions being stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above-mentioned fault stop-loss method.
[0017] The beneficial effects brought by the technical solution provided by the embodiments of the present application at least include:
[0018] By arranging the target configuration of the fault stop-loss business process corresponding to the fault scenario, including a data acquisition task set, an exception judgment rule, a fault identification rule, a fault stop-loss action, etc., the automatic stop-loss platform can collect target acquisition data from the target device by executing the fault stop-loss business process, and perform fault identification on the target acquisition data through the exception judgment rule and / or the fault identification rule. Furthermore, after identifying the target fault, the corresponding fault stop-loss action is automatically executed, realizing automatic fault stop-loss in different fault scenarios. Moreover, the process from fault discovery to fault stop-loss does not require the participation of operation and maintenance personnel, improving the fault discovery efficiency and fault stop-loss efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In the following description of exemplary embodiments with reference to the accompanying drawings, more details, features, and advantages of the present application are disclosed. In the drawings:
[0020] Figure 1Shows a flowchart of a fault loss prevention method according to an exemplary embodiment of the present application;
[0021] Figure 2 Shows a flowchart of another fault loss prevention method according to an exemplary embodiment of the present application;
[0022] Figure 3 Is a schematic diagram of the processing flow of an abnormal event provided by an exemplary embodiment of the present application;
[0023] Figure 4 Is a flowchart of a complete fault loss prevention method provided by an exemplary embodiment of the present application;
[0024] Figure 5 Is a schematic structural diagram of a fault loss prevention device provided by an embodiment of the present application;
[0025] Figure 6 Shows a structural block diagram of an exemplary electronic device capable of implementing the embodiments of the present application. Detailed implementation manners
[0026] The embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present application. It should be understood that the drawings and embodiments of the present application are only for exemplary purposes and are not used to limit the protection scope of the present application.
[0027] It should be understood that the steps recorded in the method embodiments of the present application can be executed in different orders and / or executed in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present application is not limited in this regard.
[0028] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description. It should be noted that the concepts such as "first", "second", etc. mentioned in this application are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units. It should be noted that the modifications of "one" and "multiple" mentioned in this application are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise clearly specified in the context, it should be understood as "one or more". The names of the messages or information exchanged between multiple devices in the embodiments of this application are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0029] The solutions of this application will be described below with reference to the accompanying drawings. The technical solutions provided by the embodiments of this application will be described in detail through specific embodiments and their application scenarios.
[0030] To avoid the problem of low stop-loss efficiency caused by manual stop-loss, the embodiments of this application achieve automatic fault stop-loss in different fault scenarios by arranging the fault stop-loss business process corresponding to the fault stop-loss scenario and combining multiple execution steps such as data collection, anomaly judgment, fault identification, and stop-loss actions. Please refer to Figure 1 , which shows a flowchart of a fault stop-loss method according to an exemplary embodiment of this application. This method will be described by taking its application to an automatic stop-loss platform as an example. As Figure 1 shown, this method includes:
[0031] Step 101, receive a configuration operation for the fault stop-loss business process, and determine the target configuration information corresponding to the fault stop-loss business process. The target configuration information includes at least a data collection task, an anomaly judgment rule, a fault identification rule for a target fault, and a fault stop-loss action.
[0032] In a possible implementation, for different fault scenarios, the operation and maintenance personnel can orchestrate or configure the fault stop-loss business process corresponding to the fault scenario in the automatic stop-loss platform. The fault stop-loss business process may include: at least one data collection task to be executed, an anomaly judgment rule for the data collected for each data collection task, a fault identification rule for at least one target fault in this fault scenario, and a fault stop-loss action corresponding to the target fault, etc. Correspondingly, when the automatic stop-loss platform receives a configuration operation for the fault stop-loss business process and obtains the target configuration information corresponding to the fault stop-loss business process, the target configuration information includes at least the data collection task, the anomaly judgment rule, the fault identification rule of the target fault, and the fault stop-loss action, so as to generate the corresponding fault stop-loss business process based on the target configuration information.
[0033] Optionally, the target configuration information may further include information such as the group chat, email, and phone number to which the execution result corresponding to the fault stop-loss business process is sent.
[0034] Step 102, after triggering the execution operation of the fault stop-loss business process, load the data collection task corresponding to the fault stop-loss business process, and obtain the target collection data from the target device based on the data collection task.
[0035] After pre-configuring the fault stop-loss business process, the automatic stop-loss platform discovers faults and performs automatic stop-loss by executing the fault stop-loss business process. In a possible implementation, after triggering the execution operation of the fault stop-loss business process, first load the data collection task corresponding to the fault stop-loss business process, and parse the data collection task to obtain the target device indicated by the data collection task, and then obtain the target collection data indicated by the data collection task from the target device.
[0036] Exemplarily, if the data collection task is to collect the CPU occupancy rate of each service instance (or server) corresponding to the target service, the target device is each service instance, and the target collection data is the real-time CPU occupancy rate.
[0037] Step 103, perform fault identification on the target collection data based on the anomaly judgment rule and / or the fault identification rule to obtain a fault identification result.
[0038] After collecting the target collection data based on the data collection task, the target collection data can be fault-identified according to the anomaly judgment rule and / or the fault identification rule to determine whether the target collection data meets the fault characteristics of the target fault. If it meets, it is determined that the fault identification result is that there is a target fault; if it does not meet, it is determined that the fault identification result is that there is no target fault.
[0039] Among them, the anomaly judgment rule is mainly used to judge anomalies in the collected time-series data. For example, for the collected CPU occupancy rate, the anomaly judgment rule is used to judge whether each CPU occupancy rate is abnormal; the fault identification rule is mainly used to further judge whether an abnormal event or an alarm event meets the fault characteristics of the target fault. For example, the target fault is server lag, and the fault characteristics of server lag are abnormal CPU occupancy rate and insufficient memory. That is, only when two abnormal events of abnormal CPU occupancy rate and insufficient memory (anomaly) are identified, will the target fault be triggered to perform subsequent automatic stop-loss actions.
[0040] Step 104, if the fault identification result indicates the existence of a target fault, automatically execute the fault stop-loss action corresponding to the target fault.
[0041] After the fault result indicates the existence of a target fault, since the fault stop-loss action corresponding to the target fault has been predefined, there is no need to wait for the operation and maintenance personnel to stop the loss, and the fault stop-loss action corresponding to the target fault can be directly executed automatically, improving the fault stop-loss efficiency.
[0042] In summary, the embodiment of the present application provides a fault stop-loss method: by arranging the target configuration of the fault stop-loss service process corresponding to the fault scenario, including the data collection task set, the anomaly judgment rule, the fault identification rule, the fault stop-loss action, etc., the automatic stop-loss platform can collect the target collection data from the target device by executing the fault stop-loss service process, and perform fault identification on the target collection data through the anomaly judgment rule and / or the fault identification rule. Furthermore, after identifying the target fault, automatically execute its corresponding fault stop-loss action to achieve automatic fault stop-loss in different fault scenarios. Moreover, the process from fault discovery to fault stop-loss does not require the participation of operation and maintenance personnel, improving the fault discovery efficiency and the fault stop-loss efficiency.
[0043] The data collection tasks mainly include two types. One is to obtain time-series data from service instances, that is, CPU occupancy rate and memory status; the other is that other business platforms have already processed the time-series data to generate alarm events or abnormal events, and the automatic stop-loss platform directly collects the alarm events or abnormal events from other business platforms for subsequent fault identification and fault stop-loss.
[0044] Please refer to Figure 2 , which shows a flowchart of another fault stop-loss method according to an exemplary embodiment of the present application. This method is described by taking its application to an automatic stop-loss platform as an example. As Figure 2 shown, this method includes:
[0045] Step 201: Receive the configuration operation for the fault stop-loss business process, and determine the target configuration information corresponding to the fault stop-loss business process. The target configuration information includes at least a data collection task, an anomaly judgment rule, a fault identification rule for the target fault, and a fault stop-loss action.
[0046] Step 202: After triggering the execution operation for the fault stop-loss business process, load the data collection task corresponding to the fault stop-loss business process, and obtain the target collection data from the target device based on the data collection task.
[0047] The implementation manners of Step 201 and Step 202 can refer to Step 101 and Step 102, and will not be elaborated herein in this embodiment.
[0048] Step 203: Determine the target data type of the target collection data.
[0049] Among them, the target data type includes time-series data and alarm event data. Different data processes are set for different target data types. Therefore, after the automatic stop-loss platform obtains the target collection data indicated by the data collection task from the target device, first parse the obtained target collection data, and judge whether the target data type of the target collection data is time-series data or alarm event data.
[0050] Among them, time-series data refers to the data continuously recorded as time changes, such as the CPU usage rate, memory occupancy, and network request quantity of a server. Alarm event data refers to the abnormal events or problems detected during the operation of the system, generated through predefined threshold rules or intelligent algorithms. For example: alarm information such as a certain service being unavailable or disk space being insufficient; or alarm events such as data loss, traffic anomaly, or device disconnection.
[0051] Step 204: If the target data type is time-series data, perform fault identification on the target collection data based on the anomaly judgment rule and the fault identification rule to obtain a fault identification result.
[0052] If the target data type is time-series data, first judge the target collection data according to the anomaly judgment rule. If the target collection data meets the anomaly judgment rule, generate a target anomaly event; then perform fault identification on the generated target anomaly event according to the fault identification rule. If the target anomaly event meets the fault identification rule, it is determined that the target fault is identified.
[0053] Exemplarily, if the target acquisition data collected by the data acquisition task includes CPU occupancy and memory occupancy, first, the exception judgment rule corresponding to the CPU occupancy is used to judge the exception of the CPU occupancy. If the exception judgment rule is satisfied, an exception event of too high CPU occupancy is generated; and the exception judgment rule corresponding to the memory occupancy is used to judge the exception of the memory occupancy. If the exception judgment rule is judged, an exception event of insufficient memory is generated; then, based on the fault identification rule of the target fault, the exception events of too high CPU occupancy and insufficient memory are identified for faults, and it is determined that the fault identification rule is satisfied. Finally, the target fault (the target fault is server lag) is identified.
[0054] Considering that in actual operation, the timing data of the server may have short-term fluctuations, but not real exceptions. For example, the network latency increases instantaneously but quickly returns to normal; the CPU usage rate rises briefly and then drops; if each small fluctuation is judged as an exception event, there may be many false alarms. Therefore, when generating the target exception event, a buffering mechanism of window anti-shake will be added. Before judging the exception event, observe the status for a period of time (this time period is called the "window"). When the multiple data within the window all meet the exception judgment rule, it is determined to generate the target exception event, which can filter out those short-term fluctuations and avoid false alarms. Specifically, if the target acquisition data meets the exception judgment rule, continue to obtain the newly added timing data collected within the preset duration. If the newly added timing data meets the exception judgment rule, generate the target exception event. Exemplarily, the preset duration can be set by the operation and maintenance personnel. Exemplarily, the preset duration can be 10 minutes. Taking the target acquisition data as the CPU occupancy as an example, when it is determined that the CPU occupancy of a certain server has exceeded 90% continuously for 10 minutes, it is determined to generate an exception event of CPU occupancy.
[0055] Optionally, the target exception event will also record the detailed information of the exception, such as the exception occurrence time, the exception service instance, and the specific manifestation of the exception. After generating the target exception event, it can also be used for alarm notification (such as sending text messages or emails to relevant personnel) or triggering subsequent automatic loss prevention actions.
[0056] Exemplarily, Figure 3 is a schematic diagram of the processing flow of the exception event provided by an exemplary embodiment of the present application. As Figure 3As shown, after the automatic stop-loss platform obtains the target acquisition data, it first determines whether the target acquisition data meets the abnormal conditions (the abnormal conditions are the abnormal judgment rules. The automatic stop-loss platform determines whether the current target acquisition data meets the abnormal conditions according to the set abnormal judgment rules, such as excessive CPU usage, request delay being too high, etc.). If the abnormal conditions are not met, it directly returns and waits for the newly obtained target acquisition data and the next abnormal judgment; if the abnormal conditions are met, it enters the next step, anti-shake condition filtering. Anti-shake is to avoid misjudging short-term fluctuations as abnormalities. The automatic stop-loss platform will check whether the current abnormality continuously meets a certain time window condition; if it fails the anti-shake filtering, it returns to the data judgment link and waits for the new target acquisition data; if it passes the anti-shake filtering, it enters the next step; continue to determine whether there are unclosed abnormal events. The automatic stop-loss platform will check whether there are already unclosed abnormal events currently (that is, whether the same type of abnormal events have been recorded before). If there are unclosed abnormal events, it enters the update status process; if there are no unclosed abnormal events, it enters the new abnormal event process; new abnormal event. When the automatic stop-loss platform discovers a new abnormal event, it will create a new abnormal event record and save the current abnormal information. After creation, it will enter the stage of pushing the abnormal event. Pushing the abnormal event, pushing the newly created abnormal event to relevant personnel or systems, such as sending an alarm notification, triggering automated processing, etc. Finally, update the event status and push status. If it is an existing abnormal event (there are unclosed events), update its status (such as updating the timestamp, adding new abnormal data, etc.), and at the same time record whether the event has been pushed to avoid repeated notifications. The entire process forms a closed-loop abnormal event processing system through data judgment - abnormal detection - anti-shake filtering - event recording / updating - push notification.
[0057] Step 205, if the target data type is alarm event data, perform fault identification on the target acquisition data based on the fault identification rules to obtain a fault identification result.
[0058] If the target data type is alarm event data, it means that the automatic stop-loss platform does not obtain time-series data from the server itself, but from other business platforms associated with the server. This data is an alarm event or abnormal event after processing the time-series data according to the abnormal judgment rules. Then, for alarm event data, the abnormal judgment rules do not need to be executed, and the target acquisition data can be directly fault-identified according to the fault identification rules to obtain a fault identification result. Specifically, if the target acquisition data meets the fault identification rules, it is determined that the target fault is identified; if the target acquisition data does not meet the fault identification rules, it is determined that the target fault is not identified.
[0059] Step 206: If the fault identification result indicates the existence of a target fault, obtain the target protection policy associated with the target fault.
[0060] Taking the fault loss prevention action of shielding abnormal service instances as an example, if the number of remaining service instances is small currently, directly executing the loss prevention action of shielding service instances may cause the business service to be completely inoperable. Therefore, corresponding target protection policies are also set for different fault loss prevention actions. By verifying the target protection policies, the execution safety of the fault loss prevention actions can be ensured. In a possible implementation manner, if the fault identification result indicates the existence of a target fault, the target protection policy associated with the target fault will be obtained; or, first obtain the fault loss prevention action associated with the target fault, and then obtain the target protection policy associated with the fault loss prevention action.
[0061] Step 207: If the target protection policy is met, automatically execute the fault loss prevention action corresponding to the target fault.
[0062] By verifying the target protection policy, if the target protection policy is met, the fault loss prevention action corresponding to the target fault will be automatically executed; otherwise, if the target protection policy is not met, the fault loss prevention action will not be executed.
[0063] Exemplarily, if the fault loss prevention action associated with the target fault is to shield service instances, obtain the status information of the service instances associated with the data collection task. If the status information indicates that the proportion of remaining service instances is greater than the proportion threshold, it means the target protection policy is met, and the fault loss prevention action of shielding service instances will be automatically executed; otherwise, if the proportion is less than the proportion threshold, it means the target protection policy is not met, and the fault loss prevention action of shielding service instances will not be executed.
[0064] Among them, the remaining service instances are the remaining normal service instances. The proportion threshold can be 30%.
[0065] In this embodiment, by distinguishing the target data types of the target collected data, different data processing flows can be selected, so that not only can fault identification be performed on time-series data, but also fault identification can be performed on alarm event data, expanding the fault identification scenarios; moreover, before executing the fault loss prevention action, the execution safety of the fault loss prevention action can be ensured by verifying the target protection policy.
[0066] Please refer to Figure 4 , which is the flowchart of the complete fault loss prevention method provided by an exemplary embodiment of the present application. As Figure 4As shown, the user can arrange and configure a fault stop-loss workflow (i.e., a fault stop-loss business process) on the automatic stop-loss platform, for example, including: acquisition tasks, anomaly judgment rules, fault feature recognition strategies, executed stop-loss actions, information such as groups / emails / phones reached by the results; after triggering the fault stop-loss workflow, the user-defined acquisition tasks and anomaly judgment rules can be loaded, and the task is pulled up to collect monitoring data through the acquisition module. If the collected data is time-series data, anomaly judgment is performed; if the time-series data conforms to the anomaly judgment rule, an anomaly event is generated and saved to the database, and at the same time, it enters the next process; if the collected data is alarm event data, no anomaly judgment is required, and the collected alarm event is directly saved, and at the same time, it enters the next process; fault feature recognition is performed according to the user-defined fault recognition strategy. The input of the fault feature recognition is the anomaly event generated according to the anomaly judgment rule or the collected alarm event. If it conforms to the feature strategy, it means that a fault has occurred, and automatic stop-loss needs to be performed on this fault; before performing the stop-loss action, the protection logic is first determined. The protection logic is used for stop-loss protection, for example: adding security checks to key operations, such as quota checks, etc., to prevent unauthorized stop-loss operations from being executed. If the protection logic passes, the stop-loss action is executed. The stop-loss action is user-defined, for example, shielding the business service instance, executing the script content, and the execution result is notified according to the user-defined reach channel. If the protection logic fails, the stop-loss action is not executed, and a message is directly notified. Execute the stop-loss action.
[0067] Please refer to Figure 5 , which is a schematic structural diagram of a fault stop-loss device provided by an embodiment of the present application. Exemplarily, as Figure 5 shown, the device 500 includes:
[0068] A receiving module 501, configured to receive a configuration operation on the fault stop-loss business process, and determine target configuration information corresponding to the fault stop-loss business process, where the target configuration information at least includes a data acquisition task, an anomaly judgment rule, a fault recognition rule for a target fault, and a fault stop-loss action;
[0069] An obtaining module 502, configured to load the data acquisition task corresponding to the fault stop-loss business process after triggering an execution operation on the fault stop-loss business process, and obtain target acquisition data from a target device based on the data acquisition task;
[0070] A fault recognition module 503, configured to perform fault recognition on the target acquisition data based on the anomaly judgment rule and / or the fault recognition rule to obtain a fault recognition result;
[0071] A fault stop-loss module 504, configured to automatically execute the fault stop-loss action corresponding to the target fault if the fault recognition result indicates the existence of the target fault.
[0072] Optionally, the fault identification module 503 is further configured to:
[0073] Determine the target data type of the target acquisition data;
[0074] If the target data type is time-series data, perform fault identification on the target acquisition data based on the anomaly judgment rule and the fault identification rule to obtain the fault identification result;
[0075] If the target data type is alarm event data, perform fault identification on the target acquisition data based on the fault identification rule to obtain the fault identification result.
[0076] Optionally, the fault identification module 503 is further configured to:
[0077] If the target acquisition data meets the anomaly judgment rule, generate a target anomaly event;
[0078] If the target anomaly event meets the fault identification rule, determine that the target fault is identified.
[0079] Optionally, the fault identification module 503 is further configured to:
[0080] If the target acquisition data meets the fault identification rule, determine that the target fault is identified;
[0081] If the target acquisition data does not meet the fault identification rule, determine that the target fault is not identified.
[0082] Optionally, the fault identification module 503 is further configured to:
[0083] If the target acquisition data meets the anomaly judgment rule, continue to acquire new time-series data collected within a preset duration;
[0084] If the new time-series data meets the anomaly judgment rule, generate the target anomaly event.
[0085] Optionally, the fault loss prevention module 504 is further configured to:
[0086] If the fault identification result indicates the existence of the target fault, obtain the target protection strategy associated with the target fault;
[0087] If the target protection strategy is met, automatically execute the fault loss prevention action corresponding to the target fault.
[0088] Optionally, the fault loss prevention module 504 is further configured to:
[0089] If the fault loss prevention action associated with the target fault is to shield a service instance, obtain the status information of the service instance associated with the data collection task;
[0090] If the status information indicates that the proportion of remaining service instances is greater than the proportion threshold, automatically execute the fault loss prevention action corresponding to the target fault.
[0091] In summary, the embodiments of the present application provide a fault loss prevention method: by arranging the target configuration of the fault loss prevention service process corresponding to the fault scenario, including a data collection task set, an anomaly judgment rule, a fault identification rule, a fault loss prevention action, etc., the automatic loss prevention platform can collect target collection data from the target device by executing the fault loss prevention service process, and perform fault identification on the target collection data through the anomaly judgment rule and / or the fault identification rule. Furthermore, after the target fault is identified, its corresponding fault loss prevention action is automatically executed to achieve automatic fault loss prevention in different fault scenarios. Moreover, the process from fault discovery to fault loss prevention does not require the participation of operation and maintenance personnel, improving the fault discovery efficiency and fault loss prevention efficiency.
[0092] An exemplary embodiment of the present application further provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program capable of being executed by the at least one processor, and when the computer program is executed by the at least one processor, it is used to cause the electronic device to execute the fault loss prevention method according to the embodiments of the present application.
[0093] An exemplary embodiment of the present application further provides a non-transitory computer-readable storage medium storing a computer program, where the computer program is used to cause a computer to execute the fault loss prevention method according to the embodiments of the present application when executed by a processor of the computer.
[0094] An exemplary embodiment of the present application further provides a computer program product, including a computer program, where the computer program is used to cause a computer to execute the fault loss prevention method according to the embodiments of the present application when executed by a processor of the computer.
[0095] Reference Figure 6, a block diagram of an electronic device 600 that can be a server or a client of the present application will now be described. It is an example of a hardware device that can be applied to various aspects of the present application. The electronic device is intended to represent various forms of digital electronic computer devices, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or claimed herein.
[0096] As Figure 6 shown, the electronic device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0097] A plurality of components in the electronic device 600 are connected to the I / O interface 605, including: an input unit 606, an output unit 607, a storage unit 608, and a communication unit 609. The input unit 606 can be any type of device that can input information into the electronic device 600. The input unit 606 can receive input digital or character information, and generate key signal inputs related to the user settings and / or function controls of the electronic device. The output unit 607 can be any type of device that can present information, and can include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 608 can include, but is not limited to, magnetic disks, optical disks. The communication unit 609 allows the electronic device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks, and can include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0098] The computing unit 601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 executes the various methods and processes described above. For example, in some embodiments, Figure 1 , Figure 2 the method shown can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 600 via the ROM 602 and / or the communication unit 609. In some embodiments, the computing unit 601 can be configured to execute Figure 1 , Figure 2 the method shown in any other suitable manner (e.g., by means of firmware).
[0099] The program code for implementing the methods of the present application can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0100] In the context of the present application, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0101] As used in this application, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., a magnetic disk, an optical disk, a memory, a programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0102] For purposes of providing an interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).
[0103] The systems and techniques described herein can be implemented in a computing system that includes a back-end component (e.g., as a data server), or a computing system that includes a middleware component (e.g., an application server), or a computing system that includes a front-end component (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0104] A computer system can include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship to each other.
Claims
1. A method for stopping loss due to a fault, characterized in that: The method comprises: Receive a configuration operation for a fault-stop-loss business process, and determine target configuration information corresponding to the fault-stop-loss business process, wherein the target configuration information at least includes a data collection task, an abnormality judgment rule, a fault identification rule of a target fault, and a fault-stop-loss action; After triggering the execution operation of the fault-stop-loss business process, loading the data collection task corresponding to the fault-stop-loss business process, and acquiring target collection data from the target device based on the data collection task; Performing fault identification on the target collected data based on the abnormality judgment rule and / or the fault identification rule to obtain a fault identification result; If the fault identification result indicates that the target fault exists, the fault stop loss action corresponding to the target fault is automatically executed.
2. The method according to claim 1, characterized in that The performing fault identification on the target collected data based on the abnormality judgment rule and / or the fault identification rule to obtain a fault identification result includes: Determine the target data type of the target collected data; If the target data type is time series data, fault identification is performed on the target collected data based on the abnormality judgment rule and the fault identification rule to obtain the fault identification result; If the target data type is alarm event data, fault identification is performed on the target collected data based on the fault identification rule to obtain the fault identification result.
3. The method according to claim 2, characterized in that The performing fault identification on the target collected data based on the abnormality judgment rule and the fault identification rule to obtain the fault identification result includes: If the target collected data meets the abnormality judgment rule, a target abnormality event is generated; If the target abnormal event satisfies the fault identification rule, it is determined that the target fault is identified.
4. The method according to claim 2, characterized in that: The performing fault identification on the target collected data based on the fault identification rule to obtain the fault identification result includes: If the target collected data satisfies the fault identification rule, determining that the target fault is identified; If the target collected data does not satisfy the fault identification rule, it is determined that the target fault is not identified.
5. The method according to claim 3, characterized in that: If the target collected data meets the abnormality judgment rule, generating a target abnormality event includes: If the target collected data meets the abnormality judgment rule, continue to obtain the newly added time series data collected within the preset time period; If the newly added time series data meets the abnormality judgment rule, the target abnormality event is generated.
6. The method according to any one of claims 1 to 5, characterized in that: If the fault identification result indicates that the target fault exists, automatically executing the fault stop loss action corresponding to the target fault includes: If the fault identification result indicates that the target fault exists, obtaining a target protection strategy associated with the target fault; If the target protection strategy is met, the fault stop loss action corresponding to the target fault is automatically executed.
7. The method according to claim 6, characterized in that If the target protection strategy is met, automatically executing the fault stop loss action corresponding to the target fault includes: If the fault stop loss action associated with the target fault is to shield the service instance, obtain the status information of the service instance associated with the data collection task; If the status information indicates that the proportion of remaining service instances is greater than a proportion threshold, the fault stop loss action corresponding to the target fault is automatically executed.
8. A fault-stop device, characterized in that: The device comprises: A receiving module, used to receive a configuration operation on a fault-stop-loss business process, and determine target configuration information corresponding to the fault-stop-loss business process, wherein the target configuration information at least includes a data collection task, an abnormality judgment rule, a fault identification rule of a target fault, and a fault-stop-loss action; An acquisition module, used to load the data acquisition task corresponding to the fault-stop-loss business process after triggering the execution operation of the fault-stop-loss business process, and acquire target acquisition data from the target device based on the data acquisition task; A fault identification module, used to perform fault identification on the target collected data based on the abnormality judgment rule and / or the fault identification rule to obtain a fault identification result; A fault loss-stopping module is used to automatically execute the fault loss-stopping action corresponding to the target fault if the fault identification result indicates the existence of the target fault.
9. An electronic device, comprising: processor; as well as Memory for storing programs, The program includes instructions, which, when executed by the processor, cause the processor to execute the fault loss prevention method according to any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to enable the computer to execute the failure stop loss method according to any one of claims 1-7.