Exception handling method, device, equipment and medium

By collecting and processing abnormal signals in real time, determining the level and rolling back to the exception-free stage, the stagnation problem caused by exceptions in online upgrades of distributed clusters is solved, ensuring the normal operation of the cluster.

CN115145760BActive Publication Date: 2025-08-19JINAN INSPUR DATA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210878550.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-25
Publication Date
2025-08-19
Estimated Expiration
2042-07-25

AI Technical Summary

Technical Problem

During the online upgrade of distributed clusters, an exception occurs that causes the cluster to stop upgrading, affecting use.

Method used

Abnormal signals are collected in real time, the level is determined based on the abnormal signals, the rollback command is received and the abnormality is returned to the stage where no abnormality occurs, and an alarm signal is output.

Benefits of technology

Avoid cluster stopping caused by abnormalities during online upgrade time and ensure that the cluster is in use normally.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115145760B_ABST
    Figure CN115145760B_ABST
Patent Text Reader

Abstract

The present application discloses an exception handling method, device, equipment and medium, which relate to the field of distributed technology. The exception handling method provided by the present application is applied to a distributed cluster for online upgrading, including: real-time collection of abnormal signals; determining the abnormal level according to the abnormal signal; selectively receiving a rollback instruction according to the abnormal level, the rollback instruction is an instruction to return to a stage where no abnormal signal appears; returning to the breakpoint according to the rollback instruction, and outputting an alarm signal, wherein the breakpoint is a stage that characterizes the progress of the distributed cluster. At this time, the rollback instruction can make the upgrade process return to the power-off point, and no abnormal signal will appear at the breakpoint. At this time, the alarm signal is output again, which is convenient for technical personnel to determine and resolve the abnormal signal. It avoids the occurrence of abnormal information during the online upgrade time, causing the distributed cluster to stop upgrading and affecting the use of the distributed cluster.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of distributed technology, and in particular to an exception handling method, apparatus, device, and medium. Background Art

[0002] With the continuous development of the times, the scale of distributed clusters is increasing. To ensure the continuous availability of services running in distributed clusters, distributed clusters can only be updated through online upgrades. Generally speaking, the scale of distributed clusters is large, and the online upgrade time is also long. If any exception occurs during the online upgrade time, the distributed cluster upgrade will be stopped, affecting the use of the distributed cluster.

[0003] In view of the above problems, finding out how to handle the abnormalities generated during the online upgrade time is a problem that those skilled in the art are trying their best to solve. Summary of the Invention

[0004] The purpose of this application is to provide an exception handling method, apparatus, device and medium for handling exceptions generated during online upgrade time.

[0005] To solve the above technical problems, the present application provides an exception handling method, which is applied to a distributed cluster for online upgrade, including:

[0006] Real-time collection of abnormal signals;

[0007] Determine the abnormality level based on the abnormal signal;

[0008] Select and receive a rollback instruction based on the abnormality level. The rollback instruction is an instruction to return to a stage where no abnormal signal occurs.

[0009] Return to the breakpoint according to the rollback instruction and output an alarm signal, where the breakpoint represents the progress stage of the distributed cluster.

[0010] Preferably, determining the abnormality level according to the abnormal signal includes:

[0011] Obtain abnormal nodes based on abnormal signals;

[0012] Count the number of abnormal nodes;

[0013] Determine the corresponding abnormality level according to the number and a preset first preset value, a second preset value, and a third preset value, wherein the first preset value is less than the second preset value, and the second preset value is less than the third preset value;

[0014] When the number is less than a first preset value, determining the abnormality level to be a first abnormality level;

[0015] When the number is not less than the first preset value and less than the second preset value, the abnormality level is determined to be the second abnormality level;

[0016] When the number is not less than the second preset value and less than the third preset value, the abnormality level is determined to be the third abnormality level;

[0017] When the number is greater than the third preset value, the abnormality level is determined to be the fourth abnormality level.

[0018] Preferably, determining the abnormality level according to the abnormal signal includes:

[0019] Calling a preset level relationship table, wherein the preset level relationship table contains the abnormality level corresponding to each abnormal signal;

[0020] The abnormality level is determined based on the preset level relationship table and the abnormal signal.

[0021] Preferably, receiving a rollback instruction according to the exception level includes:

[0022] When the abnormality level is the fourth abnormality level, a rollback instruction is received.

[0023] Preferably, when the abnormality level is the first abnormality level, the method further includes:

[0024] An instruction for automatically recovering from an abnormality corresponding to the first abnormality level is received and abnormal recovery is performed.

[0025] Preferably, when the abnormality level is the second abnormality level, the method further includes:

[0026] Summarize the abnormal signals of each abnormal node;

[0027] The aggregated abnormal signals are sent to the abnormality handling interface at regular intervals to facilitate manual remote processing of the abnormal signals corresponding to the second abnormality level.

[0028] Preferably, when the abnormality level is the third abnormality level, the method further includes:

[0029] receiving a pause signal indicating suspension of the online upgrade;

[0030] The abnormal signal is sent to the abnormality handling interface so as to facilitate manual remote processing of the abnormal signal corresponding to the third abnormality level.

[0031] Preferably, after determining the abnormality level according to the abnormality signal and before receiving the rollback instruction according to the abnormality level, the method further includes:

[0032] Obtain real-time online upgrade progress and distributed cluster status information;

[0033] Set breakpoints regularly based on progress and status information.

[0034] To solve the above technical problems, the present application also provides an exception handling device, which is applied to a distributed cluster for online upgrade, comprising:

[0035] Real-time acquisition module, used to collect abnormal signals in real time;

[0036] A first determining module, configured to determine an abnormality level according to the abnormal signal;

[0037] A first receiving module is configured to receive a rollback instruction according to the abnormality level, wherein the rollback instruction is an instruction to return to a stage where no abnormal signal occurs;

[0038] The return and output module is used to return to the breakpoint according to the rollback instruction and output an alarm signal, wherein the breakpoint represents the progress stage of the distributed cluster. In addition, the device also includes the following modules:

[0039] Determining the abnormality level according to the abnormal signal includes:

[0040] A first acquisition module is used to acquire abnormal nodes according to abnormal signals;

[0041] Statistics module, used to count the number of abnormal nodes;

[0042] A second determining module is configured to determine a corresponding abnormality level according to the number and a first preset value, a second preset value, and a third preset value, wherein the first preset value is less than the second preset value, and the second preset value is less than the third preset value;

[0043] a third determining module, configured to determine the abnormality level as a first abnormality level when the number is less than a first preset value;

[0044] a fourth determining module, configured to determine the abnormality level as a second abnormality level when the number is not less than the first preset value and less than the second preset value;

[0045] a fifth determining module, configured to determine that the abnormality level is a third abnormality level when the number is not less than the second preset value and less than the third preset value;

[0046] The sixth determining module is configured to determine that the abnormality level is a fourth abnormality level when the number is greater than a third preset value.

[0047] Determining the abnormality level according to the abnormal signal includes:

[0048] A calling module is used to call a preset level relationship table, wherein the preset level relationship table contains the abnormality level corresponding to each abnormal signal;

[0049] The seventh determination module is configured to determine the abnormality level according to the preset level relationship table and the abnormality signal.

[0050] When the abnormality level is the first abnormality level, it also includes:

[0051] The second receiving module is configured to receive an automatic abnormality recovery instruction corresponding to the first abnormality level and perform abnormality recovery.

[0052] When the abnormality level is the second abnormality level, it also includes:

[0053] The summary module is used to summarize the abnormal signals of each abnormal node;

[0054] The first sending module is used to send the summarized abnormal signals to the abnormal processing interface at a regular time, so as to facilitate manual remote processing of the abnormal signals corresponding to the second abnormal level.

[0055] When the abnormality level is the third abnormality level, it also includes:

[0056] A third receiving module is used to receive a pause signal indicating that the online upgrade is paused;

[0057] The second sending module is used to send the abnormal signal to the abnormal processing interface to facilitate manual remote processing of the abnormal signal corresponding to the third abnormal level.

[0058] After determining the abnormality level according to the abnormality signal and before receiving the rollback instruction according to the abnormality level, the method further includes:

[0059] The second acquisition module is used to obtain the progress of online upgrades and status information of distributed clusters in real time;

[0060] The settings module is used to set breakpoints based on progress and status information.

[0061] To solve the above technical problems, the present application also provides an exception handling device, including:

[0062] memory for storing computer programs;

[0063] The processor is used to point to the computer program to implement all the steps of the above exception handling method.

[0064] To solve the above technical problems, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of all the above exception handling methods are implemented.

[0065] The present application provides an exception handling method, which is applied to distributed clusters for online upgrades, including: real-time collection of abnormal signals; determining the abnormality level based on the abnormal signals; selectively receiving rollback instructions based on the abnormality level, wherein the rollback instruction is an instruction to return to a stage where no abnormal signal appears; returning to the breakpoint according to the rollback instruction, and outputting an alarm signal, wherein the breakpoint is a stage that characterizes the progress of the distributed cluster. At this time, the rollback instruction can make the upgrade process return to the power-off point, and no abnormal signal will appear at the breakpoint. At this time, the alarm signal is output again, which is convenient for technical personnel to determine and resolve the abnormal signal. This avoids the occurrence of abnormal information during the online upgrade time, causing the distributed cluster to stop upgrading and affecting the use of the distributed cluster.

[0066] This application also provides an exception handling device, equipment and medium, with the same effects as above. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0068] Figure 1 A flow chart of an exception handling method provided in an embodiment of the present application;

[0069] Figure 2 A structural diagram of an exception handling device provided in an embodiment of the present application;

[0070] Figure 3 A structural diagram of an exception handling device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0071] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0072] The core of this application is to provide an exception handling method, device, equipment and medium, which can handle the generated exceptions within the time of online upgrade.

[0073] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0074] Distributed computing improves efficiency by shortening the execution time of individual tasks, while clustering improves efficiency by increasing the number of tasks executed per unit time. For example, if a task consists of 10 subtasks, each of which takes one hour to execute individually, executing the task on a single server would take 10 hours. A distributed solution, with 10 servers each handling only one subtask and ignoring dependencies between subtasks, would complete the task in just one hour. A clustered solution, with the same 10 servers, would also allow each server to independently handle the task. Assuming 10 tasks arrive simultaneously, all 10 servers will work simultaneously. After 10 hours, all 10 tasks will be completed simultaneously. Overall, each task will still complete within one hour.

[0075] Distributed systems can be categorized as intra-machine systems, intra-building systems, inter-building systems, and regional systems across geographic areas. Their coupling levels, from high to low, are determined by the nature of their application domains. They can be divided into three categories: The first category is distributed parallel computer systems and distributed multi-user computer systems for computing tasks. These systems require the highest possible coupling so they can evolve to share the workload of mainframe and time-sharing computer systems. The second category is distributed data processing systems for management information. These systems can have a lower coupling level. The third category is distributed computer control systems for process control. These systems require a moderate degree of coupling, although some real-time applications may require a higher degree.

[0076] A distributed cluster typically includes a distributed cluster virtual host. Simply upload your website to a data center server, and the site content will be automatically distributed to dozens of node servers across the country. Your endpoint will automatically access the fastest server in your location. When users in different regions access the same domain name, they will ping different server IP addresses, ensuring the website is always within reach of the end user.

[0077] Figure 1 This is a flow chart of an exception handling method provided in an embodiment of the present application. Figure 1 As shown, the exception handling method is applied to a distributed cluster undergoing online upgrades and includes:

[0078] S10: Collect abnormal signals in real time.

[0079] S11: Determine the abnormality level according to the abnormal signal.

[0080] S12: Select and receive a rollback instruction according to the abnormality level.

[0081] The rollback instruction is an instruction to return to a stage where no abnormal signal has occurred.

[0082] S13: Return to the breakpoint according to the rollback instruction and output an alarm signal.

[0083] The breakpoint represents the progress stage of the distributed cluster.

[0084] In this embodiment, the representation format of the abnormal signal is not limited and can be expressed in text or data string form. When expressing the abnormal signal in text form, it can be expressed as "Exception signal is: xxx." When expressing the abnormal signal in data string form, it can be expressed using a data string of 1, 2, 4, 8, or other digits. For example, based on the digit order mentioned above, these can be "0," "01," "1000," and "01001011." It should be noted that when expressing the abnormal signal in data string form, a preset value can be set. When the binary data string is converted to decimal data, an abnormal signal is indicated if the preset value is exceeded. Alternatively, the abnormal signal can be determined based on the number of "0"s and / or "1"s in the data string. It should be understood that the above embodiments are only one or more of many and do not limit the representation format of the abnormal signal. The implementation method can be determined based on the specific implementation scenario. When the abnormality level reaches the fourth abnormality level, a rollback instruction is received. The rollback instruction is an instruction to return to a stage before the abnormal signal occurred. According to the rollback instruction, the system returns to the breakpoint and outputs an alarm signal. It should be noted that the rollback instruction can be expressed in either text or data string form. When expressed in text form, the rollback instruction can be expressed as "Rollback instruction has been issued." When expressed in data string form, the rollback instruction can be expressed using a data string of 1, 2, 4, 8, or other digits. For example, based on the digit order mentioned above, the examples include "0," "01," "1000," and "01001011." Similarly, the alarm signal can be expressed in either text or data string form. When expressed in text form, the alarm signal can be expressed as "Alarm signal has been issued." When expressed in data string form, the alarm signal can be expressed using a data string of 1, 2, 4, 8, or other digits. For example, based on the digit order mentioned above, the examples include "0," "01," "1000," and "01001011." It should be understood that the above embodiments are merely one or more of many corresponding embodiments and do not limit the representation format of the rollback instruction and alarm signal. The implementation method can be determined based on the specific implementation scenario.

[0085] In addition, it should be noted that after real-time acquisition of abnormal signals and before determining the abnormality level based on the abnormal signals, the following steps are also included:

[0086] Obtain real-time online upgrade progress and distributed cluster status information;

[0087] Breakpoints are set regularly based on progress and status information. Breakpoints represent the progress stages of the distributed cluster.

[0088] In this embodiment, the progress of the online upgrade can be represented by data with a decimal point in the numerical range of 0-1, or by a natural number in the numerical range of 0-10, or even by a natural number in the numerical range of 0-100. And the above data can all be expressed in the form of percentages, for example: 0.1%, 6%, 80%, etc. And it can be understood that the time interval for setting breakpoints can be 1s, 1min, 1h, etc. In this embodiment, the above-mentioned embodiments of the progress of the online upgrade and the time interval for setting breakpoints are only one or several of many embodiments, and do not limit the progress of the online upgrade, the time interval for setting breakpoints, and the number of breakpoints. The implementation method can be determined according to the specific implementation scenario.

[0089] In this embodiment, the abnormality levels are divided into four, namely the first abnormality level, the second abnormality level, the third abnormality level, and the fourth abnormality level.

[0090] Then, determining the abnormality level according to the abnormal signal includes:

[0091] Obtain abnormal nodes based on abnormal signals;

[0092] Count the number of abnormal nodes;

[0093] Determine the corresponding abnormality level according to the number and a preset first preset value, a second preset value, and a third preset value, wherein the first preset value is less than the second preset value, and the second preset value is less than the third preset value;

[0094] When the number is less than a first preset value, determining the abnormality level to be a first abnormality level;

[0095] When the number is not less than the first preset value and less than the second preset value, the abnormality level is determined to be the second abnormality level;

[0096] When the number is not less than the second preset value and less than the third preset value, the abnormality level is determined to be the third abnormality level;

[0097] When the number is greater than the third preset value, the abnormality level is determined to be the fourth abnormality level.

[0098] The numerical range of the first preset value, the second preset value, and the third preset value can be a value with one decimal point between 0 and 1, or a natural number between 0 and 10, or even a natural number between 0 and 100. In this embodiment, the numerical range of the first preset value, the second preset value, and the third preset value is not limited, and its implementation method can be determined according to the specific implementation scenario.

[0099] At the same time, there is another way to determine the abnormal level:

[0100] Then, determining the abnormality level according to the abnormal signal includes:

[0101] Calling a preset level relationship table, wherein the preset level relationship table contains the abnormality level corresponding to each abnormal signal;

[0102] The abnormality level is determined based on the preset level relationship table and the abnormal signal.

[0103] The present application provides an exception handling method, which is applied to distributed clusters for online upgrades, including: real-time collection of abnormal signals; determining the abnormality level based on the abnormal signals; selectively receiving rollback instructions based on the abnormality level, wherein the rollback instruction is an instruction to return to a stage where no abnormal signal appears; returning to the breakpoint according to the rollback instruction, and outputting an alarm signal, wherein the breakpoint is a stage that characterizes the progress of the distributed cluster. At this time, the rollback instruction can make the upgrade process return to the power-off point, and no abnormal signal will appear at the breakpoint. At this time, the alarm signal is output again, which is convenient for technical personnel to determine and resolve the abnormal signal. This avoids the occurrence of abnormal information during the online upgrade time, causing the distributed cluster to stop upgrading and affecting the use of the distributed cluster.

[0104] On the basis of the above embodiment, as a more preferred embodiment, when the exception level is the first exception level, it also includes: receiving the instruction of automatically recovering the exception corresponding to the first exception level and performing exception recovery. It should be noted that the instruction can be expressed in text form or in the form of a data string. When the instruction of automatically recovering the exception is expressed in text form, it can be expressed as "the instruction of automatically recovering the exception has been issued"; when the instruction of automatically recovering the exception is expressed in the form of a data string, it can be expressed using a data string of 1 bit, 2 bits, 4 bits, 8 bits, etc., and according to the order of the bits mentioned above, it can be "0", "01", "1000" and "01001011". It can be understood that the above embodiment is only one or more of many embodiments, and does not limit the representation form of the instruction of automatically recovering the exception. Its implementation method can be determined according to the specific implementation scenario.

[0105] When the abnormality level is the second abnormality level, it also includes: summarizing the abnormal signals of each abnormal node; and regularly sending the summarized abnormal signals to the abnormality processing interface to facilitate manual remote processing of the abnormal signals corresponding to the second abnormality level.

[0106] When the abnormality level is the third abnormality level, it also includes: receiving a pause signal indicating the suspension of online upgrades; sending the abnormal signal to the abnormality handling interface to facilitate manual remote processing of the abnormal signal corresponding to the third abnormality level. It should be noted that the pause signal indicating the suspension of online upgrades can be expressed in text form or in the form of a data string. When the pause signal is expressed in text form, it can be expressed as "a pause signal has been issued"; when the pause signal is expressed in the form of a data string, a data string of 1 bit, 2 bits, 4 bits, 8 bits, etc. can be used to express it. For example, according to the order of the bit numbers mentioned above, it can be "0", "01", "1000" and "01001011". It can be understood that the above embodiments are only one or more of many embodiments, and do not limit the representation of the pause signal. Its implementation method can be determined according to the specific implementation scenario.

[0107] In the above embodiments, the exception handling method is described in detail. This application also provides corresponding embodiments of the exception handling device. It should be noted that this application describes the embodiments of the device part from two perspectives: one is based on the functional module perspective, and the other is based on the hardware perspective.

[0108] Figure 2 This is a structural diagram of an exception handling device provided in an embodiment of the present application. Figure 2 As shown, the present application also provides an exception handling device, which is applied to a distributed cluster for online upgrade, including:

[0109] A real-time acquisition module 20 is used to acquire abnormal signals in real time;

[0110] A first determining module 21 is configured to determine an abnormality level according to an abnormal signal;

[0111] A first receiving module 22 is configured to receive a rollback instruction according to the abnormality level, wherein the rollback instruction is an instruction to return to a stage where no abnormal signal occurs;

[0112] The return and output module 23 is configured to return to the breakpoint according to the rollback instruction and output an alarm signal, wherein the breakpoint represents the progress stage of the distributed cluster.

[0113] In addition, the device includes the following modules:

[0114] Determining the abnormality level according to the abnormal signal includes:

[0115] A first acquisition module is used to acquire abnormal nodes according to abnormal signals;

[0116] Statistics module, used to count the number of abnormal nodes;

[0117] A second determining module is configured to determine a corresponding abnormality level according to the number and a first preset value, a second preset value, and a third preset value, wherein the first preset value is less than the second preset value, and the second preset value is less than the third preset value;

[0118] a third determining module, configured to determine the abnormality level as a first abnormality level when the number is less than a first preset value;

[0119] a fourth determining module, configured to determine the abnormality level as a second abnormality level when the number is not less than the first preset value and less than the second preset value;

[0120] a fifth determining module, configured to determine that the abnormality level is a third abnormality level when the number is not less than the second preset value and less than the third preset value;

[0121] The sixth determining module is configured to determine that the abnormality level is a fourth abnormality level when the number is greater than a third preset value.

[0122] Determining the abnormality level according to the abnormal signal includes:

[0123] A calling module is used to call a preset level relationship table, wherein the preset level relationship table contains the abnormality level corresponding to each abnormal signal;

[0124] The seventh determination module is configured to determine the abnormality level according to the preset level relationship table and the abnormality signal.

[0125] When the abnormality level is the first abnormality level, it also includes:

[0126] The second receiving module is configured to receive an automatic abnormality recovery instruction corresponding to the first abnormality level and perform abnormality recovery.

[0127] When the abnormality level is the second abnormality level, it also includes:

[0128] The summary module is used to summarize the abnormal signals of each abnormal node;

[0129] The first sending module is used to send the summarized abnormal signals to the abnormal processing interface at a regular time, so as to facilitate manual remote processing of the abnormal signals corresponding to the second abnormal level.

[0130] When the abnormality level is the third abnormality level, it also includes:

[0131] A third receiving module is used to receive a pause signal indicating that the online upgrade is paused;

[0132] The second sending module is used to send the abnormal signal to the abnormal processing interface to facilitate manual remote processing of the abnormal signal corresponding to the third abnormal level.

[0133] After determining the abnormality level according to the abnormality signal and before receiving the rollback instruction according to the abnormality level, the method further includes:

[0134] The second acquisition module is used to obtain the progress of online upgrades and status information of distributed clusters in real time;

[0135] The settings module is used to set breakpoints based on progress and status information.

[0136] Since the above-mentioned exception handling device is also applicable to the exception handling method, the method is applied to the distributed cluster for online upgrade, including: real-time collection of abnormal signals; determining the abnormal level according to the abnormal signal; selectively receiving a rollback instruction according to the abnormal level, the rollback instruction is an instruction to return to the stage where no abnormal signal appears; returning to the breakpoint according to the rollback instruction, and outputting an alarm signal, wherein the breakpoint is a stage that characterizes the progress of the distributed cluster. At this time, the rollback instruction can make the upgrade process return to the power-off point, and no abnormal signal will appear at the breakpoint. At this time, the alarm signal is output again, which is convenient for technical personnel to determine and resolve the abnormal signal. It avoids the occurrence of abnormal information during the online upgrade time, causing the distributed cluster to stop upgrading and affecting the use of the distributed cluster.

[0137] Since the embodiments of the apparatus part correspond to the embodiments of the method part, please refer to the description of the embodiments of the method part for the embodiments of the apparatus part, and they will not be repeated here.

[0138] Figure 3 This is a structural diagram of an exception handling device provided in an embodiment of the present application, such as Figure 3 As shown, the exception handling device includes:

[0139] Memory 30, for storing computer programs;

[0140] The processor 31 is configured to implement the steps of the exception handling method mentioned in the above embodiment when executing a computer program.

[0141] The exception handling device provided in this embodiment may include but is not limited to a smart phone, a tablet computer, a laptop computer, or a desktop computer.

[0142] Among them, the processor 31 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 31 can be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), and programmable logic array (PLA). The processor 31 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a central processing unit (CPU); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 31 may be integrated with a graphics processing unit (GPU), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 31 may also include an artificial intelligence (AI) processor, which is used to process computing operations related to machine learning.

[0143] The memory 30 may include one or more computer-readable storage media, which may be non-transitory. The memory 30 may also include a high-speed random access memory, and a non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In this embodiment, the memory 30 is at least used to store the following computer program, wherein, after the computer program is loaded and executed by the processor 31, it can implement the relevant steps of the exception handling method disclosed in any of the aforementioned embodiments. In addition, the resources stored in the memory 30 may also include an operating system and data, etc., and the storage method may be temporary storage or permanent storage. Among them, the operating system may include Windows, Unix, Linux, etc. The data may include but is not limited to exception handling methods, etc.

[0144] In some embodiments, the exception handling device may further include a display screen, an input and output interface, a communication interface, a power supply, and a communication bus.

[0145] Those skilled in the art will understand that Figure 3 The structure shown in the figure does not constitute a limitation to the exception handling device, and may include more or fewer components than shown in the figure.

[0146] The exception handling device provided in the embodiment of the present application includes a memory 30 and a processor 31. The processor 31 can implement an exception handling method when executing a program stored in the memory 30.

[0147] Finally, the present application also provides an embodiment corresponding to a computer-readable storage medium. The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps described in the above method embodiment.

[0148] It is understandable that if the method in the above embodiment is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium and executes all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory), ROM, random access memory (Random Access Memory, RAM), disk or optical disk, and other media that can store program code.

[0149] The above is a detailed introduction to an exception handling method, device, equipment and medium provided by the present application. The various embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same and similar parts between the various embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part description. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of this application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of this application.

[0150] It should also be noted that, in this specification, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus comprising the element.

Claims

1. An exception handling method, characterized in that: Distributed clusters used for online upgrades include: Real-time collection of abnormal signals; determining an abnormality level according to the abnormal signal; Obtain real-time online upgrade progress and distributed cluster status information; Setting breakpoints regularly according to the progress and the status information; If the number of abnormal nodes obtained according to the abnormal signal is greater than a third preset value, a rollback instruction is received according to the abnormality level, where the rollback instruction is an instruction to return to a stage where the abnormal signal does not appear; Return to the breakpoint according to the rollback instruction and output an alarm signal, wherein the breakpoint represents the progress stage of the distributed cluster.

2. The exception handling method according to claim 1, characterized in that: Determining the abnormality level according to the abnormal signal includes: Acquire an abnormal node according to the abnormal signal; Counting the number of abnormal nodes; Determine a corresponding abnormality level according to the number and a preset first preset value, a second preset value, and a third preset value, wherein the first preset value is smaller than the second preset value, and the second preset value is smaller than the third preset value; When the number is less than the first preset value, determining the abnormality level to be the first abnormality level; When the number is not less than the first preset value and less than the second preset value, determining the abnormality level to be the second abnormality level; When the number is not less than the second preset value and less than the third preset value, determining that the abnormality level is the third abnormality level; When the number is greater than the third preset value, the abnormality level is determined to be a fourth abnormality level.

3. The exception handling method according to claim 1, wherein: Determining the abnormality level according to the abnormal signal includes: Calling a preset level relationship table, wherein the preset level relationship table contains the abnormality level corresponding to each abnormal signal; The abnormality level is determined according to the preset level relationship table and the abnormality signal.

4. The exception handling method according to claim 2, wherein: When the abnormality level is the first abnormality level, the method further includes: An instruction for automatically recovering from the abnormality corresponding to the first abnormality level is received and abnormality recovery is performed.

5. The exception handling method according to claim 2, characterized in that: When the abnormality level is the second abnormality level, the method further includes: Summarizing the abnormal signals of each abnormal node; The aggregated abnormal signals are sent to an abnormality processing interface at regular intervals to facilitate manual remote processing of the abnormal signals corresponding to the second abnormality level.

6. The exception handling method according to claim 2, characterized in that: When the abnormality level is the third abnormality level, the method further includes: receiving a pause signal indicating suspension of the online upgrade; The abnormal signal is sent to an abnormality processing interface so as to manually and remotely process the abnormal signal corresponding to the third abnormality level.

7. An exception handling device, characterized in that: Distributed clusters used for online upgrades include: Real-time acquisition module, used to collect abnormal signals in real time; a first determining module, configured to determine an abnormality level according to the abnormal signal; The second acquisition module is used to obtain the progress of online upgrades and status information of distributed clusters in real time; A setting module, configured to periodically set breakpoints according to the progress and the status information; a first receiving module, configured to receive a rollback instruction according to the abnormality level if the number of abnormal nodes obtained according to the abnormal signal is greater than a third preset value, the rollback instruction being an instruction to return to a stage where the abnormal signal does not appear; The return and output module is used to return to the breakpoint according to the rollback instruction and output an alarm signal, wherein the breakpoint represents the progress stage of the distributed cluster.

8. An exception handling device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the exception handling method according to any one of claims 1 to 6 when executing the computer program.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the exception handling method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Breakpoint processing method and system based on memory database

    CN107273449A

  • Self-upgrading system of cluster management software

    CN113986287A