Application system fault monitoring and recovering method

By setting up multiple instances of fault monitoring and recovery program in the edge computing platform, the high reliability of fault monitoring and recovery programs is achieved, the system fault monitoring and recovery failure problems caused by abnormal fault monitoring and recovery programs themselves are solved, and the stability and availability of the system are improved.

CN120234181APending Publication Date: 2025-07-01SHENZHEN KAIFA TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311867445.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-29
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

In edge computing application systems, the reliability of the fault monitoring and recovery program itself is ignored, resulting in the inability to perform system fault monitoring and recovery normally in the event of a failure.

Method used

Two or more fault monitoring and recovery program instances are set up in the edge computing platform. The instances monitor each other's operating status and restart them to monitor and recover the target software application services, and adopt redundant design to ensure the high reliability of the program module.

Benefits of technology

It improves the stability and high availability of the system, maximizes the online time of edge computing platform software application services, has small performance losses and low development costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234181A_ABST
    Figure CN120234181A_ABST
Patent Text Reader

Abstract

The invention discloses an application system fault monitoring and recovering method, and belongs to the technical field of computer application. According to the scheme, two or more than two fault monitoring and recovery program instances are set to be automatically started after being started, the fault monitoring and recovery program instances mutually carry out running state monitoring and exception recovery, meanwhile, fault monitoring is carried out on the target software application service in the application system, and the faulted software application service is recovered. According to the method, redundancy is set for the fault monitoring and recovery program, so that high reliability of the fault monitoring program module is guaranteed, and the problem that system fault monitoring and recovery cannot be normally carried out due to abnormity of the fault monitoring and recovery program is avoided; therefore, the stability and high availability of long-time operation of a computer application system are better guaranteed, the online time of edge computing platform software application service is prolonged to the maximum extent, the system reliability is improved, and the scheme is small in performance loss, high in universality, low in development cost and worthy of popularization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer application technologies, and particularly to a method for monitoring and recovering application system failures. Background Art

[0002] Edge computing refers to an open platform that integrates network, computing, storage, and application core capabilities on the side close to the object or data source to provide the nearest-end services nearby. Its application programs are initiated on the edge side, generating faster network service responses and meeting the basic needs of industries in aspects such as real-time services, application intelligence, security, and privacy protection. Edge computing is located between physical entities and industrial connections or at the top of physical entities. High-reliability edge computing is a characteristic of edge computing, which enables data to be stored and processed locally. Even when the network connection is unstable or interrupted, the system can still maintain normal operation. This characteristic makes edge computing highly reliable, capable of handling large-scale data processing requirements and greatly reducing latency, and improving the response speed of the system. To ensure the normal operation of computing services, it is very necessary to monitor and recover various application software that needs to run stably for a long time.

[0003] In an edge computing application system, a fault monitoring and recovery program is often used to monitor system application failures and execute a recovery mechanism. Fault monitoring means include application service status monitoring, abnormal analysis of application service logs, process heartbeat probing mechanisms, etc. However, the reliability of the fault monitoring and recovery program itself is ignored. When the fault monitoring and recovery program itself fails or malfunctions, it will lead to the failure of the fault monitoring and recovery mechanism, and the computer application system cannot be normally monitored and recovered. Summary of the Invention

[0004] To improve the reliability of the fault monitoring and recovery program, the present invention provides a method for monitoring and recovering application system failures, and the method includes the following steps:

[0005] Step S1, set two or more instances of the fault monitoring and recovery program to start automatically when the computer boots up;

[0006] Step S2, the instances of the fault monitoring and recovery program monitor each other's running status, and restart the instances of the fault monitoring and recovery program with abnormal running status;

[0007] Step S3, the instances of the fault monitoring and recovery program monitor the target software application service for failures and recover the software application service that has failed.

[0008] Preferably, in step S1, the instances of the fault monitoring and recovery program run in the form of background service programs.

[0009] Preferably, in step S2, the fault monitoring and recovery program instances monitor the running states of each other at a first preset time interval.

[0010] Preferably, in step S2, the fault monitoring and recovery program instances query the running states of other fault monitoring and recovery program instances through the systemctl status command, and restart the processes of the fault monitoring and recovery program instances with abnormal running states through the systemctl restart command.

[0011] Preferably, in step S3, the fault monitoring and recovery program instances perform fault monitoring on the target software application service at a second preset time interval.

[0012] Preferably, in step S3, when a fault monitoring and recovery program instance detects an abnormality in a certain target software application service, it sends a synchronization message to other fault monitoring and recovery program instances, and after a first preset delay time, the fault monitoring and recovery program instance that first discovers the abnormality in the target software application service restarts the abnormal target software application service.

[0013] Preferably, the synchronization message is a broadcast message, which includes the identifier of the application service with an abnormality, the time when the abnormality is discovered, and the identifier of the fault monitoring and recovery program instance that discovers the abnormality in the software application service. The broadcast message is sent via UDP.

[0014] Preferably, the fault monitoring and recovery method further includes:

[0015] Step S4, the fault monitoring and recovery program instances analyze the system resources and load at a third preset time interval. When it detects an abnormality in the system resource load, it performs system initialization and recovery operations.

[0016] In the technical solution of the present invention, two or more fault monitoring and recovery program instances are set to start automatically when the single machine in the edge computing platform is powered on. Each fault monitoring and recovery program instance monitors the running state and performs abnormal recovery on each other. At the same time, it performs fault monitoring on the target software application services in the application system and recovers the faulty software application services. The present invention ensures the high reliability of the fault monitoring program module itself by setting redundancy for the fault monitoring and recovery program, avoiding the problem that the system fault monitoring and recovery cannot be carried out normally due to the abnormality of the fault monitoring and recovery program itself, thereby better ensuring the stability and high availability of the long-term operation of the computer application system, maximizing the online time of the software application services in the edge computing platform, improving the system reliability. The performance loss of the solution of the present invention is small, the versatility is strong, and the development cost is low, which is worthy of promotion. Description of the Drawings

[0017] The accompanying drawings described herein are used to provide a further understanding of the present invention and form a part of the present invention. The illustrative embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0018] Figure 1 It is a schematic diagram of the steps of a method for monitoring and recovering application system failures provided by an embodiment of the present invention.

[0019] Figure 2 It is a schematic diagram of system module interaction provided by an embodiment of the present invention. Detailed implementation manners

[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0021] The general idea of the present invention is: by deploying two or more independent failure monitoring and recovery program instances that can monitor and recover from each other, to ensure the high reliability of the failure monitoring and recovery program itself, and avoid the problem that the system failure monitoring and recovery cannot be carried out normally due to the abnormality of the failure monitoring and recovery program itself, so as to better ensure the stability and high availability of the long-term operation of the computer application system.

[0022] The embodiments of the present invention will be further described in detail below with reference to the accompanying drawings of the specification. It should be understood that the embodiments described herein are only used to illustrate and explain the present invention and are not used to limit the present invention.

[0023] The present invention is applicable to monitoring and recovering software application services in an application system on a single machine in an edge computing platform.

[0024] Figure 1 It is a schematic diagram of the steps of a method for monitoring and recovering application system failures provided by an embodiment of the present invention, Figure 2 It is a schematic diagram of system module interaction provided by an embodiment of the present invention. As Figure 1-2 shown, an embodiment of the present invention provides a method for monitoring and recovering application system failures, and the method includes the following steps:

[0025] Step S1, set two or more failure monitoring and recovery program instances to start automatically when the system boots.

[0026] In actual application, the number of deployed failure monitoring and recovery program instances is determined according to needs.

[0027] In an embodiment of the present invention, the application system is built based on the Linux system, and a total of two instances of the fault monitoring and recovery program are deployed, and the fault monitoring and recovery program instances run in the manner of background service programs. The specific implementation method is as follows:

[0028] (1) Write an implementation script for the fault monitoring and recovery program and create a copy of the script. The storage paths of the script and its copy are / usr / local / monitor / malfunc_monitor_1 and / usr / local / monitor / malfunc_monitor_2 respectively;

[0029] (2) Create startup service scripts malfunc_monitor_1.service and malfunc_monitor_2.service for the two instances of the fault monitoring and recovery program. The main content of the malfunc_monitor_1.service service script is as follows:

[0030] [Unit]

[0031] Description=malfunc_monitor_1Server Daemon

[0032] After=network.target

[0033] [Service]

[0034] User=root

[0035] Group=root

[0036] Type=forking

[0037] ExecStart= / usr / local / monitor / malfunc_monitor_1

[0038] Restart=always

[0039] [Install]

[0040] WantedBy=multi-user.target

[0041] The main content of the malfunc_monitor_2.service service script is as follows:

[0042] [Unit]

[0043] Description=malfunc_monitor_2Server Daemon

[0044] After=network.target

[0045] [Service]

[0046] User=root

[0047] Group=root

[0048] Type=forking

[0049] ExecStart= / usr / local / monitor / malfunc_monitor_2

[0050] Restart=always

[0051] [Install]

[0052] WantedBy=multi-user.target

[0053] Through the above script, two instances of the fault monitoring and recovery program are set to start automatically at boot and run as background programs. When the system restarts, the instances of the fault monitoring and recovery program will also start unconditionally.

[0054] It should be noted that the technical solution provided by the present invention is not limited by the type of application system platform, that is, it is applicable to various types of platforms. As described above, in the embodiment of the present invention, the Linux platform is adopted, and its corresponding boot startup configuration is applicable to most versions of Linux. On a Linux platform that does not support systemctl, or other non-Linux platforms, other methods supported by the platform can be used to implement the function of starting the fault monitoring and recovery program instance automatically at boot.

[0055] Step S2, the fault monitoring and recovery program instances monitor each other's running status, and restart the fault monitoring and recovery program instances with abnormal running status.

[0056] In the said step S2, the fault monitoring and recovery program instances monitor each other's running status at a first preset time interval. The first preset time interval is set according to specific needs.

[0057] In an embodiment of the present invention, the application system is built based on the Linux system. The first preset time interval is 30 seconds. The fault monitoring and recovery program instance queries the running status of other fault monitoring and recovery program instances through the systemctl status command, and restarts the process of the fault monitoring and recovery program instance with abnormal running status through the systemctl restart command. Under Linux, the running status of the fault monitoring and recovery program instance can also be monitored by detecting abnormalities in the running log of the fault monitoring and recovery program instance.

[0058] For other platforms, the running status monitoring, abnormal status judgment, and restart of the program instance of the fault monitoring and recovery program are performed in a manner supported by the platform.

[0059] Step S3: The fault monitoring and recovery program instance monitors the target software application service for faults and recovers the faulty software application service.

[0060] In step S3, the fault monitoring and recovery program instance monitors the target software application service for faults at a second preset time interval.

[0061] The second preset time interval needs to be set according to the acceptable software application fault time and considering the overhead of the monitoring and recovery program instance on the system.

[0062] In an embodiment of the present invention, the second preset time interval is 30 seconds.

[0063] In step S3, when the fault monitoring and recovery program instance detects an abnormality in a certain target software application service, it sends a synchronization message to other fault monitoring and recovery program instances, and after the first preset delay time, the fault monitoring and recovery program instance that first discovers the abnormality in the target software application service restarts the abnormal target software application service.

[0064] In an embodiment of the present invention, the synchronization message is a broadcast message, which includes the identifier of the application service with an abnormality, the time of discovering the abnormality, and the identifier of the fault monitoring and recovery program instance that discovers the abnormality in the software application service. The broadcast message is sent via UDP.

[0065] To prevent multiple instances of the fault monitoring and recovery program from restarting the same target software application service simultaneously or successively, when an instance of the fault monitoring and recovery program detects an abnormality in a certain target software application service, it sends a broadcast message to other instances of the fault monitoring and recovery program. The instance of the fault monitoring and recovery program that receives the message determines whether it should restart the target software application service based on the timestamp in the message. If other instances of the fault monitoring and recovery program detected the abnormality earlier, this instance will not restart the target software application service.

[0066] In an embodiment of the present invention, the first preset delay time is 10 seconds. The broadcast message is sent via UDP, and the message format is JSON. The following is an example of the broadcast message:

[0067]

[0068] Among them, the mal_module field represents the identifier of the application service with an abnormality, the time field represents the timestamp when the abnormality was detected, and the sender field represents the identifier of the instance of the fault monitoring and recovery program that detected the abnormality of the application service. The instance of the fault monitoring and recovery program that receives the above broadcast message compares it with its own detection result. If this instance of the fault monitoring and recovery program also detects the abnormality of app1.service, but the detected time is later than the time corresponding to the time field, then this instance of the fault monitoring and recovery program will not restart app1.service. The first instance of the fault monitoring and recovery program that discovers the abnormality of app1.service, after sending the broadcast message, waits for the first preset delay time. If the value of the time field in the broadcast messages received from other instances of the fault monitoring and recovery program within the first preset delay time is later than the time when it itself detected the abnormality, then after the first preset delay time, it will restart app1.service. On the Linux platform, the target software application service can be restarted through the systemctl restart command.

[0069] In addition, each instance of the fault monitoring and recovery program can also stagger the fault monitoring time by generating a random waiting time based on the second preset time interval.

[0070] Message synchronization can also be carried out between multiple instances of the fault monitoring and recovery program in other ways, and the present invention does not limit this.

[0071] The specific means for the instance of the fault monitoring and recovery program to monitor the target software application service for faults include:

[0072] (1) Check the running status of the target software application service through the systemctl status command or other commands and tools provided by the system platform.

[0073] (2) Parse the logs of the target software application service, and combine the preset fault determination rules of each target software application service to perform runtime anomaly judgment. For example, judge the anomaly of the target software application service through network communication anomalies or request response anomalies recorded in the logs.

[0074] (3) Judge whether the target software application service has an anomaly according to whether a core file or an abnormal file is generated.

[0075] (4) Send an http request to the target software application service and perform fault judgment according to the returned response value.

[0076] In short, it is necessary to select the corresponding fault judgment method in combination with the actual situation of the target software application service and the platform. The present invention does not limit the specific fault judgment method. In the prior art, there are already many mature fault judgment methods, which will not be elaborated here.

[0077] Step S4, the fault monitoring and recovery program instance analyzes the system resources and load at the third preset time interval. When the system resource load is detected to be abnormal, the system initialization recovery operation is executed.

[0078] For the Linux platform, the system resources and load can be viewed through the top or free command. When the system resources and load exceed the preset values, such as when the available memory is less than 3% and the CPU load is higher than 95%, the system is restarted through the fault monitoring and recovery program instance.

[0079] In addition, the hardware WatchDog or software WatchDog of each platform can also be combined to perform fault monitoring and recovery on the system. The WatchDog also belongs to the prior art and will not be elaborated here.

[0080] In the technical solution of the present invention, two or more instances of the fault monitoring and recovery program are set to start automatically when the single machine in the edge computing platform is powered on. Each instance of the fault monitoring and recovery program monitors the running state of each other and performs abnormal recovery. At the same time, the target software application service in the application system is monitored for faults, and the faulty software application service is recovered. By setting redundancy for the fault monitoring and recovery program, the present invention ensures the high reliability of the fault monitoring program module itself, avoids the problem that the system fault monitoring and recovery cannot be carried out normally due to the abnormality of the fault monitoring and recovery program itself, thus better ensuring the stability and high availability of the long-term operation of the computer application system, maximizing the online time of the software application service in the edge computing platform, improving the system reliability. The performance loss of the solution of the present invention is small, the versatility is strong, and the development cost is low, which is worthy of popularization.

[0081] The above is only the specific implementation manner of the present invention, and the scope of the present invention cannot be limited thereby. Equal changes made by those of ordinary skill in the art according to this creation, as well as changes well-known to those skilled in the art, should still fall within the scope covered by the present invention.

Claims

1. A method for monitoring and recovering application system failures, characterized in that, The method is used for fault monitoring and recovery of software application services in an application system on a single machine in an edge computing platform. The method includes the following steps: Step S1, set two or more fault monitoring and recovery program instances to start automatically when the machine boots up; Step S2, the fault monitoring and recovery program instances monitor the running states of each other, and restart the fault monitoring and recovery program instances with abnormal running states; Step S3, the fault monitoring and recovery program instances perform fault monitoring on the target software application services, and recover the software application services with faults.

2. The application system fault monitoring and recovery method according to claim 1, wherein In the above Step S1, the fault monitoring and recovery program instances run in the form of background service programs.

3. The application system fault monitoring and recovery method according to claim 1, wherein, In the above Step S2, the fault monitoring and recovery program instances monitor the running states of each other at a first preset time interval.

4. The application system fault monitoring and recovery method according to claim 1, wherein In the above Step S2, the fault monitoring and recovery program instances query the running states of other fault monitoring and recovery program instances through the systemctl status command, and restart the processes of the fault monitoring and recovery program instances with abnormal running states through the systemctl restart command.

5. The application system fault monitoring and recovery method according to claim 1, wherein In the above Step S3, the fault monitoring and recovery program instances perform fault monitoring on the target software application services at a second preset time interval.

6. The application system fault monitoring and recovery method according to claim 1, wherein In the above Step S3, when a fault monitoring and recovery program instance detects that a certain target software application service is abnormal, it sends a synchronization message to other fault monitoring and recovery program instances, and after a first preset delay time, the fault monitoring and recovery program instance that first discovers the abnormality of the target software application service restarts the abnormal target software application service.

7. The application system fault monitoring and recovery method according to claim 6, characterized in that, The synchronization message is a broadcast message, which includes the identifier of the abnormal application service, the time when the abnormality is discovered, and the identifier of the fault monitoring and recovery program instance that discovers the abnormality of the software application service. The broadcast message is sent via UDP.

8. The application system fault monitoring and recovery method according to claim 1, characterized in that, The fault monitoring and recovery method further includes: Step S4, the fault monitoring and recovery program instances analyze the system resources and load at a third preset time interval. When it is detected that the system resource load is abnormal, perform system initialization and recovery operations.