A middleware alarm and intelligent recovery system

By designing a middleware alarm and intelligent recovery system and utilizing AGENT programs and AI fault diagnosis, the stability and reliability issues of middleware alarm processing were resolved, enabling real-time monitoring and fault self-healing, reducing manpower workload, and ensuring long-term stable operation of the platform.

CN115664937BActive Publication Date: 2026-01-30海看网络科技(山东)股份有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211317088.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-26
Publication Date
2026-01-30
Estimated Expiration
2042-10-26

AI Technical Summary

Technical Problem

Existing middleware alarm handling methods have poor stability and reliability, and insufficient real-time performance, making it difficult for the platform to operate stably 24/7. They also require a large amount of manual work and are prone to errors or omissions, making it impossible to achieve real-time fault detection and accurate alarms.

Method used

Design a middleware alarm and intelligent recovery system. Utilize the AGENT program algorithm to collect host indicator information, and perform AI fault diagnosis and self-healing through the SERVER service system and fault recovery service system. Combine deep learning and reinforcement learning models to achieve real-time monitoring and fault handling, replacing manual processing methods.

Benefits of technology

It improves the stability and reliability of alarm processing, ensures the platform's long-term stable operation, reduces manual workload, enables real-time monitoring and feedback, and improves the accuracy of fault handling and the stability of the platform.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115664937B_ABST
    Figure CN115664937B_ABST
Patent Text Reader

Abstract

A middleware alarm and intelligent recovery system includes: a middleware client, which is pre-installed with an agent program algorithm to collect and send host indicator information and receive service operation commands and fault recovery commands from a server service system; a server service system, which establishes a communication connection with the middleware client through a Kafka message queue; a fault recovery service system, which establishes a communication connection with the server service system to perform AI fault diagnosis and AI fault self-healing operations, and after completing the fault diagnosis and processing, notifies the server service system to display the fault processing details on a web display device and notify the staff; and an alarm and fault information processor, which establishes a communication connection with the server service system and is equipped with a web display device to display alarm information and fault processing information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical fields:

[0001] This invention relates to a middleware alarm and intelligent recovery system. Background technology:

[0002] Middleware is a type of software that sits between application systems and system software. It uses the basic services provided by system software to connect various parts of application systems or different applications on the network, enabling resource sharing and function sharing.

[0003] As business scale grows and the number of server hosts increases, the scale of middleware deployed on these hosts also becomes larger. Traditional manual alarm handling for faults results in excessive manpower investment and a large amount of repetitive processing work. To reduce the large amount of manual work, computer AI-based fault handling methods are becoming increasingly important. By building a fault handling library, collecting historical fault recovery operations, and then training a large number of models through deep learning, computer AI-based fault handling has become an effective solution.

[0004] Existing manual alarm processing methods have poor stability and reliability, cannot guarantee 24 / 7 stable operation of the platform, and have poor real-time performance. At the same time, the actual workload of staff is large, and errors or omissions are very likely to occur during the monitoring process. The application scope is small, and it cannot achieve real-time fault or problem detection, thus failing to achieve intelligent and accurate alarm detection. Summary of the Invention:

[0005] This invention provides a middleware alarm and intelligent recovery system with a reasonable structural design. Based on the cooperation of multiple functional server components, it replaces the existing manual processing method, improves the stability and reliability of alarm processing, ensures the long-term stable operation of the platform, realizes real-time monitoring and feedback, and intelligently handles fault problems. At the same time, it greatly reduces the actual workload of staff, allowing them to devote more time to more important tasks, further ensuring the stable operation of the platform, and solving the problems existing in the prior art.

[0006] The technical solution adopted by the present invention to solve the above-mentioned technical problems is as follows:

[0007] A middleware alarm and intelligent recovery system, the system comprising:

[0008] The middleware client is pre-configured with an agent program algorithm to collect and send host indicator information, and receive service operation commands and fault recovery commands from the server service system.

[0009] The SERVER service system establishes a communication connection with the middleware client through the Kafka message queue. It is used for collecting and storing host indicator information, detecting abnormal indicators, connecting to the web display device to display alarms and fault transfer processing information, and sending fault information to the fault recovery service system.

[0010] The fault recovery service system establishes a communication connection with the SERVER service system to perform AI fault diagnosis and AI fault self-healing operations. After completing the fault diagnosis and processing, it notifies the SERVER service system to display the fault processing details on the web display device and notify the staff.

[0011] An alarm and fault information processor is provided, which establishes a communication connection with the SERVER service system and is equipped with a web display device to display alarm information and fault handling information.

[0012] An alarm threshold configuration device is also connected to the SERVER service system to set relevant thresholds for middleware indicator data.

[0013] The fault recovery service system is also connected to a memory, which establishes a communication connection with the SERVER service system. The memory is equipped with an alarm storage device and a fault processing storage device.

[0014] A fault handling database is connected to the fault recovery service system, and the fault handling database is an SQL database.

[0015] The web display device is a large web page screen that displays the monitoring host overview, health status, alarm information, and alarm recovery information.

[0016] The fault recovery service system employs deep learning, reinforcement learning, and complex numerical calculation models to perform fault recovery through AI fault diagnosis and AI fault self-healing.

[0017] The SERVER service system communicates with staff via WeChat notifications.

[0018] This invention employs the aforementioned structure, using a middleware client to install the agent program algorithm to collect and send host indicator information, receive service operation commands and fault recovery commands from the server service system; the server service system collects and stores host indicator information, detects abnormal indicators, connects to a web display device to display alarms and fault transfer processing information, and sends fault information to the fault recovery service system; the fault recovery service system performs AI fault diagnosis and AI fault self-healing operations, and after completing fault diagnosis and processing, notifies the server service system to display fault processing details on the web display device and notify personnel; alarm and fault information processors display alarm and fault processing information, offering advantages of stability, practicality, accuracy, and security. Attached image description:

[0019] Figure 1 This is a schematic diagram of the structure of the present invention.

[0020] Figure 2 This is a schematic diagram illustrating the process of sending and storing middleware host indicators according to the present invention.

[0021] Figure 3 This is a schematic diagram of the middleware host indicator anomaly detection process of the present invention.

[0022] Figure 4 This is a flowchart illustrating the details of the fault handling process of the present invention. Detailed implementation method:

[0023] To clearly illustrate the technical features of this solution, the invention will be described in detail below through specific implementation methods and in conjunction with the accompanying drawings.

[0024] like Figure 1-4 As shown, a middleware alarm and intelligent recovery system includes:

[0025] The middleware client is pre-configured with an agent program algorithm to collect and send host indicator information, and receive service operation commands and fault recovery commands from the server service system.

[0026] The SERVER service system establishes a communication connection with the middleware client through the Kafka message queue. It is used for collecting and storing host indicator information, detecting abnormal indicators, connecting to the web display device to display alarms and fault transfer processing information, and sending fault information to the fault recovery service system.

[0027] The fault recovery service system establishes a communication connection with the SERVER service system to perform AI fault diagnosis and AI fault self-healing operations. After completing the fault diagnosis and processing, it notifies the SERVER service system to display the fault processing details on the web display device and notify the staff.

[0028] An alarm and fault information processor is provided, which establishes a communication connection with the SERVER service system and is equipped with a web display device to display alarm information and fault handling information.

[0029] An alarm threshold configuration device is also connected to the SERVER service system to set relevant thresholds for middleware indicator data.

[0030] The fault recovery service system is also connected to a memory, which establishes a communication connection with the SERVER service system. The memory is equipped with an alarm storage device and a fault processing storage device.

[0031] A fault handling database is connected to the fault recovery service system, and the fault handling database is an SQL database.

[0032] The web display device is a large web page screen that displays the monitoring host overview, health status, alarm information, and alarm recovery information.

[0033] The fault recovery service system employs deep learning, reinforcement learning, and complex numerical calculation models to perform fault recovery through AI fault diagnosis and AI fault self-healing.

[0034] The SERVER service system communicates with staff via WeChat notifications.

[0035] The working principle of a middleware alarm and intelligent recovery system in this embodiment of the invention is as follows: based on the mutual cooperation of multiple functional server components, it replaces the existing manual processing method, improves the stability and reliability of alarm processing, ensures the long-term stable operation of the platform, realizes real-time monitoring and feedback, and intelligently handles fault problems; at the same time, it greatly reduces the actual workload of staff, allowing them to devote more time to more important tasks, further ensuring the stable operation of the platform, facilitating popularization and promotion, and having strong versatility, which is of positive significance for improving the long-term stability of the platform.

[0036] The overall solution mainly includes: a middleware client, which is pre-installed with an agent program algorithm to collect and send host indicator information, receive service operation commands and fault recovery commands from the server service system; a server service system, which establishes a communication connection with the middleware client through a Kafka message queue, for collecting and storing host indicator information, detecting abnormal indicators, connecting to a web display device to display alarms and fault transfer processing information, and sending fault information to the fault recovery service system; a fault recovery service system, which establishes a communication connection with the server service system to perform AI fault diagnosis and AI fault self-healing operations, and after completing the fault diagnosis and processing, notifies the server service system to display the fault processing details on the web display device and notify the staff; and an alarm and fault information processor, which establishes a communication connection with the server service system and is equipped with a web display device to display alarm information and fault processing information. With the coordinated communication of these multiple functional components, indicators can be collected quickly and accurately, analyzed to check for anomalies, and alarms can be promptly displayed on the web display device when anomalies occur.

[0037] The KAFKA message queue in this application is mainly used to transmit middleware indicator data, ensuring the accuracy and timeliness of indicator data transmission. After the SERVER service system detects abnormal indicators, it will synchronize the fault to the fault recovery service system. The fault recovery service system completes the fault diagnosis and handling through AI fault diagnosis and AI fault self-healing operations. After the handling is completed, it notifies the SERVER service system. The SERVER service system communicates with the fault recovery service and displays the fault handling details on the web display device.

[0038] Preferably, an alarm threshold configuration device is also connected to the SERVER service system to set relevant thresholds for middleware indicator data, thereby comparing the received indicator data with the thresholds to determine whether any abnormal conditions have occurred.

[0039] Preferably, the fault recovery service system is also connected to a memory, which establishes a communication connection with the SERVER service system. The memory is equipped with an alarm storage device and a fault processing storage device to store relevant fault processing data for easy viewing by staff.

[0040] For web display devices, which are large web page screens, information such as monitoring host overview, health status, alarm information, and alarm recovery is displayed. They can also be configured with middleware and alarm thresholds, and then display alarm information and fault handling details.

[0041] Furthermore, the AI ​​fault diagnosis and AI fault self-healing of the fault recovery service system adopt deep learning, reinforcement learning and complex numerical calculation models to complete fault recovery, ensuring the practicality and accuracy of the fault recovery service system.

[0042] It should be noted that existing communication methods such as WeChat can be used to communicate with staff, as long as staff can obtain the relevant information and data in a timely manner.

[0043] Specifically, on the web display device, the monitoring metrics item_table_size and trigger_table_size of host a are configured. item_table_size configures the amount of disk space occupied by the MySQL table table_1 in the middleware of host a, and trigger_table_size configures an alarm when item_table_size occupies more than 80% of the space. The agent client on host a sends the disk space occupied by table_1 to the message queue KAFKA.

[0044] The server system connects to the Kafka message queue. After receiving the value of the `item_table_size` indicator, it checks whether the data will trigger `trigger_table_size`, i.e., if the value exceeds 80%. If so, it performs the following operations:

[0045] 1. Save the host a index item_table_size data to the storage device.

[0046] 2. Save the alarm information for host a when the item_table_size exceeds 80% to the storage device.

[0047] 3. Display the alarm message on the display device indicating that the middleware indicator item_table_size of host a exceeds 80%.

[0048] 4. Send the host a's metric item_table_size and trigger_table_size information to the fault recovery service system AiRecovery.

[0049] After receiving the fault information, the AiRecovery fault recovery service system performs AI fault diagnosis and executes a recovery script to delete data older than 3 months from host a. Once the recovery script executes successfully, it sends a completion status message to the server service system.

[0050] After receiving the successful processing status from the AiRecovery fault recovery service system, the SERVER server performs the following two steps:

[0051] 1. Receive the latest data of host a's metric item_table_size, and check if it is less than 80% of trigger_table_size. If the data is normal, display host a's middleware alarm recovery information on the display device.

[0052] 2. Store fault recovery information to the storage device.

[0053] In summary, the middleware alarm and intelligent recovery system in this embodiment of the invention, based on the cooperative action of multiple functional server components, replaces existing manual processing methods, improves the stability and reliability of alarm processing, ensures the long-term stable operation of the platform, realizes real-time monitoring and feedback, and intelligently handles fault problems. At the same time, it greatly reduces the actual workload of staff, allowing them to devote more time to more important tasks, further ensuring the stable operation of the platform, facilitating its popularization and promotion, and demonstrating strong versatility. It is of positive significance for improving the long-term stability of the platform.

[0054] The above specific embodiments should not be construed as limiting the scope of protection of the present invention. For those skilled in the art, any alternative improvements or modifications made to the embodiments of the present invention shall fall within the scope of protection of the present invention.

[0055] Any aspects of this invention not described in detail are well-known to those skilled in the art.

Claims

1. A middleware alarm and intelligent recovery system, characterized in that, The system comprises: A middleware client, which is pre-installed with an AGENT program algorithm to collect host index information, receive service operation commands and fault recovery commands of a SERVER service system; A SERVER service system, which is connected with the middleware client through a KAFKA message queue to collect and store host index information, detect abnormal indexes, display alarm and fault transfer processing information on a web display device, and send fault information to a fault recovery service system; A fault recovery service system, which is connected with the SERVER service system to perform AI fault diagnosis and AI fault self-healing processing operations, and after completing the diagnosis and processing of the fault, informs the SERVER service system to display the fault processing details on the web display device and notify the staff; An alarm and fault information processor, which is connected with the SERVER service system and is provided with a web display device to display alarm information and fault processing information; The SERVER service system is connected with the KAFKA message queue, after receiving the value of the index item_table_size, checks whether the data will trigger trigger_table_size, i.e., the value exceeds 80%; if yes, the following operations are performed:

1. saving the host a index item_table_size data to a storage device; 2. saving the host a index item_table_size alarm information exceeding 80% to the storage device; 3. displaying the host a middleware index item_table_size alarm information exceeding 80% on the display device; 4. sending the host a index item_table_size and trigger trigger_table_size information to the fault recovery service system AiRecovery; After the fault recovery service system AiRecovery receives the fault information, performs AI fault diagnosis, and matches the execution of the recovery script: deleting the data of host a three months ago; after the recovery script is successfully executed, sends the processing completion status to the SERVER service system.

2. The middleware alarm and intelligent recovery system of claim 1, wherein: The SERVER service system is also connected with an alarm threshold configuration device to set the related threshold of the middleware index data.

3. The middleware alarm and intelligent recovery system of claim 1, wherein: The fault recovery service system is also connected with a storage, which is connected with the SERVER service system and is provided with an alarm storage device and a fault processing storage device in the storage.

4. The middleware alarm and intelligent recovery system of claim 1, wherein: The fault recovery service system is connected with a fault processing database, which is a SQL database.

5. The middleware alarm and intelligent recovery system of claim 1, wherein: The web display device is a web page large screen to display the monitoring host profile, health degree, alarm information, and alarm recovery information.

6. The middleware alarm and intelligent recovery system of claim 1, wherein: The AI fault diagnosis and AI fault self-healing of the fault recovery service system adopt deep learning, reinforcement learning, and complex numerical calculation models to complete fault recovery.

7. The middleware alarm and intelligent recovery system of claim 1, wherein: The SERVER service system adopts the WeChat notification mode to communicate with the staff.

Citation Information

Patent Citations

  • Information security monitoring system and method

    CN104052634A

  • Equipment fault repairing method and device

    CN111176879A