An alarm-based operation and maintenance method and device

Through the alarm-based operation and maintenance method, abnormal situations in the financial business system are automatically judged and handled, and the problem of low manual processing efficiency in the existing technology is solved, achieving high availability and normal operation of the system.

CN110275795BActive Publication Date: 2025-05-30WEBANK (CHINA)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201910579323.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-06-28
Publication Date
2025-05-30
Estimated Expiration
2039-06-28

AI Technical Summary

Technical Problem

In the prior art, when the financial business system is abnormal, it needs to be processed manually, which is inefficient and may affect the normal operation of the system.

Method used

An alarm-based operation and maintenance method is provided, by receiving the alarm information of the server, determining whether it is a faulty server, and generating processing measures based on the alarm type. Based on the system's operating information and resource information, determine whether the processing measures meet the high availability conditions. If so, the processing measures will be automatically implemented.

Benefits of technology

It realizes automated processing when financial business system abnormalities is abnormal, improves fault handling efficiency, and ensures high availability and normal operation of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN110275795B_ABST
    Figure CN110275795B_ABST
Patent Text Reader

Abstract

The present invention is applicable to the field of fintech, and discloses an operation and maintenance method and device based on alarms. Among them, the method includes: receiving alarm information of a first server, and according to the alarm type of the alarm information, if it is determined that the first server is a faulty server, then generating a first processing measure for the first server according to the alarm information, and combining the operation information and resource information of the system to which the first server belongs to judge whether the first processing measure meets the high-availability condition. If so, execute the first processing measure on the first server. This technical solution is used to provide an automated processing solution for abnormal situations in the financial business system and ensure the normal operation of the financial business system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the field of financial technology (Fintech), and in particular, to an operation and maintenance method and device based on alarms. Background Art

[0002] With the development of computer technology, more and more technologies are applied in the financial field. The traditional financial industry is gradually transforming into financial technology (Fintech), and operation and maintenance technology is no exception. However, due to the security and real-time requirements of the financial and payment industries, higher requirements are also put forward for this technology.

[0003] In the prior art, when a monitoring system monitors an abnormality in a financial business system, it generates an alarm message and notifies relevant staff to solve it. The staff performs manual processing on the financial business system according to the alarm message, such as restarting or shutting down the servers in the system. However, this solution not only has low efficiency but may also affect the normal operation of the entire financial business. Summary of the Invention

[0004] Embodiments of the present invention provide an operation and maintenance method and device based on alarms to provide an automated processing solution when a financial business system is abnormal and ensure the normal operation of the financial business system.

[0005] An operation and maintenance method based on alarms provided by embodiments of the present invention includes:

[0006] Receiving an alarm message from a first server;

[0007] According to the alarm type of the alarm message, determining whether the first server is a faulty server. If so, generating a first processing measure for the first server according to the alarm message;

[0008] Combining the operation information and resource information of the system to which the first server belongs to determine whether the first processing measure meets the high availability condition; the operation information of the system includes the operation information of each server in the system; the resource information of the system includes the hardware information of each server in the system; the high availability condition is a preset condition for indicating that after the first processing measure is executed on the first server, the operation information of the system is not lower than a first preset index;

[0009] If the high availability condition is met, executing the first processing measure on the first server.

[0010] In the above technical solution, after receiving the alarm information of the first server, it can first determine whether the first server is a faulty server. If so, the first processing measure for the first server will be generated to avoid unnecessary operation and maintenance processing of the server, which may affect the normal operation of the entire system. Further, after generating the first processing measure for the first server, the first processing measure is judged in combination with the operation information and resource information of the system to which the first server belongs to determine whether it meets the high-availability conditions. If so, the first processing measure is executed on the first server to ensure the high availability of the system and further avoid affecting the normal operation of the entire system due to the operation and maintenance work of one server. Moreover, an automated processing solution is adopted to improve the fault handling efficiency.

[0011] Optionally, determining whether the first server is a faulty server according to the alarm type of the alarm information includes:

[0012] When the alarm type of the alarm information is a general alarm, it is judged whether the general alarm is the first alarm. If so, the general alarm is closed and updated to the third database, and the staff is notified to enable the staff to set the processing measure for the general alarm; otherwise, the general alarm is closed and the processing measure for the general alarm is executed.

[0013] When the alarm type of the alarm information is a high-risk alarm, it is determined whether the first server is a faulty server.

[0014] In the above technical solution, after receiving the alarm information, the alarm analysis module determines whether the alarm information is a general alarm or a high-risk alarm. If it is a general alarm, it is judged whether the general alarm is the first alarm or a historical alarm. The first alarm means that the processing measure corresponding to the alarm is not stored in the measure database, and the historical alarm means that the processing measure corresponding to the alarm has been stored in the measure database. If it is the first alarm, the operation and maintenance personnel need to be notified for manual processing, and the processing measure is automatically updated to the measure database, that is, the third database, that is, the corresponding optimization item is automatically reported. If it is a historical alarm, the corresponding processing measure can be generated by integrating the processing measures in the measure database, the alarm is processed automatically, and the alarm is closed, and the operation and maintenance personnel are notified of the execution situation. For non-first alarms in general alarms, automated processing and closing are performed to improve the operation and maintenance efficiency.

[0015] Optionally, determining that the first server is not a faulty server includes:

[0016] When the alarm information of the first server indicates that the data processing information of the first server does not meet the second preset index, determine a second server located upstream of the first server according to the positioning information of the first server in the system; if it is determined that the second server is a faulty server, determine that the first server is not a faulty server; or

[0017] The determining that the first server is not a faulty server includes:

[0018] Obtain the version information of the first server from the first database according to the identifier of the first server in the alarm information; if it is determined that the alarm information is caused by the first server being in a version state change, determine that the first server is not a faulty server; the version states of each server in the system are recorded in the first database.

[0019] In the above technical solution, from the perspective of ensuring the high availability of the distributed system, a method for determining that the first server is not a faulty server is provided, avoiding performing processing measures on such faulty servers, thereby affecting the normal operation of the entire system.

[0020] Optionally, the performing the first processing measure on the first server includes:

[0021] Call the script information corresponding to the first processing measure from the second database according to the first processing measure, so as to implement the first processing measure on the first server; the script information corresponding to the processing measures for each server in the system is recorded in the second database.

[0022] In the above technical solution, the second database stores the script information corresponding to the first processing measure, and obtains the script information from the second database, thereby performing the first processing measure. The script information is the script information of the processing measures determined for each server, which is targeted and can meet the needs of different servers.

[0023] Optionally, the first processing measure is used to indicate isolating the first server;

[0024] The performing the first processing measure on the first server includes:

[0025] Send a first shutdown instruction to the message middleware corresponding to the system, where the first shutdown instruction includes the identifier of the first server and a first duration, and the first shutdown instruction is used to indicate that the message middleware isolates the first server for the first duration; or

[0026] Send a second shutdown instruction to the first server, which is used to instruct the first server to shut down the service instance running on the first server; or

[0027] Send a third shutdown instruction to the power interface of the first server, which is used to instruct the power interface to shut down, so that the first server loses power.

[0028] In the above technical solution, the first server can be isolated, thereby preventing data in the system from being sent to the first server, avoiding data loss in the system or affecting data processing in the system.

[0029] Optionally, generating a first processing measure for the first server according to the alarm information includes:

[0030] According to the identifier of the first server in the alarm information, obtain the historical fault information and corresponding processing measures of the first server from the third database; the third database is used to store the historical fault information and corresponding processing measures of each server in the system;

[0031] Combine the fault information in the alarm information, the historical fault information of the first server and the corresponding processing measures to generate a first processing measure for the first server.

[0032] In the above technical solution, by combining the current fault information and historical fault information of the first server, a processing measure for the current fault information of the first server is generated, and this processing can more efficiently solve the current fault information.

[0033] Correspondingly, an embodiment of the present invention further provides an operation and maintenance device based on alarms, including:

[0034] A receiving unit, which is used to receive the alarm information of the first server;

[0035] A processing unit, which is used to determine whether the first server is a faulty server according to the alarm type of the alarm information. If so, generate a first processing measure for the first server according to the alarm information; combine the operation information and resource information of the system to which the first server belongs to judge whether the first processing measure meets the high-availability condition; if it meets the high-availability condition, execute the first processing measure on the first server;

[0036] The operation information of the system includes the operation information of each server in the system; the resource information of the system includes the hardware information of each server in the system; the high-availability condition is a preset condition used to indicate that after the first processing measure is executed on the first server, the operation information of the system is not lower than a first preset index.

[0037] Optionally, the processing unit is specifically configured to:

[0038] When the alarm type of the alarm information is a general alarm, determine whether the general alarm is the first alarm. If so, close the general alarm, update the general alarm to the third database, and notify the staff to enable the staff to set the processing measure for the general alarm; otherwise, close the general alarm and execute the processing measure for the general alarm.

[0039] When the alarm type of the alarm information is a high-risk alarm, determine whether the first server is a faulty server.

[0040] Optionally, the processing unit is specifically configured to:

[0041] When the alarm information of the first server indicates that the data processing information of the first server does not meet the second preset index, determine the second server located upstream of the first server according to the location information of the first server in the system; if it is determined that the second server is a faulty server, determine that the first server is not a faulty server; or

[0042] The processing unit is specifically configured to:

[0043] Obtain the version information of the first server from the first database according to the identifier of the first server in the alarm information; if it is determined that the alarm information is caused by a change in the version state of the first server, determine that the first server is not a faulty server; the version states of each server in the system are recorded in the first database.

[0044] Optionally, the processing unit is specifically configured to:

[0045] Call the script information corresponding to the first processing measure from the second database according to the first processing measure to implement the first processing measure on the first server; the script information corresponding to the processing measures for each server in the system is recorded in the second database.

[0046] Optionally, the first processing measure is used to indicate isolating the first server.

[0047] The processing unit is specifically configured to:

[0048] Send a first shutdown instruction to the message middleware corresponding to the system, where the first shutdown instruction includes the identifier of the first server and a first duration, and the first shutdown instruction is used to indicate that the message middleware isolates the first server for the first duration; or

[0049] Send a second shutdown instruction to the first server, which is used to instruct the first server to shut down the service instance running on the first server; or

[0050] Send a third shutdown instruction to the power interface of the first server, which is used to instruct the power interface to shut down, so that the first server loses power.

[0051] Optionally, the processing unit is specifically configured to:

[0052] Obtain the historical fault information and corresponding handling measures of the first server from the third database according to the identifier of the first server in the alarm information; the third database is used to store the historical fault information and corresponding handling measures of each server in the system;

[0053] Generate a first handling measure for the first server by combining the fault information in the alarm information, the historical fault information of the first server and the corresponding handling measures.

[0054] Correspondingly, an embodiment of the present invention further provides a computing device, including:

[0055] A memory for storing program instructions;

[0056] A processor for calling the program instructions stored in the memory and executing the above-mentioned alarm-based operation and maintenance method according to the obtained program.

[0057] Correspondingly, an embodiment of the present invention further provides a computer-readable non-volatile storage medium, including computer-readable instructions, when the computer reads and executes the computer-readable instructions, the computer is caused to execute the above-mentioned alarm-based operation and maintenance method. Description of the Drawings

[0058] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to these drawings.

[0059] Figure 1 A schematic diagram of a system architecture provided by an embodiment of the present invention;

[0060] Figure 2 A schematic diagram of the structure of an automated processing system provided by an embodiment of the present invention;

[0061] Figure 3 A flowchart of an alarm-based operation and maintenance method provided by an embodiment of the present invention;

[0062] Figure 4 It is a flowchart showing another operation and maintenance method based on alarms provided by an embodiment of the present invention;

[0063] Figure 5 It is a structural diagram of an operation and maintenance device based on alarms provided by an embodiment of the present invention. Detailed implementation manners

[0064] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0065] Figure 1 Exemplarily shown is the system architecture applicable to the operation and maintenance method based on alarms provided by an embodiment of the present invention. The system architecture may include a financial business system 100, a monitoring system 200, an automated processing system 300, and a database system 400.

[0066] The financial business system 100 is a system for executing financial operations. The financial business system 100 may be applicable to a distributed scenario, that is, the financial business system 100 may include multiple clusters, and each cluster thereof may run multiple service instances.

[0067] The monitoring system 200 can monitor various indicators of the financial business system 100. The monitoring system 200 may include a transaction monitoring system and a resource monitoring system. The transaction monitoring system is used to monitor the access volume, request volume, etc. of the transaction services of the financial business system 100; the resource monitoring system is used to comprehensively monitor the servers, operating systems, message middleware, and applications in the financial business system 100, such as monitoring the CPU, memory, network card, etc. of the servers. The transaction monitoring system is, for example, the IMS system, and the resource monitoring system is, for example, the Falcon system.

[0068] The automated processing system 300 is used to automatically process the alarm information of the monitoring system 200, that is, after receiving the alarm information sent by the monitoring system 200, it can automatically identify the category of the alarm information and perform automated processing on the alarm information for different categories.

[0069] The database system 400 can be divided into a measure database, a version database, and a configuration management database. The measure database is used to store the historical failure information of each server in the system and the corresponding handling measures, that is, the handling measures for relevant abnormal scenarios are configured. The measure database is, for example, SOP (Standard Operating Procedure); the version database is used to store the version status of each server in the system, such as the current version information of a certain server and the version status being in the process of upgrading; the configuration management database stores and manages various configuration information of devices in the enterprise IT architecture. It is closely associated with all service support and service delivery processes, supports the operation of these processes, and gives play to the value of the configuration information. At the same time, it depends on relevant processes to ensure the accuracy of the data. The configuration management database is, for example, CMDB (Configuration Management Database).

[0070] In the embodiment of the present invention, the automated processing system 300 may include an alarm access module 310, an alarm analysis module 320, a high-availability engine module 330, a general automated processing module 340, a special automated processing module 350, an alarm information operation module 360, and a processing information notification module 370, as shown Figure 2 as follows.

[0071] 1. Alarm access module 310

[0072] The alarm access module 310 is used to receive all alarm information in production. Specifically, (1) receive alarm information regarding basic resources such as hosts, networks, and databases. This alarm information can be sent to the automated processing system 300 through the Falcon system; (2) receive alarm information regarding business generated by the financial business system 100. This alarm information can be sent to the automated processing system 300 through the IMS system. This alarm information may include the memory status of the server, the usage of the thread pool, etc.

[0073] 2. Alarm analysis module 320

[0074] The alarm analysis module 320 is used to analyze and match the input alarm information, and give an alarm judgment and corresponding operation suggestions. Optionally, the alarm analysis module 320 matches the input alarm information with relevant SOPs, and comprehensively judges the impact level and operation method of the alarm in combination with historical alarms and the current system release situation.

[0075] The embodiments of the present invention can be executed in two steps: SOP library query and historical alarm matching. Among them, SOP library query: query the fault preprocessing mechanism library of the business system, which is equivalent to the measure database, and match the corresponding fault handling steps; Historical alarm matching: After the SOP library query, analyze the historical alarm information of the business system, including the most recent alarm, the alarms in the past seven days, and the version status within three days. At the same time, combined with the handling methods recorded in the SOP library and the current health status (the comprehensive score of the mainframe, network, database, JVM, etc.), corresponding handling measures are given.

[0076] 3. High-availability engine module 330

[0077] The high-availability engine module 330 is used to ensure the high availability of the financial business system 100 in the automatic fault handling in a distributed scenario. Exemplarily, in the process of fault handling, multiple business systems may fail simultaneously or multiple instances in one system may fail. At this time, the automatic fault handling needs to ensure the high availability of the system while processing quickly, and is executed according to the optimal steps with the assistance of the high-availability engine module 330.

[0078] The high-availability engine module 330 determines whether the handling measures generated by the alarm analysis module 320 are feasible by calculating in real time the number of instances in each cluster of each system, the approximate location of the fault point during the fault occurrence, and the basic resource situation of the system. Exemplarily, the high-availability engine module 330 can judge the status of the network, database, and host resources. For example, when a certain host in the system fails, the high-availability engine module 330 will judge whether isolating this instance will meet the available state of the system by judging whether the current system instance stops.

[0079] In the embodiments of the present invention, the high-availability engine module 330 can evaluate the running information of the financial business system 100 before and after executing the handling measures, such as the throughput of business data, the time consumed for processing data, and the resource consumption, etc., and then judge whether to execute this handling measure.

[0080] 4. General automation processing module 340

[0081] The general automation processing module 340 is used for automatic fault handling in general scenarios. In a unified development environment, the restart, isolation, stop, and version rollback of most systems are specific similarities. Through the general automation processing module 340, these four types of operations can be uniformly executed for system exceptions, reducing the risks brought by inconsistent scripts and improving the resolution efficiency. This module mainly targets basic resource class exceptions, such as isolating the business instances of the faulty server when there is message congestion, and performing a health check on the system when the transaction volume is abnormal.

[0082] 5. Special automation processing module 350

[0083] The special automation processing module 350 is used for the management of automation processing scripts in some special systems and specific fault scenarios of various servers. For the inspection and processing of some non-uniform systems and specific fault scenarios, corresponding automation scripts can be configured through this module.

[0084] It should be noted that both the general automation processing module 340 and the special automation processing module 350 are set with corresponding high-availability inspection codes, so as to ensure that a comprehensive confirmation inspection of the system can be performed before the automation script is executed.

[0085] 6. Alarm Information Operation Module 360

[0086] The alarm information operation module 360 is used for the unified management and data operation of various alarm information. Specifically, for the current situation where various alarm data are messy and numerous, including various potential hidden dangers, the alarm information operation module 360 classifies and organizes the daily alarm information by mining the daily alarm information and making big data comparisons with historical data, so as to provide it to the staff. The staff updates the corresponding SOP library to facilitate subsequent automated operation and maintenance work, forming a closed loop of the entire automated operation and maintenance.

[0087] 7. Processing Information Notification Module 370

[0088] The processing information notification module 370 is used to notify various alarm information and the results after automated processing to the operation and maintenance staff, so as to enable the staff to update the system status in a timely manner.

[0089] Based on the above description, Figure 3 An exemplary flowchart of an operation and maintenance method based on alarms provided by an embodiment of the present invention is shown. This flowchart can be executed by an operation and maintenance device based on alarms.

[0090] As Figure 3 shown, this flowchart specifically includes:

[0091] Step 301, receiving the alarm information of the first server.

[0092] The first server is a certain server in the financial business system, and the alarm information can be monitored by the monitoring system and sent to the automation processing system. The alarm information of the first server may include the identifier of the first server, the fault information of the first server, etc.

[0093] Step 302, according to the alarm type of the alarm information, determine whether the first server is a faulty server. If so, generate a first processing measure for the first server according to the alarm information.

[0094] The alarm types of alarm information are divided into general alarms and high-risk alarms. When the alarm type of the alarm information is a general alarm, it is determined whether the general alarm is the first alarm. If so, the general alarm is closed and updated to the third database, i.e., the measures database, to notify the staff so that the staff can set the handling measures for the general alarm; otherwise, the general alarm is closed and the handling measures for the general alarm are executed; when the alarm type of the alarm information is a high-risk alarm, it is determined whether the first server is a faulty server.

[0095] Here, it is possible to determine whether the first server is a faulty server according to the alarm information of the first server, and filter out non-faulty servers according to preset rules. Specifically, when it is determined that the first server is not a faulty server, there are at least the following two situations:

[0096] Case 1: When the alarm information of the first server indicates that the data processing information of the first server does not meet the second preset indicator, the second server located upstream of the first server is determined according to the location information of the first server in the system to which it belongs. If the second server is determined to be a faulty server, the first server is determined not to be a faulty server. This case is explained as follows: when an upstream server fails, it will affect the processing time of the downstream server. According to the call association relationship between the servers, the downstream server will not be determined as a faulty server, but the upstream server will be automatically processed first.

[0097] Case 2: According to the identifier of the first server in the alarm information, the version information of the first server is obtained from the first database. If it is determined that the alarm information is caused by the version status of the first server changing, it is determined that the first server is not a faulty server; wherein the first database records the version status of each server in the system. The first database here is the version database.

[0098] After determining that the first server is a faulty server, a first processing measure for the first server can be generated based on the alarm information. At this time, it is necessary to consider the fault information of the first server in the historical records and the processing measures taken for each fault. The third database can be queried based on the identifier of the first server. The third database is the measure database. Specifically, based on the identifier of the first server in the alarm information, the historical fault information and corresponding processing measures of the first server are obtained from the third database, and the fault information in the alarm information of the first server, the historical fault information of the first server and the corresponding processing measures are combined to generate the first processing measure for the first server.

[0099] In a specific implementation, when receiving the alarm information from the first server, the alarm information can be rated first. The factors to be comprehensively considered include: alarm keywords, system traffic analysis, latency, whether there are associated alarms in the upstream and downstream systems, historical alarm information, etc. These considered factors are weighted to determine the alarm rating. Here, the historical alarm information can consider the number of alarms, alarm levels, handling methods, etc. of the first server in history or during a preset period in history. The alarm rating can be divided into general alarms and high-risk alarms.

[0100] For general alarms, the alarms can be automatically processed according to the handling methods in the third database. At this time, if the general alarm is the first alarm, the staff can be notified so that the staff can perform manual processing and update the processing measures to the third database; for high-risk alarms, it can be further determined whether the alarm can be automatically processed. If so, generate an automated processing measure, that is, the first processing measure; otherwise, the staff needs to be notified so that the staff can perform manual processing and update the processing measures to the third database.

[0101] Step 303: Combine the operation information and resource information of the system to which the first server belongs to determine whether the first processing measure meets the high-availability conditions.

[0102] Here, the operation information of the system includes the operation information of each server in the system, such as the business processing volume and business latency of each server. The resource information of the system includes the hardware information of each server in the system, such as the CPU, memory, and connection relationship of each server.

[0103] The high-availability conditions are preset to indicate that after the first processing measure is executed on the first server, the operation information of the system is not lower than the first preset index. For example, if the business processing volume of the current financial business system is a and a certain server fails, then it is necessary to first evaluate the relationship between the business processing volume b of the financial business system after processing the server and the business processing volume a. For example, if b > a or b is not less than (a × 90%), it is determined that the first processing measure meets the high-availability conditions.

[0104] Step 304: If the high-availability conditions are met, execute the first processing measure on the first server.

[0105] The first processing measure can be a general measure or a special measure. For general measures, a general automated processing module can be used to execute. The general automated processing module is set with script information for general measures such as restart, isolation, stop, and version rollback, which is applicable to most servers. By setting this type of script information, the risk caused by inconsistent scripts can be reduced and the resolution efficiency can be improved.

[0106] In one implementation, the first processing measure may be to isolate the first server. Specifically, the first server can be isolated or the business instances running on the first server can be deactivated. If the first server is connected to other servers through a message middleware, which is a server used to transmit messages in the system, the message middleware can receive the messages sent by the servers and send the messages to the corresponding other servers, that is, it can control whether the sender can send messages and whether the receiver can receive messages. When executing the first processing measure, a first shutdown instruction can be sent to the message middleware. The first shutdown instruction includes the identifier of the first server and a first duration. The message middleware can isolate the first server for the first duration. For example, if the first duration in the first shutdown instruction is 5 minutes, the message middleware can isolate the first server for 5 minutes and cancel the isolation of the first server after 5 minutes. If the first server is directly connected to other servers, the business instances on the first server can be directly shut down. At this time, the business instances can be logged in through the first server, and a second shutdown instruction can be sent to the first server. The second shutdown instruction is used to shut down the business instances running on the first server. When it is no longer possible to log in to the business instances on the first server, the power supply of the first server can be directly shut down. Exemplarily, a power interface is provided on the first server, and a third shutdown instruction can be directly sent to the power interface to instruct the power interface to shut down, so as to cut off the power supply of the first server.

[0107] For special measures, the fault information of each server can be used to set the corresponding script information for the server, and the script information of each server in the system can be stored in the second database corresponding to each server. In a specific implementation, according to the first processing measure, the script information corresponding to the first processing measure is called from the second database to implement the first processing measure on the first server.

[0108] To better explain the embodiments of the present invention, the operation and maintenance process based on alarms will be described below in a specific implementation scenario, as Figure 4 shown as follows:

[0109] After receiving the alarm information, the alarm analysis module determines whether the alarm information is a general alarm or a high-risk alarm. If it is a general alarm, it further determines whether the general alarm is a first-time alarm or a historical alarm. A first-time alarm means that the processing measure corresponding to the alarm is not stored in the measure database, and a historical alarm means that the processing measure corresponding to the alarm has been stored in the measure database. If it is a first-time alarm, the operation and maintenance personnel need to be notified for manual processing, and the processing measure is automatically updated to the measure database, that is, the corresponding optimization item is automatically reported. If it is a historical alarm, the processing measures in the measure database can be integrated to generate corresponding processing measures, the alarm is automatically processed, and the alarm is closed, and the operation and maintenance personnel are notified of the execution situation.

[0110] If it is a high-risk alarm, it is necessary to determine whether the high-risk alarm can be executed automatically, that is, to determine whether there is a corresponding handling measure for the high-risk alarm in the measure database. If so, it is necessary to further determine whether the handling measure meets the high-availability conditions of the high-availability engine. If so, execute the automation script corresponding to the handling measure. Otherwise, notify the operation and maintenance personnel to handle it manually. When the automation script can be executed, after the execution is completed, the corresponding optimization item can be automatically reported, and the operation and maintenance personnel can be notified of the execution status. When the automation script cannot be executed, after the operation and maintenance personnel perform manual processing, the corresponding automation script information can be improved, and the corresponding automatic optimization option can be configured.

[0111] In the above technical solution, after receiving the alarm information of the first server, it can first be determined whether the first server is a faulty server. If so, the first handling measure for the first server will be generated, avoiding unnecessary operation and maintenance processing of the server and affecting the normal operation of the entire system. Further, after generating the first handling measure for the first server, the first handling measure is judged in combination with the operation information and resource information of the system to which the first server belongs to determine whether it meets the high-availability conditions. If so, the first handling measure is executed on the first server to ensure the high availability of the system and further avoid affecting the normal operation of the entire system due to the operation and maintenance work of one server. And an automated processing solution is adopted to improve the fault handling efficiency.

[0112] Based on the same inventive concept, Figure 5 Exemplarily, the structure of an operation and maintenance device based on alarms provided by an embodiment of the present invention is shown, and the device can execute the process of the operation and maintenance method based on alarms.

[0113] The device includes:

[0114] A receiving unit 501, configured to receive alarm information of a first server;

[0115] A processing unit 502, configured to determine whether the first server is a faulty server according to the alarm type of the alarm information. If so, generate a first handling measure for the first server according to the alarm information; judge whether the first handling measure meets the high-availability conditions in combination with the operation information and resource information of the system to which the first server belongs; if it meets the high-availability conditions, execute the first handling measure on the first server;

[0116] The operation information of the system includes the operation information of each server in the system; the resource information of the system includes the hardware information of each server in the system; the high-availability condition is a preset condition for indicating that after the first handling measure is executed on the first server, the operation information of the system is not lower than a first preset index.

[0117] Optionally, the processing unit 502 is further configured to:

[0118] When the alarm type of the alarm information is a general alarm, determine whether the general alarm is the first alarm. If so, turn off the general alarm, update the general alarm to the third database, and notify the staff to enable the staff to set the processing measures for the general alarm; otherwise, turn off the general alarm and execute the processing measures for the general alarm;

[0119] When the alarm type of the alarm information is a high-risk alarm, determine whether the first server is a faulty server.

[0120] Optionally, the processing unit 502 is specifically configured to:

[0121] When the alarm information of the first server indicates that the data processing information of the first server does not meet the second preset index, determine the second server located upstream of the first server according to the positioning information of the first server in the system; if it is determined that the second server is a faulty server, determine that the first server is not a faulty server; or

[0122] The processing unit 502 is specifically configured to:

[0123] Obtain the version information of the first server from the first database according to the identifier of the first server in the alarm information; if it is determined that the alarm information is caused by a change in the version state of the first server, determine that the first server is not a faulty server; the version states of each server in the system are recorded in the first database.

[0124] Optionally, the processing unit 502 is specifically configured to:

[0125] Call the script information corresponding to the first processing measure from the second database according to the first processing measure, so as to implement the first processing measure on the first server; the script information corresponding to the processing measures for each server in the system is recorded in the second database.

[0126] Optionally, the first processing measure is used to indicate isolating the first server;

[0127] The processing unit 502 is specifically configured to:

[0128] Send a first shutdown instruction to the message middleware corresponding to the system, where the first shutdown instruction includes the identifier of the first server and a first duration, and the first shutdown instruction is used to instruct the message middleware to isolate the first server for the first duration; or

[0129] Send a second shutdown instruction to the first server, which is used to instruct the first server to shut down the service instance running on the first server; or

[0130] Send a third shutdown instruction to the power interface of the first server, which is used to instruct the power interface to shut down, so that the first server loses power.

[0131] Optionally, the processing unit 502 is specifically configured to:

[0132] Obtain the historical fault information and corresponding handling measures of the first server from the third database according to the identifier of the first server in the alarm information; the third database is used to store the historical fault information and corresponding handling measures of each server in the system;

[0133] Combine the fault information in the alarm information, the historical fault information of the first server and the corresponding handling measures to generate a first handling measure for the first server.

[0134] Based on the same inventive concept, an embodiment of the present invention further provides a computing device, including:

[0135] A memory for storing program instructions;

[0136] A processor for calling the program instructions stored in the memory and executing the above-mentioned alarm-based operation and maintenance method according to the obtained program.

[0137] Based on the same inventive concept, an embodiment of the present invention further provides a computer-readable non-volatile storage medium, including computer-readable instructions, which when read and executed by a computer, cause the computer to execute the above-mentioned alarm-based operation and maintenance method.

[0138] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a server, so that the instructions executed by the processor of the computer or other programmable data processing devices generate for implementing in the process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks the device with the functions specified therein.

[0139] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instruction means that implement the function specified in one or more of the processes and / or blocks Figure 1 in the process or processes and / or blocks Figure 1 specified in one or more of the blocks or blocks.

[0140] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the function specified in one or more of the processes and / or blocks Figure 1 in the process or processes and / or blocks Figure 1 specified in one or more of the blocks or blocks.

[0141] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made by those skilled in the art once they learn of the basic inventive concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.

[0142] It is obvious that those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.

Claims

1. An operation and maintenance method based on alarms, characterized in that, it includes: Receiving alarm information of the first server; When the alarm type of the alarm information is a high-risk alarm, determining whether the first server is a faulty server; If the first server is a faulty server, generating a first processing measure for the first server according to the alarm information; Combining the operation information and resource information of the system to which the first server belongs to determine whether the first processing measure meets the high-availability condition; The operation information of the system includes the operation information of each server in the system; the resource information of the system includes the hardware information of each server in the system; the high-availability condition is a preset one used to indicate that after the first processing measure is executed on the first server, the operation information of the system is not lower than the first preset index; If it meets the high-availability condition, execute the first processing measure on the first server; Among them, determining that the first server is not a faulty server is determined by the following method, including: When the alarm information of the first server indicates that the data processing information of the first server does not meet the second preset index, determining a second server located upstream of the first server according to the location information of the first server in the system to which it belongs; if it is determined that the second server is a faulty server, then determining that the first server is not a faulty server; or Obtaining the version information of the first server from the first database according to the identifier of the first server in the alarm information; if it is determined that the alarm information is caused by a change in the version state of the first server, then determining that the first server is not a faulty server; the version states of each server in the system are recorded in the first database.

2. The method according to claim 1, characterized in that, The determining whether the first server is a faulty server according to the alarm type of the alarm information includes: When the alarm type of the alarm information is a general alarm, determining whether the general alarm is the first alarm. If so, closing the general alarm, updating the general alarm to the third database, and notifying the staff so that the staff can set the processing measure for the general alarm; otherwise, closing the general alarm and executing the processing measure for the general alarm.

3. The method according to claim 1, characterized in that, The executing the first processing measure on the first server includes: According to the first processing measure, calling the script information corresponding to the first processing measure from the second database to implement the execution of the first processing measure on the first server; the script information corresponding to the processing measures for each server in the system is recorded in the second database.

4. The method according to claim 1, characterized in that, The first processing measure is used to indicate isolating the first server; The executing the first processing measure on the first server includes: Send a first shutdown instruction to the message middleware corresponding to the system, where the first shutdown instruction includes the identifier of the first server and a first duration, and the first shutdown instruction is used to instruct the message middleware to isolate the first server for the first duration; or Send a second shutdown instruction to the first server, which is used to instruct the first server to shut down the service instance running on the first server; or Send a third shutdown instruction to the power interface of the first server, which is used to instruct the power interface to shut down, so that the first server is powered off.

5. The method according to any one of claims 1 to 4, characterized in that the generating a first processing measure for the first server according to the alarm information includes: According to the identifier of the first server in the alarm information, obtain the historical fault information and corresponding processing measures of the first server from a third database; the third database is used to store the historical fault information and corresponding processing measures of each server in the system; Combine the fault information in the alarm information, the historical fault information of the first server and the corresponding processing measures to generate a first processing measure for the first server.

6. An operation and maintenance device based on alarms, characterized in that it includes: a receiving unit, which is used to receive the alarm information of the first server; a processing unit, which is used to determine whether the first server is a faulty server when the alarm type of the alarm information is a high-risk alarm; If the first server is a faulty server, generate a first processing measure for the first server according to the alarm information; Combine the operation information and resource information of the system to which the first server belongs to judge whether the first processing measure meets the high-availability condition; if it meets the high-availability condition, execute the first processing measure on the first server; The operation information of the system includes the operation information of each server in the system; the resource information of the system includes the hardware information of each server in the system; the high-availability condition is a preset condition for indicating that after the first processing measure is executed on the first server, the operation information of the system is not lower than a first preset index; Among them, the method for determining that the first server is not a faulty server includes: When the alarm information of the first server indicates that the data processing information of the first server does not meet the second preset index, determine a second server located upstream of the first server according to the location information of the first server in the system to which it belongs; if it is determined that the second server is a faulty server, then determine that the first server is not a faulty server; or According to the identifier of the first server in the alarm information, obtain the version information of the first server from a first database; if it is determined that the alarm information is caused by a change in the version state of the first server, then determine that the first server is not a faulty server; the version state of each server in the system is recorded in the first database.

7. The device according to claim 6, characterized in that the processing unit is specifically used for: When the alarm type of the alarm information is a general alarm, determine whether the general alarm is the first alarm. If so, close the general alarm, update the general alarm to the third database, and notify the staff to enable the staff to set the handling measures for the general alarm; Otherwise, close the general alarm and execute the handling measures for the general alarm.

8. The device according to claim 6, wherein, the processing unit is specifically configured to: According to the first processing measure, call the script information corresponding to the first processing measure from the second database to implement the first processing measure on the first server; the second database records the script information corresponding to the processing measures for each server in the system.

9. The device according to claim 6, wherein, the first processing measure is used to indicate isolating the first server; the processing unit is specifically configured to: Send a first shutdown instruction to the message middleware corresponding to the system, the first shutdown instruction includes the identifier of the first server and a first duration, and the first shutdown instruction is used to indicate that the message middleware isolates the first server for the first duration; or Send a second shutdown instruction to the first server, which is used to indicate the first server to shut down the service instances running on the first server; or Send a third shutdown instruction to the power interface of the first server, which is used to indicate the power interface to shut down, so that the first server is powered off.

10. The device according to any one of claims 6 to 9, wherein, the processing unit is specifically configured to: According to the identifier of the first server in the alarm information, obtain the historical fault information and corresponding processing measures of the first server from the third database; the third database is used to store the historical fault information and corresponding processing measures of each server in the system; Combine the fault information in the alarm information, the historical fault information of the first server and the corresponding processing measures to generate the first processing measure for the first server.

11. A computing device, wherein, comprises: a memory for storing program instructions; a processor for calling the program instructions stored in the memory and executing the method according to any one of claims 1 to 5 according to the obtained program.

12. A computer-readable non-volatile storage medium, wherein, comprises computer-readable instructions, when the computer reads and executes the computer-readable instructions, the computer is caused to execute the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Information prompting method and related device

    CN108880845A

  • Fault alarm processing method, system, and computer-readable storage medium

    CN108989132A