An information system health state monitoring method and device, and an electronic device

CN120803836BActive Publication Date: 2026-09-29360 ZONGHENG INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510746799.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2026-09-29
Estimated Expiration
2045-06-05

AI Technical Summary

Technical Problem

[0006]为解决现有技术中在复杂业务系统中无法及时发现故障、确定故障点的技术问题,本发明的目的是提供一种信息系统健康状态监测方法、装置、电子设备

Benefits of technology

本发明提供的信息系统健康状态监测方法,从主机健康状态、服务单元状态、服务健康度、业务流程健康状态、信息系统综合健康状态五个维度对所述信息系统进行度量,确认所述信息系统健康状态,相较于传统的从主机、进程、网络等完成相关系统健康状态的监测方法,更全面客观的对系统进行评价,且,本发明系统中的服务可自定义配置,通过服务基本信息,可完成对多个服务单元的状态采集,从而度量服务的健康状态,支持分布式架构,如一个服务部署多个服务单元,还能够适用于所有信息系统的综合健康监控,通用性佳;信息系统中的业务流程也可自由配置,基于业务流程包含的服务状态,有效度量整个业务流程的健康状态,通用性强;且,本发明的方案,增加有业务流程的健康状态度量,可及时有效的发现故障、确定故障点;本申请涵盖了信息系统服务、信息系统业务流程、信息系统整体健康度量,可用于一般信息系统、分布式高可用的复杂信息系统的健康度量,为系统综合健康度和业务流程健康度度量,提供了一种度量方法,可自动化发现系统故障及潜在风险,提升复杂信息系统的运维效率,有效降低故障发生概率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120803836B_ABST
    Figure CN120803836B_ABST
Patent Text Reader

Abstract

The application provides an information system health state monitoring method and device and electronic equipment, relates to the technical field of information systems, and solves the technical problem that faults cannot be found and fault points cannot be determined in time in a complex business system in the prior art.The application adopts the following scheme: a health information acquisition probe acquires state information data, uploads the state information data to a health monitoring service through an HTTP interface, and the state information data at least includes host basic information, host state information and service unit state information running on the host; according to the state information data and a preset health measurement method, the system is measured from five dimensions of host health state, service unit state, service health degree, business process health state and information system comprehensive health state, and the information system health state is confirmed; the application can automatically find system faults and potential risks, improves the operation and maintenance efficiency of a complex information system, and effectively reduces the fault occurrence probability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of information system technology, and in particular to a method, apparatus, and electronic device for monitoring the health status of an information system. Background Technology

[0002] This section is intended to provide background or context for embodiments of the invention set forth in the claims. The description herein may include concepts that may be explored, but not necessarily concepts that have been previously conceived or explored. Therefore, unless otherwise stated, what is described in this section is not prior art for the purposes of this application's specification and claims, and is not acknowledged as prior art simply because it is included in this section.

[0003] To adapt to the challenges of new technologies and new businesses, information systems are developing towards collaboration and intelligence. As the scale of systems grows and their complexity increases, system health measurement and monitoring are the foundation of operation and maintenance work. Obtaining the overall system health score and business process health score through system health measurement is an effective means to promptly detect system faults and potential risks, and to address them in a timely manner. Therefore, the daily operation and maintenance of the system is particularly important.

[0004] In existing technologies, the health status of related systems is usually monitored from the host, processes, and network. For example, the health status of information systems is assessed by monitoring CPU utilization, memory utilization, and process status. This method is only practical in small-scale business scenarios. However, when some business processes in complex information systems fail, manual judgment is required to locate the system failure point. The inability to automatically measure the overall health status and business process health status of complex business systems leads to low operation and maintenance efficiency.

[0005] Therefore, in complex business systems, how to promptly detect faults, pinpoint fault locations, and intuitively reflect the overall health status of the system and the health status of business processes is a technical problem that needs to be solved. Summary of the Invention

[0006] To address the technical problem of failing to detect and pinpoint faults in a timely manner in complex business systems in existing technologies, the present invention aims to provide a method, apparatus, and electronic device for monitoring the health status of information systems.

[0007] To address the aforementioned technical problems, in a first aspect, according to some embodiments, the present invention provides a method for monitoring the health status of an information system, comprising: A health information collection probe is used to collect status information data and report the status data to a health monitoring service via an HTTP interface. The status information data includes at least host basic information, host status information, and service unit status information running on the host. Based on the status information data and the preset health measurement method, the information system is measured from five dimensions: host health status, service unit status, service health, business process health status, and information system comprehensive health status, to confirm the health status of the information system. The host health metric is used to measure whether the server in the information system is alive. The service health status is used to confirm whether there is a fault in the current status of the service unit; The service unit status is used to determine whether the service unit will cause business process failures and / or system failures. The business process health status is used to determine whether the business process is operating normally based on the weight of the business process and the services contained in the business process. The comprehensive health status of the information system is used to comprehensively measure the health of the business processes and services contained in the information system, and to determine whether there are any faults in the business processes and / or services that would cause the information system to be usable normally.

[0008] Optionally, as one embodiment, the method for measuring the host health status includes: Read the host information to obtain the most recent heartbeat time; Determine whether the time of the most recent heartbeat is less than a preset interval. If so, mark the host as normal and end the process. If not, then the host status is temporarily abnormal.

[0009] Optionally, as one embodiment, the method for measuring the host health status further includes: For hosts identified as having a temporarily abnormal host status, perform connectivity testing; If the connection is established, it indicates that the host is in a normal state; Otherwise, the host status is marked as abnormal, and the host health status is determined to be disconnected; Optionally, as one embodiment, the method for measuring the host health status further includes: For hosts identified as having a "disconnected" health status, the health status of service units running on those hosts will be marked as "faulty"; and / or Set the alarm level of the information system to emergency alarm.

[0010] Optionally, as one embodiment, the method for measuring service health includes: Receive service unit status information and obtain the number of faulty service units and the total number of service units from it; If the number of faulty service units is 0, then the service health is marked as healthy, the service health score is set to 100, and the process ends. If the number of faulty service units is not 0, then a faulty service unit is identified.

[0011] Optionally, as one embodiment, the method for measuring service health further includes: In the case of a faulty service unit, determine whether all of the service units are faulty; If so, the service is determined to be faulty, the service is marked as having stopped running, the service health score is set to 0, and / or the alarm level of the information system is set to an emergency alarm. If not, it is marked as having more than one service unit failure.

[0012] Optionally, as one embodiment, the method for measuring service health further includes: In the presence of faulty service units, the number of faulty service units is identified, and the following determination is made: If only one of the service units is in a healthy state, the service unit is identified as having a major risk, the service health score is set to 60 points, and / or the alarm level of the information system is set to an emergency alarm. If the number of service unit failures is less than half of the total number of service units, the service health score is set to 85 points, and / or the alarm level of the information system is set to minor alarm. If the number of service unit failures is greater than or equal to half of the total number of service units, and at least two service units are in good condition, then the service health score is set to 70 points, and / or the alarm level of the information system is set to minor alarm.

[0013] Optionally, as one embodiment, the method for determining the state of the service unit includes: Collect the status information of the service unit; Obtain the number of faulty service units and the total number of service units from the service unit status information, and update the service unit status; Based on the latest service unit status information, the service health measurement method process is triggered.

[0014] Optionally, as one embodiment, the method for measuring the health status of the business process includes: Received fault and / or alarm information for a certain service unit; Obtain the business process associated with the faulty service unit; The health status of the business process is determined based on the health status and service weight of all service units included in the business process. The service weight represents the importance of a service unit in the business process. A weight of 0 means that the service unit does not affect the business process, while a weight of 1 means that the service unit must operate normally for the business process to function properly.

[0015] Optionally, as one embodiment, the method for measuring the health status of the business process further includes: Acquire and identify the health status of all service units in the business process; Determine whether the health status of all service units in the business process is healthy. If so, mark the business process as healthy and set the health score of the business process to 100. If not, the health status of the business process is determined to be a fault state.

[0016] Optionally, as one embodiment, the method for measuring the health status of the business process further includes: Determine whether at least one service unit with a weight of 1 among all service units in this business process is faulty; If so, the health status of the business process is determined to be a fault state, the health score of the business process is set to 0, and / or an emergency alarm is generated to indicate that the business process has failed and needs to be repaired immediately. If not, the health status of the business process is determined based on the risk level of all service units with a weight of 1 in the business process.

[0017] Optionally, as one embodiment, the method for measuring the health status of the business process further includes: Identify the risk level of all service units with a weight of 1 in the business process, and make the following determinations: If at least one service unit with a weight of 1 experiences only a general fault, but without significant risk, major risk, or fault, then the business process is identified as having a general risk and operating normally. The health score of the business process is set between 85 and 100 points, and / or the alarm level of the information system is set to a minor alarm, prompting the operation and maintenance personnel that the current business process has a fault but does not affect normal operation and needs to be repaired. If at least one service unit with a weight of 1 has a significant risk, but no major risks or faults, then the business process is marked as having a significant risk but not affecting its use. The health score of the business process is set between 70 and 85, and / or the alarm level of the information system is set to a minor alarm. This prompts the operation and maintenance personnel that the current business process has a significant risk but does not affect normal operation and needs to be repaired. If at least one service unit with a weight of 1 has a significant risk but no fault, then the business process service is marked as having a significant risk but is temporarily available. The health score of the business process is set to between 60 and 75 points, and / or the alarm level of the information system is set to an emergency alarm, prompting the operation and maintenance personnel that the current business process has a significant risk and needs to be repaired in a timely manner.

[0018] Optionally, as one embodiment, the method for measuring the overall health status of the information system includes: If all the health status scores of the business processes and the service health scores are 100, the overall health score of the information system is determined to be 100. If the health status of all the business processes or the health score of the service is 0, the overall health score of the information system is determined to be 0, the alarm level of the information system is an emergency alarm, and the operation and maintenance personnel are prompted that the information system can no longer work normally and needs to be repaired urgently. If all the health status scores of the business processes and the service health scores are greater than or equal to 85, it indicates that there are faults or risks that do not affect the normal operation of the information system, and the availability is high. The comprehensive health score of the information system is determined to be 85. If all the health status scores of the business processes and the service health scores are less than 85 points and greater than or equal to 70 points, it indicates that there is a fault or risk in the business processes or services in the information system. The information system can operate normally. The operation and maintenance personnel should be informed to handle the fault and potential risks in a timely manner. The overall health score of the information system is determined to be 70 points. If all the health status scores of the business processes and the service health scores are less than 70 points and greater than or equal to 60 points, it indicates that there is a significant risk in the business processes or services of the information system, but the information system can operate normally, and the comprehensive health score of the information system is determined to be 60 points. If the number of business process health status scores of 0 is greater than or equal to half of the total number of service units, or if the number of service health scores of 0 is greater than or equal to half of the total number of service units, it indicates that the information system is facing a major availability problem, and the overall health score of the information system is determined to be 0.

[0019] In a second aspect, embodiments of the present invention also provide an information system health status monitoring device, which employs the method described in any of the first aspects above, and the monitoring device further includes: a health status maintenance service and a health information collection probe; The health status maintenance service is used to provide HTTP services, provide configuration data retrieval function and status information reporting channel for the health information collection probe, obtain status information data collected by the health information collection probe through the channel, and is also used to measure the health status of the host, the status of the service unit, the health of the service, the health level of the business process, and the overall health status of the information system according to the information system health status monitoring method and the status information data. A health information collection probe is deployed on all servers in the information system to collect host process information, host status information, pull service information lists, and service unit status information running on the host, and send the collected information to the health status maintenance service through the channel.

[0020] Optionally, in some embodiments, the health status maintenance service is further configured to generate fault information and / or alarm information based on the business process and / or the alarm level.

[0021] Thirdly, according to embodiments of the present invention, an electronic device is also provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the method described in any of the first aspects above.

[0022] Fourthly, according to embodiments of the present invention, a computer-readable storage medium is also provided, on which a computer program is stored, wherein the computer program, when executed by a processor, implements the steps of the method described in any of the first aspects above.

[0023] The above-described technical solution of the present invention has at least the following beneficial technical effects: The information system health status monitoring method provided by this invention measures the information system from five dimensions: host health status, service unit status, service health, business process health status, and overall information system health status, confirming the health status of the information system. Compared with traditional methods that monitor the health status of related systems from hosts, processes, networks, etc., this method provides a more comprehensive and objective evaluation of the system. Furthermore, the services in the system of this invention can be customized and configured. By using basic service information, the status of multiple service units can be collected, thereby measuring the health status of services. It supports distributed architectures, such as deploying multiple service units for one service, and is applicable to the comprehensive health monitoring of all information systems. It has excellent versatility; business processes in information systems can also be freely configured, and the health status of the entire business process can be effectively measured based on the service status included in the business process, making it highly versatile; moreover, the solution of this invention adds health status measurement of business processes, which can timely and effectively detect faults and determine fault points; this application covers information system services, information system business processes, and overall health measurement of information systems, and can be used for health measurement of general information systems and complex distributed highly available information systems. It provides a measurement method for measuring the overall health of the system and the health of business processes, which can automatically detect system faults and potential risks, improve the operation and maintenance efficiency of complex information systems, and effectively reduce the probability of fault occurrence. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention or in the conventional art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0025] Figure 1 This is a schematic block diagram of a complex business system provided in an embodiment of the present invention.

[0026] Figure 2 This is a schematic diagram illustrating the working principle of a health status monitoring service provided in an embodiment of the present invention.

[0027] Figure 3 This is a schematic diagram illustrating the operating principle of a health status monitoring service provided in an embodiment of the present invention.

[0028] Figure 4 This is a schematic diagram of a health measurement algorithm provided in an embodiment of the present invention.

[0029] Figure 5 This is a flowchart of an information system health status monitoring method provided in an embodiment of the present invention.

[0030] Figure 6This is a flowchart of a host health status measurement method provided in an embodiment of the present invention.

[0031] Figure 7 This is a flowchart of a service health measurement method provided in an embodiment of the present invention.

[0032] Figure 8 This is one of the service health measurement example results provided in the embodiments of the present invention.

[0033] Figure 9 This is the second example result of a service health measurement provided in an embodiment of the present invention.

[0034] Figure 10 This is the third example result of a service health measurement provided in this embodiment of the invention.

[0035] Figure 11 This is the fourth example result of a service health measurement provided in an embodiment of the present invention.

[0036] Figure 12 This is a flowchart of a service unit state measurement method provided in an embodiment of the present invention.

[0037] Figure 13 This is a flowchart of a business process health status measurement method provided in an embodiment of the present invention.

[0038] Figure 14 This is one of the common business process examples provided in the embodiments of the present invention.

[0039] Figure 15 This is the second example of a common business process provided in the embodiments of the present invention.

[0040] Figure 16 This is one of the example results of business process health determination provided in the embodiments of the present invention.

[0041] Figure 17 This is the second example result of a business process health assessment provided in an embodiment of the present invention.

[0042] Figure 18 This is the third example result of a business process health assessment provided in this embodiment of the invention.

[0043] Figure 19 This is the fourth example result of a business process health assessment provided in this embodiment of the invention.

[0044] Figure 20 This is a flowchart of an information system health measurement method provided in an embodiment of the present invention.

[0045] Figure 21 This is a schematic block diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0046] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0047] Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concept of the present invention.

[0048] It should be noted that the sequence number mentioned in this application does not necessarily mean that the execution must be strictly in the correct order in the actual implementation process. The sequence number is used to distinguish each step, facilitate explanation, and prevent confusion.

[0049] Furthermore, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0050] To facilitate understanding of the information system health status monitoring method, apparatus, and electronic device provided in the embodiments of this application, the following explains the relevant terms and their specific application scenarios.

[0051] Information system: In this invention, "information system" refers to a system used for collecting, processing, storing and transmitting information.

[0052] Service: In this invention, "service" refers to the basic software unit that constitutes an information system.

[0053] Service Unit: In this invention, "service unit" refers to the process in which a service runs. A service can run on multiple servers, and the service process on each server is a service unit.

[0054] Business Process: In this invention, business process refers to the workflow within an information system consisting of multiple interrelated services, designed to achieve specific business objectives.

[0055] The complex business system of this invention refers to a system composed of multiple interdependent components, which may include hardware, software, data, business processes, etc., and typically possesses the following characteristics: 1) High interconnectivity: There is a high degree of interconnection and interaction between the various components in the system.

[0056] 2) Dynamism: The system can adapt to changes in the environment and dynamically adjust its structure.

[0057] 3) Adaptability: The system can self-adjust and optimize according to changes in external input or internal state.

[0058] 4) Nonlinearity: The behavior of the system is not a simple linear relationship; small changes may lead to large effects, and vice versa.

[0059] 5) Emergence: The overall behavior and characteristics of a system are not simply the sum of the behaviors of individual components, but emerge through the interactions between components.

[0060] 6) Uncertainty: Some behaviors and results of the system are difficult to predict, and there is uncertainty.

[0061] 7) Diversity: The system contains multiple types of components, which may have different functions, characteristics and behavior patterns.

[0062] 8) Scalability: The system may be very large, containing tens of thousands of components and processing massive amounts of data.

[0063] 9) Complexity: The structure and behavior of the system are complex and difficult to fully describe and understand using simple models or methods.

[0064] The research and design of complex information systems need to consider the system's scalability, maintainability, reliability, and security, as well as how to effectively manage and control the complexity of the system.

[0065] In this invention, a complex business system consists of multiple business processes, each business process comprises multiple services, and each service is deployed with multiple replicas in a highly available manner. Each replica is called a service unit. Examples are provided below. Figure 1 As shown, it is important to note that Figure 1 This is just an example; real systems are much more complex.

[0066] To facilitate the explanation of this invention, the components of the health status monitoring service are described.

[0067] The health status monitoring service of this invention mainly consists of two parts: a health information collection probe and a health monitoring service. The two parts communicate with each other through an HTTP interface, and the health monitoring service is responsible for providing HTTP services.

[0068] The health information collection probe and the health monitoring service form a client-server architecture. The health information collection probe is responsible for collecting status data, while the health monitoring service is responsible for processing status data and measuring health status. The health information collection probe acquires status information data and uploads the status information data to the health monitoring service through an HTTP interface.

[0069] The health monitoring service is also responsible for processing the status information data collected by the health information collection probes, and generating alarm and fault information according to business process definitions, alarm rules, etc., to complete the comprehensive status monitoring of the business process and information system as a whole, and providing an HTTP RESTful API for the health information collection probes to use.

[0070] Health information collection probes are deployed on all servers in the information system. They are responsible for collecting host basic information, host status, and service unit status information running on the host, and sending the collected information to the health monitoring service.

[0071] For an explanation of how health status monitoring services work, please refer to [link / reference]. Figure 2 As shown.

[0072] The relevant information regarding the health monitoring business process involved in this application is as follows: In this invention, a service refers to a basic software unit that constitutes an information system. A service can be deployed on multiple servers, and the service process running on each server is called a service unit. In other words, a service unit refers to a copy of a service running on different servers.

[0073] Service information should include at least the following attributes: Service Name: This application may use the service names managed by supervisord and systemd to obtain service operation information. These names must not be duplicated in the information system. Those skilled in the art can set them according to actual needs.

[0074] Service management method: supervisord or systemd.

[0075] Service information description: This refers to a detailed description of the service, used to explain its purpose.

[0076] The business process of this application mainly includes basic business process information and the service units it includes, and at least includes the following: Business process name: Defines the name of the business process, which cannot be repeated in the information system.

[0077] Business Process Description: Used to describe the purpose of the business process.

[0078] List of service units associated with the business process: This describes which service units make up the business process. Each service in the service list needs to include the following: Service Name: Consistent with the service name in the service unit information.

[0079] Service front-end node: Defines the location information of the service in the business process.

[0080] Service weight: Defines the weight information of a service in the business process. The possible values ​​for this application are as follows: 0: Indicates a service failure that does not affect the operation of the overall business process; 1: If the service fails, the entire business process will be interrupted.

[0081] In this invention, an alarm refers to an information system problem that requires attention and handling, but will not affect the normal operation of the information system if it is not handled, such as a server host losing connection or a service unit being unavailable.

[0082] In this invention, alarms can be divided into three levels: minor alarms, secondary alarms, and emergency alarms.

[0083] Alert alerts are the lowest level of alerts, primarily used to notify operations and maintenance personnel of non-critical changes to the system or applications, or states that are about to reach thresholds. These alerts generally do not require immediate response but can serve as reference information for system optimization. For example, CPU or memory usage approaching preset thresholds are examples of alert alerts.

[0084] Minor alarms refer to non-critical failures or potential risks. Although they have little impact on the current operation of the system, they may affect the long-term stability and performance of the system if left unaddressed. These alarms can be addressed later, but they still need to be logged and reviewed periodically. For example, the failure of a service unit is a minor alarm.

[0085] Emergency alerts are the highest level of alerts, indicating a serious system or application failure that may lead to service interruption or data loss. These alerts require an immediate response from the operations and maintenance team to restore normal system operation as quickly as possible. In this invention, server inactivity, service unavailability (e.g., no available service units), and business process unavailability all fall under the category of emergency alerts.

[0086] Health information collection is accomplished using a health information collection probe; its operating principle can be found in [reference needed]. Figure 3 As shown.

[0087] In this invention, the health information collection probe is used to collect the following information: Collect basic host information: When the health monitoring probe is started, it will collect basic host information once more and report it. The collected information includes server UUID, CPU information, memory capacity, disk capacity, disk partition information, system load, operating system information, IP address, etc.

[0088] Collect host operating status information: After the health monitoring probe is started, it will periodically (once every 15 seconds by default) collect and report the host operating status, including server UUID, system load, CPU utilization, memory usage information, disk usage information, etc.

[0089] Retrieve Service Information List: After the health monitoring probe is started, it will periodically (once every 15 seconds by default) retrieve a list of service information that needs to be collected from the health monitoring service, including the service name, service management method and other information for each service.

[0090] The health monitoring probe collects service unit runtime status information based on the obtained service information list and reports the collected information, including service name, service runtime status, process PID, memory usage, CPU usage, startup time, and server UUID. If the service name in the list does not exist on the server, it indicates that the service has not been installed or deployed on the server and does not need to be collected.

[0091] The health monitoring service of this invention provides a channel for health information collection probes to retrieve configuration data and report status information, and completes comprehensive health measurement and alarms for service units, services, business processes, and the system, as detailed below: Configure data retrieval interface: This interface provides the function of querying the list of all services in the information system. The health information collection probe queries the service list through this interface and completes the collection of the service unit's operating status information based on the provided service list.

[0092] Status information reporting interface: This interface is used by the health information collection probe to report the basic host information, host running status, and service unit running status collected.

[0093] Pre-defined health measurement algorithms are provided: These algorithms measure health status from five dimensions: host health status, service unit status, service health, business process health status, and overall information system health status. See the attached documentation for details. Figure 4 .

[0094] The lowest level is the host health status metric, which affects the health status of the service units running on that host.

[0095] The health status of a service unit is determined by measuring the health status of the service unit process and the host.

[0096] The measurement of service health depends on the health status of the service units it contains.

[0097] Business process health status measurement is a comprehensive measurement of the health status of a business process based on the health status of the services it contains and the weight of those services in the business process.

[0098] The system's overall health status measurement comprehensively measures the health status of the entire information system based on the health status of business processes, services, service units, and hosts.

[0099] Based on the above, embodiments of the present invention provide a method for monitoring the health status of an information system, such as... Figure 5 As shown, it includes: Step S101: A health information collection probe is used to collect status data and report the status data to the health monitoring service via an HTTP interface. The status information data includes at least host basic information, host status information, and service unit status information running on the host. Step S102: Based on the status information data and the preset health measurement method, measure the system from five dimensions: host health status, service unit status, service health, business process health status, and information system comprehensive health status, and confirm the health status of the information system. The host health metric is used to measure whether the server in the information system is alive. The service health status is used to confirm whether there is a fault in the current status of the service unit; The service unit status is used to determine whether the service unit will cause business process failures and / or system failures. The business process health status is used to determine whether the business process is operating normally based on the weight of the business process and the services contained in the business process. The comprehensive health status of the information system is used to comprehensively measure the health of the business processes and services contained in the system, and to determine whether there are any faults in the business processes and / or services that would cause the information system to be usable normally.

[0100] Specifically, as one embodiment, the method for measuring the health status of the host includes: Read the host information to obtain the most recent heartbeat time; Determine whether the time of the most recent heartbeat is less than a preset interval. If so, mark the host as normal and end the process. If not, then the host status is temporarily abnormal.

[0101] Specifically, as one embodiment, the method for measuring the health status of the host further includes: For hosts identified as having a temporarily abnormal host status, perform connectivity testing; If the connection is established, it indicates that the host is in a normal state; Otherwise, the host status is marked as abnormal, and the host health status is determined to be disconnected; Specifically, as one embodiment, the method for measuring the health status of the host further includes: For hosts identified as having a "disconnected" health status, the health status of service units running on these hosts will be marked as "faulty"; and / or, Set the alarm level of the information system to emergency alarm.

[0102] Specifically, as one embodiment, the method for measuring the health status of the host further includes: For hosts identified as having a "disconnected" health status, the health status of service units running on these hosts will be marked as "faulty"; and / or, Set the alarm level of the information system to emergency alarm.

[0103] The following is a detailed explanation, such as Figure 6 As shown.

[0104] Host health metrics are used to determine whether servers in an information system are alive, and the data comes from host status reports. This invention executes the metrics through a scheduled task. If no host status information is reported within three time periods (45 seconds), and a ping test is performed, then the host is determined to be faulty.

[0105] Step S201: Read host information and obtain the time of the last heartbeat.

[0106] Read host information from the host information data table and retrieve the "last heartbeat time" field.

[0107] Step S202: Determine if the last heartbeat occurred within 45 seconds.

[0108] The preset interval time in this embodiment of the invention is 45 seconds, which refers to the time of 3 life cycles (15 seconds for a single life cycle). Those skilled in the art can set it according to actual needs.

[0109] If the last heartbeat time exceeds 45 seconds of the current system time, it indicates that no host status information has been reported for more than 3 lifetimes, and the host may be abnormal; otherwise, it indicates that the host is normal.

[0110] Step S203: Ping test to check connectivity.

[0111] When it is determined that the host is abnormal in step 202 above, the ping command of the ICMP protocol can be used to test the host connectivity. If the connection is established, it means that the host is in normal condition and the health detection probe may be malfunctioning. The host status will be restored after the health detection probe status is automatically restored. Step S204: Determine that the host is disconnected and generate an emergency alarm.

[0112] After step S203 uses ping to detect the failure of host connectivity, it can be determined that the host is disconnected, generate an emergency alarm, and mark the health status of the service unit running on the host as not faulty.

[0113] Specifically, as one embodiment, The method for measuring service health includes: Receive service unit status information and obtain the number of faulty service units and the total number of service units from it; If the number of faulty service units is 0, then the service health is marked as healthy, the service health score is set to 100, and the process ends. If the number of faulty service units is not 0, then a faulty service unit is identified.

[0114] Specifically, as one embodiment, the method for measuring service health further includes: In the case of a faulty service unit, determine whether all of the service units are faulty; If so, the service is determined to be faulty, the service is marked as having stopped running, the service health score is set to 0, and / or the alarm level of the information system is set to an emergency alarm. If not, it is marked as having more than one service unit failure.

[0115] Specifically, as one embodiment, the method for measuring service health further includes: In the presence of faulty service units, the number of faulty service units is identified, and the following determination is made: If only one of the service units is in a healthy state, the service unit is identified as having a major risk, the service health score is set to 60 points, and / or the alarm level of the information system is set to an emergency alarm. If the number of service unit failures is less than half of the total number of service units, the service health score is set to 85 points, and / or the alarm level of the information system is set to minor alarm. If the number of service unit failures is greater than or equal to half of the total number of service units, and at least two service units are in good condition, then the service health score is set to 70 points, and / or the alarm level of the information system is set to minor alarm.

[0116] The following description uses specific examples, such as... Figure 7 As shown.

[0117] When the health monitoring service receives the service unit status information reported by the health information collection probe, if the service unit status is faulty, the service fault determination process is triggered; if the service unit status changes from the previous fault to healthy, the service fault recovery process is triggered.

[0118] Step S301: Read the service unit list.

[0119] Read the list of service units contained in the service from the database and calculate the following values: Number of faulty service units: The number of service units whose status is faulty; Total number of service units: The total number of all service units.

[0120] Step S302: Determine whether the number of faulty service units is 0.

[0121] If the number of faulty service units is 0, the service is marked as healthy, there are no faulty service units, the service status is healthy, the service health score is 100, and the service health status judgment ends; otherwise, it is marked that there are faulty service units, the service may have problems, and the process continues to step 3 to measure the service health status.

[0122] Step S303: Determine if all service units are faulty.

[0123] The current service is marked as faulty, the service has stopped running, the service health score is 0, and the information system alarm level is an emergency alarm. Step S304: Determine whether the service unit has only one state that is healthy.

[0124] If only one service unit under the service is in a healthy state while the rest are all faulty, it indicates that the service has a major risk. If it is not repaired in time, service failure may occur. The service health score is 60 points, an emergency alarm is generated, and the operation and maintenance personnel are reminded that the system has a major risk and needs to be repaired in time. Otherwise, it indicates that more than one service unit is faulty and continues to step 5 to measure the service health status.

[0125] Step S305: Determine whether the number of service unit failures is less than half of the total number of service units.

[0126] If the number of service unit failures is less than half of the total number of service units, it indicates that more than half of the service units are still able to provide services normally, with minimal impact on business operations. The service health score is determined to be 85 points, generating a minor alarm to remind operations personnel that the system has risks and needs to be repaired. Otherwise, it indicates that more than half of the service units are failures, but at least two service units are in a healthy state, and the system can be used normally, but the risks are significant. The service health score is determined to be 70 points, generating a minor alarm to remind operations personnel that the system has risks and needs to be repaired.

[0127] At this point, the service health measurement is complete.

[0128] Further explanation follows.

[0129] When service A, which is monitored by the health monitoring service, runs three service units on three servers, different health states such as Figure 8 , Figure 9 , Figure 10 , Figure 11 As shown.

[0130] As can be seen from the diagram, operations and maintenance personnel can intuitively determine the status of the current service unit through the score. Furthermore, when a server fails, operations and maintenance personnel can promptly discover the fault and pinpoint the fault location.

[0131] Specifically, as one embodiment, the method for determining the state of the service unit includes: Collect the status information of the service unit; The service unit status information is parsed, the service unit status is updated, and the latest service unit status is updated in the database; Based on the service unit status information, the service health measurement method process is triggered.

[0132] The following is an explanation through specific examples.

[0133] In complex information systems, services are deployed in a distributed manner using multiple service units. Failure of all service units will cause business process failures and system failures, triggering business process failure alarms and system failure alarms. Failure of not all service units will not affect business processes and the entire information system, but will only trigger service unit failure alarms.

[0134] For detailed processing logic, please refer to... Figure 12 As shown, Step S401: The health information collection probe reports status data.

[0135] The health status information collection probe collects the status information of the service unit and reports it to the health monitoring service.

[0136] Step S402: The health monitoring service receives and processes the status report data.

[0137] The health monitoring service receives and processes the status information reported by the health status information collection probe.

[0138] Step S403: Update the service unit status data.

[0139] The health monitoring service updates the service unit status based on the received service unit status information and updates the latest status to the database.

[0140] Step S404: Trigger the service health status measurement process.

[0141] The health monitoring service triggers the health status measurement process of the service corresponding to the service unit based on the received service unit status report information. For details, please refer to the above embodiments, which will not be listed here.

[0142] For an example of a service unit status maintenance information table, please refer to Table 1.

[0143] Table 1. Example of Service Unit Status Maintenance Information

[0144] Optionally, as one embodiment, the method for measuring the health status of the business process includes: Received fault and / or alarm information for a certain service unit; Obtain the business process associated with the faulty service unit; The health status of the business process is determined based on the health status and service weight of all service units included in the business process. The service weight represents the importance of a service unit in the business process. A weight of 0 means that the service unit does not affect the business process, while a weight of 1 means that the service unit must operate normally for the business process to function properly.

[0145] Optionally, as one embodiment, the method for measuring the health status of the business process further includes: Acquire and identify the health status of all service units in the business process; Determine whether the health status of all service units in the business process is healthy. If so, mark the business process as healthy and set the health score of the business process to 100. If not, the health status of the business process is determined to be a fault state.

[0146] Optionally, as one embodiment, the method for measuring the health status of the business process further includes: Determine whether at least one service unit with a weight of 1 among all service units in this business process is faulty; If so, the health status of the business process is determined to be a fault state, the health score of the business process is set to 0, and / or an emergency alarm is generated to indicate that the business process has failed and needs to be repaired immediately. If not, the health status of the business process is determined based on the risk level of all service units with a weight of 1 in the business process.

[0147] Optionally, as one embodiment, the method for measuring the health status of the business process further includes: Identify the risk level of all service units with a weight of 1 in the business process, and make the following determinations: If at least one service unit with a weight of 1 experiences only a general fault, but without significant risk, major risk, or fault, then the business process is identified as having a general risk and operating normally. The health score of the business process is set between 85 and 100 points, and / or the alarm level of the information system is set to a minor alarm, prompting the operation and maintenance personnel that the current business process has a fault but does not affect normal operation and needs to be repaired. If at least one service unit with a weight of 1 has a significant risk, but no major risks or faults, then the business process is marked as having a significant risk but not affecting its use. The health score of the business process is set between 70 and 85, and / or the alarm level of the information system is set to a minor alarm. This prompts the operation and maintenance personnel that the current business process has a significant risk but does not affect normal operation and needs to be repaired. If at least one service unit with a weight of 1 has a significant risk but no fault, then the business process service is marked as having a significant risk but is temporarily available. The health score of the business process is set to between 60 and 75 points, and / or the alarm level of the information system is set to an emergency alarm, prompting the operation and maintenance personnel that the current business process has a significant risk and needs to be repaired in a timely manner.

[0148] The following is an explanation through specific examples.

[0149] The health status measurement of a business process depends on the business process definition and the health status of the services it contains. The health monitoring service reads the business process definition, obtains the list of services it contains, and then combines the service units contained in the service to perform a comprehensive measurement to complete the determination of the health status of the business process.

[0150] As previously mentioned, service weight represents the importance of a service in a business process. A weight of 0 means that the service does not affect the business process, while a weight of 1 means that the service must run normally for the business process to function properly.

[0151] Methods and processes for measuring the health status of business processes, such as Figure 13 As shown.

[0152] Step S501: Read the business process definition information.

[0153] Read the business process definition information from the database and obtain the list of services included in the business process.

[0154] Step S502: Determine whether all services in the business process are in a healthy state.

[0155] The health status of all services included in the business process is determined. If all services are healthy, the business process is considered to be running healthily, with a health score of 100, and the process ends. Otherwise, it indicates that there may be a problem with the business process, and the next step is taken to determine the health status.

[0156] Step S503: Determine whether all service states in the business process are faulty.

[0157] The system determines the health status of all services included in the business process. If all services are in a fault state, the business process is considered to be unable to run, with a health score of 0. An emergency alarm is generated, prompting operations personnel to note that the business process has failed and needs to be repaired in a timely manner, and the process ends. Otherwise, the system indicates that there may be a problem with the business process and continues to the next step for further determination.

[0158] Step S504: Determine whether the health score of all services in the business process is greater than or equal to 85.

[0159] If the health status of all services included in the business process is ≥85, it indicates that the business process is running normally but there are risks that need to be repaired. If the health score is 85, a minor alarm is generated, indicating to the operation and maintenance personnel that there is a fault in the business process but it does not affect normal operation and needs to be repaired. The process ends. Otherwise, it indicates that there may be a problem with the business process and the next step of judgment is continued.

[0160] Step S505: Determine whether the service health score with a weight of 1 in the business process is greater than or equal to 70.

[0161] If the health status of all services with a weight of 1 in the business process is ≥70, it indicates that the business process is running normally but there is a significant risk that needs to be repaired in time. The health score is 70, a minor alarm is generated, indicating to the operation and maintenance personnel that there is a fault in the business process but it does not affect normal operation and needs to be repaired, and the process ends; otherwise, it indicates that there may be a problem with the business process, and the next step of judgment is continued.

[0162] Step S506: Determine that the service health score of the service with a weight of 1 in the business process is greater than or equal to 60.

[0163] If the health status of all services with a weight of 1 in the business process is ≥60, it indicates that the business process is running normally but has significant risks that need to be addressed promptly. The health score is 60, an emergency alarm is generated, and the operation and maintenance personnel are notified that the business process has significant risks and needs to be addressed promptly. The process then ends. Otherwise, it indicates that the business process has a fault, the health score is determined to be 0, an emergency alarm is generated, and the operation and maintenance personnel are notified that the business process has significant risks and needs to be addressed promptly.

[0164] Further explanation follows.

[0165] For common business process definitions, please refer to [link / reference]. Figure 14 , Figure 15 As shown. By determining the current health status of the business process, we can obtain the following: Figure 16 , Figure 17 , Figure 18 , Figure 19 The example results.

[0166] Optionally, as one embodiment, the method for measuring the overall health status of the information system includes: If all the health status scores of the business processes and the service health scores are 100, the overall health score of the information system is determined to be 100. If all the health status scores of the business processes and the service health scores are 0, the comprehensive health score of the information system is determined to be 0, the alarm level of the information system is an emergency alarm, and the operation and maintenance personnel are prompted that the information system can no longer work normally and needs to be repaired urgently. If all the health status scores of the business processes and the service health scores are greater than or equal to 85, it indicates that there are faults or risks that do not affect the normal operation of the information system, and the availability is high. The comprehensive health score of the information system is determined to be 85. If all the health status scores of the business processes and the service health scores are less than 85 points and greater than or equal to 70 points, it indicates that there is a fault or risk in the business processes or services in the information system. The information system can operate normally. The operation and maintenance personnel should be informed to handle the fault and potential risks in a timely manner. The overall health score of the information system is determined to be 70 points. If all the health status scores of the business processes and the service health scores are less than 70 points and greater than or equal to 60 points, it indicates that there is a significant risk in the business processes or services of the information system, but the information system can operate normally, and the comprehensive health score of the information system is determined to be 60 points. If the number of business process health status scores of 0 is greater than or equal to half of the total number of service units, or if the number of service health scores of 0 is greater than or equal to half of the total number of service units, it indicates that the information system is facing a major availability problem, and the overall health score of the information system is 0.

[0167] The following is an explanation through specific examples.

[0168] The comprehensive health measurement of an information system combines the health status of its business processes and services. Business process scores are represented by HB, service health scores by HS, the total number of business processes by TH, and the total number of services by TS. For the measurement method of the comprehensive health status of an information system, please refer to [link to relevant documentation]. Figure 20 As shown.

[0169] Step S601: Determine whether the health score of all business processes and services is 100.

[0170] This indicates that the entire system is fault-free, risk-free, and operating healthily, with a comprehensive health score of 100 points; otherwise, proceed to the next step of the assessment.

[0171] Step S602: Determine whether the health score of all business processes and services is 0.

[0172] This indicates that all business processes and services within the information system have failed, rendering the information system unusable. The overall health score of the information system is 0, triggering an emergency alarm that alerts maintenance personnel that the information system is no longer functioning properly and requires immediate repair. Otherwise, proceed to the next step of the assessment.

[0173] Step S603: Determine whether all HBs are greater than or equal to 85 and whether all HSs are greater than or equal to 85.

[0174] This indicates that all business processes and services in the information system are in good health and have high availability. There are some faults or risks that do not affect the normal operation of the information system, and the overall health score of the information system is 85 points; otherwise, proceed to the next step.

[0175] Step S604: Determine whether all HBs are greater than or equal to 70 and whether all HSs are greater than or equal to 70.

[0176] This indicates that there are faults or risks in the business processes or services of the information system, but the information system can still operate normally. It is necessary to handle the faults and resolve potential risks in a timely manner. The overall health score of the information system is 70 points; otherwise, proceed to the next step.

[0177] Step S605: Determine whether all HBs are greater than or equal to 60 and whether all HSs are greater than or equal to 60.

[0178] This indicates that there are significant risks in the business processes or services of the information system, such as single points of failure, but the information system can still operate normally. It is necessary to handle the fault and resolve the risk in a timely manner. The overall health score of the information system is 60 points; otherwise, proceed to the next step.

[0179] Step S606: Determine whether the number of HB equals 0 is greater than or equal to half of TH, or whether the number of HS equals 0 is greater than or equal to half of the total number of TS.

[0180] If the number of HB=0 is greater than or equal to half of TH, or the number of HS=0 is greater than or equal to half of TS, it indicates that there are some faults in the business processes or services of the information system and the number of faults is relatively large. For example, some services are unavailable or some business processes are unavailable. The information system is facing a major availability problem, and the overall health score of the information system is 0.

[0181] To further explain, the comprehensive health measurement results of the information system using the current technical solution are shown in Table 2: Table 2 Explanation of the Comprehensive Health Measurement Results of the Information System

[0182] This invention also provides an information system health status monitoring device, which employs the method described in any of the above embodiments. The monitoring device includes: a health status maintenance service and a health information collection probe. The health status maintenance service is used to provide HTTP services, provide configuration data retrieval function and status information reporting channel for the health information collection probe, obtain status information data collected by the health information collection probe through the channel, and is also used to measure the health status of the host, the status of the service unit, the health of the service, the health level of the business process, and the overall health status of the information system according to the information system health status monitoring method and the status information data. A health information collection probe is deployed on all servers in the information system to collect host process information, host status information, pull service information lists, and service unit status information running on the host, and send the collected information to the health status maintenance service through the channel.

[0183] Specifically, as one embodiment, the health status maintenance service is also used to generate fault information and / or alarm information based on the business process and / or the alarm level.

[0184] Specific implementation examples can be found in the foregoing content and will not be listed here.

[0185] According to embodiments of the present invention, an electronic device 2100 is also provided, such as... Figure 21 As shown, it includes a memory 2101, a processor 2102, and a computer program stored on the memory 2101 and executable on the processor. When the processor 2102 executes the program, it implements the steps of the method described in any of the above embodiments.

[0186] According to embodiments of the present invention, a computer-readable storage medium is also provided, on which a computer program is stored, wherein the computer program, when executed by a processor, implements the steps of the method described in any of the above embodiments.

[0187] This invention also provides a computer program product, including a computer program stored in a computer-readable storage medium; when a processor of an electronic device reads the computer program from the computer-readable storage medium, the processor executes the computer program, causing the electronic device to perform the steps of any of the methods described in the above embodiments.

[0188] Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0189] It should be understood that the specific embodiments described above are merely illustrative or explanatory of the principles of the invention and do not constitute a limitation thereof. Therefore, any modifications, equivalent substitutions, improvements, etc., made without departing from the spirit and scope of the invention should be included within the protection scope of the invention. Furthermore, the appended claims are intended to cover all variations and modifications falling within the scope and boundaries of the appended claims, or equivalent forms of such scope and boundaries.

Claims

1. A method for monitoring the health status of an information system, characterized in that, include: A health information collection probe is used to periodically pull a list of service information to be collected from the health monitoring service, collect status information data according to the service information list, and report the status information data to the health monitoring service through an HTTP interface. The status information data includes at least host basic information, host status information, and service unit status information running on the host. Based on the status information data and the preset health measurement method, the information system is measured from five dimensions: host health status, service unit status, service health, business process health status, and information system comprehensive health status, to confirm the health status of the information system. The lowest level is host health status measurement, which affects the health status of service units running on that host. Service unit health status determination: the health status of service units is measured based on the service unit process status and host health status. Service health measurement depends on the health status of the service units contained within the service. Business process health status measurement: a comprehensive measurement of the health status of the business process is performed based on the health status of the services contained within the business process and the weight of the service in the business process. The comprehensive health status measurement of the information system completes the comprehensive measurement of the health status of the entire information system based on the health status of business processes, services, service units, and hosts. The host health metric is used to measure whether the server in the information system is alive; The service health status is used to confirm whether there is a fault in the current status of the service unit; The service unit status is used to determine whether the service unit will cause business process failures and / or system failures. The business process health status is used to determine whether the business process is operating normally based on the weight of the business process and the services contained in the business process. The comprehensive health status of the information system is used to comprehensively measure the health of the business processes and services contained in the information system, and to determine whether there are any faults in the business processes and / or services that would cause the information system to be usable normally.

2. The method according to claim 1, characterized in that, The method for measuring the health status of the host includes: Read the host information to obtain the most recent heartbeat time; Determine whether the time of the most recent heartbeat is less than a preset interval. If so, mark the host as normal and end the process. If not, then the host status is temporarily abnormal.

3. The method according to claim 2, characterized in that, The method for measuring the health status of the host further includes: For hosts identified as having a temporarily abnormal host status, perform connectivity testing; If the connection is established, it indicates that the host is in a normal state; Otherwise, the host status is marked as abnormal, and the host health status is determined to be disconnected.

4. The method according to claim 3, characterized in that, The method for measuring the health status of the host further includes: For hosts identified as having a "disconnected" health status, the health status of service units running on these hosts will be marked as "faulty"; and / or, Set the alarm level of the information system to emergency alarm.

5. The method according to claim 1, characterized in that, The method for measuring service health includes: Receive service unit status information and obtain the number of faulty service units and the total number of service units from it; If the number of faulty service units is 0, then the service health is marked as healthy, the service health score is set to 100, and the process ends. If the number of faulty service units is not 0, then a faulty service unit is identified.

6. The method according to claim 5, characterized in that, The method for measuring service health also includes: In the case of a faulty service unit, determine whether all of the service units are faulty; If so, the service is determined to be faulty, the service is marked as having stopped running, the service health score is set to 0, and / or the alarm level of the information system is set to an emergency alarm. If not, it is marked as having more than one service unit failure.

7. The method according to claim 5 or 6, characterized in that, The method for measuring service health also includes: In the presence of faulty service units, the number of faulty service units is identified, and the following determination is made: If only one of the service units is in a healthy state, the service unit is identified as having a major risk, the service health score is set to 60 points, and / or the alarm level of the information system is set to an emergency alarm. If the number of service unit failures is less than half of the total number of service units, the service health score is set to 85 points, and / or the alarm level of the information system is set to minor alarm. If the number of service unit failures is greater than or equal to half of the total number of service units, and at least two service units are in good condition, then the service health score is set to 70 points, and / or the alarm level of the information system is set to minor alarm.

8. The method according to claim 7, characterized in that, The method for determining the status of the service unit includes: Collect the status information of the service unit; Obtain the number of faulty service units and the total number of service units from the service unit status information, and update the service unit status; Based on the latest service unit status information, the service health measurement method process is triggered.

9. The method according to claim 1, characterized in that, The methods for measuring the health status of the business process include: Received fault and / or alarm information for a certain service unit; Obtain the business processes associated with this service unit; The health status of the business process is determined based on the health status of all service units included in the business process and the weight of the service to which each service unit belongs. The service weight represents the importance of a service unit in the business process. A weight of 0 means that the service unit does not affect the business process, while a weight of 1 means that the service unit must operate normally for the business process to function properly.

10. The method according to claim 9, characterized in that, The method for measuring the health status of the business process also includes: Acquire and identify the health status of all service units in the business process; Determine whether the health status of all service units in the business process is healthy. If so, mark the business process as healthy and set the health score of the business process to 100. If not, the health status of the business process is determined to be a fault state.

11. The method according to claim 10, characterized in that, The method for measuring the health status of the business process also includes: Determine whether at least one service unit with a weight of 1 among all service units in this business process is faulty; If so, the health status of the business process is determined to be a fault state, the health score of the business process is set to 0, and / or an emergency alarm is generated to indicate that the business process has failed and needs to be repaired immediately. If not, the health status of the business process is determined based on the risk level of all service units with a weight of 1 in the business process.

12. The method according to claim 11, characterized in that, The method for measuring the health status of the business process also includes: Identify the risk level of all service units with a weight of 1 in the business process, and make the following determinations: If at least one service unit with a weight of 1 experiences only a general fault, but without significant risk, major risk, or fault, then the business process is identified as having a general risk and operating normally. The health score of the business process is set between 85 and 100 points, and / or the alarm level of the information system is set to a minor alarm, prompting the operation and maintenance personnel that the current business process has a fault but does not affect normal operation and needs to be repaired. If at least one service unit with a weight of 1 has a significant risk, but no major risks or faults, then the business process is marked as having a significant risk but not affecting its use. The health score of the business process is set between 70 and 85, and / or the alarm level of the information system is set to a minor alarm. This prompts the operation and maintenance personnel that the current business process has a significant risk but does not affect normal operation and needs to be repaired. If at least one service unit with a weight of 1 has a significant risk but no fault, then the business process service is marked as having a significant risk but is temporarily available. The health score of the business process is set to between 60 and 75 points, and / or the alarm level of the information system is set to an emergency alarm, prompting the operation and maintenance personnel that the current business process has a significant risk and needs to be repaired in a timely manner.

13. The method according to any one of claims 1, 5-6, and 9-12, characterized in that, The method for measuring the overall health status of the information system includes: If all the health status scores of the business processes and the service health scores are 100, the overall health score of the information system is determined to be 100. If the health status of all the business processes or the health score of the service is 0, the overall health score of the information system is determined to be 0, the alarm level of the information system is an emergency alarm, and the operation and maintenance personnel are prompted that the information system can no longer work normally and needs to be repaired urgently. If all the health status scores of the business processes and the service health scores are greater than or equal to 85, it indicates that there are faults or risks that do not affect the normal operation of the information system, and the availability is high. The comprehensive health score of the information system is determined to be 85. If all the health status scores of the business processes and the service health scores are less than 85 points and greater than or equal to 70 points, it indicates that there is a fault or risk in the business processes or services in the information system. The information system can operate normally. The operation and maintenance personnel should be informed to handle the fault and potential risks in a timely manner. The overall health score of the information system is determined to be 70 points. If all the health status scores of the business processes and the service health scores are less than 70 points and greater than or equal to 60 points, it indicates that there is a significant risk in the business processes or services of the information system, but the information system can operate normally, and the comprehensive health score of the information system is determined to be 60 points. If the number of business process health status scores of 0 is greater than or equal to half of the total number of service units, or if the number of service health scores of 0 is greater than or equal to half of the total number of service units, it indicates that the information system is facing a major availability problem, and the overall health score of the information system is determined to be 0.

14. An information system health status monitoring device, employing the method described in any one of claims 1-13, characterized in that, The monitoring device also includes: a health status maintenance service and a health information collection probe; The health status maintenance service is used to provide HTTP services, provide configuration data retrieval function and status information reporting channel for the health information collection probe, obtain status information data collected by the health information collection probe through the channel, and is also used to measure the health status of the host, the status of the service unit, the health of the service, the health level of the business process, and the overall health status of the information system according to the information system health status monitoring method and the status information data. A health information collection probe is deployed on all servers in the information system. It is used to periodically pull a list of service information to be collected from the health monitoring service, and collect host process information, host status information and service unit status information running on the host according to the service information list. The collected information is then sent to the health status maintenance service through the channel.

15. The apparatus according to claim 14, characterized in that, The health status maintenance service is also used to generate fault information and / or alarm information based on business processes and / or alarm levels.

16. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method according to any one of claims 1-13.

17. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-13.

Citation Information

Patent Citations

  • Software running state evaluation method

    CN102508771A

  • Evaluating system, information interaction system with same and evaluating method

    CN103226668A